On the TOEFL’s new 1–6 band, a good Speaking score starts at 4.0 — the minimum most graduate programs publish. In practice, the score that actually clears admissions review tends to sit higher, closer to 4.5–5.0. Hitting that bar — or pushing past it — is harder than it looks, and not for the reason most students assume.

We pulled 600+ graded Speaking practices from Toeflair learners and mapped every score onto the new 1–6 band. On the Listen & Repeat task, three out of four attempts land at 4.0 or above. On spontaneous interview questions, fewer than one in six do. Same learners, same scale — a gap of nearly 60 percentage points in who clears the bar.

The thing separating the two isn’t accent or vocabulary. It’s whether you can execute — clearly, and from memory — in real time.

What “good” means on the 2026 scale

Since January 21, 2026, the TOEFL reports Speaking on a 1–6 band in half-point steps, aligned to the CEFR (the international standard behind “B2”, “C1”, and so on). The change is worth understanding on its own — we covered it in the 2026 reform breakdown — but the short version for Speaking is this: 4.0 is the working definition of “good,” 5.0 is strong, and 6.0 is the ceiling. Each band lines up with a CEFR level, shown in the table below.

Band (1–6)
CEFR
Old 0–30
6
C2
28–30
5–5.5
C1
25–27
4–4.5
B2
20–24
3–3.5
B1
16–19
2–2.5
A2
10–15
1–1.5
A1
0–9

These bands aren’t abstract for us — we scored 600+ real Speaking attempts against them, and the average attempt lands right around the 4.0 line (the low-20s on the old scale). So 4.0 isn’t a far-off ceiling; it’s roughly where a typical learner already sits. What separates that from a competitive 4.5–5.0 is what the rest of this piece is about.

Most graduate programs anchor their minimum at 4.0. The University of Georgia asks for at least a 4 on Speaking; programs tied to teaching assistantships often want more, and some schools still list the top band. On the retiring 0–30 scale, that familiar floor was 23–26. Cutoffs vary by program and the transition is still settling in, so always check the specific school — but as a working answer, 4.0 is “good enough to apply,” and 4.5–5.0 is “competitive.”

Most learners clear the bar — on the easy task

The 2026 Speaking section has two task types:

  • Listen & Repeat — plays a sentence once and asks you to say it back.
  • Take an Interview — asks four spontaneous questions with no prep time.

We scored both, then sorted every attempt into its 1–6 band.

Grouped bar chart comparing the 1–6 band distribution for two TOEFL Speaking tasks. Listen and Repeat peaks at band 4–4.5 (35%) and 5–5.5 (30%), with 75% of attempts reaching 4.0 or above. The Interview task peaks at band 3–3.5 (52%) with another 29% at band 2–2.5, and only 16% reaching 4.0 or above — almost none reach 5.0.

The shape of the two distributions is the whole story. On Listen & Repeat, 75% of attempts reach 4.0 or higher — the task rewards learners who can hear a sentence and reproduce it. On the interview, the curve slides a full band to the left: it peaks at 3.0, and only 16% reach 4.0. Almost nobody touches 5.0.

These are the same people. What changes is the demand. One task hands you the words; the other makes you generate them. That gap — between reproducing language and producing it — is where Speaking scores are actually won and lost.

On Listen & Repeat, the ceiling is clarity, then memory

You would expect a repeat-after-me task to be a pronunciation test. It is — but pronunciation isn’t where it breaks.

Across every ability band, learners sound more fluent than they are clear. Average fluency runs 79 out of 100; average clarity (intelligibility) runs 68. The gap never closes — even top-band learners sit at 95 fluency against 87 clarity.

Grouped bar chart of fluency versus clarity across four ability bands. In every band fluency sits above clarity: Low 37 vs 22, Mid 72 vs 57, High 86 vs 75, Top 95 vs 87. The gap narrows from about 15 points to 7 but never closes.

That gap has a practical meaning: among attempts that already sound fluent (85+), one in six is still hard to understand. Smooth rhythm is masking inaccurate sounds. So which sounds? When we cluster the mispronounced words by what’s actually going wrong, the pattern is consistent — and it’s not about “hard words.”

Horizontal bar chart of pronunciation error types as a share of mispronounced words: weak-form and schwa vowels 51%, rhotic r sounds 48%, multi-syllable words 26%, final consonant clusters 20%, and linking 8% (a lower-bound, text-derived estimate). Categories overlap, so bars sum to more than 100%.

The two leaders — weak-form vowels and the American /r/ — are sounds that don’t exist in Mandarin. They’re the classic first-language transfer points, not exotic vocabulary. Even the word “for” gets mispronounced in about a quarter of attempts, against a 5% average for small function words, because its /ɔːr/ vowel has no Mandarin equivalent. The fix is sound-level, not word-list — and it’s the same drill loop we walk through in the Listen & Repeat repractice method.

But clarity isn’t the only ceiling, and the second one is sneakier. Across the full set, learners drop more words than they mispronounce — 5,134 omissions against 3,512 mispronunciations. A dropped word was never produced at all, which makes it a recall failure, not a mouth failure.

Line chart: as Listen & Repeat sentences get longer, recall accuracy falls — 87% at 5 words or fewer, 76% at 6–8, 68% at 9–11, 66% at 12+ — while the share of words dropped entirely rises — 2%, 8%, 15%, 19%.

Short sentences are conquered cleanly. By 12 words, one word in five simply never comes out — and in long sentences the dropped words cluster in the second half, exactly where short-term memory runs out. Listen & Repeat is, in part, a working-memory test wearing a pronunciation costume. Drilling phonemes won’t help the learners who can’t hold the sentence in the first place.

On the interview, the ceiling is detail, not ideas

Ask learners what’s hard about spontaneous speaking and they’ll say “I don’t know what to say.” The data says otherwise — it isn’t a blank mind. The lowest answers don’t run dry; a few even drift off-topic (asked which films their family enjoys, one learner argues for the future of cinemas). Among answers that stay on topic, two patterns stand out — one about when a score slips, one about what the score is for.

First, when. Scores don’t build as students settle in.

Line chart: overall interview score by question order on a 0–5 internal scale — Q1 narrate 2.46, Q2 opinion 2.69, Q3 agree 2.59, Q4 recommend 2.55; the opening narrative question scores lowest.

The opening question — narrate a specific experience — is the lowest-scoring of the four, even though it sounds the easiest. After a small lift on the opinion question, the line eases back down. The interview never gets easier as it runs; the “just tell me a story” opener is the hidden hardest part.

Second, what. Split each answer into the three things scored — delivery, language, and topic development — and one trails almost everywhere.

Grouped bar chart of three scoring dimensions across four question types on a 0–5 internal scale. Topic development is the weakest dimension on Narrate (2.23), Agree (2.47) and Recommend (2.37); on Opinion, language use (2.56) edges just below topic development (2.61).

All three sit in a narrow, low band — no single dimension collapses — but topic development, how far you actually take the idea, is the weakest or tied-weakest in every question type but one. Learners reliably get a position and a reason out; what the score docks is what comes after the reason. The idea is there. It just stays abstract. Two real answers to the same prompt — does exercise energize you? — show the gap:

Scored 2.5 — stays abstract: “I think being active is very fun because I can have something to do, not just reading books… to exercise is to relax myself.”

Scored 3.5 — one concrete picture: “I generally feel energized after exercise. For example, I often ride my bike after work — I can breathe fresh air, view the scenery, and strengthen my body.”

Same opinion. The second answer scores higher because it spends its sentences on one concrete image instead of restating the claim. The detail is the score.

That is what the 2026 format makes expensive. With prep time gone, the detail can’t be invented in the moment — a learner either walks in with a stock of concrete, reusable specifics to fit the prompt, or stalls and gropes for them live. The cost of groping shows up in delivery, which slips from 2.7 on the first question to 2.5 on the last as the search for something to say eats into fluency.

The hardest part of the interview isn’t having an opinion. It’s developing it with a detail you walked in carrying.

So what does it take to reach a good score?

Put the two tasks side by side and the lesson is the same. The surface skill students drill is rarely the one holding the score down.

Task
What students drill
What actually limits the score
Listen & Repeat
”hard” vocabulary
sound-level clarity (/r/, weak vowels) + holding the sentence in memory
Take an Interview
thinking of ideas
developing the idea with concrete detail — which, with no prep time, has to be brought in, not invented

So the path to a good Speaking score is less about learning more and more about executing better. Three moves do most of the work:

  1. Shadow, don’t just repeat. Speaking along half a second behind the audio trains clarity and recall at once — the two things that actually cap Listen & Repeat.
  2. Lead every interview answer with one concrete example you brought with you. A specific, pre-stocked detail (“I ride my bike after work — fresh air, quiet streets, a clear head”) develops the idea and steadies your delivery, because you’re describing something real instead of inventing under the clock. Build a small stock of these and adapt them across prompts.
  3. Trade fillers for silence. A half-second pause reads as confidence; “you know” and “uh” read as cognitive overload. Cutting them is the fastest delivery win there is.

A good TOEFL Speaking score — 4.0 and up — isn’t a vocabulary problem. It’s an execution problem. And execution, unlike talent, is trainable.

Further reading