PhonemaBlog

TOEFL Speaking Pronunciation: The 2026 Test Scores Your Words Directly

The redesigned TOEFL Speaking section opens with Listen and Repeat, where the official scoring guide drops you a point when one or two content words are ambiguous because of imprecise pronunciation.

Direct answer

The redesigned TOEFL Speaking section contains a task where a single unclear word costs you a point, in writing. In Listen and Repeat, a response that is complete and accurate still drops from 5 to 4 when, in the words of the official scoring guide, “One or two content words may be ambiguous because of imprecise pronunciation.”

No other major English exam states the price of an unclear word that plainly. It also changes what preparation looks like: the unit you practise is a word inside a sentence you heard once, not an accent and not a phoneme chart.

What changed on 21 January 2026

If your preparation materials describe a 0–30 Speaking score, an independent task about a personal opinion, and three integrated tasks with reading and lecture stimuli, they describe the previous test. The current test content is different:

  • Speaking is 11 items with a base time of 8 minutes, across two task types: Listen and Repeat, and Take an Interview.
  • Section scores are now 1 to 6 in increments of 0.5. ETS states that “A student’s overall score is the average of the four section scores, rounded to the nearest half band,” which makes Speaking exactly a quarter of your overall score.
  • The bands are aligned to the CEFR, so that “a score of 5 on any section aligns with C1 proficiency for that particular section.”

The older rubrics folded pronunciation into a column called Delivery, alongside intonation and pacing, and never scored it on its own. The replacement is blunter. Pronunciation is now named inside the descriptors of both tasks, and in one of them it is most of the task.

Listen and Repeat is a pronunciation test under another name

ETS describes the task in one line: “You will listen to short sentences and repeat them exactly as you hear them.”

The Speaking Scoring Guide scores each sentence from 0 to 5. Score 5:

The response is fully intelligible and is an exact repetition of the prompt.

Score 4 lists the two ways to fall short. One is a memory or grammar slip: a function word missing, a tense marker wrong, two words transposed. The other is this:

One or two content words may be ambiguous because of imprecise pronunciation. The speaker may self-correct, but successfully completes the response.

Read what that sentence is doing. The response is complete, on topic and self-corrected, and it still loses a point because a listener could not be sure which word you said. The rubric counts words, and it counts them by whether they were recoverable.

Score 3 adds the second failure mode, and it is not about individual sounds:

In some cases, intelligibility issues cause occasional difficulty in understanding meaning. The speaker may struggle over a word or phrase or run words together, reducing intelligibility.

Running words together is a rhythm problem, not a vowel problem. It is what happens when you rush the middle of a sentence you are afraid of forgetting. See connected speech for what should and should not merge.

Score 2 sets the floor:

Intelligibility is low; the response would be difficult to understand for a listener unfamiliar with the prompt.

That clause — a listener unfamiliar with the prompt — is the whole standard, and it is the one thing you cannot apply to yourself, because you always know what you meant to say.

Take an Interview scores pronunciation as listener effort

The second task type is a simulated interview about academic or campus situations. Its rubric weighs content, pace, grammar and vocabulary together with clarity, so pronunciation is one strand among several. It is still named at every level.

Score 5:

Pronunciation is easily intelligible; rhythm and intonation effectively convey meaning.

Score 4:

Intelligibility and meaning are not impeded by pronunciation, rhythm and intonation, although occasional words/phrases may require minor effort to understand.

Score 3:

Intelligibility is sometimes affected by inaccuracies in word-level pronunciation or stress/rhythm.

The rubric’s own term is word-level pronunciation. The scale between 5 and 3 is measured in listener effort: none, minor, sometimes affected. Accent appears nowhere in either rubric, at any score, which is the same position IELTS states explicitly and the same distinction we cover in intelligibility versus accent.

Content words are the unit to practise

Content words are the ones that carry the message: nouns, main verbs, adjectives, adverbs, numbers. Function words are the connective tissue: articles, prepositions, auxiliaries, pronouns.

The Listen and Repeat rubric treats them differently on purpose. A missing the or a wrong tense marker is a listening and memory error. An ambiguous content word is a pronunciation error. Both cost you the same point, but only one of them is fixed by saying words more clearly, and it is the one you can practise on your own.

Because the sentences are drawn from campus and academic situations, the content words that recur are predictable, and they are exactly the ones learners tend to blur:

  • Long words where the stress moves: ORiginal / oRIGinally, photograph / photography, ADministrate / adminisTRAtion. Misplaced stress makes a long word unrecoverable faster than a wrong vowel does. See word stress.
  • Voiceless /θ/ in through, thorough, method, month, anything.
  • /v/ in available, develop, advisor, review, evening.
  • /l/ and /ɹ/ in library, literature, requirement, research, enrol.
  • Long /iː/ against short /ɪ/ in deadline, seminar, field, fill in.
  • Word-final clusters that carry grammar: asked, texts, lists, months, worlds. See consonant clusters.

Do not drill that list. Build your own from recordings, because the words that fail are specific to your first language and your habits.

How to practise pronunciation for TOEFL Listen and Repeat

Repeat sentences you have heard only once

Practise with audio, not with text on screen. Play a sentence of ten to twenty words once, then say it back from memory. Practising from a written sentence removes the listening half of the task and inflates your results.

The task punishes a specific failure: you spend your attention holding the sentence in memory, and clarity collapses while you do it. That only shows up when the sentence really is gone from the screen.

Mark the content words a listener would have to guess

Play your recording back and write down every noun, verb, adjective, adverb or number that a stranger could not identify without knowing the original sentence. Those are the words the rubric is counting.

This is the hardest step to do honestly, because you know the sentence. Leaving a day between recording and listening helps a little. Feedback that marks the words for you helps more.

Fix the word alone, then put it back in the sentence

Say the word on its own until it is stable, then in its phrase, then inside the whole sentence at speaking pace. A word that only works in isolation has not been fixed, because the test never asks for it in isolation.

Phonema is built for this middle step. You type or paste the sentence you need to say, record yourself, and get feedback on which words came out clearly and which did not, so the list of words to fix comes from your own recordings instead of a generic syllabus. Everything runs on the device, so your recordings do not leave your phone. There is more in how to check your English pronunciation.

Rebuild speed until the endings survive

Repeat at conversational pace and check that word endings and the last few words of the sentence are still there. The score 3 descriptor names running words together as something that reduces intelligibility.

Most candidates lose the final third of a long sentence, not the beginning. Record the same sentence three times and compare only the last five words.

What pronunciation practice will not do for you

Listen and Repeat is roughly the half of Speaking that clarity can fix on its own. Take an Interview is not: four questions with no preparation time, on campus and academic situations, marked on elaboration, pace, grammar and vocabulary as well as intelligibility. Clear words will not rescue an answer that runs out after one sentence.

For that half, my other app, OralPrep, simulates full speaking exams and scores answers against the exam’s own criteria, which is the part a pronunciation tool cannot judge. Phonema does the narrower job: the words themselves, one at a time, until a listener who does not know the sentence can still recover it.

Common questions

Does your accent lower your TOEFL Speaking score? The current scoring guide never mentions accent. It measures how much effort a listener needs to recover your meaning.

How is TOEFL Speaking scored in 2026? Each section gets a band from 1 to 6 in half-point steps, and the overall score is the average of the four sections rounded to the nearest half band, which makes Speaking a quarter of your result.

What is the Listen and Repeat task? ETS describes it in one line: you listen to short sentences and repeat them exactly as you hear them.

Can pronunciation alone cost me a point on TOEFL? Yes. In Listen and Repeat, a response that is otherwise complete drops from 5 to 4 when one or two content words are ambiguous because of imprecise pronunciation.

Should I slow down to be clearer? Only to the point where endings survive. The interview rubric penalises “frequent or lengthy pauses” that make the pace choppy, so clarity bought with hesitation is paid for elsewhere.

Descriptors quoted from the official ETS TOEFL iBT Speaking Scoring Guide and the TOEFL iBT test content pages.

Related sounds