Quick answer: The TOEFL 2026 Speaking section has two task types with zero preparation time: Listen and Repeat (6 items, verbatim sentence repetition) and Take an Interview (5 items, spontaneous responses to an interviewer). Every response is scored 0 to 5 by the ETS SpeechRater engine across five dimensions: Fluency, Intelligibility, Language Use, Organization, and Repeat Accuracy. Rehearsed templates from the old format actively hurt scores because the 2026 rubric specifically rewards natural, spontaneous spoken English.
For the broader Speaking section format overview including timing and task-type walkthroughs, see our TOEFL 2026 Speaking Guide. For a complete overview of the 2026 test structure and all rubrics, see the TOEFL 2026 Score Guide. For TOEFL requirements at US universities including Speaking section minimums, see our reference guide.
Before understanding the rubric, understand what is being scored. The TOEFL 2026 Speaking section runs approximately 8 minutes and contains 11 scored items across two task types.
| Task Type | Items | Structure | Dimensions Scored |
|---|---|---|---|
| Listen and Repeat | 6 items | Hear a short academic sentence, repeat it verbatim. No preparation. No second attempt. | Repeat Accuracy, Intelligibility, Fluency |
| Take an Interview | 5 items | Recorded interviewer asks 5 questions across a 4-minute exchange. Zero preparation for any question. | Fluency, Intelligibility, Language Use, Organization |
Two things about this format break assumptions from the old TOEFL. First, there is no preparation time on any item. Rehearsed templates and memorized structures actively hurt scores because they produce unnatural cadence and grammar mismatched to the question. Second, the section is not adaptive: every student receives the same tasks at the same difficulty level. What varies is performance quality across the five ETS scoring dimensions.
The ETS SpeechRater engine evaluates every Speaking response across measurable speech features grouped into five scoring dimensions. Understanding what each dimension actually measures tells you exactly where your score is won or lost. For an overview of all 2026 format changes, see our hub page.
What ETS measures: Speaking rate (words per minute), frequency and duration of pauses, hesitation patterns, self-corrections, rhythm.
What high-scoring responses show: Sustained speaking rate appropriate for the task (roughly 130-160 words per minute for Take an Interview), few unfilled pauses, natural rhythm with minimal restarts.
What lowers your Fluency score: Extended silences mid-sentence, frequent "um" and "uh" fillers, restarts and self-corrections, speaking rate below 100 words per minute or above 200 words per minute.
What ETS measures: Pronunciation accuracy, stress patterns, intonation, listener effort required to understand the response.
What high-scoring responses show: Clear pronunciation of individual sounds, appropriate word stress, natural sentence intonation, no accent-related comprehension breakdowns.
What lowers your Intelligibility score: Systematic pronunciation errors that require listener effort to decode, flat intonation across sentences, incorrect word stress, mumbling or unclear articulation on key words.
Important note: An accent alone does not lower Intelligibility score. What matters is whether pronunciation errors force the listener to work to understand you.
What ETS measures: Grammatical accuracy, vocabulary range and precision, sentence structure variety, idiomatic use.
What high-scoring responses show: Complex sentence structures used accurately, precise vocabulary matching the topic, few grammatical errors, natural collocation and word choice.
What lowers your Language Use score: Systematic grammatical errors (subject-verb agreement, tense inconsistency), limited vocabulary requiring repetition, sentence fragments, incorrect word choice that changes meaning.
What ETS measures: Coherence of ideas within a response, logical progression, connection between sentences, response completeness relative to the question.
What high-scoring responses show: Clear response structure with main point followed by supporting details, appropriate transitions between ideas, complete answer that addresses the question fully before time expires.
What lowers your Organization score: Rambling responses without clear direction, ideas presented out of order, incomplete responses that trail off, failure to actually answer the question asked.
What ETS measures: Word-for-word accuracy of the repeated sentence, correct word order, correct verb forms and word endings.
What high-scoring responses show: Complete accurate repetition of the target sentence, correct pronunciation of every word, correct grammar preserved from the original.
What lowers your Repeat Accuracy score: Missing words, wrong word order, verb tense or number changes, mispronunciation that changes the target word, added or omitted articles.
Listen and Repeat is scored on a simplified scale focused on accuracy and delivery. You hear a short academic sentence (roughly 15 to 25 words) and must repeat it exactly. No preparation time. No second attempt. Six items per test.
| Score | Repeat Accuracy | Delivery |
|---|---|---|
| 5 | Verbatim reproduction with no errors or omissions | Clear pronunciation, appropriate rhythm and stress, natural intonation matching the original |
| 4 | Verbatim with at most 1-2 minor errors (article, plural marker) | Clear pronunciation with minor stress or rhythm issues |
| 3 | Main content preserved with 2-4 word-level errors or minor omissions | Pronunciation clear enough to be intelligible, some stress inconsistencies |
| 2 | Gist captured but 5+ errors, omitted phrases, or changed word order | Pronunciation may require listener effort, flat intonation |
| 1 | Fragments only, extensive omissions or errors | Pronunciation significantly impedes intelligibility |
| 0 | No attempt or response unrelated to target sentence | Not applicable |
The task rewards precise memory and accurate reproduction, not paraphrase or synonyms. Speaking rate should match the original recording rate closely. Word endings (past tense -ed, plural -s, third person singular -s) are commonly missed and cost points. Function words (articles, prepositions, auxiliary verbs) count equally to content words in the accuracy assessment. Focus attention during the listening phase rather than trying to remember individual words. For Listen and Repeat practice strategies, see our Speaking guide.
Take an Interview is the extended-response task where a recorded interviewer asks 5 spontaneous questions across a 4-minute exchange. Each response typically runs 30 to 60 seconds. Because responses are longer and more complex than Listen and Repeat, all four remaining scoring dimensions apply. Practice with 4 audio Take an Interview exercises.
| Score | Fluency | Intelligibility | Language Use | Organization |
|---|---|---|---|---|
| 5 | Sustained appropriate pace with minimal hesitation | Clear throughout with natural intonation | Complex accurate grammar, precise vocabulary | Complete response with clear structure |
| 4 | Generally appropriate pace with occasional hesitation | Generally clear with minor pronunciation issues | Mostly accurate grammar with occasional errors | Complete response with adequate structure |
| 3 | Some hesitation and inconsistent pace | Intelligible with noticeable pronunciation issues | Some grammatical errors but meaning preserved | Addresses the question but may lack full development |
| 2 | Frequent hesitation, uneven pace | Pronunciation requires listener effort | Consistent errors, limited vocabulary | Incomplete or minimally organized response |
| 1 | Extensive hesitation, unnatural pace | Pronunciation significantly impedes understanding | Severe grammatical limitations | Fragmentary or unclear response |
| 0 | Response unrelated to the question or no response given | |||
Each of the 11 Speaking items is scored 0 to 5 by the ETS SpeechRater engine. These raw scores are normalized and averaged to produce your final Speaking section band from 1.0 to 6.0 in 0.5 increments. For full band conversion details, see our TOEFL score converter.
| Speaking Band (2026) | Legacy Speaking Score (0-30) | What This Level Sounds Like |
|---|---|---|
| 6.0 | 28-30 | Near-native fluency and accuracy across all dimensions |
| 5.5 | 25-27 | Strong across all dimensions with minor issues under pressure |
| 5.0 | 22-24 | Effective communication with some noticeable errors |
| 4.5 | 19-21 | Generally clear with recurring errors that occasionally impede |
| 4.0 | 15-18 | Adequate for basic communication with noticeable limitations |
| 3.5 | 10-14 | Limited but comprehensible in familiar contexts |
| 3.0 | 5-9 | Limited communication with substantial effort required |
Understanding what each band actually sounds like helps calibrate your preparation. Below are sample transcribed responses for the Take an Interview task at bands 5.0, 4.0, and 3.0. The same prompt is used across all three to illustrate how delivery, language use, and organization change scoring.
Prompt: "Some universities require students to attend in-person classes even when online options are available. Do you think this requirement is fair? Explain your reasoning."
"The in-person attendance requirement raises an interesting tension between institutional goals and student autonomy. I think it is justified for programs where interaction is genuinely integral to learning, such as laboratory sciences, clinical training, or studio arts, because those disciplines depend on physical presence in ways that video conferencing cannot replicate. That said, universities should acknowledge that many students balance coursework with employment or family responsibilities. A more effective policy would distinguish between courses where presence adds measurable value and courses where it serves primarily administrative convenience. Blanket mandates ignore these real differences."
Commentary: Distinguished from 5.0 by more sophisticated reasoning (distinguishes between course types with specific examples), natural hedging language ("raises an interesting tension"), and a nuanced conclusion that goes beyond simple agreement or disagreement. Vocabulary is precise and varied throughout.
"I think the in-person requirement is generally fair, though it depends on the university and the field of study. For programs that emphasize collaboration or hands-on skills, like engineering or medicine, in-person attendance genuinely improves outcomes because students learn from peer discussion and lab work that online formats cannot replicate. However, for lecture-based courses where the content is primarily one-directional, requiring in-person attendance seems less justified. Universities could consider a hybrid model that respects both the value of physical presence and the flexibility that online options provide for working students."
Commentary: Complete response, clear structure (position stated, reasoning developed, nuanced conclusion), varied sentence structures, precise vocabulary, natural pace.
"I believe the requirement is fair for most cases. In-person classes allow students to participate in discussions and learn from each other, which is harder to do online. For example, in my university, group projects are much easier to manage when everyone meets face to face. However, I understand that some students have difficult schedules, especially those who work part-time. Maybe the solution is to offer some flexibility, like allowing one or two absences per semester without penalty."
Commentary: Stronger than 4.0 because it includes a concrete personal example and proposes a specific solution. Below 5.0 because vocabulary range is narrower ("difficult schedules," "much easier") and the conclusion is less developed. Grammar is mostly accurate with minor issues ("most cases" without "in").
"I think the requirement is fair but not for everybody. Some students need to work and cannot always come to campus, so online option should be available. But in-person class is better for some subjects because students can ask questions and talk with other students. My opinion is that university should offer both options and let students choose based on their situation. This is more fair for all types of students."
Commentary: Response addresses the question with some development, occasional grammatical errors ("some students need to work"), adequate vocabulary, generally clear structure with some repetition.
"I think in-person class is better than online class. First, students can ask question to teacher and get answer immediately. Second, when students sit in classroom they focus more, because at home there are many distraction like phone and television. But I know some students cannot always attend, so maybe university can record the class for them. Overall I think in-person is more effective for learning."
Commentary: Stronger than 3.0 because it follows a clear structure (position, two reasons, accommodation, conclusion) and maintains topical coherence. Below 4.0 because of consistent minor grammatical errors ("ask question," "many distraction"), limited vocabulary range, and reasoning that stays at surface level without specific examples.
"I think in-person is important because...uh...students can talk to teacher directly. Online is different because you cannot ask easily. But some students have job so is difficult come to university. Maybe university should think about this. In my country, most classes are in-person. This is normal for us."
Commentary: Basic response with limited development, consistent grammatical errors ("students have job so is difficult"), frequent hesitation, response trails off without full conclusion.
Speaking score improvement comes from targeting specific dimensions where your current responses lose points. Below are practical strategies for each of the five ETS scoring dimensions.
Practice speaking for 45-60 seconds without stopping on familiar topics. Record yourself and count filler words ("um," "uh," "like"). Focus on reducing unfilled pauses rather than eliminating all hesitation, because natural speech includes brief pauses. Use bridging phrases ("That reminds me," "Another thing is") when you need thinking time. Aim for natural conversation pace: roughly 130-160 words per minute.
Drill: 60-second sustained response. Record yourself answering a Take an Interview-style question for exactly 60 seconds. Play back the recording and count: (1) number of unfilled pauses longer than 2 seconds, (2) number of "um" and "uh" fillers, (3) approximate words per minute. Target metrics for band 5.0+: fewer than 3 unfilled pauses, fewer than 5 fillers, 130-160 words per minute. Practice this drill 3 times per week over 4 weeks. Improvement typically shows in filler reduction first, then in unfilled pause reduction, then in sustained pace.
Identify your specific pronunciation errors through recorded practice and playback. Focus on stress patterns, because English is a stress-timed language where meaning changes with stress placement. Practice intonation: rising on yes/no questions, falling on statements and wh-questions. Slow down slightly rather than mumbling through unclear words. Get feedback on which specific sounds require listener effort.
Drill: Stress-shift minimal pairs. Choose 10 word pairs where stress changes meaning (e.g., REcord vs reCORD, PREsent vs preSENT, OBject vs obJECT). Record yourself saying each pair in a sentence context. Play back and verify that the stress placement is clearly audible. Then record the same sentences at faster speaking pace and check whether stress placement is preserved under speed. Practice daily for 2 weeks. Track improvement by counting how many pairs maintain correct stress at conversation pace.
Practice using complex sentence structures naturally, not mechanically. Build topic-specific vocabulary rather than memorizing general word lists. Focus on collocation: word combinations that sound natural together ("make a decision" not "do a decision"). Practice using conditionals, relative clauses, and passive constructions in spoken responses. Reduce sentence fragments by planning your sentence completion before starting it.
Drill: Structure upgrade. Record a 45-second response using only simple sentences (subject-verb-object). Then re-record the same response, this time upgrading at least 3 sentences to include a relative clause ("which," "that," "who"), a conditional ("if," "would"), or a concessive ("although," "even though"). Compare both recordings. Track improvement by counting how many complex structures you use naturally without grammatical errors. Practice 3 times per week. Target: 4-5 complex structures per 60-second response without accuracy loss.
Structure every Take an Interview response as: position, two reasons, brief conclusion. Practice completing responses within the time limit rather than trailing off mid-sentence. Use transition phrases naturally: "First," "Additionally," "However," "Overall." Address the specific question asked rather than a related but different topic. Practice pivoting when your first idea does not develop well rather than abandoning the response.
Drill: Timed completion. Set a timer for 45 seconds. Answer a question using the structure: position (5 seconds), reason 1 with example (15 seconds), reason 2 with example (15 seconds), conclusion (10 seconds). Record 5 responses per session. On playback, mark each response as "complete" (reached a clear conclusion before time) or "incomplete" (trailed off or was cut short). Target: 4 of 5 responses complete. Practice twice per week. Most students achieve consistent completion within 3 weeks by learning to budget time across the structure.
Practice active listening rather than trying to memorize word by word. Focus on catching function words (articles, prepositions) during the listening phase, because these are the most commonly missed. Match the original speaker's pace when repeating. Do not skip word endings (-ed, -s, -ing) even under time pressure. Practice with authentic academic English at speaking pace, not slow ESL materials. For structured practice, try our Speaking guide exercises.
Drill: Shadowing with transcript check. Find a 2-minute academic lecture clip (TED-Ed, university lecture recordings) with a transcript. Play one sentence, pause, repeat it aloud, then check against the transcript. Count errors per sentence: missing words, wrong word endings, changed word order. Start with 10-word sentences and progress to 20-word sentences over 3 weeks. Target: fewer than 1 error per 15-word sentence. Practice daily for 10 minutes. Function word accuracy (articles, prepositions, auxiliary verbs) improves fastest with this method.
These are the 10 most frequent score-costing patterns observed in TOEFL 2026 Speaking responses. Each one targets a specific rubric dimension and has a concrete fix.
PrepDrills TOEFL uses Eppy, an AI grading system trained on the ETS 2026 rubrics starting in 2024. Every Speaking response submitted to Eppy returns a band-level estimate plus specific feedback on which of the five scoring dimensions cost you points.
For Listen and Repeat tasks, Eppy analyzes Repeat Accuracy (word-level accuracy against the target sentence) and Intelligibility (pronunciation and stress patterns). For Take an Interview tasks, Eppy evaluates Fluency (speaking rate, pauses, hesitations), Intelligibility (pronunciation and intonation), Language Use (grammar and vocabulary), and Organization (response structure and completeness).
Set 1 easy tasks are free with no credit card required. Full task-type coverage across all difficulty levels is available via Pro ($29.99/month) or Elite ($69.99/3 months). Try the free tier at toefl.prepdrills.com.
The ETS SpeechRater engine evaluates Speaking responses across five dimensions: Fluency (speaking rate, pauses, hesitation), Intelligibility (pronunciation, stress, intonation), Language Use (grammar and vocabulary accuracy), Organization (coherence and response completeness), and Repeat Accuracy (specific to Listen and Repeat tasks). Different dimensions apply to different task types.
Listen and Repeat is scored primarily on Repeat Accuracy and Intelligibility because the task requires verbatim reproduction of a target sentence. Take an Interview is scored on Fluency, Intelligibility, Language Use, and Organization because it requires spontaneous extended responses. Each task type has its own rubric focused on the dimensions relevant to that task.
An accent alone does not lower Intelligibility scores. The ETS rubric focuses on whether pronunciation requires significant listener effort to decode. Speakers with clear non-native accents can still achieve top bands. What matters is whether specific pronunciation errors force the listener to work to understand meaning.
HYPSM (Harvard, Yale, Princeton, Stanford, MIT) and Ivy League programs typically require overall TOEFL bands of 5.0 or higher, with Speaking section bands of at least 5.0 for programs requiring teaching assistantships. For a complete breakdown of US university TOEFL requirements including per-section minimums, see our TOEFL Requirements for US Universities 2026 resource.
The pre-2026 TOEFL Speaking section had four tasks (Independent Speaking and three Integrated tasks) each with 15 to 30 seconds of preparation time. The pre-2026 rubric evaluated each response on a 0 to 4 scale focused on Delivery, Language Use, and Topic Development. The 2026 format eliminates preparation time entirely, reduces to two task types (Listen and Repeat, Take an Interview), scores each item 0 to 5, and evaluates responses across five dimensions instead of three.
AI-graded Speaking practice provides immediate feedback on specific dimensions where your responses lose points. PrepDrills TOEFL includes Speaking tasks for both Listen and Repeat and Take an Interview, with Eppy AI grading calibrated to the ETS 2026 rubric. Free practice available at toefl.prepdrills.com.
There is no universal passing score. Speaking band requirements vary by university and program. Most competitive graduate programs require Speaking bands of 4.0 or higher. Some programs requiring teaching assistantships or heavy oral presentation require Speaking bands of 5.0 or higher. Verify specific requirements on your target program's admissions page.
For 1-on-1 TOEFL Speaking coaching, Epic Exam Prep offers online sessions worldwide with 25 years of combined expertise.