Home›Articles
Why Learn English with Real Conversations? The Research
In this article: Textbook English and real English are different languages. What 40 years of research says about hesitations, accents, noise and unscripted audio.
You have studied English for years. You can read an article, follow your teacher and pass the exam. Then two English speakers start chatting at the next table, and you catch one word in three.
Nothing is wrong with your English. What you learned from and what you are hearing are two different kinds of English. One is written down first and read aloud. The other is made up on the spot, by people who hesitate, interrupt each other and use expressions no coursebook lists.
That is why every conversation we record is unscripted. This page sets out the reasoning: what researchers have measured about real speech, where the evidence is strong, and where it is thin.
What the research says
- Textbook dialogue and real conversation are measurably different. Real talk is full of hesitations, restarts and overlaps. The typical scripted dialogue removes nearly all of them. Well established.
- Real speech is harder at first, not easier. Learners score lower on unscripted versions of the same content. Good evidence.
- It pays off when it comes with support. Over ten months, real material with tasks beat a textbook course on listening. Good evidence, from one long study.
- Noise and several voices cost learners far more than native speakers. The cause is gaps in knowledge, not in hearing. Well established.
- You get better by learning more words, more phrases and more voices. Whether noisy practice beats clean practice has not been tested directly. Well established that you improve; only early evidence on noise itself.
- Listening a lot is essential, and it works best with a transcript and a plan. Well established.
On this page
- Why can’t I understand native speakers when I understand English well?
- What is the difference between textbook English and real English?
- Does learning English with real conversations actually work?
- Why is English so hard to follow in a bar, or when several people talk at once?
- Can I get better at understanding English in noisy places?
- Does hearing different accents help me understand new ones?
- Do “um” and “er” make English harder to understand?
- How much of real English is fixed expressions?
- How much vocabulary do I need to follow a real conversation?
- Does listening to a lot of English really work?
- How are the Zapp! English recordings made?
- When is real English too hard, and what helps?
- Frequently asked questions
- Sources
Why can’t I understand native speakers when I understand English well?
Because the English you learned from was careful, scripted speech, and conversation is not. In unscripted talk, speakers hesitate, repeat or restart five or six times in every hundred words [1]. They leave about a fifth of a second between turns [2] [3]. Roughly one word in twenty loses a syllable [4]. Your English is fine. Your ear has not met this version of it yet.
When 40 learners kept diaries of their listening, the commonest problems were not rare words. They were forgetting what had just been said, and failing to recognise words they already knew [5]. You know the word. You did not hear it as that word.
The official European scale agrees. Following a lively conversation between other people, with overlapping turns and colloquial expressions at natural speed, is described at level C1. At B1 the scale expects speech that is clearly articulated [6]. So if real conversations feel beyond you at B1 or B2, you are not behind. Without support, they are above your level. The rest of this page is about the support.
What is the difference between textbook English and real English?
The typical scripted dialogue is tidier and more complete than anything people actually say. When a researcher re-enacted seven coursebook dialogues as real encounters and compared equal amounts of talk, the real version held 54 hesitations against 8 in the book. It held 30 small listener responses such as “mm” and “yeah” against 4, and 6 overlaps against none [7].
A much larger study points the same way. It set 43 school textbooks beside 11.4 million words of recorded British conversation. The textbook dialogues used softeners such as “sort of” less than half as often. They used fewer contractions and far more nouns [8]. A survey of 24 coursebooks published in Britain called their coverage of spoken grammar “at best patchy” [9].
Here is the difference in two lines. In one of our episodes, recorded on the beach, Katie asks Tom to finish the phrase “as deaf as…”:
“Erm… is this to do with doors? As deaf as a doornail? Doormouse?”
“No, I think that’s dead as a doornail.”
A hesitation, a question, two wrong guesses and a correction. No coursebook would print it. It is also how you will hear people reach for an expression. (The answer is “as deaf as a post”.)
This is not an attack on textbooks. Newer coursebooks carry more natural features than older ones [7], and clear, scripted audio has its place when you are starting out. But it cannot get you used to speech it does not contain.
Does learning English with real conversations actually work?
Yes, on two honest conditions. Real speech is harder at first: when 171 learners heard the same content either scripted or unscripted, the unscripted group scored lower [10]. And it works when it comes with support: in a ten-month university course, students taught with real material, and tasks built on it, improved more in listening than students taught from a textbook [11].
The first study tested learners of Spanish, not English, and it measured understanding on the day, not learning over time [10]. The second followed 62 students in Japan. The real-material group finished ahead on five of eight measures, including listening, vocabulary and pronunciation, though not grammar [11]. But the material did not work alone. The teacher also taught how natural speech links and shortens words. So this is evidence for real conversations plus guided work, not for audio by itself.
So we will never tell you real conversations are easier. They are the English you will actually meet. The only way to get used to it is to meet it, with help.
Why is English so hard to follow in a bar, or when several people talk at once?
Because noise and competing voices cost a learner far more than a native speaker. In one study, learners needed the speaker about five decibels louder than native listeners did to understand the same sentences against background chatter [12]. In another, students’ understanding of short conversations fell from 77% in quiet to 57% with two people talking over the recording [13].
The reason is knowledge, not ears. Two reviews of the research reach the same conclusion: the extra difficulty does not come from noise acting differently on a learner’s hearing [14] [15]. What differs is the repair. A native listener rebuilds the missing pieces from vocabulary and context. A learner has less to rebuild with. When one study compared listeners with matched vocabulary, the gap between the groups fell from nine points to five [16].
Other voices are worse than steady noise [13]. And talk you can understand is harder to ignore than talk you cannot. Across 128 bilingual listeners, background chatter in their own first language was the hardest of all [17]. The people speaking your language at the next table are the toughest background there is.
Even when learners keep up, it costs them more effort [12]. If English in a crowd wears you out, that is real.
Can I get better at understanding English in noisy places?
Yes. Research suggests the gain comes mostly from learning the words, the phrases and the voices. A learner’s level of English is closely tied to how well they cope with noise [18], and a familiar voice is easier to follow against competing talk than a stranger’s [19]. Whether practising with noisy audio beats practising with clean audio has not been tested directly.
I know this one from the inside. When I first came to Spain and was learning Spanish, two situations left me almost wanting to cry: meal times and bars. Family meals were impossible to follow. Everyone talked at once, over each other. At first I remember just zoning out, and it all became noise. But little by little, the more of those situations I sat through, the better my ear got at recognising what was being said. Getting used to those meals made my Spanish much better. For a long time I believed the noise itself had trained my ear. Here is how that story holds up.
“Impossible at first” is well supported. A table where everyone talks at once is the hardest case in the research: several competing voices, in the language you are trying to follow [13].
“You get better” is supported too, with a limit. Two things were happening at that table. My Spanish was growing, and level goes hand in hand with coping in noise [18]. And I was learning those particular voices. In one experiment, people followed their husband’s or wife’s voice against a competing talker better than a stranger’s [19]. The limit is that the gap narrows without fully closing. People who learned English before the age of six, and had no noticeable accent, matched single-language speakers in quiet and still trailed them in noise [20]. If a loud bar is harder in English than in your own language after many years, that is normal.
“The noise trained my ear” has not been shown. There is some support. Among 51 Chinese learners, those who practised English vowels in noise still showed their gains three months later, and those who practised in quiet did not [21]. Native listeners who trained against background talk improved most against background talk [22]. But Spanish learners trained on English sounds improved substantially whether they practised in quiet or in noise, and they did not become generally used to noise [23]. The field’s main review says “the role of formal training in noise is currently unknown” [15]. No study has yet taught learners with noisy conversations and with clean ones, then tested both groups somewhere new.
So here is what I can honestly say. Practice with real background sound prepares you for the situations you will meet. Laboratory studies show that practice against background talk helps most with background talk, and that a familiar voice is easier to follow in a crowd. That dinner table made me better. The evidence suggests it was mostly the Spanish I was learning and the voices I got to know, more than the noise.
Does hearing different accents help me understand new ones?
It helps with the accents you hear and, to some extent, with new speakers. Even light unfamiliar accents lowered comprehension in a study of 21,726 test-takers, and familiarity with the accent raised it [24]. Training with several voices improves how learners hear English sounds, across 79 studies [25]. Evidence that it makes brand-new accents easy is thin.
The extra voices helped upper-intermediate and advanced learners most [25]. And regional accents are less of a wall than they seem. In a 2025 study of upper-intermediate Spanish and Dutch learners, what cost them most was casual, shortened pronunciation. That cost was the same whether the speaker had a standard British accent or a Newcastle one [26].
So the remedy for an accent is familiarity with it. That is what a regular team gives you. You get used to our teachers’ voices, which is a real gain in itself. Then the guests widen it, with British and North American speakers from different regions.
Do “um” and “er” make English harder to understand?
No. For native listeners, hesitations do not get in the way, and short ones give a head start. In one experiment, people recognised the next word about 47 milliseconds faster after an “uh” than when the “uh” had been cut out of the recording [27]. Whether learners get the same benefit has not been tested, which is one reason the transcript matters.
Hesitating is not bad English. In a large collection of recorded British conversations, speakers produced about two “ums” and “uhs” in every hundred words, and some people used them many times more than others. An “uh” tended to announce a short delay and an “um” a longer one [28].
How much of real English is fixed expressions?
A great deal. English is highly phraseological: by one count, more than half of ordinary speech is built from ready-made phrases [29]. Across 35 studies, these chunks were processed faster than new combinations of words, by native speakers and by learners [30]. At natural speed you need to know the phrase, not only the words in it.
That count needs care. It rests on a small sample, judged partly by intuition [29]. And it is not a count of idioms. It covers every combination of words that speakers reuse as a ready-made unit, from everyday phrases to colourful sayings.
The colourful ones still matter, because you cannot decode them word by word. If you do not know “quid” and “loaded”, you cannot follow people talking about money. So each of our Colloquial episodes takes around ten expressions out of a real conversation and works through them one at a time.
Hear it: money expressions such as “quid”, “dosh” and “loaded”, recorded in a park while a small dog eats a plastic bottle. Free.
How much vocabulary do I need to follow a real conversation?
Fewer words than you need for reading. About 3,000 word families cover 95% of the words in unscripted conversation [31]. In listening experiments, 95% was enough for good, consistent comprehension [32]. And vocabulary goes hand in hand with listening ability: the link holds across roughly 21,000 learners [33].
A word family is a word and its close relatives: help, helps, helped, helpful.
Treat the figure as a guide, not a gate. It comes from learners hearing short stories, not conversations in a café [32]. It also counts an expression as known if you know its separate words. Knowing 3,000 word families does not mean you know the colloquial phrases built from them.
That is why our episodes are graded. Level 2 is for pre-intermediate and intermediate learners (B1). Level 3 is for upper-intermediate and advanced learners (B2 to C1). If you get lost, drop a level.
Start at your level: B1 episodes or B2–C1 episodes.
Does listening to a lot of English really work?
Yes. Listening a lot is essential, though it is not enough on its own. In Belgium, a quarter of 780 children had reached A2 level in listening from the English they met outside school, before their first lesson [34]. But listening with the text beat listening alone in a 13-week study [35], and teaching people how to listen adds a further gain across 45 studies [36].
The best-known name here is Stephen Krashen, who argued that understanding what you hear is what drives learning. Even he described understood input as “necessary (but not sufficient)” [37]. Later research agrees that guidance adds to it [36].
So listen a lot. There are more than 150 free episodes on this site, most of them between ten and twenty minutes long.
Listen free: all the podcasts.
How are the Zapp! English recordings made?
No scripts and no actors. Real people talking in real places, at natural speed, with the transcript and exercises that make it learnable.
- Real places. A park, where a parrot interrupts four people mid-conversation. The beach on a hot morning. A table outside a café. A garden with birds behind the voices.
- Two to four speakers in a conversation, who hesitate, restart, laugh and finish each other’s sentences.
- Natural speed.
- Teachers you get to know. A regular team, plus guests, with British and North American accents. Meet the team.
- Honest about what is scripted. The conversations are not. The short teaching sections that set up each task and give the answers are, so the instructions stay clear.
This is a realistic step, not the deep end. In measured conversations, about four in ten changes of speaker involve some overlap [2]. But the overlaps are short, and most of the time one person is speaking [3]. Three friends in a park are far easier to follow than a family dinner. You get the hesitations, the speed and the expressions, with a few seconds of overlap and the occasional parrot.
When is real English too hard, and what helps?
It is too hard when you cannot tell what you missed. The answer is support in the right order, not easier English. In a study of 160 students, a short introduction to the topic and a second listening helped most, and pre-teaching a vocabulary list helped least [38]. In a laboratory study, reading the words in English helped listeners tune in to an unfamiliar accent, while reading them in their own language got in the way [39].
So this is the order we recommend:
- Read the short introduction to the topic. Do not study a word list first.
- Listen straight through for the gist. No dictionary and no pausing.
- Listen again. You will catch more the second time.
- Open the transcript and listen while you read. Find the words you knew and did not hear.
- Listen once more without the text. That is the listening you are training for.
- Then do the vocabulary and the expressions, now that you have heard them in use.
The free episodes give you steps 1, 2, 3 and 5. The workbooks add what steps 4 and 6 need: the full transcript in English, a glossary, and exercises with answers, for every episode.
Get the complete courseAll workbooks and audio — €39
Frequently asked questions
Which English accent is hardest to understand?
The one you have heard least. Research does not rank accents from easy to hard. It shows that an unfamiliar accent lowers comprehension, even a light one, and that familiarity raises it [24]. Casual, shortened pronunciation is often the real obstacle, in any accent [26].
Why can I understand films with subtitles but not without?
Because the subtitles are recognising the words for you. Across 18 studies, same-language subtitles improved comprehension and vocabulary [40]. Without them you depend on catching words at speed, which is the commonest difficulty learners report [5]. Use English subtitles, not ones in your own language [39], and then practise without.
How long does it take to get used to native speakers?
Research gives no fixed number. Getting used to one voice is the quick part, and a familiar voice is easier to follow in a crowd [19]. Learning the words and phrases is the slow part, and it matters most [18]. Even people who learned English as small children find noise harder than single-language speakers do [20].
Should I start with clean audio or real audio?
Start at your level, with support. The European scale expects clearly articulated speech at B1, and places overlapping conversation at natural speed at C1 [6]. So from B1, use real conversations with a transcript and a second listening [38]. Whether practising with background noise beats practising with clean audio has not been tested directly.
Are hesitations a sign of bad English?
No. Native speakers hesitate, repeat or restart five or six times in every hundred words [1], and say “um” or “uh” about twice in every hundred [28]. For native listeners, a short hesitation even speeds up recognition of the next word [27].
Is it better to listen to one speaker or many?
Both, in that order. One familiar voice is easier to follow and a good way in [19]. Several voices train the ear more widely, and the benefit was clearest for upper-intermediate and advanced learners [25].
Sources
- Bortfeld, H., Leon, S. D., Bloom, J. E., Schober, M. F., & Brennan, S. E. (2001). Disfluency rates in conversation: Effects of age, relationship, topic, role, and gender. Language and Speech, 44(2), 123–147. doi:10.1177/00238309010440020101
- Heldner, M., & Edlund, J. (2010). Pauses, gaps and overlaps in conversations. Journal of Phonetics, 38(4), 555–568. doi:10.1016/j.wocn.2010.08.002
- Levinson, S. C., & Torreira, F. (2015). Timing in turn-taking and its implications for processing models of language. Frontiers in Psychology, 6, 731. doi:10.3389/fpsyg.2015.00731
- Johnson, K. (2004). Massive reduction in conversational American English. In K. Yoneyama & K. Maekawa (Eds.), Spontaneous Speech: Data and Analysis (pp. 29–54). Tokyo: National Institute for Japanese Language. Full text (PDF)
- Goh, C. C. M. (2000). A cognitive perspective on language learners’ listening comprehension problems. System, 28(1), 55–75. doi:10.1016/s0346-251x(99)00060-3
- Council of Europe (2020). Common European Framework of Reference for Languages: Learning, teaching, assessment — Companion volume. Strasbourg: Council of Europe Publishing. Full text (PDF)
- Gilmore, A. (2004). A comparison of textbook and authentic interactions. ELT Journal, 58(4), 363–374. doi:10.1093/elt/58.4.363
- Le Foll, E. (2021). Register variation in school EFL textbooks. Register Studies, 3(2), 207–246. doi:10.1075/rs.20009.lef
- Cullen, R., & Kuo, I.-C. V. (2007). Spoken grammar and ELT course materials: A missing link? TESOL Quarterly, 41(2), 361–386. doi:10.1002/j.1545-7249.2007.tb00063.x
- Wagner, E., & Toth, P. D. (2014). Teaching and testing L2 Spanish listening using scripted vs. unscripted texts. Foreign Language Annals, 47(3), 404–422. doi:10.1111/flan.12091
- Gilmore, A. (2011). “I prefer not text”: Developing Japanese learners’ communicative competence with authentic materials. Language Learning, 61(3), 786–819. doi:10.1111/j.1467-9922.2011.00634.x
- Borghini, G., & Hazan, V. (2018). Listening effort during sentence processing is increased for non-native listeners: A pupillometry study. Frontiers in Neuroscience, 12, 152. doi:10.3389/fnins.2018.00152
- Yang, X., Jiang, M., & Zhao, Y. (2017). Effects of noise on English listening comprehension among Chinese college students with different learning styles. Frontiers in Psychology, 8, 1764. doi:10.3389/fpsyg.2017.01764
- Scharenborg, O., & van Os, M. (2019). Why listening in background noise is harder in a non-native language than in a native language: A review. Speech Communication, 108, 53–64. doi:10.1016/j.specom.2019.03.001
- García Lecumberri, M. L., Cooke, M., & Cutler, A. (2010). Non-native speech perception in adverse conditions: A review. Speech Communication, 52(11–12), 864–886. doi:10.1016/j.specom.2010.08.014
- Schmidtke, J. (2016). The bilingual disadvantage in speech understanding in noise is likely a frequency effect related to reduced language exposure. Frontiers in Psychology, 7, 678. doi:10.3389/fpsyg.2016.00678
- Lew, E., Hallot, S., Byers-Heinlein, K., & Deroche, M. (2025). Navigating the bilingual cocktail party: A critical role for listeners’ L1 in the linguistic aspect of informational masking. Bilingualism: Language and Cognition, 28(3), 748–756. doi:10.1017/s1366728924000944
- Kilman, L., Zekveld, A., Hällgren, M., & Rönnberg, J. (2014). The influence of non-native language proficiency on speech perception performance. Frontiers in Psychology, 5, 651. doi:10.3389/fpsyg.2014.00651
- Johnsrude, I. S., Mackey, A., Hakyemez, H., Alexander, E., Trang, H. P., & Carlyon, R. P. (2013). Swinging at a cocktail party: Voice familiarity aids speech perception in the presence of a competing voice. Psychological Science, 24(10), 1995–2004. doi:10.1177/0956797613482467
- Rogers, C. L., Lister, J. J., Febo, D. M., Besing, J. M., & Abrams, H. B. (2006). Effects of bilingualism, noise, and reverberation on speech perception by listeners with normal hearing. Applied Psycholinguistics, 27(3), 465–485. doi:10.1017/s014271640606036x
- Mi, L., Tao, S., Wang, W., Dong, Q., Dong, B., Li, M., & Liu, C. (2021). Training non-native vowel perception: In quiet or noise. The Journal of the Acoustical Society of America, 149(6), 4607–4619. doi:10.1121/10.0005276
- Van Engen, K. J. (2012). Speech-in-speech recognition: A training study. Language and Cognitive Processes, 27(7–8), 1089–1107. doi:10.1080/01690965.2012.654644
- Cooke, M., & García Lecumberri, M. L. (2018). Effects of exposure to noise during perceptual training of non-native language sounds. The Journal of the Acoustical Society of America, 143(5), 2602. doi:10.1121/1.5035080
- Ockey, G. J., & French, R. (2016). From one to multiple accents on a test of L2 listening comprehension. Applied Linguistics, 37(5), 693–715. doi:10.1093/applin/amu060
- Uchihara, T., Karas, M., & Thomson, R. I. (2025). High variability phonetic training (HVPT): A meta-analysis of L2 perceptual training studies. Studies in Second Language Acquisition, 47(3), 794–827. doi:10.1017/s0272263125100879
- Verbeke, G., Mitterer, H., & Simon, E. (2025). Phonetic reduction in native and non-native English speech: Assessing the intelligibility for L2 listeners. Bilingualism: Language and Cognition, 28(5), 1266–1280. doi:10.1017/s1366728925000021
- Fox Tree, J. E. (2001). Listeners’ uses of um and uh in speech comprehension. Memory & Cognition, 29(2), 320–326. doi:10.3758/bf03194926
- Clark, H. H., & Fox Tree, J. E. (2002). Using uh and um in spontaneous speaking. Cognition, 84(1), 73–111. doi:10.1016/s0010-0277(02)00017-3
- Erman, B., & Warren, B. (2000). The idiom principle and the open choice principle. Text, 20(1), 29–62. doi:10.1515/text.1.2000.20.1.29
- Yi, W., & Zhong, Y. (2024). The processing advantage of multiword sequences: A meta-analysis. Studies in Second Language Acquisition, 46(2), 427–452. doi:10.1017/s0272263123000542
- Nation, I. S. P. (2006). How large a vocabulary is needed for reading and listening? Canadian Modern Language Review, 63(1), 59–82. doi:10.3138/cmlr.63.1.59
- van Zeeland, H., & Schmitt, N. (2013). Lexical coverage in L1 and L2 listening comprehension: The same or different from reading comprehension? Applied Linguistics, 34(4), 457–479. doi:10.1093/applin/ams074
- Zhang, S., & Zhang, X. (2022). The relationship between vocabulary knowledge and L2 reading/listening comprehension: A meta-analysis. Language Teaching Research, 26(4), 696–725. doi:10.1177/1362168820913998
- De Wilde, V., Brysbaert, M., & Eyckmans, J. (2020). Learning English through out-of-school exposure. Which levels of language proficiency are attained and which types of input are important? Bilingualism: Language and Cognition, 23(1), 171–185. doi:10.1017/s1366728918001062
- Chang, A. C.-S., & Millett, S. (2014). The effect of extensive listening on developing L2 listening fluency: Some hard evidence. ELT Journal, 68(1), 31–40. doi:10.1093/elt/cct052
- Dalman, M., & Plonsky, L. (2025). The effectiveness of second-language listening strategy instruction: A meta-analysis. Language Teaching Research, 29(3), 1039–1068. doi:10.1177/13621688211072981
- Krashen, S. D. (1982). Principles and Practice in Second Language Acquisition. Oxford: Pergamon Press (internet edition, 2009). Full text (PDF)
- Chang, A. C.-S., & Read, J. (2006). The effects of listening support on the listening performance of EFL learners. TESOL Quarterly, 40(2), 375–397. doi:10.2307/40264527
- Mitterer, H., & McQueen, J. M. (2009). Foreign subtitles help but native-language subtitles harm foreign speech perception. PLoS ONE, 4(11), e7785. doi:10.1371/journal.pone.0007785
- Montero Perez, M., Van Den Noortgate, W., & Desmet, P. (2013). Captioned video for L2 listening and vocabulary learning: A meta-analysis. System, 41(3), 720–739. doi:10.1016/j.system.2013.07.013

