Intermediate Plateau

You kept your end of the deal. The lessons got done most days, the units piled up, and for the first year the progress was real: you could order food, read simple texts, follow slow speech, survive small talk that stayed on script. Then, somewhere past the beginner material, the app kept assigning lessons and you stopped getting better.
Most people read that as a personal failure, a sign they lack the discipline or the talent to go further. The more accurate reading is structural. Mainstream language apps are engineered for the beginner stage, and the clearest evidence for that is not a critic’s opinion. It is the research the app companies have published about their own products. Here is what that evidence shows, and what the research says works once you are past the level the apps were built for.
What the app efficacy studies actually measured
When the major apps cite research showing their products work, the studies behind the claims are worth reading closely. The best known are a series run between 2009 and 2018 covering Rosetta Stone, Duolingo, Babbel, Busuu, and Italki. As a peer-reviewed 2022 review by Duolingo’s own research team documents, those studies were commissioned by the companies, published on company websites as white papers rather than in peer-reviewed journals, and used the same lead researcher across five competing products. All of them measured progress with the WebCAPE, a placement exam that tests vocabulary, reading, and grammar. None of them assessed speaking or writing.
That framing matters for a number you may have seen quoted: the 2012 finding that roughly 34 hours of Duolingo produced a placement-score gain equivalent to a first college semester of Spanish. Those 34 hours moved a score on a vocabulary, reading, and grammar test. Whether anyone could speak afterward was never measured. Independent scholars reviewing the series later noted the studies “would have benefited from more rigorous research designs,” including control of prior proficiency and time on task (same review, pp. 3–4).
To be fair to the category, the one major study in this family that did assess speaking, a 2020 university collaboration with Babbel, found learners gained 0.7 of an ACTFL sublevel in oral proficiency after roughly 12 hours of study over 12 weeks. That is a real gain, and a modest one, in college students who mostly had prior Spanish coursework.
The pattern to hold onto: the studies behind the apps’ effectiveness claims almost never asked a learner to speak, and the one that did found modest movement.
Where the courses actually stop
The ceiling is not hidden. It is in the product documentation.
"Half of B1"
how much of the intermediate band Duolingo says its Spanish and French learners have covered at the Unit 7 milestone
B2
the highest level Babbel offers for most of its major languages
B1
the top of Memrise's official course sequence, per third-party summaries
Duolingo’s own blog describes the milestone at Unit 7 of its flagship Spanish and French courses this way: “learners have completed all of A1 and A2 and half of B1.” Babbel’s course documentation runs from beginner through upper intermediate for its major languages, with more advanced content for a few. In other words, even by the companies’ own maps, the roads end at or just inside the territory where you are now stuck.
The outcome data tells the same story. In 2024, Duolingo published a whitepaper assessing learners who completed the beginner-level content of its Spanish and French courses across all four skills. Spanish course completers reached Intermediate Mid overall on the ACTFL scale, French completers Intermediate Low. Two caveats belong next to those results. ACTFL is a different framework from the European scale most learners know, and the commonly used conversions, which place these outcomes around upper beginner to low intermediate, are approximations rather than official equivalences. And these are respectable results for beginner courses; Duolingo’s researchers reasonably presented them as meeting expectations. The gap is what sits above them: the company has not published an equivalent independently assessed speaking and writing evaluation for learners who complete its intermediate content.
Why the stall is structural, not personal
Knowing where the courses stop explains the ceiling. It does not yet explain why your last six months of lessons felt like running in place. Four design choices, each sensible for beginners, account for that.
Tapping is recognition; speaking is production
The core exercise types in mainstream apps, word tiles, multiple choice, fill-in-the-blank, matching pairs, all train you to recognize language someone else has assembled. Spaced-repetition review, the engine under most of them, is excellent at keeping recognition fresh and does far less for your ability to generate a word or structure from nothing, mid-sentence, under time pressure. Knowing a grammar rule and using it fluidly are two separate cognitive processes, and app formats exercise the first almost exclusively. This is why you can clear review sessions comfortably and still stall in a real conversation: you have been getting genuinely better at the app’s task, which is not the task you care about.
App sentences are tidy; real language is not
App exercises present language in short, predictable, self-contained units, which makes them easy to design and measure, and unrepresentative of language in use. It is like practicing tennis against a ball machine that always feeds to your forehand: your forehand improves, and a real opponent still beats you, because the actual game is reading what is coming and responding to it. Real language runs in paragraphs and conversations, where pronouns refer back across sentences, tone shifts with the listener, and meaning accumulates. Intermediate ability is largely the ability to manage that extended discourse, and sentence-level drills never ask you to.
The grammar problem changes
At the beginner stage, the main cognitive job is retrieval: find the word, apply the pattern. From intermediate on, the job becomes choice: which past tense, which mood, which register for this listener, with dependencies running across clauses. Pattern matching handles “I eat an apple.” It does not handle a conditional apology to your partner’s grandmother. Practicing those choices is uncomfortable, error-filled, and repetitive in ways that engagement-optimized products are built to avoid, and that tension, between what keeps users comfortable and what intermediate learning demands, sits at the center of the stall.
The motivation engine rewards showing up, not improving
The daily streak, the apps’ signature retention feature, is built openly on loss aversion: the sting of losing a long streak keeps people opening the app. A widely observed pattern follows, with users logging in to protect the streak, tapping through a minimal lesson, and leaving. The research on these mechanics has similar conclusions: a 2023 meta-analysis found reward mechanics improve intrinsic-motivation and autonomy measures but have “minimal impact on competency.” Durable motivation, in the self-determination literature, runs on competence, the felt experience of getting better at the thing itself. A streak can carry you through the beginner stage. It cannot substitute for visible skill growth once the easy gains are gone, which is often when the habit quietly ends.
“You did not hit your ceiling. You reached the end of what the tool was built to teach.”

Beginner apps pave the road exactly as far as their curricula go: recognition, translation, sentence-level drills. The stall most learners feel around the intermediate level is the end of that pavement. The path onward is built from different materials: producing the language, getting corrected, and sustaining real conversation.
What the research supports from intermediate on
The clearest account of what the apps leave out is decades old. Merrill Swain studied French immersion students in Canada who had received years of rich, comprehensible French and still could not produce it accurately; they understood far more than they could say. Her Output Hypothesis named the reason: producing language does cognitive work that comprehension cannot. Speaking and writing force retrieval rather than recognition, show you exactly where your gaps are, and test whether the grammar in your head matches the grammar of the language. Read that list against the previous section and it is nearly point for point what app exercises never require.
Around that spine, the research converges on a practice model for this stage:
Move some hours to production. Speak and write on purpose, at and slightly beyond your comfort level, rather than waiting for conversation to happen to you. Pushed output, production a step beyond what comes easily, is where the growth is.
Get corrected, immediately and specifically. Learners who receive corrective feedback on their output show reliably greater grammatical accuracy than learners relying on input alone, and immediate feedback outperforms delayed. The research behind this is covered in our guide to how adults acquire a language; the practical version is that output only counts as practice when something flags your errors before they harden.
Seek real interaction. Conversation in which meaning has to be negotiated, where you rephrase, ask, and repair misunderstandings, is one of the core mechanisms of acquisition, and it is the one thing no self-paced exercise can simulate.
Practice at the edge of your ability. Effective practice at this level is deliberate, systematic, and feedback-driven, with difficulty you can just barely handle. Comfortable review of what you already know is the slowest possible route forward.
Keep input, aimed just above your level. Reading and listening still matter; what changes is what you ask them to do. Narrow listening, staying with one topic or speaker so vocabulary recycles, keeps authentic material comprehensible. Whether input alone can carry acquisition remains genuinely contested; what is not contested is that it will not build your speaking on its own.
From intermediate on, progress tracks what you produce and get corrected on, not what you consume or review.
Where this leaves your app, and where Dioma fits
None of this argues for deleting the app in disgust. It did the job it was designed for, and if its review sessions still earn five pleasant minutes as gamified vocabulary maintenance, keep them. Different tools serve different stages.
A working setup from this stage on usually has three parts: input from native content slightly above your level, real conversation with actual people, and structured output with correction. The third is the hardest to arrange on your own, and it is the part Dioma is built for. You speak and write against an expert-built curriculum designed for intermediate and advanced learners, every response is corrected in the moment with the rule attached rather than just the right answer, and the system tracks which patterns you personally keep missing so practice aims at your actual weak points. Dioma is one part of that three-part setup, not a replacement for the other two.
A fair way to start: pick the thing you actually want to be able to do, say, five unrehearsed minutes about your week, and make that the unit of practice instead of the lesson. Then make sure every attempt gets corrected.
The real shift
The stall you are feeling is not evidence about you. The apps’ own studies measured recognition, their own documentation marks where the courses end, and their own outcome data shows speaking trailing every other skill. You walked the paved road to its end, which is what the first year of progress was. What lies past it is built differently: language you produce, errors that get corrected, conversations that push back. The learners who keep improving from here are not the ones who found a better app for the old kind of practice. They are the ones who changed what practice means.
Frequently asked questions
Why do language apps stop working around B1?
Can Duolingo or Babbel get you to B2?
Should I stop using my language app at the intermediate level?
What should I do instead of language apps after B1?
Do streaks and daily goals actually help you learn a language?
Further reading
Jiang et al. (2022), Foreign Language Annals
Peer-reviewed assessment
The study that measured what beginner app courses actually deliver, and reviewed the commissioned white papers that came before it.
Full text (PDF)
Duolingo Research Report DRR-24-04 (2024)
Primary source
Duolingo's own four-skill proficiency assessment of learners completing its beginner content.
Whitepaper (PDF)
Dioma: How Adults Actually Acquire a Language
Related guide
The input, output, and feedback loop in full, with the research behind corrective feedback.
Read
Dioma: The Intermediate Plateau: Why You Stopped Improving
Related guide
Why progress feels like it vanished at this stage even when learning continues.
Read
Related Articles


