Zaniar Shokati
← Back to writingLanguage Learning

318,550 People Took the DTZ. The Result Still Couldn't Tell Them What to Fix.

Aug 3, 2026

In 2025, Germany's Federal Office for Migration and Refugees recorded 318,550 people with a result from the Deutsch-Test für Zuwanderer (DTZ), the language test used at the end of an integration course. Of those people, 175,193 reached B1, 102,663 reached A2, and 40,694 remained below A2.

So 55% reached B1. Another 32.2% received an A2 result, and 12.8% remained below A2.

There is an important footnote behind those numbers. Since 2018, the BAMF has counted people, not individual test attempts, and records the highest level each person achieved if they took the DTZ more than once. This is therefore not a first-attempt pass rate. And saying that the remaining 45% simply "failed" would be too blunt: A2 is a reported DTZ level, even though B1 is the language target required for successful completion of the standard integration course. The BAMF also notes that A2 is the realistic curricular target for many participants in literacy courses.

That makes the statistic more honest, but it does not make it more diagnostic. It tells us the level someone eventually reached. It cannot tell an individual learner whether the recurring problem was misunderstanding the task, missing a required content point, weak vocabulary, grammar, timing, or something else.

That gap between the result and what to change next is the problem I built DeutschGenie to work on.


Official exams give more feedback than one number

I want to get this distinction right, because German language exams are not all structured the same way.

The Goethe-Zertifikat B1 has four independently scored modules: Reading, Listening, Writing, and Speaking. Each module is passed with at least 60 points out of 100.

telc Deutsch B1 is different. Its result is broken down by subtest, but passing is decided at the level of the written and oral examinations: a candidate needs at least 60% in each. Under telc's current rules, B1 is one of the levels for which a passed partial examination can be credited when the failed part is retaken within the permitted period.

The DTZ is different again. Its result reports Listening and Reading, Writing, and Speaking separately. If someone repeats the DTZ, however, they must repeat the complete written and oral test; earlier partial results are not carried forward.

So an official result can already tell a learner much more than "your German is not B1." It can identify a weak module, component, or subtest. That is useful information.

But it usually stops one level above the question a learner asks next. "Writing was weak" does not tell you whether the main problem was task fulfillment, organization, vocabulary, or language accuracy. "Reading was weak" does not tell you whether you misunderstood the text or repeatedly chose distractors because of the same vocabulary or attention pattern.

That is the narrower claim I am comfortable making: certification exams report performance at the level their scoring system is designed to report. They are not individualized teaching sessions, and they are not designed to reconstruct every cause behind every lost point.


The cost of another attempt depends on the exam

I compressed several different fee systems into one rule the first time I worked through this argument. They do not share one.

For eligible integration-course participants, the BAMF covers the first DTZ participation. It can also cover a further attempt in defined circumstances, including when the first attempt happened before the participant used the available course hours or when an approved course repetition includes another test. A learner can repeat the DTZ more often at their own expense, but each repeat is the complete test.

Goethe works differently. In Germany, the 2026 price for the Goethe-Zertifikat B1 is €259 for the complete exam or €104 for one module. People who attended a Goethe-Institut course no more than six months before the exam receive a 20% discount, bringing those prices to €207.20 and €83.20 respectively. Those prices are specific to Germany and 2026; Goethe fees in other countries are set locally.

telc does not publish one nationwide candidate fee. Examination centers set their own prices. The retake rules also depend on the exam level, so a percentage presented as a universal "telc retake price" would be misleading.

The broader point survives the correction. Another attempt can mean another fee, another registration, and another wait for results. But the exact burden depends on the exam, the center, and whether public funding applies. That is less tidy than one dramatic number, and it is also more useful.


What the research supports — and what it does not

The argument for diagnostic feedback is not new. In 1992, Elana Shohamy proposed a diagnostic-feedback model that went beyond a general proficiency result and analyzed performance at progressively more specific levels. The basic idea still holds up: information about where a learner stands is not automatically information about what they should work on next.

There is also evidence that the source of feedback and the learning outcome do not always move together. A 2022 study compared two intact classes of Chinese university students learning English, with 35 students in each class. The teacher-feedback group produced stronger immediate revisions and rated the feedback as more useful and easier to use. The automated-feedback group scored higher on a later writing-proficiency post-test.

That is interesting, but it is not a blank cheque for automated feedback. It was one small, non-randomized study in one EFL setting, using a particular automated writing system. The authors themselves discuss possible explanations that still need verification. I would not turn it into "AI feedback is better than teacher feedback," because the study does not establish that.

What it does show is that "Which feedback did learners prefer?", "Which feedback improved this draft?", and "Which feedback improved a later piece of writing?" are different questions. A product should be judged against the outcome it actually claims to improve.


The part where AI feedback can go wrong

Diagnostic feedback is only useful when the diagnosis is credible.

A 2026 workshop paper on AI feedback in language-learning systems makes this problem concrete. Its proposed evaluation framework includes diagnostic accuracy, awareness of appropriacy, causes of error, prioritization, guidance for improvement, and support for self-regulation. The authors point out that the same visible error can have several possible causes. A wrong verb form, for example, might come from not knowing the form, misunderstanding the context, or overlooking agreement. Those causes call for different teaching responses.

The paper also warns about systems giving a definite diagnosis when several interpretations are plausible. That matters because a confident but wrong explanation can send a learner toward the wrong lesson.

There is a caveat here too: this is a workshop paper presenting an evaluation framework and preliminary observations, not a clinical-style trial proving the effectiveness or failure rate of every AI language product. I use it as a design warning, not as proof that attaching an explanation to a score automatically makes that explanation trustworthy.

That warning influenced DeutschGenie's design. The product covers practice across Reading, Listening, Writing, and Speaking, but it does not pretend those skills share one universal rubric. Reading and listening can be scored against expected answers and analyzed for recurring diagnostic patterns. Writing and speaking receive AI feedback across criteria relevant to productive language, such as task fulfillment, grammar, structure, vocabulary, fluency, pronunciation, or register where applicable. When a diagnostic tag is inferred rather than directly observed, the system can present the confidence behind that inference instead of treating every explanation as equally certain.

That is still not the same as being correct. Confidence is a communication mechanism, not a substitute for validation. It is also why DeutschGenie offers human expert review for writing and speaking feedback: automation can make detailed practice feedback available quickly, while a qualified person remains valuable when a judgment is ambiguous or consequential.


What I actually built it for

DeutschGenie is a practice tool. It does not replace Goethe, telc, ÖSD, the DTZ, or the institutions that issue recognized certificates.

Its job begins after a practice answer has been scored. It keeps the categories behind that result separate and tracks them across attempts, so a learner can look for a pattern instead of reacting to one total. Maybe vocabulary is not the recurring problem at all. Maybe the learner repeatedly misses a required point, misreads the instruction, or produces grammatically sound language that does not answer the task.

A single result tells you what happened once. Several attempts, analyzed consistently, can support a more useful hypothesis about what keeps happening.

I am careful with the word hypothesis. An automated diagnosis can be wrong, and a practice platform cannot promise the result of an official exam. The honest value proposition is smaller: give learners more specific evidence, show uncertainty where it exists, and make the next practice decision less random.

The exam tells you the level you reached. That is what a certification exam is for. DeutschGenie is built for the question that comes immediately after: what should I change before the next attempt?


Sources I leaned on