Back to Blog
forschung4 views

AI and the psychopathological assessment: what the ZI study shows and what it doesn't

A team at the ZI Mannheim pitted ten language models against 108 practising clinicians, rating all 100 items of the AMDP system. The accuracies sit close together; more revealing are the opposite error profiles: clinicians infer symptoms from incomplete information, models conservatively mark "not a

Psynex Team

A study has appeared in npj Digital Medicine that bears directly on psychopathological assessment. A team led by the Central Institute of Mental Health (ZI) in Mannheim pitted ten large language models against 108 practising clinicians. What was rated was neither a diagnosis nor a questionnaire score, but each of the 100 individual items of the AMDP system.

The summary that suggests itself is: AI performs at the level of clinicians. Put that way, it is not correct. The real finding of the paper lies elsewhere, namely that models and humans make systematically different errors.

What was studied

The material consisted of three semi-structured psychiatric interviews of 24 to 33 minutes. A trained simulated patient portrayed a depressive, a manic and a schizophrenic syndrome, without a script and after roughly one hour of clinical briefing per video. The interviews were conducted exclusively by AMDP-trained psychiatrists.

Three rating pathways were applied to this material:

  • Reference standard: a consensus panel of three experienced psychiatrists, all of whom contributed to the current edition of the AMDP manual and teach it. They first rated each video independently and then reached a unanimous consensus in discussion. This took around three hours per video.
  • Clinicians: 108 people from three psychiatric hospitals in Germany, a university hospital, a university-affiliated research hospital and a regional care centre. They were recruited between April and November 2025 during ongoing training sessions. This group watched the full video recordings with picture and sound and completed AMDP forms on paper, within the running time of the video plus 15 minutes.
  • Language models: ten models, selected by their ranking in the LMSYS Chatbot Arena as of November 2025. They received only the transcript, produced locally with Whisper large-v3, with speaker attribution, without any description of behaviour, and manually corrected against the audio.

Ratings were made on the reduced three-level AMDP scale: absent, present, not assessable. This reduction follows the AMDP recommendation for diagnostic practice, because the common classification systems only require the presence of an item for a criterion to be met. The basis was the 11th German edition with its 100 items.

The numbers

Across all three scenarios, GPT-5.1 and Gemini-3-Pro-Preview reached the highest agreement with the expert consensus, at 0.72. The clinician mean was 0.68. For the main analysis the team selected GPT-5.1, because it was numerically slightly ahead in two of the three scenarios. The calculation used a majority vote over three independent runs.

Broken down by scenario, model versus clinician mean in each case:

  • Depression: 0.81 versus 0.79 (p = 0.62)
  • Mania: 0.76 versus 0.68 (p = 0.03)
  • Schizophrenia: 0.60 versus 0.58 (p = 0.68)

None of these differences is significant after Bonferroni correction for multiple testing; the threshold was p < 0.017. Even the nominally striking value in the mania scenario falls short of it. The study therefore does not demonstrate superiority of the models, but a comparable order of magnitude. Overall, the model sits at the 64th percentile of the distribution of participating clinicians.

That distribution is wide. The accuracy of individual raters across all videos ranged from 0.36 to 0.89, and pairwise agreement among them was 65.4 percent. For putting the model performance in context, this is at least as informative as the mean: even among professionals, the same assessment comes out very differently.

Who the comparison group was

This point decides how the study should be read, and it regularly gets lost in short news items. The 108 raters were predominantly early-career professionals:

  • 71.3 percent physicians in specialist training, 11.1 percent psychologists
  • 6.5 percent board-certified psychiatrists, 3.7 percent senior consultants
  • on average 5.4 years of psychiatric experience, median 3.0 years, mode 1.0 year
  • 17.6 percent AMDP-certified

All participants used AMDP routinely, 72.4 percent as their preferred instrument, at an average of 36.3 psychopathological assessments per month. So these were practised but young users.

The paper itself spells out the consequence: how the models perform against experienced, AMDP-certified psychiatrists cannot be derived from these data, because that group was too small in the sample. In the schizophrenia scenario it was three people. A sample with a higher proportion of experienced AMDP specialists would presumably raise the accuracy of the comparison group. The sentence "AI as good as doctors" is therefore not supported. What is supported: comparable to a predominantly young comparison group, in three simulated scenarios.

The real finding: opposite error profiles

More interesting than the hit rate is the question of what each side gets wrong. Here the patterns diverge clearly.

Counting the items where the clinicians' most frequent rating deviated from the expert consensus, the clinicians made 69 errors (mania 23, depression 11, schizophrenia 35). The model deviated from the consensus in 84 cases (mania 24, depression 20, schizophrenia 40). The direction of these errors differed systematically:

  • Clinicians rated items that were not assessable according to the expert consensus as present or absent. This accounted for 44 of the 69 errors.
  • The model rated items that were actually absent as not assessable. This accounted for 28 of the 84 errors.

The contrast is sharpest for items the expert consensus had rated as absent. 14 of the clinicians' 17 errors (82 percent) consisted of assuming an item that was not there. For the model, conversely, 28 of 37 errors (76 percent) were conservative abstentions, and only 9 of 37 (24 percent) were false positives. Accordingly, the model scored above the clinician mean on absent items in all three scenarios.

From this the paper formulates a hypothesis about two different interpretation strategies: clinicians draw on context and non-verbal cues, but in doing so infer symptoms from incomplete information. Models stick more closely to what is documented in the text and to the operationalised definitions.

The same picture appears at the level of individual items: 12 items were difficult for the clinicians and not for the model, 23 the other way round, 30 for both.

Where the models clearly fail

The models' restraint is not purely an advantage. It has a concrete cause, and that cause limits the scope of use.

Nine of the 100 AMDP items are observation-dependent, meaning they cannot be inferred from what is said but only read from behaviour. In the German AMDP terminology they are: ratlos (perplexed), affektarm (blunted affect), affektlabil (labile affect), affektinkontinent (affective incontinence), affektstarr (rigid affect), motorische Unruhe (motor restlessness), Parakinesen (parakinesias), manieriert (mannerisms) and theatralisch (histrionic). The model worked from the transcript alone and marked 17 of 27 such item-video combinations (63.0 percent) as not assessable. The clinician consensus did so in 0.0 percent of cases (p < 0.001), individual clinicians in 3.9 percent.

This was most pronounced in the schizophrenia scenario: there the model rated none of the nine observation-dependent items correctly. Removing these items raises its accuracy from 0.60 to 0.66 (schizophrenia), from 0.76 to 0.81 (mania) and from 0.81 to 0.85 (depression). For items such as parakinesias, which are diagnostically significant in catatonic presentations, a purely text-based assessment is, in the authors' judgement, not sufficient in real clinical practice, even if the overall accuracy looks adequate.

The missing picture does not fully explain the restraint, however. Even on the 273 item-video combinations that are not observation-dependent, the model reached for "not assessable" more often than the clinician consensus, in 19.4 versus 11.4 percent of cases (p < 0.001). The model therefore applies a generally stricter standard for when an item counts as assessable at all.

Conversely, the schizophrenia scenario was not only hard for the models. The clinicians, too, reached their lowest accuracy here (0.58) and the highest variance, with full picture and sound. The video also received the lowest authenticity rating, 7.1 versus 8.2 for depression and 7.9 for mania.

The combination experiment

Because the error profiles are partly opposite, the team ran a post-hoc simulation of what happens when both sides work together. From the pool of ratings, 2,091 pairs of clinicians were formed. In 35.5 percent of items the two disagreed. For these cases, three resolution strategies were compared, each for depression, mania and schizophrenia:

  • Random choice between the two disagreeing ratings: 0.42, 0.46 and 0.41
  • Decision by a specialist who had rated the same video, including senior and chief consultants: 0.63, 0.68 and 0.59
  • Decision by the language model: 0.69, 0.71 and 0.52

Both supervision variants were significantly above random choice (p < 0.0002, permutation test). For depression and mania the model was somewhat stronger, for schizophrenia the specialist decision, which fits the transcript problem described above.

The weight of this calculation is limited. It is a simulation on existing data, not a prospective test, and the number of available specialist raters was small: eight for mania, nine for depression, three for schizophrenia. The authors explicitly describe the result as a hypothesis that needs prospective testing.

What the study does not show

The publication is notably clear about the limits of its work. The main limitations:

  • Three interviews, no real patients. Every statement rests on three simulated conversations with a single simulated patient. The reported confidence intervals quantify the uncertainty only for exactly these videos under exactly this rating setting. Generalisability is explicitly not given.
  • One video per diagnosis. Differences between the disorders therefore cannot be attributed to the interview, the raters or the model.
  • Unequal input. The clinicians saw picture and sound, the models read only text. This was a deliberate decision, because current multimodal models do not yet interpret non-verbal clinical signs reliably, but it means the two tasks are not interchangeable.
  • Selection effect. Testing several model families in several configurations on the same dataset and then picking the best combination overestimates accuracy. The authors speak of a "winner's curse" and read their results as an exploratory upper bound.
  • German only. This too limits transferability.

As next steps the paper names validation on interviews with real patients, extension to other languages, and prospective studies that embed model support in real workflows.

One side result is worth noting for practical use: putting the AMDP definitions into the prompt did not help consistently. Across all ten models, mean accuracy fell slightly from 0.63 to 0.61 as a result. The two strongest models, by contrast, gained marginally, from 0.70 to 0.72. More context is not automatically better.

How the authors themselves frame it

The ZI's press release is cautious in tone. Dr Esra Lenz, first author of the study and a member of the Hector Institute for Artificial Intelligence in Psychiatry (HITKIP) at the ZI, is quoted there as follows:

"Our results provide initial evidence that AI could support this process in future. But it replaces neither clinical experience nor the direct assessment by a professional."

Prof Dr Emanuel Schwarz, last author and head of HITKIP at the ZI, names the actual question:

"The decisive question is not whether AI replaces psychiatric professionals, but whether it can provide meaningful support where psychiatric diagnostics actually begins, namely in the clinical interview and the psychopathological assessment."

What follows for practice

For outpatient psychotherapy, this study changes nothing in the short term. It is a proof of concept on three simulated interviews, and its authors say so. Two points nevertheless carry over.

First, the direction in which a language model is usefully deployed. Its strength was not in detecting symptoms nobody else sees, but in cleanly marking what cannot be inferred from the material at hand. That is exactly where the clinicians were weak: they filled gaps. A system that makes visible in the assessment what was not addressed in the conversation targets a real source of error in documentation.

Second, the limit. Everything that is accessible only through observation, that is, affect modulation, psychomotor activity, facial expression, escapes text-based evaluation. What you see has to come from you. How to keep observed and reported findings cleanly apart in the assessment is something we described in our article on the psychopathological assessment according to AMDP.

What this means for psynex

psynex produces psychopathological assessments from the material of your sessions, among other places in the AMDP assessment generator and in the reports to assessors. The study describes the division of labour that is intended there: the system structures what was said and flags where no information is available. The judgement, especially everything observed, stays with you. Every generated assessment can be edited before it is adopted.

Sources

Ready to transform your documentation?

Try Psynex free for 14 days. Experience how AI-powered analysis changes your daily practice. No credit card required.

Start free trial

14 days free • No credit card • GDPR compliant