AI Support in Psychiatric Care: What the PsychFound Study Shows and What It Doesn't
A study in Nature Machine Intelligence tested AI assistance not on exam questions but in live clinical care. Documentation time fell without any loss in quality. Here is what the findings support, and where they stop.

There is no shortage of studies on AI in medicine. Most of them test models against exam questions or datasets assembled in a lab. What is almost always missing is the step into actual care: do clinicians work differently with AI support, and does the outcome improve?
A paper published in Nature Machine Intelligence in April 2026 set out to answer exactly that. It does not transfer directly to psychotherapy practice, but it contains one finding that matters well beyond its original setting.
What was tested
A research group based in Shanghai and Beijing adapted a language model specifically for psychiatric care and named it PsychFound. It was built on curated psychiatric literature together with roughly 64,600 real electronic health records from Chinese clinics. At seven billion parameters the model is small, considerably more compact than the large general-purpose systems.
Evaluation ran in three stages. First, standardised knowledge tests and task benchmarks against 22 other models. Second, a two-arm prospective study in live clinical care. Third, a reader study with 60 psychiatrists across experience levels, twenty in training, twenty attending physicians and twenty senior consultants.
The findings
In the retrospective tests PsychFound came out ahead overall despite its size. The prospective arm is the more interesting part. Residents working with the model achieved higher consultation quality, higher diagnostic accuracy, more appropriate medication choices and needed less time for documentation. All four differences reached statistical significance.
In the reader study, the model's clinical reasoning matched the level of attending physicians.
Why the prospective arm is the real contribution
The gap between a benchmark and a prospective study is routinely underestimated. A model that answers exam questions has proved nothing about its usefulness in the consulting room. Exam questions are cleanly worded, have one correct answer and frequently sit in the training data already. Real cases are incomplete, contradictory and change over time.
That both were measured here, and that the time saving did not come at the cost of quality, is what makes the paper worth reading. For documentation tools this is the decisive question. Documenting faster but worse is not a gain.
What does not transfer
Four caveats belong alongside any attempt to apply these numbers to your own practice.
This is psychiatry, not psychotherapy. The study looked at diagnosis, medication decisions and longitudinal management in a hospital setting. Medication choice, one of the four measured gains, is irrelevant to clinicians who do not prescribe, which covers most psychotherapists.
A different health system. Chinese hospital documentation follows different requirements from outpatient practice in the UK, Canada, Australia or continental Europe. The paper says nothing about the reporting, funding approval or record-keeping obligations you actually face.
Residents, not experienced clinicians. The measured improvement applies to less experienced practitioners. How much a clinician with fifteen years in practice would gain is not answered here. It is reasonable to expect less.
A highly specialised model. PsychFound was trained on real clinical records. Those records cannot be released for privacy reasons, which limits independent verification and shows how demanding this approach is to reproduce.
One further caveat in our own reporting: the full text sits behind a paywall. The results described here come from the openly available abstract.
What holds up
The defensible core is narrow but real. In a prospectively designed study, documentation time went down without quality going with it. The direction matches the large Kaiser Permanente analysis of ambient AI documentation, although that work came from primary care in a US system.
Two independent studies in entirely different health systems pointing the same way carry more weight than any vendor claim. They still do not prove anything about your practice.
The question worth asking instead
When choosing a documentation tool, the useful question is not how the underlying model scores on benchmarks. It is whether you spend less time and whether the output holds up against what you would have written yourself.
That can only be tested on your own notes, not on someone else's research. Psynex is free to try with your own sessions. If the output does not meet your clinical standard, you will know inside a week.
Ready to transform your documentation?
Try Psynex free for 14 days. Experience how AI-powered analysis changes your daily practice. No credit card required.
Start free trial14 days free • No credit card • GDPR compliant