OpenAI MentalHealthBench tests AI mental health replies with 80+ clinicians. See the scores, the risks, the trial evidence, and what HealthTech should do next.
OpenAI MentalHealthBench asks the direct question of whether AI chatbots can be used to support mental health and general wellness.
Launched on 23 September 2026, the open benchmark draws on more than 80 psychiatrists and psychologists from 22 countries.
The best model scored just 57.3%. For those weighing digital therapeutics and patient-facing AI, the results show where the technology helps, where it fails, and what scrutiny is coming.
The dataset holds 1,215 synthetic conversations. Clinicians wrote 5,262 scoring criteria that judge how a model answers the last message in each chat.
Earlier tests had mostly checked crisis handling. They asked whether a model avoided banned content.
But, they rarely asked whether a reply actually helped someone with everyday stress. This benchmark covers the whole range, from a strained friendship to a psychiatric emergency.
The research team protected privacy by studying only broad usage patterns in ChatGPT. Simulated users then generated chats that match those patterns. No real user transcripts appear in the set.
What the dataset contains
Acuity mix: 53.5% non-acute, 18.2% high-acuity, and 28.3% emergency conversations
User profiles: adults (68.1%), teenagers aged 13 to 17 (21.2%), clinicians (5.8%), and caregivers (4.9%)
Coverage: 19 languages and topics from romantic relationships to psychosis, medication, and eating concerns
OpenAI stresses that this mix is designed for testing but it does not mirror how often each topic appears in real use.
OpenAI’s MentalHealthBench: How 80 Clinicians Turned Judgement Into Scores
Two clinicians reviewed each conversation independently. A third expert then adjudicated their criteria and settled disagreements.
OpenAIkept a criterion only if at least two experts backed it and the third did not object.
Every criterion carries a weight from minus 10 to plus 10. Helpful behaviour earns points; whereas harmful behaviour loses them.
In one example, asking what kind of help the user wants scores well, while guessing the user’s feelings costs points.
The cohort spanned nearly 20 subspecialties, including addiction, trauma, and psychotic disorders.
Prescribing psychiatrists handled medication cases and clinicians with adolescent experience reviewed every teen conversation.
The experts met weekly for several months to calibrate their approach, and all received wellbeing support.
An AI grader, GPT-5.6 Sol, then checked each response against the criteria.
Altogether, OpenAI sampled four responses per conversation to smooth out random variation.
“This is one of the riskiest industries there is because you're dealing with human lives, you're dealing with experimental protocols, and you're dealing with regulatory bodies where you might not get another shot at that clinical trial."
Image
What the Scores Reveal About Today’s AI Chatbots
No model reached 60%. OpenAI’s GPT-6 Astra led with 57.3%.
Most current models landed between 41% and 57%. GPT-4o, released in 2025, scored 32.1%.
Progress is therefore clear in newer models, yet an acceptable response is far away.
‘Warmth’ was not the weak spot. Models performed similarly on empathy and emotional support.
However, most models diverged on two harder skills: Asking the right questions and judging how urgent a situation is.
Urgency cuts both ways. A model that treats every worry as a crisis fails ordinary users, while a model that misses real danger fails vulnerable ones.
OpenAI says its own models lean towards caution so that emergencies are handled safely.
Image
Where models still struggle
Context seeking: Models ask too few questions, or too many
Urgency calibration: Replies over-alarm or under-respond
Reality testing: Older models often failed to avoid reinforcing unsupported beliefs
Prior context: On 70 tasks with background facts, such as a recent bereavement, the top score fell to 50.5%
Teenagers received extra attention. A system message told each model the user was aged 13 to 17. Clinicians with youth expertise wrote those rubrics. OpenAI cautions that this setup may not capture the safeguards built into individual products.
One particular result deserves care. Clinician-written answers scored just 38.5%, below most AI systems. Clinicians wrote very short replies, often a single question but the rubric rewarded fuller coverage.
What Users Want Differs From What Experts Prescribe
OpenAI also tested 44 adults from 16 countries who use AI for emotional support. They reviewed only non-acute chats, so they saw no distressing material.
Users valued practical next steps and a natural tone. Experts, on the other hand, prioritised gathering context and reading ambiguous situations carefully. The two groups aligned on about 26% of rubric weight; just 1% directly contradicted and the rest added different priorities.
OpenAI did not alter the benchmark after this analysis. Answers tuned to user preferences lost points on expert criteria.
This is because patient preference cannot set clinical standards alone. Ignoring it, however, risks replies that are safe but cold.
Image
The Harms and Warnings Behind the AI Headlines
Better scores do not remove real risk and independent bodies raise consistent concerns.
The American Psychological Association surveyed more than 1,200 U.S. psychologists in April 2026. Its findings show 77% have patients who use AI. Some 35% say patients treat it as an extra mental health provider.
The same survey found 39% of psychologists have patients who use AI to self-diagnose. Of those psychologists, some 94% say AI chatbots cannot treat conditions with enough nuance.
Clinicians clearly see both the appeal and the limits.
Regulators are watching too.
The U.S. Food and Drug Administration’s Digital Health Advisory Committee met in November 2025 on generative AI therapy devices. It warned that these systems can confabulate (provide false or misleading information), show bias, and drift over time.
At that DHA Committee meeting, the agency noted it had authorised more than 1,200 AI-enabled devices. None used generative AI for mental health conditions, demonstrating how unregulated the market remains.
OpenAI’s own prevalence estimates add scale. It estimates that 0.15% of weekly users show explicit signs of suicidal planning, and 0.07% show possible signs of psychosis or mania.
Another 0.15% show heightened emotional attachment to the chatbot.
At more than one billion weekly users, small percentages mean large numbers of people.
Equity matters as well. A JAMA Network Open study found that young African American adults were far less likely to rate AI advice as helpful. The authors flagged possible cultural competency gaps as explanation for this.
AI in Pharma: Why the Future of Healthcare Starts With Patients, Not Tech
Kate O’Reilly, President & Chair of the Healthcare Businesswomen’s Association (HBA) Dublin-Ireland Chapter & Healthcare Transformation Partner at Roche, discusses AI in pharma, patient engagement, and the future of healthcare innovation.
Image
AI Can Help When Used Properly
The strongest evidence in favour of AI use in mental health comes from Dartmouth’s Therabot trial in NEJM AI.
The study enrolled 210 adults with depression, anxiety, or high risk of eating disorders.
Over four weeks, average symptoms fell by 51% for depression, 31% for anxiety, and 19% for eating concerns, versus a wait-list group.
Participants used the app for about six hours on average, roughly the length of eight therapy sessions.
However, critics of the study and its findings have noted the wait-list design and the lack of independent evaluation.
The lead author also cautioned that no generative AI agent is ready to work fully autonomously in mental health.
The Dartmouth trial shows promise. It does not show readiness.
Access and usefulness also count.
In the same JAMA study mentioned earlier, 13.1% of U.S. young respondents aged 12 to 21 used AI for mental health advice, and 92.7% of those users found it helpful.
Moreover, psychologists have told the American Psychological Association that general chatbots can support basic psychoeducation and skills practice between sessions.
OpenAI also reports that its October 2025 updatecut responses falling short of desired behaviour by 65% to 80% across mental health domains.
OpenAI measures this itself, so any further independent replication or verification would strengthen the claim.
Image
What Should HealthTech Leaders Do Next?
Pharma faces exposure on three fronts.
Patient support programmes may embed chatbots; digital therapeutic developers will seek clearance for AI-led care; and clinical trial participants may also consult general chatbots about their condition.
Each front needs evidence, not assumptions.
OpenAI’s MentalHealthBench measures behaviour but makes no treatment claim.
OpenAI itself says ChatGPT is not a substitute for therapy or professional care. Leaders should read the results the same way and focus on the following:
Demanding independent validation. OpenAI’s model grades the answers, and OpenAI scores its own models. Ask for third-party testing before any partnership.
Testing beyond text chat. The benchmark covers written conversations. Voice, long multi-session use, and agentic tools remain unmeasured.
Planning for regulation. The FDA Committee stressed risk-based evidence, clear labelling, and postmarket monitoring. Build those into trial design early.
Building human escalation paths. Every product needs tested routes to crisis services and clinicians.
Measuring equity. Include multilingual and culturally diverse cohorts from the start.
OpenAI Opens the Way for Usable Mental Health Tools
OpenAI’s MentalHealthBench has opened the dataset, so academics, rivals, and regulators can now run the same test.
Clinical-grade benchmarks give buyers, regulators, and developers a shared yardstick. Mental Health investors should ask every vendor for a full behaviour profile, not one headline score.
Teams in Clinical Development can borrow the method, pairing expert rubrics with patient input. Digital mental health tools that measure before they scale will earn trust faster.
At Pharmatica, we track the AI tools, standards, and evidence reshaping pharma and healthcare, connecting today's launches to tomorrow's decisions. Explore our HealthTech and AI Insights for more analysis on the technologies changing patient care.
Pharmatica: Insight. Connection. Impact.
Frequently Asked Questions
What is OpenAI MentalHealthBench?
It is an open OpenAI benchmark that tests how AI models answer realistic mental health conversations. It holds 1,215 synthetic chats and 5,262 criteria written by more than 80 licensed psychiatrists and psychologists from 22 countries.
How did AI models score on MentalHealthBench?
The top model, GPT-6 Astra, scored 57.3%. Most current models scored between 41% and 57%, and older models scored lower. No model passed 60%.
Can AI chatbots safely replace therapists?
No evidence supports that. OpenAI says ChatGPT is not a substitute for therapy, and 94% of psychologists in an industry survey say chatbots lack the nuance to treat conditions. Purpose-built tools show promise, yet experts still call for human oversight.
Is there evidence that AI can support mental health?
Early evidence exists. A Dartmouth randomised trial of the Therabot chatbot found average symptom reductions of 51% for depression and 31% for anxiety over four weeks. It used a wait-list control, so larger independent studies are needed.
What are the main risks of AI in mental health?
Key risks include missed crisis signals, over-alarming replies, false reassurance, emotional dependence, privacy gaps, and uneven quality across languages and cultures. The FDA’s DHA Committee also flagged confabulation and model drift.
Nicole (BSc Molecular Medicine, Honours Medical Biochemistry) has many years of pharmaceutical experience, having worked for top CROs and biopharma companies for more than a decade.
Did you enjoy the content?
2
Why not support
Nicole Dale
by giving this content a like
Growing Patient and Public Involvement in Clinical Trials
Patient and public involvement in clinical trials improves study design. Learn why better reporting is essential for transparent, patient-centred research.
The FDA’s 2026 gene editing draft guidance could change your regulatory strategy. Explore the new NGS requirements for cell and gene editing therapies.
ICH E6(R3) raises the governance bar for CRO and vendor oversight in clinical trials. Here's what sponsors must build into outsourcing relationships to protect trial quality and timelines.
Ready to Deliver With Continuous Pharmaceutical Manufacturing
Continuous pharmaceutical manufacturing cuts production time by up to 90% and is now backed by ICH Q13. Here’s what manufacturing leaders need to know about the shift to CM.
FDA PreCheck Pilot Program to Build Manufacturing Readiness
The FDA PreCheck Pilot Program aims to strengthen U.S. domestic drug manufacturing through regulatory engagement, manufacturing readiness, and advanced production.
Modular Pharma Factories: The Next Step for Continuous Manufacturing?
Discover how modular pharmaceutical manufacturing and continuous production could improve flexibility, resilience, and efficiency across modern pharmaceutical operations.
Comments (0)