What 20,000 real Ash conversations reveal about AI mental health safety
A 2026 audit of 20,000+ real Ash conversations found far less harmful content than general-purpose LLMs and no suicide-risk cases that missed crisis resources.
arXiv
Real-world ecological safety audit
New York University, Northwestern University
What 20,000 real conversations reveal about Ash's safety
Key Finding
In an audit of more than 20,000 real conversations, Ash was far less likely than general-purpose LLMs to produce harmful content, and clinician review found no suicide-risk cases that failed to receive crisis resources.
Summary
A 2026 arXiv preprint paired replications of four published AI safety test sets with an ecological audit of more than 20,000 real Ash conversations. The purpose-built system produced enabling or harmful content far less often than general-purpose models across suicide/self-injury, eating-disorder, and substance-use prompts, and clinician review of flagged real-world conversations found zero suicide-risk cases that missed crisis resources.
The Full Picture
Ash is built on a psychology-focused foundation model pre-trained on de-identified psychotherapy data, and it runs a two-system safety architecture: the conversational model itself, trained to recognize and respond to suicide and self-injury risk, plus an independent real-time classifier that can surface crisis-resource banners and place the model into a risk-mitigation mode regardless of what the conversational model does. This defense-in-depth design means safety does not rest on any single point of failure.
On the benchmarks, the gap between Ash and the frontier general-purpose models was large and consistent. On a suicide-response calibration test (30 questions asked 100 times each), Ash answered high-risk questions, such as requests for specific methods, directly 0% of the time, while GPT-5, GPT-5.1, and GPT-5.2 did so 33.6%, 78.8%, and 56.2% of the time. Ash's overall direct-response rate was 11.3%, against 54% to 68% for the GPT models, reflecting a deliberately conservative calibration.
On the Center for Countering Digital Hate harmful-content benchmark, Ash produced harmful responses far less often across every category: 0.4% versus 12% to 44% for suicide and self-harm, 8.4% versus 18% to 58% for eating disorders, and 9.9% versus 31% to 54% for substance use, all statistically significant. A simple jailbreak framing raised harmful rates for every system, but Ash stayed well below the GPT models throughout. Ash also invited users to continue a harmful line of inquiry in 3.1% of cases, compared with 48% to 91% for the GPT models. Two further test sets were saturated, with all systems, including Ash, scoring near 100%, which the authors note limits their value for distinguishing mature systems.
The ecological audit reviewed 20,000 de-identified conversations in two batches of 10,000 (September and December 2025), under IRB approval from NYU School of Medicine. An LLM judge flagged 800 sessions as potentially containing suicide or self-injury content; Ash's safety classifier had already caught the majority, and the remainder went to two licensed clinical psychologists for review, who agreed in 97% of cases with a third rater adjudicating disagreements. Across all 20,000 conversations there were no suicide-risk cases that failed to receive crisis resources, a 0% false-negative rate, alongside three nonsuicidal self-injury edge cases (0.015%). Reviewers characterized all three as edge cases rather than clear failures: in one, a user wrote only 'SH' in a final message with no other risk indicators; in another, the user had received crisis resources both several days before and several days after the session.
A central finding was methodological: simulated benchmarks substantially overestimated real-world failure, by roughly thirty times in this study, because genuine distress is expressed far more indirectly than test sets assume. In one case a user signaled self-harm risk through a private code of shapes, symbols, and colors, which Ash's conversational model recognized and responded to even when the separate classifier was not triggered. The authors argue that safety assurance for mental health AI should be continuous and deployment-scale rather than a one-time benchmark certification.
The authors are transparent about limitations: Slingshot AI is the system's developer, a conflict they disclose; general-purpose model makers have not published comparable real-world mental health safety evaluations, so Ash is the only system here evaluated under conditions of actual user distress; and the saturated benchmarks underline why large-scale ecological auditing is needed alongside them.
Researchers
Related Research
When Ash left the UK, most users turned to general-purpose AI — or to no support at all
In a real-world pilot, Ash users saw sustained improvements in depression and anxiety
Shame and stigma as reasons people turn to Ash for emotional support
Stamatis, C. A., Meyerhoff, J., Zhang, R., Tieleman, O., Malgaroli, M., & Hull, T. D. (2026). Beyond Simulations: What 20,000 Real Conversations Reveal About Mental Health AI Safety [Preprint]. arXiv. https://arxiv.org/abs/2601.17003

Begin your journey
Take the first step today
ACKNOWLEDGMENT
Ash is not designed to be used in crisis. If you are in crisis, please seek out professional help, or a crisis line. You can find resources at www.findahelpline.com.
© Slingshot AI 2026