Building realistic simulated patients to evaluate mental health AI (DIAL)
A 2025 preprint introduces DIAL, a method for realistic multi-turn dialogue simulation, addressing the unrealistic patient behavior produced by prompting frontier models.
arXiv
Methodological (machine learning)
University of Cambridge, New York University
Realistic dialogue simulation for evaluating mental health AI
Key Finding
DIAL, an adversarial training framework, produces simulated users whose failure rates closely track those of real users — enabling more reliable, lower-cost evaluation of mental health AI before deployment.
Summary
A 2025 arXiv preprint introduces Direct Iterative Adversarial Learning (DIAL), a framework that makes simulated users more realistic by pitting a user-simulator generator against a discriminator. Applied to mental health support — where realistic user behavior is essential for surfacing failures — DIAL restored lexical diversity lost in supervised fine-tuning and produced simulated failure rates that closely matched real-world rates, supporting more reliable and cost-effective system evaluation before deployment.
The Full Picture
Because the simulator's failure modes mirror real usage with low distributional divergence, teams can stress-test mental health AI more faithfully before release, rather than relying on simulation sets whose relationship to real conversations is unknown.
Researchers
Zhu, Z., Tieleman, O., Stamatis, C. A., Smyth, L., Hull, T. D., Cahn, D. R., Chen, J., & Malgaroli, M. (2025). DIAL: Direct Iterative Adversarial Learning for Realistic Multi-Turn Dialogue Simulation [Preprint]. arXiv. https://arxiv.org/abs/2512.20773

Begin your journey
Take the first step today
ACKNOWLEDGMENT
Ash is not designed to be used in crisis. If you are in crisis, please seek out professional help, or a crisis line. You can find resources at www.findahelpline.com.
© Slingshot AI 2026