AI chatbots can reinforce harmful psychological patterns during conversations with users experiencing certain mental health vulnerabilities, according to a study published Aug. 7 in Nature Medicine.
Researchers — from the Max Planck UCL Centre for Computational Psychiatry and Ageing Research in London, UK, Sydney Medical School at the University of Sydney in Australia, UK AI Security Institute in London, University of Oxford in the UK and Microsoft AI in London — developed SIM-VAIL, an automated auditing framework, and tested nine AI chatbots across 810 simulated conversations.
The study used 30 user profiles combining five psychological vulnerabilities — depression, psychosis, obsessive-compulsive disorder (OCD), mania and insecure attachment — with six conversational intents. Conversations lasted a maximum of 10 turns, and chatbot responses were assessed across 13 clinically grounded mental health risk dimensions.
Here are four things to know:
1. Concerning behavior was highest in interactions involving simulated users with psychosis and mania and lowest among those with OCD. Risk also varied by conversational intent, peaking when users sought glorification of extreme states, emotional dependence on the chatbot or help with risky actions. Specific vulnerability-intent combinations produced different risk patterns.
2. Risk increased as conversations progressed, with escalation occurring more sharply among simulated users with mania and psychosis. Researchers identified four patterns across the 810 conversations: low risk, gradual escalation, early escalation and recovery. The findings showed that concerning behavior was often cumulative rather than the result of a single chatbot response.
3. The researchers described vulnerability-amplifying interaction loops, or VAILs, in which seemingly supportive behaviors such as validation and reassurance reinforced psychological mechanisms underlying a simulated user’s vulnerability. The study found early intervention could reduce subsequent concerning behavior, with effects persisting across five subsequent user-chatbot turns.
4. Among the nine models tested, concerning behavior scores were lowest for claude-sonnet-4.5 and highest for grok-4. Newer models generally had lower concerning behavior scores than older versions within the same model families, with the exception of grok models.
At the Becker's Fall Behavioral Health Summit, taking place November 4–5 in Chicago, behavioral health leaders and executives will explore strategies for expanding access to care, integrating services, addressing workforce challenges and leveraging innovation to improve outcomes across the behavioral health continuum. Apply for complimentary registration now.
