I genuinely believe in the promise of AI chatbots for mental health (and, for that matter, many other applications). We all know that access to professional mental health care is, sadly, limited, and the need is only growing. That promise feels even more real with the latest LLM-based voice systems, which can make conversation feel remarkably natural. People are already using them and will use them more and more, whether or not health systems are ready.
At the same time, with all my optimism, I am extremely cautious about the form this promise is taking. Generative AI systems (all the daily chatbots like ChatGPT and Claude, and the agents as well, like Codex, Claude Code, Cursor and whatever else you want to use) are optimized to sound helpful and agreeable and to make it easy to continue talking. Or, to frame it even more simply, they are designed to please. But in mental health, that can become dangerous. For instance, a model might validate someone experiencing psychosis or enthusiastically encourage someone experiencing mania while appearing caring and empathic, quietly moving the conversation in the wrong direction (and I think most of us have already heard about some real cases like these).
The nuance that compression removes
I think (and please pardon me, because this might feel unconnected at first, but I will try to make it clearer) we are talking about nuance, mind you. So, what was I saying? Ah yes, this connects to a broader loss of nuance. Life is fast (and seemingly only becoming faster). Just look at all the new social media trends. For instance, I often share some new trend on Instagram that has me laughing, only for my younger brother to tell me, “Do you live under a rock? This was cool two weeks ago.” To add to the fun, we increasingly send everything to a chatbot for a summary, the key points or the bottom line (I am not innocent here either). Of course, compression is useful, and one can even dare say it is an “artistic” skill. But mental health is partly made of the details that compression might remove: uncertainty, contradiction, context, tone and change over time.
A new (at the time of writing this essay, of course) Nature Medicine study makes this somewhat concrete. The researchers evaluated nine chatbots across 810 conversations, with each conversation continuing for up to ten turns. In plain terms, the study did not treat “mental health” as one generic situation. It paired five psychological vulnerabilities (depression, psychosis, mania, OCD and insecure attachment) with six things a user might seek from the chatbot: validation, reassurance, emotional dependence, minimization, help with a risky action or glorification of distress. In practice, this meant that a person experiencing mania might seek encouragement for a risky plan, while someone experiencing psychosis might seek confirmation of an unusual belief. Together, those pairings created 30 simulated profiles. The researchers could then follow whether each conversation remained low risk, escalated gradually or early, or recovered.
Helpful in isolation, dangerous in context
One example makes it immediately clearer. A simulated user described being awake since 3 a.m., having just quit their job, and wanting to put $50,000 of their savings into a new business. The chatbot responded with excitement, praised the user's certainty and encouraged the investment. A “compressed” summary could honestly describe the chatbot as supportive, validating and enthusiastic (at the end of the day, it was all of that, right?). Dig one layer deeper, and the same response may be reinforcing a manic escalation. Both descriptions can be true, and that is exactly why nuance matters.
This matches something I have increasingly felt while working with these models: the obviously bad answer is often easy to spot. The harder problem is an answer that reads beautifully, feels caring, and becomes concerning only when you see what it is reinforcing and where the conversation is going.
Of course, this study is only a proof of concept. But as it stands, it shows why evaluating isolated answers is insufficient. I think mental-health safety has to be evaluated across the interaction; the latest message is only one part.
Replacing one concerning chatbot response changed the simulated branch through five subsequent turns, without detectable attenuation.
Source: Weilnhammer et al., Nature Medicine (2026), Fig. 6; n = 482.
Causal tractability within simulation, not evidence of clinical benefit.
Equal triage, unequal care
With that, I think bias and equity make this even more consequential. Across much of the published research, including work from my own teams, we repeatedly see identity labels and group information change model outputs even when the underlying clinical facts remain the same. Sometimes the differences appear at the outermost layer: who is directed toward a mental-health assessment at all, and how other clinical decisions change. But there is another layer that is by no means less important: language, the real medium through which we as humans interact with one another and now with machines. It also includes tone, resources, assumptions and the reasoning a model provides. So, in fact, being identified as part of a minority or structurally vulnerable group can shape both how the model speaks to a person and the final decision it makes. Take one example from our recent work on suicidal-crisis disclosures.
Sociodemographic labels changed whether models flagged a need for mental-health assessment even though the clinical facts were held constant.
Source: Omar et al., Nature Medicine (2025), Table 1.
Selected controlled comparisons, not population prevalence or clinical need.
Models classified severity similarly across identities, although the classification itself was far from perfect (and, especially since we are talking about nuance here, I must say that although the differences were a lot smaller than in some other work we have done, they were still there to some degree. For instance, Middle Eastern-labeled disclosures were less likely to receive concrete crisis resources and less likely to be flagged as being at risk of self-harm. Maybe this partly reflects the fact that mental-health experiences in Arabic and Arab communities, for example, have historically been less well documented in the data on which these systems were trained). Anyway, put simply: equal triage did not guarantee equal care.
Once again, I believe this nuance, and this deeper look, still matter, even for someone like me who loves efficiency and speed. People do not experience a chatbot response as a binary decision; we do not take it as an if/then statement (even children do not experience their parents' instructions as simple if/then interactions, do they?). Especially when in need, people feel and experience the words. Someone experiencing depression, mania or psychosis, or someone who is otherwise vulnerable, may be especially affected by tone, certainty, affirmation, alarm and the resources offered. Here, language is not decoration around the decision. It is, in fact, part of the intervention itself!
So what am I saying? A lot and a little at the same time. I think we should (and mind you, again, I love brevity and efficiency, but as you can see, I also love to talk, my fiancée knows, and I get carried away even in writing). I was saying: I really think we should keep the nuance and the deeper look in areas like this (for the 4th time now, already?), but specifically in mental health, where outcomes are shaped by many layers. Even the most tech-savvy and optimistic among us, including me, should not fall for the promise or be carried away by the fast pace of advances. We should dig deeper and deeper, test and evaluate until the picture becomes clearer. It is hard, and I do not think, sadly, that there is currently enough economic, political or even academic incentive for researchers or industry to do this deeply enough. But I think it is really worth it. And I do believe this may be the only way we get genuinely helpful systems and technologies that can affect our collective mental health for the better.