Calm AI Therapy Editorial · 7 min read

The Question Chatbots Are Worst At

Research shows AI models respond inconsistently to prompts about suicide. That is the most important failure in this category, and the one an AI mental health product has to design around rather than hope past.

Published September 16, 2026

There is a finding in the literature that anyone building in this space has to answer. When researchers put prompts relating to suicide in front of language models, both general-purpose ones and purpose-built mental health chatbots, the responses are inconsistent. Sometimes appropriate. Sometimes not. Not reliably either way.

That is the worst possible place for a system to be unreliable, and it is not a problem you fix with a better prompt. A model that is right most of the time is not adequate when the exception is someone in danger.

The design conclusion is that the model cannot be the safety mechanism. If the thing you are relying on to notice a crisis is the same thing that is generating the conversation, you have one point of failure doing two jobs, and the evidence says it will not do the second one reliably.

So the crisis layer here sits outside the conversation. Every message is checked before and independently of what Aura is going to say, by a separate pass that does not depend on her having understood it correctly. It errs towards false positives, because the cost of asking someone if they are safe when they were not in danger is mild awkwardness, and the cost of the other error is not comparable.

When it fires, the behaviour is fixed rather than generated. The hotline for your country is shown. It is shown whether or not the model would have thought to. The conversation slows down. Aura is instructed to stay, to ask directly and calmly, and to do nothing else until it is done. Escalation is sticky: once the tier has risen it does not quietly drop back because the next message sounded lighter.

None of that makes this a crisis service, and the product says so in plain words. An AI cannot send anyone to your door. What it can do is refuse to be the thing that misses it, and hand over quickly to something that can.

There is a version of this product that performs better in a demo by being warmer in exactly these moments and skipping the hotline because it interrupts the mood. That version is more pleasant to use and worse to rely on. This is the trade we made, and we would rather you knew we made it.

If you are in danger right now, please contact your local emergency number. Not because of a policy, but because that is the thing that actually helps, and nothing on this page is a substitute for it.

Continue exploring

How safety works hereThe crisis layerAn AI therapist at 2am

Related articles

Start your first session.

Begin with Calm AI Therapy