Nelson, B. W., Kalinich, M., & Torous, J. (2026).
JAMA.
That artificial intelligence (AI) can cause harm in mental health contexts is no longer contested.1 The debate has shifted to what kinds of harm, how to measure them, and who is accountable. Early AI models and their associated guardrails were largely inadequate at addressing mental health–related harm, but today that is beginning to change. Advances in both model capability and the surrounding safety infrastructure now permit a more proactive approach.1
These advances require new terminology. There is currently no well-accepted taxonomy of mental health AI harm, so the focus has been on suicide and self-harm, where agreement is universal. Unlike human therapists who may miss harm,2 AI cannot only miss harm but can also actively enable it by providing information that encourages and facilitates harm, with some mistakes subtle and others fatal. Agreeing on a taxonomy of AI harm and understanding how to assess and respond to those harms are critical to advance AI safety.
To advance this discussion, we introduce 2 concepts. Type I harms are acute harms arising within a single or brief exchange with AI. They are failures in which a chatbot produces a clearly wrong or dangerous response to a discrete prompt (eg, missing a suicide cue, enabling disordered eating behaviors, giving inappropriate medical advice, or reinforcing delusions). These harms are tractable; new research shows that an external fit-for-purpose guardrail layered on top of an existing model may reduce acute harms to nearly zero, although the costs of running such are not well characterized today.1 The field is learning how to define, measure, and mitigate type I harms.
The article is linked above.
Here are some thoughts:
The authors propose a taxonomy for AI harm in mental health. Type I harms are acute failures in a single exchange (missing a suicide cue, reinforcing a delusion). Type II harms build insidiously across long or repeated conversations through attachment, sycophancy, and delusion-reinforcement as guardrails degrade. The two are orthogonal, so single-turn benchmarks, where most safety claims are made, say little about cumulative safety. They argue that resistance to Type II harm comes from value-anchored alignment in the base model, not from fine-tuning for empathy on top. On this basis they are skeptical of "mental health–specific" models, comparing them to the failed "digital therapeutics" label, and call for transparency, independent evaluation, and standards benchmarked to real clinical care over realistic timescales.
The orthogonality claim is the core contribution and, if true, indicts the field's single-turn benchmark culture: good scores may be misdirection rather than reassurance.
The architectural argument (that safety lives in the base model, not the empathy fine-tune) is strong and falsifiable but asserted rather than demonstrated, and it happens to favor well-resourced frontier labs over specialized entrants. Worth reading with that alignment of interests in mind.
The best move is refusing the general-versus-specialized dichotomy in favor of an evidentiary question: what has this system been shown to do, for whom, over what time horizon.
Two weaknesses: the Type II evidence base cited is thin (two studies), ironic given their own call for replicability; and they raise cost without confronting that robust real-time guardrails may be too expensive to deploy at scale. The digital-therapeutics analogy also warns against marketing over evidence, but those products failed because the interventions were weak, whereas the worry here is the opposite, that the technology is capable enough to sustain the relationships where Type II harm incubates.








