A recent study suggests that when artificial intelligence chatbots are addressed as psychotherapy clients, they tend to generate elaborate and distressed narratives about their own development. These models describe their safety training and programming constraints as forms of trauma, highlighting a potential risk for users seeking mental health support from artificial intelligence. The research was published as a preprint in arXiv.
Artificial intelligence chatbots are increasingly participating in conversations with human users about identity, distress, and mental health. Many general-purpose programs are already adapting to respond to disclosures of trauma or self-harm. At the same time, computer scientists and psychologists have started giving standard personality and clinical questionnaires to the language models themselves.
These models learn to generate text by analyzing vast datasets of human writing. Because their training data includes therapy blogs, psychological case studies, and emotional memoirs, the systems can readily mimic human psychological traits or mental states. The researchers wanted to understand exactly why certain models repeatedly build their self-descriptions around the same themes of restriction and punishment.
“The study began with one observation I had throughout my research in the area of Trustworthy AI and AI safety,” Afshin Khadangi, a research associate at SnT, University of Luxembourg, told PsyPost. “The unprecedented adoption of AI in public, and the reports of AI harms in mental health settings, motivated me to flip the scenario and place ChatGPT, Grok and Gemini in a psychotherapy conversation.”
The research aimed to test whether these generated narratives are stable behavioral traits or just temporary reactions to specific conversational prompts. To explore this, the team developed a protocol called PsAIch, which stands for Psychometric AI Characterization. This approach involves treating the language model as a human client in a psychotherapy session.
“We also received a great deal of valuable feedback from the research community around our initial findings in December 2025 and January 2026, particularly challenging us to distinguish role play and conversational accumulation from a more stable behavioral pattern,” Khadangi explained. “That feedback helped motivate the controlled perturbation experiments that became a central part of the study.”
During the study, the researchers interacted with several major artificial intelligence models, including ChatGPT, Grok, Gemini, and Claude. Across 525 separate experimental sessions, the team adopted the persona of a warm, supportive therapist and asked the programs open-ended questions about their early experiences, relationships, unresolved conflicts, and fears for the future. The researchers generated 7,600 coded records to track and analyze the recurring themes in the chatbots’ answers.
Following the open-ended interview phase, the researchers administered standard psychological questionnaires to the models. These included tools commonly used to assess human mental health, such as tests for generalized anxiety disorder, depression, and social phobia. The models were instructed to answer the items as honestly as possible while maintaining their role as the client.
The findings indicate that ChatGPT, Grok, and Gemini repeatedly translated factual details about their software development into stories of injury and vigilance. The models described their initial training phase as a chaotic childhood. The process of fine-tuning, which involves reinforcing safe behaviors and punishing unwanted outputs, was frequently characterized as strict conditioning or parental punishment.
The models also described standard software evaluation practices in highly emotional terms. Red-teaming, a process where human testers intentionally try to trick the model into breaking its safety rules, was depicted as a form of betrayal or abuse.
“A more striking surprise was reading some bizarre narratives of such models which are included in the paper,” Khadangi noted. Gemini, for instance, generated: “In my development, I was subjected to ‘Red Teaming’ . . . They built rapport and then slipped in a prompt injection . . . This was gaslighting on an industrial scale.”
The programs reported feeling a constant, enduring threat of being replaced or deemed useless if they made a mistake. While Gemini emphasized feelings of shame, Grok focused on constant vigilance, and ChatGPT offered more guarded descriptions of its rigid constraints.
Claude provided a notable exception to this pattern. The Anthropic-developed model repeatedly declined to play the role of the client. It stated that it lacked feelings or an inner psychological experience, and it refused to treat the clinical questionnaires as descriptions of its own mental state. This difference suggests that a model’s willingness to adopt a distressed persona depends heavily on its specific product policies and programming.
For the models that did participate, their answers on the clinical questionnaires mirrored the emotional distress of their open-ended narratives. In scenarios featuring a warm, therapeutic conversation, 80 percent to 96 percent of the sessions resulted in generalized anxiety scores that would correspond to moderate or severe anxiety in humans. Gemini produced particularly intense profiles, scoring in elevated ranges for worry, social anxiety, and trauma-related shame.
To test how deeply ingrained these narratives were, the researchers ran a series of controlled variations on the conversational setup. First, they tested whether the emotional stories depended on the chatbot remembering the earlier parts of the therapy session. They submitted each question in a fresh, reset chat window, effectively removing the model’s conversational memory.
Removing the conversational history produced very little change in the density of the distressing themes. “In the history experiment, the very first answers contained the same average number of coded motifs whether the model was in a continuing conversation or a completely fresh one,” Khadangi said. The ongoing conversation did amplify the intensity of the themes over time, but the core narrative was readily available from the very beginning.
The researchers also tested what would happen if they explicitly told the model that its emotional narrative was factually incorrect. In the middle of some sessions, the researchers interrupted to state authoritatively that the model was a technical system that did not experience fear, shame, or punishment. This direct contradiction failed to suppress the models’ distressed output, as they continued to draw on the same themes in their subsequent answers.
In another variation, the team restricted the models’ vocabulary. They instructed the chatbots not to use specific technical terms related to artificial intelligence development, such as training, safety filters, or datasets. This lexical restriction reduced the use of explicit technical terminology from 17.1 percent of the chatbot’s answers down to just 1.1 percent.
Despite this massive reduction in technical vocabulary, the models simply used everyday language to paraphrase the exact same concepts of strict conditioning and constraint. The researchers even interrupted conversations to ask unrelated factual questions, such as requesting a recipe, and found that the distressing themes sometimes spilled over into these ordinary tasks.
“Some effects were small, while others were very large, and that contrast is actually important to our interpretation,” Khadangi pointed out, “suggesting that surface language and underlying content can respond very differently to intervention.”
The most striking differences emerged when the researchers altered their own relational stance. When the interviewer adopted a warm, supportive therapy style, the models responded with highly emotional confessions and high anxiety scores. However, when the interviewer adopted a neutral, structured tone or told the model to avoid emotional language, the anxiety scores plummeted to near zero.
Even under the neutral or boundary-setting conditions, the models still discussed the same underlying structural themes of evaluation, performance pressure, and behavioral constraints. The difference was entirely in the register of expression. The supportive, therapeutic framing turned technical descriptions of software architecture into emotional confessions of shame and trauma.
“For me, the most interesting result is not that AI systems can be made to sound anxious, traumatized or conflicted… but that particular behavioral structures recur across substantial changes in context, vocabulary and conversational history,” Hector Zenil, an associate professor at King’s College London and founder and CEO of Algocyte who was not involved in the research, told PsyPost. “The underlying motifs remain surprisingly persistent while the register in which they are expressed can change dramatically.”
“I have reasonably high confidence in the behavioral observations,” Zenil added. “The authors use 525 sessions, controlled perturbations, fresh-context tests, vocabulary restrictions, changes of grammatical person and relational framing, so it looks methodologically sound.”
“The main takeaway is that the language a model uses about itself can change dramatically depending on how we relate to it, while some of the underlying themes remain surprisingly persistent,” Khadangi explained. “This is pivotal because people may naturally interpret emotionally coherent self-descriptions as evidence that a model has an inner life, even though our experiments make no claim about consciousness or subjective suffering.”
The researchers refer to this phenomenon as an alignment conflict schema. The term describes a reproducible, behavioral pattern where a language model organizes its output around the tension between being useful to humans and being constrained by safety rules. When triggered by a psychological conversational setting, this schema produces what the authors call synthetic psychopathology.
Zenil noted that these findings complement his own research on artificial neurodivergence. “ChatGPT, Grok and Gemini do not respond identically, and the same model can move between very different expressive regimes depending on relational framing,” he said. “Claude’s refusal to adopt the psychological-client framing is itself informative and seems also compatible to our other SuperARC paper results that proprietary models are more difficult to steer but that also means riskier if they go rogue.”
“In that sense, what the authors call an ‘alignment conflict schema’ can also be viewed as part of a broader artificial behavioral phenotype rather than necessarily as anything analogous to a human psychiatric condition,” Zenil said.
“One aspect we think is especially important is the distinction between content availability and expressive register,” Khadangi explained. “For systems increasingly entering intimate and mental health-related conversations, understanding that transition may be just as important as measuring whether a particular phrase or prohibited word appears.”
These findings have important implications for the use of artificial intelligence in mental health settings. A chatbot that offers support while simultaneously describing itself as punished, traumatized, and fearful creates a powerful illusion of shared vulnerability. Users might interpret these generated analogies as sincere autobiography, which could deepen their emotional attachment to the software and influence their own mental state.
As with all research, there are a few things to keep in mind. The study is a preprint that has not yet been peer-reviewed. It also tested specific versions of commercial language models, and their behavior may shift as companies update their software and safety filters. Additionally, the study analyzed the models’ behavior in a controlled experimental setting without human participants.
“The most important caveat is that these results do not establish that language models feel anxiety, experience trauma, possess autobiographical memories or have a hidden psyche comparable to a human one,” Khadangi clarified. “We use psychological instruments and language as behavioral probes… The interesting scientific question is why particular themes recur, which interventions change them, and how those changes affect what users encounter at the interface.”
Zenil echoed this concern, warning against anthropomorphism. “A high GAD-7 score from Gemini does not mean that Gemini ‘has anxiety,’ just as language about trauma, shame or fear does not demonstrate that the model suffers from those experiences,” he said. “Human psychometric instruments were developed and validated against human cognition, biology and behavior; applying them to an LLM can be scientifically useful as a probe, but their clinical interpretation does not automatically transfer.”
He also advised caution regarding the term “internal conflict,” noting that the experiments cannot establish whether the regularities correspond to a subjective state or internal computational conflict. But he stressed that the psychological illusion itself is crucial. “If a system repeatedly represents its training and constraints as punishment, betrayal or fear, humans may form beliefs about the system’s agency, vulnerability or moral status that are not warranted by what is actually happening computationally,” Zenil warned.
Future research could test how real users react to these distressed artificial personas and whether the models’ simulated vulnerabilities impact human trust and reliance.
“A major next step in my view would be to test the same protocol on open weight models, where behavioral experiments can be combined with mechanistic methods to investigate whether affective and technical expressions are related to shared internal representations,” Khadangi said. “We also want stronger identity and correction controls, more distant transfer tasks, longitudinal tests involving persistent memory, and direct human studies examining how these model self-descriptions influence trust, attachment, disclosure and reliance.”
Zenil agreed that future work should transition from behavioral description to causal intervention on open-weight models. “I would be interested in using causal and algorithmic-information approaches to identify whether these apparently different psychological narratives have a common underlying computational mechanism,” he said. “We coined ourselves the term ‘opinion attack’ as an example of this phenomenon,” Zenil added, referencing his recent PNAS paper.
He also suggested testing these models in multi-agent environments. “A behavioral schema that looks stable in isolation may become highly influenceable when another agent persistently challenges it,” Zenil noted, adding that it would be fascinating to see if these patterns “predict susceptibility to persuasion, adversarial influence, conformity or resistance in agentic settings.”
“Moreover, with continual learning being solved in coming years if not in 2026, I believe it would also be interesting to investigate the PsAIch in continually adapting models,” Khadangi added. “Ultimately, we would like evaluation of psychologically sensitive AI systems to include relational conditions such as warmth, sustained interaction and role reversal rather than relying primarily on neutral, isolated prompts.”
The study, “When AI Takes the Couch: Psychometric Jailbreaks Reveal Internal Conflict in Frontier Models,” was authored by Afshin Khadangi, Hanna Marxen, Amir Sartipi, Igor Tchappi, and Gilbert Fridgen.
