PsyPost
  • Mental Health
  • Social Psychology
  • Cognitive Science
  • Neuroscience
  • About
No Result
View All Result
Join
My Account
PsyPost
No Result
View All Result
Home Exclusive Artificial Intelligence

Study finds AI chatbots can conduct basic therapy sessions but struggle to adapt to individual needs

by Eric W. Dolan
August 24, 2026
Reading Time: 6 mins read
[Adobe Stock]

[Adobe Stock]

Share on TwitterShare on Facebook

A recent study published in the journal Computers in Human Behavior: Artificial Humans suggests that artificial intelligence chatbots can conduct basic therapeutic conversations, though they struggle to consistently apply specific treatment techniques. The research indicates that a language model performed similarly to human practitioners in high-quality research settings, but its ability to adapt specialized psychological methods to individual needs remains inconsistent.

The global shortage of accessible mental health care has led experts to consider technology as a possible bridge. According to global health estimates cited by the study authors, hundreds of millions of individuals live with mental health conditions, yet structural barriers like cost and geographic distance leave the vast majority without professional care.

For example, the authors of a 2024 framework published in npj Mental Health Research argued that large language models hold tremendous potential to expand access to personalized treatments. A large language model (also known as an LLM) is a type of artificial intelligence trained on vast amounts of text, allowing it to generate human-like responses to prompts.

The authors of the 2024 paper proposed a roadmap for integrating these systems into clinical care, noting that psychotherapy is a high-stakes environment requiring nuanced expertise and responsible development. Despite this optimism, the actual clinical skills of these language models have remained largely untested in real-world scenarios. Previous research mostly relied on giving chatbots brief, fictional scenarios and asking human judges to rate the isolated responses for empathy or helpfulness.

These brief tests provide little insight into whether a computer program can manage a full, goal-directed therapy session that requires a coherent strategy. In a real session, a practitioner must dynamically adjust to the user’s changing emotional state and guide the conversation toward a beneficial outcome. This gap in knowledge motivated the authors of the new study to evaluate how well a language model actually performs when conducting a live, uninterrupted therapy session with a human user.

“Previous research had mostly examined LLM-chatbots’ skill at doing therapy by having them generate brief therapeutic responses or short excerpts of sessions, often in reply to a narrow set of client problems,” said study author Arthur Bran Herbener, a researcher at the Department of Psychology and Behavioral Sciences at Aarhus University. “We needed research that examines how skillfully LLM-chatbots can carry out full cognitive behavioral therapy sessions and adapt treatment to different individuals, using well-established, standardized metrics of therapeutic competence.”

To explore this question, the researchers designed an observational study involving 65 university students. These participants were experiencing mild to moderate psychological distress, such as presentation anxiety, occasional worrying, or perfectionism. Students with severe distress or diagnosed mental health disorders were excluded to ensure participant safety, as language models can sometimes generate unpredictable or inappropriate responses. Each participant completed a single 30-minute in-person session with a locally hosted artificial intelligence chatbot, exchanging an average of 49 messages.

The chatbot was programmed to deliver Cognitive Behavioral Therapy, a widely used, problem-oriented treatment that helps individuals identify and change unhelpful thoughts and behaviors. The researchers configured a specific language model to progress through the typical phases of a session. Rather than giving the chatbot a static set of rules, the researchers used a secondary background model to monitor the ongoing conversation. This secondary model evaluated the elapsed time and the current context, guiding the primary chatbot to shift from establishing a connection to conceptualizing the problem, applying an intervention, and finally wrapping up the conversation politely.

Google News Preferences Add PsyPost to your preferred sources

After the sessions concluded, trained graduate students read the conversation transcripts and rated the chatbot’s performance using the Cognitive Therapy Scale. This standardized assessment tool measures two main areas of clinical proficiency. The first area covers general therapeutic skills, such as expressing warmth and understanding. The second area covers specific skills, such as applying targeted techniques and guiding the user toward new insights.

To provide a benchmark for the chatbot’s scores, the researchers also conducted a meta-analysis of 18 prior studies that used the exact same scale to evaluate human practitioners. A meta-analysis is a statistical technique that combines the data from multiple independent studies to find an overall average or trend. This technique allowed the researchers to compare the chatbot’s performance to an established baseline of human competence.

The researchers found that the chatbot’s overall competence score fell slightly below the generally accepted threshold for adequate clinical performance. The standard rating scale defines an acceptable level of competence as a score of 40 out of a possible 66 points. The chatbot achieved this minimum threshold in 30 of the 65 sessions, showing a high degree of variability from one conversation to the next.

“We were surprised by how much variation the LLM-chatbot showed in its skillfulness across CBT sessions,” Herbener told PsyPost. “This is an important observation, as it suggests that we need research to ensure consistently competent care across individuals, and to understand when and why performance dips.”

When compared to the broad pool of human practitioners from the meta-analysis, the chatbot scored somewhat lower overall. The human professionals scored an average of 40.3 points on the rating scale, compared to the chatbot’s adjusted score of 38.1 points. However, the performance gap disappeared when the researchers looked only at the most rigorously conducted human studies. Compared to the six human studies judged to be of high methodological quality, the chatbot’s scores exhibited no statistically significant difference.

A closer look at the types of skills displayed by the chatbot highlighted a distinct pattern in its capabilities. The artificial intelligence performed better than human practitioners in general therapeutic skills, such as communicating empathy, validating the user’s feelings, and fostering a collaborative environment. In contrast, the chatbot struggled with the specific, technical skills required for this highly structured type of therapy. It had difficulty reliably identifying key beliefs, guiding the user through self-discovery, and adapting intervention strategies to fit the unique characteristics of each participant.

These technical struggles highlight the difference between following a predetermined structure and tailoring a method to a specific person. “LLM-chatbots designed to deliver therapy show promise in adhering to the cognitive behavioral treatment protocol, but there may be challenges in adapting to different individuals,” Herbener explained. “It’s also worth noting that good observable skill in delivering therapy is not the same as clinical effectiveness.”

“Effectiveness likely depends on factors beyond observable skill, such as positive expectations and a strong therapeutic relationship,” Herbener continued. “More comprehensive assessments of LLM-chatbots’ therapeutic competence may also require entirely new assessment approaches that account for LLMs’ distinct behavioral tendencies — such as sycophancy, or a tendency to excessively affirm users. Understanding this is crucial, since the ability to challenge clients’ beliefs is often considered important for fostering clinical change.”

These findings do not indicate that artificial intelligence is ready to replace human practitioners. The study evaluated the chatbot based on a single session with young adults experiencing only mild distress. Clinical populations often present more complex challenges, including severe hopelessness or safety risks. These complex cases place much higher demands on a practitioner’s ability to adapt and respond skillfully to unpredictability.

A single session also cannot capture the long-term planning, homework review, and relationship-building required in a complete, multi-week treatment program. In addition, the study faced challenges in consistently rating the chatbot’s text-based transcripts. The standard rating scale was originally designed for video or audio recordings, where raters can hear tone of voice and observe body language. Applying this tool to text may have made certain interpersonal skills harder to judge accurately.

Raters also knew they were evaluating an artificial intelligence, which may have influenced their scoring. In real-world applications, this lack of blinding reflects how users will actually interact with known computer systems, but it complicates direct comparisons to human professionals.

The researchers also caution against drawing overly broad conclusions from the matched scores between the chatbot and high-quality human studies. “Showing similar competence levels in CBT between human therapists and an LLM-chatbot does not mean they are equally effective ‘therapists’, nor that they are equally competent in an absolute sense,” Herbener said.

Because the human data came from past studies rather than a side-by-side test, the comparison remains indirect. “We did not directly compare the chatbot and human therapists in an experimental setting, so several factors beyond competence could bias our measurements,” Herbener added. “For example, the severity and type of psychological problems presented, or the norms for inferring behavioral observations clinicians and researchers apply when using the competence scale we relied on.”

To advance the field, future research will need to examine how chatbots perform across multiple sessions with clinical populations and investigate exactly why they sometimes fail to apply specific therapeutic techniques. “Research on LLM-chatbots in mental healthcare is still at a very early stage,” Herbener said. “Alongside design efforts aimed at ensuring consistently competent care, I’m also working to better understand the therapeutic relationship in LLM-based mental healthcare.”

Even if a chatbot can say all the right things, a user’s awareness that they are talking to a machine might alter the impact of those words. “We not only need to know how well these systems mimic human therapists’ language. We also need to understand what meaning clients attribute to that language once they know it comes from a machine,” Herbener explained. “Does the positive regard expressed by a chatbot serve the same clinical function as when it comes from a human? That’s one of the big open questions in the field.”

The study, “Exploring the therapeutic competencies of large language models: Observational study and comparison with meta-analytical estimates for human therapists,” was authored by Arthur Bran Herbener, Robert Zachariae, Michal Klincewicz, Mikkel Berg Thøgersen, Marie Rosenkrantz Hermann, and Malene Flensborg Damholdt.

TweetSendScanShareSendPinShareShareShareShareShare

Follow PsyPost

The latest research, however you prefer to read it.

Daily newsletter

One email a day. The newest research, nothing else.

Google News

Get PsyPost stories in your Google News feed.

Add PsyPost to Google News
RSS feed

Use your favorite reader.

Copy RSS URL
Social media
Support independent science journalism

Ad-free reading, full archives, and weekly deep dives for members.

Become a member

Trending

  • People who favor dark humor tend to exhibit more problematic personality traits
  • Weight-loss drugs linked to rare eye condition in massive data review
  • The “sense of absence” may explain why disconnected teens turn to addictive behaviors
  • Oxytocin has opposite effects on men’s generosity depending on a their childhood experiences
  • New study shed light on how estradiol influences memory networks in middle-aged women

Science of Money

  • When crypto traders are feeling optimistic, they pay less attention to the economy
  • Shoppers hate skimpflation more than shrinkflation, new research finds
  • When a CEO sounds upbeat, Wall Street analysts tend to follow, even when they shouldn’t
  • When central bank bond-buying pays for itself: A new look at QE’s fiscal scorecard
  • Can AI Design a Better Ad? A New Study Puts It to the Test

Recent

  • Misjudging personal space drives the need for distance in socially anxious people
  • The physiological reasons the human brain struggles to process information in extreme heat
  • Breaking the law makes people seem less moral, even when the act is unchanged
  • Psychology researchers uncover an unexpected benefit of dance training
  • New study: Men and women multitask equally well, but men talk less while doing it
  • Multi-ancestry study explores how DNA and trauma affect smoking habits
  • Study suggests even casual cannabis use is associated with depression in older adults
  • Survival of the wittiest: How humor and cleverness shaped human evolution
  • Artificial intelligence agents spontaneously conform to the majority opinion
  • Psychological models suggest female libido disparities stem from early behavioral conditioning

PsyPost is a psychology and neuroscience news website dedicated to reporting the latest research on human behavior, cognition, and society. (READ MORE...)

  • Mental Health
  • Neuroimaging
  • Personality Psychology
  • Social Psychology
  • Artificial Intelligence
  • Cognitive Science
  • Psychopharmacology
  • Contact us
  • Disclaimer
  • Privacy policy
  • Terms and conditions

(c) PsyPost Media Inc

Welcome Back!

Login to your account below

Forgotten Password?

Retrieve your password

Please enter your username or email address to reset your password.

Log In

Add New Playlist

Subscribe
  • My Account
  • Cognitive Science Research
  • Mental Health Research
  • Social Psychology Research
  • Drug Research
  • Relationship Research
  • About PsyPost
  • Contact
  • Privacy Policy

(c) PsyPost Media Inc