Mental health-related content accounted for 0.21% to 4.90% of conversations in a public ChatGPT dataset depending on the definition applied, according to a cross-sectional research letter published in
JAMA Network Open.
Researchers analyzed 620,699 conversations from WildChat-4.8M, a public corpus of human-chatbot exchanges. They developed a large language model (LLM)-based classifier that scored conversations across mental health topicality, help-seeking intent, clinical language, and affective risk. The conservative definition required an identifiable real person and explicit help-seeking for a psychological problem, while the expansive definition required mental health topicality regardless of person, focus or explicit help-seeking.
The classifier was validated in a random sample of 370 conversations coded by 2 human reviewers. Twenty-two discordant ratings were adjudicated by a third reviewer blinded to the initial ratings. Human reviewers had 94% raw agreement and interrater reliability of κ = 0.65. Against the adjudicated ratings, the classifier had 87% sensitivity, 95% specificity, a positive predictive value of 67%, and a negative predictive value of 98%.
Under the conservative definition, 1317 conversations were classified as mental health related, representing 0.21% of the dataset (95% CI, 0.20%-0.22%). The expansive definition identified 30,394 conversations, or 4.90% (95% CI, 4.84%-4.95%), a 23-fold difference. Conversations meeting the conservative definition had higher topicality and intent scores, while those meeting the expansive definition more often lacked an identifiable person focus or explicit help-seeking.
The authors said prevalence varied by more than an order of magnitude depending on definitional scope and recommended a tiered taxonomy distinguishing crisis or high-risk, clinically framed, and broad affective or interpersonal conversations. Limitations included use of a public transcript corpus without external validation, potential automated misclassification of ambiguous distress, judgment in adjudicating borderline ratings, and no assessment of whether taxonomy-based monitoring improves population health outcomes.
Source: McBain RK, Zhang LA, Burnett A, Cantor JH, Yu H. Prevalence of mental health discussions in publicly available generative AI conversations.
JAMA Netw Open. 2026;9(8):e2630635. doi:
10.1001/jamanetworkopen.2026.30635