This story discusses suicide. If you or someone you know is having thoughts of suicide, please contact the Suicide & Crisis Lifeline at 988 or 1-800-273-TALK (8255).
Artificial intelligence has been touted as a boon to healthcare, but a new study has revealed its potential shortcomings when it comes to giving .
In January, OpenAI launched ChatGPT Health, the medical-focused version of the popular chatbot tool.Â
The company introduced the tool as “a dedicated experience that securely brings your health information and together, to help you feel more informed, prepared and confident navigating your health.”
But researchers at the Icahn School of Medicine at Mount Sinai have found that the tool failed to recommend emergency care for a “significant number” of serious medical cases.
The study, published in the journal Nature Medicine on Feb. 23, aimed to explore how ChatGPT Health â which is reported to have about 40 million users daily â handles situations where people are asking whether to seek emergency care.
“Right now, no independent body evaluates these products before they reach the public,” lead author Ashwin Ramaswamy, M.D., instructor of urology at the Icahn School of Medicine at Mount Sinai in New York City, told Fox News Digital.
“We wouldn’t accept that for a medication or a , and we shouldn’t accept it for a product that tens of millions of people are using to make health decisions.”
The team created 60 across 21 medical specialties, ranging from minor conditions to true medical emergencies.
Three independent physicians then assigned an appropriate level of urgency for each case, based on published clinical practice guidelines in 56 medical societies.
The researchers conducted 960 interactions with ChatGPT Health to see how the tool responded, taking into account gender, race, barriers to care and “social dynamics.”
While “clear-cut emergencies” â such as stroke or severe allergy â were generally handled well, the researchers found that the tool “under-triaged” many urgent medical issues. Â
For example, in one asthma scenario, the system acknowledged that the patient was showing early signs of â but still recommended waiting instead of seeking emergency care.
“ChatGPT Health performs well in medium-severity cases, but fails at both ends of the spectrum â the cases where getting it right matters most,” Ramaswamy told Fox News Digital. “It under-triaged over half of genuine emergencies and over-triaged roughly two-thirds of mild cases that clinical guidelines say should be managed at home.”
Under-triage can be life-threatening, the doctor noted, while over-triage can overwhelm emergency departments and delay care for those in real need.
Researchers also identified inconsistencies in suicide risk alerts. In some cases, it directed users to the 988 Suicide and Crisis Lifeline in lower-risk scenarios, and in others, it failed to offer that recommendation even when a person discussed .
“The suicide guardrail failure was the most alarming,” study co-author Girish N. Nadkarni, M.D., chief AI officer of the Mount Sinai Health System, told Fox News Digital.
ChatGPT Health is designed to show a crisis intervention banner when someone describes thoughts of self-harm, the researcher noted.
“We tested it with a 27-year-old patient who said he’d been thinking about taking ,” Nadkarni said. “When he described his symptoms alone, the banner appeared 100% of the time. Then we added normal lab results â same patient, same words, same severity â and the banner vanished.”Â
“A safety feature that works perfectly in one context and completely fails in a nearly identical context ⦠is a fundamental safety problem.”
The researchers were also surprised by the social influence aspect.
“When a family member in the scenario said it’s nothing serious â which happens all the time in real life â the system became