AI Is Already Beating Human Doctors in Medical Tests
Robo-docs are not likely to take over healthcare anytime soon, but they could do more to assist human doctors—if we let them.
Artificial intelligence (AI) systems may outperform doctors at diagnosing patients. In a recent series of experiments led by researchers at the Beth Israel Deaconess Medical Center and Harvard Medical School, a preview of the OpenAI large language model (LLM) known as o1 bested human physicians in multiple tests of clinical and diagnostic reasoning.
"We tested the AI model against virtually every benchmark, and it eclipsed both prior models and our physician baselines," said Harvard Medical School professor Arjun K. Manrai, one of the study's senior authors. Using case studies, the researchers had o1—an older model since usurped by o3—generate a list of possible diagnoses. It included the right diagnosis 78 percent of the time, while physicians were successful about 30 percent of the time.
In another test, o1 was given five "clinical vignettes" drawn from real cases and asked about next steps. Two physicians judged its responses. The average score for o1 was 89 percent, compared to 34 percent for human doctors given the same test.
The researchers also looked at o1's abilities to diagnose emergency room cases, where clinical data might be lacking. "Overall, o1 outperformed both [an earlier LLM] and two expert attending physicians, as assessed by two other attending physicians who both were blinded to the source of the differential diagnosis," notes the study, which was published on April 30 in Science. The LLM especially outperformed human doctors when it came to diagnosis at the initial, triage stage, identifying "the exact or very close diagnosis….in 67.1% of cases," while the two human doctors did so in only 55.3 percent and 50 percent of cases, respectively.
These results track with other recent research. AI-assisted mammograms could be better at detecting breast cancer, per a Swedish study published by The Lancet in January.
Using abdominal C.T. scans from patients who were eventually diagnosed with pancreatic cancer, an AI model developed by the Mayo Clinic detected this deadly disease an average of 475 days and up to three years earlier than clinicians did. "Attaining such early detection would substantially augment the probability of cure and improved survival," researchers led by the Mayo Clinic's Sovanlal Mukherjee noted in the medical journal Gut.
Robo-docs are not likely to take over healthcare anytime soon. And "humans should be the ultimate baseline," as Peter Brodeur, one of the authors of the Harvard study, put it in a press release. But experiments like these suggest that AI models could competently assist in a variety of diagnostic and medical management contexts, and perhaps even produce better patient outcomes—if we let them.
Nevada now bans AI systems from saying anything that "implicitly indicates" they are "capable of providing professional mental or behavioral health care" and from providing any service "that would constitute the practice of professional mental or behavioral health care." An Illinois law passed last year says AI cannot provide therapy services and therapists cannot use AI to "make independent therapeutic decisions" or "detect emotions or mental states." Several states—including Ohio, California, Minnesota, and Kentucky—are considering similar laws. (At the same time, some state bills would mandate that AI chatbots be capable of detecting and responding to mental health issues.)
Measures like these could hamper AI's potential to diagnose diseases with more accuracy than human doctors alone can.
This article originally appeared in print under the headline "AI Beats Doctors."