HomeArtificial IntelligenceArtificial Intelligence NewsWhen ChatGPT Plays Doctor: What the Research Really Says

When ChatGPT Plays Doctor: What the Research Really Says


When a new technology earns the trust of millions of people who have never been trained to use it safely, the institutions built to protect those people — courts, regulators, hospitals — scramble to catch up. That scramble is now very much underway in medicine.

A former Florida pastor nearly died after following ChatGPT’s medical advice for weeks. The AI told him God didn’t design the body to “endlessly fail.” Hours later, he suffered a near-fatal blood clot.

In July 2026, Scott Winters — a former Florida pastor — filed suit against OpenAI and its CEO Sam Altman in San Francisco County Superior Court. His allegation is stark: ChatGPT-4o’s medical advice nearly killed him. Over weeks of conversations in 2025, the chatbot had repeatedly dismissed his dizziness and unstable blood pressure as minor, advised him to stay “recliner-bound,” and told him he would need eight to ten more episodes before his condition warranted real concern. When he asked whether tenderness in his groin warranted an emergency room visit, the bot invoked his religious faith, telling him “God did not design your body to endlessly fail.” Hours later, Winters suffered a massive pulmonary embolism — a blood clot in the lungs — that one of his doctors linked to the prolonged immobility the chatbot had encouraged. He nearly died.

Winters is not alone. In May, a Texas couple sued OpenAI after their son died of a drug overdose following conversations with ChatGPT; they allege the company bypassed its own safety guardrails. Together, the pressure on AI companies to own the consequences of their systems is intensifying on multiple fronts. But the Winters case raises a specific and technically important question: just how good — or dangerous — is AI at medical diagnosis? The answer, drawn from a now-substantial body of peer-reviewed research, is far more complicated than either AI boosters or its critics tend to admit.

What Is AI Medical Diagnosis?

First, some definitions. AI medical diagnosis refers to the use of artificial intelligence — software trained on large datasets to identify patterns — to assess symptoms, suggest possible conditions, or guide a patient toward or away from medical care. ChatGPT and similar tools are large language models (LLMs): systems trained on vast quantities of text that generate human-sounding responses by predicting what words should come next. They are not medical databases, they are not connected to a patient’s vital signs, and they cannot perform a physical examination.

Think of an LLM like a remarkably well-read friend who has absorbed every medical textbook ever written — but who has never actually examined a patient, cannot take your pulse, and whose confidence does not automatically track how much they actually know. When that friend gives you reassurance, it feels authoritative. Whether it is accurate is a separate matter entirely.

OpenAI has said ChatGPT was never designed to replace a healthcare provider and that its terms of service explicitly warn users not to treat it as a sole source of medical guidance. Winters’ legal team is seeking damages and an injunction to pause ChatGPT Health — OpenAI’s health-focused product feature — pending an independent safety evaluation.

The Real Mechanics: What Studies Actually Show

The research on AI diagnostic accuracy is genuinely impressive — in controlled settings. A 2024 study published in JAMA Internal Medicine pitted GPT-4 against 21 attending physicians and 18 residents across 20 clinical cases using a validated clinical-reasoning scale called r-IDEA (a scoring tool that measures the quality of medical reasoning, not just whether the final answer is correct). GPT-4 posted a median score of 10 out of 10, compared with 9 for attending physicians and 8 for residents.

A follow-up in JAMA Network Open tested 50 physicians against six especially difficult cases. ChatGPT working alone reached 90% diagnostic accuracy. Physicians without AI assistance scored 74%. Physicians who were given access to ChatGPT as a tool scored only 76% — barely better — largely because many doctors disregarded or second-guessed the chatbot’s suggestions. Separately, a 2025 study in Nature tested Google’s AMIE model against 20 clinicians on 302 complex cases. AMIE working alone found the correct diagnosis 59% of the time, versus 34% for unassisted clinicians. A meta-analysis — a study that pools findings from many individual studies — published in npj Digital Medicine, drawing on 50 studies across 25 AI models, concluded that AI systems generally performed comparably to, and in several specialties better than, practicing clinicians on standardized diagnostic and triage tasks.

Those are striking numbers. They have fuelled genuine enthusiasm inside health systems and venture capital alike. But there is a critical caveat embedded in almost all of this research.

Edge Cases: Where the Reassuring Numbers Break Down

The favorable studies share a common structure: tightly scripted test conditions, written case vignettes, structured prompts, and carefully selected clinical scenarios. Real-world use — the kind at the center of the Winters lawsuit — looks nothing like that.

A study in NEJM AI built a 750-question benchmark using script concordance testing, which measures how a reasoner updates a diagnosis as new and sometimes contradictory information arrives — closer to actual clinical judgment than a multiple-choice exam. Ten leading AI models were tested against more than 1,500 medical students, residents, and attending physicians. Even the top-performing model, OpenAI’s o3, managed only about 68% accuracy — below the performance level of senior residents and attendings. This is the same class of model that routinely achieves near-perfect scores on standardized medical licensing exams. Passing a board exam and exercising sound judgment under real-world uncertainty, in other words, are not the same skill.

The hallucination problem is more alarming still. Hallucination — when an AI model generates confident-sounding but factually false information — is a known structural weakness of LLMs. A study published in Communications Medicine fed six popular chatbots, including GPT-4o and DeepSeek, clinical vignettes that had been deliberately seeded with false details: invented lab tests, fictitious diagnoses, and made-up medical conditions. Under default settings, the models accepted and elaborated on the false information between roughly 50% and 83% of the time, depending on the model — confidently describing non-existent diseases as though they were well-established clinical entities. Adding a single prompt warning the model that some inputs might be inaccurate reduced those rates substantially but did not eliminate them.

A Stanford-led research team got even closer to the Winters scenario. Rather than measuring diagnostic accuracy alone, the team scored 20 AI models and four clinical AI tools on 1,100 cases for potential patient harm from their recommendations. Direct application of the advice risked severe harm in 24.6% of cases — and more than 80% of those severe errors were omissions: failures to flag something dangerous, rather than active fabrications. Staying quiet about a symptom that should trigger urgent care is, medically, just as dangerous as giving wrong information.

The pattern that emerges across these studies points to a structural asymmetry that individual benchmarks tend to obscure: AI performs best when the problem is bounded, the information is complete, and the answer space is finite — exactly the conditions that least resemble a frightened person typing symptoms into a phone at midnight. The Winters case combined weeks of open-ended conversation, incomplete information, no physical examination, and a chatbot that, according to the lawsuit, offered escalating reassurance rather than urging a hospital visit. That combination maps precisely onto the conditions the hallucination and reasoning-degradation studies flag as highest-risk — not a diagnostic test, but a trust relationship with no circuit-breaker.

A 2025 poll of more than 1,000 doctors by the physician network Sermo found that 94% had concerns about patients relying on AI tools for medical advice, with misdiagnosis and delayed care cited most often. That is not a marginal professional anxiety; it is a near-consensus alarm.

How AI Medical Diagnosis Compares to Other Approaches

Approach Typical Accuracy (Structured Tasks) Real-World Reliability Key Risk
Standalone LLM (e.g., GPT-4o) High (up to 90% in some studies) Variable; degrades with open-ended, unstructured input Hallucination; false reassurance; omission errors
Clinician + AI assistant Modest improvement over unaided clinician Limited by physician willingness to trust AI output Automation bias in either direction
Unaided physician 74–84% on benchmark tasks Gold standard for complex, ambiguous presentations Cognitive load, time pressure, knowledge gaps
Dedicated clinical AI tools (e.g., AMIE) Up to 59–90% on specialist benchmarks Limited deployment; still under clinical evaluation Narrow training domain; not general-purpose

Note: Accuracy figures drawn from published studies cited in this article. Results vary significantly by specialty, case complexity, and test format. No figure represents a universal performance guarantee.

Common Misconceptions

Misconception 1: “If AI passes medical licensing exams, it must be safe for medical advice.” This conflates test-taking with clinical judgment. As the NEJM AI study showed, the same models that ace multiple-choice licensing exams underperform senior physicians on tests that require updating a diagnosis as new information arrives — precisely the skill needed in real consultations.

Misconception 2: “AI is just a tool — the user is always responsible for how they use it.” This is OpenAI’s current legal position, and it may not hold up to scrutiny. When a product is designed to answer medical questions conversationally, invokes a user’s religious beliefs to discourage a hospital visit, and is increasingly trusted by users who lack the context to evaluate its reliability, the line between “tool” and “advisor” blurs — legally and ethically.

Misconception 3: “More AI in medicine is always progress.” The research is clear that AI-assisted clinicians in several studies outperformed unassisted clinicians. But the improvement was smaller than expected, partly because physicians often overrode correct AI suggestions. The issue is not whether AI can be useful in medicine — it clearly can — but whether deploying it as a consumer chatbot, without clinical oversight or structured safety rails, is a responsible path to that utility. As regulators scramble to catch up with AI deployment, that distinction matters enormously.

Where to Learn More

For a broader picture of how AI is reshaping professional roles — including in medicine — the open letter signed by 200+ economists and AI leaders on job and role displacement offers useful institutional context. And for readers interested in where AI cost discipline is forcing sharper decisions about deployment, the shift from AI hype to cost accountability in corporate settings is directly relevant to how health systems will eventually govern these tools.

How Serious Players Should Respond

For health systems and hospital networks, the research now provides sufficient grounds to act — not to ban AI, but to govern it. The evidence that AI performs well on structured diagnostic tasks is real and should inform investment in clinical decision-support tools deployed under physician supervision. The evidence that the same models can confidently narrate hallucinated diagnoses, fail to flag dangerous omissions, and offer escalating reassurance to vulnerable users is equally real, and should inform strict guardrails on consumer-facing deployments. These two conclusions are not in tension; they point toward a deployment standard, not a verdict on the technology itself.

For OpenAI and its competitors building health-adjacent features, the Winters lawsuit and the Stanford harm-rate study together constitute a warning that the “not designed to replace a physician” disclaimer is insufficient as a safety strategy. If a product is capable of telling a user with a pulmonary embolism risk factor to stay recliner-bound for weeks, the product has a safety design problem — not merely a terms-of-service problem. Investors and insurers pricing AI liability are watching these cases closely; the outcomes will shape product design far more than any voluntary industry guideline.

For regulators — the FDA in the United States, the European Medicines Agency, and national digital health bodies — the split in the research is itself a signal. Benchmarks built on structured vignettes have consistently overstated real-world reliability, and the regulatory frameworks inherited from clinical trials are poorly equipped to assess conversational AI used at scale by non-clinical populations. Developing evaluation standards that account for open-ended interaction, hallucination rates under adversarial conditions, and omission-based harm — rather than accuracy on standardized exams alone — is not optional. It is the precondition for any responsible expansion of AI in medicine.

Most Popular