The Persuasively Misinformed Patient: Why AI’s Fluency Is a Double-Edged Sword in Healthcare
Clinicians are facing an unintended consequence of the generative AI boom: the persuasively misinformed patient. The same person who uses a chatbot to draft an email or plan dinner is now using it to evaluate skin lesions and symptoms. Because the AI returns a polished, articulate explanation, the user walks away with an unearned sense of expertise.
When that diagnosis is wrong, it creates a frustrating new reality in the exam room. A patient can walk into an appointment convinced of an AI’s verdict, leaving the doctor with two jobs instead of one: figuring out what is actually going on, and gently talking the patient out of a wrong answer they already believe.
A new study in Nature Medicine, led by DBMI assistant professor Orson Xu, puts real numbers behind that dynamic — and reveals that the effect fades sharply, though not completely, with clinical training.
The Gap Between Sounding Right and Being Right
Researchers evaluated how nearly 800 participants, ranging from laypeople to primary care physicians, interacted with an AI dermatology tool that explained its reasoning in plain language. For laypeople, the AI’s fluency cut both ways: a persuasive explanation made them more likely to trust the AI’s call, regardless of whether that call was accurate.
“In our study, LLM explanations were a double-edged sword for laypeople: when the AI was right, they produced the largest improvement in diagnostic accuracy—about 13%,” Xu said. “But when the AI was wrong, they also produced the largest decline—about 21%. A fluent explanation can make an answer feel more reasonable without making it more correct.”
Physicians, notably, did not follow this pattern. Their diagnostic accuracy held steady even when the AI delivered bad advice accompanied by a convincing rationale.
Expertise as a firewall
The resilience physicians demonstrated wasn’t just a byproduct of general medical background—it scaled directly with real-world clinical experience. When the researchers gave the same diagnostic task to medical students, who possess clinical knowledge but far less time in practice, the students were noticeably more deferential to incorrect AI advice than experienced doctors.
The difference, Xu explained, comes down to how an explanation is used.
“Experienced primary care physicians were much more resilient when the AI made mistakes,” Xu said. “For an expert, an explanation can be evaluated against an existing body of knowledge; for a less-experienced user, the explanation itself may become part of the evidence that convinces them.”
In other words, a physician uses an AI’s explanation to check their own thinking. A layperson — or a less experienced clinician — is more likely to let the explanation become their thinking.
The firewall has a weak point
That clinical resilience is not absolute. The sequence in which information is presented turned out to matter almost as much as who was evaluating it. When clinicians saw the AI’s suggestion before forming their own independent judgment, they became noticeably more likely to defer to it — a pattern the researchers observed in laypeople too but hadn’t expected to see so clearly in experienced clinicians.
“We found that when people saw the AI recommendation first, they tended to become more deferential to it — a form of automation bias,” Xu said. “The effect was visible not only among laypeople but even among physicians, where LLM explanations presented first produced particularly strong deference.”
The takeaway isn’t that doctors are easily swayed. It’s that their ability to catch AI errors depends on thinking for themselves first. Because no algorithm is 100% accurate, the moment a tool reveals its answer in the clinical workflow can be the difference between an error caught and an error followed.
Designing for the handoff
For Xu, the solution is not to make AI less capable, but to fundamentally rethink its purpose in consumer hands. That shift starts with building systems comfortable with communicating their own limitations.
“A good health AI should sometimes be able to say, very clearly, ‘You should not rely heavily on this answer,'” Xu said. “Today, generative AI is largely optimized to always give us a coherent response. In medicine, that can be dangerous because the fluency of the answer can hide the uncertainty underneath it.”
Rather than delivering a confident-sounding verdict a patient carries into the exam room, Xu argues consumer health AI should be built to prepare patients for a conversation, not replace one.
“The ideal system should not send a patient into the clinic saying, ‘The AI told me I definitely have this,'” Xu said. “It should help them arrive saying, ‘Here is what I noticed, here is what I learned, here is what remains uncertain, and here are the questions I want to discuss with you.’ Designed that way, AI could become a bridge into the doctor–patient relationship rather than a competing authority—and potentially make the clinical conversation better informed than it was before.”
That reframing transforms the doctor’s visit back into a single collaborative task: instead of untangling a patient’s misplaced certainty before getting to the real diagnosis, the clinician can build upon what the patient and the AI explored together.
More Information
The article Divergent impacts of explainable AI for dermatological diagnosis on clinicians versus lay people was published Aug. 4, 2026, by Nature Medicine.
Xuhai ‘Orson’ Xu is the lead author. Will Ke Wang, Apoorva Mehta, Alejandro A. Gru, Noémie Elhadad & Lena Mamykina are also affiliated with Columbia University.
The full list of authors is as follows: Xuhai ‘Orson’ Xu, Haoyu Hu, Haoran Zhang, Will Ke Wang, Reina Wang, Luis R. Soenksen, Omar Badri, Sheharbano Jafry, Elise Burger, Lotanna Nwandu, Apoorva Mehta, Erik P. Duhaime, Asif Qasim, Hause Lin, Janis Karleen Pereira, Jonathan Hershon, Paulius Mui, Alejandro A. Gru, Noémie Elhadad, Lena Mamykina, Matthew Groh, Philipp Tschandl, Roxana Daneshjou, and Marzyeh Ghassemi.