A radiologist in Sweden opens a mammogram and finds that the software has already flagged the spot she was about to look at twice. Somewhere else the same evening, a man with a tight chest types his symptoms into a chatbot at one in the morning and gets told to rest and see how he feels.
Both of those things are happening right now. One of them is backed by a randomized controlled trial with tens of thousands of women in it. The other one sits at the top of a national list of the most dangerous technology in medicine.
That gap is the honest story of AI in healthcare. It isn’t one technology with one verdict. Some of it has cleared the hardest bar medicine has. Some of it is a slide deck with a login screen. Plenty of it sits in between, quietly useful in one hospital and quietly useless in the next.
So this article does the sorting. What’s working, with published evidence behind it. What’s failing, sometimes badly enough to hurt people. And what you can reasonably expect the next time you’re sitting in a waiting room.
Three very different things get called the same name
Most of the confusion comes from one word doing three jobs at once. Separate them and the picture gets a lot clearer.
Regulated medical devices. Software that the FDA has reviewed and authorized for a specific clinical purpose. As of the agency’s early 2026 update, roughly 1,500 AI enabled devices have been cleared or approved, and about three quarters of them are in radiology. These are narrow tools. One reads a chest scan. One measures a heart chamber. None of them chat with you.
Administrative software. The scribes, schedulers, billing engines, and inbox helpers. This layer touches your care but is usually not regulated as a medical device, because it isn’t officially diagnosing anything. It’s also where hospitals are spending most of their money.
General chatbots. The ones you already use. Not built for medicine, not validated for medicine, not regulated as a medical device, and used for medical questions by an enormous number of people anyway.
When a headline says AI beat doctors, it’s almost always talking about the first bucket. When something goes wrong in the news, it’s usually the third. Keep the three separate and most of the hype evaporates on its own.

Where AI in healthcare is genuinely working
These are the areas where the evidence is real, published, and not funded entirely by the company selling the product.
Reading mammograms
The MASAI trial in Sweden is the strongest evidence anyone has produced so far. It’s the first randomized controlled trial of AI in breast cancer screening, and it’s the largest randomized study of AI in cancer screening of any kind.
AI supported reading found 6.1 cancers per 1,000 women screened. Standard double reading by two radiologists found 5.1 per 1,000. That’s about 29% more cancers found, with no increase in false positives. Recall rates were 2.2% with AI and 2.0% without, and the false positive rate was 1.5% in both groups.
The extra cancers weren’t trivial finds either. They skewed toward small, lymph node negative invasive tumors, including more of the aggressive subtypes that tend to show up between screenings.
There’s a second win buried in that trial. The screen reading workload dropped by roughly 44%. In a system where radiologists are scarce and screening volume keeps climbing, that number may end up mattering as much as the detection rate.
Writing the notes
Ambient scribes listen to your appointment and draft the clinical note, so your doctor can look at you instead of a keyboard. This is the least glamorous use of AI in healthcare and probably the most useful one today.
A quality improvement study published in JAMA Network Open followed 263 physicians and advanced practice practitioners across six health systems. After 30 days with an ambient scribe, burnout among the ambulatory clinicians fell from 51.9% to 38.8%. Other real world studies found clinicians spent about 8.5% less total time in the medical record and over 15% less time composing notes.
Two honest caveats. Appointment length didn’t shrink, so the time comes back to the clinician, not to your visit. And the results aren’t universal. At least one study found the documentation time dropped with no measurable change in burnout at all.
Screening eyes without an eye doctor
Autonomous AI for diabetic retinopathy was the first system cleared to give a clinical result without a specialist reviewing the image. A camera in a primary care clinic takes a picture of the retina and the software says whether you need to see an ophthalmologist.
In real world use it behaves in a specific and useful way. A Swiss study of 1,141 patients found the system consistently overestimated disease severity, which means a negative result is highly trustworthy while a positive result needs a human look. That’s an acceptable trade for a screening tool.
The catch is boring and physical. In one 2026 study of 875 patients, 26.1% of the images simply couldn’t be analyzed, thanks to pupil size, patient age, and who was operating the camera. Great software still loses to a bad photo.

What hospitals are actually buying
Adoption has moved fast, and faster than in most of the wider economy. A February 2026 industry survey found about 75% of health systems are using or planning to use at least one AI application, and half now run three or more.
The leading use case isn’t diagnosis. It’s clinical notetaking and ambient listening, at roughly 68% adoption. Documentation improvement sits second. In other words, the money is going into paperwork, not into the dramatic stuff.
More than half of the systems that could actually measure return on investment reported at least double their money back. That “could actually measure” qualifier is doing heavy lifting, and it’s the sort of detail vendor case studies tend to leave out.

Where it fails, and the failure nobody advertises
The clearest cautionary tale in this field is a sepsis prediction model built into one of the most widely used hospital record systems in the country. Sepsis kills fast, so early warning is exactly the kind of problem AI should be good at.
Then researchers at Michigan Medicine tested it on their own patients and published the results in JAMA Internal Medicine. The model missed roughly two thirds of the sepsis cases it was supposed to catch. A later external validation across two county emergency departments was worse still, reporting sensitivity of 14.7% and a positive predictive value of 7.6% for sepsis within six hours of the alert.
A tool that misses most of the cases and cries wolf on the ones it flags doesn’t just fail quietly. It trains nurses and doctors to ignore alerts, which makes the next alert worth less than nothing.
The lesson generalizes far beyond sepsis. A model that performs beautifully in the hospital where it was trained can fall apart in yours, because your patients, your coding habits, and your workflows are different. Anyone buying clinical AI should be asking for local validation, not a national brochure.

The chatbot problem is the biggest one
ECRI, the nonprofit that publishes an annual ranking of the biggest health technology hazards, put the misuse of AI chatbots at number one on its 2026 list. Ahead of every device, alarm, and infusion pump.
The reasoning is simple. These tools sound like experts, aren’t regulated as medical devices, haven’t been validated for clinical use, and are now consulted by patients, clinicians, and hospital staff every single day. ECRI notes that more than 40 million people a day turn to one chatbot alone for health information.
The audits back up the worry. A 2026 study in BMJ Open reviewed 250 health responses across five chatbots and rated 49.6% of them problematic, with 19.6% considered highly problematic or potentially harmful. Separate research has found high rates of confidently invented detail in medical answers, including citations that don’t support the claim they’re attached to.
And people aren’t just double checking their doctors with these tools. Survey work in 2026 found roughly one in seven people had used AI instead of seeing a health provider at all. That’s the scenario the safety people are actually worried about.
None of this makes chatbots useless for health. They’re genuinely good at translating jargon, helping you write down questions before an appointment, and explaining what a term on your test result means. They’re bad at deciding whether your symptom is an emergency, because they can’t see you, don’t know your history, and have no idea what medications you take.
The part of AI in healthcare with money on the other side
Not every algorithm in your care is trying to help you get better. Some of them sit at the insurer.
A class action against UnitedHealth Group, originally filed in 2023 by the families of two deceased Medicare Advantage members, alleges that an algorithm used to review post acute rehabilitation coverage cut care short and sometimes overrode the judgment of treating physicians. In early 2025 a federal court declined to dismiss the breach of contract and good faith claims, letting the case move forward, and a 2026 discovery order required the company to hand over a broad set of documents.
These are allegations, not findings, and the case is still in progress. But it’s a useful reminder that the same word covers a tool that finds your tumor and a tool that denies your rehab. When someone tells you a hospital or insurer is “using AI,” the only sensible next question is what for.
Adoption is rising while trust is falling
Here’s the tension that will shape the next few years. Hospitals are buying more AI every quarter. The public is getting less comfortable with it.
A 2026 national survey found only 42% of Americans were open to AI being used as part of their care, down from 52% in 2024. Roughly 72% said they were uncomfortable with AI systems accessing large amounts of their personal data.
The same people are using it anyway. About 44% reported using AI to help explain a test result or a diagnosis. That’s not hypocrisy. It’s what happens when a tool is genuinely helpful for one job and genuinely unsettling in another.
What this actually means for you
You don’t need to have an opinion about the industry. You do need a few habits.
Ask what the AI did, not whether it was used. “Did software help read this scan” and “did software decide my coverage” are completely different conversations, and both are fair questions to ask out loud.
Treat a chatbot as a translator, never as triage. Ask it what a word means or what to ask your doctor. Don’t ask it whether your chest pain can wait until Monday.
Take a second opinion seriously when a tool flags something. Screening AI is often tuned to overcall, which is the right design. A flag is a reason to look closer, not a diagnosis.
And if a coverage decision goes against you, appeal it. In the case above, the plaintiffs allege that the overwhelming majority of appealed denials were eventually reversed. Whatever the courts conclude about that specific number, appealing is free and algorithms are not infallible.
Key takeaways
- AI in healthcare covers three separate things: regulated devices, back office software, and unregulated chatbots. Only the first has been reviewed for a clinical purpose.
- The strongest evidence is in imaging. The MASAI trial found about 29% more breast cancers with AI support and no rise in false positives, while cutting screen reading workload by roughly 44%.
- Ambient scribes are the quiet win. Burnout dropped from 51.9% to 38.8% in one multisite study, though results vary and your appointment doesn’t get longer.
- Failures are real. A widely deployed sepsis model missed about two thirds of cases in an outside test, which is why local validation matters more than a vendor demo.
- Chatbots top ECRI’s 2026 list of health technology hazards. Around half of audited chatbot health answers were rated problematic.
- Hospital adoption is near 75% while public comfort has fallen to 42%. Expect that gap to drive the regulation coming next.

Frequently asked questions
Is AI actually diagnosing patients yet?
In a few narrow places, yes. Autonomous diabetic retinopathy screening delivers a result without a specialist reviewing the image. Almost everywhere else, AI produces a flag, a measurement, or a draft, and a clinician makes the call.
Can I trust ChatGPT with a medical question?
For understanding words, preparing questions, and general background, it’s useful. For deciding whether something is urgent, it isn’t. Audits keep finding a meaningful share of health answers that are wrong or potentially harmful, and safety groups now rank chatbot misuse as the top health technology hazard.
Does my doctor have to tell me AI was used?
Rules vary by state and by country, and they’re changing quickly. Many practices ask for your consent before an ambient scribe records a visit. You can always ask, and you can decline the recording.
Will AI replace doctors?
Nothing in the current evidence points that way. The clearest wins are AI as a second reader on images and AI as a typist for notes. Both keep a human accountable for the decision, which is also what regulators and insurers currently require.
How do I know if a health AI product is legitimate?
Look for FDA authorization for the specific claim being made, published results from somewhere other than the vendor, and evidence that it was tested on patients like the ones it will be used on. If a company can’t produce all three, treat the claims as marketing.
This article is for general information only and is not medical, legal, or financial advice. Talk to a qualified clinician about your own care, and don’t use a chatbot to decide whether a symptom is an emergency. If you think you’re having one, call your local emergency number.