On airBulletin №14 · AI in Crisis: The Unintended Consequences of Chatbots|Up next · AI Music Revolution: Artistry or Algorithm?
From the journal

Nobody Checked the Medical AI: The Clearance Doctors Assume Does Not Exist

Published
23 July 2026

It is a little after two in the morning, and somewhere in a hospital the pattern repeats. A resident stands at a workstation with a patient who does not fit the textbook: an unusual drug interaction, a rare presentation dressed up as a common one, a question the ward round never resolved. On the screen is a clinical decision-support tool, licensed by the hospital, embedded in the workflow, bearing the name of a publisher clinicians have trusted for a generation. In the resident's coat pocket is a general-purpose chatbot with no medical pedigree whatsoever. The resident types the question into both.

Which one is better, and does anyone actually know? The resident assumes the licensed tool is the safer bet: the hospital bought it, it lives inside the electronic record, somebody senior must have checked. That assumption is entirely reasonable. It is also, on the evidence, resting on nothing.

On 12 June 2026, Nature Medicine published online a head-to-head benchmarking study that put the assumption to the test. Researchers at NYU Langone Health and the University of Texas at Austin took three general-purpose frontier chatbots that no medical regulator has evaluated for clinical decision support: OpenAI's GPT-5.2, Google's Gemini 3.1 Pro, and Anthropic's Claude Opus 4.6. They pitted them against two specialist clinical decision-support tools marketed for exactly this job: OpenEvidence and Wolters Kluwer's UpToDate Expert AI. Then they measured all five against real, unstructured questions physicians had actually asked, alongside two established public benchmarks. The frontier models outperformed the specialist tools in all three evaluations. On the messiest and most realistic measure, the specialist tools performed no better than the artificial-intelligence summary that appears at the top of an ordinary Google search.

Read that again slowly, because the discomfort is in a detail most commentary missed. The obvious story is that regulated products lost to unregulated ones. That story is wrong, and the truth beneath it is worse. Neither OpenEvidence nor UpToDate Expert AI holds clearance from the United States Food and Drug Administration. Search the agency's public 510(k) database for either company, or for Wolters Kluwer, and it returns nothing. Both operate outside the medical-device framework entirely, under a statutory carve-out requiring no premarket evaluation of any kind. So do the chatbots. The study did not compare a vetted tool against an unvetted one. It compared two unvetted tools against three others, and the ones with the medical branding lost.

The stamp everyone assumed was there is not there. And the absence is invisible to precisely the people staking clinical decisions on it.

What the Study Actually Measured

The design matters, because the finding is easy to caricature. This was not a stunt in which someone asked a chatbot to diagnose a celebrity. It was a structured, three-part evaluation, with twelve clinicians performing blinded review that generated some 1,800 annotations.

The first component was MedQA, a benchmark drawn from medical licensing examinations, testing whether a system has internalised the formal knowledge base of medicine. Across five hundred questions, Gemini 3.1 Pro scored 97.4 per cent, GPT-5.2 94.2 per cent, Claude Opus 4.6 90.2 per cent. OpenEvidence reached 89.6 per cent, UpToDate Expert AI 88.4 per cent. Not humiliating, but the wrong side of a line nobody expected them to be on.

The second was HealthBench, an open benchmark released by OpenAI in 2025 and built with 262 physicians from 60 countries writing rubric criteria across dozens of specialities. It does not test trivia. It scores multi-turn conversations against fine-grained, physician-authored standards for accuracy, completeness, communication, hedging, and the crucial clinical skill of asking for missing context rather than charging ahead. Its hardest tier is punishing enough that, at release, no model scored above roughly a third of the points. Across five hundred items on a hundred-point scale, GPT-5.2 scored 88, Gemini 3.1 Pro 79.3, Claude Opus 4.6 77. OpenEvidence managed 62.6 and UpToDate Expert AI 61.3. Here the gap is a chasm of some twenty-five points, on the benchmark designed by physicians to reward the behaviours physicians value.

The third component gives the study its teeth: a real clinical queries benchmark of one hundred de-identified questions that practising physicians had submitted, in their own words, through NYU Langone Health's HIPAA-compliant instance of a general-purpose model, in live clinical use. Not tidied-up vignettes. The actual half-formed, context-poor, urgent questions clinicians type when a patient is in front of them. This is the terrain on which decision-support tools live or die, and precisely the terrain structured benchmarks avoid because it is so hard to score.

Into that evaluation the researchers inserted a control that reads, in retrospect, like a provocation: Google Search's AI Overview, the summary that materialises above the blue links. It was included because it is genuinely part of the clinical information environment; clinicians encounter it constantly. On the real clinical queries, the specialist tools performed comparably to it. The purpose-built, subscription-funded, health-system-deployed products were level with a consumer search feature nobody markets as a clinical instrument at all.

This is not a permanent verdict on any product; a benchmark is a snapshot, not a prophecy. The durable finding is structural. The institutional standing of a clinical AI tool, the trusted publisher, the enterprise contract, the integration into the record, told you nothing about which system would answer a real physician's real question better. Authority and performance had come apart, and no mechanism existed to notice.

The Route That Requires No Test

Here is where the story turns, and where much coverage went astray. Several accounts described the losing products as “FDA-cleared clinical AI,” framing the result as regulated technology beaten by unregulated technology. It is an appealing narrative. It is also false, and the falsehood conceals something worse.

Neither product went through the FDA. Neither had to. Section 3060 of the 2016 21st Century Cures Act amended the Federal Food, Drug, and Cosmetic Act by inserting section 520(o), listing categories of software function excluded from the definition of a medical device altogether. Subparagraph (E) is the clinical decision support carve-out. Software that displays or analyses medical information, supports a health-care professional's recommendation about prevention, diagnosis, or treatment, and is designed so the professional can independently review the basis for that recommendation rather than relying primarily on it, is not a device. Not a low-risk device. Not an exempt device. Not a device at all.

This is not a loophole in the pejorative sense. It was a deliberate judgement that software functioning as a reference aid for a trained clinician, who remains the decision-maker, does not warrant device regulation. The logic is coherent. The consequence is that the tools sitting under it are engineered, explicitly, to stay there.

OpenEvidence positions itself accordingly: as evidence retrieval and clinical knowledge synthesis, an informational tool surfacing and summarising the peer-reviewed literature so a physician can evaluate it. Every element of that positioning maps onto the criteria for staying outside the device definition. UpToDate Expert AI, launched in September 2025 and now deployed at more than fifty major United States health systems, occupies the same ground for the same reasons.

So the comparison the study ran was not regulated versus unregulated. It was two products that took the non-device route against three that were never candidates for the device route at all. All five occupy substantially the same regulatory position: nobody evaluated any of them before they reached a clinician. The specialist tools did not fail a weak test. They took a path requiring no test.

What Clearance Would Have Certified

To see the shape of the gap, look at the regime these products did not travel through. It is what people picture when they imagine a vetted medical AI product, and even it would not have answered the resident's question.

Most software-based medical devices reach the American market through 510(k) premarket notification. The mechanism at its heart is not a test of excellence but a test of resemblance. To obtain clearance, a manufacturer need not prove its device is good, or effective, or better than anything. It must demonstrate “substantial equivalence” to a device already legally marketed, known as a predicate. As the FDA's own guidance makes explicit, substantial equivalence supports an assessment that a device introduces no new hazards and works at least as well as its predicate. It does not require proof of superiority. It was never meant to.

This is why the vocabulary matters. The FDA reserves “approved” for a far more demanding process: premarket approval, or PMA, applied to the highest-risk devices, where a manufacturer must submit clinical evidence of safety and effectiveness in its own right. A third route, De Novo, exists for genuinely novel low-to-moderate-risk devices with no predicate to point at. But roughly nine in ten devices subject to premarket review come through 510(k). They are cleared, not approved.

The distinction is not pedantry. A federal courtroom has had to rule on whether 510(k) clearance speaks to safety and effectiveness at all, precisely because the pathway is comparative rather than absolute. What clearance certifies is a relationship: this device is like an existing device once judged acceptable, which was like an earlier device, in a chain stretching back decades. Critics have a name for what happens along that chain, predicate creep, in which each device is equivalent to the last by a small margin until the margins accumulate and a modern product is justified by resemblance to an ancestor it barely resembles. Reviews of the programme, including the Institute of Medicine's, have warned for years that substantial equivalence is too easily equated, by clinicians and courts and the public, with a positive judgement of safety and effectiveness.

Even had these two products gone through 510(k), then, clearance would have certified only that each resembled some prior device closely enough. At no link does anyone ask what the sleepless resident cares about: is this the best available tool, or merely an adequate one? The stricter regime contains no comparative-effectiveness requirement either, only a resemblance requirement. The gap the study exposed is not one clearance would have closed. It runs through both sides of the line.

The Specification a Manufacturer Writes for Itself

Artificial intelligence complicates the regulated picture in a way the FDA has genuinely tried to address. Traditional devices are static. A machine-learning system is not; the thing cleared on Monday may behave differently by Friday, and freezing it defeats the purpose.

The answer arrived on 3 December 2024, in a guidance titled “Marketing Submission Recommendations for a Predetermined Change Control Plan for Artificial Intelligence-Enabled Device Software Functions.” The Predetermined Change Control Plan, or PCCP, is a clever instrument. It lets a manufacturer specify in advance the modifications it intends to make, the protocol by which it will validate them, and an assessment of how those changes might affect safety and performance. Once the FDA authorises the plan, the manufacturer can implement the pre-agreed changes without a fresh submission each time. The final guidance broadened the concept to all AI-enabled devices and required labelling to disclose when a device had been authorised with such a plan.

But notice what the PCCP measures against. It measures a device against the specification its own manufacturer defined. The modification protocol asks whether changes stay within the envelope the sponsor drew; the impact assessment asks whether the sponsor's own performance claims still hold after the sponsor's own planned updates. The apparatus is internally referential: a rigorous test of whether a product does what its maker promised, held to the standard its maker set. That is not a flaw. It is the PCCP working as intended.

What the plan cannot do, because it was never built to, is look sideways. No clause asks whether a freely available chatbot, updated last week by a company that filed nothing with the FDA, now answers clinical questions more accurately than the authorised device. The regulatory frame has no peripheral vision. It looks at the device and its own past self, not at the field. And this is the best case, inside the system. The two tools the study measured are outside it, where there is no specification and no plan at all.

The Safe Harbour Widens

If the timing were fiction, an editor would strike it as too neat. In the same year researchers were demonstrating that specialist clinical AI underperforms general-purpose models on physicians' real questions, the FDA was making the non-device carve-out those tools rely on broader and easier to sit inside.

The agency issued a revised Clinical Decision Support Software guidance on 6 January 2026 and published it in final form on 11 March 2026, superseding the September 2022 version. The new document takes a more explicitly risk-based approach and expands the circumstances in which software qualifies as non-device CDS. The most consequential change concerns the criterion that had caused developers most difficulty: that the software not provide a specific preventive, diagnostic, or treatment output or directive.

Under the 2022 guidance, the FDA held that software offering a single recommendation, rather than options for the clinician to weigh, inherently supplants clinical judgement and therefore falls inside the device definition. The 2026 guidance reverses that stance. Where only one clinically appropriate recommendation exists, and the software otherwise satisfies the statutory criteria, the agency has said it does not intend to enforce device requirements against it. Illustrative examples indicate that software generating a recommended treatment plan, or producing a differential diagnosis for a practitioner to review and finalise, may fall outside device regulation. Software analysing a medical image or physiological signal to generate a diagnostic recommendation remains a regulated device, and the discretion is not blanket: a tool predicting a time-critical event, a cardiovascular collapse within twenty-four hours, say, stays squarely within oversight. Classification remains highly fact-specific; small changes to design, labelling, or workflow can move a product across the line either way. But the direction of travel in 2026 was unmistakable, and it was outward. The regulatory bar described it, approvingly, as cutting red tape.

Set the two events side by side and the irony is almost unbearable. In January and March, the safe harbour sheltering generative clinical decision support from premarket review got wider. In June, the first rigorous head-to-head evidence arrived showing that the flagship products sheltering there were being outperformed, on physicians' actual questions, by consumer chatbots and a search box. The system did not fail to act on the evidence. It moved, in good faith and for defensible reasons of proportionality, in the opposite direction, months before the evidence existed. What neither the carve-out nor its expansion engages with is whether the tools inside the harbour are any good.

The Counterfactual Nobody Runs

This is the gap the study drove a lorry through. Every decision in this space, whether a clearance, a change-control plan, or a determination that a product is not a device at all, answers a question of the form “does this thing meet the criteria for its category?” None answers “compared to what?” The counterfactual is absent from the assessment, and for non-device tools there is no assessment for it to be absent from.

In some corners of medicine this would be unthinkable. A pharmaceutical company seeking approval for a new drug is increasingly expected, by payers if not always by the FDA, to show not merely that the drug beats a placebo but how it stacks up against existing therapy: comparative effectiveness research. Software has escaped this comprehensively. A cleared clinical AI device must be as good as its predicate. A non-device clinical AI tool must be nothing at all, in performance terms, to anyone. Neither must ever be as good as the best thing a doctor could otherwise reach for, and in the age of the general-purpose chatbot that alternative is extraordinary, improving fast, and free.

The result is a category error hiding inside institutional furniture. When a hospital procurement committee evaluates a generative decision-support product from a long-established medical publisher, already deployed at fifty peer institutions, integrated with the electronic record and sold with an enterprise contract and a security review, it reads all of that as evidence of vetting. And it is evidence of something: of commercial diligence, of information-governance compliance. It is not evidence that anyone measured whether the tool answers clinical questions better than the alternatives. The authority is institutional and reputational, not regulatory, and certainly not empirical.

The committee is not being deceived. Nobody told it the product was cleared. But nobody told it the product was not, either, and no label, no field on a procurement form, would prompt the question. Absence does not announce itself. A missing clearance looks exactly like a clearance nobody happened to mention. The researchers put the alternatives in the same room; that is the whole contribution, and it is why the study reads like an exposed foundation.

The Doctors Are Already in the Room

If this were a theoretical gap, it could wait. It is not, because physicians have already voted with their keyboards.

The American Medical Association surveys physician sentiment towards what it carefully calls augmented intelligence. Its 2026 edition, published in March and drawn from 1,692 physicians surveyed in the opening weeks of the year, found that more than four in five, 81 per cent, now use AI in practice. That is more than double the 38 per cent recorded in 2023, and it continues a trajectory the 2024 survey caught mid-flight, when the figure stood at 66 per cent. The average number of distinct use cases per physician has risen from 1.1 to 2.3. Confidence has climbed alongside adoption: more than three-quarters now believe AI improves their ability to care for patients, up from 65 per cent in 2023. And the most common uses reported are summarising medical research and supporting clinical documentation, which is to say: exactly the task the benchmark study measured.

Here the two findings lock together into something sharper than either alone. Four in five physicians are already using AI, most commonly to digest the literature. The tools sold to them for that purpose carry no evidence of comparative superiority and, on the best available evidence, have none. A large and growing share of clinical AI use is therefore happening with no reliable relationship between the standing of a tool and its performance, and with no way for the clinician to know which system on their screen is better for the question in front of them. The doctor reaching for the licensed product because the hospital bought it, and the doctor reaching for the free chatbot because it is faster, are both flying blind with respect to the only thing that matters at the bedside: which answer is more likely to be right.

Automation bias sharpens the edge. A substantial literature documents the human tendency to defer to a confident machine, and a tool wrapped in the authority of a trusted publisher is exactly that. The branding does not merely fail to guarantee superior performance. It borrows trust it has not earned, shaping which answer the clinician believes rather than which is correct.

Why the Specialist Tools Might Lose

It is tempting to assume the specialists lost because they are cruder than the frontier chatbots, but the likely reasons cut against the intuition that a purpose-built medical tool should beat a generalist.

Both products are, at bottom, generative systems layered on curated medical knowledge. OpenEvidence markets itself on grounding answers in a vast corpus of peer-reviewed literature; UpToDate Expert AI inherits the authority of a resource clinicians have trusted for decades. Their pitch is precisely their specialism: constrained, cited, grounded in vetted sources, less prone to invention than a free-roaming chatbot. That specialism is real. But it can cut against raw performance. A system constrained to a particular corpus, tuned for caution, and updated on the slower clock of an established product may reason less fluently across the cross-domain messiness of a real clinical question than a frontier model trained on a far larger slice of the world.

There is a second constraint, and the regulatory story illuminates it. Staying outside the device definition is not free; it shapes the product. A tool engineered to remain a non-device must present information the clinician can independently review, must avoid crossing into directive output, and must keep its framing on the safe side of a line whose exact position is, as the FDA keeps saying, highly fact-specific. Caution of that kind is legally load-bearing. It is also, on a benchmark rewarding decisive, complete, well-calibrated clinical reasoning, expensive. The general-purpose models labour under no such constraint, because nobody has ever suggested they are medical products, which is exactly why nobody has asked them to behave like one.

There is also a brutal asymmetry of cadence. The frontier laboratories spend sums that dwarf the budgets of clinical-content companies and ship flagship models on a cycle measured in months. UpToDate Expert AI launched in September 2025; by June 2026 it faced models that did not exist when it was designed.

None of this makes the specialist tools worthless. Grounding answers in cited, vetted literature is a genuine safety property, and a fluent model that invents a plausible-sounding reference is dangerous in a way a constrained tool is not. But specialism plus institutional authority does not add up to superiority. The tool may be worth choosing for its citations, its integration, its accountable vendor. It is not worth choosing because it feels official, because that feeling corresponds to no evaluation anyone ran.

The Precedents That Should Have Warned Us

Medicine has been here before, in forms that ought to have inoculated the field against confusing standing with real-world superiority.

The pulse oximeter is the cleanest example, and instructive precisely because it was, unlike the tools in this study, a genuinely cleared device. For decades a ubiquitous and trusted instrument, the finger-clip oximeter was shown, in research by Sjoding and colleagues published in the New England Journal of Medicine during the COVID-19 pandemic, to systematically overestimate blood-oxygen levels in patients with darker skin, missing dangerous occult hypoxaemia roughly three times as often in Black patients as in white ones. The device met its specification. It had been cleared. It was used everywhere. And it was quietly failing a whole class of patients, because its clearance had never required anyone to test what actually mattered across the population that actually used it. The FDA updated its testing guidance only after the harm became impossible to ignore.

The lesson generalises in both directions. A clearance certifies conformance to a specification, and if the specification omits the question you care about, the clearance is silent on it. That is the oximeter. The harsher corollary for the AI case: if there is no specification at all, there is nothing even to be silent about. The oximeter at least had a document somebody could later go back and criticise. For a non-device clinical decision-support tool there is a marketing page and a licence agreement.

What is genuinely new is the availability of the alternative. In the oximeter case there was no free, superior device in every clinician's pocket; the harm was a gap between the tool and reality. Here the superior alternative is neither hypothetical nor scarce. It is the chatbot the resident already has open in another tab, updated last week, costing nothing, evaluated by no one, and, on the evidence, better.

What Authority Communicates, and to Whom

All of which forces the sleepless resident's question into the open. If nobody has measured whether a marketed clinical AI tool performs better than the free alternative already in widespread use, what is the clinician actually relying on, and should they be told plainly what has and has not been checked?

The honest answer to the first is that they are relying on a bundle of proxies, each real and none of them performance. A trusted publisher's reputation, built over decades on a product that was not this product. A procurement decision that assessed cost, security, and integration. Fifty peer institutions having made the same decision, which is informative mainly about the sales team. An interface that cites sources, which establishes that sources exist, not that the synthesis of them is sound. These proxies are not worthless; an accountable vendor is genuinely better than a chatbot answering to nobody. But every one is a statement about the circumstances surrounding the tool, not about its rank among the options a clinician could actually choose.

The honest answer to the second is yes, and the case for candour is strengthened by the FDA having already conceded the principle in an adjacent context. The December 2024 PCCP guidance added a labelling requirement precisely so users would be told when a device had been authorised with a change-control plan, on the reasoning that a clinician deserves to know the epistemic status of the tool in their hands. The logic extends further. If a user deserves to know a device may change under a pre-agreed plan, they surely deserve to know when a tool in their workflow underwent no premarket review whatsoever, and that a freely available alternative may match or beat it. The infrastructure for honest disclosure exists. What is missing is any obligation to use it outside the device framework, where it would matter most.

None of this argues against the safe harbour. But the frontier chatbots are magnificent on a benchmark and entirely ungoverned in a hospital: no manufacturer accountable for a clinical failure, no labelling, no change control, no disclosure when a model is silently updated in a way that degrades its medical reasoning, no obligation to ground an answer in a real citation rather than a fluent invention. The lesson is not that regulation failed. It is that on the question clinicians most need answered, which tool is better, the regulatory system was never asked to speak, and its silence is being mistaken for a verdict.

Closing the Peripheral-Vision Gap

The fix that suggests itself is not to drag every chatbot through a 510(k). Forcing every clinical reference tool through premarket review would be futile against models that update monthly, and would freeze the field in amber. It is to attach a comparative-evidence obligation to the act of marketing a tool for clinical use, wherever that tool sits relative to the device line. A product sold to health systems as clinical decision support could be required to publish its performance against the relevant general-purpose baseline on standardised, physician-authored benchmarks of exactly the kind this study used, and to keep publishing as both the tool and the baseline evolve. This is not a novel imposition so much as an import: comparative-effectiveness thinking, long established for drugs, brought at last to software that reasons about drugs. The question shifts from “does this thing fall inside or outside a statutory category?” to “does it earn its place against the free alternative a clinician would otherwise use?”

Such a regime would have teeth precisely because the benchmarks exist and the comparison is cheap to run. HealthBench and its kin demonstrate that physician-graded evaluation is feasible at scale, and two university health systems with twelve blinded reviewers managed the whole thing at a cost that would not register on a procurement budget. The obstacle is that comparative testing would occasionally return the answer no vendor wants, that its marketed product is being beaten by a free tool, and that a hospital paying the licence might reasonably ask why. That is not a reason to avoid the test. It is the reason to run it.

Until something like that exists, the burden falls where it always falls when a system withholds a truth it could easily tell: on the individual, exercising judgement in the dark. The resident at two in the morning, typing the same question into the licensed tool and the free one, is running a private, unrecorded version of the Nature Medicine study on every difficult case, and drawing conclusions no regulator will ever see. The least the system owes them, and owes the patient whose care hangs on which answer the resident believes, is honesty about what has and has not been examined. It is not that the stamp certified too little. It is that there is no stamp, and the absence looks, from where the resident is standing, exactly like its presence.

That is the gap: not the failure of any one company or regulator, but the distance between an evaluation that never happened and the confidence a frightened, time-poor human being reads into a familiar name at the bedside. Closing it does not require dismantling the safe harbour. It requires the people selling into it to say out loud what they have measured, and against what, before the next question is typed and the next answer trusted.

Previous

Facial Recognition and Wrongful Arrest

AI in Crisis: The Unintended Consequences of Chatbots
Bulletin №14