FDA Wants to Evaluate AI Devices the Way It Evaluates Doctors
- Sharmila Bhatt
- 3 days ago
- 5 min read
Yesterday, FDA's Digital Health Center of Excellence published something it was careful to call a discussion paper rather than guidance: a 30-page document laying out considerations for regulating generative AI-enabled medical devices, opening a public comment docket that runs until October 19. The agency has been explicit that this isn't draft policy — it doesn't propose regulatory expectations, doesn't address whether new legal authority would be needed, and doesn't commit to a specific framework. What it does do is pose 26 detailed questions, with multiple sub-questions apiece, that map out exactly where FDA's current thinking is unsettled — and reading those questions closely tells you more about where this is headed than a premature guidance document would.
The timing matters. FDA has authorized more than 1,200 AI-enabled medical devices, but as CDRH Director Michelle Tarver noted at a Digital Health Advisory Committee meeting last November, none of those authorizations involve generative AI — the category of model that can produce novel, open-ended output rather than a bounded classification or prediction. That's the gap this paper is trying to address before it becomes a backlog of submissions the agency doesn't have a framework to evaluate consistently.
The idea worth paying attention to: competency assessment
The paper's most substantive proposal is a premarket evaluation model built around competency assessment—explicitly inspired, at a high level, by how physicians are trained and evaluated rather than how traditional medical devices are cleared. The approach combines non-clinical device benchmarking with clinical confirmation, aiming to establish whether a GenAI-enabled device performs as intended before it reaches patients, accounting for the device's ability to generate novel outputs rather than select from a fixed, pre-specified set of responses.
That's a genuinely different mental model than FDA has historically applied to software as a medical device. Traditional SaMD evaluation, even for adaptive AI/ML systems, has generally worked by establishing a bounded performance envelope and verifying the device stays inside it — the logic behind Predetermined Change Control Plans, which let a device evolve within pre-authorized limits. A generative model's core value proposition is often the opposite of a bounded, predictable output space; it's designed to handle novel inputs with novel outputs. Evaluating something for its judgment and adaptability, rather than its adherence to a fixed specification, is closer to how a licensing board evaluates a physician than how an inspector verifies a device meets its cleared specifications. The paper's own framing — a "nimble regulatory approach that employs least burdensome principles" — signals FDA is trying to avoid simply forcing generative AI devices into the PCCP framework built for a different kind of system.
The two-axis risk framework and what it implies about foundation models
The paper proposes evaluating risk along two axes rather than the single risk-classification scale that's governed device regulation for decades. The specifics of both axes weren't fully detailed in the initial coverage, but the shift to a multi-dimensional risk model is itself a signal: a single risk axis works reasonably well when risk is primarily a function of clinical consequence and diagnostic centrality — the logic behind Class I, II, and III device categories. It works much less well when a device's risk profile also depends on characteristics like model transparency, the degree of autonomous action it takes without a clinician in the loop, or how frequently its underlying foundation model gets updated by a third party the device manufacturer doesn't fully control. The paper explicitly raises foundation models and agentic AI systems as distinct considerations—devices built on a general-purpose model developed by someone other than the device sponsor, and devices that take autonomous action rather than simply generating a recommendation for a human to act on. Both categories complicate the traditional assumption that a single, identifiable sponsor controls and can fully characterize everything relevant to a device's behavior.
Postmarket monitoring gets equal billing, not an afterthought
Unlike some past FDA discussion documents on AI, this one gives real weight to postmarket monitoring rather than treating premarket clearance as the primary event and postmarket surveillance as a secondary compliance obligation. That's consistent with a separate, related request for comment FDA has open — on measuring and evaluating real-world AI-enabled device performance — which asks pointed questions about detecting performance drift, balancing human expert review against automated monitoring, and what data sources (electronic health records, device logs, patient-reported outcomes) are actually usable for ongoing evaluation. Read together, the two documents suggest FDA is trying to build premarket and postmarket frameworks for generative AI devices in parallel, rather than finalizing premarket expectations first and treating postmarket monitoring as a problem for later — a sequencing choice that matters given how much a generative model's real-world behavior can diverge from its premarket testing environment.
Practical steps worth taking during the comment window
For a device company already working on a generative AI product, waiting for a finalized framework before engaging with this discussion paper is likely a missed opportunity rather than a cautious default. A few concrete things worth doing before the October 19 deadline: map your device's current validation approach against the competency-assessment model described in the paper, and identify specifically where your existing evidence generation plan would and wouldn't satisfy a "clinical confirmation" standard modeled on physician evaluation rather than a fixed-specification standard. If your device relies on a foundation model developed by a third party, start documenting now how you monitor and characterize changes to that underlying model, since the paper explicitly raises this as an open question FDA hasn't resolved — a sponsor with a clear, defensible answer already in hand is in a stronger position than one improvising a response after a framework locks in around a different assumption. And if your device includes any agentic capability — taking action without a clinician actively reviewing and approving each step — treat that as the single highest-scrutiny element of your submission strategy going forward, since it's the characteristic the paper singles out most explicitly as breaking the traditional assumption that a device simply generates a recommendation for a human to act on.
What this means for device makers building in this space now
Nothing in this discussion paper is binding, and FDA has been unusually explicit that it isn't meant to signal what a future guidance document will say. But for any company building or planning a generative AI-enabled device, there's real value in engaging with the 26 questions directly rather than waiting for a more finalized framework. The competency-assessment model, if it becomes the agency's actual direction, would require device sponsors to think about validation evidence differently than they have for prior AI/ML-based SaMD submissions — less "does this system stay within its validated performance envelope" and more "can this system be trusted to exercise sound judgment across a class of situations it wasn't explicitly tested against," which is a harder, more expensive, and less familiar kind of evidence to generate. Sponsors with generative AI devices already in development have a real opportunity, through the comment docket, to shape which side of that evidentiary line the eventual framework lands on — and the docket closing October 19 gives a concrete, near-term deadline for engaging rather than treating this as a someday problem.
The broader signal for the device industry generally: FDA explicitly frames this effort as a potential model for regulators elsewhere, following a similar discussion paper the UK's MHRA released in June on AI in healthcare regulation. Whatever framework eventually emerges from this docket is likely to become a reference point well beyond U.S. borders, which raises the stakes of getting the underlying questions right during this comment period rather than after a framework is already locked in.
Sources: U.S. Food and Drug Administration, Considerations for the Regulation of Generative AI-Enabled Medical Devices (Discussion Paper), August 18, 2026; FDA Digital Health Advisory Committee Meeting Summary, November 6, 2025; FDA Request for Public Comment, Measuring and Evaluating Artificial Intelligence-Enabled Medical Device Performance in the Real-World; Axios; Medical Device Network; AuntMinnie.



Comments