FDA Releases Discussion Paper on Generative AI-Enabled Medical Devices: What Manufacturers and Investors Need to Know

08.26.2026

On August 18, 2026, the FDA's Digital Health Center of Excellence (DHCoE), part of the Center for Devices and Radiological Health (CDRH), released Considerations for the Regulation of Generative AI-Enabled Medical Devices, a discussion paper and request for feedback under docket FDA-2026-N-7874. Comments are due October 19, 2026.

This is not a draft or final guidance. It proposes no binding policy changes and does not address whether the approaches discussed fall within FDA's existing legal authorities. But make no mistake—the paper's 26 numbered discussion questions are a roadmap for where CDRH intends to go. Companies building AI-enabled products that incorporate generative AI, large language models, or foundation models should pay close attention. The feedback FDA collects here will shape the evidence standards that reviewers apply to future submissions.

Why This Paper Matters: A Departure from the Existing AI/ML Framework

FDA has been steadily building its regulatory framework for AI/ML-enabled medical devices since the 2021 AI/ML SaMD Action Plan, followed by the Good Machine Learning Practice Guiding Principles (October 2021), the Predetermined Change Control Plan (PCCP) Final Guidance (December 2024), Transparency Guiding Principles (June 2024), and the Draft AI-Enabled SaMD Lifecycle Guidance (January 2025). That existing framework, however, was designed for traditional AI/ML devices—products with defined inputs, deterministic or statistically bounded outputs, and algorithms that learn from structured training data. Generative AI-enabled devices are fundamentally different. They produce variable, open-ended outputs; they may evolve through interactions with users; and they can operate with degrees of autonomy that the current framework was not built to address. CDRH acknowledges as much in the paper, noting that new methodologies may need to be developed to evaluate the performance of these products.

This discussion paper represents FDA's first dedicated effort to develop a regulatory approach specifically for generative AI-enabled medical devices—a category that, as of today, has no established pathway to market.

Key Developments

1. A New Two-Axis Risk Framework

The paper proposes a two-axis risk framework to calibrate regulatory expectations for generative AI-enabled devices. One axis scores how independently a device function acts—ranging from providing non-directive information, to action-directing information, to supervised action, to fully autonomous action. The other axis scores the severity of consequences if the device produces an incorrect output, ranging from limited harm to severe.

This is a significant conceptual advance. Evidence expectations scale with the device's position on the grid, not simply with the presence of an LLM or generative model in the architecture. Importantly, FDA judges directiveness on substance: adding a disclaimer like "talk to your doctor" does not make an action-directing output less directive. Patient-facing functions may also move up the consequences axis because patients lack the clinical training to independently verify the output.

For manufacturers, this means your intended-use statement now must articulate both how independently the function acts and how consequential an incorrect output would be—and you must defend both placements in your risk file.

2. A Competency-Based Approach to Premarket Evaluation

Perhaps the most novel element of the paper is CDRH's proposal to evaluate generative AI-enabled devices the way medicine evaluates physicians: through standardized benchmarking, clinical confirmation, and ongoing assessment. CDRH cites academic literature proposing competency-based regulation for AI.

The approach has two components:

  • Device Benchmarking. High-throughput, non-clinical testing of the device in its deployed configuration across up to ten elements spanning safety behaviors, clinical proficiency, generalizability, and agentic conduct. This includes testing for adversarial prompting and prompt injection, scope maintenance and boundary adherence, calibration and clinical deferral, robustness across demographics and dialects, and—for agentic systems—planning within a safety envelope and recognition of tool errors. FDA is considering shared, standardized benchmarks rather than bespoke test sets per sponsor, which would give reviewers comparable scores across submissions.
  • Clinical Confirmation. A range of approaches to generate evidence of real-world performance, with explicit relief that "clinical confirmation might not require a prospective clinical study in every case." FDA outlines five approaches in escalating rigor, from retrospective evaluation to prospective studies, with the appropriate level driven by the device's position on the risk grid.

For devices generating open-ended outputs with no single "correct" answer, FDA suggests the comparator could be "a panel of qualified clinicians whose consensus reflects the applicable standard of care, or a median clinician in practice." FDA also signals that an LLM itself may serve as an expert adjudicator for benchmarking if it meets independence and qualification standards.

3. Expanded Postmarket Monitoring Expectations

CDRH signals that it may require a greater reliance on postmarket monitoring to ensure the ongoing safety of generative AI-enabled devices, in part because the variable nature of these outputs limits the degree to which premarket evaluation alone can provide a comprehensive assessment. The premarket competency-based assessment would become the baseline, and a modified device would be "re-benchmarked against the same capabilities."

PCCPs—the mechanism finalized in December 2024 for managing iterative AI/ML software updates—would extend to generative AI-enabled devices, though with unique challenges, particularly for updates initiated not by the device manufacturer but by a third-party foundation model developer.

4. Foundation Model Device Master Files

The paper also introduces the concept of a voluntary Foundation Model Device Master File (MAF), leveraging FDA's existing Device Master File program. A foundation model developer would confidentially submit information—including model architecture, training data provenance, healthcare-specific failure modes, subgroup benchmark results, guardrails, update notification commitments, and audit log availability—that device sponsors could reference with permission.

A MAF would not authorize the model for any medical use; the device sponsor would still bear the regulatory burden of proving its device is safe and effective. But it could streamline submissions for sponsors who build on well-documented foundation models. The open question is whether foundation model developers, who have limited regulatory incentive to disclose, will voluntarily participate.

5. Agentic AI Gets Its Own Competency Element

The paper addresses agentic AI systems: generative AI-enabled systems that autonomously plan and execute multi-step tasks, use external tools, or take actions across a sequence of steps. CDRH acknowledges that some agentic AI functions may meet the definition of a medical device, for example where an agentic system's action sequences result in control of another medical device.

Agentic competency would probe whether the system plans within its safety envelope, recognizes tool errors, and maintains human checkpoints before irreversible actions. This signals that FDA views agentic architectures as subject to heightened scrutiny, particularly where multi-step autonomy and tool use are involved.

Open Questions for Industry

The 26 discussion questions embedded in the paper deserve careful attention. Several stand out as particularly consequential for strategic planning:

  • Q7: Is the competency-based approach appropriate? This is a foundational question about whether the entire evaluation paradigm proposed in the paper is the right one.
  • Q10–Q12: Benchmark validity, clinical confirmation approach, and sample sizing. These questions will determine how much non-clinical testing can substitute for traditional clinical evidence—and how clinical evidence should be structured when outputs are open-ended.
  • Q18: The premarket-to-postmarket trade. This may be the single most consequential question. It asks whether the evidence burden for this category should shift from premarket to postmarket—a question that directly affects time-to-market and ongoing compliance costs.
  • Q24: Third-party foundation model changes. When the foundation model vendor updates its model, who bears the regulatory responsibility, and what triggers a new submission?
  • Q25: What would make a voluntary Foundation Model MAF useful? This probes whether the incentive structure for foundation model developers can make the MAF concept viable.

What You Should Do Now

  • Submit comments by October 19, 2026. FDA accepts partial responses—you do not need to address all 26 questions. Focus on the questions most relevant to your product, your business, and your regulatory strategy.
  • Review your intended use statements. The two-axis risk framework will reshape how FDA evaluates directiveness and consequence. If your product interacts with patients or provides clinical recommendations, assess where it falls on the grid and whether your current intended-use language adequately captures both dimensions.
  • Evaluate your foundation model dependencies. If your product relies on a third-party foundation model, the paper's discussion of vendor-initiated model changes (Q24) and the MAF concept (Q25) should inform your vendor contracting and change-management strategies now—before formal guidance arrives.
  • Engage with trade associations and standards bodies. FDA's interest in shared, standardized benchmarks suggests an industry role in developing the evaluation infrastructure. Early participation positions your organization to influence these standards rather than simply comply with them.
  • For investors, this paper signals that FDA is working to create a viable—if rigorous—path to market for generative AI-enabled medical devices. The competency-based approach, the postmarket monitoring trade, and the Foundation Model MAF concept all suggest that CDRH is actively seeking frameworks that enable innovation rather than block it. Companies that engage constructively in this comment period and begin building the benchmarking and monitoring infrastructure the paper envisions will be better positioned when formal guidance follows.

For questions about how these developments affect your products or regulatory strategy, please contact the Maynard Nexsen Health Care & Life Sciences Practice Group.

About Maynard Nexsen

Maynard® is a nationally ranked, full-service law firm with more than 600 attorneys nationwide, representing public and private clients across diverse industries. The firm fosters entrepreneurial growth and delivers innovative, high-quality legal solutions to support client success.

Media Contact

Tina Emerson

Chief Marketing Officer
TEmerson@maynardnexsen.com 

Direct: 803.540.2105

Photo of FDA Releases Discussion Paper on Generative AI-Enabled Medical Devices: What Manufacturers and Investors Need to Know
Jump to Page