Why Regulatory Reviewers Reject Black-Box AI Outputs, and What Documentation They Actually Want
explainable AI platform enterprise

Why Regulatory Reviewers Reject Black-Box AI Outputs, and What Documentation They Actually Want

Srinivas Padmanabharao

Author

Srinivas Padmanabharao

Published : 01 Aug 2026

Key Takeaways :

Explainable AI is no longer optional in pharmaceutical regulatory submissions. FDA, EMA, NICE, and G-BA increasingly expect AI systems to be transparent, traceable, and verifiable from the design stage, with clear documentation, human oversight, and claim-level source attribution. Rather than relying on black-box outputs and retrospective compliance efforts, organisations should adopt AI platforms built for explainability by design. KnolForge addresses these requirements through architectural transparency, complete audit trails, and inspection-ready documentation, helping enterprise pharma teams meet evolving regulatory expectations with confidence.

Human oversight in pharma AI is not a bolt-on control layer added at deployment. According to the FDA-EMA joint Guiding Principles published January 14, 2026, it is a design property of the AI system that must be established at the specification stage and validated throughout the lifecycle. If a sponsor cannot show at the concept stage how oversight will operate, later attempts to add human review to a black-box output are not going to satisfy a well-prepared reviewer. [1] This single sentence from the 2026 joint principles encapsulates why regulatory reviewers at NICE, G-BA, FDA, and EMA are increasingly rejecting AI-assisted submissions that lack adequate explainability documentation, and why the documentation gap is not closed by adding a reviewer sign-off box after an opaque model has already generated its output.

At Pienomial, we built KnolForge as the explainable AI platform enterprise pharma organisations need precisely because we understood that the explainability requirement is architectural, not procedural. Explainability cannot be retrofitted to a system built on black-box generation. It must be the foundational property from which every output is derived. This post explains why regulatory reviewers reject enterprise AI without black-box models alternatives, what specific documentation FDA, EMA, NICE, and G-BA actually require, and how KnolForge satisfies those requirements by design rather than through compliance documentation exercises.[9]

1. The Regulatory Rejection Problem: Why Black-Box Outputs Fail Review

The New England Journal of Medicine and STAT News editorial boards have both warned explicitly about black-box AI in healthcare and regulatory contexts. [4] The regulatory concern is not hypothetical and it is not primarily about accuracy. A black-box model can produce outputs that are accurate on average and still be rejected by a regulatory reviewer, because the reviewer cannot assess the basis of any specific output without the model's reasoning being visible. The accuracy track record of a model in testing does not tell a reviewer whether this specific output, applied to this specific regulatory question, is reliable.

This is the structural problem with black-box AI in regulatory submissions. A NICE technical reviewer examining a systematic literature review evidence table cannot verify that the AI model correctly applied the inclusion criteria to the papers it screened if the screening decision process is opaque. A G-BA scientific reviewer examining an indirect treatment comparison cannot verify that the AI correctly handled a methodological edge case in the network if the analytical steps are not documented. An FDA reviewer examining a clinical dossier cannot assess whether an AI-generated summary correctly represents the underlying source data if there is no traceable link from summary statement to source document.

The regulatory requirement is therefore not that AI be accurate. It is that AI be verifiable. And verifiability requires explainability documentation that black-box models structurally cannot provide.[3]

2. What the FDA's January 2025 Draft Guidance Requires for Explainability

The FDA's January 2025 draft guidance, 'Considerations for the Use of Artificial Intelligence to Support Regulatory Decision-Making for Drug and Biological Products', is the primary US regulatory document establishing explainability documentation requirements for pharma AI. The guidance expects drug sponsors to treat AI tools as part of their Quality Management System, not as productivity aids exempt from validation requirements. [3] Within the seven-step credibility assessment framework, explainability requirements appear at multiple steps.

At Step 2, Context of Use specification, the guidance requires sponsors to document precisely what the AI model will do, in what analytical context, and with what relationship to the human decision-maker's ultimate judgment. An AI model whose context of use is documented as 'assist with evidence synthesis' without specifying how the model's output will be evaluated by a qualified human, and what methodology that human will apply to assess the output's reliability, does not satisfy the COU requirement.[8]

At Step 3, Model and Analytical Process documentation, the guidance requires documentation of the model's architecture, training data sources, preprocessing decisions, and the specific analytical steps from input to output. For a generative AI model, this documentation must explain what the model was trained on and what mechanisms it uses to generate outputs, not merely that it uses a large language model. [3] The FDA has explicitly signalled that expectations for explainable AI models specifically covering generative AI and LLMs will be clarified in the final guidance expected in Q2 2026, because the January 2025 draft barely addressed these model classes.

At Step 6, Credibility Evidence Documentation, the FDA expects explicit documentation of model limitations and known failure modes, what the model does and does not do reliably, and how human review is designed to catch errors that the model's known limitations make likely. A model card, the standardised format for documenting model purpose, performance characteristics, known limitations, and appropriate use cases, is specifically cited as an expected element of regulatory acceptance. [2]

3. What NICE Requires for AI Used in HTA Evidence Submissions

NICE's 2024 position statement on AI in evidence generation is the primary HTA explainability documentation requirement for submissions targeting the UK market. The statement has three core requirements that directly address the black-box AI problem.[5]

First, AI use must be made explicit. Submissions may not use AI for any component of evidence generation without declaring that use. Undisclosed AI assistance in a literature review, evidence synthesis, or indirect treatment comparison is not acceptable, regardless of the quality of the output.[5]

Second, methods must be explained fully, including identified risks and mitigation steps, in language accessible to non-experts. This requirement excludes documentation that describes the AI methodology in technical terms a clinical evidence reviewer cannot evaluate. NICE explicitly expects that explainability be accessible to reviewers who are not AI specialists, which effectively requires that the AI system's operation be describable in terms that map to familiar evidence standards rather than requiring knowledge of machine learning methodology to assess.[6]

Third, NICE references two specific documentation frameworks that sponsors using AI are expected to apply. The PALISADE Checklist, developed by ISPOR's Machine Learning Task Force, provides structured guidance for ML methods transparency in HEOR, covering how models were developed, how they were validated, and what their limitations are. The TRIPOD+AI Checklist provides reporting standards for studies involving prediction modelling with AI components. Together these frameworks define the minimum documentation standard NICE expects before an AI-assisted submission will be accepted without challenge.[5]

4. What G-BA and IQWiG Expect: Reproducibility as the Core Standard

G-BA and IQWiG apply explainability standards through the lens of reproducibility: the methodology documented in a dossier must be sufficient for an independent assessor to reproduce the evidence synthesis and reach the same conclusion. This standard has significant implications for AI-assisted evidence generation.[9]

A systematic literature review conducted with AI assistance must document the search strategy with sufficient specificity that the search could be independently re-executed and verified to produce the same initial record set. An abstract screening process that used AI classification must document the specific inclusion and exclusion criteria the AI applied, with the precision required for a human reviewer to confirm whether a specific paper would correctly be included or excluded by the documented criteria. A data extraction process that used AI must document the specific data fields extracted, the source locations within each included study, and the handling of any ambiguous or conflicting data points.

When an AI system's operation is not sufficiently documented to satisfy this reproducibility standard, G-BA and IQWiG classify the evidence synthesis as methodologically inadequate, not merely underdocumented. The consequence is that the evidence is either excluded from the benefit assessment or assigned a lower evidence grade, directly affecting the reimbursement outcome regardless of the clinical quality of the underlying trial data the synthesis was based on.[9]

5. The Pharmacovigilance Explainability Standard: What Inspectors Look For in 2026

For AI systems used in pharmacovigilance, the 2026 inspection standard is the most operationally specific of any regulatory context. The joint FDA-EMA Guiding Principles published in January 2026 made explicit that AI governance in pharmacovigilance must be explainable, traceable, and inspection-ready, applying the same standards as any other GxP-regulated system. [7] When FDA or EMA inspectors examine an AI-enabled pharmacovigilance system, they are evaluating specific documented capabilities, not generic assertions of AI governance.

The pharmacovigilance process owner must possess a comprehensive understanding of the AI at a process level, must be capable of communicating its operation as it relates to patient safety and risks, and must be able to explain the AI to non-expert regulators to provide assurance about its appropriate use. Be prepared to explain in plain language how the AI makes decisions, without necessarily exposing proprietary algorithms, and what human review steps follow AI-generated outputs before they are used for regulatory reporting.[7]

At minimum, inspection-ready pharmacovigilance AI documentation must include: a system description within the Pharmacovigilance System Master File, validation records demonstrating that the AI performs within defined parameters for the specific safety use case, a description of the human review steps that follow AI output, a record of any AI-generated outputs that have been overridden by human reviewers with the rationale for each override, and a post-deployment performance monitoring record demonstrating ongoing performance within the validated parameters.[7]

6. The Documentation Package Regulators Actually Want: A Structured Checklist

Drawing on FDA, EMA, NICE, G-BA, and pharmacovigilance inspection requirements, the following structured documentation package represents what regulators across major pharma markets now expect for AI-assisted evidence generation used in regulatory submissions.[8]

1. AI System Description Document: What the AI system does, what it does not do, its context of use, the specific workflow steps it performs, and the interface points between AI execution and human review. This document must be accessible to a non-expert regulatory reviewer.

2. Training Data and Data Governance Documentation: Sources of training or knowledge base data, preprocessing decisions, bias assessment across relevant subgroups, representativeness analysis relative to the deployment population, and update and maintenance procedures for the data foundation.

3. Model Card or Equivalent Performance Documentation: Standardised documentation of model purpose, performance characteristics including known limitations and failure modes, benchmark performance on tasks relevant to the regulatory context of use, and appropriate use boundaries that define where the model should not be applied.[2]

4. Methodology Documentation Satisfying PALISADE and TRIPOD+AI: For HTA submissions, completed PALISADE and TRIPOD+AI checklists demonstrating transparency in ML methods development and findings, and in prediction modelling study reporting. This is not optional documentation for NICE submissions; it is a stated expectation.[5]

5. Complete Audit Trail: A timestamped, tamper-evident log of every AI action from input to output: which sources were queried, which records were retrieved, which screening decisions were made, which data extractions were performed, and which outputs were generated. This log must be reviewable by an external regulatory assessor.

6. Human Oversight Procedure Documentation: Documented procedures specifying what AI outputs are reviewed by which human role with what qualification, before what downstream action, and with what documented rationale when a human reviewer modifies or overrides an AI output.[1]

7. Post-Deployment Monitoring and Lifecycle Management Plan: Performance monitoring methodology, drift detection thresholds, the CAPA workflow integration for threshold breaches, and the change control procedures governing model updates and revalidation scope.[3]

7. Why Claim-Level Source Attribution Is the Key Explainability Feature for Submission Use Cases

Among all explainability capabilities, claim-level source attribution delivers the most direct value for the specific regulatory contexts where AI is being used in pharma submissions, because it satisfies the verifiability requirement at the level that regulatory reviewers actually care about: the individual factual claim.[9]

A NICE technical reviewer examining a systematic literature review evidence table does not need to understand how the AI model works. They need to be able to verify, for any specific cell in the evidence table, what primary source the value came from and where in that source document it was located. Claim-level attribution provides exactly this: a link from each claim in the output to the specific primary source, the specific section or table, and the specific extraction from that location.

Document-level attribution, where the output references a list of sources but does not specify which source supports which specific claim, does not satisfy this requirement. A reviewer who discovers that three of the sources in a document-level reference list do not support the specific claim they were examining cannot determine from document-level attribution whether the claim came from one of the other sources or was generated without primary source support at all.[9]

This is the specific explainability gap that most generic AI tools deployed in pharma evidence synthesis cannot close, because their generation process does not maintain claim-to-source links at the individual claim level. KnolForge closes this gap architecturally: every output claim is generated from a retrieved, validated, source-attributed knowledge graph triple, and the link from claim to source is a structural property of the output rather than a manually added annotation.

8. The Common Documentation Mistakes That Trigger Regulatory Rejection

Based on the FDA's review experience of over 500 AI-containing submissions and NICE's post-position statement submissions, several common documentation mistakes consistently trigger rejection or formal clarification requests.[4]

Mistake 1, Describing the AI system in marketing language rather than regulatory language: A submission that states the AI uses 'advanced natural language processing to synthesise evidence' without specifying what that means at the methodological level provides no basis for a reviewer to assess the system's reliability. Regulatory documentation must describe what the system does, step by step, not what it claims to achieve.

Mistake 2, Asserting accuracy without context-specific validation evidence: Citing general benchmark performance metrics that were not measured on tasks representative of the submission's specific context of use does not satisfy the FDA's context-of-use requirement. Accuracy on general biomedical NLP benchmarks does not predict performance on the specific extraction task the system performed for this specific regulatory application.[3]

Mistake 3, Treating human sign-off as a substitute for explainability: Adding a reviewer signature to an AI-generated output does not make that output explainable. NICE and FDA both require that the human reviewer can actually assess the reliability of the AI output, which requires that the AI's reasoning and sources are visible to that reviewer. A signature on an opaque output is not oversight; it is endorsement without evaluation.[1]

Mistake 4, Omitting documentation of AI use entirely: NICE's statement requires that AI use be made explicit. Any submission where AI assisted evidence generation without disclosure risks rejection on this ground alone, regardless of the quality of the underlying evidence.[5]

Mistake 5, Providing a single AI documentation section rather than claim-level attribution throughout: Regulatory reviewers are increasingly checking specific claims in submissions against their stated sources. An AI documentation section that explains the system's general operation does not help a reviewer who is trying to verify whether a specific efficacy figure cited in the evidence table was correctly extracted from the stated source.[9]

9. How Fast Can Your Team Deploy Inspection-Ready Explainable AI with KnolForge?

Building the explainable AI platform enterprise documentation package that FDA, EMA, NICE, and G-BA now require does not need to be a multi-month documentation exercise when the AI system is architected for explainability from the start. KnolForge generates the required documentation as a standard operational output rather than a retrospective compliance exercise.[9]

Sprint 1, Weeks 1 to 2, AI System Description Document and Audit Trail live: KnolForge's system description documentation, covering context of use, workflow step documentation, and human oversight interface points, is reviewed and confirmed for your primary use cases. The audit trail is active from the first session, generating a complete, timestamped log of every KnolAI action that satisfies both FDA ALCOA+ and NICE documentation requirements.

Sprint 2, Weeks 3 to 4, PALISADE, TRIPOD+AI, and Model Card documentation completed: Your HEOR and regulatory team completes the PALISADE and TRIPOD+AI checklists using KnolForge's platform documentation as the source material for the methodology sections. The model card for KnolForge's knowledge graph retrieval architecture is generated for inclusion in regulatory submission documentation packages.[5]

Sprint 3, Weeks 5 to 6, Claim-level attribution integrated into all submission outputs and human oversight procedures documented: KnolAI output formatting for submission use cases is confirmed to include claim-level source attribution for every evidence table and synthesis document. Human oversight procedures are documented in the required format for FDA QMS inclusion and PSMF update. Your organisation's submission documentation satisfies the explainability requirements across FDA, NICE, G-BA, and EMA simultaneously from this sprint forward.[9]

Conclusion

The regulatory rejection of black-box AI outputs in pharma is not a temporary caution that will ease as regulators grow more comfortable with AI. It is the expression of a permanent principle: outputs that influence patient safety and regulatory decisions must be verifiable, and verifiability requires explainability documentation that black-box generation cannot provide. The FDA, EMA, NICE, and G-BA have each communicated this principle clearly, in different regulatory languages but with the same underlying requirement.

At Pienomial, we built KnolForge as the explainable AI platform enterprise pharma organisations need because we believe the documentation regulators want should not require a separate documentation project layered on top of the AI system that produced the output. Claim-level source attribution, complete audit trails, and context-of-use documentation should be structural properties of every output, generated automatically by the architecture that produced it. That is what KnolForge delivers, and it is what turns AI from a regulatory liability into a regulatory asset. [9] 

CTA: See how KnolForge produces inspection-ready explainable AI documentation for every submission. Book a demo with the Pienomial team today.

Frequently Asked Questions

[1]  Sakara Digital (2026). Human-in-the-Loop Pharma AI: FDA and EMA Requirements. FDA-EMA 10 joint Guiding Principles published January 14, 2026. Human oversight is a design property of the AI system, not a bolt-on control layer. If a sponsor cannot show at the concept stage how oversight will operate, later attempts to add human review to a black-box output are not going to satisfy a well-prepared reviewer.  https://sakaradigital.com/blog/human-in-the-loop-requirements-pharma-ai-fda-ema/

[2]  Alignmt.ai (2026). What FDA's AI Guidance Really Demands. FDA expects model cards: standardised documentation of model purpose, performance characteristics, known limitations, and appropriate use cases as a condition of regulatory acceptance. High-risk AI systems face mandatory pre-market conformity assessment, post-market monitoring, human oversight, and technical documentation.  https://www.alignmt.ai/post/what-fda-s-ai-guidance-really-demands

[3]  IntuitionLabs (2026). Pharma AI Validation Packages for FDA and EMA Compliance. FDA January 2025 draft guidance outlines risk-based credibility assessment framework. AI-specific elements include pre-specified context of use, data provenance, transparency, bias control, and traceable documentation. FDA expects drug sponsors to treat AI tools as part of their Quality Management System.  https://intuitionlabs.ai/articles/pharma-ai-validation-evidence-fda-ema

[4]  IntuitionLabs (2026). FDA and EMA Good AI Practice Guide for Drug Development. FDA has reviewed over 500 drug submissions with AI components since 2016. NEJM publications have warned about black-box AI in healthcare. New England Journal and STAT News editorial emphasise transparent, evidence-based AI use.  https://intuitionlabs.ai/articles/fda-ema-good-ai-practice-drug-development-2

[5]  BeaconOne Healthcare Partners (2025). NICE Opens Door to Use of AI in HTA Submission. NICE requires submissions to make AI use explicit, explain methods fully including risks and mitigations, in language accessible to non-experts. References PALISADE Checklist from ISPOR ML Task Force and TRIPOD+AI Checklist for reporting prediction modelling studies.  https://beacononehcp.com/2025/02/11/nice-opens-door-to-use-of-ai-in-hta-submission/

[6]  PMC (2025). RWE Ready for Reimbursement: Developments in Real-World Evidence Relating to HTA Part 17. NICE requires submitting organisations to clearly justify AI use and outline assumptions, with more explainable methods presented as the first instance compared to less transparent approaches. Justification can use ISPOR's PALISADE framework.  https://pmc.ncbi.nlm.nih.gov/articles/PMC11650383/

[7]  Clinevotech (2026). AI Governance in Pharmacovigilance: 2026 Inspection Guide. Joint FDA-EMA principles early 2026 made explicit: AI governance in pharmacovigilance must be explainable, traceable, and inspection-ready, no different from any other GxP-regulated system. Documentation must include system description in PSMF, validation records, and plain-language explanation of how AI makes decisions.  https://www.clinevotech.com/blog/ai-governance-pharmacovigilance-2026/

[8]  IntuitionLabs (2026). FDA's AI Guidance: 7-Step Credibility Framework Explained. FDA-EMA January 2026 joint principles cover 10 areas including human-centric design, risk-based approach, clear context of use, data governance and documentation, model design, risk-based performance assessment, lifecycle management, and clear essential information.  https://intuitionlabs.ai/articles/fda-ai-drug-development-guidance

[9]  Pienomial (2025). KnolForge: Explainable AI Platform for Enterprise Pharma Submissions. Knolens claim-level attribution architecture satisfying FDA, EMA, NICE, and G-BA explainability documentation requirements.  https://www.pienomial.com/products/knol-forge

Connect With Us

Related Posts