|

Security and Governance Evidence for AI Features 

Diagram showing a digital workflow passing through a central trust boundary and human review checkpoint before approval.

Cyber Alchemy × Mantra Systems  |  Episode 4 

This article is published in partnership with Mantra Systems. Cyber Alchemy focuses on cybersecurity, helping teams develop and evidence security for software-enabled and connected medical devices. Mantra Systems specialises in regulatory strategy and clinical evidence for UKCA and EU MDR/IVDR pathways. Together, we are producing a practical series for MedTech teams: what to build, what to defer, and how to avoid avoidable rework when moving between UK, NHS procurement, and EU routes. 

The guidance has not fully settled. The questions have not waited.

If you are adding AI to a medical device right now, you are working slightly ahead of the rulebook. Specific, settled guidance for AI in medical devices is still forming, and that can feel like an awkward place to build from. Here is the reassuring part: reviewers, notified bodies and NHS procurement teams are not waiting for perfect guidance, and nor should you. In practice they apply the expectations that already exist (clear boundaries, risk traceability, verification evidence, and post-market monitoring) and then ask a handful of extra questions where AI changes the threat model and the governance surface. 

That is good news, because it means the evidence you need is largely knowable today. You do not have to predict the final shape of every framework to prepare well. You have to assemble security and governance evidence that is auditable, testable and maintainable, in a form a reviewer can follow. 

This article is the Cyber Alchemy perspective on how to do that. We use an ambient scribe as the main worked example, with decision support and anomaly detection as two further archetypes, and we show how the same evidence maps across the regimes you are likely to meet. The companion article from Mantra Systems covers the clinical and regulatory evidence narrative when guidance lags. It is written for early and mid-stage MedTech teams who are building AI features now and want to avoid leaving the evidence until late. 

Key Takeaways

  • You can prepare now without waiting for settled AI rules. Reviewers apply existing expectations (boundaries, traceability, verification, post-market monitoring) and ask AI-specific questions on top. Those questions are knowable today. 
  • Keep AI evidence in one place: an AI Security and Governance Annex inside your wider Security Evidence Pack. It stops AI evidence being scattered across notebooks, vendor PDFs and slide decks. 
  • Three things decide most AI reviews: a clear AI boundary, real auditability (you can reconstruct what the model did and what the clinician did next), and a tested rollback or safe state. Weakness in any one generates delay. 
  • Treat untrusted input as untrusted. For language-model features, prompt and data injection is a real control area. Define what the AI is allowed to do, then test those boundaries.
  •  Build the evidence once and it travels. The same annex maps onto the EU AI Act, FDA expectations (including a PCCP and GMLP), EU MDR and NHS DTAC, with packaging differences rather than rebuilds. 

    Companion perspective (Mantra Systems)

    This article focuses on the cybersecurity and governance evidence for an AI feature: boundaries, abuse cases, auditability, rollback and monitoring. The companion piece from Mantra Systems covers the clinical and regulatory side: how to tell a credible safety and performance story for AI when specific guidance is still catching up, and how that fits your technical documentation. 

    ►  Read Mantra Systems’ expert perspective

    ►  Book a joint Cyber Alchemy × Mantra Systems review: contact us 

    Why this matters now

    It helps to see where the frameworks actually are, because reality is quite fragmented. The detail differs by regime, but they are converging on the same expectations. 

    The regulatory picture is firming up, not absent

    Under the EU AI Act, an AI feature inside a medical device is treated as a high-risk AI system where the device itself requires third-party conformity assessment, which catches most of the market above Class I self-certification. That brings expectations around data governance, transparency, human oversight, risk management and post-market monitoring. The original timeline applied these obligations to AI that is itself part of a regulated product (such as a medical device) from August 2027, and a proposed “Digital Omnibus” package announced in late 2025 is expected to adjust the dates and streamline how AI Act and EU MDR/IVDR assessments fit together. The exact deadlines are still moving; the direction is not. 

    In the United States, the FDA finalised its guidance on Predetermined Change Control Plans (PCCPs) for AI-enabled device software functions in December 2024, and followed it with draft lifecycle-management guidance in January 2025. A PCCP lets you pre-specify how a model may change, the methodology you will use to validate those changes, and an assessment of their impact, so the device can improve without a fresh submission each time. The groundwork for that came earlier: in October 2023 the FDA, Health Canada and the UK’s MHRA jointly published five guiding principles for PCCPs, which must be focused and bounded, risk-based, evidence-based, transparent, and grounded in a total product lifecycle perspective. Alongside this sit the ten Good Machine Learning Practice (GMLP) principles, first issued by those three regulators and finalised by the International Medical Device Regulators Forum (IMDRF) in January 2025. They are non-binding, but increasingly expected in submissions. 

    In the UK, the MHRA has run its AI Airlock, a regulatory sandbox for AI as a Medical Device (AIaMD), since 2024, with a second phase running through 2026 focused on risk classification, change management and bias and fairness. And the existing EU MDR expectations already apply: a state-of-the-art lifecycle including information security (GSPR 17.2), minimum security requirements for safe operation (GSPR 17.4), the cybersecurity guidance in MDCG 2019-16, and cybersecurity risks reflected in your ISO 14971 risk file. 

    Put together, that is a lot of acronyms pointing at one idea: show what your AI does, prove you have thought about how it could go wrong, evidence that your controls work, and demonstrate that you keep watching after launch. Build evidence for that and you are ready for most of what any of these regimes asks. 

    Where early-stage teams get caught

    The usual trap is focus rather than negligence. Teams pour their energy into making the AI feature genuinely good, which is exactly where it should go, and leave the evidence until a procurement questionnaire or a notified body asks for it. That late-stage scramble tends to produce the same three problems: an unclear AI boundary, weak auditability, and no credible rollback plan. Each one generates delay and difficult questions. The fix is the same work done earlier, in a way you can maintain. 

    What good looks like

    A pile of machine-learning artefacts does not count as governance evidence on its own. What a reviewer needs is a readable annex that answers a short list of questions: what the AI does and does not do, what data flows exist, what credible abuse cases there are, what controls address them, how those controls were verified, and how the feature is monitored and controlled after launch. 

    The simplest effective pattern is an AI Security and Governance Annex that sits inside your wider Security Evidence Pack (the single, versioned evidence base introduced earlier in this series). Keeping AI evidence in one named, versioned place stops it being scattered across notebooks, vendor documents and slide decks, and it means a reviewer can find any answer in minutes rather than reconstructing it from fragments. 

    Figure1d - security and governance evidence for ai features 

    One annex, several regimes

    Because the regimes are converging, one well-built annex answers most of them. The truth stays the same for every audience; only the packaging changes. 

    Regime What it expects for AI Where your annex answers it 
    EU AI Act (high-risk) Data governance, transparency, human oversight, risk management, post-market monitoring AI boundary and data-governance statement; human-in-the-loop design; AI threat model; auditability; monitoring plan 
    FDA (PCCP + GMLP) Pre-specified, validated change control; good ML practice across the lifecycle Rollback and version control plus a change plan; testing evidence; drift and performance monitoring 
    EU MDR (GSPR 17.2/17.4, MDCG 2019-16, ISO 14971) State-of-the-art secure lifecycle; security risks in the risk file; verification AI threat model mapped to controls and verification; residual risk recorded in the ISO 14971 file 
    NHS DTAC Organisational posture and product controls; data protection Data-governance statement; access control and audit logging; testing evidence summary 

    One AI Security and Governance Annex, packaged for each regime. Build the truth once, then tailor the presentation. 

    The minimum artefacts list

    For most AI-enabled MedTech features, the minimum set of security and governance artefacts is short. Each one answers a question a reviewer will ask, so aim for one named, versioned artefact per question rather than the same content spread across ten documents. 

    • AI boundary and data-flow diagram. Training, inference and logging, plus retrieval or tool calls if your feature uses them. What is in scope, and what is not. 
    • AI threat model summary. Credible abuse cases for your intended environment, each mapped to a control and to verification evidence. 
    • Data-governance statement. Data lineage, access control, retention, segregation and the tenant model, plus where models are hosted and who approves changes. This is the same material we work through with clients building AI governance from scratch; writing it down for a reviewer is the easy part once the decisions exist. Be explicit about what is captured and why. 
    • Testing evidence summary. What “testing the AI” means for your feature, including prompt and data injection checks where a language model is involved. 
    • Auditability design. What is logged and why, retention, access control, and how tampering is detected. You should be able to reconstruct which model version produced an output, what context it used, and what the clinician did next. 
    • Rollback, kill switch and safe-state plan, with evidence it works: feature flags, version rollback, and a communications plan. 
    • Post-market monitoring plan. Security signals together with performance and drift signals, and the triage workflow that connects them to a decision. 
    • Incident response addendum. How AI-specific incidents are handled and documented, as an extension of your existing incident plan. 

    If those eight artefacts exist, are versioned and cross-reference each other, most of the questions a reviewer will ask already have answers waiting. 

    Common failure patterns

    AI evidence tends to fall short in predictable, fixable ways. We recognise these patterns because we keep meeting them in the AI systems we test and the governance work we do, from vetting platforms handling sensitive data to clinical tools. None of these mean the AI is bad. They mean the evidence has not caught up with it yet. 

    • The AI boundary is implicit. No one can point to which systems and data are actually in scope, so controls look incomplete to a reviewer and excessive to you. 
    • Auditability is weak. You cannot reconstruct what model and version produced an output, what context it used, and what the clinician did next. From an evidence point of view, an undescribed log is the same as no log.
    • There is no credible rollback. No safe-state behaviour and no tested mechanism to disable the AI quickly if something goes wrong. 
    • Security testing is vague. “We tested the platform” but not the AI workflows themselves: prompt and data injection, tool misuse, and permission boundaries. 
    • Data retention and access control are unclear, especially where logs or prompts may contain sensitive clinical information. 

    Practical steps: three AI archetypes

    Below are worked examples for three common AI archetypes: an ambient scribe (language-model heavy), decision support, and anomaly detection or triage. The aim is not to boil the ocean. It is to produce evidence that is auditable, testable and maintainable, using controls that are proportionate to the real risk. 

    Worked example 1: ambient scribe, injection and auditability 

    Ambient scribe workflows often take in untrusted input: spoken content, imported documents, copied-and-pasted text. If that input can influence retrieval or trigger tool calls, you need explicit controls and explicit tests. The recognised reference point here is the OWASP Top 10 for Large Language Model Applications, where prompt injection sits at the top of the list. You do not need to treat it as exotic; you need to treat input as untrusted and prove you have. We have tested this boundary on a live AI scribe, and the findings that mattered sat not inside the model but where the AI’s reach met the platform’s access control. We also built a deliberately vulnerable demonstrator, Dr Llama, shown live at Digital Health Rewired 2026, to make the point in public: bolt an LLM onto a system that does not validate or enforce access control underneath it, and the AI inherits every gap. – Dr llama being broken by attendees 

    Minimum controls typically include a strict permission boundary (what the AI is allowed to do, and what it is not), input sanitisation and policy constraints, segregation of tenants and data, and audit logging that records the model version and the key decision points. In evidence terms: 

    • Define the permission boundary: the specific actions the AI can trigger, and the ones it cannot.
    •  Test the injection paths: untrusted text that attempts to override instructions or exfiltrate data, with the test set referenced by a stable ID. 
    • Log provenance without over-collecting: model version, retrieval sources, key events and clinician overrides, with sensitive content minimised. 

    The abuse-case mapping is the single most useful page in the annex. Every row links a credible abuse case to its control and the evidence that the control works, in the same shape your wider risk file uses: 

    Abuse case Impact Control Verification Monitoring signal 
    Injection via dictated or imported text overrides instructions Unintended action or data leak Input treated as untrusted; strict permission boundary; policy limits on tool/retrieval calls Injection test set (OWASP LLM01), V-AI-01 Spike in blocked or anomalous tool calls 
    Sensitive data over-collected in prompts or logs Confidentiality and data-protection risk Minimise captured data; redact identifiers in logs; tenant segregation Log-content review, V-AI-02 Alerts on identifier patterns in logs 
    Output altered between inference and the saved note Integrity risk to the record Integrity-checked output path; access control on configuration Integrity test, V-AI-03 Integrity-check failures 
    Model version changes behaviour silently Safety or performance drift Release-tied model version; PCCP-style change control; drift monitoring Drift report, V-AI-04 Drift threshold breached; version mismatch 

    Example AI abuse-case mapping for an ambient scribe. Stable IDs, one line per column, references that resolve. 

    Worked example 2: decision support, integrity and automation bias 

    Decision support features raise integrity and safety-adjacent questions, because the output influences a clinical decision. Your security evidence should show how outputs are protected from manipulation, how the system behaves in a degraded state, and how you keep auditability when a recommendation feeds into care. 

    From a security perspective, emphasise access control on the model configuration, protection of input-data integrity, audit trails for each recommendation, and monitoring for unusual patterns. From a human-oversight perspective (which the EU AI Act expects explicitly), show how the design keeps a clinician meaningfully in the loop and guards against automation bias, the tendency to over-trust a confident-looking output. A short statement of how the interface presents uncertainty, and how an override is captured, does a lot of evidential work here. 

    Figure2b 1 - security and governance evidence for ai features 

    Worked example 3: anomaly detection and triage, monitoring and rollback 

    For anomaly detection, post-market monitoring is not optional, because model performance and drift are part of the risk story rather than an afterthought. Pair drift monitoring with security monitoring: unusual inputs, data-pipeline anomalies, and model-version changes all belong on the same dashboard. 

    The key piece of evidence is a tested rollback plan: feature flags, a defined safe-state behaviour, and a documented decision path for disabling the feature and communicating the change. A rollback you have rehearsed is credible. A rollback you have only described is not, and reviewers can tell the difference. 

    Figure3 1 - security and governance evidence for ai features 

    From experience: the strongest AI features we assess treat auditability and rollback as first-class controls, designed in from the start, rather than afterthoughts bolted on before review. Added late, they are hard to evidence credibly, because the logs and the safe-state behaviour were never built to be shown. Added early, they are some of the easiest evidence in the whole pack. 

    Free downloads

    System Boundaries: Who Controls What

    Defining the boundary is the highest-leverage first step for any AI feature, because it sets the scope for everything else. This one-page reference covers device and firmware, mobile applications, cloud and backend, integrations, and shared responsibilities, with the explicit assumptions that need to be written into your documentation. 

    10 Core Procurement Artefacts

    Your AI annex sits inside the wider Security Evidence Pack. This checklist covers the ten artefacts most frequently requested across DTAC, EU MDR and private procurement, with a suggested cadence for keeping them current. 

    Companion perspective and next step

    Mantra Systems' companion article covers the clinical and regulatory evidence narrative for AI features when specific guidance lags: how to tell a credible safety and performance story, and how it fits your technical documentation. Read both to get the full picture of AI evidence as something you assemble deliberately rather than scramble for late. 

    ►  Read the companion perspective from Mantra Systems: 

    Next step: book a joint review 

    If you want a clear, evidence-first view of where your AI feature stands against the EU AI Act, FDA expectations, EU MDR and NHS DTAC, and a sense of what to build first, book a joint review with Cyber Alchemy and Mantra Systems. In 30 minutes we will: 

    • Clarify your AI boundary and intended environment, and what your documentation currently implies. 
    • Identify your top three cybersecurity and governance assurance gaps, and how they connect to the wider evidence story. 
    • Agree the minimum artefacts and verification evidence you need, so you do not overbuild for markets you may never enter. 
    • Turn that into a short, prioritised action plan you can execute over the next two to four weeks. 

    ►  Book a joint review: contact us 

    If you want to go deeper on the cybersecurity side specifically, our AI security services page sets out how we help teams secure and evidence AI features. 

    FAQs

    Does the EU AI Act apply to the AI in my medical device?

    If your device includes an AI feature, it is generally treated as a high-risk AI system under the EU AI Act, which brings expectations around data governance, transparency, human oversight, risk management and post-market monitoring. The application dates are still moving (a proposed Digital Omnibus package is expected to adjust the timeline and align AI Act and EU MDR/IVDR assessments), so confirm the current position for your route, but the substance of what to prepare is already clear. 

    What is a PCCP and do I need one?

    A Predetermined Change Control Plan lets you pre-specify how an AI model may change, the methodology you will use to validate those changes, and an assessment of their impact, so the device can improve without a new submission each time. The FDA finalised its PCCP guidance in December 2024. It is most relevant for FDA routes, but the underlying discipline (planned, validated, version-controlled change) is good practice everywhere and maps well onto EU expectations. 

    What does it actually mean to "pen test" and AI feature?

    It means testing the AI workflow, not just the platform around it. For a language-model feature that includes prompt and data injection (the top entry in the OWASP Top 10 for Large Language Model Applications), tool misuse, and permission-boundary checks: can untrusted input make the AI do something it should not? The output is a scoped test summary with findings, retests and stable evidence IDs, the same shape as any other security test evidence. We have delivered exactly this, including penetration testing of an AI-enabled ultrasound analysis feature as part of an FDA 510(k) submission. And the findings are rarely where teams expect: in one AI assessment, input the model was never designed to handle took down the software around it, a denial of service the platform's own testing had no chance of catching because nobody had tested the AI workflow as a workflow. 

    What should we log for an AI feature?

    Enough to reconstruct what happened without over-collecting sensitive data: the model version, the inputs or retrieval sources used, the key decision points, and any clinician override. Define retention, who can read the logs, and how tampering is detected. An undescribed log counts as no log from an evidence perspective, so the one-page description matters as much as the logging itself. 

    Do we really need a kill switch?

     You need a credible, tested way to return to a safe state: a feature flag or version rollback, a defined safe-state behaviour, and a communications plan. It does not have to be elaborate. It does have to be rehearsed, because a rollback you have only described is not evidence that you can actually recover. 

    We have already built the AI feature. Is it too late?

    No. The artefacts here are the same ones you would assemble during a corrective action plan, and a pre-emptive evidence pass before the questions arrive is almost always faster and cheaper than reacting to them. The most common situation we see is good AI with thin evidence, and that gap is very closable. 

    Related reading

    • Security Risk Controls That Read Well in a Technical File: how to present controls as traceable evidence (Episode 3). 

    Similar Posts