Prove It to a Regulator: The Evidence an AI Audit
Level 3 · Regulated — Reading time: 12 minutes Last reviewed: August 2026
The request rarely arrives with much warning. A customer’s procurement team sends a questionnaire with an AI section that did not exist last year. A notified body asks for the technical documentation. A regulator opens an inquiry about a specific decision your system made about a specific person, eleven months ago.
At that moment, one question decides how the next six weeks go: does the evidence already exist, or are you about to create it?
Organizations that have to create it can usually produce something. What they produce is a set of documents authored after the request, describing decisions taken before it, signed by people reconstructing their reasoning from memory. Auditors recognize this immediately, and it is worse than having less documentation, because it raises a question about everything else you have submitted.
This guide covers what evidence actually means in an AI audit, which artifacts the EU AI Act and ISO/IEC 42001 expect, and how to build an evidence set that accumulates as a by product of running your program rather than as a project you undertake under pressure.
Assurance is a claim. Evidence is a record.
Most AI governance material stops at assurance — the policy exists, the committee meets, the risk framework is adopted. Those are claims about how your organization operates. An auditor’s job is to test whether the claims match what happened.
Four properties separate evidence from documentation:
Contemporaneous. Created at the time the decision was made, not afterward. A risk
assessment dated three days before deployment carries weight. The same document
produced during the audit carries almost none.
Attributable. A named person made the decision. “The committee reviewed it” is weaker than a record showing who was present, what was presented, and who accepted the residual risk.
Complete for its scope. If you claim every high-risk system gets an impact assessment, the auditor will sample. One system without one turns a control into an exception, and exceptions get expanded testing.
Retrievable. You can produce it within the response window without a search operation. Evidence that exists but cannot be located within the deadline fails in the same way as evidence that does not exist.

That last one is more often the failure than people expect. The material is usually somewhere — in an email thread, a Slack channel, someone’s local drive, a ticket nobody thought to tag.
What the EU AI Act expects from a high-risk system
The obligations differ substantially by role. Providers of high-risk systems carry the heaviest set; deployers carry a lighter but non-trivial one. Establish your role first, because building provider-grade documentation when you are a deployer wastes months, and the reverse leaves you exposed.
For providers of high-risk systems, the documentation set centers on:
Technical documentation. A structured description of the system, its intended purpose, design choices, and performance — detailed enough that a competent authority can assess conformity. Annex IV sets out what it must contain, and it is more prescriptive than most teams expect on first reading.
A risk management system. Not a one-time assessment but a continuous, documented process across the lifecycle: identify risks, evaluate them, adopt mitigations, test, and revisit as the system and its context change. The evidence is the sequence of records over time, not the framework document.
Data governance records. Training, validation, and testing data — where it came from, how it was prepared, what examination for bias was performed, and what gaps or limitations are known.
Automatic logging. High-risk systems must record events over their lifetime to a degree that supports traceability. This is a design requirement, not a policy one; retrofitting logging into a deployed system is expensive, so it belongs in your intake checklist rather than your remediation plan.
Human oversight design. Documented evidence that the system was built so a person can understand its output, monitor it, and intervene — and that the person assigned has the competence and authority to do so.
Accuracy, robustness, and cybersecurity. Declared performance levels and the testing that supports the declaration.
A quality management system, plus conformity assessment and registration steps before the system is placed on the market.
Post-market monitoring. A plan and the records showing it operates — including serious
incident reporting.
Deployers carry a narrower set: use the system per instructions, assign competent human oversight, monitor operation, retain logs, inform affected people where required, and in some public-sector and essential-services contexts complete a fundamental rights impact assessment.
Article and annex references shift as consolidated texts and guidance are published. Verify specific citations against the current official text before relying on them in a submission.
What an ISO/IEC 42001 auditor samples
Certification audits follow a familiar management-system pattern, which is good news if you have been through ISO 27001. The auditor is testing whether the system operates, and they test it by sampling records.
Expect requests for:
- The AI policy, with evidence of approval and communication
- A defined scope, and a rationale for what sits outside it
- Roles and responsibilities, with evidence the people named know they hold them
- The AI system inventory, current as of the audit date
- Risk assessments and treatment decisions, with dates and owners
- Objectives, and measurement showing whether they were met
- Competence records for people in defined roles
- Supplier and third-party controls, applied to actual suppliers
- Internal audit records, including findings — an internal audit that never found anything is
- itself a finding
- Management review minutes, with decisions and follow-through
- Nonconformities and corrective actions, closed with evidence
The pattern is consistent: for every claim, a record; for every record, a date and a name.
The three failures that cost the most
Retrospective documentation. Producing artifacts after the request, dated to lookcontemporaneous. Beyond the audit consequence, this can become a serious problem in a regulatory context. Late documentation is recoverable; misrepresenting when it was created is not.
Orphaned artifacts. A risk assessment exists but references a system that is no longer in the inventory under that name. A policy references a committee that stopped meeting. Auditors follow these threads, and a broken chain suggests the program stopped running at some point.
Undated decisions. A document saying a risk was accepted, without recording who accepted it or when. This is the single most common gap, and the easiest to fix: add an approver and a date field to every template you use.
Structuring the evidence pack
Organize by obligation, not by document type. When a request arrives citing a specific requirement, you want to open one folder rather than assemble from six.
A structure that holds up:

Two habits make this work. Keep quarterly snapshots of the inventory, so you can show what you knew at a point in time rather than only what you know now. And record decisions in the same place every time, so nobody has to remember whether the approval was in email or a ticket.
What good looks like under pressure
- A program is audit-ready when it can answer these within a week, without a project:
- Which of our AI systems are high-risk, and what is our role for each?
- Show the risk assessment for this system, and the decision to accept the residual risk.
- Who provides human oversight here, and what evidence shows they can override the
- output?
- What did the inventory look like eighteen months ago?
- What incidents occurred, and what changed as a result?
- When did the committee last review this, and what did it decide?
If any answer requires reconstructing rather than retrieving, that is where to focus.
A ninety-day path
Weeks 1–3. Confirm your role for each system. Classify against the risk tiers. This determines everything downstream, and getting it wrong is the expensive error.
Weeks 4–7. Build the evidence structure. Move what already exists into it. Mostorganizations find they have more than they expected, scattered across mail and drives.
Weeks 8–10. Fix the gaps in the highest-risk systems only. Resist the urge to work across the whole inventory — an auditor samples from the top of your risk ranking.
Weeks 11–13. Run a rehearsal. Pick two systems, ask a colleague to request the evidence as an auditor would, and time it. What you cannot retrieve in a week, you cannot retrieve under real conditions either.
The rehearsal is the step people skip, and it is the one that surfaces the retrievability problem while there is still time to fix it.
Where this leads
Audit readiness is not a phase at the end of a program. It is a property of how the program records what it does. If the inventory is current, decisions carry names and dates, and artifacts live in one predictable structure, readiness is largely a filing question rather than a documentation project.
Earlier guides in this path cover building the AI system inventory and what AI governance
covers — both are prerequisites for the work described here.
Get the toolkit
Our EU AI Act & ISO 42001 Toolkit includes the gap assessments, the AI impact assessment, the evidence pack structure above, and the document templates with approver and date fields already built in — tailored to your sector, size, and risk profile.
Facing a questionnaire, a notified body, or a customer audit with a deadline? Book a free 30-minute call. Thirty minutes, no slides — you leave with the two or three things worth doing
first.
References
- Regulation (EU) 2024/1689 (EU AI Act) — high-risk obligations and Annex IV
- ISO/IEC 42001:2023 — AI management system requirements
- NIST AI Risk Management Framework 1.0 — Govern and Manage functions
Advisory content, not legal advice. Verify specific article references against the current
official text.
