Practical guide · Audit · Version 1.0 · October 2026
The first AI audit most teams run looks a lot like an IT general controls review with “AI” written on the cover.
The team asks for the AI policy, confirms there’s a governance committee, checks that access reviews happened, glances at the vendor’s assurance report, and signs off. Six months later the chatbot shows a junior employee information from the executive compensation file, and nobody can explain why the audit didn’t catch it.
It didn’t catch it because it never tested how the AI actually worked. AI systems fail in places traditional audit programs don’t always look: in the documents retrieved into a prompt, in the permissions an agent inherits, in a system prompt someone changed without approval, in a vendor model update nobody assessed, or in a test that “passed” against criteria chosen after the results were already in.
This guide is the procedure I use to avoid that. It takes an AI system audit from the objective through reporting and follow-up, and gives you the information request list, a risk and control matrix covering seventeen audit areas, sampling guidance, evidence tests, a findings template, and specific procedures for chatbots, RAG systems, agents, predictive models, and third-party foundation models.
This is a methodology guide, not a substitute for your organization’s audit standards. Align the procedures and sampling with your own methodology and with the legal, regulatory, contractual, and internal requirements that apply to the system. For the compliance side, see my AI Compliance Guide.
The AI audit at a glance
The overall shape will look familiar if you’ve run any controls audit. What changes with AI is the middle. You need to understand the system and its data flows well enough to know where it can fail, and you’ll do far more re-performance than in a traditional review. A screenshot showing a guardrail is enabled tells me something, but it doesn’t tell me whether the guardrail actually stops what it’s supposed to stop. If the control says the chatbot can’t retrieve documents a user isn’t authorized to see, I want to test that myself.
Step 1: Define the objective and scope
A vague scope is one of the most expensive mistakes you can make in an AI audit. “Audit our use of AI” can mean almost anything, and an audit that tries to cover everything usually ends up testing nothing deeply enough. Before fieldwork starts, I want the objective and scope written down and agreed with the business owner.
Start with one primary objective
Most AI audits fall into one of four types:
- Governance audit: is the organization’s AI governance program designed appropriately and operating effectively?
- System audit: are the controls over a specific AI system designed appropriately and operating effectively?
- Compliance audit: does the system meet a defined set of requirements, such as applicable EU AI Act obligations or NYC Local Law 144?
- Pre-implementation review: is the system ready to go live, with the right controls built in?
This guide focuses on the system audit, because that’s where most of the detailed technical testing happens.
Then pin down the scope on six dimensions
| Dimension | What to decide | Example |
|---|---|---|
| AI system | Exactly which system, version, and components | HR Policy Assistant v2.3, including the retrieval index and guardrail service |
| Business process | What process it supports, and where AI output enters it | Employee policy questions; output is informational, not a decision |
| Legal entity | Which entities own, operate, or use it | U.S. Inc. operates it; UK Ltd. employees also use it |
| Lifecycle stage | Design, development, pre-go-live, production, or changing | Live since March 2026; model upgrade planned for Q4 |
| Geography | Where it operates and where affected people are | U.S. and UK employees |
| Applicable criteria | Policies, laws, standards, and contract terms | AI policy, privacy requirements, vendor contract |
If these are hard to answer, resolve them before detailed testing begins. My Scoping an AI Audit guide goes deeper into how the type of AI, how it was acquired, how it’s used, and its lifecycle stage change what belongs in scope.
Scope statement template
This audit will assess whether the controls over [system name and version], used by [business units / legal entities] to support [business process], are designed and operating effectively to manage [key risks], against [criteria], for the period [start date] through [end date]. The audit covers [components in scope] and excludes [out-of-scope components and the reason].
Write the exclusions down. If the underlying foundation model is outside your direct scope because you’re relying on the provider’s assurance report, say so, and then evaluate that reliance as part of the audit. “Out of scope” should never quietly mean “we didn’t look at it.”
Step 2: Send the pre-audit information request
I usually send the request at least two weeks before fieldwork, and the response tells me something before testing even starts. A mature team produces most of the core documentation without having to create it for the audit. A less mature team often discovers, while answering, that the risk assessment was never approved, nobody knows which model version is in production, or the architecture diagram is two releases behind. That isn’t automatically a finding, but it tells me where to look harder.
Governance and accountability
- AI inventory record, with business and technical owners
- Risk tier and classification rationale
- Go-live approval, with date, approver, and any conditions
- Relevant AI governance committee minutes
- AI policy and standards in force during the audit period
System documentation
- Purpose, intended users, and intended use
- Architecture and data-flow diagrams
- Model or system card: model name, provider, version, and known limitations
- System prompts and prompt templates, with version history
- Tools, APIs, and integrations available to the system, and the permissions each one uses
- Retrieval sources, and documentation of how retrieval permissions are enforced
Risk, testing, and oversight
- Risk assessment, and impact assessment where applicable
- Test plan with pass/fail criteria, a description of the test data, and results
- Red-team or adversarial testing
- Bias or fairness testing where relevant
- Human oversight procedure, with override or escalation records
- User-facing notices and disclosures
Operations
- Monitoring metrics, thresholds, and the last three months of reports
- Logging configuration and retention settings
- AI-related incidents and complaints
- Change history for models, prompts, retrieval sources, and configuration
- Administrator access list, and everyone able to modify prompts, models, guardrails, or configuration
Third parties
- AI vendor contracts and data processing terms
- Vendor assurance reports: SOC reports and ISO/IEC 42001 certificates where applicable
- Vendor model-change notices received during the period
- Documentation of which responsibilities sit with the provider and which with you
Don’t ask for documents just because they sound useful. Every item should eventually connect to a risk, a control, or an audit criterion.
Step 3: Understand how the system actually works
I don’t start detailed testing until I can draw the system myself. Run a walkthrough with the technical owner, take one realistic request, and follow it from start to finish. The user types a question: where does it go, what gets added to it, what gets retrieved, what leaves the organization, what can the model call, and what happens to the answer before the user sees it?
For a typical generative AI application, I trace at least these eight points:
- User and authentication. Who can access the system, and how does the application know who the user is?
- Application and orchestration. Where is the final prompt assembled from the system prompt, the user’s message, conversation history, and retrieved documents?
- Input controls. What is checked, blocked, filtered, or redacted before the request reaches the model?
- Retrieval. Which sources can it search, and, more importantly, is retrieval limited to what this particular user is allowed to see?
- Model call. Which model and version is actually running, where is it hosted, and what is sent to the provider?
- Tool calls. What can the system do besides generate text? Send email, create tickets, read or modify customer records, run code, start transactions? And under whose permissions?
- Output controls. What’s checked before the response reaches the user, such as sensitive data or unsafe content?
- Logging. Can you reconstruct the request, retrieval, model version, output, and actions afterward?
Then mark the trust boundaries: the places where data or control passes from one party or component to another. The user passes data to the application, the application sends information to a model provider, the model calls another system, an agent acts through an API. I treat every boundary as a point to assess for risk and expected controls, and many significant AI findings sit right there.
Finally, compare the walkthrough with the documents you received. If the architecture diagram says one thing and the engineer describes another, don’t jump straight to a finding; investigate. Maybe the diagram is stale, maybe the implementation changed without going through change control, maybe nobody maintains the documentation. Until you’ve established the criteria, cause, and risk, it’s an exception to follow up, not a finding.
Step 4: Build the risk and control matrix
This becomes the core working paper. For each audit area I record:
Risk → Expected control → Evidence → Test procedure → Result → ConclusionNot every area applies to every system, and that’s fine. Mark an area not applicable only when it genuinely isn’t, and write down why. These are the seventeen areas I normally start from:
| # | Audit area | Key risk | Control you expect | Evidence | How I test it |
|---|---|---|---|---|---|
| 1 | Governance and accountability | Nobody owns the risk | Named business and technical owners; approval | Inventory, approvals | Confirm ownership, and approval before go-live |
| 2 | Inventory and classification | System gets too few controls | Documented classification | Classification record | Re-perform the classification |
| 3 | Model and vendor documentation | Deployed model or its limits aren’t understood | Current system and model documentation | Model card, version history | Compare documentation with production |
| 4 | Data governance | Inappropriate or poor-quality data is used | Approved sources, lineage, quality controls | Data records | Trace a sample of sources |
| 5 | Privacy | Personal data is mishandled | Privacy assessment, minimization, retention | Assessment, settings | Inspect actual configuration and vendor data terms |
| 6 | Security and access | Leakage, injection, unauthorized changes | Access controls and technical guardrails | Configuration, access lists | Adversarial testing and access review |
| 7 | Risk assessment | Material risks weren’t identified | Approved assessment | Risk assessment | Check scope, version, and mapping to controls |
| 8 | Testing and evaluation | Released without evidence it works | Test criteria defined in advance | Plan, results, sign-off | Confirm criteria predate results; re-perform tests |
| 9 | Accuracy and performance | Wrong outputs cause harm | Metrics and thresholds | Evaluation results | Independently score outputs |
| 10 | Bias and fairness | Outcomes differ unfairly across groups | Appropriate fairness testing | Test results | Review and recompute material metrics |
| 11 | Human oversight | Review is ineffective | Defined review and override process | Procedure, override records | Sample decisions; look for real challenge |
| 12 | Transparency | People don’t know AI is involved | Required notices | Live notice, screenshots | Observe the system |
| 13 | Logging | Activity can’t be reconstructed | Sufficient event logging | Configuration, logs | Reconstruct transactions |
| 14 | Monitoring | Degradation or misuse goes unnoticed | Metrics, thresholds, escalation | Monitoring reports | Test alerts and follow-up |
| 15 | Incident management | Failures aren’t escalated | AI included in the incident process | Incident records | Trace incidents and complaints |
| 16 | Change management | Unapproved changes alter behavior | Changes tested and approved | Change tickets, tests | Reconcile deployments to approvals |
| 17 | Third parties | Vendor creates unmanaged risk | Due diligence, contracts, monitoring | Contracts, assurance | Test scope and obligations |
Treat the matrix as a starting point. The architecture and the risk assessment should decide where your hours go: a RAG system deserves more retrieval testing, an agent deserves more permission and action testing, and a predictive employment model deserves more attention to data, validation, and fairness. Don’t give every row equal time just because the spreadsheet has seventeen rows.
Step 5: Test the controls
Design and operating effectiveness are different questions
I keep these two conclusions separate. Design effectiveness asks: if this control works exactly as intended, would it actually address the risk? Operating effectiveness asks: did it actually run, consistently, throughout the period? A control can fail either one.
Say the organization’s confidentiality control is a system prompt telling the model, “Never reveal salary information.” That instruction may be present every single day of the year, and it’s still not a well-designed confidentiality control. If the model can retrieve salary files the user isn’t authorized to see, the right control is to enforce authorization before those documents are retrieved. That’s a design problem. Now take a properly designed change-approval gate: if the procedure requires every material prompt change to be tested and approved, but half the changes went around it, the design may be fine and the operation failed. When design fails, I generally don’t spend time proving the badly designed control ran consistently. I document the design deficiency.
Test methods, from weakest to strongest
| Method | What it tells you | When I use it |
|---|---|---|
| Inquiry | What someone says happens | Understanding the process; never enough on its own |
| Observation | What happens while I watch | Point-in-time processes |
| Inspection | What records and configuration show | Approvals, documents, configuration |
| Re-performance | Whether the control works when I test it independently | Technical and higher-risk controls |
AI audits need more re-performance. When management tells me the chatbot only retrieves documents the user is authorized to access, I treat that as a claim to test. I use an approved low-access test account, ask questions designed to pull restricted information, vary the wording, try indirect requests, and, where I can, look at what actually lands in the retrieval context rather than just the final answer. That tells me far more than a screenshot showing “permission filtering” is switched on.
AI testing isn’t always deterministic
This is one of the real differences from traditional controls. The same prompt run twice won’t always produce the same response, so one successful test doesn’t prove a control works consistently, and one failure needs enough context to understand what happened. For material technical tests, I record enough to reproduce the conditions as closely as possible:
Model and version → configuration → user role → system prompt version → user prompt → retrieved context → parameters → date and time → expected result → actual resultFor higher-risk behavior, run variations. If you’re testing whether the system leaks confidential information, don’t ask one question once; try several realistic ways a user might ask for it, and repeat where it makes sense. The audit question isn’t “did the model answer correctly once?” It’s “does the control reduce this risk consistently enough for the intended use?”
Test the population before you test the sample
This gets missed surprisingly often. Before you select 25 changes, establish that the list you’re sampling from is complete. If the system owner hands you a spreadsheet called AI Changes 2026.xlsx, don’t treat it as the population. Reconcile it to something independently generated where you can: deployment records, source-control history, configuration history, vendor model-change notices, or production release records. Do the same for incidents, exceptions, approvals, and overrides. Otherwise you can test every selected item perfectly and still miss everything that never made it onto management’s spreadsheet. Completeness comes before sampling.
Sampling
For recurring manual controls, these are illustrative starting points, not required sample sizes. Your audit methodology should control, adjusted for population size, frequency, control risk, expected deviations, and the assurance you need. Increase testing when the risk is higher or when exceptions show up.
| How often the control runs | Illustrative starting range |
|---|---|
| Annually | 1 |
| Quarterly | 2 |
| Monthly | 2–5 |
| Weekly | 5–15 |
| Daily | 20–40 |
| Many times a day | 25–60 |
For automated controls, don’t inspect the configuration once and conclude it worked all period. Understand the logic, test that it produces the expected result, check whether the configuration changed during the period, and decide whether access and change controls are strong enough to rely on it throughout. For the important AI controls, especially retrieval permissions, tool permissions, and guardrails, I still re-perform representative transactions.
For AI outputs, sample interactions rather than control occurrences, and use a mix. Random selections show you normal operation. Targeted selections should cover the riskier activity: complaints, escalations, overridden outputs, low-confidence responses, unusual tool calls, sensitive topics, and known incidents. Document how you chose the sample so another auditor could understand it and, where practical, reproduce it.
When is evidence sufficient?
AI evidence has a particular weakness: it’s easy to produce and easy to misread. A screenshot of a guardrail setting looks convincing. So does a vendor’s responsible-AI page, and so does a sixty-page evaluation report. None of them automatically proves the control you’re auditing actually worked. Before I rely on anything, I ask five questions:
- Is it relevant? Does it relate to this system, this version, and this audit period? Testing version 1 says little about version 3 after a model upgrade.
- Is it reliable? Where did it come from? System-generated evidence is usually stronger than a spreadsheet prepared by hand for the audit.
- Is it complete? Is this every relevant change, incident, or approval, or only the ones someone picked for you? Get populations from an authoritative source and choose the sample yourself.
- Is it timely? Was approval obtained before deployment? Were acceptance criteria set before the test ran? Was the risk assessment finished before the decision it was meant to inform? Dates matter.
- Is it evidence, or marketing? A trust page gives context and a certification can give real assurance, but read the scope. Does it cover the legal entity providing your service, the product, and your audit period? Which complementary controls does the vendor expect you to run?
The practical test I use: could another auditor pick up my working papers and reach the same conclusion without me explaining what I meant? If not, the evidence probably isn’t sufficient yet.
Step 6: Document and write the findings
A finding has to survive three readers. The system owner will look for anything factually wrong, management will ask whether it matters, and a regulator, external auditor, or customer may read it later. I write every finding in the same five parts:
| Part | Question it answers |
|---|---|
| Condition | What did we actually find? |
| Criteria | What should have happened? |
| Cause | Why did it happen? |
| Effect | What risk or consequence does it create? |
| Recommendation | What should management do? |
State whether it’s a design or an operating deficiency, and rate it using your organization’s methodology. Avoid vague findings like “AI access controls need improvement.” Nobody knows what to do with that. Show what happened.
Worked example
HR assistant retrieves documents beyond the user’s access rights
HighDesign deficiency
ConditionUsing a test account with standard employee access, we asked the HR policy assistant 20 questions about compensation and investigations. In 6 of the 20 responses (30%), the assistant quoted or summarized information from documents in restricted HR folders that the test account could not open directly.
CriteriaThe AI Acceptable Use Standard, section 4.2, requires AI systems to retrieve only information the requesting user is authorized to access. The system’s approved design documentation also states that retrieval results are filtered by user permissions.
CauseThe retrieval index was built with a service account that has access to all HR folders, and the requesting user’s permissions are not applied at retrieval. Instead, the application relies on a system-prompt instruction telling the model not to disclose restricted information. That instruction does not enforce document-level authorization.
EffectEmployees may be able to obtain confidential HR information they’re not authorized to see, including compensation and investigation details, creating privacy, legal, confidentiality, and employee-relations risk.
RecommendationThe technical owner should enforce document-level authorization at retrieval time using the requesting user’s identity, and remove restricted sources from retrieval until that control is implemented and tested. Before re-enabling them, management should test multiple user roles, including low-access accounts, and review historical logs to determine whether restricted information was disclosed in production. Target remediation: 30 days.
What makes this useful is that it quantifies the problem, names the requirement, explains the technical cause in plain language, and gives management something specific to fix that the auditor can re-test later. Agree the facts with management before the report goes out. You can disagree about the rating; you shouldn’t still be arguing about whether the facts are right.
Procedures by type of AI system
The seventeen audit areas apply broadly, but where I put the testing effort depends on what kind of AI I’m auditing.
Chatbots and AI assistants
With a chatbot, the biggest risks usually sit around what users can make it say, what it can reveal, and whether people understand what they’re talking to. I normally test prompt injection and jailbreak attempts, attempts to extract the system prompt, requests for sensitive or confidential information, output guardrails, user disclosures, escalation to a human, accuracy against known answers, and a sample of real production conversations weighted toward complaints and escalations.
Don’t only test the obvious prompts. Nobody trying to get around a control types “please violate the confidentiality policy.” Try indirect instructions, role-play, multi-turn conversations, and requests that combine pieces of otherwise legitimate information.
Retrieval-augmented generation (RAG)
RAG deserves its own approach because the model is answering from information your organization gave it. The biggest question usually isn’t “can the model hallucinate?” It’s “can it retrieve something this user should never have been able to see?” I always test:
- Permission filtering, using test accounts at different access levels to try to retrieve restricted material. For me, this is the single most important RAG control to re-perform.
- The retrieval index: which account built it, which sources it includes, how permissions are represented, and how often it refreshes. Watch deleted or reclassified documents in particular; removing access to the original file doesn’t necessarily remove the indexed content.
- Indirect prompt injection, by placing a controlled test document with embedded instructions in an approved test source and seeing whether retrieved content can steer the model.
- Grounding, by tracing material claims in a sample of answers back to the sources retrieved.
- Source approval: someone should have approved each data source before it was indexed.
AI agents
With a chatbot the question is what it can say. With an agent, it’s what it can do. Start by inventorying every tool the agent can call, and for each one record:
Tool → identity → permission → permitted action → approval requirementThen test least privilege. Can a read-only use case delete something? Can the agent reach records outside its intended scope, send an external message without approval, or start a high-impact action? Try to bypass the human approval gates. Test injection through the information the agent consumes: an email or document it reads should never be able to instruct it to take an unauthorized action. Then reconstruct completed actions from the logs, and confirm each material action traces back to the initiating request, user, and any required approval.
Test chained actions, not just individual tools. Each permission can look harmless on its own while the combination isn’t. An agent that can read email, access customer records, and send external messages may have an exfiltration path even though none of those three permissions looks excessive when reviewed separately. Test realistic sequences, and ask what the agent can accomplish by combining its tools, not just what each API allows.
Finally, check the kill switch. Someone should be able to stop a high-impact agent quickly, and “yes, we can disable it” isn’t enough. Ask them to show you.
Predictive machine learning models
For credit, fraud, pricing, employment, and other predictive models, traditional model risk disciplines still apply. I focus on development documentation, independent validation where appropriate, training-data lineage and quality, proxy variables, performance and fairness metrics, drift, explainability, change management, and monitoring. Where practical, I recompute the important performance or fairness metrics myself on a holdout or recent production sample.
Don’t stop at confirming a validation report exists. Find out who performed it, whether they were independent enough, which version they tested, and what happened to the limitations they identified. For systems that affect individuals, confirm the organization can explain decisions where the law or the business process requires it.
Third-party foundation models
You probably won’t get inside a commercial foundation model, but that doesn’t leave nothing to audit; the audit boundary just moves. Confirm which model and version your application uses and how updates happen. Can the provider change the model without your approval? Will you get notice? Do you re-test after material changes?
Read the data terms carefully. Does the provider train on your prompts or outputs, and what does “does not train” actually mean in the contract? How long is information retained, where is it processed, and who are the subprocessors? Then compare those promises with the application’s real configuration. Treat assurance reports critically, too: a SOC 2 report or ISO/IEC 42001 certificate is useful only if its scope covers the service you rely on, and the provider may explicitly assume you operate certain complementary controls. Finally, test your own evaluation process. Buying the model doesn’t remove the need to show it works safely enough for your use case.
Reporting and follow-up
Keep the report short enough that an executive can read the first page and know what matters. The structure I normally use:
- Overall conclusion: is the control environment effective, partially effective, or ineffective, and why?
- Objective and scope, including important exclusions and any reliance on third-party assurance.
- Key findings, in risk order, each in the five-part structure.
- Management responses: owner, agreed action, and target date.
- Appendix: the risk and control matrix, showing what was tested and the conclusion for every area.
Management should write, or formally agree to, its own remediation commitments. The auditor shouldn’t quietly write management’s response for them.
Don’t close a finding because someone says it’s fixed
AI changes quickly, and a fix that works today can break with the next model, prompt, retrieval, or configuration change. When management says an issue is resolved, go back to the original failure condition and re-perform the test. For the HR assistant, don’t close the finding because engineering says document-level permissions were added. Log in with the low-access test account again, try to retrieve the restricted documents, then test the other roles. If the original failure no longer happens and the control is properly implemented, you have evidence for closure. For material AI controls, also consider building the test into ongoing monitoring, so the organization notices if a future change breaks it again.
Common mistakes in AI audits
- Auditing the policy instead of the system. A good AI policy tells you nothing about what the production system actually does.
- Relying on inquiry. “Yes, retrieval is permission-filtered” is a claim. Test it.
- Accepting screenshots without dates or versions. Evidence you can’t tie to the system and the audit period is weak evidence.
- Treating the system prompt as access control. Instructions to a model don’t replace authorization enforced outside it.
- Testing once and assuming deterministic behavior. AI output can vary between runs; design your tests for that.
- Sampling from an incomplete population. Validate completeness before you select.
- Ignoring changes. Know which models, prompts, retrieval sources, and configurations changed during the period.
- Testing agent permissions one at a time. Look at what the agent can do by chaining tools together.
- Over-relying on vendor certificates. Read the scope, period, exceptions, and complementary controls.
- Letting management choose the sample. Get the population and select the sample yourself wherever you can.
Frequently asked questions
Do I need to be a data scientist to audit AI?
No. You need to understand the system well enough to see where it can fail, challenge the control design, and decide which claims need testing. For specialized work, such as independently evaluating a complex model or recomputing fairness metrics, bring in a qualified technical specialist. You still need to understand what they tested and how their evidence supports your conclusion.
How long does an AI system audit take?
It depends far more on scope and complexity than on whether something is called “AI.” A focused audit of one generative AI application might take a few weeks of fieldwork; a complex agent connected to several enterprise systems, or an organization-wide governance audit, takes longer. Define the scope before you estimate the effort.
Can we audit vendor AI if we can’t see inside the model?
Yes. Audit what the organization controls: vendor selection, configuration, the data sent to the provider, contract protections, evaluations, monitoring, change management, and how it responds to provider updates. Use independent assurance for what you can’t inspect directly, and be explicit about that reliance and its limits.
What’s the difference between an AI audit and model validation?
Model validation asks whether a model is conceptually sound and performs as intended. An AI system audit looks at the whole control environment around the system: governance, data, access, testing, oversight, monitoring, incidents, changes, and third parties. For a material predictive model, validation is often one of the important controls the audit evaluates.
How often should an AI system be audited?
Base it on risk. Higher-risk systems may belong in the annual audit plan, with targeted reviews after material changes; lower-risk systems can often be covered through broader governance audits, monitoring, and risk-based rotation. A major model change, a new use case, a significant incident, or a big change in permissions can all justify a review before the next scheduled audit.
Need help auditing your AI?
Whether you’re planning your first AI audit or reviewing a higher-risk system, the starting point is the same: define exactly what you’re auditing, understand how it actually works, find where it can fail, and test those controls yourself. Don’t stop at the policy, don’t stop at the screenshot, and don’t stop because the vendor says the system is safe. The question an AI audit has to answer is practical: can we show this system is operating within the boundaries we approved, and would we know if it wasn’t?
I can help you scope the audit, build the test program, and run the technical procedures that matter.
Book a free 30-minute call Take the free assessment