# AI Governance Checklist: 40 Items to Audit Your Program Against

_A 40-item AI governance checklist organized by the four NIST AI RMF functions. Score each item in place, partial, or absent to find where your program is thin._

An AI governance checklist is a gap-analysis instrument. You take a list of what a working [AI governance](/ai-governance) program contains, hold it against what your organization actually has, and mark the difference. It certifies nothing, and it does not replace an AI governance framework like the [NIST AI RMF](/nist-ai-rmf) or [ISO 42001](/iso-42001). Its advantage over both is speed: in an afternoon, it shows a compliance lead, a CISO, or an internal auditor where the program is thin before someone outside the company finds out.

**The short version: the 40 items below cover the four NIST AI RMF functions (Govern, Map, Measure, Manage). Most programs we see score well on the first two and thin out on the last two, because Measure and Manage require something running against live traffic rather than a document in a shared drive.**

## How to use this AI governance checklist

Score each item one of three ways: in place, partial, or absent. Resist the urge to collapse this to yes and no. Binary scoring hides the most common state, which is "we have a policy that says this, and nobody has checked whether it happens." That middle column is the point of the exercise.

An item is in place only when you can produce evidence for it on request: a dated document, a log, a ticket, a report that reached a named person. Partial means the intent exists and the evidence does not. Absent means nobody owns it.

Work through the list with the people who would produce the evidence, not the people who wrote the policy. Two hours with the platform team and the model owners will surface more partials than a week with the governance committee.

One note on wording. NIST's own language for many of these items is "controls." We use guardrails, safeguards, and measures throughout and mean the same thing.

## Govern: policy, roles, and accountability

Govern is the function everything else hangs on. It is also the one most programs can evidence, because most of it is documents, and documents are what governance teams produce well. Watch items 6 through 10, where the document is necessary but the evidence has to be a record of something that happened.

- **1. Written AI governance policy**, approved by leadership, versioned, with a named owner and a review date.
- **2. Defined roles and decision rights** for AI: who approves a deployment, who can pause one, who signs off on accepting a risk.
- **3. A named accountable executive** for the AI program, plus a named owner for every system in the inventory.
- **4. AI acceptable use policy** for employees, covering approved tools, prohibited data classes, and what happens on violation ([template here](/post/ai-acceptable-use-policy-template)).
- **5. Risk appetite statement** for AI that says which use cases are prohibited outright, which need review, and which are pre-approved.
- **6. Role-based training** delivered and logged: general awareness for staff, deeper material for builders, reviewers, and approvers.
- **7. Third-party and vendor requirements** written into procurement: model provenance, data handling, change notification, right to audit.
- **8. Board or executive reporting** on AI risk at a fixed cadence, with the last report on file.
- **9. A documented exceptions process**, including who can grant one, for how long, and where the register lives.
- **10. Governance program review** on a schedule, with evidence that the last review changed something.

## Map: what you run, what it touches, what applies to it

Map is where you find out whether the program describes the AI you actually run or the AI you approved. The two diverge fast. The inventory comes first because every item after it is only as good as the inventory it draws from.

- **11. AI inventory** listing every model, agent, and embedded AI feature in use, including SaaS features your vendors switched on and tools staff adopted without asking ([how to find shadow AI](/post/how-to-detect-shadow-ai)).
- **12. Inventory refresh process** that catches new systems within a defined window, as opposed to an annual survey.
- **13. Risk-tier classification** for each use case (for example prohibited, high, limited, minimal), with the criteria written down.
- **14. Intended use and known limits** documented for each system: what it is for, what it is not for, and where it is known to fail.
- **15. Data lineage** for each system: training sources, retrieval sources, and what leaves the organization at inference time.
- **16. Data class mapping** showing which systems touch PII, PHI, payment data, privileged material, or trade secrets.
- **17. Stakeholder and impact assessment** for high-tier use cases, covering who is affected by a wrong output and how they would find out.
- **18. Regulatory mapping** per system: EU AI Act tier, applicable state laws (Colorado, Texas, California), and sector rules.
- **19. Sector-specific obligations** identified and assigned. For carriers, that means the [NAIC model bulletin](/naic-model-bulletin-ai) and whichever state adoptions apply.
- **20. Agent-specific mapping** for systems that take actions: which tools they can call, which systems they can write to, and which other agents they can delegate to ([more on agent governance](/ai-agent-governance)).

## Measure: testing, thresholds, and evidence

This is where partial starts to dominate. Nearly every program tests before launch. Far fewer define what passing looks like in advance, and fewer still keep testing once the system is live.

- **21. Pre-deployment evaluation** against a defined test set, with results stored alongside the version that was tested.
- **22. Acceptance thresholds** set before testing begins, so the result decides the outcome and not the other way around.
- **23. Bias and fairness testing** for any system that affects individuals, using the protected classes relevant to your jurisdiction.
- **24. Red teaming** of high-tier systems, with findings tracked to closure.
- **25. Prompt-injection and jailbreak testing** for any LLM system that reads untrusted content or holds tool access.
- **26. Drift monitoring** in production, with a threshold that triggers review and a record of the last time it fired.
- **27. Human review sampling** of live outputs at a set rate, with reviewer findings logged and fed back to the owner.
- **28. Performance metrics reported** to the system owner and the governance function on a schedule.
- **29. Re-evaluation triggers** defined: model version change, prompt change, data source change, vendor notification.
- **30. Evaluation methods documented** well enough that a second team could reproduce the result.

## Manage: guardrails, response, and lifecycle

Manage is where the gaps concentrate, and the reason is structural. Every item here requires something to exist at the point where the AI system meets the user or the data. A policy can require a kill switch, but a policy cannot stop a request.

- **31. Runtime guardrails at the point of use** that apply policy to each request (data leaving, prohibited uses, output filtering) rather than relying on training and good intent.
- **32. Scoped permissions for agents**, so each agent has its own identity and the narrowest credentials its task needs.
- **33. Kill switch and rollback** for every production system, tested within the last review period.
- **34. AI incident response plan** with AI-specific triggers, severity levels, and a named on-call ([plan template](/post/ai-incident-response-plan)).
- **35. Audit trail per request** capturing prompt, response, policy decision, user, and system version, retained for the required period ([what belongs in it](/ai-audit-trail)).
- **36. Vendor change notification** handled: a process for what happens when a provider swaps a model version underneath you.
- **37. Escalation paths** from the system to a human, with the triggering conditions built into the system rather than described in the manual.
- **38. Decommissioning procedure** that revokes credentials, archives records, and removes the system from the inventory.
- **39. Periodic review cadence** per risk tier (for example quarterly for high, annually for minimal), with the last review dated.
- **40. Remediation tracking** for every gap this checklist surfaces, with an owner and a date.

## Using this as an AI audit checklist

Internal audit and regulators read the same 40 items differently from the people who built the program. For them, every item is a request for evidence, and "documented" has a specific meaning: a dated artifact that existed before the audit started, produced by the process the policy describes. AI governance documentation written the week before the review counts as absent, and experienced auditors can tell.

Three habits turn the list into an AI audit checklist:

- **Ask for the artifact.** The question "do you monitor for drift?" gets a yes. Asking for the last drift alert and what happened after it gets the actual state of the program.
- **Sample.** Pick three systems from the inventory at random and walk items 13 through 39 for each. Consistency across the sample matters more than the best-documented system.
- **Trace one request end to end.** Take a single production interaction and follow it: which policy applied, what the guardrail decided, where the log lives, who could have reviewed it. Wherever the trace breaks, that is the finding.

For a program working toward ISO 42001 certification, the items line up with the standard's Annex A groupings closely enough to serve as a pre-assessment.

## What most programs are missing

Across the programs we have reviewed, the scoring pattern repeats. Govern and Map score well on paper. Measure comes back mostly partial: testing happened once, the threshold was set after the fact, and monitoring lives on a dashboard nobody is assigned to watch. Manage is where the absents cluster, because items 31 through 38 only exist if something is running in the request path, and most governance programs were assembled by policy teams, who do not deploy gateways.

That runtime layer is where Swept sits. The [gateway](/offering/governance) applies policy to each request as it happens, scopes what agents can reach, and writes the per-request audit trail as a byproduct of operating, which covers most of the Manage column and the production half of Measure. If your checklist comes back strong on the first twenty items and thin on the last twenty, that is the conversation worth having. [Talk to us](/offering/governance), or start with the [AI governance maturity model](/post/ai-governance-maturity-model) to see where the gaps usually sit by stage.