AI model governance is the application of model risk management discipline to artificial-intelligence systems that inform or produce compliance decisions. SR 11-7, the supervisory guidance on model risk, supplies the framework: sound development, independent validation, and documented oversight. An AI compliance tool is defensible when the institution can explain how the model works, show that it was validated independently, and identify the qualified person who reviewed and attested to the output. Accountability rests with that named person rather than with the model.
AI model governance is the application of model risk management discipline — sound development, independent validation, and documented oversight — to artificial-intelligence systems that inform or produce compliance decisions. In U.S. banking supervision the controlling framework is SR 11-7, the interagency Supervisory Guidance on Model Risk Management issued by the Federal Reserve and the OCC in 2011 (OCC Bulletin 2011-12). Its definition of a model reaches AI systems that score risk, draft SAR narratives, tune monitoring thresholds, or read regulatory change.
This guide covers what SR 11-7 requires, why it reaches AI compliance tools, what validation looks like in practice, the function of explainability and the audit trail, and the questions an institution puts to a vendor selling an AI compliance tool.
Supervisory posture toward AI in compliance
Compliance teams have adopted AI to score risk, draft SAR narratives, tune monitoring, and read regulatory change, and supervisory attention has followed. The examination question concerns how an institution controls the model it uses rather than whether it uses one. Documented control over the model is the condition on which supervisory acceptance turns; an AI treated as an unexamined vendor black box does not satisfy it.
What SR 11-7 is
SR 11-7 is the interagency Supervisory Guidance on Model Risk Management, issued by the Federal Reserve and the OCC in 2011 (OCC Bulletin 2011-12). It became the reference standard for how a regulated institution manages the risk of relying on a model. It defines a model broadly: a quantitative method that turns inputs into estimates, scores, or decisions. The definition is deliberately wide, and modern AI systems fall within it.
The guidance rests on the premise that models are useful and can also be wrong, so an institution that depends on a model must manage the risk that the model is wrong. The degree of rigor scales with the weight of the reliance and the consequence of an error.
The three elements SR 11-7 expects
| Element | What it requires |
|---|---|
| Development, implementation, and use | A sound design built on appropriate data and method, documented well enough that someone other than the builder can understand it, and used only for the purpose it was built for. |
| Validation | An effective, independent challenge to the model: is it conceptually sound, does it still perform, and do its outcomes hold up. Performed by people with distance from the developers. |
| Governance, policies, and controls | Clear ownership, written policies, an inventory of models, defined roles, and board and senior-management oversight of the whole thing. |
SR 11-7 is the supervisory backbone, and a newer companion speaks directly to generative tools. The NIST AI Risk Management Framework (AI RMF 1.0, 2023) and its Generative AI Profile (NIST-AI-600-1, July 2024) give a common vocabulary for AI-specific risks such as confabulation and information integrity. They are voluntary rather than binding, but they are the reference an examiner and a model-risk team increasingly expect to see mapped alongside SR 11-7 when the model is generative.
What validating an AI compliance model means
SR 11-7 frames validation in three parts, each of which maps onto an AI system:
- Conceptual soundness. Is the design fit for purpose and is the data behind it appropriate and representative? For an AI model, this includes the data it learned from and the assumptions baked into it.
- Ongoing monitoring. Does the model still perform as conditions change? Models drift as customer behavior, products, and typologies move, and monitoring is the control that detects the drift.
- Outcomes analysis. Do the results stand up when measured against benchmarks or back-tested against known answers? Outcomes analysis converts an asserted accuracy rate into documented evidence of accuracy.
Independence is the operative term. Validation performed by the same people who built the model, without separation or effective challenge, carries limited weight in examination.
Explainability and the audit trail
A model that produces a score or a narrative with no traceable reasoning gives a reviewer nothing to inspect, and an attribution of the outcome to the model itself does not identify a decision-maker. Two elements address this condition. Explainability means the institution can describe, in terms a reviewer follows, how the model reached its result. The audit trail means every step is logged, so the path from input to output can be reconstructed long after the fact. Together they make an otherwise opaque output inspectable.
The human in the loop
Under this control a qualified person reviews the AI's work and attests to it before that work is used: the model produces a draft and the reviewer decides whether to adopt it. The control matters because accountability cannot sit with a model. It sits with a named person who can be asked to explain a filing or a decision. A design that removes the human from that loop removes the element that makes the output defensible.
Questions to ask an AI compliance vendor
An institution evaluating a tool that places AI near a compliance decision puts the following questions to the vendor and weighs the answers:
- How does the AI reach a result, and is that reasoning auditable?
- How was the model validated, and by someone independent of the people who built it?
- What ongoing monitoring detects drift as conditions change?
- How is bias tested, and how often?
- Where in the workflow does a human review and attest before output is used?
- What audit trail could an examiner inspect, and how far back does it reach?
- Can the vendor produce documentation mapped to the elements of SR 11-7?
A vendor that cannot explain how its AI reaches a result leaves the institution without the documentation an examination requires. Governing an AI system on the same terms as any other model preserves the institution's ability to support the result.