Commentary

Model Validation's Four Questions, Asked of Agentic AI

C. Erik Larson · October 9, 2026

The revised interagency model risk guidance issued in April (SR 26-2, OCC Bulletin 2026-13) replaces the prior guidance (SR 11-7, OCC Bulletin 2011-12) with a materiality-based approach, and it places generative and agentic AI outside its scope. The agencies have said a separate request for information on banks' use of AI will follow.

That has been read in some quarters as suggesting that the model risk management framework is not adequate for these systems. I read it somewhat differently.

Four questions

For me, model validation has always come down to four questions, recognizable to some degree in supervisory guidance dating back even to OCC 2000-16:

  1. What business problem is the model intended to solve?
  2. Why is this design appropriate for that objective, relative to the alternatives that could have been chosen (i.e. what is the conceptual soundness)?
  3. Has the model been developed and implemented consistently with that design (i.e. what is the developmental evidence)?
  4. What evidence, from ex-ante benchmarking or ex-post outcomes analysis, shows that it performs acceptably, and against what thresholds (i.e. what is the performance analysis)?

These questions were never meant to be confined to a single model. When my colleagues and I examined credit decisioning systems, our conclusion was that the thing being validated was not the scorecard alone. The interesting business decision usually was based upon a combination of scorecard, credit loss forecast, capital allocation and risk-based pricing models, supplemented by policy overlays and override behavior, all run up against a risk-adjusted return objective. Each of those models deserved validation in its own right, but each was an input to something larger.

Indeed, Question 1 could not be answered until the bank stated what the system as a whole was attempting to achieve, or optimize. Similarly, Question 4 could not be answered until that objective had a metric, for instance SVA, which would mean comparing the return expected at origination against the return actually realized.

The important point is that nothing in the four questions presumes a technology, and nothing confines them to a single component. The questions apply to a standalone logistic regression, a foundation model, or, let me claim, an AI agent that plans and acts.

What is not new

Asked of agentic systems, the four questions hold. What changes is how difficult they are to answer, and some of what is presented as a new requirement is not new at all.

A recent IBM Promontory paper holds that validating what an agent can do should take precedence over validating how accurately it thinks. But asking whether an agent can only do what its design permits is asking whether it was built the way it was specified, with the subject being its permissions rather than its code. That is Question 3. The permissioning problem is not a new requirement; it is an old one that has become harder to discharge.

What genuinely changes

My view is that two things do genuinely change. First, Question 3 changes shape. With a conventional model you can confirm that what was built matches what was specified: the code implements the documented method, the data arrives as designed. An agent chooses its own sequence of steps as it runs, so there is no single procedure to check it against. What has to be established instead is that the limits hold whatever sequence it chooses. The test is not that the agent took the intended path, but that no path available to it leads somewhere it should not go.

Second, Question 4 must apply to the route taken, and not only the result. An agent can arrive at an acceptable answer by an unacceptable route, and outcomes analysis that ignores the route will pass systems that should fail. That, rather than novelty, is the argument for a decision ledger.

Containment is not validation

Containment (blocklists, rate limits, fail-secure defaults) is not a substitute for validation. The override policies in the credit decisioning systems were guardrails too, and supervisors asked for override tracking precisely because the existence of a guardrail was never evidence that the system performed. A design is not appropriate for its objective unless it is appropriate given what happens when it is wrong. Reversibility and how far the damage spreads are design properties, which is to say they belong to Question 2.

And Question 2 is the one most agentic deployments cannot readily answer: Why specify an agent rather than a deterministic workflow? Why support model inference at this step rather than use a lookup, a rule, or a computation? In an interesting Substack post, Agus Sudjianto describes a "verification" hierarchy, which makes the same point from the design side. This is where some real work has to be done, and I expect it will require significant changes to how first line developers interact with second line challengers in the validation process. More to come on this topic.

In closing

I think that the frameworks are in better shape than much of the commentary suggests. What is missing sits in the deployments, most of which cannot state an objective in measurable terms or produce evidence against it. That is a first-line deficiency rather than a gap in model risk management. And while the current guidance does not cover these systems, the principles behind it do. At least that is my reading, though not necessarily the agencies'.

← Back to Thought Leadership