Skip to content
Qofi
← all insights
EssayAug 2026 · 4 min read

when the model is not a model

Model risk frameworks assume versioned code, specified inputs, and replicable outputs. An agent has none of these. The institutions that translate the framework’s intent now — rather than force-fitting its ritual — will end up writing the standard everyone else adopts.

The first agentic system to reach a bank’s validation queue arrives with the same intake template as everything before it, and every field resists filling in. The template asks for the input space; the honest entry is “any document an employee chooses to paste, plus whatever the agent retrieves on its own.” It asks for the output specification, and the answer is a paragraph, different each time. It asks what constitutes a material change, and nobody in the room can say whether a reworded system prompt qualifies, or a vendor’s quiet weight update, or a new tool added to the agent’s reach. The team is not being difficult. The form was written for a thing this is not.

That form encodes fifteen years of model risk practice — the SR 11-7 lineage in the US, its supervisory cousins elsewhere — and the practice worked because its subject held still. A pricing model or a PD model is a bounded artifact: defined inputs, a processing component someone in-house built and can explain, outputs you can score against realized outcomes. Versioned code, so the thing you validated is the thing in production. Replicable, so a finding can be reproduced by the person challenging it. And changes arrive through a release process the institution controls, which is what makes “revalidate on material change” an enforceable sentence rather than a wish.

the assumptions break one by one

An agentic system violates each of these quietly. Behavior changes through prompts, not code — a product owner rewording an instruction block is a meaningful model revision that never touches a repository or triggers a change ticket. The vendor can move the weights underneath you, so the system you validated in March is not the system running in June, and no gate you own was involved. The output is a distribution, not a value: a test case that passes tells you about one sample from that distribution, and the same test tomorrow may tell you something else. And the framework’s first pillar, conceptual soundness, asks the validator to judge the theory and design of a system nobody in-house trained, on data nobody in-house has seen, with a parameter count that makes “inspect the processing component” a category error.

The tempting responses are both wrong. One is to force-fit the ritual — run the template anyway, file a validation report whose every section is an approximation, and let the paperwork claim more certainty than anyone has. The other is to declare the framework obsolete and wait for supervisors to write a new one. The first produces documents that will not survive contact with an examiner who asks a second question. The second leaves the institution running consequential systems under no discipline at all.

A validation you cannot rerun is an opinion with a timestamp.

an inventory is not a row count

Start with the question every model risk office asks first: what goes in the inventory? Is each prompt a model? Each agent? Each combination of base model, retrieval corpus, and tool configuration? Counting artifacts misses the point, because the artifact is not where the risk lives. The unit that matters is the deployed system in its context of use — this model, with these instructions, these tools, this data, feeding this decision. Two agents on the same weights, one drafting marketing copy and one screening transactions, are not two rows of the same model; they are different risks that happen to share a component. The honest inventory is a materiality tiering for stochastic systems: what decisions the system touches, how much autonomy it has been given, what happens on the day it is confidently wrong. That is harder than counting, and it is the version an examiner can actually use.

monitoring becomes the load-bearing wall

The deeper shift is in where the assurance weight sits. The classical framework front-loads it: exhaustive pre-deployment validation, then annual review, with ongoing monitoring as the dutiful third pillar. When the input space is unbounded and the output is a distribution, pre-deployment testing cannot be exhaustive even in principle — it can establish that the system behaves acceptably on a curated evaluation set, and nothing stronger. So the weight moves to production: outcome monitoring against ground truth where it exists, sampled human review where it does not, evaluation suites rerun on every upstream change — prompt, corpus, tool, or vendor version — and challenge that happens continuously rather than on an anniversary. Monitoring stops being the afterthought and becomes the load-bearing wall. The eval set becomes what the validation report used to be: the durable, versioned artifact you can point to.

What survives the translation is the framework’s intent, and it survives almost intact. Effective challenge by people with standing to say no. Independence between the builders and the reviewers. Documented intended use, and documented limitations that someone accountable has read. None of that depends on determinism; all of it transfers.

Supervisors have so far mostly said the existing guidance applies, without saying how — which is a deferral, not an answer. It means the first credible translation will be written inside an institution, not a regulator, and the institutions that write theirs now will be defending a considered position while everyone else is defending a template. The framework’s letter was written for models. Its intent was written for exactly this.

← all insightsstart a conversation →