Regulatory

AI agents are out of scope on both sides of the Atlantic

The US has taken AI agents out of its model-risk guidance. The UK has kept them in a framework built for something else. Two central banks have now said the human in the loop does not scale. The gap that opens is the same in both jurisdictions, and it is the one nobody has written down.

Summary. Does model-risk regulation cover AI agents? In the US, no: SR 26-2 removed generative and agentic AI from scope in April. In the UK, nominally yes, under SS1/23, but that framework validates a fixed model, not an agent that composes tools at runtime. Neither says what evidence a firm should hold that an agent’s controls are operating in production, and that evidence is the same on both sides of the Atlantic.

Key points

  • SR 26-2 (Fed, OCC, FDIC, 17 April 2026) explicitly excludes generative and agentic AI as “novel and rapidly evolving” and leaves them to broader risk management practices. Examiners are still asking.
  • SS1/23 applies to AI and the PRA has named it a 2026 priority, but its unit of analysis is a model: validated once, monitored for drift. An agent is not that object.
  • The Bank of England (June) and the New York Fed (September) have both said, in different words, that the supervision-to-output ratio existing controls assume will not hold.
  • The evidence that matters comes from the running system, not the guidance: what the agent read, what it could call, what it changed, and whether the limits held. It is testable today in either regulatory vocabulary.

The US stepped back

On 17 April the Federal Reserve, the OCC and the FDIC issued SR 26-2, the first rewrite of model-risk guidance since SR 11-7 in 2011. It is shorter, more principles-based, and narrows what counts as a model. Buried in the scope section is the line that matters for anyone running agents inside a bank: generative and agentic AI are “novel and rapidly evolving” and are not within the scope of the guidance. Institutions are directed to govern them through their broader risk-management practices instead.

Governor Bowman confirmed the boundary in a speech on 1 May: the revised guidance applies narrowly to traditional models and basic AI applications. The Fed is now soliciting input on how agentic systems should be governed.

Read one way, this is honest. Regulators had spent fifteen years watching banks stretch a framework written for pricing models over things it was never designed for, and chose not to stretch it over the newest one. Read another way, it is a retreat at the wrong moment. The most consequential AI in the institution, the kind that executes rather than recommends, is now the only kind without a supervisory template. Examiners have not stopped asking how it is governed; they have simply stopped telling you what a good answer looks like.

The US commentary since April has been almost entirely from vendors explaining the carve-out to their own customers. The practitioner question underneath it has not been answered: if effective challenge no longer formally applies to the agent, what does?

The UK leaned in, with the wrong tool

The UK went the other way. SS1/23, the PRA’s supervisory statement on model risk management, applies to AI, and the PRA has named AI a supervisory priority for 2026. A UK bank running an agent is expected to bring it inside the model-risk framework, with a model owner, a tier, validation, and ongoing monitoring proportionate to materiality.

The problem is not the intent. It is the unit of analysis. SS1/23 governs a model: an object with defined inputs, a defined transformation, and defined outputs, which you validate before use and then watch for drift. That description fits a credit-scoring model or a VaR engine. It does not fit an agent.

An agent is a loop, not a function. On each turn it reads context, decides which of its tools to call, calls them, reads the result, and decides again. The transformation is not fixed; it is composed at runtime from the tools the agent has been given and the data it can reach. Two runs on the same input can take different paths and both be correct. Validating the model inside the loop tells you almost nothing about what the loop will do, because the behaviour that matters lives in the permissions, the tool definitions, the retry logic and the fallback state, none of which are the model and none of which SS1/23 asks you to test.

So the UK firm that has dutifully registered its agent in the model inventory, tiered it, and validated the underlying LLM has satisfied the letter of the framework and evidenced very little about the thing that will actually cause a loss. The US firm has no framework and the same gap. They have arrived at the same place by opposite routes.

Two central banks, one ratio

If the frameworks are silent, the central bankers are not. Two speeches this year, one on each side of the Atlantic, have said the same thing from different directions.

On 30 June, Sarah Breeden, Deputy Governor of the Bank of England, told the ECB’s Sintra forum that the Bank’s frameworks “were not built to contemplate autonomous agents”, and that relying on a human in the loop for every agent action is unlikely to be realistic. The reasoning is arithmetic: an agent that makes decisions faster than a person can read them cannot be supervised one decision at a time. A reviewer who is notionally in the loop and practically not is worse than no reviewer, because the control is recorded as present.

On 24 September, Mihaela Nistor, Chief Risk Officer of the Federal Reserve Bank of New York, approached it from the other end. Speaking in a personal capacity at Risk Live, she described autonomous agents as one of three “acceleration pathways” risk functions should stress-test, because they change the ratio between human supervision and output that governance and control structures currently assume stays roughly constant. Then she went further. The junior work that agents absorb, building the reports, walking the process maps, was the training mechanism for judgment. Automate it and the people who would have caught the error in seven years’ time are never formed. In her words, that erosion will not trip a control.

Put the two together and the human-in-the-loop control fails twice. Breeden: the human cannot keep up today. Nistor: the human who could have kept up will not exist tomorrow. Both SS1/23 and SM&CR assume a competent, accountable individual at the point of review. Both central banks have now said that assumption is wearing out.

The gap both left

Here is what neither regulator has said, and what neither speech quite reached: what a firm should be able to show that its agent’s controls are operating, in production, right now.

Not the policy. Every firm has the policy; it says the agent has a budget, a permitted tool set, a fallback state and an escalation route. Not the pre-deployment evaluation either; that tests what the agent did in a sandbox with the tools it had that day. The evidence gap is about the running platform: the configuration that is live, the permissions that are actually enforced at the tool boundary, the log that shows what the agent read before it acted, and the proof that when a limit was hit, the agent stopped.

The US has delegated this to broader practice without saying what practice. The UK has delegated it to a framework that asks for validation of the model and monitoring of its outputs, which is necessary and does not touch it. In both, the honest answer to the examiner’s question “how do you know the agent is doing what the policy says?” is currently: we tested it before go-live and we have a human watching. And both central banks have just explained why the second half of that answer is not a control.

This is not a criticism of either regulator. Model-risk guidance was never the right home for a runtime question, and the people who write supervisory statements are not the people who can tell you whether a tool boundary is enforced in a specific deployment. But the effect is that a firm operating in both jurisdictions, and most of the ones that matter do, faces the same unwritten expectation twice, in two vocabularies, with no template on either side.

What the evidence looks like, in either vocabulary

The good news is that the evidence does not depend on the regulator. It comes from the running system, and it is the same system whether the examiner works for the OCC or the PRA. Four things, each of which can be tested by someone outside the team that built the agent.

What it read. For any action the agent took, the context it was given at that moment: the documents, the records, the prior turns. If that cannot be reconstructed after the fact, no reviewer, human or otherwise, can say whether the action was reasonable. This is a data-layer property, not a model property; most platforms overwrite the state the agent acted on and lose it.

What it could call. The tool set and permissions in force at the time, enforced at the boundary between the agent and the tool, not described in the prompt. A prompt that says “do not transfer funds” is an instruction; a tool interface that has no transfer capability is a control. Testing this means attempting the prohibited action and showing it fails.

What it changed. Every write the agent made, as an append-only record with the time the change was made and the time it took effect. Corrections are new entries, not overwrites. This is what lets a firm answer, months later, what the agent did and what it believed at the time.

Whether the limits held. The budget, the retry cap, the escalation trigger and the fallback state, each exercised in production or in a production-equivalent environment, with the result recorded. A fallback that has never been triggered is a hypothesis.

In US terms this is effective challenge applied to the deployment rather than the model, and it produces exactly the artefacts an examiner asks for under broader risk management practices. In UK terms it is the evidence SS1/23’s ongoing-monitoring principle implies but does not specify, and it is what a Senior Manager needs to hold before signing a statement of responsibility that covers an agent. Same four things. Different cover sheet.

What to do with this

If you run agents in a US bank, the carve-out means the burden of designing the governance has moved to you, and the Fed is asking for input on what it should look like. Send them the four tests above; they are what your examiner will accept in the meantime.

If you run agents in a UK firm, the model inventory entry is necessary and not sufficient. Before the Senior Manager signs, ask for the four artefacts, and if the platform team cannot produce them, that is the finding.

If you run agents in both, which is most of the firms this is written for, the same evidence pack serves both regulators. That is the one advantage of a gap nobody has filled: you get to fill it once.

The FSB’s final report on responsible AI adoption is expected in October and will be the first global text to address agent oversight directly. I will go through it against these four tests when it lands.

Questions this piece answers

Does SR 26-2 apply to AI agents?

No. SR 26-2 explicitly excludes generative and agentic AI from its scope and directs institutions to govern them through broader risk-management practices. Examiners still expect that governance to exist.

Does SS1/23 apply to AI agents?

Yes, in that the PRA applies its model-risk principles to AI. But SS1/23 is written for models with fixed inputs and outputs, and does not address the runtime behaviour of an agent that composes tools dynamically.

What should a firm be able to show a supervisor about an AI agent?

Four things from the running system: what the agent read before acting, what tools and permissions it had enforced at the boundary, an append-only record of what it changed, and evidence that its limits and fallbacks have been exercised.

Is a human in the loop an adequate control for AI agents?

Both the Bank of England (June 2026) and the New York Fed’s Chief Risk Officer (September 2026, personal view) have said the supervision-to-output ratio it assumes will not hold at agent scale.

Sources


Antony Coppellotti is founder and CTO of Gordion Solutions, which puts independent, tested controls around AI agents in regulated firms and builds the systems that need them.

Take the Readiness Check · See what a Health Check covers · Get in touch · info@gordionsolutions.co.uk