Executive Leadership
When the Model Changes Underneath You: Change Control for an AI Agent Fleet
Your engineers need two approvals to change one line of billing code. Somebody swapped the brain inside forty agents last Tuesday with a config edit and a shrug.
Every few months a model provider ships a new version. It is cheaper, or it scores better on some benchmark with a name like a license plate, and within a week somebody on your team has pointed a production agent at it because the old one felt slow.
That is a change to how your company works. In most companies I see, it goes through no review at all.
I run my own fleet of agents across content and operations work, and I have rolled a new model version across most of it more than once this year. The first time, I treated it as a routine upgrade. Two agents got noticeably better. One got worse in a way that took eleven days to notice, because it kept producing confident, well-formatted work that happened to skip a verification step the old model had always performed without being told.
Nobody had written that step down. The old model just did it.
The quietest change in the building
Software teams have spent twenty years building change control for code. Pull requests, staging environments, feature flags, release notes, a rollback button that somebody has actually pressed. It is boring infrastructure and it catches an enormous amount of damage before customers see it.
Model swaps slip past all of it. The agent’s code and prompts stay exactly the same, and its tests (if it has any) usually check that the output parses, which a new model will happily satisfy while reasoning quite differently about what to put inside the brackets. From the point of view of the deployment pipeline, nothing happened. From the point of view of the customer who got a refund denied that would have been approved last week, a great deal happened.
So the first governance decision is a definitional one. A model version change is a production change, with an owner and a way back.
What I require before a model moves
This list has grown every time something went wrong, which is how all good checklists grow.
- A frozen evaluation set per agent. Fifty to two hundred real past tasks, with the output a human judged correct, stored somewhere the agent cannot edit. The new model runs the whole set before it touches live work, and the owner reads the diffs. Reading the scores alone is how the eleven-day problem happens.
- A written list of behaviors the agent performs without being asked. This one is tedious and it is the one that would have saved me. Watch the current agent work for a day and write down every check it runs and every request it declines on its own. Then look for each one in the new model’s output.
- A cost and latency comparison, because a cheaper model that needs three retries is more expensive.
- A named rollback owner who can revert the fleet without a meeting.
Notice what is absent. There is no requirement that the new model win on a public benchmark. I have stopped caring about those for operational work. They measure something real, but that something is rarely the thing my agents do all day.
Pick the canary on purpose
After the evaluation set, one agent goes first on live work for at least a week. Most teams pick the easiest agent, the one with low stakes and simple tasks, because it feels safe.
That is backwards. The easy agent will pass on almost any model, so its canary week tells you nothing, and you will roll the upgrade to the harder agents with false confidence. I pick the agent whose work is most dependent on judgment and whose errors are still recoverable, which in practice is usually something in tier two of the delegation of authority matrix, where mistakes are visible to customers but can still be corrected. If the new model holds up there, the rest of the fleet is a much smaller bet.
Agents with authority to commit the company never go first. They go last, after everyone else has run clean for two weeks.
Rollback has a clock
Keep the previous model version configured and callable for thirty days after the last agent moves. Providers retire old versions on their own schedule, sometimes with less notice than you would like, so the thirty days is a floor you control inside a ceiling you do not.
I also set a trigger in advance. If the error rate a reviewer finds in daily samples rises by more than a set margin over the old baseline, the rollback owner reverts first and investigates second. Deciding the number before the upgrade matters, because after the upgrade everybody who championed it will have a reason why this particular spike does not count.
The last number
Write down how well the old model did on every evaluation set before the vendor retires it, because one day that log will be the only surviving witness to how good your operations used to be, and someone in a quarterly review will insist they were always this way.
What the board should see
Directors should know the company has a process for changing models, since a single upgrade can alter the behavior of every customer-facing agent at once, and that is the kind of correlated operational risk audit committees exist to ask about. It fits naturally alongside the concentration questions in what boards should ask about AI.
Once a quarter, the report I would want as a director is short. Which models the fleet runs on and which version changes happened, with any rollback explained in a sentence. One page. If the answer to the second question is “we are not sure,” that is the finding.
The CEO’s job here is small and specific.
Declare that model changes are production changes. Then make sure the next upgrade someone proposes on a Friday afternoon waits until Monday, when there is time to watch it land and a person awake to pull it back if the eleven-day problem shows up on day two.
This article is part of the Executive Leadership cluster, focused on board governance and the operating discipline required to run AI systems responsibly at the executive level.