Model governance that survives contact with the second model
Date Published

The first model through a governance process is always fine. Someone senior cares, the documentation is fresh, and the reviewers have time. The second is where the process is actually tested, and the fortieth is where it either works or quietly stops being followed.
That last outcome is the common one, and it rarely announces itself. Nobody decides to abandon governance. The reviews simply take longer, the queue lengthens, exceptions become routine, and eventually the register describes a process that no longer resembles what the organization does.
Thoroughness does not scale; by-products do
The instinct when review quality slips is to make review more thorough. This is precisely backwards. A more demanding process applied to a growing population produces a longer queue, and a longer queue produces exceptions, and exceptions are how governance becomes decorative.
What scales is not effort per model but the amount of evidence that arrives without anyone assembling it. Lineage captured as a by-product of training costs nothing per model and is always current. Lineage written up afterwards costs several days per model, is stale on arrival, and is the first thing dropped when the queue grows.
The same logic applies to challenger baselines. A challenger produced on request is a research exercise. A challenger produced automatically whenever a candidate model is registered is a property of the pipeline, and it is available at review time whether or not anyone remembered to ask.
Three things that have to hold
Lineage that is captured rather than written. Which data, which code, which parameters, recorded by the system that ran the training, not by the person who remembers having run it.
Challenger baselines that exist before anyone asks for them. If producing a comparison is a task, it will be skipped under pressure. If it is a side effect of registration, it is simply there.
An owner named on the model rather than in a slide. Not a team, not a function — a person, recorded against the artefact, who is accountable for its behaviour and empowered to withdraw it. Ownership recorded anywhere else evaporates at the first reorganisation.
Cadence is the actual deliverable
Once those three hold, something changes in the character of the process. Review stops being an event that has to be scheduled, resourced and survived, and becomes a cadence that runs whether or not anyone is paying attention to it.
That shift is what a regulator is actually looking for, though it is rarely how the requirement is written. The question behind the question is never "was this model reviewed" — it is "would you notice if this model started behaving differently, and how long would it take you". A quarterly cadence answers that. A thorough one-off review, however good, does not.
It also changes the economics. The organizations we have seen do this well are not spending more on governance than their peers. They are spending it earlier, on plumbing rather than on review capacity, and the review capacity they do have is aimed at the small number of models where judgement genuinely matters.