Silent failures tend to get promoted. Somewhere between the first time an agent does something nobody can explain and the conference panel about it, the problem moves upstairs and becomes a governance problem: a line in an AI risk register, a section in a responsible-use policy, a question for whoever owns model risk. The vocabulary moves with it, to oversight, auditability, controls.
That framing is right about where the problem ends up. I think it’s often wrong about where it starts, and for a small team the difference decides whether it can be dealt with at all.
Recently I wrote about the operator gap, the point where a workflow makes more decisions than the team running it can look at. That piece was about who goes missing. This one is about what that person would have produced if they’d been there, and why it turns out to be what governance asks for later.
Where it shows up first
Here is a composite, built to show the shape rather than report a case.
A B2B software company of a few dozen people runs a support agent. One of the things it can do is issue a goodwill credit when a customer was affected by a service incident, inside a limit the founder approved when the feature shipped. The instruction is plain: if the customer was affected by a confirmed incident, offer a credit.
The week after a bad outage, a customer on a renewal call mentions that a contact at another account got a credit for it. They didn’t. They wrote in the same morning, about the same problem.
The person who hears this is the head of support. She isn’t a governance function. She’s the operator, the one who owns the workflow the agent runs and who has to explain this to the customer before the week is out. Her question is why did it do that?, or more exactly, why for one of them and not the other.
She pulls both conversations. The first customer wrote, “your outage took down our dashboards.” The second wrote, “our dashboards have been blank since this morning, is something wrong on your end?” The agent connected the first message to the confirmed incident and offered a credit. It handled the second as a troubleshooting request, walked the customer through clearing a cache, and closed the ticket when the incident resolved and the dashboards came back. From where the agent sat, the second ticket went well.
By the afternoon her working theory is that the agent was keying on whether a customer named the incident rather than whether they were affected by it. That’s plausible. It’s also an inference, since nothing she can see shows which part of either message the agent actually weighed. She edits the instruction, credits the second customer by hand, and leaves a few lines in the support channel. Then the harder question arrives on its own: how many others? She can search that day’s tickets for “dashboard” and “blank,” which finds the customers who happened to use her guessed words and nobody else.
Now look at who else in the company knows any of this happened.
The founder sees credit spend as a line in the monthly numbers, and this month it came in lower than you’d expect after an outage, because the agent under-issued. Nothing in that view looks wrong. There’s no one whose job is AI policy. There is a policy page, written last spring to get through a security questionnaire.
A few weeks later a larger customer’s vendor review asks the company to describe how it oversees automated decisions that affect customer accounts, and how it would detect inconsistent treatment. That’s a governance question, arriving the way governance often reaches a company this size, from outside, on someone else’s schedule. Whatever honest answer exists to it lives in the head of support’s afternoon, if it lives anywhere.
Governance is a question asked of a history
So take the research question seriously. What would have to exist for that vendor review to get a real answer? Break the review into the questions it’s actually asking and look at what each one presupposes.
| What the review asks | What has to exist to answer it | Who was first in a position to produce it |
| ------------------------------------------------------------ | --------------------------------------------------------------------------------------- | ----------------------------------------------------------- |
| Are similar customers treated consistently? | Individual decisions you can compare, with what each one rested on | The operator, holding two tickets side by side |
| Can you show the basis for a specific decision? | What the decision was made from, with what was seen kept apart from what was guessed | The operator, at the start of the investigation |
| When something went wrong, what did you do, and did it hold? | The problem, the change, and what happened after it, attached to the decisions involved | The operator, the afternoon she edited the instruction |
| Would you notice if this behavior changed? | What this kind of decision looked like before the change | Anyone who looked at these decisions before something broke |
Every row comes back to the same object. What governance needs is the artifact the operator needed, many times over, kept, and comparable with one another.
I’ve argued here before that one decision isn’t behavior, and that behavior emerges once those decisions accumulate. Governance sits on the same line. Rather than a separate tier added on top of the stack, it’s a set of questions an organization asks about a population of decisions, often because someone outside it has started asking. The population is assembled from the thing a single operator needed on the first bad day: the decision, what it rested on, what was done about it, and whether that held.
Starting at governance fails a small team twice
A common order of operations is to establish the governance frame first and let it determine what gets captured. In regulated or model-risk-heavy organizations that’s often the right call, and governance shapes the system from the start, because there’s a person whose job is to ask and a budget for building the capture their questions require. On a small team, governance that arrives before any decision history fails in two separate ways.
There’s no one to own it. In the company above, the nearest thing to a policy owner is the founder, who has a dozen other jobs and saw nothing, because the only number in their view moved in the reassuring direction. A governance program presupposes someone whose job is to ask questions about the system. Small teams have someone whose job is to answer them, the operator, and she only asks when a customer forces the question.
The second failure goes deeper. There’s nothing to govern yet. “Credits must be applied consistently” is a sentence. To check it you need decisions you can compare, from before the outage and after it, from customers who named the incident and customers who didn’t. Writing that policy is reasonable. With no retained decisions behind it, though, it can’t be verified: the team has a standard and no history to hold it against. When the standard is finally tested, by the vendor review, the answer gets built from what does exist, which is execution records: logs are retained, a human reviews exceptions. I wrote in the accountability piece about how that kind of answer can satisfy a records requirement while saying nothing about why anything was chosen. The narrower point here is that a governance effort with no decision history to draw on ends up governing the logs, because the logs are what’s there.
There’s also a quieter cost to the framing. Once silent failures are filed as governance risk, they get triaged on the governance calendar, meaning the quarterly review and the next audit. The head of support’s afternoon doesn’t fit that calendar. Her customer is on the phone this week.
What the first bad day has to leave behind
If the operator’s investigation is where the governance record originates, it only becomes one if it survives. The afternoon above ends with an edited instruction and a few lines in a channel. A month later nobody could say which tickets were affected, what the working theory was, or whether the edit fixed it. I’ve written separately about what an investigation loses when nothing keeps it, so I’ll stay on the narrower question this piece raises: what does an operator’s answer have to contain so that, later, it can be counted?
It has to be about specific decisions, the two tickets and whichever others turn out to match, not about “the agent.” It has to keep its basis honest, separating what was observed from what the agent or the application reported about itself and from what the operator inferred. Her theory about named incidents is a good one, and on file it should still read as a theory. It has to record what was changed, and leave room for what happened next. And it has to be comparable with the next case, so that the next inconsistent credit is recognizably the same kind of event and not a fresh mystery.
Each of those properties earns its place by making the operator’s own next investigation faster, before anyone above her asks for it. That’s the alignment I find most interesting: the record that serves her the next time a customer calls is the same record a review would need a quarter from now. Nobody has to be talked into keeping it for compliance reasons, because it already pays for itself before compliance ever asks.
Order matters
Silent failures do become governance problems eventually. In larger organizations, and as the external questions get sharper, the policy owner and the audit and the review will all be real, and they’ll ask about populations of decisions over time. Nothing here replaces them. The operator’s record is the material they work from, and the argument is about sequence: that material has to exist before anyone can govern with it.
The operator’s question comes first in time, and it’s the one that produces the object. One investigation produces a record. Records accumulate into a history. A history is what governance asks its questions of. Run it in the other order and you get policy with no evidence underneath it, which on a small team means a policy page and a search box.
My expectation is that the teams that end up with credible answers for their first vendor review or their first audit won’t be the ones that started with a governance program. They’ll be the ones that took why did it do that? seriously the first time a customer asked, and kept what they found. Governance asks what a system has been doing. The only honest way to answer is to have asked, one decision at a time, why it did that.
That first question is the one I’m building Nalyqor to answer, and to keep.

