The Last Reviewer — Part 2: Who Polices the Machines?
The Last Reviewer — Part 2: Who Polices the Machines?
Published: March 2026
Part 1: When Humans Can't Keep Up with Their Own Machines
The governance instinct is right. The sequencing is wrong. And there's a harder conversation underneath both.
Last month's article on the software factory got more traction than I expected — and more pushback than I expected, which is a better sign.
The concerns that came back weren't the reflexive ones. Nobody seriously argued that human review should stay exactly as it is. The sharpest critics were asking a more interesting question: yes, but how do you build the safety net before you let go of the trapeze? Several people raised the same framework independently — call it verification-before-velocity. Build the assurance infrastructure first. Then open the throttle.
It's a responsible position. I want to engage it properly. Because I think it's right about the principle and quietly wrong about what follows from it — and underneath that disagreement is a structural confusion about what human review in governance systems actually is.
The Prerequisite That Never Gets Met
The "build assurance first" argument sounds like discipline. In practice, it behaves like a lock with no key.
Here's the mechanics: verification infrastructure doesn't get built in a vacuum. It gets built because something is producing output that needs verifying. No generation means no urgency to build assurance. No urgency means no budget. No budget means no assurance. No assurance means no generation. You've drawn a dependency loop and called it a policy.
I've heard the same logic at DHL roughly four hundred times, in a different register: "We'll migrate to the cloud once we've documented all our legacy systems." Nobody has ever documented all their legacy systems. Nobody ever will. And yet somehow, the cloud migrations that actually happened didn't wait. They used the migration itself as the forcing function for documentation. The lab work got done alongside the lectures, not before them.
This isn't recklessness. This is how capability actually develops. You don't learn what verification you need by theorising about it. You learn it by watching what generation produces, stress-testing your assurance pipeline against real output, and improving both simultaneously. Sequential is intellectually tidy. Parallel is how things actually work.
And here's the deeper problem with "class must repeat before you advance": what if the class itself is failing? Level 1–2 human review is what most enterprises have right now — and it's already broken. The review theatre I described in Part 1? That is the current state. Humans nominally own it. The quality is nominally garbage. Telling an org to "stay here and master it" is telling them to keep repeating a class they're already flunking.
The Policy You Don't Announce
There's a second failure mode that the "cap the agent" position tends to ignore, and it's more dangerous than moving too fast.
Tell teams they cannot use AI agents until governance catches up, and they will use agents anyway. Just without guardrails, without logging, without anyone knowing. You don't get compliance. You get shadow AI — the 2026 equivalent of shadow IT, shadow cloud, shadow everything that happened every time a centralised policy tried to suppress a capability that was genuinely useful.
We know how that story ends. People don't stop doing the useful thing. They just stop doing it visibly. And then you lose both the guardrails and the oversight that would tell you guardrails are needed.
The organisations that managed cloud adoption badly weren't the ones who moved fast. They were the ones who announced a hard stop, watched it be ignored, and then had no governance infrastructure in place when they finally accepted reality. Cap the agent is the new block the cloud. The intent is responsible. The outcome is the same.
Rules Are Human. Enforcement Isn't.
Here's the part that took me a while to articulate clearly, and I think it's the conversation the governance community is avoiding.
There is a categorical difference between setting policy and enforcing it. We've been treating them as the same job, assigned to the same people, and that confusion is the root of most of what's broken in enterprise governance.
Policy is a human job. Deciding what the rule is, why it exists, where it applies and where it doesn't — that's judgment. That's values. That's context. Humans own that. Full stop.
Enforcement is not a human job. Not at scale. It never was. It just looked like one because we had no alternative.
What we call "human review" in most governance frameworks is not enforcement. It's interpretation dressed up as enforcement. And interpretation at scale is structurally inconsistent — shaped by who's reviewing, how tired they are, whether they understand the domain well enough to catch what they're supposed to catch, whether they have seventeen other things open on their screen. You're not getting the rule applied uniformly. You're getting one person's experience of the rule, filtered through their current state, on a particular afternoon.
That's not a criticism of individual reviewers. That's what human review structurally is when deployed as a quality gate at volume. And it's what most organisations have been calling governance.
Tools don't have difficult mornings. They apply the same rule at 3am on a Sunday as they do at 10am on a Tuesday. Across a thousand PRs, twelve teams, three continents — that consistency isn't a nice-to-have. It's the entire point of having a rule at all. The moment you ask a human to enforce the same constraint a thousand times, you no longer have a rule. You have a suggestion that some people take seriously.
The Boundary That Actually Matters
So the question isn't whether tools should have authority. They already do — every CI pipeline, every security scanner, every CMDB compliance check that blocks a deployment regardless of seniority is a tool exercising enforcement authority. We just haven't been honest with ourselves about what that means.
The question is whether we've drawn the boundary correctly.
Humans: own the policy. Define the criteria. Set the thresholds. Decide what "good enough" means and why. Review the tool's decisions at the edges — the ambiguous cases, the novel patterns, the places where the rule meets a situation it wasn't designed for.
Tools: enforce the policy. Consistently. At scale. Without emotion, without fatigue, without the quiet social pressure not to block a senior engineer's PR on a Friday.
[Kevin: insert the specific DHL example — the pipeline or compliance check that enforces regardless of who's asking — that makes this concrete]
This isn't a future state. It's already the architecture of every mature governance system. We're just reluctant to say it plainly because it sounds like we're removing humans from the loop. We're not. We're putting humans in the right part of the loop — the part that actually requires them.
The Question I Keep Coming Back To
The software factory will happen. The dark factory already exists at Fanuc. What matters is what kind of governance we build around it — and whether we build it honestly.
The honest version acknowledges that "a human reviewed it" stopped being a quality guarantee long before AI agents arrived. The honest version distinguishes between the judgment required to write a policy and the consistency required to enforce one. The honest version builds verification infrastructure alongside generation, not as a prerequisite that conveniently never gets met.
And the honest version asks the harder question: if we accept that tools are the enforcement layer — what does that mean for the humans who built their identity around being the ones who checked the work?
That's not a governance question. That's the same grief problem I described in Part 1. And I still don't have a clean answer for it.
Tags: #EnterpriseFuturewright #AIGovernance #SoftwareFactory #DigitalAlchemy #TechContrarian