
Læs dette blogindlæg på dansk
Go find the colleague in your organization who approves output from an AI agent, and ask one question:
How many AI outputs have you said no to in the past month?
If the answer is zero, you don't have oversight. You have a person clicking approve. And that's a more dangerous position than having no check at all, because the organization believes it's covered.
It took me a year and a half of running things, and one deeply uncomfortable afternoon, to understand that difference.
"Human-in-the-loop" almost always gets presented as a concession. Something we're stuck with for now. A temporary tax on autonomy that disappears as the AI technology matures.
Notice what that framing assumes: that the goal is zero humans, and that every person in the flow is a step in the wrong direction.
Organizations that actually run AI agents in production have stopped talking that way. In practice, the discussion is rarely autonomous versus non-autonomous — it's risk-managed autonomy: how much the system gets to do on its own, weighed against the consequences if it gets something wrong.
There's been a shift most organizations still talk too little about. We treat AI as a tool each person needs to learn to use, but in practice it's rearranging who does what, and at which point in the process someone needs to check the work.
I've spent the past couple of years standing in the middle of that shift, and in my experience, one question gets overlooked: at what point in the flow should an output be checked? That step is exactly the vulnerable one, because it easily looks like pure friction. Which is why it's often the first thing to get cut.
The conversation about human oversight points toward zero humans. Practice points toward the right person in the right place.
In my experience, one question gets overlooked: at what point in the flow should an output be checked?
The playbook is familiar. You add an approval step, write "with human validation" into the process doc, and consider the question closed.
But a human in the loop only has value if three conditions are met at the same time:
Miss just one of the three, and what you have in practice is an alibi, not oversight. Which is why the rejection-count question is so useful: it doesn't measure the principle, it measures whether someone actually has the capacity, the understanding, and the mandate to use it.
Zero rejections doesn't mean the system is good. It means nobody checked.
The human wasn't what made us slow. The human was what let us dare to move fast.
We ran four product teams at a SaaS company where most of the product development and customer dialogue was digitally connected. Case management, behavioral data, the codebase, and the design system were all linked, and AI agents had gradually become part of nearly every core workflow.
The rule I landed on was boring enough to actually stick: put the human as late in the flow as possible, and make the task a yes or a no.
Put the human first, and you're spending a person on instructing a machine. That's slow, and it wastes the exact judgment the person is there for. Put the human last, and you're spending a person on deciding whether something finished is good, right before it becomes irreversible. That takes thirty seconds.
In practice, it looked like this for me: emails that came in during meetings sat as drafts by the time I got back — not sent, but ready for my approval. A blocked story got investigated before I even knew someone was stuck, but I still had to confirm the message to the developer. A bug turned into a finished pull request that the development team weighed in on.
None of those steps slowed anything down. They just moved where a human's attention got spent — from the start of the work to the end.
One concrete case showed just how fast that could go: an IT partner needed a security benchmark analysis for customer meetings. Less than thirty minutes later, the need was analyzed, the solution validated, built, and shipped to production. A human sat inside that AI flow the whole time.
The human wasn't what made us slow. The human was what let us dare to move fast.
We also tried flows where the AI agent got to carry work all the way through, with no human checkpoint along the way.
That's where the problem showed up.
An AI agent was set up to update product descriptions from data in a spreadsheet. One column had shifted, but the agent kept going as instructed. The error sat live for four hours before a customer called to ask if the offer could really be right.
It ended up costing nothing. But there was no step in the flow where a human had to approve the change before it went live. That was luck, not oversight.
The agent did, in principle, exactly what it was set up to do. The flaw was in the flow around it.
Nobody had decided the changes needed approval before publishing. But nobody had actively decided they didn't, either. The decision just never got made.
That's how a black box shows up in practice. Not necessarily because the technology fails, but because nobody decided where the oversight should sit.
The result is an unintended incident: something happens not because anyone decided it should, but because nobody decided it shouldn't.
The exercise takes an hour, a whiteboard, and one flow where you're already using AI.
Worth noting: none of the three steps require you to understand the technology. They require someone to take a position. And that position can't be worked out from a meeting room alone — it has to happen with the people who have their hands in the work every day. They're the ones who know how often an output lands per hour, and whether anyone actually has time to read it.
Decisions that don't get made still get made — just by accident.
The autonomy conversation is often held as if it's about how mature, safe, or capable the technology is. It isn't. It's about how much consequence the organization can accept if something goes wrong — and whether anyone has actually sat down and worked that out.
Human-in-the-loop isn't what's left over when automation runs out of reach. It's the decision that lets you dare to let automation reach as far as it does.
Decisions that don't get made still get made — just by accident.

Vibe Coding for Product Managers

Read blog post