"Product ops is moving to agents. Almost nobody is writing down what the agents aren't allowed to do – and when the line lands in the wrong place, nothing looks like an error." – Chase W. Hughes, Founder of ProAI

Nia Brown's From dashboards to agents argues that product operations has to stop being a reporting layer downstream of the work and become a real-time coordination system: signals captured from the work itself, agents surfacing blockers and proposing updates, humans keeping judgment. 

She’s right. And the most important phrase in the piece goes by fastest. Agents, she writes, "route and coordinate workflows within defined boundaries." That's the piece most teams still haven't worked out. In most teams I've seen, the boundary isn't defined at all; it's inherited from whatever the framework made easy to wire up, and discovered later by somebody outside the company.

So, here’s the rule I would defend. The deterministic boundary goes between things that have a correct answer and things that have a good answer. Anything with a correct answer belongs to code, permanently, however capable the model gets. Anything with a good answer – a judgment, a ranking, a guess about what matters this week – belongs to the model, and no prompting will make it reproducible.

I’ve been on both sides of this equation inside the same product.

What the handoff actually looks like

In 2021, I was generating financial projections on a language model, when hallucination was not a risk to mitigate but a certainty to design around. The model could write a paragraph about a business that a banker would find plausible. However, it could not be trusted to carry out simple additions or subtractions.

So, the model never produced a number. It read the user's description of their business and created a labeled scenario – industry, revenue model, headcount ramp, and a growth assumption expressed as a tier like "aggressive" or "conservative," not a percentage. A deterministic rules engine converted the tier into an actual rate, ran every calculation, and enforced the constraints: gross margin could not exceed 100%, cash could not go negative without a financing line, headcount cost had to tie to the payroll assumption.

What it rejected most often was the model implying a growth rate the industry tier did not allow – it wanted 40% month over month for a services business. The engine clamped it to the tier ceiling, and the user never saw the original.

That’s the boundary, in one handoff. Note what the rules engine is not: a safety net bolted on after a bad demo. It was the product, and the model was the interface to it. You are not adding guardrails to an agent – you’re building a deterministic system and using a model to make it usable by people who cannot specify what they want. 

It shipped to more than 300,000 businesses and institutions and held, for a boring reason: the model was never in a position to be wrong about anything verifiable.

Where I put the line in the wrong place

In the same product, I let the model own document assembly – deciding which sections went into a plan and in what order. It seemed like a judgment call. It wasn't. Certain sections were required for a bank submission, and when a user's description was thin, the model would occasionally drop it because nothing in the input suggested it was needed.

Here’s the part worth sitting with. It read as a reasonable editorial choice every single time. No exception, no malformed output, no confidence score to put a threshold under – a slightly shorter document with a sensible section order isn’t a malfunction, and there was nothing for a monitor to fire on. We found out when a user's lender bounced the document back for a missing section. The detection came from outside the system entirely, from a third party with a checklist we had never encoded.

Section inclusion had a correct answer. It was a rule about the submission target, not a judgment about the business. I handed it to the model because it felt like writing.

Air Canada made the same mistake in public. Its website chatbot told a passenger he could claim a bereavement discount retroactively; the airline's own policy page, linked from inside the chatbot's answer, said the opposite. 

When Air Canada argued it could not be liable for what its chatbot said, a tribunal called that submission remarkable: "While a chatbot has an interactive component, it is still just a part of Air Canada's website." The correct policy already existed on a page they controlled. Someone put a generator in front of it and let the generation be the answer.

The second line: Propose versus commit

The category rule gets you most of the way. It does not cover the agent doing something with no correct answer but real consequences: reprioritizing a backlog, closing tickets as duplicates, flipping a flag, or emailing a customer. These kinds of judgment calls cannot safely be handed over to an agent.

The second boundary is the write. An agent may read anything, compute anything, propose anything. The commit is executed by deterministic code, as a typed action from a fixed set, with preconditions re-checked at the write rather than at the plan. That last clause is where multi-step agents come apart: the agent verifies a condition in step two, acts on it in step nine, and the world moved in between.

That’s not a preference – it’s how the most realistic public benchmark I know of for tool-using agents scores them: τ-bench drops an agent into a customer-service domain with real APIs, a simulated user, and a written policy, then throws the conversation away and compares the final database state against the ground truth. The write is the only thing that counts. 

When it was published, GPT-4o with function calling solved about 61% of its retail tasks and 35% of the airline ones – and pass^8, the chance of getting the same task right on eight consecutive runs, fell below 25%. Consistency collapses faster than capability, which is why a demo tells you nothing about a system that runs the same decision ten thousand times a week.

Supervision that isn't

The other direction is less discussed and, in product operations, more common. Route everything through a human and call it supervision.

There are three decades of evidence on what happens next. In a 1993 experiment run by NASA-affiliated researchers, people monitored an automated system while also doing manual work. When its reliability held constant they caught 33% of its failures; when it visibly varied, 82%. With monitoring as their only task, they caught about 97%. 

The review synthesizing this literature is blunt: complacency "cannot be overcome with simple practice," and automation bias "cannot be prevented by training or instructions." It is not laziness; it is attention going where it is needed.

A reviewer signing off on output that is right 98% of the time is not a reviewer by week three; they are a latency cost with a mouse. Supervision has a throughput budget, far smaller than most designs assume. Spend it on what is irreversible, expensive, or visible to a customer, and constrain everything else so the bad output cannot be expressed.

"This is just software engineering"

The strongest objection is that none of this is new. Typed actions, input validation, transactions, preconditions checked at commit, an audit log: competent backend engineering from thirty years ago in agent vocabulary.

The discipline isn’t new; what’s new is that teams abandon it because the model appears to make it unnecessary. 

Nothing else in your stack fails the way a model does. A broken service fails loudly: it returns a 500 error, with a stack trace pointing to what went wrong. My document assembler returned a clean, well-ordered, plausible plan with a required section missing. Fluency defeats inspection, and inspection is the control almost everyone is quietly relying on. The fix is structural, not procedural: not "we will check the output," but "that output cannot be produced."

Before the agent ships, go action by action and write down three things: whether it has a correct answer or a good one, whether it reads or writes, and whether the write can be undone. Three columns, an afternoon. That’s the system specification.

Brown is right that the goal is not autonomous operations, and right that judgment stays with people. What I’d add is that judgment does not stay with people by default. It drains toward whichever component is willing to produce an answer, and a language model is always willing.

So, be most suspicious of the actions you want to hand over because they involve words. That’s where I lost, and it’s the whole rule in one line: "this feels like language" is not the same as "this has no correct answer."