IP / READ / 2026

The most expensive resource in an agentic system

A mistake at the boundary between agent and human is not visible, because the system keeps working correctly.


Almost every team rolling out agents right now takes the same route. First there is a demo where everything works. Second there is a pilot that confirms the demo, and the team concludes that the work really can be handed to a machine. Third there are a few months in which it turns out that not all of the work can be handed over, not in that way, and not at once, and the date for the next stage slips a second time.

The simplest explanation is “we underestimated the complexity.” It is correct and almost useless, because it does not say what exactly was underestimated.

The same thing gets underestimated almost every time, and it is not the model, the tools, or the integrations. It is the boundary between the agent and the human. The boundary covers what the system decides on its own and at what moment it calls a person. It also covers what the agent brings when it calls, and what happens to that person’s work after part of it has been taken away.

A mistake at the boundary costs more than a technical one. A technical mistake is visible, because you can point to the wrong model, the wrong tool, or the wrong way of holding state. Somebody finds it and fixes it. A mistake at the boundary between agent and human is not visible, because the system keeps working correctly. It simply does the wrong thing in the wrong place. You can do everything else well and still get nothing out of the system. You can also do everything else badly and still end up with a system that delivers value, because the people around it can see what it does and have time to correct it.

The simplest example is an operator who has confirmed the agent’s correct action twenty times in a row, and who then makes the twenty-first confirmation without reading it. The behaviour is not carelessness, and it is not a question of discipline. It is how attention works. When a signal turned out to be a false alarm nineteen times, the twentieth one is noise rather than a signal.

Either the twenty-first action is also correct, and nobody notices anything, or it is wrong, and the postmortem carries the line “the human confirmed it.” The line is formally true, and in reality there had been no human in that decision for hours.

We have learned to count latency, memory, inference cost and error budgets. We have not learned to count what we spend every time the system turns to a person.

What has become cheaper and what has not

Machine execution has become cheap enough to change how software is built. Reading a hundred thousand lines of context, working through a dozen hypotheses, running a check three times, and generating three variants in order to throw two away now cost so little that nobody economises on them.

What has not become cheaper is the part that stays human. A person still has to judge an ambiguous case, take responsibility, understand a goal nobody articulated, and notice that the context has changed so the earlier agreement no longer holds.

Everything else in the text follows from one claim.

The most expensive resource in an agentic system is qualified human attention.

A less obvious consequence follows as well. If a resource is scarce, you cannot spend it on a leftover basis, and a leftover basis is exactly how we treat it. The question of when the system turns to a person gets decided at the end, together with screen layout, and the team calls it UX. It is an architectural decision, and it has to be made at the same time as the decisions about transaction boundaries and what to do with partial failure.

We provisionally call the class of decisions Agent Operator Experience, or AOX.

Assumptions about the operator

Most designs of the human-agent boundary carry a hidden model of the person. In that model she is attentive and rested, she remembers the task context, she reads what she is shown, and she has time.

That person does not exist. The operator who does exist is in a hurry because she has four more tickets, or she is on the eighth hour of a shift. She leaves the obvious unsaid because to her it is obvious, and she mistakes the plausible for the correct, because the agent writes with confidence. She has also learned already not to read the monitoring, because the monitoring lies three times a week.

The list is not a list of bad operators. The items are operating conditions, in the same sense that a network that drops packets or a disk that will eventually fill up are operating conditions. With the network we act accordingly, and nobody designs a distributed system on the assumption that packets do not get lost. With people we act differently, and we are surprised by the result every time.

Human-in-the-loop is a topology, not an architecture

The simplest answer to the risks of autonomy is to put a human in the loop.

The answer says whether there is a human in the system, and it says nothing else. It does not say why the human is there, at what moment she is called, or what exactly she is shown. It also does not say how much context she has to reconstruct before she can judge, or what happens if she is called twenty times in a row without need.

A system that escalates every low-risk decision spends the resource it will need at the moment of genuine uncertainty. It spends the resource invisibly, because in the charts the pattern looks like exemplary oversight, with a hundred percent of actions confirmed by a human.

Human-in-the-loop is a topology. Human oversight is a design problem.

The attention budget

In architecture we have long worked with budgets. A latency budget says how many milliseconds you can spend before the user starts to notice. An error budget says how many failures are allowed before the team stops shipping and goes to fix things. Both work because they turn a vague complaint into a number you can argue about at planning, and their precision is not what makes them work.

Agentic systems need one more budget, for human attention. The budget makes these questions meaningful:

  • How many times a day the agent interrupts the operator.
  • What share of those interruptions genuinely required human judgement.
  • How much context the person has to reconstruct before she can decide.
  • How many signals she will see before she stops reading them.
  • Which decisions can be accumulated and shown in a batch instead of twenty interruptions.
  • Which checks the agent is obliged to run before turning to a person.

For agentic systems the quantities have not yet been assembled into a set you can use at planning. The related measures in human factors have been studied for decades, and they include cognitive load, situation awareness, and false-alarm fatigue. Versions of them carried over to an agentic system and fit for daily use do not exist, and the absence is not a reason to stop asking. The latency budget also started with someone asking for the first time how much latency they actually had, when nobody held the answer.

A budget holds more than costs, and without the other entry it reads wrongly. An operator whom the agent has fully freed from routine will, a year later, be judging the package the agent prepared for her rather than the system itself. She can no longer check the agent, because the only thing she sees of the system is what the agent showed her. So part of the attention has to be spent deliberately and seemingly pointlessly, in order to preserve the ability to spend it well. Nobody knows how much. The question is honestly open, and I am not offering a principle here.

The attention budget also counts one channel of loss out of several. A person in a process performs informal functions as well, and they disappear along with the automated work, while the attention budget has no entry for them. The loss is a separate problem, and there is a separate text about it, “The number they ask us for.” Here it is enough not to confuse saved attention with the full price of the transition.

One transition that worked

A support organisation hits the same wall at a certain size. Answering a request needs both the system and the specific customer’s situation, and the two live in different heads. Add several markets, several languages, a range of hardware and rules that differ by country, and consistency stops being a question of effort. It becomes a question of how much one person can hold at once. Until recently the only answer was more people.

One vendor chose to work the problem in steps instead of replacing the function. What follows is their sequence, without figures.

First we built something small, assistants over the knowledge base that found and suggested the right article to the operator. Second we built a copilot that pulled together everything about the client and the history of interaction with them, translated it into the client’s language, and left the decision to the person.

The copilot was where the boundary became visible. With the copilot in place you could see which requests the first line closes mechanically, with the same question, the same context, and the same answer. You could also see where the operator had to work the request out herself every time, because the configuration, the regulation, or the client’s history made the case unrepeatable. Guessing the split up front did not work, and we arrived at it by trial.

Only then did they rebuild the structure of support. The model took the first line. The people moved to the decisions that need judgement and to the link with engineering teams. Quality levelled out, as far as we could observe, and I cannot give figures.

The work around the boundary changed as well, because the knowledge base is filled differently now and the interaction with engineering is arranged differently. The change is a separate topic, and it is covered in the adjacent text.

On our side, the transition worked because we had experience with transitions of the same kind. On the client’s side, it worked because they were willing to change processes and understood what exactly they were starting. Each condition is uncommon, and together they are rarer still.

The copilot was the probe we used to find the boundary, and we reached the boundary through iterations that were long and exhausting for people. Part of the iteration is unavoidable. Discipline does not cancel it. What discipline does is make the iterations fewer, and keep the transition from resting on the coincidence of two rare conditions.

Earlier work on the same problem

Everything written above has been known for a long time, and other people knew it first.

Lisanne Bainbridge described the ironies of automation back in eighty-three. The better a system is automated, the less the operator practises, and the worse prepared she is for the moment when she is needed. The same line of work continues. Endsley wrote on the situation awareness of a human taken out of the loop. Parasuraman and Riley wrote on the three things people actually do with automation, which are to over-trust it, to ignore it because of false alarms, and to deploy it without asking what will happen to the human on the other side. Hollnagel and Woods wrote on the joint cognitive system, instead of treating human and machine separately. Aviation, energy and control rooms have lived with the problem for decades.

The old body of work is being carried over to a new setting, where an agent acts for a long time, acts in parallel, accumulates state, and takes hundreds of steps without us. The setting has no new laws of human behaviour in it.

Existing practical guidance

Practical recommendations also exist, and you should acknowledge them before proposing anything of your own. The following is the best material there currently is, and you should read it before you read me:

  • Guidelines for Human-AI Interaction, Amershi et al., CHI 2019: eighteen rules, distilled from the literature and industry guidance and validated on twenty products;
  • Microsoft HAX Toolkit: the same rules plus a workbook and a pattern library;
  • NIST AI RMF: Agentic Profile, Cloud Security Alliance, March 2026, draft: four levels of autonomy and an oversight boundary, with an explicit description of the moments in which an agent is obliged to stop and ask;
  • OWASP LLM06, Excessive Agency: the same thing from the security side, where excessive autonomy stands next to excessive permissions and excessive functionality.

The last two items deserve separate attention. Risk management and security came from different directions to the same conclusion, which is that the moment at which an agent stops and calls a human has to be described explicitly.

Why the existing guidance is not enough

First, the guidance sits at the level of principles. “Help the user calibrate trust” is true and nearly unachievable without ten more decisions that every team makes its own way. You cannot assemble a reproducible process from such a set, and without reproducibility there is no evaluation and no transfer of experience between teams.

Second, and the reason weighs heavier, the guidance describes the steady state, in which the levels are defined, the thresholds are set, and the operator knows her role. People wrote about the transition itself long ago, under the names adaptive automation, adjustable autonomy, dynamic function allocation, and human-autonomy teaming. That work contains function allocation methods and step-by-step procedures for determining the level of autonomy. It was written for aviation, defence unmanned systems and nuclear plants, and it did not become standard engineering practice for an ordinary product team. The missing piece is a process by which an ordinary product team gets from pilot and copilot to a working boundary, and measures along the way how much attention the boundary costs.

What the transition looks like

First there is a demonstration of capability. Somebody shows what the model can do, and it becomes obvious to everyone that part of the work can be handed over. At the demonstration stage the estimates are the most accurate and the most wrong at the same time, because the step really is cheap, and what lies behind it is not yet visible.

Second a copilot appears. Nobody plans the copilot, and it appears because full automation does not hold straight away. A copilot is an artificial arrangement in which the person stays in place and the machine prompts, and the arrangement is where you can see where the boundary actually runs. It is imperfect by definition. Its value is in letting you find the boundary while the functions have not yet been removed and you can still see what breaks, and the value is not productivity.

Third the longest part begins. By then you can see where the boundary runs and also what to do with it. A repeatable decision goes into policy. A decision that needs judgement every time forces the boundary back, closer to the person.

Every shift of the boundary has to be measured by something, because otherwise the next iteration will be as much of a guess as the first. Then the loop runs again, and again. Quality grows neither immediately nor linearly.

The process is poorly predictable, and it is fair to say that nobody will make it predictable. The question to ask instead is how to make the process fit for planning costs. Predictability and cost planning are different tasks, and confusing them is expensive, because the board hears “we do not know how long this will take” and reads it as incompetence rather than as a property of the task.

What happens without the loop

When there is no design discipline but there is pressure, autonomy starts being designed from capability rather than from the process.

The logic is simple. The business sees the potential, and the engineers see agents, tools, and models. Between them there is no established way to design the joint work of human and machine, so they automate what can be automated, rather than the places where automation creates value.

The result is “an agent because agent,” “multi-agent because scalability,” and “a confirmation gate because safety.” The organisation copies the external form of someone else’s automation without copying the reasons it worked in another context.

The pattern is agentic cargo cult, and you recognise it by one sign. Nobody can say what decision the person actually makes at the gate, or what happens if she clicks without thinking.

The boundary can be designed explicitly

The templates below are not a checklist. They will become one when enough completed cases accumulate, and nobody has them yet. Take a template, cut it in half, and rework it for your own process.

First template: what to ask the system at an architecture review. Not about the agent, about the boundary.

  • Decision class. Which class of decisions the agent closes on its own, rather than what the agent can do. What happens if it closes one wrongly, and how long before anyone notices.
  • The call. At what moment the agent calls a person and on what signal. How many such calls per day fall on one operator.
  • Share of judgement. What part of the calls genuinely required a human decision, and what part was the designer’s insurance.
  • The material. What the agent brings at the moment of the call. What the person has to reconstruct herself, and how long until her first meaningful decision.
  • The decision itself. What the person decides here, in one sentence with a verb. If the sentence will not write, the gate stands for the report rather than for control.
  • The loss. What the person did outside the function, and who picks it up after her.

None of the questions has an answer you can copy. What they do at a review is show that the team has not thought about the boundary.

Second template: the form for recording a decision about the boundary. One boundary, one record, seven lines. The example is filled in for a first-line support case.

Line of the record Example
Decision class A ticket that closes with a known solution from the knowledge base
The agent does on its own Classifies, answers, closes, if confidence is above the threshold and the client is not on the exception list
Hands to a human Lower confidence, a repeat request from the same client, any mention of money or of cancellation
Arrives with The client’s history, its own hypothesis, a list of what it has already checked, and the place where exactly it is unsure
The human decides Whether this is a case the knowledge base does not know
The boundary is in the wrong place if The operator almost always answers “the base knows this”, which means the agent is calling for nothing. Or she closes calls faster than she can read them
Adjacent outputs Retrospective input and topping up the knowledge base. Picked up by a weekly review of tickets the agent closed, with a separate person, separate time, and separate owner

The last line is the weakest, because I do not have a metric that will show in a quarter that the replacement worked. The line stands in the form anyway, because without it the record looks complete.

Where the transition leads

At the end of the transition there is a system with several recognisable properties. Attention in it is budgeted rather than spent on a leftover basis, and autonomy is allocated by decision class. Escalation costs the machine work, because before calling a person the agent does everything it can itself and arrives with material rather than with a question. Context before handoff is reconstructed, and an autonomous action is reversible, observable and limited in how much it can affect. Trust is calibrated in both directions, and a repeatable decision sooner or later becomes policy.

The list is itself at the level of principles, and my complaint about other people’s lists applies to it just the same. It gives a direction rather than a checklist.

The second note weighs heavier. Almost every item saves attention and at the same time moves the person further from the raw material. The better such a system works, the less ground the operator has for her own judgement. Compensating for the loss has to be deliberate, and I do not yet know how.

What is still missing

The body of work for the discipline exists, and it is listed above. The other half does not.

There is no synthesis that puts forty years of human factors and the guidance of the last two years in one order, saying what follows from what, what outweighs what, and what survives the move to an agent that acts for a long time, acts in parallel, and accumulates state. There is also no settled name that an engineer can search for, the engineer who is designing a confirmation gate today and does not know that his question was answered in the eighties.

The AI era has not made the older work less deserving of attention. It has made the work more urgent, because a mistake spreads faster than anyone can notice it, and it costs differently than it cost in a control room.

What we called Agent Operator Experience at the start is so far a working name for the gap, and it is not the discipline itself. Nobody can assemble the discipline alone, and we do not claim that we can. The task is in plain view, and anyone who has already started the transition can see it.

Designing a system of human plus agent

We are designing a system of “human plus agent.”

The boundary of interaction with the operator is not an interface detail to be drawn in after the architecture is ready. It is itself part of the architecture, and decisions about it cost exactly as much as decisions about transactions, queues and fault tolerance.

A good agent is not one that never asks a person. A good agent knows when the person should be interrupted.

EXACT PREFIX · ENTER TO GO · ESC TO CLOSE