IP / READ / 2026

The number they ask us for

A single confident number does not mean there is no risk. It means the risk was not counted.


We are a technical contractor. Clients come to us with the decision already made. They have decided that the work goes to a model, and the date is already in the plan. What they want from us is an estimate.

We have to give the estimate, because saying “it is all complicated here” and refusing to name a number is not a professional position.

The difficulty is elsewhere. We will count what the process description covers, and we will count it correctly. More will be taken from the person than the description covers, and on the day we name the number, nobody sees the difference.

The client has a problem of their own, because every way of answering costs them something. First, if they name a small number, in six months they will be explaining why it was not enough. Second, if they name a large one, they will be asked where it came from. Third, if they give a range, it reads as “he does not know himself.”

The number does not follow from the process description

The usual way to count works like this. First, you take a process that has a regulation and metrics. Second, you run a pilot on a small volume. Third, you multiply the result.

The pilot does not lie, because it measures what is described, and it measures honestly.

But a person does more than what is written in her job description. The rest of what she does does not make it into any table we look at. The reason is not that it cannot be described, because in mature companies feedback, knowledge management and process improvement are all documented well enough. They are simply not tied to a specific role as something that disappears along with it.

An example from a first line of support

Take a first line of support.

The regulation says that she accepts a ticket, assigns a category, answers within the deadline, and escalates what is complex. All of it is measurable, and that is why automation starts there.

Every day she also does work that the regulation does not mention. First, she notices that the third client this week is tripping over the same thing, and she says so at the retrospective. Second, she knows the place where the instruction lies, because she has worked from it herself. Third, she passes on to a developer a complaint the client dropped in passing.

Let us call the three of them outputs, meaning what a person gives outward beyond her main work. Each output is needed by somebody. Without the first, the retrospective gets shorter; without the second, the knowledge base stops being topped up; without the third, the developer stops hearing what irritates people.

The outputs have no owner, no metric and no line in the process description. They exist only because a particular person holds the role.

They are not in the process diagrams either. On the diagram the role stands in one process. The same person is drawn on two more diagrams, and nothing connects the three pictures.

Support here is just an example. The same will happen with an accountant who calls the supplier before an invoice goes into rejection, because she remembers whose VAT comes in wrong. Or with a lawyer who stops at a clause in a contract, because last year the same clause cost money in another deal. The question in every case is the same. How many people outside her own process did this person bring something to? If the answer is nobody, the client loses little. If the answer is four, the client gets four problems, and each one will be noticed separately and will not be connected to our project.

Earlier descriptions of the same problem, without artificial intelligence

Clients often hear the argument as a complaint about artificial intelligence. The shortest answer is a story with no artificial intelligence in it at all.

In 1996 Julian Orr published a book about the technicians who repaired Xerox copiers. He observed their work. The main thing he saw was that they fixed the hardest breakdowns from the stories rather than from the manual. They would gather and retell each other the cases they had already dealt with. The manuals existed, and they were not enough.

The ideas below are worth knowing because they give you something to answer the client with.

The first idea is the substitution myth, which Christoffersen and Woods named in the early two-thousands. We think a machine simply takes a person’s place and the rest stays as it was. In reality everything in a system is fitted to everything else. Remove one part, and the neighbouring ones start working differently. The mistake is not in how we counted the work. The mistake is in how we imagined the system.

The second idea comes from manufacturing. Toyota has a rule called jidoka, which is one of the two pillars of their production system. A machine stops itself the moment something goes wrong, and a person stops the line the same way. First you notice the deviation and stop, then you fix it, then you find the cause and put in a guard so it does not repeat. Automating work and losing the ability to notice deviations is considered a mistake there.

What changed when language models arrived

Before, what got automated was work that has a form, such as reconciliations, moving data and calculations. Conversation could not be automated, so everything a person noticed during a conversation stayed with her.

First, the model can do conversation, so automation now takes exactly the work during which the person was doing the noticing.

Second, the cost of automating language operations and operations that require judgement fell sharply. Integration, security and process rebuilding did not get cheaper. The threshold for the first decision dropped, and more such decisions are now being made.

Third, and worst, the model writes an answer you would not be ashamed to show a client. Before, a broken process was visible from a bad result. Now the result looks fine, and nobody brings what the person used to bring on the side.

The formal project will close on time

Working formally, the task is solvable. You take the described work, you describe the target state, you give an estimate, and you hold the deadline. The metrics go where they should, and the project closes with a sign-off.

A formally successful project is the worst of the possible outcomes, and it is worst for us.

Our estimate was truthful about the part that was described, and about the rest it was silent. The decision, though, was made about the whole person. A few months later something will start falling apart in a neighbouring department, and nobody will connect it to our project. Formally we did everything right.

The gap does not get caught at planning, and the reason is simple. One set of people described the processes in the company, and we do the automating. They do not know what the model can already do, and we do not know whom this person brought something to beyond her main work. Each side is right about its own half, and we are the ones who bring the number to the plan.

A list of outputs makes the estimate possible

A small number, a large one and a range are equally bad, because all three are an opinion. A number is defended by where it came from, so what we need is something that can be enumerated.

What can be enumerated is the outputs, meaning whom this person brings something to beyond her main work.

Asking the person herself is not enough, and the reason is not that she is hiding anything. For her it is not work, but something self-evident. Asked “what else do you do here,” she will recite her job description.

What works is different:

  • watch how she actually works;
  • ask about several recent concrete cases;
  • look at whom she wrote to and what came of it;
  • ask the people downstream who will notice if this role disappears.

The last one gives the fastest result. You ask the team lead, the product manager and whoever runs the knowledge base, and each of them names their own.

The list is ready, but it is not yet a cost estimate.

Three outputs can mean two hours of work on an automatic digest, or three months of rebuilding how the company stores knowledge. They can also mean a separate person at half time, or a decision to accept the loss and do nothing about it. The number of outputs does not equal the volume of work.

The list changes what the client walks into the defence with. Before the list, they carried a feeling. With the list, they carry items that are each estimated separately and in the ordinary way.

The range does not disappear, but now it has a date and a reason. At the defence it sounds like this. The first phase costs this much; we will name the second after the first is done, when there is a list; the list will be ready on this date.

What the parallel period shows

You collect the list during the parallel period, when the model is already working and the person is still in place.

Usually a parallel run is done with one purpose, which is to check whether the machine copes with the old work. It also gives you something else. For several weeks you can see what the person fills the freed time with, and what she fills it with is the list. The parallel period gives what a process description does not, because it shows the outputs already in the new working schedule.

The period saves nothing, and you pay for it twice, because for a while both the model and the person are working. It is the first thing cut when the effect has to be shown sooner.

Only someone working inside the process can compile the list of outputs, and only we can say how much it costs to close each of them. Neither of us manages it alone.

The list does not solve the problem

For every output left without a replacement, you have to design a mechanism that gives the same result in the new schedule. Building the mechanism is design work, and it is no longer listing.

Retrospective input can be replaced by having the agent gather recurring requests once a week and bring them to people. Updating the knowledge base can be assigned to a specific role, with time in the calendar and a name in the list of duties. But the fact that the person knew which engineer to go to with a strange case lends itself poorly to replacement, and the honest answer here is to record the loss and decide whether it is acceptable.

A contractor’s own part ends here. We will build the mechanism, but it will only start working when daily habits change, and by habits we mean roles, the order of meetings, and who gets asked when something is unclear.

The full picture does not live in a document

In one of our projects we built a chatbot over a corporate learning platform. The requirements were agreed in detail, down to data formats and the boundaries of each side’s responsibility. We formally did what we had agreed on.

Then the client took it to their own customers, into their real knowledge bases.

Part of the learning material is written ambiguously. What a person reads without trouble is ambiguous for search. In many customers the materials are also mixed by language, and where the languages mix, the formal agreements about language stop working the way they were written. And the answer had to depend on where the reader already was, which is not something either side could have written down in advance. Nobody knew it mattered until real material was in front of the system.

None of the three problems was a mistake in the requirements. The full picture simply existed only in real customers’ materials, and not in the document.

We fixed it by iterating rather than by renegotiating scope. Our engineers looked for workarounds, and the client agreed to change what they had at first considered unchangeable. We arrived at a solution that suited everyone.

Where the argument becomes an excuse

The same argument works just as well in the opposite direction. “You do not understand, everything here rests on the informal” is a standard phrase of resistance to change, and it sounds just as convincing as everything above.

Part of the invisible work is necessary, because somebody downstream depends on it. Part of it exists only because the process is bad, and a person has to fix by hand what should not have been breaking. Automation that removes the second kind is doing the right thing.

We do not have a clean criterion. The best test we have is that a real output has somebody downstream who will notice its disappearance. If over a full cycle of the process nobody the output fed into noticed, it was most likely patching a bad process. The cycles here differ, from a week in one place to a year in another, so naming a single period would mean inventing precision.

What we take on and what we cannot

There is no method here. Selling it as a method would be the same trick the whole text is written against. What we have instead is a working schedule.

We give an estimate. Along with it we give a list of outputs and say which of them are left without a replacement. Then we design and build the compensation mechanisms. Everything that can be done technically is our part.

The client’s own part cannot be handed to us, and without it nothing listed above works. First, we need feedback from reality, because the full picture lives in their materials and in their people rather than in a document. Second, and more important, we need a willingness to change what at first seemed unchangeable, along with the habits of the people who live in that process.

AI transformation is a rebuilding of how people work, and it is not the same as delivering software. The rebuilding is the half of the work a contractor will not do for the client.

One more thing usually goes unsaid. A single confident number does not mean there is no risk. It means the risk was not counted.

EXACT PREFIX · ENTER TO GO · ESC TO CLOSE