AI Won’t Transform Work One Engineering Ticket at a Time
When the people who know the work can’t correct the agent themselves, every routine fix has to wait for an engineer.
Our grading rubrics lived in Google Drive. They held the most valuable knowledge in the education company I founded: what examiners looked for, why a technically correct answer could still be poor. The examiners could not change how the AI assessed an answer. They explained their judgment to us, we wrote specifications, engineers wrote software, and only then could they see whether the system did what they meant. Often, seeing it work gave them something else to explain.
As a consultant I built a service-desk AI system for a Norwegian managed-services provider. It worked. By handover I already regarded parts of it as legacy, because changing it had become uncomfortably expensive. We would show a workflow to a domain expert. That is not quite how we do this, they would say. We would change the flows and bring it back, and the revised version would reveal another judgment nobody had articulated. Back we went.
Decagon’s founder Jesse Zhang describes a customer that built roughly three customer-service workflows in a year while changes went through engineers, then seven in about a month after moving to a product its own staff could inspect and change. That is his account of one customer, and the pattern I recognize.
Everyone in these stories was competent. Some knowledge only became explicit when there was an attempt to react to. A 2024 study by Shankar and colleagues found the same difficulty in evaluation: inspecting language-model outputs helped practitioners discover and revise their criteria. Repeated correction is part of finding out what the system needs to do.
Making engineers the compulsory route for routine changes that practitioners can test and judge is now an obsolete design choice. I would no longer deliver a system on that basis.
A different way to deliver AI
The queue is built into the arrangement whenever engineering remains the only route to changing the agent. A delivery contract that treats handover as completion can leave the client with a working system but no way to improve it. And treating engineering control as governance makes even a correction to business instructions wait for technical work. Those choices are enough to preserve the queue, however capable the people in it.
Delivery roles are already changing. Accenture describes forward-deployed engineers translating requirements into AI applications; Kyndryl asks engineers to own delivery from scoping to production; Anthropic includes reusable agent skills among the things they deliver. These are different versions of an important job. The question for each is what the client can change once that job is done.
Some firms are already changing what they build around. Capgemini is hiring an engineer to develop the agent skills behind its own offer-creation process, working with the people whose expertise shapes those offers. Decagon describes redesigning its delivery system so Agent Development Managers can change and test customer policies in plain English. It reports cutting custom engineering per agent by 80 percent. One example concerns a consultancy’s internal work, the other a vendor’s delivery team. Both make reusable knowledge and continuing improvement part of the system being built.
I think an emerging philosophy is taking shape here: build AI systems around the people who will keep improving them. I would take that principle all the way to the client. The people who know the process should be able to correct their agents, test the change and keep a better way of working. That governed ability to improve belongs in the product the client buys, without another call to the delivery firm for every routine correction.
Let the people doing the work improve their agents
The tools are making this practical. An agent can read standing instructions in a file such as AGENTS.md, and use skills that describe how a particular kind of work gets done. OpenClaw can turn corrections into reusable skills; Hermes can create and revise its own procedures. Practitioners can work with the agent on those instructions without translating every change into application code.
The business user is in the driving seat. They work with their own agent, explain what they want, react to what it does and steer it towards better performance through conversation. The agent can save that feedback as instructions for future work. Krista Letz describes doing this with Grok Bot, asking it to add her corrections to its skills. The user should be able to try the change on examples they can judge and see whether it improves the result. They should only need to converse with the agent and judge its work; they should not even need to know what a skill is.
The obvious objection is governance. A correction that suits one user may be wrong for the team. The person responsible for the process should decide what becomes part of the shared instructions. Engineering should provide the integrations, access controls, evaluation tools and mechanisms for approving, releasing and reversing changes. Practitioners can then improve behaviour within those boundaries. Changing what the agent is allowed to access remains a different responsibility from correcting how it does an already permitted task.
The pace of activity makes this worth taking seriously now. In the four full weeks ending September 13, OpenRouter’s coding-agent category recorded 84 percent more tokens than in the preceding four weeks. It includes Hermes Agent, Claude Code, Codex and OpenClaw. That is explosive growth in agent traffic through one gateway; it does not tell us how many workers gained control over their work.
OpenRouter coding-agent category · trillion tokens
84% more coding-agent tokens passed through OpenRouter
Weekly tokens, January 5–September 13, 2026. The category includes Hermes Agent, Claude Code, Codex, OpenClaw, Kilo Code, Cline, pi and other apps.
Selected tools · trillion tokens routed through OpenRouter
Hermes Agent, Claude Code and Codex all grew on OpenRouter
Four-week token volumes rose 53%, 106% and 271%, respectively, from July 20–August 16 to August 17–September 13, 2026.
For evidence of adoption beyond developers, look at who is using Codex. In June, OpenAI reported more than five million weekly active users, up more than sixfold since the desktop app’s February launch. Knowledge workers accounted for about a fifth of users and were growing more than three times as fast as developers. By August 31, Thibault Sottiaux was reporting 25 million active users in an announcement about Codex and ChatGPT Work, without specifying the activity period. The June knowledge-worker figures are the clearer evidence that people outside engineering are taking these tools into their own work.
Reported adoption milestones · 2026
Codex expanded beyond developers as its audience grew
OpenAI reported more than 5 million weekly active users in June; knowledge workers represented about 20% and were growing more than three times as fast as developers.
Knowledge workers represented about a fifth of users and were growing more than three times as fast as developers.
Announcement accompanying Codex and ChatGPT Work usage resets; activity window and product breakdown unspecified.
The architecture is moving in the same direction. On September 10, OpenAI introduced the Agents API in public beta, offering the system that runs Codex agents as a managed service. Businesses can build on that foundation instead of assembling every part themselves. I read the combination of rapid activity, adoption beyond developers and this shared infrastructure as a clear signal of where demand is going. I would design around people directing and improving their own agents as part of doing their work.
Ask who can make the next correction
As an AI practitioner, I have to be willing to revise my priors. And the evidence is clearly telling us that it’s about time we do so. I would approach those projects differently today: get useful agents into clients’ hands faster and more cheaply, and give their people much more agency over how those agents do the work.
I expect OpenAI’s DevDay on September 29 to give us another signal of this direction, and Anthropic and Google to keep pushing it too. Delivery firms should be building that offer now.
The decisive measure is correction latency: how long it takes a practitioner to turn a mistake into a tested correction that improves subsequent work. Ask the vendor to demonstrate that process with the people who will own it. Getting the first agent into production quickly and cheaply matters. So does what happens the next time someone says, that is not quite how we do this.
“If a routine correction still means raising an engineering ticket, you have bought a queue.”
— Nico Escudero
Before buying, ask: after handover, who can change the agent’s behaviour, and how long does a tested correction take? If a routine correction still means raising an engineering ticket, you have bought a queue.