01Usage climbed, value didn't
The whole market is learning the same lesson.
In December 2025, Uber gave around 5,000 engineers access to Claude Code. By March 2026, 84% were agentic coding users and around 95% were using AI tools every month. Average spend ran at $150 to $250 per engineer a month, with heavy users at $500 to $2,000, and the chief technology officer reported spending $1,200 in a single two-hour session. By April, the company had used its entire planned 2026 AI budget, and its CTO said Uber was back to the drawing board. An internal leaderboard had ranked teams by how much they used AI tools. Its chief operating officer later said the link between that usage and more useful consumer features was not there yet.
In the same month, Microsoft opened Claude Code to its Experiences and Devices division, the group behind Windows, Microsoft 365, Outlook, Teams and Surface. It became popular quickly. By May 2026, an internal memo directed engineers to move to Microsoft's own GitHub Copilot CLI by 30 June, the end of Microsoft's financial year. Microsoft framed the change as toolchain unification; reporting pointed to token-based costs running well ahead of budget as well.
Engineers loved both tools. That was the problem. Adoption climbed, the spend climbed, and the value only shows once the work changes and someone measures it. In maturity terms, both were still at Level 2.
Research points the same way. MIT's Project NANDA found in preliminary 2025 findings that 95% of generative AI pilots produce no measurable financial impact. BCG puts roughly 70% of the challenge of AI in people and process, with the algorithms and technology accounting for the rest.
02The technology has moved. Most people are still catching up.
Before giving a talk on AI agents to a group of 20 capable professionals, Thomas Paule, founder of the Australian firm Visory, surveyed the room. 58% named data privacy and security as their top concern, 52% were worried about hallucinations, and many said they didn't know where to start. Only one or two of the twenty could explain what an AI agent was in practical terms. The rest were still using AI as a chatbot.
That is not a sign of a less capable group. It is what the wider market looks like. Most people, and most organisations, have not yet seen what an agent can do.
The technology has not waited. In November 2025, Austrian developer Peter Steinberger released a weekend project: an open-source AI agent that later became OpenClaw. It runs on a person's own machine, talks to them through the messaging apps they already use, and works on goals around the clock. It became the fastest-growing project in GitHub history, passing 200,000 stars within three months, and developers bought Mac minis to run always-on agents from their desks. In February 2026, Steinberger joined OpenAI.
OpenClaw went viral because it showed people what an agent feels like: a digital worker that takes a goal and gets on with it. Agentic tooling, the Model Context Protocol that connects agents to business systems, and orchestration frameworks are all working in production today. What most organisations lack is a plan to climb from where they are.
03From chatbots to agents
For the past few years, most people have used AI the same way. You open a chat window, ask a question, get an answer, copy it somewhere, then come back with another question. You are the glue between every step: you prompt, you interpret, you move the output, you prompt again.
An agent works differently. You give it a goal and the context, and it operates end to end. It plans, reasons through the problem, calls tools, reads data, writes outputs, checks its own work and delivers a completed result. People set the direction and review the outcome; the agent does the work in between.
An agent is built much like a new member of staff:
- The model is its brain: the underlying AI that reasons. Different models suit different work, and the fastest is sometimes better than the most powerful.
- Instructions are its role and standards: who it is, how it behaves, and what it must never do, such as approving payments above a threshold.
- Skills are its experience: packaged know-how and process for a specific kind of work, which the agent can take on in seconds.
- Tools are its hands: discrete actions it can take, such as searching, reading a database, drafting a document or sending an email, connected to business systems through APIs and the Model Context Protocol.
- Memory is what it learns on the job: the context it keeps about clients, preferences and past work.
Tools carry no judgement about when to use them. The skills and instructions do. That is why designing the work, the roles and the controls matters more than choosing the tool.
04The five levels
Ad Hoc Experimentation “Button mashing”
- What it looks like
- staff using ChatGPT or Copilot on their own. No governance, no strategy, no measurement. Shadow AI everywhere.
- The reality
- sensitive data going into public tools, pockets of productivity, and no organisational leverage.
- Enterprise impact
- high risk, low return.
Assisted Productivity “Following the tutorial”
- What it looks like
- AI is sanctioned, but people are still the glue. You prompt, shape the output, paste it somewhere and prompt again. Every step still needs you.
- The reality
- better outputs, the same workflows, and often more effort. Teams draft emails and summarise documents faster, but a tool has been added to the work and the work itself hasn't changed.
- Enterprise impact
- individual productivity gains, with no structural change. Usage is the only thing to measure. This is the trap.
Embedded AI Workflows “Learning the combos”
- What it looks like
- AI wired into business processes. It is part of the system, working in the background.
- The reality
- service queues classified before a person touches them, compliance checks running in the background, data extraction on autopilot. Work accelerates, errors fall and friction drops.
- Enterprise impact
- the return becomes measurable, because the process is redesigned around AI and value is baselined. The operating model is still the same one: what exists has been optimised.
Agentic Operations “Boss level”
- What it looks like
- you stop giving AI tasks and start giving it goals. Agents plan, reason, call tools, check their own work and deliver completed outcomes end to end.
- The reality
- work completes itself, people supervise and set direction, and output matters more than activity.
- Enterprise impact
- process steps disappear, and roles change. Workforce design becomes part of the AI program.
Orchestrated Intelligence “New game+”
- What it looks like
- many agents working together across functions and systems, improving continuously, with people at the control points. In a transformation program, one agent analyses stakeholders, another maps dependencies, a third drafts the business case, and a human director steers.
- The reality
- end-to-end orchestration, continuous improvement, and systems that optimise themselves.
- Enterprise impact
- the organisation moves differently, and the gap to competitors at Levels 1 and 2 becomes very hard to close.
Shares are Transformativ analysis of the Gartner AI Maturity Model, McKinsey State of AI 2025, Writer Enterprise AI 2026 and the ServiceNow AI Maturity Index 2026.
05Where most organisations sit
| Level | Name | Nickname | Approximate share |
|---|---|---|---|
| 1 | Ad Hoc Experimentation | Button mashing | ~45% |
| 2 | Assisted Productivity | Following the tutorial | ~30% |
| 3 | Embedded AI Workflows | Learning the combos | ~20% |
| 4 | Agentic Operations | Boss level | ~4% |
| 5 | Orchestrated Intelligence | New game+ | <1% |
Shares are Transformativ analysis of the Gartner AI Maturity Model, McKinsey State of AI 2025, Writer Enterprise AI 2026 and the ServiceNow AI Maturity Index 2026.
Three in four organisations are at Level 1 or Level 2. The technology is ready for Level 4 and beyond.
06From assisting the work to doing the work
The move from Level 2 to Levels 4 and 5 is a redesign. AI stops assisting people through their existing work and starts carrying out the work itself, with people directing it. That changes three things:
which roles exist, how work flows, and where people's judgement is needed.
process steps disappear, and the cost of serving each customer falls.
work that took weeks moves in days.
It is rarely about saving twenty minutes on an email. It is about entire process steps disappearing. The organisations that move early improve, and they also pull away.
07How to climb
Three moves carry an organisation from Level 2 to Level 3, and set it up for Levels 4 and 5.
- 1Choose one area and baseline it.
Pick a workstream where value is visible, and measure its current cost, effort, quality and AI spend before anything changes.
- 2Give every task a verdict.
Work with the people who do the work. Each task is automated, augmented or preserved, and the people involved help design what replaces it.
- 3Build measurement in from day one.
Design the target operating model, roles, AI agents and technology together, with a measurement framework that proves value against the baseline.
What you can measure is a function of the level you're at. Catalyst+, our AI transformation product, is built for this climb: a fixed six-week Define sprint that produces a board-ready blueprint, a live Activate stage with value measured from day one, and wave-based Scale. Value stays provable at every step.
Explore Catalyst+Notes and sources
- Uber's AI budget, adoption and per-engineer costs: Forbes, May 2026
- Uber CTO confirmation: The Information, April 2026
- Uber's leaderboard and COO comments: Fortune, May 2026
- Microsoft's move to GitHub Copilot CLI: Windows Central, May 2026
- Microsoft and token-based costs: Forbes, June 2026
- The survey of 20 professionals, and how agents are built (model, instructions, skills, tools, memory): Thomas Paule, "WTF are AI Agents?", AI from the Inside, March 2026
- OpenClaw's origins, growth and Mac mini demand: AWS Builder Center, February 2026
- OpenClaw's creator and OpenAI: Fast Company, February 2026
- 95% of generative AI pilots with no measurable financial impact: MIT Project NANDA, The GenAI Divide 2025 (preliminary findings), via Fortune's coverage
- 70% of the AI challenge in people and process: BCG, 2025
- Level shares: Transformativ analysis of the Gartner AI Maturity Model, McKinsey State of AI 2025, Writer Enterprise AI 2026 and the ServiceNow AI Maturity Index 2026.


