Tool-shaped goals
- ai
- adoption
- metrics
Subscribe
New essays and episodes, sent when there is something worth reading. No noise, unsubscribe anytime.
Subscribe
New essays and episodes, sent when there is something worth reading. No noise, unsubscribe anytime.
Salesforce engineers have a widget on their Macs that refreshes every fifteen minutes with the amount of money they have spent on AI tools. The Pragmatic Engineer reported that the widget also displays a minimum expected spend. The week the story broke, the target was a hundred US dollars on Claude Code and seventy on Cursor. The meter is not there to keep the number down.
And it is not one company being strange. A Disney employee interacted with Claude 460,000 times in nine days. At Meta, an employee built a leaderboard ranking more than eighty-five thousand colleagues by token consumption, with titles like Token Legend at the top. Nobody in leadership asked for it, and Meta said so pointedly once journalists noticed. By then the metric had soaked in deeply enough that the workforce was gamifying itself, unprompted. Engineers already have a name for the behaviour these systems reward: tokenmaxxing. Point an agent at work you would finish faster by hand, ask the model questions you already know the answer to. The leaderboard does the rest.
The logic behind all this deserves a fair reading, because it is not stupid. The capability jump is real. Adoption inside a large company is uneven; one team rebuilds its whole workflow while the team next to it waits for the fad to pass. Leaders who believe the technology matters cannot tell, from the outside, who is transforming and who is stalling, and waiting for organic diffusion feels like watching competitors pull away. So they reach for the oldest lever in management: what gets measured gets managed. Microsoft told managers that AI use was no longer optional and to weigh it when evaluating performance. Amazon employees admitted to running AI on tasks that did not need it, purely to lift their usage scores. If uptake is the thing you can see, uptake is the thing you push.
Everyone reading already has the diagnosis loaded. Goodhart's Law: "when a measure becomes a target, it ceases to be a good measure." Charles Goodhart wrote the original version in 1975 about monetary policy, and the phrasing everyone quotes came from the anthropologist Marilyn Strathern in 1997, watching the same disease work through the audit culture then spreading across British universities. Metric becomes target, target gets gamed, everyone nods.
But the Goodhart diagnosis is too comfortable, because it explains the wrong end of the problem. It explains the engineers. It does not explain the executives. The people who set minimum spend targets and wrote AI use into performance reviews manage incentive systems for a living. They have watched sales teams sandbag quarters and call centres hang up on customers to protect handle-time targets. Nobody at that level is innocent of Goodhart. So the real question is not how the metric got gamed. It is why people who knew better set it anyway.
Part of the answer is that adoption is countable and value is not. An adoption percentage travels well, into a board pack or an investor call, and I have written before about how the countable thing grows because it is countable. But legibility only explains why adoption was available as a metric. It does not explain why it got promoted into a goal.
The stronger explanation has a name in the management accounting literature: surrogation. Michael Harris and Bill Tayler described it in Harvard Business Review in 2019. Strategy is abstract, metrics are concrete, and under pressure people mentally replace the strategy with the measure until the measure is the strategy. Their example was Wells Fargo, where a strategy of deepening customer relationships collapsed into a cross-selling target, and employees opened 3.5 million unauthorised accounts, incinerating the exact relationships the number was supposed to stand for. Surrogation is usually told as a story about staff corrupting leadership's pure intent. The AI version runs upstream. Leadership surrogated first. Somewhere between the strategy document that said AI could transform how the company serves customers and the OKR sheet that said drive adoption, the transformation collapsed into the technology. The goal arrived on employees' desks already tool-shaped, and at Meta the substitution had run so deep that the employees built the scoreboard themselves.
A tool-shaped goal is easy to spot, because it has a technology's name in it. Increase AI adoption. Get every team on Copilot. Nobody sets forklift adoption targets; warehouses filled with forklifts because pallets moved per hour was the goal and forklifts moved it. A tool earns its place in a business by shortening the distance to an outcome. When the tool appears inside the goal itself, the reasoning has inverted, and the tool is borrowing authority from an outcome nobody wrote down.
Under a tool-shaped goal, tokenmaxxing stops looking like misbehaviour. The engineer running redundant prompts to stay off the bottom of a leaderboard is executing the strategy faithfully, exactly as encoded. Goodhart's Law describes good measures corrupted by targeting. These measures were never good. They were born as surrogates, and the gaming just made the substitution visible.
There is an obvious pushback: early on, doesn't raw exposure have value? It does. People cannot find use cases in a tool they have never touched, and a structured push to get hands on the thing is legitimate. But that is an argument for training, and training has a natural end. The difference between "everyone spends two days learning what this can do" and "your token count is in your performance review" is the difference between education and surrogation. One builds capability and expires. The other renews the confusion of means with ends every review cycle.
The version without the tool in the goal runs roughly in reverse order to the leaderboard playbook. Find where the technology fits the processes the business already runs, and accept that in some of those processes it currently fits nowhere. The people who run those processes get trained on their own work. Then the expectations on those teams go up, denominated in the units the process was already judged by: turnaround time, error rate, cost per case, time to answer a customer. Paul David's study of factory electrification is the standing reminder of where technology value actually lives. Electric power was available from the 1880s, and the productivity surge arrived in the 1920s, once factories stopped bolting motors onto steam-era layouts and rebuilt production around what small motors made possible. The redesign was the value, and the redesign took forty years. The dynamo-era version of a token leaderboard would have ranked factories by kilowatt-hours, and it would have started the same fires.
Raised expectations are the honest form of adoption pressure. A team that has the tool and the training should be asked for more, and a team asked for more will pull the tool in without a leaderboard, the way warehouses pulled in forklifts. If the gains exist, they show up in numbers the business was already tracking, numbers a customer would recognise as service getting better. If they show up nowhere except the usage dashboard, that is the metric confessing there was nothing else to show.
Which is what the widget was displaying all along, every fifteen minutes, with perfect accuracy: a number going up that no customer will ever see.