AI Tokenomics – the new line item CFOs can't ignore

Profile Pic of Charlie Ponsonby

Charlie Ponsonby

Co-founder & CEO

|

AI Tokenomics – the death of tokenmaxxing and why AI ROI matters

A career-defining moment for CTOs

Over 90% of AI token spend tends to be consumed by technology – particularly software engineering – teams within organizations. As token usage scales, these costs are becoming a meaningful technology cost line, making the economics of token use increasingly important.

Board members therefore need to understand how AI spend is translating into engineering output, particularly as agentic software engineering becomes more widespread, and where additional spend stops delivering meaningful returns.

It was all about ‘tokenmaxxing’

CTOs are navigating a period of incredibly rapid change, with no clear playbook and new tools and methods emerging weekly.

As a result, the approach has understandably focused on ‘test and learn,’ with ‘tokenmaxxing’ emerging as one strategy – encouraging engineers to jump in, experiment with AI tooling, and use tokens aggressively. Tokens are a common unit of consumption for AI models and, particularly for API-based services, an important driver of cost.

This has undoubtedly led to some spectacular jumps in learning and productivity. But it has also produced striking examples of wasted effort and cost. At Rippling, for example, one engineer was reportedly spending $50,000 per month on AI tokens as the company’s overall AI spend accelerated rapidly.

By July 2026, most major organizations had begun moving away from the ‘spray and pray’ approach to tokenmaxxing. Indeed, both Uber and Tesla announced controls on token usage across their engineering teams over the summer.

Suddenly, it’s all about ‘tokenomics’

With aggregate token spend rising rapidly in some organizations – and often little clarity about the additional value being generated – ‘tokenomics’ is becoming an increasingly important topic for technology leaders and boards.

In this context, tokenomics means understanding the economics of AI token consumption: where tokens are being spent, what they cost by activity or team, and how that spend relates to changes in engineering output and productivity.

This is not as easy as it sounds because it opens up a very old and controversial topic in software engineering – the measurement of productivity and value.

Measuring output and ROI in software engineering has always been controversial

It is worth first distinguishing between output and value creation. Software engineering output is itself difficult to measure, but ‘output’ clearly differs from ‘value.’ You can build a perfectly functional piece of software that is never monetized or creates little meaningful business value.

Measuring output in software engineering has always been controversial. Unlike building cars, where you can count clearly defined units of output, software engineering is more complex. Lines of code are meaningless as a measure of productivity – a book is not better because it contains more words.

However, observable units of engineering activity and output do exist, including completed work items and pull requests (PRs). Used carefully, normalized appropriately, and interpreted alongside measures of speed and quality, these can provide useful proxies for changes in engineering output.

With the improved data available from the latest generation of Developer Productivity Insight (DPI) tools, it is now possible to measure changes in software delivery output and efficiency – and use those changes to build a practical proxy for AI ROI.

Why you have to act now to measure AI ROI

As AI adoption scales, aggregate token spend can rise extremely quickly. At the same time, significant differences are emerging between teams that use tokens efficiently and those where additional consumption produces little additional output – and may introduce unnecessary cost or risk.

Now is therefore the time to establish effective token economics across your organization.

Fail to act, and AI costs can escalate while pockets of inefficient or risky token use proliferate, increasing both cost and AI agent risk. Managing token consumption is also part of a broader set of AI agent guardrails that should be in place across every organization to safely accelerate the transition toward agentic ways of working.

A simple approach to tracking ‘tokenomics’ – and AI ROI – that works at scale

You can now measure software engineering output using a DPI tool, but directly measuring the business value of everything produced is far harder and, in our view, not viable at scale.

The value of a piece of software output may be relatively easy to calculate in a limited set of circumstances, such as a change to a feature that can be A/B tested – for example, changing a webpage call to action.

But much software output is either not directly customer-visible – such as technical debt reduction or foundational work – or contributes to customer outcomes alongside many other variables. Revenue may be affected by pricing, seasonality, competition, marketing, and numerous other factors. You therefore cannot reliably infer that a particular piece of software generated the revenue that subsequently accrued from its use.

Our recommended approach to assessing the returns from AI use in software engineering therefore adopts a marginal-returns approach. Rather than attempting to attribute revenue directly to each unit of software output, it examines changes in the cost of production relative to changes in measured engineering output.

We can then calculate how much engineering output is being delivered for every dollar spent on input costs, including AI tokens – and identify the point at which further increases in cost are no longer producing a material increase in output.

One unit of software output we use is the pull request (PR) – a proposed change to a codebase that is submitted for review and merge as part of the software delivery process.

The Pull Request Quotient combines merged PR output, time to merge, and engineering team size to provide a normalized view of PR output per engineer. As with any PR-based measure, it should be interpreted alongside measures of quality and other delivery signals rather than treated as a standalone measure of value.

The Pull Request Quotient – a measure of output in software engineering

PR Quotient formula
  • This measures code flow efficiency per engineer over a quarter

  • Time to Merge captures review bottlenecks that raw PR counts mask

  • As agent use grows, a rising quotient shows the system enables more code to be merged relative to the team size

Once we have a consistent measure of software engineering output that can be surfaced across teams using a DPI tool such as Plandek, we can observe how that output changes relative to changes in production costs – including additional AI costs.

From this, we can create an ROI-Adjusted PR Quotient that shows the marginal impact of AI spend on engineering output – in other words, the relationship between additional AI cost and additional output.

This provides a practical and scalable way to assess token economics across engineering teams and identify where AI spend is producing additional output – and where it is not.

The ROI-Adjusted Pull Request Quotient for calculating tokenomics in software engineering

Further reading – now you can measure tokenomics, how do you improve it?

The next challenge is to improve tokenomics – and that means carefully managing the factors that determine how effectively agents work. As organizations move further into agentic software development, this increasingly means looking beyond the model itself at the context, controls, and execution environment around it.

Coding agents are highly capable, but their effectiveness depends heavily on the context and environment in which they operate. Poor or outdated context, weak controls, inadequate tooling, and ineffective AI agent feedback loops can all lead to poor output, rework, and wasted token spend.

Two increasingly important disciplines are therefore context engineering and harness engineering. I explained these concepts here.

Context engineering is the discipline of ensuring AI coding agents have accurate, relevant, current, and appropriately governed information about the codebase, business rules, and engineering process before they act. ‘Freshness’ is one useful measure here – are agents working from up-to-date information rather than stale documentation or assumptions?

If context engineering is what the agent knows, harness engineering is the environment that determines how the agent works. The agent harness includes the tools, permissions, guardrails, test gates, feedback loops, and execution controls used to guide and validate agentic work before it reaches production.

This also makes testing AI agents, agent evaluation, and agent observability increasingly important. Organizations need to understand not only whether an agent completed a task, but whether it followed the right process, used the right context, and produced output that meets the required quality and governance standards.

Written by

Profile Pic of Charlie Ponsonby
Profile Pic of Charlie Ponsonby

Charlie Ponsonby

Co-founder & CEO

Charlie Ponsonby is CEO and Co-founder of Plandek, the leading Developer Productivity Insight (DPI) platform that helps software engineering teams drive productivity and transition to AI-led engineering. He writes widely on the opportunities and challenges inherent in the transition to the agentic SDLC. Prior to founding Plandek, Charlie founded Simplydigital, which grew to become the UK's largest broadband and digital services comparison business before being acquired by Europe's largest consumer electronics retailer. He started his career at Accenture and has held senior leadership roles in retail and telco. Charlie holds a degree from the University of Cambridge.

See how your engineering efforts translate into measurable business impact

Measure delivery performance, AI impact, and engineering productivity with hundreds of metrics, OOTB dashboards and custom configurations.