LLM model selection: a key discipline in AI FinOps 

Ruth Dillon-Mansfield | Growth Partner at Plandek

Ruth Dillon-Mansfield

—

Growth Partner

|

LLM Model Selection in DevFinOps

As tokenmaxxing gives way to valuemaxxing, a key question being asked is: what really are the key drivers of improved tokenomics?

The tokenomics challenge

Token spend has increased ten times in the last 12 months and is forecast by Goldman Sachs to grow by another 24 times by 2030.  At the same time, Plandek’s Q4 2026 AI Adoption and Impact Benchmarks taken from data collected from over 2,500 teams, shows the difficulty in translating that increased AI token spend into real engineering productivity improvement.

The most agentic teams achieved a 50% improvement in output (measured in normalized PRs) over the last six months, but a 27% increase in bugs and a 9% reduction in focus on value creating work.  And Lead Time to Value (measuring the total time from ideation to deployment into the live environment) showed only a marginal fall of 4%, reflecting the difficulty in removing underlying bottlenecks across the PSDLC.

In a nutshell, token spend is rising almost exponentially and it is not being matched by similar increases in real output/productivity.

This raises two key questions:  How do you increase underlying engineering productivity as you move to agentic ways of working? And how do you optimize token spend for the tasks in hand?

LLM Model Selection, if done well, makes a big impact on both of these questions.

What is LLM Model Selection?

AI model selection is the practice of matching model capability and cost to the complexity, risk, and value of the task being performed.  

There is clearly a big difference between model capabilities, but there is also a >100x difference in token cost between models and many thousand models to choose from.  Indeed open-source (open-weight) models have no token costs at all (but do have other costs associated with their use) - so the impact of model selection on total AI costs and hence AI ROI is very significant.

LLM Model selection’s role within harness engineering

An AI harness is the wider environment that determines what an AI agent refers to (its context window), what it can do (guardrails), which models and tools it can use, and how its work is controlled and evaluated.

Model selection is a key element of a well structured agent harness and is regarded as one of the biggest drivers of agent effectiveness and cost (and hence tokenomics).

For each class of work, there is a decision to make about the level and type of model capability required. In practice, the options include:

  • Frontier models: the most capable models available at a given point in time, suited to complex reasoning, sophisticated agentic work, or high-risk tasks where the additional capability justifies the inference cost.

  • Smaller or efficient models: lower-cost, lower-latency models that may be sufficient for routine or well-defined work such as classification, extraction, summarization, and routing.

  • Distilled models: models created through model distillation – smaller models trained using the outputs or knowledge of a larger “teacher” model, allowing them to retain strong performance on specific tasks with lower inference cost and latency.

  • Open-weight models: models whose trained weights are available for organizations to run, adapt, or fine-tune themselves, subject to their license. They can provide greater control over deployment and economics.

  • Locally deployed or self-hosted models: models run on an organization’s own or dedicated infrastructure rather than consumed entirely through an external model API. At sufficient scale, this can improve the economics of high-volume workloads and provide greater control over privacy and data residency.

These categories overlap: a smaller model may also be distilled and open-weight, while an open-weight model may be run locally or by a third-party provider. 

Capability and inference cost now vary enormously between models, while the work agents perform ranges from simple classification and extraction through to complex reasoning and autonomous software development.

If every task is routed through frontier-level inference, token spend will rise far faster than the value of the work being produced. 

So, do we push everything to the cheapest possible model? This isn’t the right approach either and headline model pricing can be misleading. Microsoft Research recently found that in 22% of model-pair comparisons, the model with the lower listed price actually incurred the higher total inference cost, with some reversals reaching 28x. Token consumption, not simply the advertised price per token, materially changed the economics. 

The reason is relatively simple – a low-cost model will fail repeatedly and produce output that can’t pass your quality gates, leading to escalation and rework. It ends up being more expensive, in practice.

So the useful economic measure is not cost per token, but cost per successful outcome, in order to optimize AI tokenomics based on real productivity gains.

Learn about harness engineering in software development.

Model selection and model routing

AI workflows tend to make the model selection decision once: choose the most capable model available, wire it into the workflow, and leave it there.  This is clearly inefficient and doesn’t work economically at scale.

In contrast, AI model routing is the real-time, dynamic selection of the most appropriate model for a task based on factors such as complexity, cost, quality requirements, latency, privacy, and risk.

Routing infrastructure can select models according to the work being attempted. Routine, high-volume activity can be sent to smaller models. More demanding work can be escalated to stronger ones. Frontier inference can be reserved for tasks where its additional capability justifies the cost. The optimization problem is increasingly a choice across frontier models, smaller proprietary models, distilled models, and open-weight or locally deployed models.

For high-volume, latency-sensitive, or privacy-constrained workloads, that can materially change the economics.

AI evals are part of the model control system

Routing only works if you know what “good enough” looks like. You cannot safely move work onto a cheaper model without evidence that it can perform that work to the required standard.

AI evaluations, or evals, are systematic tests used to determine whether a model or agent can perform a defined task to an acceptable level of quality, reliability, and safety.

LLM observability is the continuous monitoring of how models perform in production, including their cost, latency, errors, retries, and output quality. Together, evals and observability provide the evidence needed to decide when a lower-cost model is genuinely good enough.

That makes evaluation a core part of AI FinOps. For each meaningful class of work, teams need to define an acceptable quality threshold, test candidate models against it, and then measure what happens in production: retries, rejection, escalation, human intervention, quality failures, successful completion, and the impact on engineering output downstream.

Closing the loop: from model spend to engineering impact

There is one further key consideration. When a well-selected model operates inside a broken engineering workflow, you’re still going to end up with the same downstream constraints – slowed reviews, deployment bottlenecks, and so forth – and hence limited improvement in overall productivity and value creation.

Optimized model selection will come to nothing, if the underlying PSDLC remains inefficient due to problems with your foundational engineering health and DevEx.

This is where Developer Productivity Insight (DPI) tools like Plandek are key.  They provide an overall metrics framework to enable you to track, drive (and benchmark) your underlying engineering productivity (see Plandek’s ProductivityRadar™) – as well as measure the effective roll-out of AI tooling, including model selection.

This integrated metrics framework and related AI transition methodology is provided in the Plandek AI Transition Playbook.

Plandek connects AI usage and cost data to what happens downstream across the SDLC, so teams can see whether model choices are actually producing better engineering outcomes.

In practice, that means being able to answer four questions:

  • Where are the tokens going? By model, team, and type of work.

  • How effectively were they spent? Accepted output, rejected output, retries, human intervention, and work that actually made it into the codebase.

  • Did the token spend actually result in improved engineering productivity? As measured holistically with Plandek’s 4 Pillars of Engineering Productivity which measure focus; speed (including a normalized measure of engineering output); quality; and predictability.

  • What is the AI ROI (tokenomics)? By comparing AI and engineering cost with normalized engineering output, to accurately define real tokenomics expressed in terms of improvement in ‘engineer FTE equivalents’.

In summary, model selection and routing is key to optimize inference cost, but Plandek actually shows whether that optimization translates into improved engineering economics - so that your teams can refine model choice, routing, context, and guardrails based on real delivery outcomes, and ultimately get more useful engineering output from every dollar of AI spend.

For practical guidance on managing your AI transition, download the Software Engineering AI Transition Playbook. Built from lessons across 2,500+ engineering teams, it covers everything from context and harness engineering to identifying constraints, measuring AI impact and proving ROI.


AI Transition Playbook for AI Adoption by Plandek

Written by

Ruth Dillon-Mansfield | Growth Partner at Plandek
Ruth Dillon-Mansfield | Growth Partner at Plandek

Ruth Dillon-Mansfield

Growth Partner

Ruth has 10 years' experience in growth leadership in the tech industry. She has served as COO at a FTSE-listed company, and works with tech start-ups and scale-ups to help their customers understand how to realize value from their products.

See how your engineering efforts translate into measurable business impact

Measure delivery performance, AI impact, and engineering productivity with hundreds of metrics, OOTB dashboards and custom configurations.

  • G2 Fall 2026 Easiest To Do Business With Enterprise
  • G2 Fall 2026 Europe High Performer
  • G2 Fall 2026 High Performer
  • 4.5 Rating from software advice
  • 4.5 rating from  capterra
  • AICPA SOC badge — SOC for Service Organizations compliance certification
  • Crozdesk Happiest Users badge — High User Satisfaction 2023
  • Software Advice Best Customer Support award badge 2022
  • G2 Fall 2025 High Performer Enterprise badge
  • SD Times 100 award badge 2023
  • G2 Fall 2025 Easiest To Do Business With Enterprise badge
  • G2 Fall 2025 Best Support Enterprise badge
  • G2 Fall 2026 Easiest To Do Business With Enterprise
  • G2 Fall 2026 Europe High Performer
  • G2 Fall 2026 High Performer
  • 4.5 Rating from software advice
  • 4.5 rating from  capterra
  • AICPA SOC badge — SOC for Service Organizations compliance certification
  • Crozdesk Happiest Users badge — High User Satisfaction 2023
  • Software Advice Best Customer Support award badge 2022
  • G2 Fall 2025 High Performer Enterprise badge
  • SD Times 100 award badge 2023
  • G2 Fall 2025 Easiest To Do Business With Enterprise badge
  • G2 Fall 2025 Best Support Enterprise badge
  • G2 Fall 2026 Easiest To Do Business With Enterprise
  • G2 Fall 2026 Europe High Performer
  • G2 Fall 2026 High Performer
  • 4.5 Rating from software advice
  • 4.5 rating from  capterra
  • AICPA SOC badge — SOC for Service Organizations compliance certification
  • Crozdesk Happiest Users badge — High User Satisfaction 2023
  • Software Advice Best Customer Support award badge 2022
  • G2 Fall 2025 High Performer Enterprise badge
  • SD Times 100 award badge 2023
  • G2 Fall 2025 Easiest To Do Business With Enterprise badge
  • G2 Fall 2025 Best Support Enterprise badge