WRWriting

Right-Sizing Intelligence: The New Economics of Enterprise AI

AI FinOps starts with matching model capability to workload demand and measuring cost per successful outcome—not merely tokens or cost per request.

AI / FinOpsAugust 10, 202618 min read

Right-sizing intelligence by routing simple workloads to efficient models and escalating complex work to more capable tiers.

Cloud taught us not to run every workload on the biggest server available. Enterprise AI is about to learn the same lesson with models.

For years, cloud architects have repeated a simple principle:

Use the right resource for the workload.

You would not normally run a lightweight batch process on the most powerful GPU instance available.

You would not provision the largest database tier for an application serving 50 users.

You would not reserve premium compute for workloads that can comfortably run on commodity infrastructure.

We call that right-sizing.

Yet as enterprises move generative AI into production, many are doing the equivalent of exactly what cloud architecture taught us not to do.

They choose a flagship model.

Then everything goes to it.

Classification. Summarization. Extraction. Coding. Customer support. Document analysis. Research. Reasoning. Agent tasks.

The logic is understandable. The strongest model is more capable, so using it everywhere feels like the safest architectural choice.

But at scale, it can become economically irrational.

AWS's Agentic AI Well-Architected guidance recommends matching model capability to task demand and using the smallest model that still meets the workload's quality bar rather than sending every task to a heavyweight model.

That principle deserves a name.

I call it:

Right-sizing intelligence.

And I think it will become one of the central FinOps disciplines of enterprise AI.

Tokens are becoming a new unit of infrastructure consumption

Cloud computing trained finance and technology teams to understand infrastructure through units such as CPU hours, gigabytes, requests, storage, and network transfer.

Generative AI adds another increasingly important consumption unit:

Tokens.

But tokens behave differently from conventional infrastructure units.

Two requests that appear identical from a business perspective can have dramatically different costs depending on:

  • the model selected;
  • prompt and output length;
  • reasoning effort;
  • context size and caching;
  • service tier;
  • tool use and agent steps;
  • retries;
  • provider.

Current model catalogs make the spread visible.

Anthropic currently lists Claude Haiku 4.5 at $1 per million input tokens and $5 per million output tokens, while Claude Opus 5 is priced at $5 and $25 respectively. OpenAI's GPT-5.4 nano is priced at $0.20 per million input tokens and $1.25 per million output tokens.

Those models are not equivalent, and comparing price alone would be meaningless.

But that is precisely the point.

Different levels of intelligence have materially different economics.

At prototype scale, those differences can seem trivial.

At enterprise volume, they are not.

Consider an illustrative workload processing one million requests per month, each consuming an average of 1,000 input tokens and 300 output tokens.

At GPT-5.4 nano's published rates, that token volume would cost roughly $575.

At Claude Opus 5's published rates, the same token volume would cost roughly $12,500.

That is about a 22× difference in inference cost before considering caching, tool calls, service tiers, discounts, or different model behavior.

Again, this is not an argument that the smaller model can replace the larger one for the same workload.

The financially important question is:

How many of those million requests actually require the larger model?

That is where AI FinOps begins.

Cost per successful outcome

There is an important distinction between price and economics.

Suppose a lower-cost model answers a task correctly 85% of the time and a premium model succeeds 98% of the time.

If failures trigger human review that costs $20 each, the cheaper model may ultimately be far more expensive.

Conversely, if the task is simply classifying support tickets into five categories and both models achieve the required quality threshold, paying a 10× or 20× inference premium may create no additional business value.

So the metric should not be cost per token.

It should increasingly become:

Cost per successful outcome.

A useful simplified formula is:

Cost per successful outcome =
Total AI execution cost ÷ Outputs meeting the acceptance threshold

For agentic systems, I would expand it further:

Total execution cost =
model cost + tool cost + retries + human review + downstream error cost

That last component matters.

A cheap model that causes expensive mistakes is not cheap.

A premium model used on trivial work is not efficient.

The objective is neither minimum cost nor maximum intelligence.

It is the lowest total cost that reliably produces the required business outcome.

This is why evaluation becomes part of FinOps

Traditional FinOps can often optimize a resource by examining utilization.

CPU usage is 8%. The instance is oversized. Right-size it.

AI is harder.

A smaller model may cost less, but you cannot safely move the workload until you know whether quality survives the change.

That means AI cost optimization requires something FinOps teams have historically not had to own:

Evaluation data.

Microsoft's model-selection guidance recommends testing candidate models on representative data using criteria that include accuracy, output quality, speed, and cost. Its Model Router guidance similarly advises benchmarking routing against a baseline on quality, cost, and latency before moving production traffic.

This produces a new optimization loop:

PRODUCTION WORKLOAD
        |
        v
   Current Model
        |
        v
   Measure Cost
        |
        v
  Evaluate Output
        |
   +----+----+
   |         |
Passes?    Fails?
   |         |
   v         v
Try a      Escalate
cheaper    capability
model
   |         |
   +----+----+
        v
Cost per successful outcome

Now FinOps and AI evaluation begin to overlap.

That is an important organizational change.

Not every prompt deserves the same intelligence tier

Imagine an enterprise running the following workloads.

Classification

“Which department should handle this support ticket?”

This may require little reasoning and highly structured output. An efficient model may be the correct choice.

Extraction

“Extract the invoice number, supplier name, date, and amount.”

Again, the problem is bounded and easily evaluated. Potentially a small model.

Summarization

“Summarize these meeting notes into five actions.”

Perhaps a balanced general-purpose model.

Complex analysis

“Examine these four financial scenarios and identify hidden assumptions that could materially change the acquisition model.”

Now deeper reasoning may justify substantially higher inference cost.

Cybersecurity investigation

“Correlate these IAM, CloudTrail, endpoint, and network events and produce competing hypotheses for the attack path, with evidence and uncertainty.”

Quality matters enormously. A premium reasoning model may be entirely justified.

Restricted-data processing

“Analyze these confidential customer documents.”

Here, the primary routing factor might not be price or capability at all. It could be security policy requiring private-cloud or local inference.

The economic model therefore looks more like:

Classification ──────────► Efficient Model
Extraction ──────────────► Efficient Model
Routine Summarization ───► Balanced Model
Complex Reasoning ───────► Premium Model
High-Risk Decisions ─────► Premium + Human Review
Restricted Data ─────────► Private / Local Model

This is not merely model selection.

It is workload placement for intelligence.

The intelligence right-sizing ladder maps efficient, balanced, premium, and private model tiers to workload requirements.

The cloud analogy is stronger than it first appears

When cloud adoption accelerated, organizations frequently made a predictable mistake.

They migrated workloads without changing the operating model.

An application that once ran on oversized physical infrastructure simply moved onto oversized virtual infrastructure.

Cloud offered elasticity. Organizations initially consumed it inefficiently.

FinOps emerged because the cloud made consumption easier than governance.

AI may repeat the pattern.

Inference APIs make intelligence extraordinarily easy to consume. A developer can add a model call with a few lines of code. An agent can call a model repeatedly. One agent can invoke other agents. Those agents can retrieve more context, call tools, fail, retry, reflect, and invoke another model.

Each decision can consume more tokens.

AWS's 2026 Agentic AI Lens warns that a single request can trigger multiple inference calls, tool invocations, memory retrievals, and inter-agent communications, each adding cost, latency, and failure surface.

The danger is not simply that one prompt is expensive.

It is that intelligence becomes programmatically consumable at enormous scale.

That should sound familiar to anyone who watched the early years of cloud.

Agentic AI makes the economics more interesting

A conventional chatbot may involve:

one user request → one model response

An agentic workflow may look like:

User Request
     |
     v
Planner Model
     |
     +──► Research Agent
     |       +──► Model Call
     |       +──► Search
     |       +──► Model Call
     |       └──► Model Call
     |
     +──► Data Agent
     |       +──► SQL Tool
     |       └──► Model Call
     |
     └──► Review Agent
             +──► Model Call
             └──► Revision

The user sees one answer.

Finance may see ten or twenty inference operations underneath it.

This is why per-seat pricing alone will not tell enterprises what their AI workload actually costs.

The relevant unit may eventually become:

  • cost per case resolved;
  • cost per code issue completed;
  • cost per customer conversation;
  • cost per report generated;
  • cost per incident investigated;
  • cost per autonomous task completed.

Microsoft's Agent Monitoring Dashboard tracks system-level outcomes such as token usage, latency, success rates, and evaluation results for production agent traffic.

That is exactly where enterprise AI economics need to go.

The first lever is model routing

Right-sizing begins with the simplest question:

Can a less expensive model complete this task at the required quality?

Increasingly, platforms can make that decision dynamically.

Amazon Bedrock intelligent prompt routing predicts which supported model can deliver the desired response quality while balancing cost, then sends each request accordingly.

Microsoft Foundry's Model Router evaluates prompt complexity and can optimize routing for cost, quality, or a balance of the two.

OpenRouter's Auto Router similarly evaluates attributes such as prompt complexity, task type, and model capability, while its provider-routing controls can prioritize price, latency, throughput, or fallback behavior.

This is the architectural foundation established in Article 1: Beyond One Model.

Once model selection is abstracted away from the application, economics can become part of routing policy.

Routing should escalate, not simply downgrade

There is a subtle but important difference between cost optimization and cost cutting.

A crude optimizer says: route everything to the cheapest model.

A mature one says: start with the least expensive approved intelligence likely to succeed, then escalate when evidence indicates greater capability is required.

Incoming Request
      |
      v
Efficient Model
      |
      +── High confidence + validation passes ──► RETURN
      |
      v
Balanced Model
      |
      +── Evaluation passes ────────────────────► RETURN
      |
      v
Premium Reasoning Model
      |
      +── High-stakes decision? ────────────────► HUMAN REVIEW
      |
      v
    RETURN

This creates something analogous to tiered storage or autoscaling.

Intelligence scales up when the workload demands it.

That is very different from permanently provisioning every task at the highest capability tier.

Caching may be the easiest AI savings opportunity

Not every optimization requires switching models.

Many enterprise prompts contain large blocks of repeated context: system instructions, policy documents, tool definitions, coding conventions, product catalogs, and long conversation histories.

If the same prefix is sent repeatedly, paying to process it from scratch every time can be wasteful.

Amazon Bedrock prompt caching is explicitly designed to reduce both inference latency and input-token cost by reusing supported prompt prefixes. OpenRouter prompt caching similarly uses provider-sticky routing to preserve caches when doing so produces a cost benefit.

For a high-volume enterprise workload, caching strategy can become as relevant to AI economics as application caching became to web architecture.

The FinOps question becomes:

Why are we repeatedly paying the model to read something that has not changed?

Batch and latency tolerance have financial value

Not every AI request is interactive.

A user waiting for a chatbot response values latency. A nightly workflow classifying 500,000 documents may not.

That distinction creates another right-sizing dimension.

OpenRouter service tiers let supported inference trade latency and availability for lower price. Microsoft similarly distinguishes among standard, batch, priority, and provisioned deployment approaches based on latency and capacity requirements.

Urgency has a price.

If the workload can wait, it should not necessarily pay interactive-workload economics.

Prompt design is also an infrastructure decision

There is a temptation to regard prompting solely as application logic.

At enterprise scale, it is also a cost decision.

A 20,000-token system prompt invoked ten million times is infrastructure consumption.

So is unnecessary chat history.

So is returning 4,000 tokens when the application needs 400.

AWS's Agentic AI cost guidance recommends prompt compression, explicit token budgets, caching, and documented routing policies for high-volume systems.

That makes AI cost optimization partly a software-engineering discipline.

The cheapest token remains the token you did not need to process.

Local inference changes the equation again

Article 1 introduced three execution environments: public AI, private-cloud AI, and local AI.

Cost makes that choice more nuanced.

Hosted APIs have compelling economics when demand is low, variable, experimental, or difficult to predict. The enterprise consumes exactly what it needs without owning GPU capacity.

Self-hosted inference introduces a different model. Now costs may include:

  • GPU acquisition or leasing;
  • utilization and power;
  • networking and operations;
  • model serving and engineering;
  • redundancy.

At low utilization, self-hosting can be wasteful.

At sustained high utilization, the economics may become more interesting.

But the correct comparison is not API tokens versus GPU price.

It is:

Fully loaded cost per successful inference outcome.

This should include the people and infrastructure required to operate the environment.

That is another lesson FinOps already learned from cloud:

Owned infrastructure is not free merely because there is no per-request invoice.

From tokens to business outcomes

Imagine a team reports: “We reduced AI token spend by 40%.”

That sounds excellent.

Then you discover customer-service resolution quality dropped, escalations doubled, and employees are manually correcting model output.

The infrastructure number improved.

The business economics deteriorated.

This is why the right AI FinOps dashboard cannot stop at token consumption, model spend, and cost per request.

It needs to include quality, completion rate, human escalation, retries, latency, and business outcome.

A useful metric hierarchy might look like:

LEVEL 1 — CONSUMPTION
Tokens · Calls · Tool invocations

LEVEL 2 — UNIT COST
Cost/request · Cost/session · Cost/agent run

LEVEL 3 — QUALITY
Success rate · Evaluation score · Escalation rate

LEVEL 4 — BUSINESS ECONOMICS
Cost per successful outcome
Cost per case resolved
Cost per customer served
Cost per engineering task completed

The farther down that stack an organization can measure, the more meaningful its optimization becomes.

AI FinOps maturity progresses from token consumption through unit cost and quality to cost per successful business outcome.

Finance cannot manage AI costs alone

There is an organizational implication to all of this.

Finance knows what was spent.

Engineering knows how the application works.

AI teams know why particular models were chosen.

Product knows what outcome matters.

Security knows where the data may go.

No single one of those groups can optimize AI economics independently.

The operating conversation needs to become:

Finance: What does the workload cost?

Engineering: What drives that cost?

AI/ML: What is the minimum model capability that meets the quality requirement?

Product: What outcome are we actually paying for?

Security: Which execution environments are permissible?

That begins to look like AI FinOps, not merely cloud cost management.

And just as traditional FinOps created a shared language between engineering and finance, AI FinOps will need a shared language around intelligence consumption.

The control plane needs an economic policy

This connects directly back to Article 1.

If the AI control plane knows task type, data classification, approved models, model quality history, latency target, and provider status, then it can also know request budget, cost ceiling, expected business value, and acceptable escalation path.

The policy can become something like:

WORKLOAD: Support ticket classification

Allowed models:
  Efficient tier
  Balanced tier

Quality threshold:
  >= 96%

Latency:
  < 2 seconds

Maximum cost:
  $0.002 per ticket

Escalation:
  Balanced model if confidence < threshold

Human review:
  Required only for unresolved category

Now finance is no longer trying to optimize a monthly AI invoice after the fact.

Economics have moved into the execution policy itself.

That is a major architectural shift.

The AI FinOps feedback loop routes, executes, evaluates, measures cost, and updates routing policy using production outcomes.

Budgets may eventually become runtime controls

Cloud budgets historically began as reporting.

Then came alerts.

Then automated policies.

AI could move faster.

Imagine an application operating under a monthly intelligence budget.

As utilization changes, the router might:

  • increase caching;
  • move low-priority jobs to batch;
  • route simple workloads to cheaper models;
  • cap unnecessary output;
  • delay nonurgent jobs;
  • preserve premium models for high-value workloads.

Not indiscriminately.

Within predefined quality and security boundaries.

That is the important condition.

A budget should constrain waste, not intelligence required for the business outcome.

The goal is governed efficiency.

Not artificial scarcity.

What enterprises should start measuring now

Before attempting sophisticated routing, organizations need visibility.

I would start with the following measurements:

Spend by application, model, and business unit

Which applications consume the most? Where is premium intelligence being used? Who owns the consumption?

Spend by task type

Classification? Coding? Research? Summarization? Agent execution?

Average model calls per user request

Especially important for agents.

Input-to-output token ratio

Are enormous contexts being sent for small answers?

Latency by model and workflow

Is additional cost actually improving user experience?

Success or evaluation score by model

Which cheaper models meet the same acceptance threshold?

Cost per successful outcome

Ultimately, the metric that matters.

Without those measurements, AI optimization becomes guessing.

What I would not optimize first

There is also a danger of trying to FinOps AI too early.

If an organization has one experimental application generating $500 per month in API charges, building an elaborate routing platform probably costs more than it saves.

Optimization should follow materiality.

Start simple:

  1. Understand consumption.
  2. Tag it by owner and workload.
  3. Identify the large recurring workloads.
  4. Build an evaluation set.
  5. Test smaller models.
  6. Introduce routing where the economics justify it.

AWS's Well-Architected guidance similarly frames model selection and cost optimization as workload-specific rather than assuming that every system needs the same optimization techniques.

FinOps exists to improve economics.

Not to create architecture for architecture's sake.

Right-sizing intelligence is not about using weaker AI

This distinction is worth making explicit.

The premise of this article is not: use cheap models.

It is:

Use expensive intelligence where expensive intelligence creates value.

If a frontier reasoning model prevents a multimillion-dollar error, the token bill is irrelevant.

If the same model is categorizing employee expense receipts, there may be a better economic choice.

Good AI FinOps preserves premium capability.

It simply stops wasting it.

My takeaway

The first phase of enterprise AI was about proving that generative models could create value.

The next phase will be about producing that value economically.

That requires moving beyond the idea that every AI request deserves the same model.

Cloud infrastructure taught us an important lesson:

Scarce, expensive resources should be matched to workloads deliberately.

AI intelligence should be treated the same way.

The smallest model that meets the required quality threshold may be the right model.

The premium model should remain available when complexity warrants it.

Batch workloads should not pay real-time economics when they do not need to.

Repeated context should be cached.

Agent loops should be measured by outcomes, not hidden behind a single user request.

And ultimately, organizations should stop asking:

How much are we spending on tokens?

and start asking:

How much intelligence did we need to produce a successful business outcome?

That is a much harder question.

It is also the one that matters.

We learned to right-size compute.

Now we need to right-size intelligence.

Next in the series

Article 3 — The AI Governance Gap: Who Decides Which Model Gets Your Data?

Multi-model architecture gives the enterprise choice.

Choice creates a new problem.

Who determines which providers are approved? Which data can leave the organization? When must inference stay private? How do we know why a router selected a particular model—and can we prove it six months later?

The next article will explore why dynamic model routing inevitably becomes a governance problem.

References