Learn Hub | G2

How Enterprises Measure LLM Cost: 53% Still Have No Formal Metric

Written by Emilie Audubert | Oct 9, 2026, 2:10:09 PM

Most enterprise leaders don’t evaluate a large language model (LLM) solely on its cost-per-token pricing. Instead, they consider its potential savings. G2 interviewed 102 US enterprise leaders for its AI Custom Research, and the results back this up. Of the 98 who described how they compare models on cost, 60% look at ROI and labor savings rather than price per million tokens. Yet when asked about cost per outcome, which is a formal way to measure the price of one finished piece of work, 53% of the 96 who answered said they have no formal metric for it.

Weighing savings against cost is a sensible way to make a purchase decision. But G2’s analysis suggests many organizations are making that call on judgment more than measurement. Judgment becomes harder to sustain once the bill reaches finance.

This article walks through what LLM cost per outcome means, how the leaders who track it do the math, where enterprise LLM spending starts to block scale, how teams keep costs in check, and whether enterprises measure ROI.

What is cost per outcome for LLMs, and why does it matter?

Cost per outcome is what a company pays to get one specific result from an LLM, such as a resolved support ticket, a processed document, or a closed case.

A usage bill tells finance how much text was processed by a model. Cost per outcome tells them how much a finished piece of work costs, so they can set it against the cost of doing the same work by hand.

Buyers rarely start there. Of the 102 enterprise leaders G2 interviewed, 68 used the word “token,” most often when asked about the bill and about scaling. Tokens are the small units of text vendors use to meter every prompt and response. Yet only one leader reasoned from the advertised price per million tokens, so the number on the pricing page is rarely the one buyers use.

G2’s analysis points to a simple logic: the lowest price per million tokens isn’t the best deal if the cheaper model needs a rerun or a person to fix its output.

How do enterprises measure the cost of LLMs?

G2 Data shows enterprise leaders mostly measure LLM cost by what a model gives back, not by what it charges. Of the 98 leaders asked how they compare models on cost, 60% look at ROI and labor savings, while 23% watch total spend and token efficiency instead.

The math is simplest when a team has a unit to count. A senior director of AI and automations at a SaaS company pays $1 per ticket an AI support product touches, against at least $50 for a human interaction once all costs are counted. The decision, the director said, was “a no-brainer for us.” An enterprise SaaS team takes a similar approach: it compares its current cost per ticket with what a ticket cost before AI, to confirm it’s on the right path.

Fewer than a third of leaders have a unit like that. G2 asked 96 of them whether their organization measures what one completed piece of work costs:

53%

of enterprise leaders have no formal way to track the cost of a finished piece of work.

 

Source: G2 Data

Buyers have chosen the right comparison, payoff against price, but most make it by judgment and not by measurement. Only the 29% who price each unit can put a firm number on their AI ROI.

When does LLM cost start to block enterprise scale?

LLM costs start to cap scale when usage grows faster than the budget supporting it, and about one in four leaders G2 interviewed has already reached that point. Of the 91 leaders who discussed cost per outcome, 26% described having to cap spending, scale back a project, or move to a cheaper model because of what the AI token costs added up to.

G2’s interviews show how fast it can happen. One enterprise SaaS team scaled up its AI use, and usage spiked across its developers and employees. The team described an “unexpected ramp up in cost,” capped spending for the year, and plans to add limits for each user and each team.

Others pulled back. A technology CEO said they “closed down several projects because they were running above cost.” A senior director of managed cloud delivery at an IT services firm described a change of mind: the company is moving from using AI for everything to using AI where it makes sense, because “the bill has gone very high very quickly as we push that out.”

Those are individual cases. The 26% is the measured share.

Many leaders won’t see the limit coming. About half have no threshold at which the numbers would tell them to expand a rollout or stop one. When the bill does get attention, it lands with finance, the CFO, or the CIO, for about three-quarters of the leaders who named an owner. A spike in usage is what gets noticed first.

How do enterprises control and reduce LLM costs?

Enterprises hold LLM costs down in three ways: by limiting how much gets used, and by paying less for the same work, and by changing how they are billed. Most of the examples below came from single teams, so G2’s interviews show a patchwork more than a playbook, which is what LLM cost management looks like today.

Some teams hold costs down by using less:

  • Spend caps: The enterprise SaaS team described above capped its total budget for the year and plans limits for each user and each team.
  • Pulling the plug: A construction IT director turned off some AI capabilities based on token usage, one of them because it “became a significant burden.”

Others try to pay less for the same work:

  • Trading down: An enterprise SaaS team moved coding and planning work to a cheaper tier and found it delivered the same value.
  • Tuning: The SaaS team that tracks cost per ticket adjusts its model choice, caching, and the context it sends.

And some buyers want to be billed differently:

  • Asking for subscriptions: Some buyers push vendors for subscription-style plans in place of counting tokens.

One measured figure fits the picture: of the 81 leaders asked what they’ve built on top of their vendors’ controls, about a quarter have added spend limits and human oversight.

Cost is the reason behind a minority of model changes, though. Of 84 leaders who described their most recent change, 57% said performance or capability prompted it, and 15% said cost did. G2’s analysis suggests even a major price drop would leave the bigger constraints in place: security, reliability, and governance.

Do enterprises measure ROI on LLMs?

Most don’t. Only 29% of the 96 leaders G2 asked put a price on each completed unit of work, and 18% use a rough stand-in, such as time saved. Enterprise leaders judge LLM value more often than they calculate it.

Who you ask matters. Of the 18 leaders who own the budget or approve spending, half measure cost per task. Among the 61 who run the model tests, a quarter do, and more than half have no formal metric at all. The budget-owner group has fewer than 30 people, so treat the gap as a direction, not a precise figure. It isn’t a clean split, either: about a third of those budget owners have no formal metric.

"When 53% of enterprise leaders lack that metric, they manage inputs rather than outcomes."

Godard Abel
CEO, G2.

Where does that leave enterprise LLM buyers?

Token prices matter. For a quarter of enterprise leaders, the experiment has already met finance: they have capped spending, traded down, or shut something off. This is compounded by leaders’ weighing LLMs by what they deliver without setting up the bookkeeping to prove it.

If you’re buying an LLM, start small. Pick one task you can count, such as a resolved ticket or a processed document, and work out what it costs today without AI. Then set thresholds for when to expand the rollout and when to stop it. About half of the leaders G2 interviewed haven’t set one. When you talk to vendors, ask what a finished task costs and what the spending ceiling is. G2’s analysis suggests labs that publish both would give finance a number it can sign off on.

Frequently asked questions

Q1. What is cost per outcome for LLMs?

Cost per outcome is the price of one finished unit of work done with an LLM, such as a resolved support ticket, a processed document, or a closed case. It lets a buyer compare AI spending with the cost of doing the same work by hand.

Q2. What’s the difference between cost per token and cost per outcome?

Cost per token is the vendor’s price list: what a model charges for each chunk of text it reads or writes. Cost per outcome is what a result costs once the work is done. Tokens come up constantly in interviews, but leaders judge a model by what a finished result costs, and a cheaper model can cost more overall if its answers need a second pass.

Q3. How do you measure AI ROI for an LLM?

The leaders who do it follow a simple method: pick a unit of work, put a price on the AI side, and compare it with what the same work costs before AI. A pharmaceutical governance lead weighs an automated report that costs a thousand dollars against four hours of a specialist paid $400 an hour. A SaaS team compares its current cost per ticket with its pre-AI cost.

Q4. Do enterprises hit token cost ceilings?

Some do. Of 91 leaders who discussed cost per outcome, 26% described capping spending, scaling back a project, or moving to a cheaper model. Many won’t see it coming: about half of the leaders G2 interviewed have no threshold at which they would expand or stop a rollout.

Q5. Does cost make enterprises switch LLMs?

Not often on its own. Of the 84 leaders who described their most recent model change, 15% said cost drove it, while 57% said performance or capability did.

Research methodology

We ran 102 AI-led, open-ended interviews with US enterprise leaders in August and September 2026. Participants included CIOs, CTOs, CFOs, and heads of AI who evaluate, select, or fund LLMs. Of the 100 who described their role, 61% are hands-on in choosing models, 21% hold governance or recommendation roles, and 18% own the budget.

 

Of the 93 who named an industry, 49% work in healthcare, education, the public sector, or industrial companies. Another 33% work in technology, software, and professional services, and 17% in financial services, insurance, and banking.

 

Not every leader answered every question, so each figure shows how many people it’s based on. With 102 interviews, percentages show direction more than precision, and results from groups smaller than 30 are directional. All respondents are in the United States, so the findings describe these buyers and not the whole market.