AI spend benchmarkLLM API costsAI FinOpsAI gross margincost per task

How Much Should You Spend on AI APIs? Benchmarks and Red Flags (2026)

Everyone wants a benchmark, a percentage of revenue that tells them their AI bill is normal. There isn't one, because a normal-looking bill can be almost entirely waste and a scary-looking one can be perfectly efficient. So this piece gives you the real reference points that exist, and then the three red flags that mean you are overspending regardless of what any benchmark says.

By Roman Rose, Founder, Parity Layer11 min read

Key takeaways

  • There is no single correct number for AI API spend, because a bill that looks normal against a benchmark can still be mostly waste, and a high bill can be perfectly efficient. Benchmarks tell you what others spend, not whether your spend bought the cheapest model that still held quality.
  • The real reference points: enterprise foundation-model API spend reached $12.5B in 2025 as part of $37B total generative-AI spend (Menlo Ventures), average enterprise LLM budgets sit around $7M and are expected to grow roughly 65% next year (a16z), and 98% of FinOps teams now manage AI spend with FinOps-for-AI the number-one forward priority (FinOps Foundation).
  • AI is a gross-margin story, not just a cost line: AI product gross margins average around 52% (ICONIQ) versus the 80 to 90% SaaS standard, and bolting an AI feature onto an $80 seat can add roughly $15 of variable cost and drop that seat's margin from 80% to about 65%.
  • Three red flags mean you are overspending regardless of any benchmark: a frontier model wired into every prompt, no cost-per-task attribution, and no proof that any model was actually the right one for the job.
  • The fix is not a cheaper model chosen on vibes, it is proving a cheaper model matches or beats your current one on your own prompts before you route to it. Prove it offline on a JSONL export first, and none of this applies to coding agents.

There is no single correct amount to spend on AI APIs, and anyone who hands you a tidy percentage of revenue is selling comfort, not truth. Here is the honest version: your AI bill is fine if you can attribute it to specific tasks, tie it to value, and prove that the model you are paying for is the cheapest one that still holds quality on those tasks, and it is a problem if you cannot, no matter how normal the total looks against an industry average.

That is the whole answer, and the rest of this piece is the two halves nobody gives you. First, the real reference points that actually exist, traced to primary sources rather than SEO aggregator maths, so you have something to anchor to. Second, the three red flags that mean you are overspending regardless of any benchmark, because a bill that looks perfectly average can be half waste, and a scary-looking one can be perfectly efficient.

The one sentence to remember

The number that decides whether you are overspending is not your AI bill as a percentage of revenue. It is whether you can prove each model was the right one for the task it is running, which is the job Parity is built around.

Is there a benchmark for how much a company should spend on AI?

Not a useful one, and it is worth being clear about why before you go looking for it. A "percentage of revenue" benchmark assumes every business uses AI at the same intensity, which is nonsense. A support tool that runs one classification per ticket and a product that generates long documents on every click have wildly different AI-to-revenue ratios, and both can be perfectly well run. The average of the two tells you nothing about either.

So the useful reference points are not ratios, they are directional anchors: how fast the whole market is spending, what a typical enterprise line looks like, and what AI does to your margins. Those give you a sense of scale and trajectory. What they cannot do, and what no benchmark can do, is tell you whether your specific spend bought the right model. That part is on you, and it is the part the red flags below are about.

The real reference points (traced to primary sources)

Here is what is actually true, from the primary reports rather than the blogs that reprint them.

The market is spending enormous and fast-growing sums. Enterprise generative-AI spend reached $37 billion in 2025, which Menlo's own figures put as a tripling from $11.5 billion the year before (Menlo Ventures 2025 report announcement), and within that the foundation-model API layer alone was $12.5 billion, per Menlo Ventures' 2025 State of Generative AI in the Enterprise. Zoom in on the API line specifically and the growth is even steeper: enterprise LLM API spend roughly doubled in six months, from $3.5 billion in late 2024 to $8.4 billion by mid-2025, per Menlo's mid-year update.

At the individual-company level, a16z's CIO survey puts average enterprise LLM spend around $7 million, up from roughly $4.5 million two years earlier, with leaders expecting budgets to grow another 65% to about $11.6 million next year. And the discipline is catching up to the spend: 98% of FinOps teams now manage AI cost, up from 63% the year before, with "FinOps for AI" ranked the number-one forward-looking priority in the State of FinOps 2026 report. Everyone is spending more and, finally, everyone is worried about it.

The part that reframes the whole question is margin. AI is not just a bigger cloud bill, it sits in cost of goods sold and compresses your gross margin directly. AI product gross margins now average around 52%, against the 80 to 90% that defined the last decade of software (ICONIQ and Bessemer figures, summarised here). In concrete terms, bolt an AI assistant onto an $80-per-month seat and the inference plus supporting infrastructure can add roughly $15 of variable cost, dropping that seat's margin from 80% to nearer 65% overnight (The SaaS CFO). That is the real reason the absolute number matters less than the efficiency behind it: every wasted dollar of API spend comes straight off a margin that is already structurally thinner than the SaaS you are used to.

Reference pointWhat the primary source saysWhat it tells you
Total enterprise genAI spend, 2025$37B, up from $11.5B a year earlier; foundation-model APIs $12.5B (Menlo Ventures, report announcement)The market is huge and accelerating, so rising spend alone is not a red flag
Enterprise LLM API spend growth$3.5B to $8.4B in six months (Menlo mid-year)The API line specifically is the fastest-moving part, worth watching closely
Average enterprise LLM budget~$7M now, up from ~$4.5M, expected +65% to ~$11.6M (a16z)A rough sense of scale for a large org, not a target to hit
AI product gross margin~52% average vs 80 to 90% SaaS standard (ICONIQ / Bessemer)AI spend lands in COGS, so waste hits margin directly
AI-feature margin drag~$15 variable cost on an $80 seat, 80% to ~65% margin (The SaaS CFO)A useful per-unit lens: what does AI cost per thing you sell?
FinOps adoption of AI cost98% now manage AI spend, top forward priority (FinOps Foundation)You are late to worry, not early, if you have not started

Read those as scale and direction, not as a target. None of them answers the only question that decides whether your bill is healthy, which is whether the money bought the right model. That is what the red flags are for.

The three red flags that mean you are overspending anyway

Forget the benchmark for a moment. These are the signs that you are burning money regardless of what your total looks like against any industry average, and I have hit all three in my own companies, which is how I know they are the ones that matter.

Red flag one: a frontier model wired into every prompt

The single most common cause of an inflated AI bill is a top-tier model used as the default for everything, including the high-volume, low-difficulty prompts that never needed it. The price gap inside a single vendor's line-up is large, often around 5x between the frontier tier and the cheap tier, so paying frontier prices on a prompt that a much cheaper model handles cleanly is pure waste repeated thousands of times a day. Most teams pick the strongest model once, during the exciting demo phase, and never revisit it once the same model is quietly running the boring classification job at scale. If you have one default model and it is your most expensive one, you have found your first leak. The triage of which prompts genuinely need the expensive model is a whole exercise in itself, and I wrote up the full cost-optimisation guide around exactly that decision.

Red flag two: no cost-per-task attribution

Ask most teams which of their prompts is responsible for the largest share of the bill and they cannot tell you, because the invoice arrives as one undifferentiated number. That is the FinOps gap the whole industry is now scrambling to close, and it is a genuine blind spot: without cost-per-task attribution you cannot find the leak, so you cannot fix it, and you certainly cannot decide where a cheaper model would pay off. Total spend is a symptom. Cost per task is the diagnosis. If you cannot break your bill down to "this prompt type costs us this much per thousand runs", you are flying blind, and the practical playbook for closing that gap is in FinOps for AI.

Red flag three: no proof that any model was the right one

This is the one nobody talks about and it is the most expensive. Almost every model choice in production was made on a public benchmark, a vendor's marketing, or a gut feel during that first demo. None of those is evidence about your traffic. Public benchmarks get contaminated and gamed, and your support macros and extraction schemas look nothing like MT-Bench anyway, so "we picked the best model" almost always means "we picked the model that looked best on someone else's test". The result is a bill full of choices you cannot defend: you are paying frontier prices with no proof the frontier model was necessary, and you would switch to something cheaper if only you could trust it. That trust gap is the actual product problem, and it is a proof problem, not a pricing one.

So what should you actually do about it?

You stop trying to hit a benchmark number and you start proving each choice. The benchmark question ("is my bill normal?") is the wrong question because it can only ever tell you what other people spend. The right question is "have I proven the cheapest model that holds quality on each of my tasks, and am I running that one?" Answer that and the total takes care of itself.

Concretely, that means taking the high-volume prompts that are actually burning your bill, the red-flag-one prompts, and for each one testing a cheaper model against the model you run today, on your own real inputs, until you can see whether the cheaper one genuinely holds. You route the ones that pass, you leave the ones that fail on the expensive model, and you keep the expensive model armed as an instant fallback for anything that drifts later. On the prompts that pass, the proven saving typically lands in the 30 to 60% range. It is a band and not a hero number because it genuinely varies prompt by prompt, which is the entire reason you prove each one rather than swinging your whole bill at a single cheap model on faith.

The part that makes this trustworthy rather than another vibe-based switch is the standard you judge against. You cannot call a cheaper model worse until you know how much your own expensive model already disagrees with itself, because ask it the same prompt twice and the answer moves. So you measure your baseline's own run-to-run consistency first, and a cheaper model passes when it disagrees with your baseline no more often than your baseline disagrees with itself, judged blind, on three axes: format as a hard exact-match gate, categorical differences re-judged blind, and semantic as a diagnostic. The standard is your model's, not a vendor's and not a leaderboard's. If you want the full mechanics of that measurement, it lives in the cost-optimisation guide and the wider cut-your-AI-bill-without-dropping-quality write-up.

Two honest caveats, stated plainly. First, none of this is for coding agents. Long-horizon agentic coding is exactly the broad, high-variance work where cheaper models still lose, and no honest bar will pretend otherwise. This is for the high-frequency, well-defined jobs a business runs all day, the classification, extraction, summarisation, qualification and generation off structured data, which is also where most of a bill is hiding. Second, you do not have to trust any of this on my word: the offline path lets you export a JSONL of past requests, upload it, and get the proof on your own historical prompts before you change a line of code. Start there. If you want a fast read on whether your current bill is carrying waste, see how much you're overspending on your AI API bill first, then prove the fix on your own data.

Frequently asked questions

How much should a company spend on AI APIs?

There is no single right figure, and any blog that gives you one number is selling you a false comfort. The honest answer is that your AI spend is fine if you can attribute it to specific tasks, tie it to value, and prove the model you are paying for is the cheapest one that still holds quality on those tasks. If you cannot do those three things, the absolute number is irrelevant, because you have no way to know whether it is half waste. Spend as much as passes that test and not a token more.

What is a normal AI API bill as a percentage of revenue?

People want this number and it does not exist in a useful form, because AI intensity varies wildly by product. A more useful lens is gross margin: AI product gross margins now average around 52% against the 80 to 90% that defined the last decade of SaaS (ICONIQ, Bessemer), so if an AI feature is dragging a seat's margin from 80% down toward 65%, that is expected, not alarming. The question is not what percentage of revenue you spend, it is whether that spend is buying the right model at the right price for each task.

How much do enterprises actually spend on LLM APIs?

Enterprise foundation-model API spend reached $12.5B in 2025, part of $37B in total enterprise generative-AI spend that tripled year on year (Menlo Ventures). At the company level, a16z's CIO survey puts average enterprise LLM spend around $7M, up from roughly $4.5M two years earlier, with budgets expected to grow about 65% to $11.6M next year. So the direction is steeply up, which is exactly why the efficiency of each dollar matters more, not less.

How do I know if I am overspending on AI?

Look for three red flags rather than a benchmark. First, a frontier model wired into every prompt, including the cheap high-volume ones that never needed it. Second, no cost-per-task attribution, so you cannot say which prompt is burning the bill. Third, no proof that any model was actually the right one, meaning every model choice was made on a benchmark or a vibe, not on your own traffic. If any of those is true, you are almost certainly overspending, whatever your total looks like against an industry average.

Should I just switch everything to a cheaper model to cut the bill?

No, because a blanket switch trades a visible cost problem for an invisible quality one, and your customers find the degraded output before your dashboard does. The saving is only real on the prompts where a cheaper model genuinely matches or beats your current one, and that varies prompt by prompt, which is why you prove each one rather than swinging the whole bill at a single cheap model. Proven savings on the prompts that pass typically land in the 30 to 60% range. This does not apply to coding agents, where cheaper models still lose.

Sources

  1. 1.Menlo Ventures: 2025 The State of Generative AI in the Enterprise ($37B total, $12.5B foundation-model APIs)
  2. 2.Menlo Ventures 2025 report announcement: enterprise AI investment tripled from $11.5B to $37B in one year
  3. 3.Menlo Ventures: 2025 Mid-Year LLM Market Update (LLM spend $3.5B to $8.4B in six months)
  4. 4.Andreessen Horowitz (a16z): Leaders, gainers and unexpected winners in the Enterprise AI arms race (CIO budget survey)
  5. 5.FinOps Foundation: State of FinOps 2026 (98% now manage AI spend; FinOps for AI the top forward priority)
  6. 6.FinOps Foundation: The State of FinOps 2025 Report (63% managing AI spend, doubled year on year)
  7. 7.Avante Ventures: AI startup gross-margin benchmark 2026 (ICONIQ ~52%, Bessemer ~65%)
  8. 8.The SaaS CFO: Your AI Feature Is Quietly Destroying Your Gross Margin ($80 seat, ~$15 added cost)

Prove it on your own prompts

See whether a cheaper model matches or beats your output for 30-60% less. Up to 10 prompts free, no credit card.

Keep reading