THINXSTER
Blog/AI Automation
AI Automation9 min readJuly 30, 2026

AI Infrastructure Cost Trends: Why Your Bill Went Up While Prices Fell 90%

Token prices collapsed and AI spending has never been higher. The mechanism behind that paradox — and the only cost metric worth tracking.

RK
Ryan Korsz
Founder & CEO, Thinxster

TL;DR

Token prices collapsed and AI spending has never been higher. The mechanism behind that paradox — and the only cost metric worth tracking.

→ See how this applies to your business (free 30-min call)

Here's a fact that confuses almost every business owner running AI: the cost of a unit of AI capability has fallen off a cliff, and their monthly AI bill has gone up anyway.

Both things are true, and the mechanism connecting them determines how you should budget, how you should price if you sell AI-powered services, and where you should actually look for savings. Spoiler: it's almost never in the per-token price.

What Actually Happened to Prices

For a fixed level of capability, the cost of inference has dropped by roughly an order of magnitude per year for the last couple of years. A task that cost meaningful money to run in 2023 costs a rounding error now. Three forces drove it:

  • Better hardware utilization — batching, quantization, speculative decoding, and inference-specific chips.
  • Smaller models reaching prior-generation capability. The model that costs a fiftieth as much today matches what the flagship did eighteen months ago on ordinary commercial tasks.
  • Genuine price competition among providers who all need volume.
  • The practical consequence: if you built a cost model for an AI feature more than a year ago and haven't revisited it, your assumptions are stale by a large multiple. That's real money sitting on the table.

    Why Your Bill Went Up Anyway

    Three things offset the price decline, and they're all consequences of the technology getting good.

    Reasoning consumes far more tokens. Models that think before answering can burn many times the tokens of a direct completion for a single response. Better answers, dramatically more tokens. The per-token price fell; the tokens-per-answer rose.

    Agents make many calls where a chatbot made one. A single agentic task — check the calendar, look up the customer, decide, draft, verify — is five to fifty model calls. You didn't get more expensive per call. You started making fifty times as many.

    Success expands scope. This is the biggest one and nobody puts it in a spreadsheet. The AI agent works on inbound calls, so you add outbound follow-up. That works, so you add reactivation campaigns on your dormant list. Your bill tripled because the thing worked and you fed it more work. That's not cost inflation, that's a functioning investment — but it looks identical on an invoice.

    Falling prices don't lower your bill. They raise the ceiling on what's worth automating.

    Where the Money Actually Goes in a Voice Agent

    Since voice is the highest-value AI application for most service businesses, here's the real cost decomposition of a five-minute qualifying call, ordered by share of spend:

    1.

    Text-to-speech. Usually the single largest line item for premium, natural-sounding voices. Voice quality directly affects whether callers stay on the line, so this is the least sensible place to cut.

    2.

    Telephony. Per-minute carrier costs plus number rental. Steady, predictable, and largely non-negotiable below serious volume.

    3.

    Speech-to-text. Streaming transcription. Meaningful but has fallen substantially and continues to.

    4.

    The language model. Frequently the smallest component of a voice call, which surprises people who assume the AI is the expensive part. A five-minute conversation is a modest number of tokens.

    5.

    Orchestration platform fees. The vendor's margin on top. Often the largest single line and the least examined.

    The strategic implication is blunt: optimizing your model choice to save on tokens is usually optimizing the smallest number on the bill. The leverage is in voice provider selection, call duration, and platform fees.

    The Only Cost Metric That Matters

    Cost per minute, cost per token, cost per conversation — all inputs. The metric that belongs on your dashboard is cost per booked appointment, or whatever your equivalent outcome is.

    Run the arithmetic once and the framing changes permanently. Take a business handling 400 inbound leads a month with a $3,000 average ticket and a 30% close rate on qualified appointments:

  • AI system all-in cost: roughly $2,000 a month.
  • Qualified appointments produced: around 100.
  • Cost per booked appointment: about $20.
  • Revenue per booked appointment at a 30% close rate: about $900.
  • At those numbers, a 20% reduction in AI costs saves $400 a month. A 5-point improvement in qualification rate produces roughly $4,500. The optimization energy belongs entirely on the second one — and yet most teams spend their attention on the first, because a bill is concrete and a conversion rate is abstract.

    62%
    average lead qualification rate across client accounts

    Where to Actually Find Savings

    If you do want to cut costs, cut them here, in descending order of impact:

    Shorten the conversation, not the model. A qualifying call that reliably reaches an outcome in three minutes instead of six halves your per-call cost across every component simultaneously. Better conversation design saves more than any model swap.

    Stop calling leads you shouldn't call. Filtering obvious non-fits before the agent dials cuts spend and improves your contact-rate metrics at the same time.

    Route by difficulty. Use a cheaper, faster model for simple turns and escalate to a stronger one for complex reasoning. Meaningful savings, modest engineering effort, no quality loss if the routing rule is sane.

    Cache the repeated. Your system prompt, business details, and FAQ content are identical on every call. Prompt caching cuts the cost of that repeated context substantially and improves latency at the same time — a rare case where cheaper is also better.

    Renegotiate at renewal. If your vendor's pricing was set against costs from eighteen months ago, their input costs have fallen and yours haven't. Bring the numbers to the conversation.

    How to Actually Forecast Next Year's Bill

    Most AI budgets are built by taking last month's invoice and multiplying by twelve, which is wrong in both directions simultaneously — it assumes flat prices and flat scope, and neither holds.

    A better method, and it takes twenty minutes:

    1.

    Split your current bill into per-unit costs and fixed costs. Platform fees and number rentals are fixed. Model, speech, and telephony charges are per-unit. You need these separated because they move in opposite directions.

    2.

    Project your volume, not your spend. How many conversations do you expect next year? Base this on lead volume growth, not on last year's bill.

    3.

    Apply a per-unit cost decline to the variable portion. Assuming a 30-50% reduction in per-unit inference cost over twelve months has been conservative recently. Apply a much smaller decline — or none — to telephony and platform fees.

    4.

    Add a scope line, explicitly. What are you likely to automate next that you aren't automating today? Outbound follow-up, reactivation, review requests. Estimate each as its own volume line rather than burying it in a growth percentage.

    5.

    Convert the whole thing to cost per outcome and compare to revenue per outcome. That ratio is what you actually manage. If it's improving, a rising total bill is good news.

    The output usually surprises people: variable costs roughly flat despite significant volume growth, fixed platform fees becoming a larger share of the bill, and the biggest single driver being scope expansion you chose. Which reframes the budget conversation from cost control to investment sizing.

    What to Assume Going Forward

    For planning purposes over the next two to three years:

  • Per-unit inference costs keep falling. Don't sign multi-year deals priced on today's usage assumptions without a step-down clause.
  • Total AI spend rises anyway, because scope expands. Budget for that as growth investment, not overrun.
  • Voice and telephony fall more slowly than text. They're closer to physical infrastructure.
  • Platform fees are the softest number on your bill. They're margin, and margin is negotiable.
  • The capability ceiling rises faster than costs fall. Things that are uneconomical to automate today become obviously worth it within a year. Revisit your "not worth automating" list every six months.
  • $102M+
    client revenue generated on systems where cost per booked job is the metric that gets watched

    The Reframe

    Stop asking what AI costs. Ask what it costs per outcome, and compare that to what the outcome is worth. At Thinxster, the number we hold ourselves to isn't cost per minute — it's whether the client's AI callers, responding to every inbound lead within 90 seconds, produce more booked revenue than they consume. Across client accounts that's been $102M+ in tracked revenue at a peak ROAS of 9.2×, and the AI bill was never the interesting line on the P&L.

    If your AI spend is a number you look at without knowing what it produced, that's the fixable problem. [Book a free strategy call](/book) and we'll build the cost-per-outcome math for your business.

    Free Weekly Briefing

    One AI Marketing Tactic.
    Every Tuesday. Free.

    What's actually working across our client accounts right now — ROAS moves, follow-up sequences, creative angles. The stuff that isn't in any blog post yet.

    No spam. Unsubscribe anytime. 1,200+ business owners already in.

    Ready to Deploy

    SEE THIS IN
    YOUR BUSINESS.

    30 minutes. We scope the exact systems that apply to your situation and give you a plan.

    ★★★★★ Trusted by 47+ local service businesses

    BOOK A STRATEGY CALL →