TL;DR
Most AI purchases are made on a demo and regretted at renewal. The diligence checklist — data, compliance, exit terms, pilot design — that prevents it.
→ See how this applies to your business (free 30-min call)AI procurement fails in a predictable way. Someone sees a demo that's genuinely impressive, a pilot gets approved on enthusiasm, the pilot has no defined success criteria, and eleven months later the renewal lands on a desk with nobody able to say whether the thing worked. It renews, because cancelling requires proving it failed and nobody measured anything.
The fix isn't more caution. It's buying AI the way you'd buy any other operationally critical system — with defined outcomes, real diligence, and an exit that doesn't require a rebuild. Here's the checklist.
Start With the Outcome, Not the Category
The worst procurement briefs name a technology. "We need an AI chatbot." "We should get AI agents." That framing guarantees you'll evaluate vendors on features instead of results.
Write the brief as a measurable business outcome with a baseline:
The second version tells you exactly what to measure, makes vendor comparison trivial, and — importantly — leaves open the possibility that the answer isn't AI at all. Sometimes it's a routing rule and a phone rota. Finding that out during procurement is a win, not a failure.
The Diligence Questions That Matter
Most vendor questionnaires are 80 questions long and ask nothing important. These twelve do the work.
On data:
Where is our data stored, in what jurisdiction, and for how long?
Is our data used to train your models or any shared model? Get this in writing — verbal assurances change with the terms of service.
Which subprocessors touch our data? Every AI vendor is a wrapper around at least one model provider. You need the full chain.
On termination, what do we get back, in what format, and how long do you retain a copy?
On the model:
Which models power this, and what happens when one is deprecated? Deprecation is not hypothetical — providers retire models on 6-12 month cycles.
Can we see version history for the model and system prompt? If they change the underlying model without telling you, your carefully tuned workflow silently changes behavior.
How is accuracy measured, and can we see the evaluation set? "99% accurate" is meaningless without knowing on what.
On operations:
What's the documented uptime, and what's the remedy when it's missed? Service credits on a $2,000 monthly contract are not a remedy — the real question is whether you can operate without it.
What's the escalation path when the AI can't handle something? An agent with no graceful handoff to a human is a liability, not an asset.
Who is accountable for output quality — you or us? This gets uncomfortable and it's the most important question on the list.
On commercials:
What does this cost at 3× current volume? Model it. Per-unit pricing that's attractive at pilot volume is often indefensible at scale.
What's the notice period and what triggers an auto-renewal? Twelve-month auto-renew with 90-day notice means you have a three-week window each year to make a decision.
The Compliance Surface People Miss
AI systems that talk to customers carry obligations that ordinary software doesn't. In rough order of how often they're overlooked:
Disclosure. A growing set of jurisdictions require disclosing that a caller or chat participant is an AI, sometimes proactively. Even where it isn't required, the reputational math is one-sided: getting caught concealing it costs far more than disclosing it ever did. Build the disclosure into the opening line and stop thinking about it.
Call recording consent. Two-party consent states require both parties to agree before recording. If your AI agent records every call for quality and training — and it should — that consent has to be captured at the top of the call, and your vendor needs to support it per-jurisdiction, not globally on or off.
Outbound calling and messaging rules. Time-of-day restrictions, do-not-call list scrubbing, prior express written consent for automated calls, and A2P 10DLC registration for SMS. Automation does not lower the bar here; it raises the volume of your exposure. Confirm the vendor enforces these rather than trusting you to.
Sector rules. Healthcare, financial services, and legal each add their own layer — a BAA for anything touching patient data, disclosure and suitability rules for financial advice, unauthorized practice of law risk for legal intake. If your vendor can't discuss the relevant framework fluently, they haven't sold into your sector before.
If a vendor's compliance answer is "our customers handle that," you just learned who carries the risk.
Design the Pilot to Produce a Decision
A pilot without predefined success criteria isn't a pilot, it's a free trial with extra meetings. Structure it like this:
Fixed duration. 60 or 90 days. Long enough for a real sample, short enough to force a decision.
One named metric with a baseline and a threshold. "Lead-to-appointment rate over the pilot period must beat our trailing 90-day baseline of 24% by at least 5 points." Write the baseline down *before* you start, from your own data.
A control if you can get one. Half your lead volume through the new system, half through the existing process. This is the difference between knowing it worked and believing it did.
A named owner who has time allocated, not just assigned. Pilots die from lack of attention more than from bad technology.
A pre-committed decision rule. "If we hit the threshold we expand to all locations. If we miss by less than 2 points we extend 30 days. If we miss by more, we stop." Agree to this in writing before the pilot, when nobody is emotionally invested.
That last step is the one that saves organizations from years of zombie subscriptions.
The Total Cost Nobody Quotes
The license fee is rarely the biggest number. Budget for:
Buy Outcomes, Not Access
The cleanest structural test of a vendor: are they selling you access to a capability, or accountability for a result?
Access-based vendors sell seats and minutes and are indifferent to whether you succeed. Outcome-based partners tie their reporting — and ideally some of their compensation — to the metric you actually care about. Ask directly whether they'll be measured on your booked appointments or closed revenue rather than usage. The reaction tells you which kind of company you're dealing with.
This is the standard we hold ourselves to at Thinxster. Clients get AI caller agents that respond to every inbound lead in 90 seconds, GoHighLevel pipelines wired so every ad dollar traces to a booked job, and reporting on the number that matters — revenue, not usage. That accountability is why the systems we've built have carried $102M+ in tracked client revenue at a peak ROAS of 9.2×.
Who Belongs in the Room
AI procurement gets routed to whoever is most enthusiastic, which is rarely the right group. Four perspectives need to be present before signing, even informally at a small company:
Fifteen minutes with these four perspectives kills more bad purchases than any amount of vendor questionnaire.
The One-Page Summary
Before you sign anything: define the outcome and baseline, confirm data ownership and exit terms in writing, verify the compliance surface for your sector, model the cost at 3× volume, design a pilot with a pre-committed decision rule, and name the person who owns it.
Six items. Most failed AI purchases skipped at least four of them.
If you want a second opinion on an AI proposal sitting on your desk — including an honest read on whether it's worth buying at all — [book a free strategy call](/book) and we'll go through it with you.
Free Weekly Briefing
One AI Marketing Tactic.
Every Tuesday. Free.
What's actually working across our client accounts right now — ROAS moves, follow-up sequences, creative angles. The stuff that isn't in any blog post yet.
No spam. Unsubscribe anytime. 1,200+ business owners already in.