TL;DR
Almost every company claims to be AI-powered. Four tests separate the real ones from the wrappers — the same four you should run on any vendor pitching you.
→ See how this applies to your business (free 30-min call)Somewhere around 2024, "AI company" stopped meaning anything. Every software product added a chat box, every agency added the words to its homepage, and a term that used to describe a specific technical posture became a marketing adjective.
That's a problem if you're buying. So here's a definition that survives contact with reality, built as four tests. A company that passes all four is doing something structurally different. A company that passes none is a normal business with an API key, which is fine — it just shouldn't be priced like the former.
Test One: Does the Product Fail Without the Model?
The cleanest test, and the first one to run.
Remove the AI. If the product still does its core job — a CRM still stores contacts, a design tool still lets you draw, a project tracker still tracks projects — then AI is a feature. Valuable, possibly excellent, but a feature.
If removing the model leaves nothing, the company is built on it. A voice agent that answers your phone and books appointments is not a phone system with an AI add-on; there's no product underneath.
This isn't a judgment about quality. Plenty of great companies fail this test and should. It's a judgment about what you're actually buying and how it should be priced. Feature-level AI is worth a feature-level premium.
Test Two: Is There a Proprietary Data Loop?
The question that separates durable advantage from a temporary one.
Everyone has access to the same frontier models. What's not equally available is data generated by your own operation that makes your system better in a way competitors can't replicate. A real loop has four properties:
It's generated by usage, not purchased. Bought data is available to whoever else buys it.
It's labeled by outcome. Not just what happened, but whether it worked — did the conversation convert, did the extraction turn out correct, did the routing decision hold up.
It feeds back into the system. Retraining, prompt refinement, retrieval corpora, evaluation sets. A data lake nobody uses isn't a loop.
It compounds. Month twelve is meaningfully better than month one because of accumulated data, not because someone swapped in a newer model.
Ask a vendor to describe their loop concretely. Most can't, because most don't have one. That answer is genuinely diagnostic.
Access to a model is a commodity. What you learn from running it at scale, on outcomes only you can see, is not.
Test Three: Is There Evaluation Infrastructure?
The most reliable technical signal, and the one nobody markets.
Companies that seriously build on AI have a way to answer "did that change make it better?" That means a test set of real cases, automated scoring, regression runs on every change, and tracked quality metrics over time. It's unglamorous, expensive to build, and never appears on a homepage.
Companies that don't have it are shipping prompt changes on intuition. You can identify them by asking one question: "How do you know your system got better last quarter?"
The strong answer is a number and a methodology. The weak answer is a feature list. The disqualifying answer is confusion about why you're asking.
Test Four: Do They Own Their Cost Curve?
At scale, unit economics become the whole business. A company that treats model spend as an unexamined line item hasn't hit the scale where it matters, or isn't paying attention.
Signals that they own it: they route different tasks to different model tiers rather than sending everything to the most expensive option. They cap agent loops. They know their cost per completed task and can tell you whether it's rising or falling. They've made a considered decision about which workloads run on hosted APIs versus their own infrastructure.
This matters to you as a buyer for a direct reason: a vendor whose costs are out of control will eventually pass them to you, restrict usage, or fail.
The Honest Taxonomy
Four categories, and each is legitimate at the right price:
Model developers. They train frontier models. A very small number of companies globally. Not you, and not your vendor.
AI-native products. Fail test one — the product doesn't exist without the model. Usually pass two and three if they're any good. This is where most genuinely interesting AI companies sit.
AI-enhanced products. Real software with real AI features that make it meaningfully better. Most good software in 2026. Nothing wrong with this; it just isn't a different category of company.
Wrappers. A thin interface over someone else's API with no data loop, no evaluation, and no cost discipline. Some are useful — packaging and distribution have real value. But there's no moat, and the pricing should reflect that.
The pitch decks that annoy people are wrappers presenting as AI-native. The tests above sort them in about ten minutes.
The Uncomfortable Self-Assessment
It's worth applying this to ourselves rather than only to other people.
We're a marketing agency that operates AI systems. On test one, our client-facing systems fail cleanly without the model — an AI caller that responds to inbound leads within 90 seconds and runs a qualifying conversation has no non-AI fallback that does the same job. On test two, we have a real loop: every call across every client account produces a transcript, a qualification decision, and an eventual outcome — booked, closed, or lost — and that labeled data is what has moved our qualification rate to where it is.
On test three, every qualification decision is logged and scored against what actually happened, which is why we can talk about that number at all. On test four, we tier models by task and cap loops, because voice minutes at volume are the kind of cost that punishes carelessness quickly.
What we are not is a model developer, and we'd be lying if we implied otherwise. We're an operator of AI systems with a proprietary outcome loop in one specific domain: converting inbound demand into booked revenue for local service businesses.
Three Patterns You'll See in Pitches
Once you have the four tests, the failure patterns become easy to spot.
The capability tour. The deck lists everything the underlying model can do — summarize, generate, classify, converse — as though those were the company's capabilities. They're the model provider's capabilities, available to everyone including your existing vendors. The tell is that nothing in the deck is specific to their domain.
The proprietary-data claim without a loop. "We've trained on millions of data points." Ask what the outcome labels are and how the data feeds back. Frequently the answer is that they have a lot of stored records and no mechanism connecting them to whether anything worked. Volume isn't a loop; outcome labeling is.
The benchmark that isn't yours. Impressive accuracy figures on a public benchmark or a curated internal test set that bears no resemblance to your data. The right response is to ask them to run it on a hundred of your real, messy examples before you sign anything. Companies with genuine evaluation infrastructure find this easy and interesting. Companies without it find reasons why it's not possible.
None of these mean the vendor is dishonest. Most often it means they're a good product company using the industry's vocabulary because everyone else is. Price accordingly and the relationship works fine.
The Four Questions to Ask Any Vendor
Compressed into what you should actually say on a call:
"If the model went away tomorrow, what's left of your product?"
"What data do you generate that a competitor couldn't buy?"
"How do you know your system got better last quarter — with a number?"
"What's your cost per completed task, and is it going up or down?"
Vendors who've thought about this answer all four within a couple of minutes. Vendors who haven't will redirect to their client list.
The Point
"AI company" isn't a badge, and it isn't a moat by itself. What's defensible is the loop — the accumulated, outcome-labeled data from running real systems for real customers, and the measurement discipline to turn it into improvement.
Judge vendors on that rather than on the adjective. And judge results on revenue, which is the one metric that was never in question.
If you want to see what an outcome loop looks like applied to your inbound demand — and what it would produce — [book a free strategy call](/book).
Free Weekly Briefing
One AI Marketing Tactic.
Every Tuesday. Free.
What's actually working across our client accounts right now — ROAS moves, follow-up sequences, creative angles. The stuff that isn't in any blog post yet.
No spam. Unsubscribe anytime. 1,200+ business owners already in.