TL;DR
Deploying an AI agent is a project. Managing one is a role. Here's what the work actually involves week to week — and why unmanaged agents quietly decay.
→ See how this applies to your business (free 30-min call)There's a pattern I've watched play out at more than a dozen companies now. An AI agent launches. The first month is great — response times collapse, the team is thrilled, someone posts about it on LinkedIn. Month three is fine. By month six, booking rates are quietly below where they started, nobody can say when the decline began, and the prevailing theory is that "AI doesn't really work for our business."
The AI worked fine. Nobody managed it.
This is the gap almost every AI deployment has: budget for building, budget for licensing, zero budget for the ongoing operational work of keeping an agent good. That work is a real job with real practices, and it's emerging as its own service category. Here's what it consists of.
Why Unmanaged Agents Decay
An AI agent isn't software in the traditional sense. Traditional software does the same thing on day 400 as day 1. An agent's performance is a function of things that all drift:
Your business changes. New service, new pricing, a location closed, a promotion ended. The agent is still confidently telling callers about the old one.
Your market changes. A competitor launches an offer and suddenly every third caller asks a comparison question your agent has never seen and handles badly.
The model changes. Providers update models, deprecate versions, and adjust behavior. A prompt tuned against one model version can behave measurably differently on the next.
Edge cases accumulate. Every week produces conversations nobody anticipated. Unhandled, they compound into a growing share of calls that end badly.
Small changes go unmeasured. Someone tweaks a line in the prompt to fix one complaint, and it degrades three other paths. Without an evaluation set, nobody knows.
None of these are failures of the technology. They're failures to staff the operating role.
An AI agent is closer to a new hire than a new tool. Nobody expects a new hire to perform for two years with zero feedback.
What the Work Actually Is, Week to Week
Here's the real cadence of managing a production agent. It's less exotic than it sounds and more disciplined than most teams expect.
Daily (10 minutes): Check the exception queue. Every conversation the agent escalated, abandoned, or ended without a clear outcome. You're not fixing anything yet — you're spotting whether today looks like yesterday.
Weekly (60-90 minutes), the core loop:
Read 20 transcripts. Not summaries. Not a dashboard. Actual conversations — 10 that converted, 10 that didn't. This is where every real insight comes from, and it's the step everyone skips because it isn't automatable.
Categorize the failures. Each lost conversation gets a tag: wrong information, misunderstood intent, poor handoff, caller frustration, out of scope. After four weeks you have a ranked list of what's actually costing you bookings.
Change one thing. Fix the top category. One change, so you can attribute the effect.
Run the evaluation set before shipping it. Which requires having one — see below.
Log the change with the date and the reason. Six months from now, when performance shifts, this log is the only way to find out why.
Monthly: Review the trend metrics, refresh the knowledge base against actual business changes, and check unit cost per outcome.
Quarterly: Re-examine whether the agent's scope is still right, and rebuild the evaluation set from recent conversations so it reflects the business you have now rather than the one you had at launch. Usually scope should expand — the things you were nervous about in month one are now obviously safe. Sometimes it should contract.
The Metrics That Actually Predict Revenue
Most AI agent dashboards report volume metrics: calls handled, messages sent, minutes used. Those measure activity, not value. The ones that matter:
The Evaluation Set: The Practice That Separates Serious Teams
If you take one thing from this article, take this. An evaluation set is 50-100 real conversations from your own history, each with a documented correct outcome. Before any prompt or model change goes live, you replay the set and compare.
Building one takes about a day. Pull real transcripts, weight toward tricky cases, and for each one write down what the agent should have done. Include the nasty ones: the caller who changes their mind mid-conversation, the one who asks about a service you don't offer, the one who's angry about something unrelated.
Without this, every change is a guess, and you find out you were wrong from a revenue report six weeks later. With it, you find out in ten minutes. Teams that run evaluation sets ship changes weekly with confidence. Teams that don't eventually freeze — too scared to touch a system they can't verify.
Who Does This Job
Three models, in order of how common they are:
Bolted onto an existing role. The ops manager or marketing lead adds it to their week. Works up to a point. Fails when it gets busy, which is exactly when the agent is handling the most volume.
A dedicated AI operations role. Emerging at companies with several agents in production. The skill profile is unusual and specific: comfortable reading transcripts at volume, good at pattern recognition, decent at writing, able to reason about metrics. It is not a machine learning role — almost nothing about it requires training a model.
Outsourced to a managed partner. The vendor or agency owns the weekly loop and reports on outcomes. This is what the "AI agent management business" actually refers to as a category, and it exists because the work is real, recurring, and unglamorous enough that most companies would rather not staff it.
The honest framing: if you can't name the person doing this work and see time on their calendar for it, it isn't happening — regardless of what anyone believes.
What This Looks Like Done Right
At Thinxster, the weekly management loop is the product, not a support function. The build gets a client's AI callers responding to every inbound lead within 90 seconds and writing into a GoHighLevel pipeline. What actually compounds the results is what happens after: transcripts read every week, failure categories ranked, one change at a time, measured against a real evaluation set, with the number being tracked as booked revenue rather than call volume.
That loop is the difference between systems that hold a 9.2× peak ROAS and systems that look brilliant in month one and get quietly turned off in month eight. Across client accounts, the compounding version has carried $102M+ in tracked revenue.
Start This Week
You don't need a role or a budget to begin. You need ninety minutes:
Pull 20 transcripts from last week — 10 that converted, 10 that didn't.
Read them and tag each failure with a category.
Fix the single most common one.
Write down what you changed and why.
Put the same ninety minutes on your calendar for next week.
Do that for a month and you will find things about your business — objections you didn't know existed, questions you've never answered on your website, a competitor you didn't know you were losing to — that no dashboard would ever have shown you.
If you'd rather have that loop run for you, with reporting on booked revenue instead of usage stats, [book a free strategy call](/book) and we'll show you what managed looks like.
Free Weekly Briefing
One AI Marketing Tactic.
Every Tuesday. Free.
What's actually working across our client accounts right now — ROAS moves, follow-up sequences, creative angles. The stuff that isn't in any blog post yet.
No spam. Unsubscribe anytime. 1,200+ business owners already in.