Skip to content

· · AI for Sales  · 10 min read

How to Buy an AI Sales System Without Getting Burned

Demos are the vendor's best case; what you're buying is the Tuesday-afternoon case. This final post is the checklist for seeing it before you pay: five contract-level questions, the red flags, a fair pilot design — and, since I sell these systems, the scenario where the honest answer is don't buy from me.

Demos are the vendor's best case; what you're buying is the Tuesday-afternoon case. This final post is the checklist for seeing it before you pay: five contract-level questions, the red flags, a fair pilot design — and, since I sell these systems, the scenario where the honest answer is don't buy from me.

The AI sales market in 2026 sells demos, and a demo is the best case. The vendor picked the script, rehearsed the happy path, and ran it on a quiet afternoon. What you’re actually buying is a Tuesday three months from now: the caller is angry, the question is one nobody scripted, your prices changed last week, and the person who configured everything is on vacation. The first post in this series called this the gap between the demo and the Tuesday-afternoon reality, and said the gap comes out of your pocket. This last post is about seeing the Tuesday before you pay for it.

The disclosure that opened this series matters more here than anywhere else: I build and sell these systems. A vendor writing the buyer’s guide is an old trick: you publish a checklist, and every box happens to describe your own product. So judge this one by a different standard: whether it can be used against me. Everything below applies to what I sell, and one section ends with the conclusion that you shouldn’t buy from me. If a buyer’s guide can’t produce that sentence, it’s an ad.

Five questions that expose a weak vendor

The receptionist post closed with six questions for that specific product. Two of them generalize so well they reappear here; the other three are new, and they’re about the contract more than the software.

1. “Which number does this move, and how do we measure it from my data?” If you ran the stopwatch audit from the last post, you’re holding the answer to the second half: three response-time numbers, dated, measured before any vendor entered your life. A good vendor should be glad that baseline exists, because it’s the cheapest proof of value either of you will ever get. A vendor who resists measuring from your data, and steers you toward their dashboard’s definition of “engagement” instead, is planning to sell you a feeling.

2. “What happens when it doesn’t know?” Carried over from the receptionist post because it’s still the fastest filter. The acceptable answer involves the system saying so and routing to a person. The wrong answer is fluent improvisation. Follow it with the refusal list: which questions will it decline on purpose? A system pointed at a clinic or a law office that has no refusal list has not really met a clinic or a law office.

3. “Show me a real escalation transcript.” Also carried over, because it’s still the single most revealing artifact. Not a demo: a real conversation where the system hit its limit and handed off, with the caller’s context traveling to the human. Any vendor with customers has these transcripts. A vendor who won’t show one either has no customers or has escalations they’d rather you not read, and both of those are answers.

4. “If we part ways in a year, what leaves with me?” The question buyers ask least, at the moment it’s cheapest to ask. Your phone number. The WhatsApp number your customers have saved. The chat history. The contact and appointment records the system accumulated. By default, you find out what you own on the day you try to leave, which is the most expensive possible time to learn it. Get the answer in writing before signing; an honest vendor will put it there without flinching.

5. “What breaks when my business changes, and who fixes it?” Prices change. Hours go seasonal. You drop a service and add another. The system that answered perfectly in March can be confidently quoting January’s prices in June. A wrong confident answer, the receptionist post argued, costs more than no answer at all. So: who makes the change, how fast does it land, and what does it cost? “You edit it yourself” and “tell us and it’s live in a day” are both fine answers. “That requires a change order” is a price you want to hear before signing, not after.

Red flags that should end the meeting

“Fully autonomous sales agent.” The whole argument of post one. Anyone selling replacement in 2026 is selling ahead of what the technology reliably does, and the gap gets priced into your invoice, not theirs.

Pricing that punishes success: metered per-conversation pricing sounds fair until the system starts working. Run the math at the volume you’re hoping for, not the volume you have today. The red flag isn’t metering itself; it’s any pricing where the system underperforming would come as a relief to your bookkeeping.

No handoff design: ask where the human enters the conversation. If the answer is a shrug, or “it handles everything,” the vendor either hasn’t met the four failure modes from the receptionist post or is hoping you haven’t.

No failure case on offer: the receptionist post admitted that one of those failure modes has no clean fix. A vendor who claims to have none has stopped admitting things, and you’ll be the one paying the difference during the pilot, on your own leads.

Your identity as their asset: if the phone number or the WhatsApp identity your market knows is registered to the vendor, leaving them means re-teaching every customer how to reach you. This is question four wearing its worst-case outcome, and it’s common enough to deserve its own flag.

What a fair pilot looks like

Here the audit from the last post stops being homework and becomes leverage. A fair pilot has five properties, and you can insist on all of them.

After-hours only, to start. It’s the natural first slice: it’s where your audit number hurts, it competes with voicemail instead of with your team, and it fences the risk to hours when nothing human was going to answer anyway.

A baseline measured before launch: the three stopwatch numbers, written down and dated before the system goes live. Without the before picture, the after picture proves nothing; every system in the world “works” against no baseline.

Thirty to sixty days: long enough for real Tuesdays to happen — the angry caller, the ambiguous question, the week your prices change.

One number, agreed in advance, that decides renewal: median response time to after-hours leads, or appointments booked from after-hours contacts. Which number matters less than when it’s chosen: before launch, in writing, because after launch everyone’s judgment (the vendor’s, yours, mine) bends toward the sunk cost.

A clean exit: if the number doesn’t move, you leave with everything question four covered, at no cost beyond the pilot. A vendor confident in the product agrees to this quickly. Hesitation here is information.

A fair pilot process can also return two verdicts before it starts, and you should want a vendor willing to say both. “Don’t buy this at all”: if your audit came back in minutes, you already won the race, and the last post said it plainly — you didn’t need me. “Don’t buy this yet”: if your leads arrive in single digits per month, a thirty-day pilot measures noise, and no agreed number can honestly decide anything. Your bottleneck is lead flow, which is a marketing problem, mostly not an AI one. Fix that first and come back when there’s something to measure. Any vendor, me included, has to be willing to hand you those verdicts; one who takes your money anyway wants your subscription more than your outcome.

What it should cost

No dollar figure here, and be suspicious of one that arrives before questions about your business do. What I can give you honestly is the shape of the market and the shape of the decision.

Two pricing shapes dominate. The platform subscription: monthly fee, low setup, you configure it yourself. Cheaper and faster to start, template-shaped, and the configuration burden, including the refusal list, is yours. Built for you: a setup fee plus a smaller monthly, and someone else owns the configuration and the maintenance from question five. More up front, and it’s what businesses buy when their reality doesn’t fit a template: two languages in one phone line, refusals a regulated field requires, integration with the calendar you already run on.

Whichever shape, the same things push cost up: voice rather than chat, integrations with your records and scheduling systems, doing two languages properly instead of nominally, and the engineering that makes a system refuse reliably instead of improvising.

Then do the only comparison that matters: the system against the silence it replaces, not one vendor against another. Your audit told you what an after-hours lead currently experiences. You know what an average booked job, patient, or case is worth to your business. Multiply honestly, with your own numbers. I’m deliberately not handing you industry averages, because the last post showed what happens to numbers after a decade of vendor slide decks. If the coverage costs more than the leads it could rescue are worth, don’t buy it, from anyone. That arithmetic outranks every argument in this series.

The part where you shouldn’t buy from me

One more question, and this one cuts against me specifically: “How many people can maintain this system, and what happens to it if you disappear?”

A large platform survives losing any one person. A system built by a small operation — and mine is small — is as durable as its builder, its documentation, and its exit terms, and you’re allowed to price that risk. If what your business needs most is institutional continuity (a support desk, a service agreement, the certainty that the vendor still exists in five years), buy from an established platform, not from a solo builder like me, and accept the template fit that comes with that choice. If what you need most is a system shaped to your actual business (the bilingual line, the refusal list your field requires, the calendar you already use), a builder is what that takes, and you protect yourself the boring way: question four in writing, and documentation a successor could pick up.

I’d rather name that trade than get caught avoiding it. The checklist above only means something if it can point away from the person who wrote it.

The whole series in one paragraph

AI will not replace salespeople. It replaces the logistics inside sales — the 9:40pm lead, the twelve questions, the follow-up that never went out, the first qualification pass — because that work has checkable done states, while promises and judgment stay with people. At the front desk, that principle becomes a real job description with real failure modes — anger, refusals, callers who need a person, ambiguity — which is why the handoff to a human is the design decision worth inspecting hardest. One claim in the whole pitch requires no trust at all: speed to lead, a number you can measure with a stopwatch this week and that no staffing plan can fix. And buying the fix safely looks like this: baseline first, pilot small, one number decides renewal, own your exit, and make every vendor pass the same checklist, including the one who wrote it.

If you ran the audit and your number is Monday: applying this evaluation to your pipeline is a conversation I do. Bring the checklist. Ask me question one and question four first, and watch whether I flinch.

    Share:

    Enjoyed this post?

    Get new posts on web dev, AI and SEO straight to your inbox. No spam, unsubscribe anytime.

    No spam. By subscribing you agree to the privacy policy .

    Back to blog

    Related Posts

    View All Posts »
    Speed to Lead: The Sales Metric AI Actually Fixes
    EN

    Speed to Lead: The Sales Metric AI Actually Fixes

    You can't verify most AI sales claims without buying first. Speed to lead is the exception: how long a new lead waits for a real answer is a number, and you can measure yours this week with a stopwatch. I build these systems for a living, and this is the one part of my pitch you don't have to take on trust.

    What an AI Receptionist Actually Does All Day (and Where It Breaks)
    EN

    What an AI Receptionist Actually Does All Day (and Where It Breaks)

    Vendors sell the term 'AI receptionist' on warmth. Here's the shift log instead: what one actually handles in a day for a clinic or a contractor, what a good handoff to a human looks like, and the four places where it breaks. I build these systems, so the failure modes come first.