Skip to content
Otobiz
Sign in

AI for customer support6 min read

How to evaluate an AI customer messaging platform

Twelve questions that separate products in this category, plus what to do during a trial. Most of the differences only surface after you have committed.

Every product in this category has a shared inbox, an AI that replies, broadcasts, and a page of integration logos. The demos look similar because the demos are showing the same twelve features. The differences show up three months in, and by then you have moved your number.

These are the questions that surface those differences early. None of them are hostile, and a vendor who has thought about their product will enjoy answering most of them.

1. Who owns the WhatsApp Business Account?

Ask first, because it decides what everything else is worth. If the WhatsApp Business Account sits inside your own Meta business, the number, the approved templates, and the opt-in history are yours, and you can connect them elsewhere. If a provider holds it, changing vendor usually means starting again.

Follow-up questions: can we see the account in our own Meta Business Manager today? Do we pay Meta directly or do you resell messages to us? What is the exit process, in steps?

2. What does "supports Instagram" mean here?

Channel lists are the least reliable part of any vendor site. The word "supports" covers everything from a full two-way inbox with delivery status to a form that posts a message somewhere.

For each channel you actually need, ask four things: can we receive, can we reply, do we see delivery status, and what happens when the channel's own rules block a send. The last one is where the differences live, because every Meta channel has a 24-hour window and only WhatsApp has templates as a way around it.

3. What does the AI answer from?

Covered at length in what grounding means, and it comes down to one demand: show me an answer with its sources, then show me what happens on a question my material does not cover.

Then ask the second half, which most evaluations skip. Can it read my order records, or only my documents? "Where is my order" is the message that arrives most, and a product that can only answer from a policy page will deflect far less than the demo suggested.

4. What can it do without a person?

There is a large difference between a product that drafts and one that acts, and both are legitimate. What matters is that the boundary is explicit and that you set it.

  • Which categories send automatically, and can we turn that off per conversation and workspace-wide?
  • What is held for approval, and does an approval bind the exact text that will send? An approval that commits to a draft which can then change is not a control.
  • Can a person take over a thread mid-conversation, and does the AI stop when they do?

Ask to see the consent record for a contact. What you want to find: the wording shown, where and when it was collected, which purpose it covers, and the history of changes including opt-outs.

Then ask where suppression is checked. At send time is the correct answer; at audience-build time means someone can opt out between the two and still receive the message. And ask what happens when Meta returns a failure meaning the recipient opted out. That should suppress the contact automatically rather than land in a report.

6. Can we see quality and template health?

Numbers get restricted because of recipient behaviour, and the signals arrive per template. Ask whether you can see quality per template rather than only per number, and whether a paused template is surfaced before the next campaign or after it fails.

7. Does reporting show cost, or only sends?

Meta reports the billed category on the delivery receipt for every message. A platform that captures it can show you what a campaign actually cost by category. One that does not will show you send counts and let you multiply.

Ask to see a campaign report. If it has delivered, read, and failed counts but no billed categories, you will be reconciling invoices by hand.

While you are there, ask what the failure breakdown looks like. Grouping every failure as "failed" hides the difference between a temporary throttle, a closed window, and an opt-out, and those need three different responses.

8. How is our data separated from other customers'?

The answer to look for is structural rather than procedural. Workspace isolation enforced underneath the software means one business's records are unreachable from another's even when the software above has a bug. "Every query filters by account id" is a promise about code review.

Then: how are the logins to your channels stored? Encryption at rest with the key held separately is the standard answer. And where does the data physically live, which matters if you have a residency requirement.

9. Is there an audit trail, and can it be edited?

Ask what an audit trail records here: every send, every approval, every settings change, with who and when. Then ask the question that separates a log from a record: can past entries be edited? A record nothing can change after the fact, checked for tampering, gives you an answer when a customer disputes what they were sent. One the product can quietly rewrite proves less.

10. What do the integrations actually do?

Integration counts are the least meaningful number on any vendor site, because a listing can mean a full two-way sync or a single trigger.

Name the two or three systems you genuinely need before you look at any list. For each, ask exactly what is read, what is written, and what is kept in sync. Then ask whether there is a real API available to customers, and whether outbound webhooks exist so your systems learn about events without polling.

11. What is the pricing model, precisely?

Establish Meta's own per-message rates for your markets first, then work out what the vendor adds. Then get specific about their unit.

If it is monthly active contacts, ask whether an inbound-only contact counts, whether the count resets, and what happens when you exceed it. If it is seats, ask whether a read-only seat is billed. If it is credits, ask what a credit buys and whether unused ones expire.

Then price your own expected volume through each vendor's model rather than comparing headline plans. The order usually changes.

12. How do we get our data out?

Ask before you sign, not after you are unhappy. Conversation history, contacts with their consent records, and templates. In what format, through what mechanism, and how long it takes.

A vendor confident in their product finds this an easy question.

What to do during the trial

Vendor demos are rehearsed. Twenty minutes of your own testing tells you more than an hour of theirs.

  • Ask the AI something your content does not cover. Phrase it like a customer would. Watch whether it admits the gap or fills it.
  • Ask it something your content covers badly. Two documents that disagree is the normal state of any real knowledge base. See which one wins and whether you can tell.
  • Send yourself a message and let the window close. Then try to reply free-form the next day. A good product tells you the window closed and offers the template. A weak one shows a generic failure.
  • Opt out, then try to send. Reply "stop" from a test number, then attempt to include it in a broadcast. It should be excluded automatically, everywhere, not just in that campaign.
  • Have two people work one thread. Assignment, ownership, and whether the AI keeps drafting after a person takes over.
  • Read a report. Not the dashboard tour. An actual campaign report, and see whether it answers "what did this cost" and "why did these fail".
  • Break a template. Send with a variable missing and see what the error tells you. "Send failed" and "the template expects three variables, two were supplied" are different products.

The questions that are not worth asking

Some evaluation criteria feel rigorous and predict nothing.

Total integration count. Covered above. Three that work beat ninety listed.

Number of AI features. Sentiment tags, summaries, and suggested replies are cheap to add and mostly decoration. What matters is whether the AI answers correctly from your material and stops when it cannot.

Deflection rate quoted as a benchmark. Deflection is measured differently by every vendor and can be improved by making it harder to reach a person. Ask how they define it, and whether a resolved conversation and an abandoned one are told apart.

Customer logos. They tell you a company signed a contract. They tell you nothing about whether it works for a business your size, in your market, with your volume.

The useful signal is narrower than any of that: what the product does when it does not know the answer, what it refuses to send without you, and who owns the account when you leave.

Published by Otobiz on .