Skip to content
Start free trialFree trial

AI for customer support8 min read

How to evaluate an AI customer messaging platform

Twelve questions that separate products in this category, plus what to do during a trial. Most of the differences only surface after you have committed.

A messaging-platform assessment board inspected with a violet magnifying lens.A messaging-platform assessment board inspected with a violet magnifying lens.

To choose a WhatsApp AI chatbot or customer messaging platform, start with the work your team needs to complete: qualify a request, arrange an appointment, answer a support question or follow up on a purchase. Similar feature lists can conceal different permissions, integrations, costs and review behavior.

Use these questions to shortlist platforms, then repeat the same scenarios during each trial.

1. Who owns the WhatsApp Business Account?

Ask first, because it decides what everything else is worth. If the WhatsApp Business Account sits inside your own Meta business, the number and approved templates are visible in WhatsApp Manager and remain under your control. Meta says businesses onboarded through Embedded Signup own their WhatsApp assets. Consent records are separate data held by the software platform, so ask how those are exported too. If a provider owns the account, changing vendor may depend on a migration that provider controls.

Follow-up questions: can we see the account in our own Meta Business Manager today? Do we pay Meta directly or do you resell messages to us? What is the exit process, in steps?

2. What does "supports Instagram" mean here?

Channel lists are the least reliable part of any vendor site. The word "supports" covers everything from a full two-way inbox with delivery status to a form that posts a message somewhere.

For each channel you actually need, ask four things: can we receive, can we reply, do we see delivery status, and what happens when the channel's own rules block a send. The last one is where the differences live. WhatsApp has a 24-hour customer service window, while Meta's Messenger and Instagram policy uses a 24-hour standard messaging window with narrower exceptions. WhatsApp templates provide a structured way to reach people outside its window.

3. What does the AI answer from?

Covered at length in what grounding means, and it comes down to one demand: show me an answer with its sources, then show me what happens on a question my material does not cover.

Then test the customer-specific part of your workflow. Can it read the correct appointment, property request, enrollment, invoice, or order record? Check how it verifies the customer and handles a record it cannot access. Choose the examples from your own inbox rather than assuming one type of question dominates every business.

4. What can it do automatically?

There is a large difference between a product that drafts and one that acts, and both are legitimate. What matters is that the boundary is explicit and that you set it.

  • Which categories send automatically, and can we turn that off per conversation and brand-wide?
  • What is held for approval, and does an approval bind the exact text that will send? An approval that commits to a draft which can then change is not a control.
  • What happens when confidence is low: does it ask a clarifying question, hold the reply for approval, or send anyway?

Ask to see the consent record for a contact. Meta's current opt-in guidance requires a clear agreement to receive communication from the named business and compliance with local law. What you want to find in the product is more useful: the wording shown, where and when it was collected, which purpose it covers, and the history of changes including opt-outs.

Then ask where suppression is checked. At send time is the correct answer; at audience-build time means someone can opt out between the two and still receive the message. And ask what happens when Meta returns a failure meaning the recipient stopped marketing messages. That should suppress future marketing automatically rather than land in a report.

6. Can we see quality and template health?

Numbers can be restricted when recent recipient feedback stays poor, and templates have their own quality and status signals. Ask whether you can see quality per template rather than only per number, and whether a paused template is surfaced before the next campaign or after it fails. Meta documents these states in its template status reference.

7. Does reporting show cost, or only sends?

Meta includes billing information in outbound message status webhooks. Those receipts help identify billable messages and categories. Currency cost also requires the applicable rate card, volume tier, and provider charges; reconcile estimates with the invoice.

Ask to see a campaign report that distinguishes sends, deliveries, billable events, estimated cost, and invoiced cost. Check which components are included and which still need reconciliation.

While you are there, ask what the failure breakdown looks like. Grouping every failure as "failed" hides the difference between a temporary throttle, a closed window, and an opt-out, and those need three different responses.

8. How is our data separated from other customers'?

Ask how brand isolation is enforced across storage, search, integrations, and staff access. Request evidence that one business cannot retrieve another’s records through the relevant workflows. A design description is useful, but it is not an absolute guarantee against every implementation defect.

Then ask how channel credentials are protected, who can access them, and how access is revoked. Confirm where data is stored and processed if location or residency requirements affect your business.

9. What does the activity history record?

Ask what the activity history records: every send, approval, and important settings change, with who and when. Then ask whether it is tamper-evident. A normal history helps the team trace events. A compliance archive needs stronger controls that make later changes detectable. Otobiz currently provides the former, not the latter.

10. What do the integrations actually do?

An integration listing can mean a full two-way sync or a single trigger. Check the work it completes rather than the number of logos.

Name the two or three systems you genuinely need before you look at any list. For each, ask exactly what is read, what is written, and what is kept in sync. Then ask whether there is a real API available to customers, and whether outbound webhooks exist so your systems learn about events without polling.

11. What is the pricing model, precisely?

Establish Meta's current per-message rates for your markets first, then work out what the vendor adds. Then get specific about their unit.

If it is monthly active contacts, ask whether an inbound-only contact counts, whether the count resets, and what happens when you exceed it. If it is seats, ask whether a read-only seat is billed. If it is credits, ask what a credit buys and whether unused ones expire.

Then price your own expected volume through each vendor's model rather than comparing headline plans. Include the same billing term, message volume, team size, branches, and required integrations in each estimate.

12. How do we get our data out?

Ask before you sign, not after you are unhappy. Conversation history, contacts with their consent records, and templates. In what format, through what mechanism, and how long it takes.

What to do during the trial

Use a sandbox or a dedicated test number and contacts you control. Keep test audiences separate from real customers. Agree which checks are previews and which make an actual provider request before running them.

  • Ask the AI something your content does not cover. Phrase it like a customer would. Watch whether it admits the gap or fills it.
  • Ask it something your content covers badly. Two documents that disagree is the normal state of any real knowledge base. See which one wins and whether you can tell.
  • Send yourself a message and let the window close. Then try to reply free-form the next day. A good product tells you the window closed and offers the template. A weak one shows a generic failure.
  • Opt out, then try to send. Reply "stop" from a test number, then attempt to include it in a broadcast. It should be excluded automatically, everywhere, not just in that campaign.
  • Review one thread together. Ask a colleague who has not seen it to find the owner, last promise and next action on the device they normally use. Check whether agent drafting follows your review rules.
  • Try a repeat contact across channels. Use test contacts to check how records are matched and who owns the next action. Two different people with the same name must not have their histories combined.
  • Read a report. Not the dashboard tour. An actual campaign report, and see whether it answers "what did this cost" and "why did these fail".
  • Check template validation. In a draft or sandbox, omit a required variable and inspect the feedback before any send. If an actual provider failure must be tested, use the agreed test account and recipient.

Treat headline metrics as supporting evidence

Feature counts and customer logos can help form a shortlist. They do not establish that a workflow works for your business.

Total integration count. Covered above. Three that work beat ninety listed.

Number of AI features. Test whether a summary preserves important details, a suggested reply uses current evidence, and automation stops when information or authority is missing. Count useful outcomes rather than feature labels.

Deflection rate quoted as a benchmark. Deflection is measured differently by every vendor and can be improved by making it harder to reach a person. Ask how they define it, and whether a resolved conversation and an abandoned one are told apart.

Customer logos. Check the relationship and the scope of the work. A relevant case study should explain the workflow, starting conditions, measurement period, and results; a logo alone does not establish those facts.

The useful signal is narrower than any of that: what the product does when it does not know the answer, what it refuses to send without you, and who owns the account when you leave.

Keep a workflow scorecard

For each scenario, record demonstrated, partly demonstrated, or not demonstrated, with the evidence and any missing step.

ScenarioEvidence to request
A service appointment changesCorrect availability, an authorized update, and the saved appointment record
A property request needs a viewingCorrect lead ownership and a verified next step
A learner asks about enrollmentCurrent course information and access only to the correct learner’s record
A distributor receives a quote requestCaptured requirements and review before an unsupported price commitment
A customer needs a personA visible owner, conversation context, and automation yielding to the team

Select the scenarios that match your business. A working text reply is only one part of each outcome.

Repeat one scenario after the underlying record changes: move the test appointment or replace a course date, then ask again. The next answer should use the current record. Also ask the vendor to show an unavailable integration using agreed test data. The conversation should explain what remains unresolved and who will act next, rather than announce a result the system could not verify.

Put your own assumptions to the test

Use the missed-opportunity calculator to explore a scenario with your own conversation volume and sales assumptions. It estimates potential sales value, not profit or a promised conversion lift. Then use the cost calculator to see the plan payment and message-credit budget separately.

For a practical first step, personalize a message and compare it with how your team currently responds. A useful evaluation should end with a concrete conversation to improve, not just a feature checklist.

Published by Otobiz on . Last updated .

Ready to put this to work?

Start with one customer journey and build from there.

7-day free trial. Nothing is charged when you start. Cancel before it ends and pay nothing.