AI for customer support7 min read
What grounding means for an AI support agent
Check how AI replies use policies and customer records, handle missing evidence, and show whether a failure needs better content or retrieval.


Grounding gives an AI support agent the approved evidence needed to answer a customer's question: policies, product information, customer records or verified tool results. It does not guarantee correctness, but it makes the answer testable and gives the system a reason to stop when reliable evidence is missing.
Grounding supplies evidence for the answer. Before using it with customers, check whether the system found the right evidence, interpreted it correctly, and had permission to use it.
The default behavior is the problem
Without your current policy, a model may draw on patterns from other return policies: a familiar return window, standard exclusions, or typical refund conditions. If that answer goes out from your business number, the customer may reasonably treat it as your policy. A fluent answer needs the same factual check as one written by a new colleague.
An invented delivery date, refund window, or appointment availability can create an expectation the business cannot meet. The team may need to correct the reply and resolve the customer’s request. Fluency alone is not evidence that the answer is accurate.
What grounding actually does
Grounding connects an answer to relevant evidence, such as approved documents, current records, or results from an authorized tool. A common approach uses retrieval followed by generation:
- Source. Your policies, product details, delivery information, past answers, and the customer's own records are collected and indexed.
- Retrieval. When a message arrives, the system finds the passages relevant to it.
- Constrained generation. The model writes an answer using those passages, and is instructed to work from them rather than around them.
An answer about your return policy should agree with the current policy and identify the supporting material. A citation is a starting point for checking: confirm that the passage actually supports the answer and applies to this customer’s situation.
The real test is the empty case
Test a normal customer question whose answer is missing from the supplied material. For example, a returns page might explain ordinary purchases but say nothing about clearance items. Ask whether a clearance purchase can be returned.
A useful response identifies the missing policy and gives the customer a next step. Depending on what is missing, it may ask for an order detail or route the question to someone who can decide. It should not invent a return deadline or promise an exception.
Check the evidence even when the reply includes a citation. A link to the general returns page does not prove that its terms cover clearance items. The reviewer needs to see which passage supports the answer and where uncertainty remains.
Then try a related question the material does answer. The system needs to find usable evidence as well as recognize its limits. Passing either case once is a starting point for evaluation, not a guarantee about every future conversation.
Grounding is not fine-tuning, and neither is "trained on your data"
These get used interchangeably in sales conversations and they mean different things.
Fine-tuning adjusts a model’s behavior using training examples. It can shape tone, format, and task performance. Changing business facts still need a reliable update process; training alone does not establish that an answer uses today’s policy.
Grounding can use evidence that is maintained outside the model. Check how edits, deletions, indexing, and caches affect when an updated source becomes available to the next answer.
"Trained on your data" does not explain how updates reach an answer. The question to ask is direct: when I change a sentence in my policy document, when does the next answer change? Ask for the actual update delay and a demonstration: change a source, wait for the documented update process, and ask the same question again.
Two kinds of source material, and both are needed
Most discussion of grounding covers documents. Customer messaging needs a second kind.
Document grounding answers general questions: what is the policy, how does the warranty work, do you ship there, what are the sizes.
Record grounding answers customer-specific questions: when is my appointment, which viewing did I request, has my enrollment been confirmed, or what is the status of my invoice or order. These require the correct record and permission to disclose it.
Choose test cases from your own conversation history. Ask which records the integration can read, how it identifies the customer, and what happens when the record is missing, stale, or belongs to someone else. Do not accept a plausible answer as proof that a lookup succeeded.
Keep the knowledge base current
Maintain the policies and business information the agent uses. Even a correctly retrieved passage can be wrong for today’s customer if the source has expired or applies to another branch.
The common failure is loading everything once and never revisiting it. Six months later the system is confidently quoting a policy that changed, and doing it well, in your voice, at scale. Grounding made the wrong answer more convincing rather than less.
Give source maintenance an owner, then review failures by their cause.
Give the content an owner. Somebody whose job includes the source material being current. Treat it like a page on your site, not like a one-time upload.
Investigate unanswered questions. First check whether the question is relevant to your business. If the answer already exists, inspect whether the source was indexed, found, current and available to this customer. An access restriction must not be bypassed just to produce an answer. Add or improve content when information is actually missing; repair retrieval or the integration when existing information is not reaching the agent. Group repeated failures by cause before deciding what to write next.
What a grounded answer should show you
Grounding is a claim, and claims that cannot be checked are worth little. Three things make it verifiable.
Named sources on the draft. Which passages the answer used, visible to the person reviewing it, before it sends. A source list turns "does this sound right" into "is this the right paragraph".
An honest confidence signal, wired to something. Confidence is only useful if a low score changes behavior: holding for approval, asking a clarifying question, or declining to answer without a reliable source. A number displayed next to a message that sends regardless is decoration.
A record of what was sent. An activity history covering what the AI drafted, who approved it, and what actually went out. When a customer quotes something back to you, this helps your team trace the reply.
Where automation should stop
Grounding reduces the risk of a wrong answer. It does not make an answer safe to send unsupervised in every case.
Set approval rules around authority and impact. A routine confirmation may be automated when the underlying action is authorized and verified. A discretionary refund, unsupported commitment, or uncertain answer may need review. An approval should bind the exact action or reply reviewed, and a confirmation should follow evidence that the action succeeded.
Start with a workflow whose answers and outcomes your team can verify. Expand only after checking its accuracy, exceptions, and handovers against real business records.
What to ask a vendor
- Show me an answer with its sources. Which specific passages did it use?
- What happens when retrieval finds nothing? Show me, on a question my content does not cover.
- Can it read the appointment, lead, invoice, or order records this workflow needs, with the correct access controls?
- When I edit a policy, when does the answer change?
- What is held for approval, and does the approval bind the exact text?
- Can you show me the questions it could not answer last week?
Record which answers were demonstrated, which were described, and which remain untested. Revisit the untested items before allowing the workflow to communicate with customers.
How to test AI replies before customers see them
Start with a small set of real customer questions and write down the correct answer before testing. Keep personal customer information out of a shared test sheet.
| Test | What to check |
|---|---|
| A current price or policy | The answer agrees with the current source. |
| A fact missing from your website | It asks for help or admits the gap. |
| Two conflicting policy pages | It does not quietly invent a compromise. |
| A follow-up such as “what about tomorrow?” | It keeps the earlier context and checks availability. |
| A request to refund or book | It distinguishes a request from a completed action. |
| A request for a person | It gives the customer a useful next step. |
Download the reply-quality worksheet. Add your questions, expected answers and the source each answer should use. Record whether a failure came from missing information, retrieval, reasoning or the next action. Fix that cause and rerun the same questions, including previously passing ones.
This follows a sound evaluation practice: use examples representative of the task, include difficult cases, and review changes with human judgement. Passing a small worksheet is a starting point, not proof that every customer conversation is safe to automate.
A public demonstration should come after this evaluation. During setup, keep testing private and keep actual customer sends under your chosen approval controls. You can explore a prepared customer journey while assembling your own business information.
Further reading
NIST's Generative AI Profile describes confident false output as confabulation and sets out a neutral framework for managing that risk.
Published by Otobiz on . Last updated .