Direct answer. Compare the reply with the help centre, the order system and the price list before it is sent. Gemini 3 Flash and GPT-5.4 list customer service as a best-fit job, with Truth Scores of 80 and 82. Those scores do not detect a bad ticket. A mismatch with a system of record does.
Customer-support hallucinations show up as tickets, not as a leaderboard. A reply quotes a refund rule that is not in the help centre, invents an order status, or states a price the catalogue does not charge. Detecting that is a review of the message against systems of record. It is not the same task as reducing hallucinations in general, and it is not a ranking of which model fabricates most on a public benchmark. The published dataset does name models whose best-fit jobs include support. Gemini 3 Flash lists customer service bots, with speed 95, cost efficiency 89 and Truth Score 80. GPT-5.4 lists customer service among its best-fit jobs, with Truth Score 82 and task 92. Claude Sonnet 4.6 does not list customer service as a best-fit job. Its Truth Score is 96, which matters when the queue needs caution more than raw throughput. None of those scores detect a bad ticket on their own. Detection is a set of checks a person or a monitor can run before the reply is sent.
Detection is a ticket check
A support hallucination is a sentence the customer might act on: a refund ceiling, a delivery date, a plan name, a policy exception. The check is whether that sentence exists in a system you control. General reduction methods, including retrieval, are documented on how to reduce hallucinations. Use them to lower how often drafts fail. Use this page to notice the draft that already failed.
Signals in the reply
- A policy number, exception or refund amount that the help centre does not state.
- An order status or tracking event that the order system does not show.
- A price, discount or plan name that the catalogue does not sell.
- A citation or help-article title that does not open. See unsourced citations.
- A confident apology that promises a remedy no agent is allowed to grant.
Models aimed at support in the dataset
| Model | Support on the best-fit list? | Truth | Speed |
|---|---|---|---|
| Gemini 3 Flash | Yes — customer service bots | 80 | 95 |
| GPT-5.4 | Yes — customer service | 82 | 86 |
| Claude Sonnet 4.6 | No — coding, writing, analysis, API, agents | 96 | 82 |
Figures from /data/ai-comparison-2026.json. Last verified 2026-06-26. Scoring version v1.0. Affiliate links are off.
Speed 95 on Gemini 3 Flash is why a bad reply can leave the queue quickly. Pair any of these models with the customer service guide and the tone guide. Tone does not fix a fabricated policy.
When a signal fires
Hold the reply. Correct it from the system of record, or send it to a person. Do not ask the same model to "confirm" its own sentence and treat agreement as proof. If the queue is about to let the model grant refunds or change accounts, that is a final decision and belongs on blocking final decisions.
Picking the support model? The match engine can weigh speed, cost and privacy. Detection still happens in the ticket, not in the score.
FAQ
Which published models list customer service as a best fit?
Gemini 3 Flash lists customer service bots. GPT-5.4 lists customer service. Claude Sonnet 4.6 does not.
Does a higher Truth Score detect a bad ticket?
No. Detection is a comparison between the reply and systems of record. Truth Score is a model-level mark, not a per-ticket alarm.
Is this the same as reducing hallucinations?
No. Reduction techniques are a different page. This page is the check you run on a support reply before the customer sees it.