Hallucination decisions ยท High-stakes questions

Do legal and medical questions make AI hallucinate more?

Direct answer. These questions are riskier than an average Truth Score implies, because they demand a specific case, dose or statute. HIPAA true on Claude Sonnet 4.6, GPT-5.4 and Gemini 3.1 Pro is a data-handling flag, not permission to give the final legal or clinical answer. Keep the model on drafts a qualified person owns.

Legal and medical questions ask a model to recall a specific authority: a case, a dose, a contraindication, a statute. That is a different job from the average Truth Score, which compresses many tasks into one number. This page does not republish the sector rate table. It decides whether a higher overall truth mark, or a HIPAA flag, is enough to let the model give the final answer. It is not. Claude Sonnet 4.6 leads the published Truth Scores at 96 and is marked HIPAA true. GPT-5.4 has Truth Score 82 and HIPAA true. Gemini 3.1 Pro has Truth Score 84 and HIPAA true. Gemini 3 Flash has Truth Score 80 and HIPAA false. DeepSeek V3 has Truth Score 68, HIPAA false, and GDPR false. HIPAA in this dataset is a data-handling flag. It does not mean the model was cleared to practise law or medicine. The decision is to keep the model on drafts and admin text, and to keep the final legal or clinical answer with a qualified person.

An average Truth Score is the wrong instrument

A single Truth Score blends many tasks. Legal and medical questions ask for a named authority. Missing that authority is the failure mode, even when everyday summaries look fine. Sector ranges and the mitigation column stay on hallucination by industry. This page does not rebuild that table. It asks a narrower question: which published flags would a buyer misuse as permission?

HIPAA is a data gate

ModelTruthHIPAAGDPRUse on this question
Claude Sonnet 4.696truetrueStrongest truth mark. Still a draft model.
Gemini 3.1 Pro84truetrueEligible on the data flag. Not a clinician.
GPT-5.482truetrueEligible on the data flag. Not counsel.
Gemini 3 Flash80falsetrueDo not use for medical-record text.
DeepSeek V368falsefalseHosted API fails both flags in this dataset.

Figures from /data/ai-comparison-2026.json. Last verified 2026-06-26. Scoring version v1.0. Affiliate links are off.

The decision: drafts, not the final answer

Let a model summarise documents a lawyer or clinician already supplied. Do not let it invent the case, the dose or the filing. Pair that rule with unsourced citations and with when to block final decisions. Privacy handling for any record you do submit is on the privacy checklist.

Research workflows that still need sources belong on best AI for research. The methodology page explains why safety is only 10 percent of the weighted score, which is another reason the overall number cannot clear a legal answer.

Choosing a draft model under a privacy constraint? Weight safety and privacy in the match engine. The engine does not replace the qualified reviewer.

FAQ

Does HIPAA true mean a model is safe for medical advice?

No. In this dataset HIPAA is a compliance flag on the record. It is not a clinical validation.

Which profiled model is HIPAA false?

Gemini 3 Flash is marked HIPAA false, with Truth Score 80. Do not send medical-record text to it on the basis of this dataset.

Where are sector hallucination ranges discussed?

The industry guide owns that table. This page decides whether truth and HIPAA flags are enough to let the model answer.