Direct answer. Block AI when the output would bind the company: a payment, a contract, an employment decision, a clinical or legal result, or a safety change. Weighted scores, including Claude Sonnet 4.6 at 87 and Truth Score 96, are evidence about drafts. They are not authority to make the final decision.
A final decision is one the business cannot cheaply undo: money leaves, a contract is signed, a person is hired or fired, a patient or customer is told a result is authoritative, or a safety control is changed. Blocking AI from that seat is a decision-rights rule. It is not the same page as the five ongoing guardrails, which cover approved tools, data handling, review, access and escalation after AI is already in use. No weighted score in the published set is a permission slip to bind the company. Claude Sonnet 4.6 at 87 overall and Truth Score 96 is still a drafting model. GPT-5.4 at 85 overall and Truth Score 82 is still a drafting model. Gemini 3 Flash at 86 overall is fast, and speed makes an unreviewed send more likely, not more legitimate. The block list below is the set of outcomes that stay with a named human even when the draft is excellent. The match engine can still choose the model that prepares the draft.
Final means the company is bound
A draft, a summary and a suggested reply can be wrong and still be cheap to discard. A payment, a signed term, a hiring decision or a safety change is not. The standing controls for tools and access are on AI governance basics. This page only draws the line where the model must not be the one who decides.
No score in the dataset is a licence
| Model | Weighted | Truth | What the number allows |
|---|---|---|---|
| Claude Sonnet 4.6 | 87 | 96 | A stronger draft. Not the signature. |
| Gemini 3 Flash | 86 | 80 | A faster draft. Speed is not authority. |
| GPT-5.4 | 85 | 82 | A widely integrated draft. Not the approval. |
Figures from /data/ai-comparison-2026.json. Last verified 2026-06-26. Scoring version v1.0. Affiliate links are off.
Outcomes that stay with a named human
- Money: refunds, payouts, pricing exceptions, credit.
- Contracts and filings, including anything a lawyer must own. See legal and medical questions.
- Employment: hiring, firing, discipline, performance ratings that affect pay.
- Safety controls, access revocation, and irreversible changes to production systems.
- Any customer message that grants a remedy. Detection of bad drafts is on support hallucination checks.
What the model may still prepare
It may summarise the file, propose options, and draft language a person will edit. What AI can't do is the wider limit. A small business choosing which jobs to assist, rather than which jobs to abdicate, should start from best AI for small business and the scoring method.
The draft still needs a model. Use the match engine for that choice. Do not use it as a signature.
FAQ
Does the highest Truth Score remove the need for a human decision?
No. Claude Sonnet 4.6's Truth Score of 96 is still a drafting mark. Binding outcomes stay with a named person.
Is a customer-support refund a final decision?
Yes, if the reply moves money or changes the account without a person. Drafting a suggested reply is not the same step.
Where are the day-to-day guardrails?
Approved tools, data handling, review, access and escalation are on the governance basics page. This page is only the block list.