04 / LLM EVALUATION

Check the claims, not the confidence

A fluent answer is not an evidence trail. Evaluate AI-generated claims against inspectable sources, with uncertainty and human review built in.

Neon Check the LLM typography with claim and evidence checks and an unresolved assessment.

Define what your check is allowed to conclude

Checking an LLM answer against supplied documents is a different task from researching it against additional sources. An answer can faithfully repeat a flawed document, and a generally correct statement can be unsupported by the particular evidence supplied. Name that boundary in the result.

For a document-grounded workflow, ask whether each important assertion is supported by the approved collection. For a wider research workflow, record where the system searched and which limitations remain. Do not describe a restricted search as comprehensive verification of the public record.

Evaluate at the claim level

The FActScore research paper describes decomposing generated text into atomic facts and evaluating support from a reliable knowledge source. That principle is useful for detecting mixed support within a long answer. It does not establish the present-day accuracy of any service described elsewhere.

Preserve the original answer and extract statements without losing attribution, dates, or qualifiers. Associate each assessment with a passage. Review whether the conclusion follows from that passage instead of treating shared keywords as proof. An explanation that introduces additional unsupported reasoning is not a repair for weak evidence.

Diagnose the stage that failed

A proposed pipeline contains extraction, retrieval, assessment, and presentation. Check them separately. A correct assessor may return insufficient evidence because the retrieval stage missed the relevant document. An accurate source may be attached to the wrong claim because extraction resolved an entity incorrectly.

Use targeted tests for date mismatches, changed denominators, negation, missing information, and causal overstatement. Keep human rationales with expected labels. Where reviewers disagree about the rubric, clarify the policy before treating every disagreement as a model failure.

Make your scorecard useful

Track support quality, claim coverage, unresolved cases, and reviewer reversals separately. Define denominators and evaluate important slices such as numbers, quotations, and time-sensitive assertions. An answer that avoids difficult facts may look precise while being incomplete. A checker that abstains from everything may be safe from confident errors but operationally unhelpful.

When comparing configurations, use the same claims and evidence boundary. Inspect examples whose labels changed, especially new supported results. Include review effort and correction work in an operational comparison rather than deciding solely from an average automated score.

Keep a person in the decision loop

For consequential material, show the claim, source, passage, and limitations to a reviewer with appropriate domain knowledge. Preserve corrections and model-version information. A second model’s agreement is not an independent source of evidence.

Read the complete LLM evaluation guide, the evidence-grounded citation workflow, and the service-selection checklist. FactAPI.com publishes independent guidance; it does not promise a live AI fact-checking service or automated certification of arbitrary answers.

Follow the claim. Find the evidence.

Go from a better question to a clearer evidence trail—one fieldnote at a time.

Explore the journal