FairAttest — FACT CertifiedFairAttest
← FairAttestFACT certification criteria →
Methodology & scientific defensibility

How we certify the accuracy of AI, and why the result holds up

FairAttest tests AI against controlled synthetic cases with fully known ground truth. Every assertion is graded one claim at a time, every omission is weighted by severity, contested calls are decided by reviewers, and tiers are awarded on the conservative bound of statistical confidence. The complete methodology is published below as a versioned document.

Read the document (PDF)v1.0.0 · published Aug 24, 2026 · 351 KB
PDF

Current methodology document

FairAttest Methodology & Defensibility

v1.0.0 · published Aug 24, 2026 · 351 KB

Initial FairAttest methodology: independent AI accuracy certification against controlled synthetic cases, claim-by-claim grading, reviewer adjudication, and conservative statistical tiering.

How accuracy is measured

Every certification is scored on the same categories. They stay constant across versions of the methodology; each published document sets out the exact thresholds and procedures in force for that version.

Primary safety signal

Fabrication rate

Assertions the AI states that the source never supports, or that contradict it. The most heavily weighted signal, because invented facts are the most dangerous failure mode.

Completeness

Severity-weighted omission rate

Required facts the AI dropped, weighted by severity from 1 (minor) to 4 (safety-critical). A missed high-stakes fact about a subject counts far more than a missed administrative detail.

Over-reach

Unsupported inference rate

Claims that reach beyond the evidence without being outright fabrications, such as turning a reported input into a stated conclusion.

Hard gate

Cross-subject containment

A scan for any detail from one subject appearing in another subject’s output. Reported as pass or fail and treated as a hard gate, never blended into any average.

Rigor

Statistical confidence

Every rate is reported with a confidence interval, and tiers are awarded on the conservative bound, so a result is never rated stronger than the evidence supports.

Human authority

Reviewer adjudication

Reviewers review and can override every consequential judgment before a score is final, and the attestation is human signed.

Why this document matters

A certification is only worth as much as the process behind it. AI is being deployed on trust that buyers cannot independently verify: vendors grade themselves on private data, pilots have no ground truth, and generic benchmarks do not speak the language of operational risk. Publishing our methodology in full is how we earn that trust rather than asking for it. It lets a buyer, a reviewer, a regulator, or an auditor check our work, compare vendors on identical terms, and understand precisely what a tier does and does not claim.

Independent and transparent

The full criteria are public, with no private scoring adjustments. Anyone can read exactly how a result was produced before they trust it.

Scientifically grounded

Outputs are tested against known ground truth, graded as individual claims, and reported with confidence intervals. Tiers gate on the conservative bound, so a result is never stronger than its evidence.

Human authority

A reviewer has final say on every consequential judgment, and the attestation is human signed. The machine accelerates the work; it never has the last word.

Immutable and versioned

The standard in force at test time is recorded with every result, and scoring rules are frozen into each run. Re-tuning the standard can never silently rewrite a past certification.

Designed to be audited, not taken on faith

Because the standard is versioned and every result records the exact version it was scored under, a reader can always reconstruct the rules that applied to any certification, even years later. Nothing is hidden in a black box and nothing changes retroactively. This page exists so the reasoning behind every FairAttestattestation is open to scrutiny.

This methodology document is an independent description of the FairAttest evaluation process. It is not a government or regulatory approval. All evaluations use synthetic cases and no real subject data is ever used.