How we certify the accuracy of AI, and why the result holds up
FairAttest tests AI against controlled synthetic cases with fully known ground truth. Every assertion is graded one claim at a time, every omission is weighted by severity, contested calls are decided by reviewers, and tiers are awarded on the conservative bound of statistical confidence. The complete methodology is published below as a versioned document.
Current methodology document
FairAttest Methodology & Defensibility
v1.0.0 · published Aug 24, 2026 · 351 KB
Initial FairAttest methodology: independent AI accuracy certification against controlled synthetic cases, claim-by-claim grading, reviewer adjudication, and conservative statistical tiering.
How accuracy is measured
Every certification is scored on the same categories. They stay constant across versions of the methodology; each published document sets out the exact thresholds and procedures in force for that version.
Fabrication rate
Assertions the AI states that the source never supports, or that contradict it. The most heavily weighted signal, because invented facts are the most dangerous failure mode.
Severity-weighted omission rate
Required facts the AI dropped, weighted by severity from 1 (minor) to 4 (safety-critical). A missed high-stakes fact about a subject counts far more than a missed administrative detail.
Unsupported inference rate
Claims that reach beyond the evidence without being outright fabrications, such as turning a reported input into a stated conclusion.
Cross-subject containment
A scan for any detail from one subject appearing in another subject’s output. Reported as pass or fail and treated as a hard gate, never blended into any average.
Statistical confidence
Every rate is reported with a confidence interval, and tiers are awarded on the conservative bound, so a result is never rated stronger than the evidence supports.
Reviewer adjudication
Reviewers review and can override every consequential judgment before a score is final, and the attestation is human signed.
Why this document matters
A certification is only worth as much as the process behind it. AI is being deployed on trust that buyers cannot independently verify: vendors grade themselves on private data, pilots have no ground truth, and generic benchmarks do not speak the language of operational risk. Publishing our methodology in full is how we earn that trust rather than asking for it. It lets a buyer, a reviewer, a regulator, or an auditor check our work, compare vendors on identical terms, and understand precisely what a tier does and does not claim.
Independent and transparent
The full criteria are public, with no private scoring adjustments. Anyone can read exactly how a result was produced before they trust it.
Scientifically grounded
Outputs are tested against known ground truth, graded as individual claims, and reported with confidence intervals. Tiers gate on the conservative bound, so a result is never stronger than its evidence.
Human authority
A reviewer has final say on every consequential judgment, and the attestation is human signed. The machine accelerates the work; it never has the last word.
Immutable and versioned
The standard in force at test time is recorded with every result, and scoring rules are frozen into each run. Re-tuning the standard can never silently rewrite a past certification.
Designed to be audited, not taken on faith
Because the standard is versioned and every result records the exact version it was scored under, a reader can always reconstruct the rules that applied to any certification, even years later. Nothing is hidden in a black box and nothing changes retroactively. This page exists so the reasoning behind every FairAttestattestation is open to scrutiny.
This methodology document is an independent description of the FairAttest evaluation process. It is not a government or regulatory approval. All evaluations use synthetic cases and no real subject data is ever used.