Trust center / model evidence

Model evidence, without a universal ranking.

Use reputable evaluations to answer a defined workflow question. This registry links to source methods and boards; it does not copy current rank values, merge unlike metrics, or feed a quality score into the money formula.

Return to workbench

Admission contract

01Question before order

Every source must name the decision question it can and cannot answer.

02Method before metric

Prompts, judges, categories, refresh behavior, and uncertainty travel with the link.

03Rights before display

Until exact data reuse is cleared, the product links to the operator instead of reproducing the board.

Global evidence registry

Credible sources, kept dimension-specific.

Current status is deliberately conservative. “Link only” means the source can add discovery value without pretending the workbench owns or redistributes its results.

Question
How do people compare responses in blinded pairwise preference tests?
Method
Crowd-sourced pairwise human preference with evolving categories and confidence-aware ranking research.
Version
Live board / category-specific
Limitation
Preference, style, sampling, and user-selection effects do not establish workflow economics or factual reliability.
Health
Monthly method and operator review
Rights
Link only
Question
What do transparent, reproducible scenario evaluations show?
Method
Academic evaluation with published prompts, scenarios, metrics, and result context by board.
Version
Board and release specific
Limitation
Coverage and freshness differ across boards; an older narrow board cannot stand in for current general capability.
Health
Release and board review
Rights
Link only
Question
How do systems perform on refreshed objective questions with ground-truth answers?
Method
Dated multi-category releases designed to reduce contamination through periodic refresh.
Version
Release specific
Limitation
Difficulty and coverage move between releases, so cross-release order is not directly interchangeable.
Health
Release and repository review
Rights
Link only

Artificial Analysis

Open official source ↗
Question
How do independent capability, speed, latency, and price measurements differ?
Method
Published methodology and defined API metrics with explicit attribution requirements.
Version
Continuously updated operator dataset
Limitation
Composite indices and provider mapping require interpretation; external redistribution needs separate permission.
Health
Monthly method and terms review
Rights
Link only
Question
How do coding systems perform on newly released contest problems?
Method
Continuously updated coding evaluation using newer contest problems and multiple code capabilities.
Version
Dataset release specific
Limitation
Contest coding does not represent repository maintenance, product architecture, or long-running agent work.
Health
Candidate / repository review
Rights
Candidate

Product boundary

Evidence informs a user-defined threshold. It never becomes an invisible multiplier in the economics.

When a report cites a source later, it should preserve the dated release, method, mapping confidence, and known limitations alongside the cost decision.