Trust center / model evidence
Model evidence, without a universal ranking.
Use reputable evaluations to answer a defined workflow question. This registry links to source methods and boards; it does not copy current rank values, merge unlike metrics, or feed a quality score into the money formula.
Return to workbenchAdmission contract
01Question before orderEvery source must name the decision question it can and cannot answer.
02Method before metricPrompts, judges, categories, refresh behavior, and uncertainty travel with the link.
03Rights before displayUntil exact data reuse is cleared, the product links to the operator instead of reproducing the board.
Global evidence registry
Credible sources, kept dimension-specific.
Current status is deliberately conservative. “Link only” means the source can add discovery value without pretending the workbench owns or redistributes its results.
- Question
- How do people compare responses in blinded pairwise preference tests?
- Method
- Crowd-sourced pairwise human preference with evolving categories and confidence-aware ranking research.
- Version
- Live board / category-specific
- Limitation
- Preference, style, sampling, and user-selection effects do not establish workflow economics or factual reliability.
- Health
- Monthly method and operator review
- Rights
- Link only
- Question
- What do transparent, reproducible scenario evaluations show?
- Method
- Academic evaluation with published prompts, scenarios, metrics, and result context by board.
- Version
- Board and release specific
- Limitation
- Coverage and freshness differ across boards; an older narrow board cannot stand in for current general capability.
- Health
- Release and board review
- Rights
- Link only
- Question
- How do systems perform on refreshed objective questions with ground-truth answers?
- Method
- Dated multi-category releases designed to reduce contamination through periodic refresh.
- Version
- Release specific
- Limitation
- Difficulty and coverage move between releases, so cross-release order is not directly interchangeable.
- Health
- Release and repository review
- Rights
- Link only
- Question
- How do independent capability, speed, latency, and price measurements differ?
- Method
- Published methodology and defined API metrics with explicit attribution requirements.
- Version
- Continuously updated operator dataset
- Limitation
- Composite indices and provider mapping require interpretation; external redistribution needs separate permission.
- Health
- Monthly method and terms review
- Rights
- Link only
- Question
- How do coding systems perform on newly released contest problems?
- Method
- Continuously updated coding evaluation using newer contest problems and multiple code capabilities.
- Version
- Dataset release specific
- Limitation
- Contest coding does not represent repository maintenance, product architecture, or long-running agent work.
- Health
- Candidate / repository review
- Rights
- Candidate
Product boundary
Evidence informs a user-defined threshold. It never becomes an invisible multiplier in the economics.
When a report cites a source later, it should preserve the dated release, method, mapping confidence, and known limitations alongside the cost decision.