Model evidence

Model evidence for the decision in front of you.

Cost, response speed, and task fit answer different questions. This local interface demonstrates the comparison structure without redistributing rights-uncleared external scores.

Illustrative local preview — not current provider data
01

Cost efficiency

Lower is better
  1. 1Fast utility$0.28
  2. 2Balanced generalist$0.42
  3. 3Reasoning frontier$0.63
  4. 4Open deployment$0.68
02

Task fit

Scenario evidence
  1. 1Reasoning frontierStrong
  2. 2Balanced generalistStrong
  3. 3Open deploymentMapped
  4. 4Fast utilityMapped
03

Response speed

Faster is better
  1. 1Fast utility0.38s
  2. 2Balanced generalist0.62s
  3. 3Open deployment0.74s
  4. 4Reasoning frontier0.88s

Interface preview only. Production values require approved attribution and redistribution rights.

SOURCE LEDGER

Evidence sources stay independent

The production tracker will preserve each source's method, version, date, limitations, mapping confidence, and rights state. It will not synthesize a universal best-model score.

Artificial Analysis

Independent intelligence, speed, latency, and pricing evidence

Open source

LMArena

Crowd-sourced human preference evidence

Open source

Stanford HELM

Transparent multi-scenario academic evaluation

Open source