Make it yours
Make a custom ranking reflect your deployment priorities. Capability evidence drives the score; provider access, runtime constraints, and token cost remain separate.
Published custom ranking is unavailable.
Entered quality weight total: 90. The formula divides by this total; weights are not silently redistributed.
Set these axes to zero, or explicitly apply equal weights to published axes.
Chart: top 20 plus explicitly selected eligible models. PNG exports both charts with weights and source context; CSV includes every matching row.
Observed runtime constraints
Absent runtime measurements are unobserved. Hiding unobserved/outside-SLA rows requires both qualified measurements to meet the selected thresholds.
TTFT (seconds)
Output speed (tok/s)
Exact SLA measurements
| Model | TTFT seconds | TTFT assessment | Throughput tok/s | Throughput assessment | Combined SLA |
|---|
Weighted score vs. cost
Prices shown were recorded with these benchmark results; current prices may differ. Exact route and catalog revision are included in the source receipt.
Blended price = (3 × input price + output price) / 4, in USD per 1M tokens: an explicit 3:1 input/output assumption, excluding cache, batch and long-context adjustments. The blend is rounded upward to one microUSD per 1M tokens (less than $0.000001 above the exact blend); source integer rate precision is preserved in the receipt. 0 ranked rows lack a verified route with both token prices. These models remain eligible for quality ranking; their cost frontier is unavailable.
Score frontier
Verified input and output token prices are not reported for these rows. Quality ranking remains available independently.
Cheapest-first score ranking
Exact score and token-cost values (cheapest 20)
| Cost order | Model | Provider | Weighted score | Blended $ / 1M tokens USD | Frontier |
|---|
| Select / rank | Model | Provider | Weighted | Blended $ / 1M | Input $ / 1M | Output $ / 1M | TTFT s | Throughput tok/s | SLA | Frontier |
|---|
Methodology & source receipt
Rank-percentiles are derived from a published rank and exact cohort: 100 × (cohort size − rank) / (cohort size − 1). Only matching published capability definitions are compared. Agentic, Coding, Reasoning, Math and Multimodal each require their own exact source binding. Throughput is a separately qualified runtime measurement and never enters the quality score.
Exact weights, axis definitions and publication lineage
{
"weights": {
"agentic": 20,
"coding": 20,
"reasoning": 20,
"mathematics": 15,
"multimodal": 15
},
"revision": null,
"cacheRevision": null,
"tokenCostBlend": "3:1 input/output; ceiling to microUSD per 1M tokens",
"benchmarkCatalogRevision": null,
"currentCatalogRevision": null,
"axes": null
}Current output page: per-model capability, price and runtime conditions
[]
