Key implications
Insufficient shared-metric coverage. Each finding is tied to a published metric, price, context, or modality fact.
- Input API price: GPT-5.6 Terra has the lower verified rate ($2.5 / 1M tokens vs $5 / 1M tokens).
- Output API price: GPT-5.6 Terra has the lower verified rate ($15 / 1M tokens vs $25 / 1M tokens).
- Context window: GPT-5.6 Terra has the larger published context window (1,050,000 tokens vs 1,000,000 tokens).
Shared metric view
A radar is shown only when at least four compatible supported score metrics are published.
Comparable metric detail
- AgenticClaude Opus 4.8: 61.78 · GPT-5.6 Terra: 58.86
- CodingClaude Opus 4.8: 71.98 · GPT-5.6 Terra: 69.29
- KnowledgeClaude Opus 4.8: 86.8 · GPT-5.6 Terra: 84.2
- MathClaude Opus 4.8: 66.8 · GPT-5.6 Terra: 97
- MultimodalGroundedClaude Opus 4.8: 85.4 · GPT-5.6 Terra: 71.4
- ReasoningClaude Opus 4.8: 68.3 · GPT-5.6 Terra: 78.1
- OverallClaude Opus 4.8: 76.6 · GPT-5.6 Terra: 72.95
Source metrics
Friendly metric names and published units stay visible. Missing measurements remain unavailable rather than becoming a score.
Source metric comparison| Metric | Unit | Claude Opus 4.8 | GPT-5.6 Terra |
|---|
| Agentic | score | 61.78 | 58.86 |
|---|
| Coding | score | 71.98 | 69.29 |
|---|
| Knowledge | score | 86.8 | 84.2 |
|---|
| Math | score | 66.8 | 97 |
|---|
| MultimodalGrounded | score | 85.4 | 71.4 |
|---|
| Reasoning | score | 68.3 | 78.1 |
|---|
| Overall | score | 76.6 | 72.95 |
|---|
Agentic
- Unit
- score
- Claude Opus 4.8
- 61.78
- GPT-5.6 Terra
- 58.86
Coding
- Unit
- score
- Claude Opus 4.8
- 71.98
- GPT-5.6 Terra
- 69.29
Knowledge
- Unit
- score
- Claude Opus 4.8
- 86.8
- GPT-5.6 Terra
- 84.2
Math
- Unit
- score
- Claude Opus 4.8
- 66.8
- GPT-5.6 Terra
- 97
MultimodalGrounded
- Unit
- score
- Claude Opus 4.8
- 85.4
- GPT-5.6 Terra
- 71.4
Reasoning
- Unit
- score
- Claude Opus 4.8
- 68.3
- GPT-5.6 Terra
- 78.1
Overall
- Unit
- score
- Claude Opus 4.8
- 76.6
- GPT-5.6 Terra
- 72.95
Pricing and context
Verification is shown beside each selected route. Missing facts remain Not verified.
Route pricing and context comparison| Field | Unit | Claude Opus 4.8 | GPT-5.6 Terra |
|---|
| Input API price | USD / 1M tokens | $5 | $2.5 |
|---|
| Cached input API price | USD / 1M tokens | Not verified | $0.25 |
|---|
| Output API price | USD / 1M tokens | $25 | $15 |
|---|
| Route context | tokens | 1,000,000 | 1,050,000 |
|---|
| Input modalities | published list | Not verified | Not verified |
|---|
| Output modalities | published list | Not verified | Not verified |
|---|
Input API price
- Unit
- USD / 1M tokens
- Claude Opus 4.8
- $5
- GPT-5.6 Terra
- $2.5
Cached input API price
- Unit
- USD / 1M tokens
- Claude Opus 4.8
- Not verified
- GPT-5.6 Terra
- $0.25
Output API price
- Unit
- USD / 1M tokens
- Claude Opus 4.8
- $25
- GPT-5.6 Terra
- $15
Route context
- Unit
- tokens
- Claude Opus 4.8
- 1,000,000
- GPT-5.6 Terra
- 1,050,000
Input modalities
- Unit
- published list
- Claude Opus 4.8
- Not verified
- GPT-5.6 Terra
- Not verified
Output modalities
- Unit
- published list
- Claude Opus 4.8
- Not verified
- GPT-5.6 Terra
- Not verified
Evidence provenance
Source records, route identity, timestamps, and methodology are consolidated here without declaring either model a winner.
- Publication time
- Sep 14, 2026, 12:37 AM UTC
- Freshness
- Stale — Published benchmark evidence includes benchlm content observed outside the 8-day evidence window.
- Methodology
- benchlm: benchlm_raw_composite
- Model records
- Claude Opus 4.8 — source benchlm · artifact models · model claude-opus-4-8
- GPT-5.6 Terra — source benchlm · artifact models · model gpt-5-6-terra
- Selected price routes
- Claude Opus 4.8 — route benchlm:claude-opus-4-8 · source benchlm · provider anthropic
- GPT-5.6 Terra — route benchlm:gpt-5-6-terra · source benchlm · provider openai
Other reviewed matchups from the same published revision, ready to open without changing this result’s evidence.
Switch model pair
Choose from this result’s current and reviewed related models. Switching opens a reviewed comparison; it does not change this result’s evidence.