TokenMix Research Lab · 2026-07-14

Claude Opus 5 Review 2026: $5/$25, 61 Score, Real API Catch
Last Updated: 2026-07-25 Author: TokenMix Research Lab Data verified: 2026-07-25 - Anthropic launch post, Claude Platform model and pricing docs, Opus 5 migration notes, system card, Artificial Analysis measurements, Axios reporting, and the live TokenMix model catalog
Claude Opus 5 launched on July 24 at $5 per million input tokens and $25 per million output tokens. It adds a 1M-token context window, 128K maximum output, five effort levels, and thinking by default. The upgrade is substantial, but migration is not a model-ID-only change.
Anthropic positions Opus 5 as its everyday premium model for complex coding and enterprise work, below the $10/$50 Fable 5 tier. The official launch says Opus 5 more than doubles Opus 4.8 on Frontier-Bench at a lower cost per task, while independent testing places max effort at 61 on the Artificial Analysis Intelligence Index versus 56 for Opus 4.8. The catch is operational: max effort nearly doubles the independent evaluation spend versus high effort, and disabling thinking at xhigh or max now returns HTTP 400. These claims are kept separate as Confirmed specs, Confirmed independent measurements, and Confirmed vendor claims throughout this review (Anthropic launch, Claude Platform migration notes, Artificial Analysis).
Table of Contents
- Quick Verdict
- Release, Model ID, and Core Specifications
- Pricing: Standard, Cache, Batch, Fast, and US-Only
- Benchmarks: Independent Results vs Vendor Claims
- Effort Levels: The Real Cost Lever
- Opus 5 vs Opus 4.8 vs Fable 5
- Cost per Workload
- API Changes and the HTTP 400 Catch
- Availability: Anthropic, Clouds, and TokenMix
- Migration and Use Case Matrix
- Where Opus 5 Loses
- Final Recommendation
- FAQ
Quick Verdict
Opus 5 is a real generational upgrade, but high effort is the default production starting point; max effort is an escalation tier, not a sensible universal setting.
| Claim | Status | Source |
|---|---|---|
| Claude Opus 5 launched on July 24, 2026 | Confirmed | Anthropic launch |
The API model ID is claude-opus-5 |
Confirmed | Claude Platform docs |
| Standard API pricing is $5 input and $25 output per MTok | Confirmed | Anthropic pricing |
| Context is 1M tokens and maximum output is 128K | Confirmed | Claude Platform docs |
| Thinking is enabled by default | Confirmed | Migration notes |
| Max effort scored 61 on the AA Intelligence Index | Confirmed independent measurement | Artificial Analysis |
| Opus 5 more than doubles Opus 4.8 on Frontier-Bench | Confirmed vendor claim | Anthropic launch |
| Opus 5 is always cheaper per task than Opus 4.8 | False as a universal claim | Token use and effort determine task cost |
Disabling thinking at xhigh or max returns HTTP 400 |
Confirmed | Migration notes |
| Opus 5 is currently listed in TokenMix's public catalog | False as of 2026-07-25 | Live TokenMix catalog |
| Opus 5 should replace Fable 5 for every workload | False | Fable remains Anthropic's highest-capability tier |
Release, Model ID, and Core Specifications
Claude Opus 5 is generally available with a 1M context window and 128K maximum output, not a limited preview.
| Field | Claude Opus 5 | Status |
|---|---|---|
| Release date | 2026-07-24 | Confirmed |
| Claude API ID | claude-opus-5 |
Confirmed |
| Context window | 1,000,000 tokens | Confirmed |
| Maximum output | 128,000 tokens | Confirmed |
| Input modalities | Text and image | Confirmed |
| Output modality | Text | Confirmed |
| Thinking default | On | Confirmed |
| Effort levels | low, medium, high, xhigh, max |
Confirmed |
| Default effort | high |
Confirmed |
| Prompt cache minimum | 512 tokens | Confirmed |
| Model weights | Closed | Confirmed |
The model ID is unusually simple: claude-opus-5. Anthropic says 1M tokens is both the default and maximum context size, so there is no smaller context variant to select. The model supports text and image input but produces text output; parameter count and training details remain undisclosed (model overview).
The 128K output ceiling is meaningful for long reports, code generation, and multi-agent traces, but it is not a target. Because thinking tokens and visible response tokens share max_tokens, a configuration copied from Opus 4.8 can terminate early after thinking becomes active by default.
Pricing: Standard, Cache, Batch, Fast, and US-Only
The sticker price did not change from Opus 4.8, but the cheapest valid rate ranges from $0.50 cache hits to $50 fast-mode output per MTok.
| Billing mode | Input / MTok | Output / MTok | Notes | Status |
|---|---|---|---|---|
| Standard | $5.00 | $25.00 | Same as Opus 4.8 | Confirmed |
| 5-minute cache write | $6.25 | $25.00 | 1.25x input write | Confirmed |
| 1-hour cache write | $10.00 | $25.00 | 2x input write | Confirmed |
| Cache hit | $0.50 | $25.00 | 90% input discount | Confirmed |
| Batch API | $2.50 | $12.50 | 50% discount, asynchronous | Confirmed |
| Fast mode | $10.00 | $50.00 | Claude API only, research preview | Confirmed |
| US-only inference | $5.50 | $27.50 | 1.1x multiplier | Confirmed calculation |
The official Claude API pricing table confirms that cache, batch, fast mode, and geography modifiers can materially change the bill. Fast mode cannot be combined with Batch API. Prompt-cache and US-only multipliers can stack with supported modes.
Caching has a lower entry threshold on Opus 5: 512 prompt tokens versus 1,024 on Opus 4.8. A five-minute cache write pays back after one cache read because the first write costs 1.25x base input and each hit costs 0.1x. A one-hour write pays back after two reads.
Benchmarks: Independent Results vs Vendor Claims
Independent testing puts Opus 5 max at 61, but Anthropic's benchmark claims should remain labeled as vendor results until reproduced.
| Evaluation | Opus 5 result | Comparison | Evidence class |
|---|---|---|---|
| AA Intelligence Index, max | 61 | Fable 5: 60; GPT-5.6 Sol: 59; Opus 4.8: 56 | Confirmed independent measurement |
| GDPval-AA v2 | 1861 Elo | More than 100 Elo ahead of Fable 5 and GPT-5.6 Sol | Confirmed independent measurement |
| AA-Briefcase | 1720 Elo | 146 Elo ahead of Fable 5 | Confirmed independent measurement |
| Frontier-Bench v0.1 | More than 2x Opus 4.8 | Lower cost per task | Confirmed vendor claim |
| CursorBench 3.2, max | Within 0.5% of Fable 5 peak | About half Fable's cost per task | Confirmed vendor claim |
| ARC-AGI 3 | 3x next-best model | Novel problem solving | Confirmed vendor claim |
| AutomationBench | About 1.5x next-best pass rate | Same cost per task | Confirmed vendor claim |
| OSWorld 2.0 | Above Fable 5's best result | Just over one-third of Fable's cost | Confirmed vendor claim |
Artificial Analysis evaluated Opus 5 before release and reports a 61 Intelligence Index score at max effort. Its Opus 5 analysis also reports leadership on GDPval-AA v2 and AA-Briefcase. Those are independent measurements, although they still represent one evaluator and one harness.
Anthropic's launch charts cover Frontier-Bench, CursorBench, ARC-AGI 3, AutomationBench, and OSWorld 2.0. The relative gains are useful, but the provider selected the tasks, prompts, effort levels, and cost framing. Treat them as Confirmed vendor claims, not independent proof that every production workload improves by the same amount.
The honest negative result is speed. Artificial Analysis measured roughly 53 output tokens per second and a long time to first token at max effort. Opus 5 is a reasoning-heavy model; it is not the obvious choice for low-latency chat or high-volume classification.
Effort Levels: The Real Cost Lever
Moving from high to max added 2 Intelligence Index points while the reported evaluation spend rose about 94%.
| Effort | AA Intelligence Index | Reported evaluation cost | Output tokens in suite | Interpretation |
|---|---|---|---|---|
| Low | 51 | Not reported in cited snapshot | Not reported | Cheap routing candidate; test carefully |
| Medium | 56 | $1,114.96 | 29M | Strong value for routine difficult work |
| High | 59 | $1,973.77 | 52M | Default production starting point |
| Xhigh | 60 | $2,909.91 | 76M | Expensive one-point gain over high |
| Max | 61 | $3,835.51 | 100M | Deepest reasoning, highest cost and latency |
These totals are Artificial Analysis evaluation-run costs, not a forecast of your monthly invoice. They are still useful because the same evaluator ran the same suite across effort levels. Based on the published totals, max cost about 94% more than high for a two-point gain, while xhigh cost about 47% more than high for a one-point gain (medium, high, xhigh, max).
This makes effort a routing decision. Start at high, as Anthropic recommends. Step down to medium where your acceptance rate holds. Reserve xhigh and max for failed high-effort tasks, high-value research, difficult debugging, or work where one correct answer is worth more than latency.
Opus 5 vs Opus 4.8 vs Fable 5
Opus 5 replaces Opus 4.8 on capability at the same price, but it does not eliminate Fable 5's highest-capability role.
| Field | Opus 4.8 | Opus 5 | Fable 5 |
|---|---|---|---|
| Standard input/output | $5 / $25 | $5 / $25 | $10 / $50 |
| Context | 1M | 1M | 1M |
| Maximum output | 128K | 128K | 128K |
| Thinking default | Off unless enabled | On | On |
| Full effort ladder | More limited behavior | Low to max | Low to max |
| AA Intelligence Index, max | 56 | 61 | 60 |
| Best fit | Stable legacy Opus route | Daily premium agents | Highest-capability long-running agents |
| Migration risk | Baseline | Default-thinking behavior change | Higher cost and stricter workload economics |
For a new premium agent, Opus 5 is the better default than Opus 4.8 because the price is unchanged and the measured capability is higher. Existing Opus 4.8 users should still canary the migration because default thinking, longer outputs, more progress narration, and more frequent subagent delegation can change cost and integration behavior.
Fable 5 remains a separate tier. Anthropic calls it the model for the highest available capability, while Opus 5 is the recommended starting point for complex agentic coding and enterprise work. Our Fable 5 review explains why its $10/$50 rate still needs a success-rate advantage to justify production routing.
Cost per Workload
A 100K-input, 20K-output agent run costs $1.00 standard, $0.55 with a cache hit, $0.50 in batch, or $2.00 in fast mode.
| Workload | Token shape | Standard | Optimized route | Optimized cost |
|---|---|---|---|---|
| Repository coding run | 100K input + 20K output | $1.00 | 100K cache hit + standard output | $0.55 |
| Overnight evaluation | 100K input + 20K output | $1.00 | Batch API | $0.50 |
| Interactive incident response | 100K input + 20K output | $1.00 | Fast mode | $2.00 |
| US-only agent run | 100K input + 20K output | $1.00 | inference_geo: "us" |
$1.10 |
| Full-context report | 1M input + 64K output | $6.60 | Cache hit after first run | $2.10 |
| Maximum-output request | 1M input + 128K output | $8.20 | Batch API | $4.10 |
The repository example is simple:
Standard: (0.10 x $5) + (0.02 x $25) = $1.00 per run
Cache hit: (0.10 x $0.50) + (0.02 x $25) = $0.55 per run
Batch: $1.00 x 50% = $0.50 per run
Fast: (0.10 x $10) + (0.02 x $50) = $2.00 per run
At 1,000 identical runs per month, those routes cost $1,000 standard, $550 with stable cache hits, $500 in batch, or $2,000 in fast mode. Fast mode only makes sense when the time saved is worth more than the extra $1,000. The Claude API cost calculator can be used to substitute your own input, output, cache-hit, and call-volume assumptions.
These calculations do not predict effort-level token use. Opus 5 can produce longer default responses and use more thinking tokens than a non-thinking Opus 4.8 request. Log actual input, cache, thinking, and output usage before projecting a monthly bill.
API Changes and the HTTP 400 Catch
Changing the model ID is easy; preserving Opus 4.8 behavior is not, because thinking now defaults on and some disabled-thinking requests fail.
| Change | Opus 4.8 behavior | Opus 5 behavior | Migration action |
|---|---|---|---|
| Model ID | claude-opus-4-8 |
claude-opus-5 |
Change explicit ID |
| Thinking | Off unless adaptive thinking is set | On by default | Revisit max_tokens and budgets |
| Disable thinking at high or below | Accepted | Accepted | Keep effort at high or lower |
| Disable thinking at xhigh/max | Previously independent | HTTP 400 | Remove disabled thinking or lower effort |
| Tool list between turns | Fixed in ordinary flow | Can change with beta header | Test cache preservation |
| Server-side fallback | Explicit list | New default beta mode |
Audit fallback identity |
| Cache minimum | 1,024 tokens | 512 tokens | Shorter prompts can now cache |
| Verification prompts | Often useful | Can cause over-verification | Remove redundant instructions |
The minimum migration is:
import anthropic
client = anthropic.Anthropic()
response = client.messages.create(
model="claude-opus-5",
max_tokens=64000,
output_config={"effort": "high"},
messages=[
{"role": "user", "content": "Review this repository change and verify the tests."}
],
)
Thinking is already on, so no thinking field is required. If you send thinking: {"type": "disabled"} with effort: "xhigh" or "max", the API returns HTTP 400. Anthropic also warns that disabling thinking can occasionally make the model print a tool call as text or expose internal XML tags, so lower effort is the cleaner cost control (Opus 5 migration notes).
Two beta features require dated headers. Mid-conversation tool changes use mid-conversation-tool-changes-2026-07-01; the new default fallback mode uses server-side-fallback-2026-07-01. Beta headers are API contracts, not decorative flags. Pin them, test them, and log the final model that handled each request.
Availability: Anthropic, Clouds, and TokenMix
Opus 5 is available from Anthropic and three major cloud channels, but TokenMix had not listed it at the July 25 verification time.
| Channel | Model identifier | Availability | Status |
|---|---|---|---|
| Claude API | claude-opus-5 |
All customers | Confirmed |
| Amazon Bedrock | anthropic.claude-opus-5 |
Available | Confirmed |
| Google Cloud | claude-opus-5 |
Available | Confirmed |
| Microsoft Foundry | Platform deployment | Available | Confirmed |
| Claude paid plans | Strongest model on Pro; default on Max | Available | Confirmed |
| TokenMix | No public Opus 5 row as of 2026-07-25 | Not yet listed | Confirmed catalog observation |
The official model overview lists Claude API, AWS, Google Cloud, and Microsoft Foundry availability. Fast mode is narrower: it is currently limited to the first-party Claude API and is unavailable on Bedrock, Google Cloud, or Microsoft Foundry.
TokenMix's live public catalog still showed Opus 4.8, Opus 4.7, Sonnet 5, and Fable 5 when checked on July 25, but no Opus 5 entry. Do not use an invented TokenMix route or assume silent remapping. Check the model catalog before deployment; this article will be updated when a verified route appears.
Migration and Use Case Matrix
Migrate premium agent workloads through a canary, while keeping Sonnet 5 for volume and Fable 5 for the hardest unresolved tasks.
| Use case | Recommended model/effort | Why | Confidence |
|---|---|---|---|
| Complex coding agent | Opus 5 high | Best balance of measured capability and effort cost | Confirmed recommendation |
| Hard debugging escalation | Opus 5 xhigh, then max | Higher effort can improve difficult tasks | Likely; validate locally |
| High-volume support classification | Sonnet 5 or cheaper tier | Opus price and latency are unnecessary | Confirmed economic judgment |
| Long legal or financial analysis | Opus 5 high with cache | Strong vendor and customer evidence; cache repeated context | Likely |
| Maximum-capability autonomous research | Fable 5 canary vs Opus 5 max | Fable remains top product tier | Confirmed positioning |
| Overnight batch evaluation | Opus 5 Batch API | 50% token discount | Confirmed |
| User-facing urgent response | Opus 5 fast only with SLA value | 2x token price | Confirmed pricing judgment |
| Compliance-sensitive US workload | Opus 5 US-only | 1.1x cost for US inference | Confirmed |
Use a 100-300 task canary from your own production distribution. Record pass rate, retries, human corrections, tool-call errors, output tokens, total latency, and cost per accepted result. A five-point composite benchmark gain is valuable only if it reduces failures or labor in your workflow.
The Sonnet 5 review remains relevant because Sonnet's introductory $2/$10 price is 60% below Opus 5 through August 31. A router that tries Sonnet first and escalates failures to Opus can beat an Opus-only stack on total cost.
Where Opus 5 Loses
Opus 5 loses on latency, low-value volume, and workloads where max effort adds token spend without improving acceptance.
| Weakness | Evidence | Better choice |
|---|---|---|
| Slow interactive response | Independent testing reports long reasoning latency | Sonnet 5 or a faster small model |
| Expensive output | $25/MTok standard; $50/MTok fast | Batch, lower tier, or tighter output limits |
| Max-effort diminishing returns | 59 to 61 AA points from high to max; evaluation cost +94% | High effort by default |
| More verbose default behavior | Anthropic documents longer deliverables and progress narration | Prompt for concise output; cap tokens |
| Migration can break disabled-thinking calls | HTTP 400 at xhigh/max | Keep thinking on or lower effort |
| Proprietary weights | No local deployment or weight audit | Open-weight model where required |
| Not yet on every gateway | TokenMix public catalog had no Opus 5 row on July 25 | Use a verified listed route |
The largest mistake is setting max globally because it won a chart. Independent measurements show a steep cost curve after high effort. Route by task value and failure history, not by model prestige.
Final Recommendation
Migrate Opus 4.8 premium workloads to an Opus 5 high-effort canary, not directly to max.
Opus 5 keeps the $5/$25 price while improving measured capability, context behavior, and agent work. The migration risk is real: thinking defaults on, output can grow, and disabled thinking at xhigh/max returns HTTP 400. Use medium or high for normal production, then escalate only failed high-value tasks.
FAQ
Is Claude Opus 5 officially released?
Yes. Anthropic released Claude Opus 5 on July 24, 2026. It is generally available through the Claude API and supported cloud channels, not a rumor or waitlist-only model.
How much does Claude Opus 5 cost?
Standard pricing is $5 per million input tokens and $25 per million output tokens. Cache hits cost $0.50 per million input tokens, Batch API gives a 50% discount, and fast mode costs $10 input and $50 output per million.
What are the Claude Opus 5 context and output limits?
Claude Opus 5 has a 1M-token context window and a 128K-token maximum output. One million tokens is both the default and maximum context size; Anthropic does not offer a smaller Opus 5 context variant.
What is the Claude Opus 5 API model ID?
The first-party Claude API model ID is claude-opus-5. Amazon Bedrock uses anthropic.claude-opus-5, while Google Cloud uses claude-opus-5.
Is Claude Opus 5 better than Opus 4.8?
Yes on the published evidence, but the production gain is workload-specific. Artificial Analysis scored max effort at 61 versus 56 for Opus 4.8, while Anthropic reports larger gains on agentic coding and knowledge work. Both models retain the same $5/$25 sticker price.
Is Opus 5 cheaper than Fable 5?
Yes per token. Opus 5 costs $5/$25, exactly half Fable 5's $10/$50 standard rate. Cost per successful task can differ, so difficult long-running agents still need a direct canary comparison.
Why does Opus 5 return HTTP 400 when thinking is disabled?
Opus 5 accepts disabled thinking only at high effort or below. Combining thinking: {"type": "disabled"} with xhigh or max returns HTTP 400. Remove the disabled-thinking field or lower the effort level.
Should every Opus 5 request use max effort?
No. Artificial Analysis measured only a two-point Intelligence Index gain from high to max while its evaluation spend rose about 94%. Start at high, test medium for routine work, and reserve max for expensive failures or unusually valuable tasks.
Can I use Claude Opus 5 through TokenMix today?
Not yet according to the public catalog checked on July 25, 2026. TokenMix listed Opus 4.8, Opus 4.7, Sonnet 5, and Fable 5, but no Opus 5 route. Verify the live model catalog before sending production traffic.
About TokenMix
TokenMix.ai is an AI API relay that routes Claude, OpenAI, Gemini, DeepSeek, Qwen, and other large language models through a single OpenAI-compatible endpoint at https://api.tokenmix.ai/v1. Current model availability and per-token rates are listed on the pricing page and the model catalog. Integration uses the standard OpenAI SDK; details are in the OpenAI compatibility reference.
Sources
- Anthropic: Introducing Claude Opus 5
- Claude Platform: What's New in Claude Opus 5
- Claude Platform: Models Overview
- Claude Platform: API Pricing
- Anthropic: Claude Opus 5 System Card
- Artificial Analysis: Opus 5 Analysis
- Artificial Analysis: Opus 5 Max
- Artificial Analysis: Opus 5 Medium
- Artificial Analysis: Opus 5 High
- Artificial Analysis: Opus 5 Xhigh
- Axios: Anthropic Releases Opus 5
- TokenMix Public Model Catalog