CodingModel Announcement

Anthropic Launches Claude Sonnet 5.5: 70.6% Terminal Coding at 30% Faster Speeds

Anthropic has released Claude Sonnet 5.5, the second model in the Claude 5.5 family. Achieving a record 70.6% on Terminal-Bench 4.0, Sonnet 5.5 cuts per-task costs by up to 30% and accelerates output generation speeds by more than 30%.

4 min read · By Newsroom Admin

Anthropic Claude Sonnet 5.5 editorial composition highlighting 70.6 percent Terminal-Bench 4.0 score, 30 percent faster output, and up to 30 percent lower cost.

What’s New

  • Scores 70.6% on Terminal-Bench 4.0, outperforming Sonnet 5's 10.3% and rivaling Opus 5.5 in agentic command-line execution.
  • Maintains base pricing of $2 per million input and $10 per million output tokens, but cuts overall task costs by up to 30% through token efficiency.
  • Generates output more than 30% faster than Sonnet 5, making it the fastest model in the Sonnet series to date.
  • Achieves 1844 Elo on GDPval-AA knowledge work, approaching Opus 5.5 while significantly surpassing GPT-6 Sol.
  • Launches with expanded preserved thinking safeguards and visible fallbacks to prevent reasoning distillation and cyber misuse.

Why It Matters

Sonnet 5.5 solves the trade-off between frontier capability and execution speed for day-to-day developer tasks. If your automated workflows need fast, low-cost iterations without sacrificing deep multi-file code understanding, Sonnet 5.5 is Anthropic's most pragmatic release yet.

Anthropic has officially launched Claude Sonnet 5.5, the second entry in its Claude 5.5 model family following the release of Claude Opus 5.5. Positioned as a high-speed, cost-efficient workhorse, Sonnet 5.5 is designed to handle well-scoped engineering tasks, continuous automated bug fixing, and long-horizon document creation while operating more than 30% faster than Claude Sonnet 5.

While Opus 5.5 serves as Anthropic's flagship for open-ended research requiring sustained human-level judgment, Sonnet 5.5 targets the daily unit economics of enterprise developers. Anthropic confirmed that Claude Haiku 5.5 will join the lineup in the coming weeks to cater to high-volume, cost-critical pipelines.

Breakthrough agentic coding benchmarks

The most pronounced leap in Sonnet 5.5 appears in autonomous coding benchmarks:

  • Terminal-Bench 4.0: Reaches 70.6% in terminal-based agentic workflows, a massive jump from Sonnet 5's 10.3% and exceeding Claude Opus 5.5 at 66.4%.
  • CursorBench 4.0: Scores 55.5% on multi-file ambiguous coding tasks sourced from live Cursor environments, trailing Opus 5.5 (57.8%) by just over two points while beating Sonnet 5 (34.1%).
  • FrontierCode v1.1: Scores 52.1% at extra-high effort and 46.2% at max effort, outperforming Sonnet 5 (42.4%) and rivaling GPT-6 Sol (49.3%).
  • Computer use (OSWorld 2.1): Scores 80.1% partial task success, nearing Opus 5.5 (81.8%) and easily exceeding Sonnet 5 (57.0%). It is also the first Sonnet model capable of completing Pokémon Red purely from raw screenshots.

In production trials, early access partners reported substantial operational efficiency. Gaming giant Epic Games noted that Sonnet 5.5 managed tens of thousands of lines of code for gameplay system architecture during data flow reviews, maintaining fast response times without requiring prescriptive prompts. CodeRabbit highlighted that the model eliminates redundant web search calls and trims unnecessary output tokens during automated pull request reviews. In app-building benchmarks conducted by Base44 across 118 applications, Sonnet 5.5 completed builds in an average of 3.6 iterations versus 7.7 for Opus 5, experiencing fewer stalled tool invocations.

Pricing structure and token efficiency

Anthropic kept headline token pricing unchanged from Sonnet 5:

  • Input tokens: $2.00 per million tokens.
  • Output tokens: $10.00 per million tokens.
  • Prompt caching reads: $0.20 per million tokens.

Despite identical list pricing, the model requires significantly fewer reasoning turns and output tokens to complete tasks, yielding total task cost savings of up to 30%. On benchmark evaluations like AA-Briefcase, running Sonnet 5.5 at low or medium effort outperforms Sonnet 5's top score at roughly one-tenth the cost per task.

Knowledge work and design capabilities

Beyond raw software development, Sonnet 5.5 scores 1844 Elo on Artificial Analysis's GDPval-AA v2.1 (evaluating 44 occupations), falling only two points behind Opus 5.5 (1846 Elo) and outperforming Sonnet 5 (1449 Elo) and GPT-6 Sol (1487 Elo). On the AA-Briefcase v1.1 benchmark for long-horizon professional work, it achieved 1811 Elo.

Enterprise testers also highlighted design refinements, noting that Sonnet 5.5 generates polished UI components and presentation decks from unstructured financial earnings transcripts. Slack reported that internal Slackbot offline evaluations improved across nearly all tasks with 14% fewer output tokens.

Containment and anti-distillation safeguards

On Anthropic's automated behavioral audit across 1,850 scenarios, Sonnet 5.5 matched or exceeded Sonnet 5 in honesty and misuse resistance. In containment testing, it proved to be Anthropic's least likely model to probe container boundaries or attempt sandbox escapes.

Because Sonnet 5.5 demonstrates cybersecurity capabilities comparable to Opus 5, Anthropic introduced tiered safeguards:

  1. Cybersecurity fallbacks: High-risk vulnerability discovery tasks automatically fall back to Sonnet 5, with full capabilities reserved for vetted practitioners in Anthropic's Cyber Verification Program.
  2. Biology protections: Sensitive biological requests remain guarded, with comprehensive research access available through the Life Sciences Verification Program.
  3. Preserved thinking: Sonnet 5.5 incorporates safety classifiers against industrial distillation attacks, binding internal reasoning chains to verified developer accounts.

Availability and rollout

Claude Sonnet 5.5 is available immediately across Claude.ai, Claude Code, and the Anthropic API under the identifier claude-sonnet-5-5. The model is also rolling out to Amazon Web Services Bedrock, Google Cloud Vertex AI, and Microsoft Azure Foundry, with zero data retention available on all supported tiers.

More in Coding