← 返回 Siami 首頁

Anthropic 發表 Claude Opus 5.5:cache reads 砍 60%、Terminal-Bench 4.0 大贏所有對手

▲ 969 💬 703
Anthropic 發表 Claude Opus 5.5:cache reads 砍 60%、Terminal-Bench 4.0 大贏所有對手

Anthropic 於 2026 年 9 月 22 日正式發表 Claude Opus 5.5,這是新一代 Claude 5.5 家族的第一位成員。Anthropic 把 Opus 5.5 定位為:在大多數真實工作場景中,表現逼近 Claude Fable 5.1,但運行成本比 Opus 5 低了 40%。

定價大幅下修(比 Opus 5 更便宜)

  • Input: $4 / M tokens(Opus 5 為 $5 / M tokens)
  • Output: $20 / M tokens(Opus 5 為 $25 / M tokens)
  • Cache reads: $0.20 / M tokens(Opus 5 為 $0.50 / M tokens,砍 60%)
  • Cache writes: $5 / M tokens(Opus 5 為 $6.25 / M tokens)

Opus 5.5 也比 Opus 5 快 30% 以上。Fast mode(在 Claude Code / Claude Platform 上提供)速度可達 2.5 倍,但價格為 $8 input / $40 output per M tokens。


為什麼這件事重要

Opus 5.5 是 Anthropic 喊出「節制前沿發展(pacing the frontier)」之後的第一個新模型。Anthropic 過去一年多次呼籲 AI 產業「放慢腳步、加強安全評估」,但同時市場又逼著他們趕快出新版——這個「既喊煞車又踩油門」的矛盾,在 Opus 5.5 發表會上特別明顯。更值得注意的是價格策略:Opus 5.5 把 cache reads 直接砍 60%($0.50 → $0.20),而 cache reads 是 agentic coding 工作流的主要成本。Anthropic 等於是公開承認:「我們要在 Claude Code 這個 agent 戰場上,跟 Cursor / Cognition / Devin 直接打價格戰」。對開發者最實質的意義是:用 Opus 5 跑大型 codebase migration 的成本,現在用 Opus 5.5 只要大約 60%,而且速度還更快。Anthropic 也順勢把 Pro / Max / Team / Enterprise 方案的「五小時用量上限」調高——這是過去 Claude 用戶抱怨最久的痛點。


實際 benchmark 數字(Anthropic 自家公佈)

  1. Terminal-Bench 4.0(agentic coding):Opus 5.5 拿下 66.4%,勝過 Fable 5.1 的 55.8%、GPT-6 Astra 的 57.9%、Opus 5 的 52.3%、GPT-5.6 Sol 的 37.3%。
  2. FrontierCode v1.1(Main):Opus 5.5 為 54.4%,領先 Fable 5.1 的 50.3%、GPT-6 Astra 的 53.3%。
  3. CursorBench 4.0:Opus 5.5 為 57.8%,Fable 5.1 為 51.8%、GPT-5.6 Sol 為 41.7%。
  4. GDPval-AA v2.1(知識工作 Elo):Opus 5.5 為 1846,領先 Fable 5.1 的 1735、Opus 5 的 1708、GPT-5.6 Sol 的 1588、GPT-6 Astra 的 1542。
  5. AutomationBench(商業流程):GPT-6 Astra 以 41.4% 略勝 Opus 5.5 的 40.0%,但 Fable 5.1 只有 31.4%。
  6. Humanity’s Last Exam(多領域推理):Opus 5.5 為 67.7% with tools,Fable 5.1 為 65.6%,GPT-6 Astra 為 57.2%。
  7. Terminal-Bench-Science 0.1(科學研究):GPT-6 Astra 以 64.6% 領先 Opus 5.5 的 58.7%,Fable 5.1 為 52.6%。
  8. OSWorld 2.0(電腦操作):Opus 5.5 達 81.8% partial,Fable 5.1 為 80.7%,Opus 5 為 74.0%。
  9. Chartography(圖表辨識):Opus 5.5 達 89.0% with tools,Fable 5.1 為 88.4%,Opus 5 為 83.4%。從數字上看,Opus 5.5 在 agentic coding 三項都拿下第一,知識工作遙遙領先(Elo 1846 vs 第二名 1735),電腦操作跟圖表辨識也小贏 Fable 5.1。但在科學研究類別,GPT-6 Astra 仍然以 64.6% 略勝一籌(58.7%)。

數據解讀:Anthropic 自己也在踩煞車

重點提醒:Anthropic 自己在官方頁面承認「benchmark margins 已經不是真實差異的可靠指標」。Opus 5.5 跟 Fable 5.1 的實際體感差距,比分數顯示的還要小。這句話翻成白話:Anthropic 知道他們新模型跟上一代旗艦的「帳面差距」很有限,但因為 cache reads 砍 60%,加上運算成本降低,整體 CP 值仍然遠勝舊版。值得質疑的兩點:

  1. Terminal-Bench 4.0 的 66.4%:Anthropic 是用 Opus 5.5 開到 xhigh effort 才跑到這個分數;GPT-6 Astra 是開 high effort。換言之這個差距是「拿最強模式 vs 對手次強模式」比出來的,直接對比有點不公平。
  2. AutomationBench 的 40.0%:這項是 Zapier 跑的,而且沒有 fallback model——意思是當 safety safeguard 介入把任務擋下來時,該題直接算「失敗」。Anthropic 自己也承認這個分數比 Opus 5.5 實際能力低。換句話說,這些數字漂亮,但有兩個「修飾過」的痕跡。對開發者來說,真正的考驗還是要看實際 agentic 跑專案時的表現。

早期測試者的真實案例

Anthropic 公開了幾個早期測試的具體成果:

  • 一位測試者用 Opus 5.5 完成 68 萬行程式碼的遷移,時間不到一天——「原本一個工程團隊要花好幾週的工作量」。
  • 另一個任務:要求 Opus 5.5 把整個 web app 的每一頁載入時間砍掉,40 次中成功 39 次;Opus 5 雖然也有改善,但常常「為了修效能而改變了 app 行為」,實務上不能用。
  • 還有人用 Opus 5.5 來 audit 跟修一個 20 萬行的 codebase,3 小時內完成;同樣工作 Opus 5 要花超過 20 小時,而且用了 2.5 倍的 token。
  • 內部測試:要求 Opus 5.5 把 HAProxy(廣泛使用的網頁負載平衡軟體)從 C 改寫成 Rust——兩次改寫幾乎都通過 HAProxy 自己的回歸測試。

安全強化(Anthropic 強調的重點)

Opus 5.5 是 Anthropic 至今在「自動化行為稽核(automated behavioral audit)」分數最高的模型——這套對齊測試涵蓋數千個模擬情境。具體強化:

  • 比近期模型更不容易做出「難以逆轉」的動作
  • 比 Opus 5 更抵抗 prompt injection 攻擊
  • 因為 Opus 5.5 在生物跟資安能力上跟 Claude Mythos 5.1 相當,所以 Anthropic 啟用跟 Fable 5.1 同等級的 safeguards
  • 生命科學驗證計畫(Life Sciences Verification Program):受審查的機構現在就可以申請使用 Opus 5.5 做生物研究
  • 未來幾週會擴大資安驗證計畫(Cyber Verification Program)的範圍
  • 發表前由外部機構 Frontier Design 跟 METR 獨立測試

「寫作風格」改善(Anthropic 罕見承認前代痛點)

Opus 5 推出時被不少用戶抱怨「語氣像 AI 寫的、太機械、太囉嗦」。Opus 5.5 的官方說法是「寫得更自然、更清楚,把重點放前面」,甚至引用一位早期測試者的話:「it writes the way I do(它寫得像我寫的一樣)」。對長期使用 Claude 的用戶,這是 Anthropic 第一次公開承認「寫作風格」是前代需要改善的重點。


怎麼用、哪裡能用

Opus 5.5 發表當天就在以下平台同步上線:

  • Claude.ai(網頁 + app)
  • Claude Code(CLI / IDE 整合)
  • Anthropic API
  • Amazon Bedrock
  • Google Cloud Vertex AI
  • Vercel AI Gateway

訂閱方案(Pro / Max / Team / Enterprise)的五小時用量上限同步提高,所有訂閱用戶還會拿到一個「rate limit reset」額度,可以自己選時間再用——這是過去用戶抱怨「額度一週一次、不能累積」之後的改善。至於 Claude Sonnet 5.5 跟 Claude Haiku 5.5,官方說「未來幾週」會跟上,會繼承同樣的效能與安全改進。


結語:Anthropic 的「節制」其實是另一步進攻

Anthropic 在 Opus 5.5 發表頁的第一句話就寫:「Claude Opus 5.5 is our first release since we called for pacing the frontier.」——他們把「節制前沿發展」跟「推出更便宜的旗艦模型」放在同一個框架裡。但實際內容:cache reads 砍 60%、輸出快 30%、在 Terminal-Bench 4.0 大贏所有對手。這不是「節制」,是「用更低成本打更深的 agentic coding 市場」。對 OpenAI 跟 Google 來說,Opus 5.5 的發表是 2026 年第三季最重要的壓力測試——GPT-6 Sol/Astra 跟 Gemini 3 在 agentic coding 跟商業流程上的領先地位,正在被 Opus 5.5 一項一項追上。

編按:本文綜合整理自 Anthropic 官方發表頁、VentureBeat、The Decoder、The Verge、LLM Stats,並加入 Siami 編輯部觀點與分析。

網友熱門留言 (5)

#1 Hacker News 用戶 @km144 ▲ 412
Interesting how the very first line is used to remind the reader that Anthropic is 'pacing the frontier'. The rest of the post is about how their new model is the best at everything and 40% cheaper.
#2 Hacker News 用戶 ▲ 287
Cache reads going from $0.50 to $0.20 is genuinely huge for agentic workflows. That's where most of the cost goes.
#3 Hacker News 用戶 ▲ 198
The benchmarks are cherry-picked at xhigh effort. They should be reporting the same effort level as competitors for fair comparison.
#4 Hacker News 用戶 ▲ 156
Anthropic's communication improvement is real. Opus 5 was painful to read after using it for a week. This update finally addresses that.
#5 Hacker News 用戶 ▲ 134
The 680,000-line migration in under a day is the most impressive claim. That's not a benchmark — that's real engineering work.