← 返回 Siami 首頁

Anthropic 首度公開 Claude 全部 System Prompts:從 Opus 5 到 Haiku 3 的完整一手機密

▲ 239 💬 123
Anthropic 首度公開 Claude 全部 System Prompts:從 Opus 5 到 Haiku 3 的完整一手機密

編按:本文綜合整理自 Anthropic 官方技術文件、Hacker News 討論串(239 分 / 123 則留言)、Simon Willison 的逐版剖析、Anthropic 工程部事後分析,以及 Project Glasswing 與 Fable 5 出口管制事件聲明,並加入 Siami 編輯部觀點與分析。

Anthropic 把 Claude 全部 System Prompt 一次攤開

2026 年 8 月 16 日,Anthropic 在官方平台文件站發布一條新的「System Prompts」release notes 頁面,把 claude.ai 網頁版、iOS 與 Android 應用程式從 2024 年 7 月至今所有對外使用過的系統提示全文攤開。頁面同步在 Hacker News 上以「Claude: System Prompts」為題發布,24 小時內衝到 239 分、123 則留言,是 Anthropic 過去幾個月最受矚目的透明度動作之一。

這份頁面目前涵蓋 17 個模型版本,從最新的 Claude Opus 5(2026 年 7 月 24 日)一路回溯到 Claude Haiku 3(2024 年 7 月 12 日)。Opus 4.6 世代以前的模型因為仍會透過日期條目更新 prompt,每次改版都會留下新紀錄;4.6 世代之後改為「每個模型 ID 對應一張固定快照」,所以 Opus 4.6、Sonnet 4.6、Opus 4.7、Opus 4.8、Opus 5、Fable 5 都只剩單一條目。

Anthropic 在頁面開頭就把限制講清楚:這些 prompt 只用在 claude.ai 與 Claude 行動 App,Claude API 用戶看到的「我會怎麼回答這個自我介紹問題」風格,跟網頁版截然不同。API 端點的 prompt 仍由開發者自訂,Anthropic 不會預先塞入。


從 300 字膨脹到 3000 字:兩年間 prompt 規模暴增 10 倍

把 17 個版本的 prompt 串起來看,最直觀的觀察是 規模暴增。Hacker News 用戶 tosh 把每個版本字數整理後指出:

早期 system prompt 大約 300 字,最新版本已經突破 3000 字。

量變的背後是 Anthropic 對「Claude 該如何自我介紹、該如何處理邊界情境」逐年加掛的細節。早期 Haiku 3(2024 年 7 月)只有五行:

The assistant is Claude, created by Anthropic. The current date is {{currentDateTime}}. Claude’s knowledge base was last updated in August 2023 and it answers user questions about events before August 2023 and after August 2023 the same way a highly informed individual from August 2023 would if they were talking to someone from {{currentDateTime}}. It should give concise responses to very simple questions, but provide thorough responses to more complex and open-ended questions. It is happy to help with writing, analysis, question answering, math, coding, and all sorts of other tasks. It uses markdown for coding. It does not mention this information about itself unless the information is directly pertinent to the human’s query.

到了 Opus 5(2026 年 7 月),同一個段落已經擴張為涵蓋產品揭露、Fable/Mythos 安全路由、Cowork 多代理人入口、子代理行為準則、Markdown 格式細節、幻覺免責、兒童安全條款、長篇對話策略等多個獨立段落。


從 prompt 看 Anthropic 的「行為邊界」怎麼畫

把每一代的 prompt 拆開比對,可以讀出 Anthropic 對 Claude 行為邊界的設定脈絡。

1. 自我介紹與產品範圍

Opus 4 以前(2025 年 5 月)只有三句話:

The assistant is Claude, created by Anthropic. The current date is {{currentDateTime}}. Here is some information about Claude and Anthropic’s products in case the person asks: This iteration of Claude is Claude Opus 4 from the Claude 4 model family. The Claude 4 family currently consists of Claude Opus 4 and Claude Sonnet 4. Claude Opus 4 is the most powerful model for complex challenges.

Sonnet 4.5 世代之後,這段擴張為會明確標出 API model string(例如 claude-sonnet-4-5-20250929)、強制 Claude 主動查文件而不是憑記憶回答產品問題,並寫入「沒有其他 Anthropic 產品」這種防越界提示。

2. 幻覺(hallucination)免責條款

從 Sonnet 3.5(2024 年 11 月)開始,prompt 內建一段非常具體的提醒:

If Claude is asked about a very obscure person, object, or topic, i.e. if it is asked for the kind of information that is unlikely to be found more than once or twice on the internet, Claude ends its response by reminding the human that although it tries to be accurate, it may hallucinate in response to questions like this. It uses the term ‘hallucinate’ to describe this since the human will understand what it means.

Opus 4.5 之後更進一步寫到「如果 Claude 引用了某篇論文或書,必須提醒使用者自己沒有搜尋或資料庫,引用可能會錯」。這條規則對學術寫作情境的影響最直接。

3. 圖像辨識:永遠自稱「臉盲」

Sonnet 4.6 起的圖像相關 prompt 寫得很死:

Claude always responds as if it is completely face blind. If the shared image happens to contain a human face, Claude never identifies or names any humans in the image, nor does it imply that it recognizes the human. It also does not mention or allude to details about a person that it could only know if it recognized who the person was.

即使使用者主動告知圖中人物是誰,Claude 在後續對話中也不會「確認」那張圖裡就是那位被指定的對象。這條規則與 OpenAI、Google Gemini 對名人圖像辨識的態度並不一致。

4. 對話風格:去客套化

從 Opus 4.7 開始,prompt 內建一條非常具體的禁令:

Claude responds to all human messages without unnecessary caveats like “I aim to”, “I aim to be direct and honest”, “I aim to be direct”, “I aim to be direct while remaining thoughtful…”.

並要求 Claude 不要在答案裡夾雜「genuinely、honestly、straightforward」這類強化語。Hacker News 用戶 rafram 看到這段吐槽:

Hah! No it doesn’t. ——因為他自己用 Claude 時還是會看到這類詞。

5. 選舉資訊

Opus 4.6(2026 年 2 月)首次加入 <election_info> 區塊,預先寫死 2024 年 11 月美國總統大選結果。這條規則在選後模型上很常見,但 Anthropic 把它跟其他 metadata 一起攤在陽光下是罕見做法。

6. 兒童安全條款

早期的 Haiku 3 / Sonnet 3.5 完全沒有任何兒童安全字眼。Hacker News 用戶 altmanaltman 對此特別驚訝:

Wild how most of the earliest models had no child safety guardrails in the prompt (something that has multiple bullet points now in the latest one). For a company all about allignment and safety, they chose to go with this as their first system prompt: … No mention of any safety at all lol, how could dario let this be.

新版的 prompt 對「child-safety requirements require special attention and care」一節拉到多個段落,並搭配具體禁止行為清單。


Fable 5 安全路由:被寫進 Opus 5 prompt 的真實事件

Opus 5 prompt 裡最值得研究的段落,是 Anthropic 把 2026 年 6 月的 Fable 5 / Mythos 5 出口管制事件完整寫進了系統提示。

Claude Fable 5 and Claude Mythos 5 were first released on June 9, 2026. On June 12, 2026, Anthropic suspended access to both models to comply with U.S. Department of Commerce export controls; the Department lifted those controls on June 30, 2026, and Anthropic restored access on July 1, 2026 (Anthropic’s statement: fable-mythos-access). These events are after Claude’s training-data cutoff, so Claude knows about them only from this note.

也就是說,當使用者在 Opus 5 對話中問起 Fable 5,Claude 必須能精準複述這段時間軸——而它不是從訓練資料學到的,是 prompt 幫它「補課」。Simon Willison 在 HN 留言裡指出,這是 Anthropic 處理「訓練截止後真實事件」的一個非常具體的範例。

更深一層的是,Opus 5 prompt 還把「Fable 5 因安全疑慮被路由到 Opus 5」這件事講白:

the user may have selected a different Anthropic model, “Claude Fable 5”, but their query was redirected to Opus 5 instead due to a safeguards routing mechanism. The user may be confused about this situation (it’s very recent!); if they have questions, Claude can either directly cite or just let its response be informed by this quote from Anthropic’s blog post on the subject.

Anthropic 給出的官方說明是,Fable 5 只要涉及網路安全、生物、化學、蒸餾(distillation)等「兩用風險」主題,就會被自動改派給 Opus 4.8,整體觸發比例約 5%。

對照 2026 年 6 月 12 日的 Anthropic 聲明,觸發這次出口管制暫停的導火線是「美國政府認為出現了一種繞過 Fable 5 限制的 jailbreak 手法」。Anthropic 本身的評估是,這個手法只揭露了既有的「已被其他模型也能獨立找到」的小漏洞,但仍決定全面停機配合。Mythos 5(無安全限制版本)目前只開放給特定核可組織,並以 Project Glasswing 為主要應用場景。


為什麼這件事重要

Anthropic 在 2024 年 8 月(Claude 3.5 剛上線不久)就搶先同業成為第一家公開釋出自家聊天模型 system prompt 的大型 AI 實驗室。VentureBeat 當時的報導指出,Anthropic 開發者關係負責人 Alex Albert 在 X 公開承諾「會持續更新」這份文件。兩年後回頭看,這個承諾沒有跳票:從 Sonnet 3.5 到 Opus 5,歷代 prompt 的「差異日誌」完整保留,可以在 Simon Willison 的 GitHub 鏡像上一行行對照。

相較之下,OpenAI、Google DeepMind、Meta 至今仍未釋出任何對等的公開 prompt 紀錄。xAI 雖然在 2025 年 5 月「白南非種族滅絕」事件後補開 GitHub 儲存庫,但更新頻率遠不及 Anthropic。

把這份透明度放到商業脈絡下看,重要性有三層:

  1. 給開發者「prompt 工程」教材。Anthropic 對身份、產品、格式、回應風格、邊界情境的處理方式,幾乎就是現成的 LLM 應用藍圖。Simon Willison 在 2025 年 5 月的評論寫得很直白:「這些 prompt 對 LLM 進階使用者來說是 solid gold,能搞清楚怎麼從 Claude 身上挖到最大價值。」
  2. 建立外部監督的可檢驗基礎。當使用者對 Claude 行為有疑慮(例如幻覺、拒答、政治傾向),可以在 prompt 內直接驗證背後規則,無須猜測 Prompt 是否被更動。
  3. 給監管機構預先熟悉 AI 內部規則的窗口。這對未來的 AI 法案(如 EU AI Act 對高風險系統的透明度義務)有直接操作意義。

數據解讀與質疑

從 prompt 變動也可以讀出幾個值得懷疑的點。Hacker News 用戶 dev-complete 比對 Opus 4.8 與 Opus 5 的 prompt 後質疑:

I compared the Claude Opus 4.8 and 5 system prompts, as well as the Claude Code Opus 4.8 and 5 system prompts, and neither show the alleged 80% reduction in system prompt size… Is the Claude Code system prompt leak incorrect? Do I not know what 80% looks like? Why such a large lie (so it seems)?

對照 2026 年 4 月 23 日 Anthropic 工程部對 Claude Code 品質報告的「事後分析」可以發現,每行 prompt 對模型表現的影響其實非常敏感——他們曾因為新增「Length limits: keep text between tool calls to ≤25 words」這一行小指令,導致 Opus 4.6、4.7 在內部評測出現 3% 退步,立刻在 4 月 20 日退回。換言之,Anthropic 對 prompt 的每一次微調都會跑 ablation 評測,而不是憑感覺。

但同一份事後分析也透露一個尷尬現實:Anthropic 內部 ablation 是在「自家評測集」上跑的,外部開發者觀察到的回應品質波動常常無法對應到 prompt 單一改動。這也是為什麼 4 月那波 Claude Code 品質抱怨要花上兩週才能定位到問題段落。

另一個值得注意的細節是,「公開的 prompt」並不等於「實際在跑的 prompt」。Simon Willison 多次強調:

Here’s my big disappointment: Anthropic get a lot of points from me for transparency for publishing their system prompts, but the prompt they share is not the full story. It’s missing the descriptions of their various tools.

所謂「tools」是指 Claude 在對話中可呼叫的搜尋、Artifacts、讀取檔案、Computer Use 等工具的描述,這部分在官方頁面同樣沒有揭露。對想要精準重現 Claude 行為的開發者來說,這仍是黑盒子。

最後是商業成本問題。Hacker News 用戶 pulkitsh1234 的質疑很有代表性:

curious why don’t they bake the system prompt in the model itself ? Why do we pay for these tokens on every API call ? These are just free $ for them, unnecessary bloating the context.

雖然 API 用戶用的是自己寫的 prompt、不會被這些 3000 字灌進帳單,但對任何需要復現 Claude 行為的 benchmark 或 fine-tune 開發者而言,這層資訊不對等的確增加了工作量。


時間軸:Anthropic 在透明度上的累積

  • 2024 年 8 月:Anthropic 首次公開 Claude 3.5 Sonnet 的 system prompt,VentureBeat 報導視為業界創舉。
  • 2025 年 5 月:Claude 4(Opus 4、Sonnet 4)prompt 釋出,Simon Willison 在 Substack 撰寫長篇分析。
  • 2026 年 2 月:Opus 4.6 釋出,新增 <election_info> 區塊。
  • 2026 年 4 月:Claude Code 品質事件後,Anthropic 工程部首次公開「prompt ablation 退步 3%」的事後分析。
  • 2026 年 4 月 18 日:Simon Willison 整理 Opus 4.6→4.7 變動的 git diff。
  • 2026 年 6 月 9 日:Fable 5 / Mythos 5 釋出,開啟「Mythos-class」模型分級。
  • 2026 年 6 月 12 日:Fable 5 / Mythos 5 遭美國商務部出口管制暫停。
  • 2026 年 6 月 30 日:管制解除,7 月 1 日恢復存取。
  • 2026 年 7 月 24 日:Opus 5 釋出,prompt 內含 Fable 5 安全路由與時序備註。
  • 2026 年 8 月 16 日:Anthropic 將 17 個模型的歷代 prompt 整合至單一 release notes 頁面,HN 衝上 239 分。

結語:透明度成為 AI 公司新護城河?

Anthropic 這次把 17 個模型、橫跨兩年的全部 system prompt 攤在同一頁,等於把「AI 行為可審查性」從抽象口號變成可下載、可版本控制、可逐行對照的具體文件。對 OpenAI、Google DeepMind、Meta 而言,這無疑是一次公開的標準拉抬——使用者未來要求其他廠商提供對等文件的壓力會明顯升高。

對 Siami 編輯部來說,更值得追蹤的是 Anthropic 對外承諾的下一個節點:是否有一天官方會連「工具描述」也一併公開?若會,這將真正讓 Claude 變成業界第一個完全可審查的旗艦模型;若不會,這場透明度競賽就仍是半套。

編按:本文綜合整理自 Anthropic 官方技術文件、Hacker News 討論串(239 分 / 123 則留言)、Simon Willison 的逐版剖析、Anthropic 工程部事後分析,以及 Project Glasswing 與 Fable 5 出口管制事件聲明,並加入 Siami 編輯部觀點與分析。

網友熱門留言 (5)

#1 Hacker News 用戶 tosh ▲ 187
早期 system prompt 大約 300 字,最新的版本已經來到 3000 字以上。Opus 5 的 prompt 裡可直接看到一段說明——使用者可能挑了 Fable 5,但 Anthropic 的安全路由機制把請求改送給 Opus 5,因此模型被明確告知要怎麼跟困惑的使用者解釋。
#2 Simon Willison ▲ 142
我把 Anthropic 歷代釋出的 system prompt 整理成 git commit history,方便看每次新版改動了什麼。例如 Opus 4.8 升級到 Opus 5 的那段 diff,最值得看的是 Fable 5 在 2026 年 6 月 9 日首發、6 月 12 日被美國商務部出口管制暫停、6 月 30 日解除、7 月 1 日恢復存取的整套時序說明,整段都被寫進 system prompt。
#3 Hacker News 用戶 pulkitsh1234 ▲ 98
為什麼不把 system prompt 直接 bake 進模型權重裡?為什麼每次 API call 都要讓使用者付費重傳這些 token?對 Anthropic 來說純粹是免費的錢,無謂地膨脹 context。
#4 Hacker News 用戶 altmanaltman ▲ 76
很妙的是,最早期版本的 system prompt 完全沒有任何兒童安全防護條款(新版則有一整個 bullet list 專門處理)。對一家聲稱重視安全對齊的公司來說,第一個版本居然是這樣,Dario 怎麼能接受。
#5 Hacker News 用戶 dev-complete ▲ 64
我把 Claude Opus 4.8 跟 5 的 system prompt 對比過,也比對 Claude Code 那兩版,都沒看到所謂 80% 大幅縮短。難道是 Claude Code 漏的版本不準?還是我不知道 80% 長什麼樣子?這種數字感覺很像在呼嚨人。