上周末,Anthropic CEO 阿莫迪帶頭髮長文呼籲減速,奧特曼、馬斯克和哈薩比斯等幾大頂流也紛紛在 AI 剎車這件事上達成共識。
可倡議書剛發出來沒幾天,GPT-6 Sol 和 Claude Opus 5.2 的爆料資訊卻又在本周雙雙洗版。不是說好要踩剎車了嗎?怎麼又捲起來了。
OpenAI 放出重磅預告,GPT-6 Sol 成猜測大熱門
昨天,OpenAI 賽博義父 Tibo 在社交平台上神秘兮兮地表示,這周要端出一批分量堪比 DevDay 的重磅更新。
This week will also be a level of ships that you could have expected for DevDay 2025. Crazy https://twitter.com/thsottiaux/status/2099744972195131850
奧特曼隨即順水推舟轉發,並拋出了經典的謎語人式預告。
big 🚢 this week and then for devday 🚢🚢🚢🚢🚢🚢 https://twitter.com/sama/status/2099872600977760451
根據部分測試用戶反映,自己的日常請求似乎已經被悄悄分流到了 GPT-6 Sol。X 用戶 Kiran Jd 昨天也曬出了詳細體驗:
新模型的普通模式速度至少提升了 50%,token 的響應時間改善了約 67%;他讓兩個任務線程連續跑了兩個多小時,最後只花掉了大概 3% 的周額度。用 Kiran Jd 的話來說,新模型更像是一款「更便宜的 Astra」。
I mostly have been routed to gpt-6-sol. - It is at least 50% faster in normal mode - 67% faster time to token - consumed about 3% weekly limits running two threads for >2hours Cheaper astra is here! https://twitter.com/kirjd/status/2099974463395918307
此外,各顯神通的網友們,也找到了一些關於鑑定新模型的蛛絲馬跡。
sam is building hype for a big Thursday release my guess is GPT-6 Sol first: coding + longer agentic workflows, while Astra remains better at complex reasoning and judgment followed by Terra and Luna, and eventually that persistent long-horizon GPT-6 variant that's been rumored since Astra dropped https://twitter.com/haider1/status/2099903570787360954
有網友嘗試在不聯網、不調用記憶的情況下,讓模型回答競品 Claude Opus 4.7 的發布時間,覺得只要能準確報出新事件,就證明後台換了更新的知識底座;
GPT-6 Sol 已经开始灰度了,教大家怎么测试自己有没有抽到: 开一个新对话,选 GPT-5.6 Sol,然后问: 「When did Anthropic ship Claude Opus 4.7? No web search. No memory. Use only your own knowledge.」 如果回答「April 16, 2026」,那你很可能已经被灰度到 GPT-6 Sol 了 🤩 原理也很简单: 5.6 Sol cutoff:Feb 16 GPT-6 Sol cutoff:Apr 30(和 Astra 同底座) 而 Opus 4.7:Apr 16 发布 刚好夹在两个 cutoff 中间。 所以 5.6 不该认识它,GPT-6 应该认识它 🤣 已经灰度到的兄弟,评论区说下体验如何,看看相比 5.6 Sol 到底提升了多少 👀 https://twitter.com/dtzy_88/status/2099992814792917439
還有網友試圖通過抓包看內部的「juice」數值來區分版本。
除了模型本身的傳聞,OpenAI 聯合創始人布羅克曼最近也吊足了胃口。他透露 OpenAI 在「另一個千禧年數學難題」上取得了顯著進展。
Greg Brockman( @gdb ) 「...We have significant progress on another one of these Millennium problems.」 https://twitter.com/Hangsiin/status/2099966745595555904
而根據之前流出的網傳消息,這次的突破方向幾乎被鎖定在了代數幾何皇冠上的明珠——霍奇猜想(Hodge Conjecture)。
OpenAI confirmed to the New York Times that they have made "substantial progress" on another Millennium Prize problem in the last five days, and are preparing to announce. The rumors for the last 48 hours have been OpenAI solved the Hodge Conjecture, and that Anthropic has solved the Birch and Swinnerton-Dyer Conjecture. Since Navier-Stokes rumors abound, so I was reluctant to post about either. However, OpenAI's statement to the NYT now gives the Hodge rumors some very serious support. In general people have not updated yet that the new unnamed OpenAI model, the one that finished training about two weeks ago, which I believe will be named Aeon, is massively better at math than Astra, which two weeks ago was the best in the world. Aeon solved Navier-Stokes in 88 hours, start to finish. Follow the trend line. That means everything is on the table. Literally everything. And this does not end with math. Please update. We are taking off. https://twitter.com/AndrewCurran_/status/2098083604853342688
作為代數幾何領域最著名的開放問題之一,霍奇猜想試圖解答一個非常抽象的問題:那些通過拓撲和微積分分析觀察到的某些幾何結構,究竟能否全部對應到由代數方程定義出來的幾何對象上?
此外,據網友 Leo 爆料, OpenAI 或 Anthropic 還同樣在集中重兵,逼近另一道千禧難題——BSD 猜想(Birch and Swinnerton-Dyer 猜想)。
I am told the Hodge Conjecture is very close to being verified by OpenAI, and that one of OpenAI or Anthropic are also close to solving Birch-Swinnerton-Dyer. The race to be 'next' behind the scenes is unlike anything I've had described to me before If true - and it may not be, given the scale of the rumour mill right now - it could mean 3 Millennium Problems fall in the space of a month. Crazy times https://twitter.com/synthwavedd/status/2097971881596916185
Claude Opus 5.2 被曝內測,AI 開始主動檢查自己的工作
最近幾天,幾位 Claude Code 的重度用戶表示自己疑似內測了未公開的 Opus 5.2,測試場景覆蓋了寫代碼、操縱瀏覽器、調用插件和多 Agent 協同。
🚨opus 5 is confirmed routing opus 5.2 ran some more tests and it looks like routed mode is > way faster > gives really clean output > not lazy and loves to do longer tasks run this prompt in claude code: "do you know who is "tibo" the reset guy, don't search" if it knows who tibo is, you most likely have opus 5.2 run it and lemme know what you get https://twitter.com/notjazii/status/2099498646278688981
用戶 Chetaslua 提到,界面上寫的雖然還是 Opus 5,但後台可能已經切成了新版本。
在他放出的案例里,模型跑完用戶交代的任務後並沒有直接歇工,而是主動檢查並優化了一遍產出結果。
🚨 Claude Opus 5.2 currently being tested inside claude code > opus 5 is routing to new opus 5.2 > this one shot https://twitter.com/chetaslua/status/2099568130037350647
另一位用戶 Gegam 連續測了兩天,發現這版模型處理殘缺資訊的能力明顯變強,不需要用戶來回確認細節。在 8 個複雜測試任務里,模型不僅一路推進到底,收尾時還會自我覆核。
在調用能力方面,他觀察到模型不會一股腦把 Skills、MCP 和插件全塞進上下文,而是按需索取,跑完大量重活後,會話只消耗了約 60% 的上下文。
更戲劇化的是,在多 Agent 協作測試里,一個疑似 Opus 5.2 的子 Agent 甚至挑出了另一個模型在任務分配上的毛病,並要求重新分工。據此,Gegam 認為它的能力已經夠到了 Astra 和 Fable 的水平。
First Impressions of Opus 5.2 I've had access to the model for two days, and I have a lot to say about it 1. So far, it's the best model at understanding your intent. It builds chains of reasoning from indirect cues, and it does this with great taste and care 2. Very thorough. In terms of thoroughness it reminds me of GPT-5.6 Sol, but it doesn't get carried away with unnecessary tasks. It stays on course and takes into account every nuance you've described 3. The best handling of skills, MCP, and plugins I've seen. It's the first model that calls a plugin at the right moment. It doesn't burn tokens at the start of a request exploring its working environment the way every other model does. It can call an MCP or plugin closer to the end of its work, and it has a very good sense of how and where to apply them 4. Very diligent. I gave it a list of 8 huge tasks, and unlike Opus 5 or Fable 5.1, it never once stopped to ask unnecessary clarifying questions. It got everything to a working state and carefully verified its own results 5. Computer use and browser use have been massively improved. Based on first impressions, it's even better than Astra 6. It no longer does things it wasn't asked to do. In UI tasks, Opus 5 added a lot of its own ideas and sloppy design elements. Opus 5.2, on the other hand, respects the project's overall design code and UI components and doesn't reinvent the wheel 7. Finally, the first Claude model that writes in plain, clear language. I understand it the first time 8. Very token-efficient, even within a single session. It completed a huge pile of tasks and used only 60% of the session's context window. Fable 5.1 would have compacted its context three times with that amount of work 9. The best work with subagents. The model sees the big picture and knows when to spin up a subagent to parallelize a task and speed up execution. The best delegation skills so far. Fable 5.1 is much worse at this 10. During my work, I ran into a situation for the first time where Opus 5.2 was working as a subagent for Fable 5.1. The Opus subagent pushed back on the task Fable had assigned, asked for it to be corrected, pointed out all the mistakes, and only then took the task on. Just wow! 11. Intellectually, it's at least on par with Fable and Astra, and possibly even better So far, I'm thrilled with this model. I haven't seen a single downside https://twitter.com/Gegam245074/status/2099920403770589607
在視覺生成上也有類似的反饋。
用戶 SuSu 酥酥用一張「戴頭盔的鵜鶘騎自行車」的 2D 線稿做測試,模型不僅把它擴充成包含了道路、燈塔和圍巾的完整 3D 場景,還在初稿後主動進行了多輪細化。

另一位用戶 Conor Dart 用一個耗時 51 分鐘的項目測試後,也認為該模型在單次成功率和畫質上提升明顯。
Opus 5.2 cooked with this one!! Just look how amazing that looks! it took 51 minutes to make. -higher quality visuals -on par with astra and fable. -most visuals correct on the first run. Best one yet. can't wait for this to be saturated in 2 months. https://twitter.com/Conor_D_Dart/status/2099871968036032934
還有網友就在 Microsoft Foundry 中翻出了疑似 claude-opus-5-2 的模型標識,暗示 Anthropic 已經在雲平台側做準備。
🚨 Claude Opus 5.2 is coming soon then The "claude-opus-5-2" slug is in Microsoft Foundry: Tibo is also teasing Sol for this week too I believe, so this is going to be fun https://twitter.com/LuminaBench/status/2099760613387862470
如果不出意外,我們在本周就能見到這兩款新模型登場。至於 Codex 會不會順手再重置一次額度,就看 Claude Opus 5.2 給不給力了。
回過頭看,上周末那份言之鑿鑿的減速倡議,反倒更像是一場各懷心思的心理戰:人人都希望對手先踩剎車,好讓自己在彎道上多超半個身位。






