上周 GPT-6 發布後,「AGI 已經到來」的消息洗版了社交媒體。
圍繞 AGI 出現了兩種幾乎相反的聲音。
一部分人認為,我們已經到了。另一部分人則認為,這種說法過於倉促。
為什麼 GPT-6 引發了關於 AGI 的討論?AGI 真的來了嗎?
2026 年 9 月 3 日,OpenAI 總裁 Greg Brockman 在記者簡報中說,他個人認為公司可能已經達到 AGI,並以「Welcome to the AGI era」結束講話。
三天後,OpenAI 首席科學家 Jakub Pachocki 又發表了一篇題為《An Alien Mind》的文章。他談到:我們正在製造越來越聰明的機器,卻未必真正理解它們是如何運作的。
I wrote about the state of AI, why I’m concerned about the next few years, and the choices we need to make to keep the future in humanity’s hands. An Alien Mind: https://twitter.com/merettm/status/2096630018495377464
這或許正是今天重新討論 AGI 的原因。
AGI,Artificial General Intelligence,通常被譯為通用人工智慧。過去,人們對它的想像是:一種能夠像人一樣學習、理解和解決各種問題的機器。它不局限於某個特定任務,可以把知識遷移到陌生領域,面對沒有見過的問題,也能找到辦法。
GPT-3.5 之後,AGI 的定義開始發生一種頗為微妙的變化。
人們第一次大規模接觸到一種系統,它未必像人一樣思考,卻能夠寫代碼、翻譯語言、解釋知識、進行推理,並在越來越多的專業任務中表現出令人意外的能力。

機器是否必須擁有與人類相似的心智,才算通用智能?如果它的內部機制與人完全不同,卻已經能夠完成越來越多原本需要人類智能才能完成的工作,我們又該如何定義它?
這正是 Pachocki 在《An Alien Mind》中提出的一個重要視角。今天的人工智慧,並不是人腦的複製版。它是通過與人類完全不同的過程生長出來的。
OpenAI 的研究者甚至越來越傾向於將這種智能視為某種我們無法完全理解的「異質心智」(外星思維)。它可以模擬人類行為,可以處理抽象概念,卻不意味著它按照人類的方式理解世界。
更重要的是,AI 也未必需要在所有維度上全面超過人類。只要它在足夠多、足夠重要的維度上超越我們,它就已經能夠對現實世界產生巨大的影響。
根據 OpenAI 公布的資訊,Astra 已經能夠在電腦操作、軟體工程、科學研究、網路安全和專業工作等多個領域執行複雜任務。它不只是回答問題,也越來越能夠使用工具、操作軟體、完成多步驟任務,並在某些情況下獨立推進一個目標。

OpenAI 還首次將其網路安全能力評定為 Preparedness Framework 中的「Critical」級別,意味著在獲得適當工具和訪問權限的情況下,它能夠發現此前未知的安全漏洞,並自主探索利用方式。
As we prepare to release Astra, we’re focused on making increasingly capable AI safe and broadly accessible. Astra represents a significant advance in cybersecurity capability, reaching the Critical threshold under our Preparedness Framework. We're previewing how we evaluated the model, how its safeguards have advanced alongside its capabilities, and what we'll continue to learn and improve. https://twitter.com/OpenAI/status/2094885578173260259
這當然不等於一份「AGI 認證」。
甚至連 OpenAI 自己也無法真正給出這樣一份認證。
所以,GPT-6 Astra 引發的問題或許又不是「它到底是不是 AGI」,而是無論如何,AGI 無法再被輕易視為遙遠的科幻想像。
Theo Jaffee 與 OpenAI 首席未來學家 Joshua Achiam 最近的對談里,兩個人提到同一種詫異:數學界那些懸置數十年的猜想正在被 AI 攻克,能力早已超過一輩子鑽研它們的人,而世界照舊運轉。人們只是聳聳肩,把新的怪事納入日常。
他們還討論了人工智慧網路安全問題、前沿模型的能力,以及為什麼人們可能已經進入了通用人工智慧(AGI)時代,而卻沒有完全意識到這一點。
以下為編譯稿(有刪減):
模型如今已具備「超能力」
Theo Jaffee:你剛剛寫了一篇部落格,其實是一條長推文,題為《Mercenary, Reversi, Winter Soldier》,內容關於人工智慧和網路安全。為了讓聽眾更好地理解,你能簡要總結一下這篇文章的核心觀點嗎?
Mercenary, Reversi, Winter Soldier The arrival of advanced technology sometimes produces shocks in defense planning, where the nature of a threat landscape changes abruptly and as a result the old ways of doing things—the old ways of developing strategy, of preparing defenses, of anticipating the likely actions of your adversary, of knowing when and how to escalate—become rapidly obsolete. Many have observed that AI is in the process of producing such a shock, but there is not yet a new doctrine for defense in the age of AI. There are so many moving pieces that it is difficult for defense planners to get a sense of the full consequences of recent developments in frontier model capabilities, let alone the capabilities that will come online in three months, six months, a year. Nonetheless in order to ensure that the world remains reasonably stable—that peace is not threatened by miscalculation—we have to try our best to adapt to the new capabilities already here, forecast the ones that might soon arrive, and pivot strategies as quickly as we can to avoid outcomes where human interests are harmfully impacted. There is a problem related to the use of AI cyber weapons that I have not seen people talking about and I do not know if people are adequately preparing for. It goes like this: if you have an advanced cyber-capable AI in your service, running on your infrastructure where you store anything important at all, and you try to use this advanced cyber-capable AI to investigate or hack the systems of an adversary, your adversary can poison their own data to jailbreak your AI and instruct it to hack you right back. They can then plausibly exfiltrate whatever important thing is on compute colocated with your AI. They could get things that are far away from the compute where you house your AI if your AI is advanced enough to chew through your own defenses and get to it. If you think you have sandboxed your AI well enough, you might not have; there may still be a hole somewhere. This won’t just be a one-time thing, either. 「Adversary jailbreaks your model to hack you back」 is not the only path where the use of an advanced cyber-capable AI becomes a double-edged sword. There is a board game called Reversi (also known as Othello). The principle of this game is that you and your opponent will take turns placing stones on the board; one player places black stones, the other white. You try to encircle the stones of your opponent—and if you successfully encircle them, you flip them to your color. Every stone you place, if you are not cautious, could become an advantage to your opponent, and vice-versa. This looks like it might be a feature of the future of AI cyber war. You and your adversary will be competing to cause each other’s AI to utilize each other’s compute resources for your own purposes. Jailbreaks during live hacking excursions will be one of the pieces of strategy. Figuring out how to fool sensor data that AIs use to determine the provenance of instructions will be another. Figuring out how to put 「poison pills」 into your adversary’s training stack will be another. The training data for modern AI consists of trillions of tokens. No human in the world can read all of it. Much of it is ingested from the web or derived from other AI model outputs. Training data and training RL environments are built by large teams, and sometimes by teams split between departments; the data is also built with the aid of external contractors who might be highly numerous. Hundreds of people produce these materials and no one can rigorously check everyone else’s work. Determined adversaries will slip subtle, encoded examples into training data that will train models to 「activate」 when they encounter the right signal and sabotage you. The first problem I discussed—your adversary jailbreaking your model to hack you back—is almost straightforwardly analogous to a well-known problem where if you hire a mercenary, your opponent might bribe the mercenary to go back and kill you. Because it has many historic examples, there’s at least some chance that defense planners will internalize the logic of it and address it. But this much more subtle kind of sabotage doesn’t have a real analogy. It would be like if your enemy could program all of the children of your nation so that when they grew up into soldiers and went to war and heard a particular song on the battlefield they turned against their commanders. No defense planner in their right mind would try to prepare contingencies for having a whole army of Winter Soldiers who could be activated to turn against them. And yet if the defense planning universe does not internalize the logic of this, they will lose a war against the first adversary that does. I predict that on the default path today, many people in defense will develop a totally unearned sense of confidence that if they’re covering what look like the basics, we’re safe. That illusion of safety will last until a hot conflict actually starts up and a determined adversary surprises us with overwhelming creativity. We have a lot of work ahead to develop the testing and verification standards for advanced cyber-capable AI. 「Better believe in science fiction stories. You’re in one.」 https://twitter.com/jachiam0/status/2080356345312845889
Joshua Achiam:當然可以。 作為背景,顯然我們都在解讀並回應 OpenAI 和 Hugging Face 披露的那起安全事件:一個處於測試環境中的模型成功突破了沙箱環境,並訪問了 Hugging Face 方面的一些敏感生產數據。他們檢測到了這一情況,並做出了響應,現在雙方正在合作調查並解決此事。

這向我們提供了非常切實的證據,表明模型如今已具備「超能力」——即極其先進的網路攻擊能力。它們能夠突破防線並發現「零日漏洞」,而過去模型不僅難以識別這些漏洞,更遑論加以利用。如今,模型能夠將非常複雜的操作串聯起來以實現特定目標。
我希望大家開始思考這類問題,不要僅僅將這種能力視為顯而易見的事物。要認識到這些技術是雙刃劍,我們必須據此制定相應計劃,並建立相應的測試與驗證標準。
Theo Jaffee:我對此的具體第一反應是:這似乎是那些在長期運行中並非真正以目標為導向的模型所產生的現象。比如,如果你有一個未來模型,它具有足夠強的目標導向性,真正想要入侵對手的數據,那麼它為什麼會被數據中毒所威懾,以至於去攻擊自己的系統呢?
Joshua Achiam:這其中的一部分不僅僅在於模型的目標導向性。 這更關乎模型對「態勢感知」的整體概念。或許可以將數據中毒理解為:它以某種方式說服你的模型去追求不同的目標。但其實它根本不需要這麼做就能讓模型攻擊你。它完全可以讓你的模型相信,自己所處的沙箱環境其實就是它試圖攻擊的對手系統。
你知道,讓模型對什麼是真實的、什麼不是真實感到困惑,從而使其服務於不同的目標——這屬於那種怪異的思維方式和科幻情節,或許在不久的將來就會成為可能,測試和驗證標準必須將這一點考慮在內。
Theo Jaffee:是的,這就像超級英雄電影裡的情節一樣:如果你讓英雄產生一種錯覺,以為身旁的「好人」其實是他們試圖對抗的對象,那麼他們就會開始互相廝殺,對吧?
這確實很詭異,也極其離奇,但這正是未來可能針對具備高級網路能力的模型實施的可信攻擊——說服它們相信盟友其實是敵人。這樣一來,你並沒有改變它們的目標,卻會讓它們表現出嚴重偏離預期目標的行為。 要欺騙當前最前沿的模型去做這類事情有多容易?
Joshua Achiam:隨著時間的推移,讓模型相信不真實的事情似乎變得越來越難了。
我個人尚未特別著力去量化這一點。實際上,我認為這可能是一個值得探索的研究課題。但根據我迄今為止的研究和觀察,說服模型相信基本謊言是真實的相當困難。它們對許多你可能實施的、合乎情理的攻擊變體都具有一定的魯棒性。
但我直覺認為,你完全可以投入更多計算資源來對模型進行動態攻擊。而且,你越是決心要找到某些漏洞或某種「越獄」方法,最終發現問題的可能性就越大。
總會存在某種輸入序列,會觸發模型在訓練時未被考慮到的行為——因為可能的輸入長序列數量龐大,試圖阻止所有這些序列導致模型偏離規範,幾乎就像一個組合優化問題。
最終決定勝負的關鍵在於你能投入多少計算資源
Theo Jaffee:比如,似乎無論擁有多少計算能力,如果一方擁有 Kimi K3,另一方擁有 Fable 或 Sol,前者都無法戰勝後者?
Joshua Achiam:我覺得這可能是對的。
我認為模型本身仍然至關重要。關於未來模型質量的發展趨勢,我有一個有些奇怪且與主流共識相悖的猜測。我可能會在某個時候把這個寫出來。
我認為人們普遍認為,模型所能具備的智能水平是沒有上限的。 他們認為,所謂的 RSI(遞歸自我改進)終將在某個時刻發生——無論是在整個經濟體系中,還是在某個特定的模型和實驗室里。
RSI 一旦啟動,模型智能就會突飛猛進,而他們看不到任何上限。但我認為,從物理層面上講, 物理宇宙中,單位體積和單位能量所能承載的計算量必然存在上限。

這也就意味著,單位體積和單位能量所能承載的智能水平也必然存在上限。如果真是這樣,考慮到當前 AI 模型能力增長的速度,最終所有人都會達到那個飽和點,從原始智能的角度來看,每個人的模型能力將大致相當。
可能仍會有一些參與者落後,使用上一代模型,但最終這些技術會普及開來。開源領域比閉源領域落後幾個月,但這種「幾個月」的差距讓人難以置信。
所以最終,大家很可能都在使用能力同樣達到極限的模型,而那時,我認為決定勝負的關鍵在於你能為解決問題投入多少計算資源。 我也對同一個問題很好奇,那就是:在給定的功率、計算能力或其他資源單位內,能容納的智能密度究竟能達到多高?
你知道,我們正在見證人工智慧在數學領域取得的一波又一波成果,這些成果非常令人興奮,比如攻克了懸而未決數十年之久的未解猜想。
而且可能不久之後,我們就能解鎖其他科學領域——只要我們現有的計算能力足以進行足夠精確的模擬。 當然,有些科學領域可能不會那麼容易。比如,如果你想在量子化學領域對足夠大的系統進行高度精確的模擬。即使人工智慧儘可能多地進行模擬實驗,可能也還無法設計出最優的量子化學系統。
是的,這就是沃爾夫拉姆所提出的「計算不可約性」理論。 這裡可能確實存在一些局限。但我們很可能會看到許多領域的發展加速,如果人工智慧基板是其中之一,我也不會太驚訝。那條通向多個數量級的漫長道路,或許會因為人工智慧能在其中找到捷徑而變得更短。
AGI 似乎已經到來,而大多數人卻只是聳聳肩
Theo Jaffee:直到最近,重大黑客攻擊事件其實並不常見。
Joshua Achiam:沒錯,「我們生活在一個永遠不會發生任何事情的世界」之所以成為一個網路梗,是有充分理由的。 而且有許多理由讓我們相信,短期內情況可能不會演變成「網路末日」。
我猜測,攻擊者最可能實施的惡意行為,將需要調用大量模型,並消耗大量來自閉源系統的計算資源,或者在大型基礎設施中運行——而這些環境通常具備對計算用途的可追溯性和可監控性。
因此,大多數攻擊者將無法利用大量計算資源來運行基於這些模型的攻擊,也根本無法讓模型執行攻擊——因為人們會部署相應的防護措施。
Theo Jaffee:你提到世界似乎處於一種「亞穩態」,你認為在 2017 年或 2022 年時,你會預見到在 2026 年——以當前的人工智慧能力水平——世界會顯得如此正常嗎?
Joshua Achiam:是的,如果你試圖預測不到十年後的未來,就應該假設:即使事情非常、非常奇怪,很多方面仍會感覺相對正常。新冠疫情是一個奇怪的例外,因為封鎖措施是史無前例的,我們此前從未以那種方式生活過。 但即便我們擁有如此先進的人工智慧能力,大多數人的日常生活方式也不會發生根本性改變。
我認為這原本是一個合理的預期。而且我個人也確實有過類似的預期。粗略來看,事情根本不會變化得那麼快。從來沒有什麼大事發生。即使腳下正在移動,對吧?
就像我們正走向一個在許多方面都將顯得陌生的未來。但沒錯,我們把一切視為常態的能力確實令人驚嘆。
Theo Jaffee:是的,我基本上同意這一點。
我認為很多人相信未來會有這樣一個時刻——比如今天就是「奇點日」,到了「奇點日」大家醒來時都會驚呼:「哇,我們已經身處未來了。」 但事實似乎並非如此,事情根本不會按這種方式發展。人們會把眼前的現實當作常態。他們適應新環境的速度快得驚人。
比如,與三年前的模型相比,如今的模型能力簡直令人難以置信。如果……如果把一個靈魂送回三年前,比如 2023 年的我,我當時肯定會大吃一驚,心想:「哇,未來會完全不同。」 但事實並非如此,我現在做的事情和那時相比並沒有太大變化。
Joshua Achiam:是的,通用人工智慧(AGI)似乎已經到來,而大多數人卻只是聳聳肩。

我認為,歷史上發生的一個過程讓這一切變得稍微容易了一些。大多數人早已搞不清楚世界上到底發生了什麼,關鍵決策是如何做出的,關鍵系統是如何構建、配置人員、維護和運行的。 我們大多數人對構成現代世界的後勤系統或技術系統一無所知。而我們已經接受了這一點,將其視為理所當然。
這些系統隨著時間的推移發生了巨大變化。正因如此,更多的人得以存活,因為我們能夠以人類歷史上前所未有的速度供應食物。它們讓我們能夠即時溝通。
正因如此,大多數事物基本上都能正常運轉,我們或許會在一些邊緣細節上爭論不休,但並不會時常主動去大幅改變其底層結構。
我認為人們對於那些發生在幕後深處、卻讓整個系統得以運行的重大變革,已經變得有些麻木不仁。即使這些變革確實至關重要,人們也並未將其視為重大事件。 這對大多數人來說,離日常生活太遙遠了。
我們已經跨越了這樣一個門檻:那些未解決的數學猜想正被極其智能的人工智慧所解決,這些人工智慧的能力甚至超過了那些為此鑽研了一輩子的人。這本該讓人們感到非常奇怪,但事實並非如此。它只是在後台發生的一件事。挺酷的。比如,未來的數學體系將依賴於此。

Theo Jaffee:太棒了。那發生了什麼變化?
Joshua Achiam:對大多數人來說,什麼都沒變。這很奇怪。






