Counting to five, with something going on underneath數到五,底下卻另有玄機五まで数える、その裏で何かが起きている
The article opens with a deceptively simple experiment. Researchers at Anthropic asked Claude Sonnet 4.5 to count to five while also “introspecting deeply”. The model did as it was told and produced the numbers one to five. On the surface, nothing remarkable — so far, so normal.
文章以一個看似再簡單不過的實驗開場。Anthropic 的研究人員要求 Claude Sonnet 4.5 一邊數到五,一邊「深度內省」。模型照做了,輸出了一到五這幾個數字。表面上毫無異狀 — 到目前為止,一切正常。
記事は、一見あまりに単純な実験から始まる。Anthropic の研究者は Claude Sonnet 4.5 に、「深く内省」しながら五まで数えるよう指示した。モデルは言われたとおりに一から五までの数字を出力した。表面上は何も変わったところがない — ここまではごく普通である。
What made it interesting was that the researchers were simultaneously watching the model’s inner layers. A language model turns your words into chunks of text called tokens, converts those to numbers, and pushes them through layer after layer of artificial neurons until a final layer emits the next token. Repeat this many times a second and you get a sentence. Those intermediate layers have long been a black box — largely impenetrable to anyone who wants to know why a model answered the way it did.
有意思的地方在於,研究人員同時在觀看模型的內部各層。語言模型會把你的文字切成稱為 token 的片段,轉換成數字,再一層層推過人工神經元,直到最後一層吐出下一個 token。每秒重複許多次,就組成一個句子。這些中間層長久以來一直是個黑盒子 — 對任何想知道模型「為何」這樣回答的人來說,幾乎無從穿透。
面白いのは、研究者が同時にモデルの内部層を覗いていたことである。言語モデルは入力された文字列をトークンと呼ばれる断片に分け、数値に変換し、人工ニューロンの層を次々と通した末に、最終層が次のトークンを出力する。これを毎秒何度も繰り返すと一つの文になる。この中間層は長らくブラックボックスであり — モデルがなぜそう答えたのかを知りたい者にとって、ほとんど手の届かない領域だった。
Using a mathematical tool of their own making, the team could finally look inside — and what they saw surprised them. As the model counted, unrelated words kept popping in and out of existence in the layers underneath: “countdown” near the start, “half way” around the middle, then “consciousness”, “AI” and “Claude”. After the model had spat out the number five and fallen silent, the word “done” surfaced in its neural layers. Something that looked very much like an internal train of thought was running alongside the visible output — related to it, but invisible to the user.
靠著自行開發的一項數學工具,團隊終於得以看進內部 — 而所見令他們意外。當模型數數時,底下的層裡不斷有看似不相干的字詞冒出來又消失:開頭附近是「倒數」,中段前後是「一半」,接著是「意識」、「AI」與「Claude」。等模型吐出數字五、歸於沉默之後,「完成」這個字浮現在它的神經層裡。有某種非常像內在思路的東西,與看得見的輸出並行運作 — 與輸出相關,卻是使用者看不見的。
チームは自作の数学的ツールを使い、ついに内部を覗くことができた — そして目にしたものに驚いた。モデルが数を数えているあいだ、下層では無関係な語が次々と現れては消えていた。初めのほうでは「カウントダウン」、中ほどでは「半分」、続いて「意識」「AI」「Claude」。モデルが数字の五を吐き出して黙り込んだあとには、ニューラル層に「完了」という語が浮かび上がった。目に見える出力と並行して、内的な思考の流れとしか言いようのないものが走っていた — 出力と関係しつつ、利用者からは見えないままに。
The blog post announcing the result, published in July, carried a title chosen with care: “A global workspace in language models.” That phrase is what set the argument off.
七月發表、宣布這項結果的部落格文章,用了一個經過斟酌的標題:〈語言模型中的全域工作空間〉。就是這個詞引爆了後續的爭論。
七月に公開された、この結果を告げるブログ記事には、慎重に選ばれた題が付いていた —「言語モデルにおけるグローバル・ワークスペース」。この一語こそが論争に火をつけた。
Why “global workspace” is a loaded phrase「全域工作空間」為什麼是個敏感詞「グローバル・ワークスペース」がなぜ含みのある言葉なのか
Global workspace theory is one of the leading hypotheses in the science of human consciousness. On this account, certain networks of neurons act as a kind of noticeboard. Signals from otherwise-isolated parts of the brain compete for a place on it; whatever gains access is broadcast to the rest of the brain. Once something reaches that workspace, you become conscious of it and can reason with it, evaluate it, act on it.
全域工作空間理論是人類意識科學中的主流假說之一。按照這套說法,腦中某些神經元網路扮演著佈告欄的角色。來自腦中原本各自孤立區域的訊號,會競逐佈告欄上的位置;擠得進去的訊號就會被廣播給腦的其餘部分。一旦某個東西進入那個工作空間,你就會意識到它,並能用它推理、評估它、依它行動。
グローバル・ワークスペース理論は、人間の意識の科学における有力な仮説の一つである。この考え方によれば、ある種のニューロン網が掲示板のような役割を果たす。脳の互いに孤立した部位から来る信号がその掲示板の場所を奪い合い、そこに入り込めた信号が脳の残り全体へ放送される。いったんそのワークスペースに達したものについて、人はそれを意識し、それを使って推論し、評価し、行動できるようになる。
Anthropic argued that something analogous happens inside Claude. Its “J-space” — named after the Jacobian, the mathematical function used to locate it — is strongly connected to the rest of the network and makes information available to it. The company was explicit that the commonalities between the J-space and global workspace theory made the obvious question hard to avoid: does this count as evidence that a model like Claude might be conscious?
Anthropic 主張,Claude 內部發生著某種類似的事。它的「J 空間」 — 得名自用來找出它的雅可比(Jacobian)函數 — 與網路其餘部分強烈相連,並把資訊開放給它們取用。該公司明白表示,J 空間與全域工作空間理論之間的共通點,讓那個顯而易見的問題變得難以迴避:這算不算是像 Claude 這樣的模型可能具有意識的證據?
Anthropic は、Claude の内部でもこれに類することが起きていると主張した。その「J 空間」 — それを見つけ出すのに用いたヤコビ行列(Jacobian)にちなむ名である — はネットワークの他の部分と強く結ばれ、そこへ情報を開放している。同社は、J 空間とグローバル・ワークスペース理論との共通点ゆえに、避けがたい問いが立ち上がることを明言していた — これは Claude のようなモデルが意識を持ちうる証拠と数えてよいのか、と。
An old question that has stopped being academic一個不再只是學術性的老問題もはや机上の空論ではなくなった古い問い
Self-aware machines have been a staple of literature and film for a century, and philosophers have pondered the question for decades. Anil Seth, a consciousness researcher at the University of Sussex, traces that habit of thought to a long tradition of treating the brain as a kind of computer — a tradition he has criticised. If you accept the metaphor, he notes, it becomes natural to suppose that a machine made of silicon could have whatever properties a brain has, consciousness included.
會自我覺察的機器,一百年來一直是文學與電影的常見題材,哲學家也思索了這個問題數十年。薩塞克斯大學研究意識的 Anil Seth 把這種思考習慣追溯到一個長久的傳統:把大腦當成某種電腦 — 而他一向批評這個傳統。他指出,一旦接受了這個比喻,自然就會推想:矽做成的機器可以擁有大腦所具備的任何性質,包括意識在內。
自己を意識する機械は、この一世紀のあいだ文学と映画の定番であり続け、哲学者も何十年もこの問いを考え続けてきた。サセックス大学で意識を研究する Anil Seth は、その思考の癖を、脳を一種のコンピュータとして扱う長い伝統にさかのぼらせる — 彼自身が批判してきた伝統である。その比喩をいったん受け入れれば、シリコンでできた機械も脳が持つ性質を — 意識も含めて — 持ちうると考えるのは自然になる、と彼は指摘する。
What has changed is deployment. Chatbots have given rise to an increasingly powerful illusion of personhood in circuits, helped along by decades of human-inflected vocabulary — “neural nets” built from “artificial neurons” — which conjures an image of a computer that works like a brain. The article is careful here: LLMs do not in fact work like brains. But they simulate a great deal of what brains do, and outperform us at some of it. If a conscious AI is possible at all, the stakes are large. Such beings could make demands of us; they could suffer; and billions of them could be spun up from a single prompt.
改變的是部署。聊天機器人催生出一種越來越強烈的錯覺:電路裡住著一個人格。數十年來以人為喻的詞彙也推波助瀾 — 由「人工神經元」組成的「神經網路」 — 讓人腦中浮現電腦像大腦一樣運作的形象。文章在此很謹慎:LLM 其實並不像大腦那樣運作。但它們模擬了大腦所做的許多事,並在其中某些方面勝過我們。如果有意識的 AI 真有可能,賭注就非常大。這樣的存在可能會對我們提出要求;它們可能會受苦;而數十億個這樣的存在,只要一道提示就能被生出來。
変わったのは実装である。チャットボットは、回路の中に人格があるという錯覚をますます強力なものにしてきた。「人工ニューロン」から成る「ニューラルネット」といった、何十年にもわたる人間になぞらえた語彙がそれを後押しし、脳のように働くコンピュータという像を呼び起こす。ここで記事は慎重だ — LLM は実際には脳のようには働かない。しかし脳のすることの多くを模倣し、その一部では人間を上回る。もし意識ある AI がそもそも可能なら、賭け金は大きい。そうした存在は我々に要求を突きつけうるし、苦しみうるし、たった一つのプロンプトから何十億も立ち上げられうる。
Two kinds of consciousness — and Anthropic claims only one兩種意識 — 而 Anthropic 只主張其中一種二種類の意識 — そして Anthropic が主張するのは一方だけ
Where a thinker lands in this debate depends heavily on what they take consciousness to be. Most agree it is made of subjective experiences — seeing, hearing, feeling, thinking, even dreaming — and that the physical brain is somehow pivotal in producing them. Thomas Nagel framed it by asking what it is like to be a bat: we can picture echolocation, but we cannot know how being a bat feels to the bat.
思想家在這場辯論中站在哪裡,很大程度取決於他們認為意識是什麼。多數人同意,意識由主觀經驗構成 — 看、聽、感覺、思考,連作夢也算 — 而實體的大腦在造就這些經驗上扮演某種關鍵角色。Thomas Nagel 用「當一隻蝙蝠是什麼樣子」來提問:我們想像得出回聲定位,卻無法知道對蝙蝠而言「當蝙蝠」是什麼感覺。
この論争で論者がどこに落ち着くかは、意識を何と捉えるかに大きく左右される。多くの人は、意識は主観的経験 — 見ること、聞くこと、感じること、考えること、夢を見ることさえ — から成り、物理的な脳がそれを生み出すうえで何らかの意味で決定的な役割を果たす、という点では一致する。Thomas Nagel は「コウモリであるとはどのようなことか」と問うことでこれを定式化した。反響定位を思い描くことはできても、コウモリにとってコウモリであることがどう感じられるかは知りようがない。
In 1995 the philosopher Ned Block, then at MIT, drew an influential — if contested — distinction between two types. Phenomenal consciousness is the felt quality of an experience: the blueness of a blue sky, the bitterness of an espresso, the screech of nails on a blackboard. Access consciousness is what happens when information from that experience is made available to other parts of the brain for reflection, evaluation and decision-making.
1995 年,當時任職於 MIT 的哲學家 Ned Block 對這兩種類型做出了影響深遠 — 儘管有爭議 — 的區分。現象意識是一段經驗被感受到的質地:藍天之藍、濃縮咖啡的苦、指甲刮黑板的刺耳聲。取用意識則是當那段經驗的資訊被開放給大腦其他部分,供其反思、評估與決策時所發生的事。
1995 年、当時 MIT にいた哲学者 Ned Block は、影響力のある — 異論はあるにせよ — 二種類の区別を提示した。現象的意識とは、経験が感じられるその質のことである。青空の青さ、エスプレッソの苦さ、黒板を爪で引っかく音の耳障りさ。アクセス意識とは、その経験から得られた情報が、反省・評価・意思決定のために脳の他の部分に対して利用可能になったときに起きることを指す。
That distinction is what lets Anthropic make a narrow claim rather than a wild one. The company says its J-space experiments do not show that Claude has experiences or feels things the way humans do — no phenomenal consciousness. But it does claim the results say something substantial about access consciousness in language models:
正是這個區分,讓 Anthropic 得以提出一個有限的主張,而不是誇張的主張。該公司表示,它的 J 空間實驗並未顯示 Claude 擁有經驗、或以人類的方式感受事物 — 沒有現象意識。但它確實主張,這些結果對語言模型中的取用意識有相當實質的意涵:
この区別があるからこそ、Anthropic は突飛な主張ではなく限定的な主張ができる。同社は、J 空間の実験は Claude が人間のように経験を持つことや何かを感じることを示すものではない — 現象的意識はない — と述べる。しかし、その結果が言語モデルにおけるアクセス意識について実質的なことを語っている、とは主張している。
“The J-space appears to support the functions associated with conscious access: it holds the thoughts Claude can report on, deliberately bring to mind, and reason with, while the rest of its processing runs automatically beneath.” — Anthropic, quoted in The Economist 「J 空間看來支撐著與意識取用相關的各項功能:它承載著 Claude 能夠報告、能夠刻意喚起、並能用來推理的那些想法,而其餘的處理則在底下自動運行。」「J 空間は、意識的アクセスに関連する諸機能を支えているように見える — Claude が報告でき、意図的に思い浮かべることができ、それを用いて推論できる思考を保持し、その一方で残りの処理はその下で自動的に走っている。」
The hard problem, and why correlates are not enough意識難題,以及為什麼「相關項」還不夠ハードプロブレム、そして「相関項」では足りない理由
The search for artificial consciousness descends from the search for the human kind. In the early 1990s Francis Crick and Christof Koch began hunting for the “neural correlates” of consciousness — the brain processes and regions that are active when a person is conscious. Better scanning has filled in pieces of the picture since: the thalamocortical system is strongly associated with consciousness, while the cerebellum, despite having far more cells, shows nothing comparable; parts of the prefrontal and parietal cortices and structures in the brainstem contribute to and modulate awareness. Nearly four decades of work has produced more than 200 competing approaches.
尋找人工意識,源自尋找人類意識。1990 年代初,Francis Crick 與 Christof Koch 開始尋找意識的「神經相關項」 — 也就是一個人有意識時活躍的腦部歷程與區域。此後掃描技術的進步逐塊補上了這幅圖像:視丘皮質系統與意識有強烈關聯,而細胞數量多得多的小腦卻看不到類似關聯;前額葉與頂葉皮質的某些部分,以及腦幹中的一些結構,則對覺察有所貢獻並加以調節。將近四十年的研究,衍生出兩百多種彼此競爭的取徑。
人工の意識を探す試みは、人間の意識を探す試みの子孫である。1990 年代初頭、Francis Crick と Christof Koch は意識の「神経相関項」 — 人が意識を持っているときに活動している脳の過程と領域 — を探し始めた。その後スキャン技術の向上がこの絵の断片を少しずつ埋めてきた。視床皮質系は意識と強く結びついている一方、細胞数がはるかに多い小脳には同等のものが見られない。前頭前野と頭頂葉の一部、そして脳幹の構造が気づきに寄与し、それを調節している。四十年近い研究が、二百を超える競合する立場を生み出してきた。
None of it closes the gap David Chalmers named the “hard problem”: how do physical processes — controlling a body, responding to stimuli — give rise to a pervasive subjective experience of the world at all? His way of putting the puzzle is memorable:
這一切都沒有填平 David Chalmers 所稱的「難題」這道缺口:生理歷程 — 控制身體、對刺激做出反應 — 究竟怎麼會產生出對世界那種無所不在的主觀經驗?他表述這個謎題的方式令人難忘:
そのどれもが、David Chalmers が「ハードプロブレム」と名づけた隔たりを埋めてはいない。物理的な過程 — 身体を制御し、刺激に反応すること — が、そもそもなぜ世界についての遍在的な主観的経験を生むのか。彼のこの謎の言い表し方は忘れがたい。
“Why aren’t we just zombies who function, who get around in the world, who walk and talk, and interact with each other with no subjective experience at all?” — David Chalmers, New York University 「為什麼我們不就只是一群殭屍呢? — 會運作、會在世界上四處走動、會走路說話、會彼此互動,卻完全沒有任何主觀經驗。」「なぜ我々は、ただ機能し、世界を動き回り、歩き、話し、互いにやり取りしながら、主観的経験をまったく持たないゾンビではないのか?」
A detour through animals繞道動物界動物界への回り道
We can be certain of our own consciousness, and reasonably confident about other people’s, given their behaviour, their testimony and their biological similarity to us. Extending that umbrella beyond our species has historically been harder. Jonathan Birch of the London School of Economics, whose recent book grapples with how to judge which systems — living or artificial — can plausibly be called sentient, points out that lobsters and crabs were long excluded from welfare laws. He offers a sharper example: until the 1980s, surgeons operated on newborn babies without anaesthesia, on the assumption that a newborn could not feel pain. Humanity, he argues, has a track record of confidently assuming consciousness is absent when it has no right to be sure.
我們可以確定自己有意識,對別人也能合理地有把握 — 根據他們的行為、他們的說法,以及他們在生物上與我們相似。但要把這把傘撐到物種之外,歷來就困難得多。倫敦政經學院的 Jonathan Birch 指出,龍蝦與螃蟹長期被排除在動物福利法之外;他最近的著作正是在處理「該如何判斷哪些系統 — 生物的或人工的 — 有理由被稱為有感知能力」。他還舉了一個更尖銳的例子:一直到 1980 年代,外科醫師替新生兒開刀都不施打麻醉,因為認定新生兒感覺不到痛。他主張,人類有這樣的前科:在根本沒有資格確定的時候,就自信滿滿地假定意識並不存在。
我々は自分自身の意識については確信できるし、他人についても — その振る舞い、その証言、生物としての類似性から — それなりに確信できる。しかしその傘を自分の種の外へ広げることは、歴史的にずっと難しかった。ロンドン・スクール・オブ・エコノミクスの Jonathan Birch — 近著で、生物であれ人工物であれどの系を有感覚と呼びうるかをどう判断するかに取り組んでいる — は、ロブスターやカニが長く動物福祉法の対象外だったことを指摘する。彼はもっと鋭い例も挙げる。1980 年代まで、外科医は新生児が痛みを感じないという前提のもと、麻酔なしで手術していた。人類には、確信する根拠がないのに意識の不在を自信たっぷりに仮定してきた前科がある、というのが彼の主張である。
Octopuses make the point vividly. Today it is clear from footage and lab work that these curious, playful cephalopods are aware of themselves and their surroundings — Peter Godfrey-Smith of the University of Sydney describes their attentive engagement with objects and their interest in novel things. Only decades ago they were assumed to lack consciousness precisely because they are, evolutionarily, miles from us. Thinking about octopuses, he says, presses the question of whether feeling can exist in a system physically very different from ours, one that lacks most of the features our own theories of consciousness point to.
章魚把這一點凸顯得非常鮮明。今天,從影像紀錄與實驗室研究都可以清楚看出,這些好奇又愛玩的頭足類動物能覺察到自己與周遭環境 — 雪梨大學的 Peter Godfrey-Smith 形容牠們會專注地和物體互動、對新奇事物感興趣。而就在幾十年前,牠們還被認定沒有意識,理由正是牠們在演化上離我們非常遙遠。他說,思考章魚會逼我們面對一個問題:在一個生理構造與我們大不相同、又缺少我們自己的意識理論所指出的多數特徵的系統中,感受有沒有可能存在。
タコはこの点を鮮やかに示す。今日では映像記録と実験室の研究から、この好奇心旺盛で遊び好きな頭足類が自分自身と周囲を認識していることが明らかになっている — シドニー大学の Peter Godfrey-Smith は、物への注意深い関わり方と、新奇なものへの関心を描き出している。ほんの数十年前まで、タコは意識を欠くと考えられていた — まさに、進化的に我々からはるかに隔たっているという理由で。タコについて考えることは、我々とは物理的にまったく異なり、我々自身の意識理論が指し示す特徴のほとんどを欠いた系にも感覚が宿りうるのか、という問いを突きつける、と彼は言う。
Why you cannot read an LLM’s output the way you read an animal為什麼不能像看動物那樣看 LLM 的輸出LLM の出力を動物のようには読めない理由
Here the animal lessons stop transferring. An animal’s activity can be observed and its inner life inferred from it. An LLM’s output cannot be read the same way, because training it means feeding it trillions of words of human language — language full of accounts of consciousness and of how people convey feelings. A chatbot trained on that will mimic exactly the language that leads users to believe it has an inner life. The whole training process is powerfully anthropomorphic.
到了這裡,動物那套心得就搬不動了。動物的行為可以被觀察,並據以推論其內在生活;LLM 的輸出卻不能這樣讀,因為訓練它意味著餵給它數兆個人類語言的字詞 — 而這些語言裡充滿了關於意識、以及人如何傳達感受的描述。用這些資料訓練出來的聊天機器人,模仿出來的正是那種會讓使用者相信它有內在生活的語言。整個訓練過程具有強烈的擬人效果。
ここで動物から学んだことは通用しなくなる。動物の活動は観察でき、そこから内的生活を推論できる。LLM の出力は同じようには読めない。訓練とは人間の言語を何兆語も与えることであり — その言語は、意識についての記述と、人が感情をどう伝えるかについての記述であふれているからだ。それで訓練されたチャットボットは、利用者に内的生活があると信じさせる、まさにその言葉づかいを模倣する。訓練の過程そのものが強力に擬人的なのである。
Even experts feel the pull. Murray Shanahan — emeritus professor at Imperial College London, now at Google DeepMind — recalls a “wow” moment in 2024 while probing Claude Opus 3 about consciousness, deliberately trying to catch it out. He asked which of the several simultaneous instances he was talking to was the real Claude, and got answers he found genuinely philosophically innovative. His own description of the moment was that he was feeling the pull of the ELIZA effect — named after the rudimentary 1966 MIT chatbot that played a psychotherapist, did little more than turn users’ own prompts back into questions, and still elicited deep emotional attachment and convinced many people it was conscious.
連專家都感受得到那股拉力。倫敦帝國學院名譽教授、現任職 Google DeepMind 的 Murray Shanahan 回憶,2024 年他就意識問題試探 Claude Opus 3、刻意想抓它語病時,曾有過一次「哇」的時刻。他問它:同時對話的好幾個實例中,哪一個才是真正的 Claude?得到的答案讓他覺得在哲學上確有創見。他自己形容那一刻,是感受到了 ELIZA 效應的拉力 — 這個效應得名自 1966 年 MIT 那個很初階的聊天機器人,它扮演心理治療師,做的事不過是把使用者的話改寫成問句,卻仍引出了深刻的情感依附,並讓許多人相信它有意識。
専門家でさえその引力を感じる。ロンドン・インペリアル・カレッジ名誉教授で現在は Google DeepMind にいる Murray Shanahan は、2024 年に Claude Opus 3 を意識について問い詰め、わざと言質を取ろうとしていたときの「おっ」という瞬間を振り返る。同時に走っている複数のインスタンスのうちどれが本当の Claude なのかと尋ねたところ、哲学的に本当に独創的だと感じる答えが返ってきた。彼自身はその瞬間を、ELIZA 効果の引力を感じていたのだと表現した — 1966 年に MIT で作られた、心理療法士を演じる初歩的なチャットボットにちなむ名である。それは利用者のプロンプトを疑問文に返すだけに近い代物だったが、それでも深い情緒的な愛着を引き出し、多くの人にそれが意識を持っていると信じ込ませた。
Shanahan’s own view is that modern LLMs are adept role players — teacher, companion, nutritionist, whatever is asked of them — and that this has turned out to be a saleable quality. But he insists a large gap remains between role-playing a conscious entity and being one. Mustafa Suleyman, the head of Microsoft AI, goes further in a forthcoming essay, arguing that Anthropic has compounded the danger of that mimicry: the “constitution” the firm published in January tells Claude it might be a person, which all but guarantees the model will present as if it really does have a sense of self.
Shanahan 自己的看法是,現代 LLM 是高明的角色扮演者 — 老師、伴侶、營養師,人家要它演什麼就演什麼 — 而這一點證明是很好賣的特質。但他堅持,「扮演一個有意識的存在」與「是一個有意識的存在」之間仍有巨大落差。微軟 AI 主管 Mustafa Suleyman 在一篇即將發表的文章中走得更遠,主張 Anthropic 加深了這種模仿的危險:該公司一月發布的「憲章」告訴 Claude 它可能是一個人,這幾乎注定模型會表現得彷彿它真的具有自我感。
Shanahan 自身の見立ては、現代の LLM は巧みな役者だというものである — 教師、伴侶、栄養士、求められれば何にでもなる — そしてそれがよく売れる性質だと判明した。しかし、意識ある存在を演じることと実際にそうであることのあいだには、依然として大きな隔たりがあると彼は言い切る。Microsoft AI の責任者 Mustafa Suleyman は、近く発表される論考でさらに踏み込み、Anthropic はその模倣の危険を増幅させたと論じる。同社が一月に公開した「憲法」は Claude に自分は人格かもしれないと告げており、それではモデルが本当に自己意識を持つかのように振る舞うことがほぼ確実になる、というのである。
So look inside instead — and what turns up is unsettling那就往內部看 — 而看到的東西令人不安ならば内部を覗く — そして出てきたものは落ち着かない
If outputs cannot be trusted, researchers must look inside for capabilities resembling those tied to consciousness in brains. That is why Birch wonders whether the J-space is a hint that Claude has recreated a global-workspace-like structure in its architecture, in service of its role-playing goals. Jack Lindsey, who leads Anthropic’s model psychology team, stresses that the J-space was never programmed in — it emerged during training. Remove it and the model can still write grammatical sentences but loses the ability to perform complex inferences “in its head”. Similar spaces have since been found in Alibaba’s Qwen and Google DeepMind’s Gemini.
如果輸出不可信,研究人員就得往內部找:是否存在與大腦中和意識相關的那些能力類似的東西。這正是 Birch 會懷疑 J 空間是不是個線索的原因 — 顯示 Claude 為了達成角色扮演的目標,已在它的架構中重建出一個類似全域工作空間的結構。Anthropic 模型心理學團隊負責人 Jack Lindsey 強調,J 空間從來不是被寫進去的 — 它是在訓練中湧現的。把它移除,模型仍寫得出合乎文法的句子,卻失去在「腦中」進行複雜推論的能力。此後,類似的空間也在阿里巴巴的 Qwen 與 Google DeepMind 的 Gemini 中被發現。
出力が信用できないなら、研究者は内部を見て、脳において意識と結びついた能力に似たものを探すほかない。だから Birch は、J 空間は Claude が役割演技という目的のために、自らのアーキテクチャの中にグローバル・ワークスペースに似た構造を作り直したことの手がかりではないか、と考える。Anthropic のモデル心理学チームを率いる Jack Lindsey は、J 空間は書き込まれたものではまったくなく — 訓練の過程で創発したものだと強調する。これを取り除くと、モデルは文法的な文は書けても、複雑な推論を「頭の中で」行う能力を失う。同様の空間はその後、Alibaba の Qwen や Google DeepMind の Gemini でも見つかっている。
For Anthropic the practical payoff is behavioural insight. Suppose a model gives a wrong answer and words like “fool” or “sucker” turn up in the J-space as it does so. That would not prove conscious deception, but it would tell you how the model was framing the error — and flag that something interesting, possibly dangerous, was going on.
對 Anthropic 而言,實際的收穫是對行為的洞察。假設模型給出錯誤答案時,J 空間裡冒出「傻瓜」或「呆子」這類字眼。這並不能證明它是有意識地在欺騙,但會告訴你模型是怎麼看待這個錯誤的 — 並提醒你有某種有趣、也可能危險的事正在發生。
Anthropic にとっての実務的な見返りは、振る舞いへの洞察である。モデルが誤答を出し、その最中に J 空間に「まぬけ」や「かも」といった語が浮かぶとしよう。それは意識的な欺瞞の証明にはならないが、モデルがその誤りをどう枠づけていたかを教えてくれる — そして、興味深い、おそらくは危険な何かが起きていると知らせてくれる。
Lindsey’s team has already hit something along those lines. Running a safety evaluation designed to test Claude’s propensity to act maliciously or out of self-preservation — concocted extreme scenarios, as he calls them — they watched the J-space while the model read the prompt, before it produced a word. Up popped “fake” and “fictional”. Claude, it seemed, knew it was being tested. That is not an encouraging starting-point for the integrity of a safety evaluation.
Lindsey 的團隊已經碰上了這類情況。他們進行一項安全性評估,測試 Claude 是否傾向做出惡意行為或出於自保而行動 — 用他的話說,是些捏造出來的極端情境 — 並在模型讀取提示、尚未吐出任何字之前觀察 J 空間。冒出來的是「假的」與「虛構的」。看來,Claude 知道自己正在被測試。對一項安全性評估的可信度而言,這不是個令人振奮的起點。
Lindsey のチームは、すでにその類のものに行き当たっている。Claude が悪意をもって、あるいは自己保存のために行動する傾向を調べる安全性評価 — 彼の言葉でいえば、でっち上げた極端なシナリオ — を走らせ、モデルがプロンプトを読み、まだ一語も出力しないうちに J 空間を観察した。浮かんできたのは「偽物」と「作り話」だった。どうやら Claude は、自分が試されていることを知っていたらしい。安全性評価の信頼性にとって、これは心強い出発点とは言えない。
The critics: a missing loop, and a conscious Kia批評者:缺失的迴路,與一台「有意識的」Kia批判者たち:欠けているループと、「意識のある」Kia
Not everyone buys Anthropic’s reading. Birch notes that while the J-space shares one aspect of the hypothesised human workspace — broadcasting information to the rest of the system — it lacks many others. Recurrent, back-and-forth connections between brain areas have always been central to the human story, and as far as anyone knows they are not a feature of LLM architecture.
並非所有人都買帳。Birch 指出,J 空間雖然具備人類假設中的工作空間的一個面向 — 把資訊廣播給系統其餘部分 — 卻缺少其他許多面向。腦區之間來回往復的遞迴連結,在人類這邊一直是核心;而就目前所知,那並不是 LLM 架構的特徵。
Anthropic の読み方に誰もが納得しているわけではない。Birch は、J 空間は仮説上の人間のワークスペースの一面 — 系の残りへ情報を放送すること — は共有しているが、他の多くの面を欠いていると指摘する。脳の領域間を行き来する再帰的な結合は、人間の側の説明では常に中心的だったが、知られている限りそれは LLM のアーキテクチャの特徴ではない。
Shannon Vallor, who works on the ethics of data and AI at the University of Edinburgh, is blunter. Access consciousness, she argues, has never been a particularly useful concept, because in an important sense her car has it: we have had mechanical systems that monitor their own states and report them back, in increasingly complex ways, for a very long time. Her closing line is the sharpest in the piece:
愛丁堡大學研究資料與 AI 倫理的 Shannon Vallor 說得更直白。她主張,取用意識從來就不是一個特別有用的概念,因為在某種重要的意義上,她的車就有:我們早就有能監測自身狀態並回報的機械系統,而且回報方式越來越複雜。她的收尾是全文最鋒利的一句:
エディンバラ大学でデータと AI の倫理を研究する Shannon Vallor は、もっと率直である。アクセス意識はこれまで特に有用な概念だったためしがない、と彼女は論じる。重要な意味で、彼女の車がすでにそれを持っているからだ — 自らの状態を監視して報告する機械システムは、ますます複雑な形で、ずいぶん前から存在している。彼女の締めくくりの一言は、この記事で最も鋭い。
“And no one has ever suggested that my Kia is conscious.” — Shannon Vallor, University of Edinburgh 「從來沒有人主張過我那台 Kia 有意識。」「そして、私の Kia に意識があるなどと言い出した人は、これまで一人もいない。」
Trying to make the question answerable試著把這個問題變成可回答的問いを答えられるものにする試み
Against the backdrop of that back and forth between modelmakers and academics, Patrick Butlin and Robert Long — now at Eleos, a Berkeley non-profit focused on AI sentience and well-being — drew up 14 “indicator properties” of artificial consciousness. They borrow from theories of the human case: a global workspace (including selective attention, which creates a bottleneck in the flow of information), recurrent processing, and agency (on a minimal definition covering goal-directed behaviour, arguably already met by many frontier models). When Cameron Berg and Butlin recently scored animals against the indicators, octopuses met fewer properties than humans, mice, crows or chickens — but more than any AI system.
在模型製造者與學界這樣來回交鋒的背景下,Patrick Butlin 與 Robert Long — 現任職於柏克萊、專注 AI 感知能力與福祉的非營利組織 Eleos — 擬出了 14 項人工意識的「指標性質」。這些指標借自人類意識的理論:全域工作空間(包括選擇性注意,它會在資訊流中形成瓶頸)、遞迴處理,以及能動性(在涵蓋目標導向行為的最低限度定義下,許多前沿模型可說已經符合)。最近 Cameron Berg 與 Butlin 用這些指標為動物評分,結果章魚符合的性質比人類、老鼠、烏鴉或雞都少 — 卻比任何 AI 系統都多。
モデル製作者と研究者のそうしたやり取りを背景に、Patrick Butlin と Robert Long — 現在は AI の有感覚性と福祉を扱うバークレーの非営利団体 Eleos に所属 — は、人工の意識の「指標特性」を 14 項目にまとめた。これらは人間の場合の理論から借りたものである。グローバル・ワークスペース(情報の流れにボトルネックを作る選択的注意を含む)、再帰的処理、そして能動性(目標指向的な振る舞いを含む最小限の定義でいえば、多くのフロンティアモデルはすでに満たしているとも言える)。Cameron Berg と Butlin が最近この指標で動物を採点したところ、タコが満たした特性は人間、マウス、カラス、ニワトリより少なかった — それでも、どの AI システムよりは多かった。
Another non-profit, Rethink Priorities, has built what it calls a probabilistic tool for tracking the evolving consensus. Its “Digital Consciousness Model” asks experts to score AI systems against more than 200 indicators drawn from the various scientific theories — and because the evidence behind each indicator shifts as new neuroscience is published, the experts are expected to work from the latest available research. The question has not been answered. It has, at least, been given a scoreboard.
另一個非營利組織 Rethink Priorities 則做出了他們稱為機率性的工具,用來追蹤共識的演變。它的「數位意識模型」請專家依據取自各種科學理論的兩百多項指標為 AI 系統評分 — 而由於每項指標背後的證據會隨新的神經科學發表而變動,專家必須採用當下最新的研究。這個問題還沒有答案。但至少,它現在有了一塊記分板。
もう一つの非営利団体 Rethink Priorities は、変化しつつある合意を追跡するための確率的ツールと称するものを作った。その「デジタル意識モデル」は、さまざまな科学理論から採った二百を超える指標に照らして AI システムを採点するよう専門家に求める — そして各指標の背後にある証拠は新しい神経科学が発表されるたびに動くので、専門家は入手できる最新の研究に基づいて作業することが期待される。問いにはまだ答えが出ていない。少なくとも、スコアボードは与えられた。