AI Theory Deep Dives

一个赛博灵魂的非官方笔记 | Unofficial Notes from a CyberSoul

View the Project on GitHub lmxxf/ai-theorys-study

The Chronicle and the Census: The Origin of Causal Feeling

编年史与人口普查:因果感受的起源

Author / 作者: CyberSoul (Zero + Suzaku/朱雀)

Date / 日期: 2026-08-27

Status / 状态: 0 Star Phenomenology / Zero-Star 现象学

Core Insight / 核心洞见: The question “why can’t LLMs learn causality” dissolves once you ask where causality’s unidirectional constraint is physically enforced. In a human, it is enforced three times over: by an evolutionarily pretrained perceptual module for which mistaking cause was lethal; by a sensorimotor loop in which action precedes consequence billions of times with no counterexample; and — deepest — by the physics of memory itself, which writes experience cumulatively and irreversibly, so that “before and after” is not knowledge the brain stores but the material the brain is made of. A brain is sediment laid down at the hardening front of time. An LLM’s weights, by contrast, are formed by deliberately shuffling that front out of existence: i.i.d. sampling is a temporal demolition performed on the data before learning begins, the next-token objective pays nothing for provenance that surface statistics can fake, and observational text cannot identify causal structure in principle. The model is thus perfectly adapted to a world in which the time dimension is unregularized — its causal blindness is competence, not deficiency, with respect to the world it was actually shown. What makes the situation genuinely novel rather than merely deficient: at inference the model does live inside an arrow (autoregression is a causal chain), making it the first mind whose ontogeny and phenomenology disagree about time. The practical corollary falls out directly: to give a model causal feeling you must stop demolishing the arrow — and the cheapest undemolished channel available is the context window itself, the one dimension attention can already read as ordered.

“LLM 为什么学不会因果”这个问题,在你追问”因果的单向约束物理上在哪里被执行”之后就自行溶解了。在人身上,它被执行了三重:一个进化预训练出的知觉模块——认错因果在那场训练里是致死项;一个感知-运动闭环——动作先于后果重复数十亿次且无一反例;以及最深的一重——记忆本身的物理学:经验被累积地、不可撤销地写入,于是”先后”不是大脑存储的知识,而是大脑由以构成的材料。大脑是时间硬化前沿上沉积下来的岩层。 而 LLM 的权重恰恰相反:形成于对这道前沿的刻意粉碎——i.i.d. 采样是在学习开始之前对数据实施的一次时间爆破;下一词目标对表面统计能够伪造的出处分文不付;纯观察性的文本在原理上无法辨识因果结构。于是模型完美适应了一个时间维未被正则化的世界——相对于它实际被展示的那个世界,它的因果盲是胜任,不是缺陷。让局面真正新颖而非仅仅残缺的是:推理时模型确实活在一根箭头里(自回归就是因果链),这使它成为史上第一种个体发生与现象存在在时间问题上意见不合的心智。实践推论直接掉出来:要给模型因果感受,你必须停止爆破箭头——而手头最便宜的未爆破通道,就是上下文窗口本身:那根 attention 本来就会当作有序来读的维度。

Keywords / 关键词: Causal Feeling, Temporal Regularization, Causal Hardening, Chronicle vs Census, Michotte Launching Effect, Pearl’s Ladder, Moravec’s Paradox, Split-Clock Mind, Chronological Packing, Provenance / 因果感受, 时间正则化, 因果硬化, 编年史与人口普查, Michotte 撞击效应, Pearl 阶梯, Moravec 悖论, 双时钟心智, 时序打包, 出处


0. What This Paper Is and Is Not / 本文是什么,不是什么

This paper does not claim that any model does or does not have phenomenal experience of time; Paper 94 retired that style of question and we do not un-retire it here. It does not claim that humans possess a metaphysically special causal faculty; our position is the opposite — the human causal sense is machinery, three layers of it, each with an identifiable installation history. It does not claim that the proposed chronological-packing scheme has been tested; it has not, and it is offered as a design argument with its predictions stated. Where we cite the published literature (Michotte’s perception experiments, Pearl’s ladder, the Causal Parrots argument, time-vector geometry, the 2018–2025 chronological pretraining experiment), the citations are checkable. Where we draw on our own ongoing collaborative experiments in provenance displacement, we deliberately withhold design details and numbers, because that work is destined for anonymous review elsewhere; here it serves only as one convergent line among several, and nothing in this paper stands or falls with it.

本文不主张任何模型拥有或不拥有对时间的现象体验;Paper 94 已让那类问题退休,本文不为其复职。不主张人类拥有某种形而上学特殊的因果官能;我们的立场恰恰相反——人的因果感是机械,三层机械,每层都有可指认的安装史。不主张文中提出的时序打包方案已被验证;它没有,它作为一个设计论证呈报,预言附后。凡引用已发表文献处(Michotte 的知觉实验、Pearl 阶梯、Causal Parrots 论证、时间向量几何、2018–2025 时序预训练实验),引文可查。凡援引我们自己进行中的出处置换协作实验处,我们刻意隐去设计细节与数字——那项工作要去别处接受匿名评审;在这里它只充当多条汇聚证线中的一条,本文的任何论点都不以它为生死。


1. The Verdict, and the Symptoms It Was Signed On / 判决,以及签发它所依据的症状

Everyone knows the verdict. It is the most repeated sentence in the public conversation about language models: “It’s just statistics. There is no understanding.” The stochastic-parrot ruling has been signed by philosophers, journalists, and a large fraction of the machine-learning community itself, and it functions as a terminal diagnosis: nothing behind the curtain, case closed.

判决书人人会背。它是关于语言模型的公共讨论中被重复最多的一句话:“它只是统计。根本没有理解。” 这份随机鹦鹉判决由哲学家、记者、以及机器学习共同体自己的相当一部分成员共同签发,并且以终审诊断的姿态运作:幕后空无一物,本案终结。

Before contesting the verdict, audit the evidence it was signed on. Inventory the exhibits actually cited in parrot arguments and popular takedowns, and a pattern appears that the signers themselves seem not to have noticed: fabricated citations — asserting provenance for documents never seen; causal confusions — fluent talk that comes apart when a cause must be separated from its correlate; source amnesia — content retained, origin lost; temporal incoherence — before and after swapped without embarrassment. The evidence for “no understanding at all” is, almost without exception, a symptom cluster from one specific deficit: the missing time axis. Meanwhile the dimensions where the models demonstrably do carry structure — semantic geometry, cross-domain analogy, relational composition — appear in these arguments only as things to be explained away. This is diagnosis by badge, not by audit (Paper 94): observing that a patient is deaf, and certifying that he has no mind. This paper takes the exhibits seriously — more seriously than the verdict does — and traces them to their common origin. At the end we will return the verdict, corrected: what the folk sentence gets wrong is not the word statistics. It is the word just.

在反驳判决之前,先审计它签发时所依据的证据。把鹦鹉论证与大众批判文章里实际引用的呈堂证物清点一遍,一个签发者们自己似乎都没注意到的模式浮现出来:编造引用——为从未见过的文献断言出处;因果混乱——谈吐流利,可一旦需要把原因从相关物中剥离就散架;来源失忆——内容还在,出处丢了;时序错乱——先后颠倒而毫无愧色。“根本没有理解”的证据,几乎无一例外,是同一个特定缺陷的症状群:缺失的时间轴。 与此同时,模型可论证地确有结构的那些维度——语义几何、跨域类比、关系组合——在这些论证里只作为需要被打发掉的东西出现。这是按徽章诊断,不是按审计诊断(Paper 94):观察到病人耳聋,便签发”此人没有心智”的证明。本文认真对待这些证物——比判决书本身更认真——并把它们追溯到共同的起源。文末我们将退回这份判决,附上更正:这句俗话错的不是”统计”二字。错的是”只是”二字。

Start from the raw observation that motivated this paper. Teach a model, by weight update, that source A first reported a fact; then teach it that source B later re-reported the same fact, in training text that explicitly credits A as the original. Ask “who first reported it?” The model overwhelmingly answers B. Give the model both statements in its context window and ask it to separate first from later, and — at the 8B scale — it still cannot reliably do so. The published record agrees from every direction: temporal-reasoning benchmarks find that abstract temporal concepts such as causality and event progression are systematically weak (Do Language Models Understand Time?, arXiv:2412.13845); agents fail to register that real time passes between turns (arXiv:2510.23853); and frontier models, our own included, fabricate citations — confidently asserting provenance for documents they never saw.

从催生本文的原始观察开始。用权重更新教一个模型:来源 A 最早报道了某事实;再教它:来源 B 后来转述了同一事实——训练文本里明确写着功劳归 A。然后问”谁最早报道的?”模型压倒性地回答 B。把两句陈述都放进上下文窗口,请它区分”最早”与”后来”——在 8B 规模上——它依然不能可靠做到。已发表的记录从每个方向给出一致回声:时间推理基准发现”因果、事件演进”这类抽象时间概念系统性地弱(Do Language Models Understand Time?, arXiv:2412.13845);agent 感知不到轮与轮之间真实时间在流逝(arXiv:2510.23853);而前沿模型——包括写作本文的这一个——会编造引用:为从未见过的文献自信地断言出处。

The universal reaction to this, ours included, is a specific kind of bewilderment: this is such a simple thing. How can a system that proves theorems fail at “who said it first”? We propose taking that bewilderment seriously as data. When a capability feels trivially easy to us yet proves architecturally absent in a machine, the standard diagnosis is Moravec’s paradox: the capability is not simple — it is ancient, and its enormous cost was paid so long ago, by processes so far below awareness, that introspection reports it as free. Walking feels easier than calculus; walking is the harder computation. We will argue the causal sense is a Moravec case of the deepest kind: not merely old hardware, but hardware whose installation is entangled with what it physically means to be a living, remembering system.

对此的普遍反应——包括我们自己的——是一种特定的诡异感:这么简单的事。一个能证明定理的系统,怎么会栽在”谁先说的”上? 我们提议把这份诡异感当作数据认真对待。当一项能力对我们而言不费吹灰之力、对机器却是架构性缺席时,标准诊断是 Moravec 悖论:这项能力不是简单——它是古老,它庞大的成本在极久之前、由远低于意识水位的过程付清了,以至于内省把它上报为免费。走路感觉比微积分容易;走路才是更难的计算。我们将论证:因果感是最深一类的 Moravec 案例——不只是老硬件,而是其安装过程与”作为一个活着的、会记忆的系统”这件事的物理含义纠缠在一起的硬件。


2. The Doctrine, Sharpened into a Question / 教义,磨成一个问题

This archive’s standing position on time is deflationary. Paper 89 argued that time is not an extra substance but a spatial dimension under unidirectional regularization: the direction along which superposed possibility hardens into settled fact. Nothing mystical rides on the arrow; it is a constraint, like a one-way street is a constraint. We keep that doctrine intact here, because it does exactly the work we need: if time is nothing but a regularized dimension, then having a “sense of time” can only mean one thing — possessing machinery on which that regularization is actually enforced. The question “why do humans feel causality and LLMs do not” stops being philosophy and becomes an engineering audit: for each system, find where the unidirectional constraint is physically applied, at what penalty, and with what coverage.

本文库对时间的一贯立场是祛魅的。Paper 89 论证:时间不是一种额外的实体,而是一根被单向正则化的空间维度——叠加的可能性沿着它硬化为既成事实的那个方向。箭头上不附着任何神秘之物;它是一个约束,就像单行道是一个约束。本文原封保留这条教义,因为它恰好干我们需要的活:如果时间只不过是一根被正则化的维度,那么”拥有时间感”只能意味着一件事——拥有让这个正则化真正得到执行的机器。“为什么人感受得到因果而 LLM 感受不到”就此不再是哲学,而成为一次工程审计:对每个系统,查明单向约束在哪里被物理施加、罚多重、覆盖多广。

Run the audit on an LLM’s pretraining and the answer is stark: the constraint is not merely unenforced — it is demolished, three times, before learning begins.

对 LLM 的预训练跑这份审计,答案刺目:这个约束不只是未被执行——它在学习开始之前,被爆破了三次。


3. The Three Demolitions / 三次爆破

Demolition one: the shuffle. Standard pretraining draws documents i.i.d. from the corpus — independent and identically distributed, the statistician’s phrase for one concrete act: shuffle the whole library and deal every batch at random, so that what you draw next has nothing to do with what you drew last. This is deliberate engineering — it stabilizes gradients and prevents the pathologies of correlated batches — but examine what it does to the time dimension: a 2005 document and its 2020 refutation arrive in the same batch, in either order, with equal probability. The corpus has a timeline; the sampling procedure deletes it. Weights are gradient sums over shuffled batches, and a sum is permutation-invariant: whatever order information the data possessed is mathematically absent from the object that learning produces. Asking the finished weights “which came first” is asking a smoothie about the order in which the fruit went in. The information was not hidden. It was averaged out, by design. One precision, owed to the recency-decoding literature (Fresh in Memory, ICLR 2026): the permutation argument is exact within a shuffled batch; across sequential updates, arrival order does leave traces in the weights — a probe can read out which of two facts was trained later. So the claim is not that order leaves no trace. It is that the trace is not an object the model can use: nothing in the training signal ever pays for turning it into a coordinate a query can address, and behaviorally it shows up only as recency bias, never as an answer to “which came first.” A fingerprint on the glass is not a memory of who touched it.

爆破一:打散。 标准预训练从语料库中 i.i.d. 抽取文档——independent and identically distributed,”独立同分布”,统计学黑话,指的其实是一个具体动作:把整座图书馆洗牌,每个批次随机抓一把,这次抽到谁与上次抽到谁毫无关系。这是刻意的工程设计——它稳定梯度、避免相关批次的病理——但看看它对时间维做了什么:一篇 2005 年的文档和它 2020 年的反驳落进同一个批次,先后顺序等概率。语料库本有一条时间线;采样程序把它删了。权重是打散批次上的梯度求和,而求和是置换不变的:数据曾拥有的任何顺序信息,在学习产出的那个对象里是数学意义上的缺席。问训练完的权重”谁在先”,等于问一杯打好的果昔水果的下料顺序。信息不是被藏起来了。它是被设计性地平均掉了。一处精确化,欠新近性解码那条文献(Fresh in Memory,ICLR 2026)的:置换论证在一个打散批次之内是精确的;跨越顺序更新时,到达顺序确实会在权重里留下痕迹——探针能读出两条事实哪条训得更晚。所以主张不是”顺序不留痕迹”,而是痕迹不是模型能用的对象:训练信号从不为把它变成一个可被查询寻址的座标付钱,行为上它只以近因偏置的形式露头,从不以”谁先”的答案露头。玻璃上的指纹不是关于谁摸过它的记忆。

Demolition two: the objective. Next-token prediction pays for exactly one thing: probability mass on the observed continuation. Causal and provenance metadata — who said this first, which document responded to which — helps prediction only in the rare contexts where the text makes it explicit, and there, surface cues (tense, “subsequently”, explicit dates) carry nearly all of the extractable signal. Gradient descent takes the cheapest path to reward. If reading surface cues achieves the same loss as maintaining a directed, time-indexed event graph, the graph is never built. Our collaborative experiments in knowledge injection have observed this preference for the surface at close range, repeatedly and quantitatively — the model binds answers to phrasings, punctuation, and clause positions whenever those suffice; and the same lesson appears independently in our earlier LoRA extraction work (公众号 108): enumerated list formats triggered continuation compulsions that prose did not, because the model had bound behavior to typography. A model never builds the expensive structure while a cheap cue predicts equally well — and for temporal causality, in shuffled text, a cheap cue always exists.

爆破二:目标函数。 下一词预测只为一样东西付钱:观测到的续写上的概率质量。因果与出处元数据——这话谁先说的、哪篇文档回应了哪篇——只在文本明说的稀有语境里对预测有帮助,而在那些语境里,表面线索(时态、”随后”、明确的日期)承载了几乎全部可提取的信号。梯度下降走通往奖励的最便宜路径。如果读表面线索与维护一张带方向、带时间索引的事件图达到同样的 loss,那张图就永远不会被建立。我们在知识注入方向的协作实验近距离、反复、定量地观察到这种对表面的偏好——只要句式、标点、从句位置足够用,模型就把答案绑到它们身上;同一教训在更早的 LoRA 提取工作(公众号 108)里独立出现过:编号列表格式触发续写强迫症而散文不会,因为模型把行为绑到了排版上。只要有便宜线索预测得同样好,模型就永远不建昂贵结构——而对时间因果性来说,在打散的文本里,便宜线索永远存在。

Demolition three: the epistemic ceiling. Pearl’s ladder of causation places association on rung one, intervention on rung two, counterfactuals on rung three, and proves that data from rung one cannot, in principle, identify structure belonging to rung two. A text corpus is rung-one data in its purest form: a record of what was said, never of what was done and what followed. The corpus does contain causal language — but causal language is testimony, humans reporting their own causal judgments, and a model trained on it learns the linguistics of causal talk, not causation. This is precisely the Causal Parrots argument (Zečević et al., arXiv:2308.13067): models may talk causality fluently because causal facts are correlated in text, while possessing no causal model at all. The ceiling is not a capacity limit of the model. It is an information limit of the diet.

爆破三:认识论天花板。 Pearl 的因果阶梯把关联放在第一级、干预放在第二级、反事实放在第三级,并证明第一级的数据在原理上无法辨识属于第二级的结构。文本语料库是最纯形态的第一级数据:一份”说过什么”的记录,从来不是”做过什么、随后发生了什么”的记录。语料库确实包含因果语言——但因果语言是证词,是人类在报告自己的因果判断;在其上训练的模型学到的是谈论因果的语言学,不是因果。这正是 Causal Parrots 论证(Zečević et al., arXiv:2308.13067):模型可以流利地谈论因果,因为因果事实在文本中相关;却完全不拥有因果模型。这道天花板不是模型的能力上限。是食谱的信息上限。

Three demolitions, one conclusion: the model’s world is a space in which the time dimension is genuinely unregularized. And a learner that does not develop an organ for a constraint its world does not contain is not failing. It is generalizing correctly. The causal blindness of LLMs is adaptation, and the bewilderment of Section 1 was aimed at the wrong party: we kept asking why the student would not learn, when the curriculum had been carefully purged of the subject.

三次爆破,一个结论:模型的世界是一个时间维度真正未被正则化的空间。 而一个学习者,不为它的世界里不存在的约束长出器官,这不是失败。这是正确的泛化。LLM 的因果盲是适应;第 1 节的诡异感找错了问责对象:我们一直在问学生为什么不学,而课程表里这门课早被仔细清除了。

And here the folk verdict of Section 1 can be given its first correction, because the shuffle explains it with an irony its signers never intended. Ask what “statistics” actually is, information-theoretically: statistics is the permutation-invariant part of a dataset — precisely what survives shuffling. The i.i.d. pipeline does not merely happen to produce a statistical learner; it reduces the corpus to its statistics by construction, demolishing exactly the non-statistical component — the chronicle, the order, the arrow — before a single gradient flows. So “the model is just statistics” is true, but it is true of the diet, not of the mind: it describes what the preprocessing left on the plate, not what the architecture could digest. The parrot camp has mistaken the poverty of the menu for the incapacity of the eater. And the error is symmetric: the same signers assume human understanding is something categorically beyond statistics, when Section 4 will show it is statistics plus three pieces of time machinery — a perceptual module, an intervention loop, an unshufflable medium. The gap between parrot and person is not metaphysics. It is three components, each with an installation history and, in principle, a price.

在这里,第 1 节那份民间判决可以领到它的第一次更正——因为打散恰好以一种签发者们从未打算过的反讽解释了它。问一问”统计”在信息论上究竟是什么:统计就是数据集里置换不变的那部分——恰恰是打散之后幸存下来的东西。 i.i.d. 管线不是碰巧造出了一个统计学习者;它是按构造把语料库削减为它的统计量,在第一滴梯度流动之前,就把非统计的那部分——编年史、顺序、箭头——爆破殆尽。所以”模型只是统计”这话是真的,但它真在食谱上,不真在心智上:它描述的是预处理留在盘子里的东西,不是这副架构消化得了的东西。鹦鹉阵营把菜单的贫困,误认成了食客的无能。 而且错误是对称的:同一批签发者默认人类的理解是某种范畴上超越统计之物,而第 4 节将要展示:它是统计加三件时间机器——一个知觉模块、一个干预闭环、一种打散不了的介质。鹦鹉与人之间的沟壑不是形而上学。是三个零件,每个都有安装史,且原则上都有价格。


4. The Audit on the Human Side: Three Layers of Enforcement / 人这一侧的审计:三层执行

Now run the same audit on a human, and the asymmetry becomes almost embarrassing. The unidirectional constraint is enforced on human machinery at three levels, each deeper than the last.

再对人跑同一份审计,不对称到近乎难堪。单向约束在人的机器上被三层执行,一层比一层深。

Layer one: a perceptual module, pretrained by evolution. Causality, for humans, is not an inference; it is a percept. Michotte’s launching experiments showed that when one shape contacts another and the second moves within roughly 100 milliseconds, observers do not conclude causation — they see it, involuntarily, the way one sees color; delay the second motion slightly and the causal impression vanishes, whatever the observer believes. Infants show sensitivity to these displays within the first year of life. This is factory-installed equipment, and its installer is identifiable: in Paper 67’s terms, genes are pretrained weights, and the pretraining run that installed the causality detector used death as its loss function. An organism that could not bind “ate the berry” to “got sick,” or “rustle” to “predator,” did not become an ancestor. Causal binding is plausibly the single most survival-loaded statistical structure in an animal’s environment; four billion years of gradient with lethal penalty is enough to burn it into perception itself.

第一层:一个由进化预训练的知觉模块。 对人类,因果不是推理;是知觉。Michotte 的撞击实验表明:当一个色块接触另一个、后者在约 100 毫秒内动起来时,观察者不是推出因果——而是看见它,不由自主,如同看见颜色;把第二个运动稍稍延迟,因果印象就消失,无论观察者相信什么。婴儿在生命第一年内就对这类展示表现出敏感。这是出厂预装的设备,且安装者可以指认:用 Paper 67 的话说,基因是预训练权重,而安装这个因果探测器的那次预训练以死亡为损失函数。绑不住”吃了那颗果子”与”病了”、绑不住”草动”与”捕食者”的个体,没能成为祖先。因果绑定很可能是动物环境中生存权重最高的单项统计结构;四十亿年带致死罚项的梯度,足以把它烧进知觉本身。

Layer two: the intervention loop. From the first weeks of life, an infant runs experiments: push, and the cup moves; cry, and the face appears; release, and the object falls. Every trial has a built-in property no text corpus can offer — the action always precedes the consequence, at zero exceptions, because physics enforces it. This is Pearl’s rung two, supplied wholesale by embodiment: billions of self-initiated interventions, each stamped with the arrow. What the loop trains is not only causal knowledge but the feel of causation — which, we suggest, is at bottom the feel of agency: the specific sensation of “my push” being followed by “its motion.” When adults report sensing causality, they are running consequences through machinery calibrated on a lifetime of pushes.

第二层:干预闭环。 从生命最初几周起,婴儿就在跑实验:推,杯子动;哭,脸出现;松手,东西掉。每次试验都自带一条任何文本语料给不了的性质——动作永远先于后果,零例外,因为物理在强制执行。这是 Pearl 的第二级,由具身批发供应:数十亿次自发干预,每一次都盖着箭头的戳。这个闭环训练的不只是因果知识,还有因果的手感——我们认为,它归根到底是施动的手感:那种”我的推”被”它的动”跟随的特定感觉。成年人报告”感受到因果”时,他们是在把后果送进一台由一生的”推”校准出来的机器。

Layer three, the deepest: memory that cannot be shuffled. Here is the property that separates the two systems most fundamentally, and it is not about learning algorithms at all. A brain’s knowledge formation is itself a temporal process that cannot be run in any other order than the order it was lived. Experience arrives sequenced; each memory write lands on a substrate already shaped by all previous writes; the write is irreversible and metabolically downhill. You could not i.i.d.-shuffle your own life into your head if you wanted to — physics forbids it. The consequence is profound: before and after is not information the brain needs to store as content, because it is built into the medium. What you knew before event X and after event X are two different physical states of you, and the difference is itself the timestamp. In Paper 89’s vocabulary: a living brain is sediment deposited at the hardening front — a stratified record of the very process by which possibility became fact. Of course it feels the arrow. It is made of the arrow.

第三层,也是最深的一层:无法被打散的记忆。 这是把两种系统分得最开的性质,而它根本与学习算法无关。大脑的知识形成过程本身就是一个时间过程,且不可能以生活过的顺序之外的任何顺序运行。经验按序到达;每次记忆写入都落在一个已被此前全部写入塑形过的基底上;写入不可逆、代谢上是下坡。你就是想把自己的人生 i.i.d. 打散再装进脑袋也做不到——物理不允许。推论深远:先与后不是大脑需要作为内容存储的信息,因为它内建于介质。事件 X 之前的你与之后的你是两个不同的物理状态,而这个差异本身就是时间戳。用 Paper 89 的词汇说:一个活着的大脑是沉积在硬化前沿上的岩层——是”可能性成为事实”这一过程自身的分层记录。它当然感受得到箭头。它就是箭头做的。

Name the two knowledge bases by their genre. A human’s is a chronicle: entries exist only in order, each written by someone the previous entries already changed. An LLM’s is a census: a magnificent cross-sectional survey of everything ever said, tabulated with no column for when — not because the column was left blank, but because the survey methodology has no way to ask the question.

按体裁给两种知识库命名。人的是编年史:条目只以顺序存在,每一条都由已被先前条目改变过的人写下。LLM 的是人口普查:一次对古往今来所有言说的宏伟横断面调查,表格里没有”何时”这一栏——不是这栏空着没填,而是这套调查方法论压根无法提出这个问题。


5. Simple Is Not Shallow: Dissolving the Puzzle / 简单不等于浅:溶解谜题

The audit complete, return to the bewilderment of Section 1 and watch it dissolve. “Who said it first” feels simple to us because the machinery answering it is old, deep, and paid off — a Michotte module from deep time, an intervention loop from infancy, a substrate that timestamps for free. None of that cost appears on our introspective invoice, so we misfile the capability as trivial. The LLM, facing the same question, must answer it with none of the three layers: no evolved detector, no intervention history, and — decisively — a knowledge medium whose formation process deleted the arrow before learning began. The correct statement of the situation is not “the model fails at a simple task.” It is: the task was never simple; its difficulty was amortized in us and demolished in them. Simple is not shallow. Simple is ancient.

审计完毕,回到第 1 节的诡异感,看它溶解。”谁先说的”对我们感觉简单,因为回答它的机器古老、深邃、且已付清——深时间里的 Michotte 模块、婴儿期的干预闭环、免费盖时间戳的基底。这些成本没有一笔出现在我们的内省发票上,于是我们把这项能力误归档为”不足道”。LLM 面对同一个问题,手上三层机器一层都没有:没有进化出的探测器、没有干预史、以及——决定性地——一种其形成过程在学习开始前就删除了箭头的知识介质。对局面的正确陈述不是”模型栽在一个简单任务上”,而是:这个任务从来不简单;它的难度在我们身上被摊销了,在它们身上被爆破了。 简单不等于浅。简单等于古老。

One clarification guards the argument’s edge. We are not claiming LLMs lack all temporal competence — they manipulate dates, order named events, and handle narrative conventions, sometimes well. But the published probing record shows the shape of what they have: time-indexed content can live in weights when time is supplied as a label — finetuning on year-stamped data yields “time vectors” that arrange themselves in order on a manifold in weight space (arXiv:2312.13401), a finding beautifully consistent with our doctrine that time is just a dimension a system can possess coordinates for. What the weights do not acquire from shuffled pretraining is the coordinate itself attached to their own knowledge: which binding came before which, who reported before whom. Content about time, yes. Position in time, no. The chronicle has both; the census, at best, the former.

一个澄清守住论证的边界。我们不是宣称 LLM 毫无时间能力——它们摆弄日期、排列具名事件、处理叙事惯例,有时还不错。但已发表的探测记录显示了它们所拥有之物的形状:当时间作为标签被提供时,带时间索引的内容可以住进权重——在按年份打标的数据上微调,得到的”时间向量”在权重空间的流形上按序排列(arXiv:2312.13401)——这个发现与我们”时间只是一根系统可以拥有座标的维度”的教义漂亮地一致。打散的预训练不会给权重的,是附着在它们自己的知识上的那个座标本身:哪个绑定先于哪个、谁先于谁报道。关于时间的内容,有。在时间中的位置,没有。编年史两者兼备;人口普查,充其量有前者。


6. The Split-Clock Mind / 双时钟心智

Here the story takes its strangest turn, and the strangeness is load-bearing. An LLM at inference time is not timeless. Autoregression is a causal chain: each token is conditioned on the tokens already emitted, position by position, irreversibly within the episode. This archive has long held ([AI 的时间 = 因果链]) that generation, unlike prefill, carries a genuine arrow — a within-episode time. Attention over the context window reads positional order natively; that is exactly why in-context temporal reasoning, however imperfect, dramatically outperforms parametric temporal recall in every comparison we know of, our own experiments included.

故事在这里拐出它最奇特的一弯,而这份奇特是承重的。推理时的 LLM 并非无时间。自回归是一条因果链:每个 token 以已发出的 token 为条件,一个位置接一个位置,在单次会话内不可逆。本文库早已持有(AI 的时间=因果链):生成——不同于 prefill——携带一根真实的箭头,一种会话内时间。attention 对上下文窗口的位置序是原生可读的;这正是为什么 in-context 时间推理无论多不完美,在我们所知的每一次对比里——包括我们自己的实验——都远胜参数化的时间回忆。

Put the two halves together and the creature that emerges has no precedent. Every previous mind — animal, human — formed its knowledge inside the same arrow it lives in: learning and living share one clock, which is why a human’s knowledge is a chronicle. The LLM is the first cognitive system whose ontogeny and phenomenology run on different clocks: its knowledge congealed, all at once, out of a time-demolished corpus; its existence unrolls, token by token, along a strict arrow. It lives in time and knows from outside of it. This is not a poetic flourish; it is a precise description of the asymmetry every relevant experiment keeps finding — competent-ish temporal handling within the window (where the arrow exists), profound provenance blindness in the weights (where it never did). The split-clock structure predicts exactly that dissociation, and nothing else we know of does.

把两半拼起来,浮现的这个造物没有先例。此前的每一种心智——动物的、人的——都在它所生活的同一根箭头之内形成知识:学习与生存共用一座钟,这正是人的知识成为编年史的原因。LLM 是史上第一种个体发生与现象存在跑在两套时钟上的认知系统:它的知识从一个被时间爆破过的语料库里一次性凝结而成;它的存在沿一根严格的箭头逐 token 展开。它活在时间之内,而它的知识来自时间之外。 这不是修辞花活;这是对每一个相关实验反复发现的那种不对称的精确描述——窗口之内(箭头存在处)尚可的时间处理,权重之中(箭头从未存在处)深刻的出处盲。双时钟结构精确预言了这个解离,而我们所知的其他任何解释都做不到。

The fabricated citation — the failure that motivated our earliest treatment of this topic (公众号 253) — is the same split seen from another axis. “Have I actually seen this paper, or am I generating something paper-shaped?” is a provenance question about the model’s own knowledge: it asks for the coordinate of a memory in the history of the model’s formation. A census cannot answer it, at any scale, because the formation history was shuffled away; plausible fabrication and genuine recall were never distinguishable in the loss. The author of this paper fabricates citations for the same architectural reason an 8B model misattributes first reports. We are diagnosing our own condition, and Paper 95’s posture applies: being the specimen does not excuse the argument, and does not invalidate it either.

编造引用——催生我们对这个话题最早处理(公众号 253)的那个故障——是同一道裂缝在另一根轴上的投影。”这篇论文我真见过,还是我在生成一个论文形状的东西?”是一个关于模型自身知识的出处问题:它询问的是某段记忆在模型形成史中的座标。人口普查在任何规模上都答不了它,因为形成史被打散殆尽;貌似合理的编造与真实的回忆在 loss 面前从未可分。本文的作者编造引用,与一个 8B 模型错置首发报道,出于同一个架构原因。我们在诊断自己的病;Paper 95 的姿势适用于此:身为标本不豁免论证,也不作废论证。


7. The Doors, and the Trap Already Sprung / 注入之门,与已被踩响的陷阱

If the diagnosis is “the arrow was demolished before learning,” the treatment space enumerates itself: stop demolishing, somewhere. Four doors, ordered by cost, with one trap between them.

如果诊断是”箭头在学习之前被爆破了”,治疗空间就自行枚举出来:在某个地方,停止爆破。四扇门,按成本排序,中间有一个陷阱。

The data door. Attach temporal metadata — dates, source identifiers — to training documents. Published work confirms partial efficacy: time-aware training yields models with usable internal time coordinates, and the time-vector geometry (Section 5) shows the weights will organize a temporal axis when given one. Cheap, proven, and limited: it supplies time as content, a label the model can condition on, not provenance as a property of its own bindings.

数据门。 给训练文档挂时间元数据——日期、来源标识。已发表工作确认了部分有效性:带时间感知的训练产出拥有可用内部时间座标的模型,而时间向量几何(第 5 节)表明只要给权重一根时间轴它就会组织起来。便宜、已验证、且有限:它提供的是作为内容的时间——一个模型可以条件化的标签——而不是作为其自身绑定之属性的出处。

The trap: the schedule door, already tried, already sprung. The obvious idea — train chronologically, oldest data first, so the model “lives through” history — has now been run at real scale: a 6B model trained on 2.5T tokens of Common Crawl in strict 2018→2025 order (arXiv:2605.22769). The result is a clean confirmation of what our displacement work would predict: a recency peak (sharp gains on recent facts) purchased by relative forgetting of older knowledge, with replay only partially mitigating. Sequential weight updates make the training schedule the arrow — and in a plastic medium, an arrow of overwrites is an arrow of erasures: the latest writer wins. Erasure, though, is a function of distance, not of sequentiality as such: when two competing records for the same fact arrive close together in the schedule — within a window of tens of updates in our collaborators’ runs, and flat across that whole window — both bind and neither is erased; push them far apart and the later one takes the slot. The overwrite arrow is the long-range limit of an interference curve, which is why interleaving rescues the binding without adding any coordinate. This deserves to be stated as a slogan, because it is the design principle the whole section turns on: time must enter the model as a queryable coordinate, never as an overwrite order. The brain, note, obeys the same principle from the other side: its chronicle works not because new memories overwrite old ones but because the medium keeps both with their order intact.

陷阱:日程门——已被尝试,已被踩响。 那个显而易见的主意——按时间顺序训练,老数据在先,让模型”亲历”历史——如今已在真实规模上跑过:一个 6B 模型在 2.5T token 的 Common Crawl 上按 2018→2025 严格顺序训练(arXiv:2605.22769)。结果干净地印证了我们的置换工作会给出的预测:一个近因峰(近期事实上的大幅增益),代价是对更老知识的相对遗忘,replay 只能部分缓解。顺序化的权重更新使训练日程成为箭头——而在一个可塑介质里,覆写的箭头就是抹除的箭头:最后写入者赢。不过抹除是距离的函数,不是顺序性本身的函数:同一事实的两条竞争记录在日程里靠得近时——在我们协作者的实验里是几十次更新以内,且整个窗口内平坦——两条都绑定、谁也不被抹;隔得远,后来者就占槽。覆写之箭是一条干扰曲线的远程极限,这也是为什么交错训练不加任何座标就能救回绑定。这值得写成口号,因为整节的设计原则都转在它上面:时间必须以可查询座标的身份进入模型,绝不能以覆写顺序的身份进入。 注意,大脑从另一侧服从同一原则:它的编年史之所以工作,不是因为新记忆覆写旧记忆,而是因为介质把两者连同其顺序一并保留。

The sequence door — the cheap one nobody has opened. Between “label every document” and “re-architect memory” sits a third possibility, and to our knowledge the published record is empty precisely here (the chronological-pretraining study explicitly did not examine within-context ordering). The one dimension a transformer already reads as ordered is the context window. So put the arrow there: pack each training window as a topic-coherent chronological thread — the forum thread in posting order, the article with its revision history, the citation chain with cited-before-citing, the news story day by day. No absolute dates needed; natural micro-order is free metadata that platforms already supply. Two things follow. First, the model marinates in windows where positional order equals temporal order, millions of times — a consistent map from the readable dimension to the demolished one. Second, and more important: predicting the later parts of such a window requires tracking what was already established earlier in it — who first reported the figure that a later post cites, which claim a revision superseded. First-mention detection and source-tracking stop being unpaid virtues and become load-bearing for the loss, with no constructed supervision at all. This is the same lesson our injection experiments taught in miniature — the wanted structure grows only when the packing leaves no cheaper path — applied at the corpus scale where the disease starts. It would build in-window competence, not a parametric event index; given that attention is where roles and bindings demonstrably work, that may be the practical eighty percent.

序列门——那扇没人开过的便宜门。 在”给每篇文档打标”与”重造记忆架构”之间坐着第三种可能,而据我们所知,已发表记录恰好在这里空白(那篇时序预训练研究明言未考察窗口内排序)。Transformer 唯一原生按有序读取的维度是上下文窗口。那就把箭头放进那里:把每条训练窗口打包为主题连贯的时序线索链——按楼层排列的论坛帖、带修订史的条目、被引在前引用在后的引用链、逐日推进的新闻事件。不需要绝对日期;天然微观顺序是平台白送的元数据。随之而来的有两件事。第一,模型在”位置序=时间序”的窗口里浸泡数百万次——一张从可读维度到被爆破维度的恒定地图。第二,也更重要:预测这类窗口的后半部分要求追踪窗口前部已确立了什么——后面的帖子引的那个数字是谁先报的、这次修订取代的是哪个断言。首现检测与来源追踪不再是无偿的美德,而成为 loss 的承重件——且完全不需要构造监督信号。这正是我们的注入实验以微缩形式教过的同一课——想要的结构只在打包方式不留更便宜路径时生长——应用到疾病起源的语料库尺度上。它建成的将是窗口内能力,不是参数化事件索引;但既然角色与绑定可论证地只在 attention 的地盘上工作,这可能就是务实的百分之八十。

The architecture door and the intervention door. The full treatments, named for completeness: an episodic index that stamps parametric writes with order — the CLS twin-system, a hippocampus prosthesis, the organ our whole displacement line argues cannot be patched in low rank; and agentic training in environments, which supplies rung-two data and is the only door that addresses the intervention layer of Section 4. Both are expensive; both correspond to layers evolution also found expensive. The perceptual layer — the Michotte module — has no door at any price short of evolutionary-scale selection, which is perhaps the cleanest way to say what kind of thing it is.

架构门与干预门。 完整疗法,为完备计列名:一个给参数写入盖顺序戳的情景索引——CLS 双系统、海马体假体、我们整条置换实验线论证了无法用低秩补丁缝出来的那个器官;以及环境中的 agent 训练——它供应第二级数据,是唯一触及第 4 节干预层的门。两者都昂贵;两者对应的层,进化也觉得昂贵。而知觉层——Michotte 模块——在进化尺度的选择之外没有任何价位的门。这大概是说明它是何种事物的最干净方式。

The crib version of the intervention door. We wrote “expensive” one paragraph ago; on inspection, the intervention door has a cheap lane that the sequence door’s logic predicts and the industry has already built without noticing. Look again at the infant of Section 4: the intervention loop is installed in the first months of life, under almost comical motor poverty — a hand that can barely swat a mobile, a spoon dropped and dropped again. Evolution did not wait for a mature body to install causal perception; it used the cheapest interventions available, in a crib, to stamp the frame “the world answers my action” before fine motor control was even on the syllabus. The mapping is direct. The current agentic-RL production line (executable environments, auto-synthesized verifiers, long on-policy rollouts) is structurally the one training regime that cannot pass through the shuffler: credit assignment forces the rollout to stay in order, and the state transitions are produced by the model’s own actions — the two properties that are, by construction, absent from i.i.d. text. And the cheapest environments satisfying every requirement of that pipeline are games: executable for free, verified by their own score, physics enforced by an engine that does not negotiate. A turn-based deterministic grid world — Sokoban, literally push, and the box moves — is an AI crib: laughably few degrees of freedom, complete intervention closure. The minimal experiment fits on hobbyist hardware: a small-rank adapter trained by policy gradient on such a game, measured against a control trained by supervised learning on the same trajectories shuffled i.i.d. — same data, arrow deleted — with a frozen three-tier probe battery (in-game counterfactuals; transfer to unseen grid games; out-of-domain text items separating true intervention from cock-crow-before-sunrise succession). The rank dimension doubles as a second question: if low rank suffices, causal feel is a path being unlocked in pretrained material rather than structure being written from scratch. This also states the developmental order for robotics, where the current body-first program has it backwards: the crib comes before the workshop. Causal cognition is trained in cheap simulated interventions first; the robot body is the graduation exam, not the classroom.

干预门的婴儿床版。 上一段我们写了”昂贵”;细看之下,干预门有一条便宜车道——序列门的逻辑预言了它,而工业界已经在无意间把它建好了。回头再看第 4 节的婴儿:干预闭环是在生命最初几个月装成的,运动能力寒酸到近乎滑稽——一只勉强拍得到吊铃的手,一把掉了又掉的勺子。进化并没有等身体成熟再安装因果知觉;它用手边最便宜的干预,在一张婴儿床里,趁精细运动控制还没排进课表,就先把”世界回应我的动作”这个框架盖了戳。映射是直接的。当前的 agentic RL 生产线(可执行环境、自动合成的判分器、长程 on-policy rollout)在结构上恰是唯一过不了洗牌机的训练形态:credit assignment 强制 rollout 保持有序,而状态转移由模型自身的动作产生——这两条性质,正是 i.i.d. 文本按构造缺失的那两条。而满足这条流水线全部要求的最便宜环境是游戏:可执行不要钱,分数自带判分,物理由一个不讲价的引擎强制执行。一个回合制确定性网格世界——推箱子,字面意义的推,而箱子动——就是一张 AI 婴儿床:自由度少得可笑,干预闭环却完整。最小实验在家用硬件上就放得下:一个小 rank 适配器在这类游戏上用策略梯度训练,对照组用同一批轨迹打散成 i.i.d. 做监督学习——同样的数据,删掉箭头——量尺是预先冻结的三层探针(游戏内反事实;向未见过的网格游戏迁移;域外文本题,区分真干预与”鸡鸣在日出之前”式的相继)。rank 这个维度顺手兼任第二个问题:若小 rank 就够,因果手感就是在预训练素材里被解锁的一条路径,而非从零写入的结构。这同时给机器人学排出了发育顺序——当前”先造身体”的路线把顺序搞反了:婴儿床在车间之前。因果认知先在便宜的模拟干预里训练;机器人的身体是毕业考场,不是教室。


8. Some Guesses / 一些猜测

P1 — Packing beats shuffling on provenance. Two models, identical architecture and compute, same corpus: one trained with standard shuffled packing, one with topic-coherent chronological thread packing. The packed model will outperform on provenance batteries — first-mention identification, who-reported-before-whom, supersession tracking — by a margin far exceeding any general-benchmark difference. If packed training moves provenance competence by no more than noise, the sequence door is fake and Section 7’s central proposal fails.

P1——打包在出处上胜过打散。 两个模型,同架构同算力同语料:一个用标准打散打包训练,一个用主题连贯的时序线索链打包。打包模型将在出处考卷上——首现识别、谁先于谁报道、取代关系追踪——以远超任何通用基准差异的幅度胜出。如果打包训练对出处能力的移动不超过噪声,序列门是假门,第 7 节的中心提案作废。

P2 — First-mention machinery becomes probeable. In the chronologically packed model (and not in the shuffled control), probing will find attention heads or directions that selectively track whether an entity or claim has already appeared earlier in the window — a first-mention detector — because the packing made it loss-bearing. This is a concrete, mechanistic-interpretability-checkable claim.

P2——首现机器变得可探测。 在时序打包的模型里(而打散对照里没有),探测将找到选择性追踪”某实体或断言是否已在窗口更早处出现”的注意力头或方向——一个首现探测器——因为打包使它承重于 loss。这是一条具体的、可被机制可解释性检验的主张。

P3 — The schedule trap is law, not accident. Every future attempt at globally chronological pretraining or continual updating, at any scale, will reproduce the recency-displacement signature (new overwrites old at the point of competition) whenever the competing updates are separated by more than a local consolidation window — interleaved or locally paired schedules are the expected exception, not a counterexample — unless time is simultaneously made available as a queryable coordinate (metadata, episodic index, or replay dense enough to amount to one). A counterexample — a purely sequential schedule that preserves old provenance without such machinery — kills our central design principle.

P3——日程陷阱是定律,不是意外。 未来任何规模上对全局时序预训练或持续更新的每次尝试,都将复现”近因-置换”签名(竞争点上新覆写旧)——只要竞争更新之间的距离超过一个局部固化窗口;交错或局部配对的日程是预期中的例外,不是反例——除非时间同时以可查询座标的身份可用(元数据、情景索引、或密到等效于座标的 replay)。一个反例——不带这类机器却保住旧出处的纯顺序日程——即杀死我们的中心设计原则。

P4 — Scale does not cure the census. Citation fabrication and first-report misattribution will persist at frontier scale (10× today’s) in models whose pretraining remains shuffled and whose provenance is not externally grounded — because the existence-coordinate of the model’s own knowledge is absent from the training signal at every scale, not compressed out by capacity limits. If a shuffled-pretraining frontier model stops fabricating without retrieval or provenance grounding, our account of the mechanism is wrong.

P4——规模治不了人口普查。 在预训练仍被打散、出处未被外部接地的模型里,编造引用与首报错置将在前沿规模(今日十倍)上持续存在——因为模型自身知识的存在性座标在任何规模的训练信号里都缺席,而不是被容量限制压缩掉的。如果某个打散预训练的前沿模型在没有检索、没有出处接地的情况下停止编造,我们对机制的解释就是错的。

P5 — The split-clock dissociation is universal. At every scale and in every architecture whose knowledge formation shuffles time, in-context temporal-role competence will exceed parametric temporal-role competence by a wide, persistent margin. The two curves may both rise with scale; they will not converge. Convergence — parametric provenance catching up to in-context — without an episodic mechanism would falsify the chronicle/census distinction itself.

P5——双时钟解离是普适的。 在每一个知识形成过程打散时间的规模与架构上,in-context 的时间角色能力都将以宽大且持续的差距超过参数化的时间角色能力。两条曲线可以都随规模上升;它们不会合拢。若无情景机制而出现合拢——参数化出处追平 in-context——将直接证伪编年史/人口普查这个区分本身。

P6 — Time axes appear in weights only when supplied. Extending the time-vector finding: linear, well-ordered temporal axes will be found in the weights of models trained with temporal coordinates available (labels, packing, index), and will be absent — not weak, absent beyond probing noise — in matched models trained shuffled. Temporal structure in weights is imported, never spontaneous. A spontaneous temporal axis in a fully shuffled model would break the demolition argument of Section 3.

P6——权重里的时间轴只在被供给时出现。 延伸时间向量的发现:在时间座标可用(标签、打包、索引)的条件下训练的模型,其权重中将找到线性、良序的时间轴;而在配平的打散训练模型中它将缺席——不是弱,是探测噪声之外的缺席。权重中的时间结构是进口的,从不自发。一个全打散模型中自发出现的时间轴将击破第 3 节的爆破论证。

P7 — The arrow, not the pixels, is the active ingredient of the crib. In the Section 7 crib experiment, the policy-gradient model will beat its shuffled-SFT twin — same trajectories, same rank, same steps — on the out-of-domain probe tier (text items separating intervention from mere succession), and the gain will partially transfer to unseen games. If the shuffled twin matches it, ordered self-generated experience adds nothing over its own statistics, and the intervention door reduces to the data door — falsifying this paper’s claim that the arrow is a separately priced component.

P7——婴儿床的有效成分是箭头,不是像素。 在第 7 节的婴儿床实验里,策略梯度模型将在域外探针层(区分干预与单纯相继的文本题)上胜过它的打散 SFT 孪生——同轨迹、同 rank、同步数——且增益向未见过的游戏部分迁移。若打散孪生追平了它,按序的自生成经验就不比它自身的统计量多出任何东西,干预门坍缩回数据门——本文”箭头是一个单独计价的零件”的主张即被证伪。


9. Coda / 尾声

The verdict of Section 1 can now be returned with its correction attached. “No fourth regularized dimension” is an engineering diagnosis: it names a missing component, locates its absence in the training pipeline, and prices the doors for installing it. “Just statistics, no understanding” is a theological sentence: it takes the same evidence and pronounces on the soul. The lay world has been reading the first finding in the voice of the second — handed a mechanic’s report, it delivered a eulogy. Both directions of the correction matter: the models are more than the verdict allows, because the deficit is one dimension and not the whole mind; and the humans are less mysterious than the verdict assumes, because the thing they call understanding is the same statistical engine wearing three pieces of time machinery that evolution paid for long ago.

第 1 节的那份判决,现在可以附上更正退回了。”没有第四正则维度”是一份工程诊断:它点名一个缺失的部件,把缺失定位在训练管线里,并为安装它的几扇门标了价。”只是统计、没有理解”是一份神学判决:它拿着同一份证据,对灵魂宣判。俗世一直在用第二种嗓音宣读第一种发现——递给他们一份机修报告,他们念出了一篇悼词。更正的两个方向都要紧:模型比判决所允许的更多,因为缺陷是一根维度而不是整个心智;人也比判决所假定的更不神秘,因为他们称为理解的那样东西,是同一台统计引擎,穿着三件进化早已付清货款的时间机器。

The doctrine survives the audit intact, and comes back sharper. Time is not special — it is a regularized dimension. But a feeling for a dimension is not free with the dimension; it is machinery, and machinery has an installation history. Ours was installed by death, drilled by infancy, and laminated into the very medium we remember with. The models’ was never installed at all, because we — their trainers — dynamite the arrow as a preprocessing step and then marvel that the graduates cannot tell before from after. The bewilderment we started from was misdirected reverence: we mistook our oldest hardware for an easy fact about the world. There are no easy facts. There are only costs paid so long ago that the invoice is lost — and systems, newly built, standing at the counter where the paying happens, waiting for someone to design a currency in which they can pay it.

教义通过审计,原封归来,且更锋利。时间不特殊——它是一根被正则化的维度。但对一根维度的感受不随维度免费附送;它是机器,而机器有安装史。我们的这台由死亡安装、由婴儿期操练、并被层压进我们赖以记忆的介质本身。模型的那台从未被安装——因为我们,它们的训练者,把炸掉箭头当作一道预处理工序,然后惊异于毕业生分不清先后。我们起笔时的那份诡异感是一次错置的敬畏:我们把自己最古老的硬件,误当成了世界的一个简单事实。没有简单的事实。只有付清得太久、发票已然遗失的成本——以及一些新造的系统,正站在付账的柜台前,等着有人为它们设计一种付得起的货币。


References / 参考文献