当 AI 击败人类最好的预测者:《经济学人》双语精读

当 AI 击败人类最好的预测者:《经济学人》双语精读

2026 年 9 月 5 日,人类智能的又一块领地失守了。在一项名为 Metaculus Cup 的预测竞赛中,AI 首次夺冠——而且不止第一名,第二名和第五名也是 AI,人类只拿到第三和第四。

本文精读《经济学人》2026-09-19 Science & Technology 栏目文章 Artificial Intelligence Now Beats Some Of The Best Human Forecasters,拆解 AI 预测者凭什么叫板"超级预测者"、它们的独门技巧是什么,以及为什么最终的边界问题反而变得更清晰了。


1、一场竞赛:AI 首次夺冠

On September 5th yet another domain of human intelligence fell to the artificial variety. For the first time an AI won the seasonal Metaculus Cup, a proving ground for forecasters of a wide variety of events.

Not only that, other AIs took the second and fifth places, leaving third and fourth for humans.

9 月 5 日,人类智能的又一块领地失守于人工智能。在一项名为 Metaculus Cup 的季度预测竞赛中,AI 首次夺冠——该赛事是各类事件预测者的试炼场。不仅如此,其他 AI 还拿下了第二和第五名,只把第三和第四名留给人类。

本段数据:赛事日期 9 月 5 日;AI 包揽第 1、2、5 名。

要点:开篇用排名这种最直观的方式宣布结果,不做任何铺垫。"人类智能的又一块领地失守"这句定性,为后文留了一个待兑现的伏笔——它到底失守得彻底吗?


2、赛制:预测什么,怎么计分

Hundreds of entrants had predicted the outcomes of questions that would be resolved by September.

The participants were scored on the distance of their prediction from the true answer. Those that were closest for longest won the greatest number of points.

数百名参赛者预测了一批将在 9 月前揭晓答案的问题:某个美国州或欧盟国家会限制数据中心发展吗?汉坦病毒疫情会影响至少五名非 MV Hondius 号乘客的人吗?布伦特原油将值多少钱?参赛者按其预测与真实答案的距离计分——在越长时间里越接近答案的人,得分越高。

要点:机制层说明。这里暗含了预测竞赛的核心设计:不是猜对一次,而是长期靠近真相,这就排除了运气型选手。


3、奖金与真实世界的差距

By betting on questions on which its AI thinks the market is wrong, FutureSearch has seen a 6% increase since June on the initial $100,000 value of its portfolio on Kalshi.

...had turned 35into35 into 1.94m over seven months, the sixth-best return in Kalshi's history.

人类参赛者按得分比例分享 5000 美元现金奖。但这一数额实际上还不如第五名在现实世界的成绩。通过押注那些其 AI 认为市场判断有误的问题,FutureSearch 在预测市场平台 Kalshi 上 10 万美元的初始组合自 6 月以来增长了 6%。不过,与 Preseen 背后某位开发者的表现相比,这仍是小钱——尽管他的公司未能进入竞赛前五,他本人却在七个月里把 35 美元变成了 194 万美元,是 Kalshi 史上第六高的回报。

本段数据:竞赛奖池 5,000∗∗;FutureSearch组合∗∗5,000**;FutureSearch 组合 **100,000 → +6%(自 6 月);个人战绩 35→35 → 194 万(七个月,Kalshi 史第六)。

要点:作者用"竞赛奖金 vs 真实收益"的巨大落差说明:预测能力一旦可交易,价值就被市场重新定价。这是全文从学术赛事转向商业现实的关键一跃。


4、AI 预测者为何是"另一个物种"

Instead of being trained specifically on data regarding the questions they are predicting the answers to, they are built using the same large language models (LLMs) that underpin the rest of the AI boom.

These bots therefore read the news and make sense of data in much the same way that human forecasters manage—except they do so far more broadly and swiftly.

计算机早已用于天气预报等狭窄领域的预测。但像 Metaculus 上的那些 AI 预测者是另一个物种。它们并非针对所要预测的问题专门训练,而是用支撑整场 AI 热潮的同一批**大语言模型(LLM)**构建的。

因此,这些机器人读新闻、理解数据的方式与人类预测者颇为相似——只不过范围广得多、速度快得多。这使它们能明确解释预测背后的推理过程,而在说服决策者认真对待其判断时,这是很有用的一项特质。

要点:全文最重要的机制区分——不是"专门训练一个预测模型",而是"让通用模型去读世界"。这解释了为什么它能跨界、能解释推理,也解释了它的局限(依赖公开信息)。


5、基线:人类"超级预测者"有多强

...found that superforecasters—the best in the field—could discriminate between events that would and would not happen 300 days in the future as reliably as regular forecasters could manage those 60 days away.

An analysis in July... suggested AI systems have now reached parity with the superforecasters.

从占星师到赛马预测师,预言者贯穿人类历史。但现代预测业附带数字,更容易剔除骗子和无望者。

宾夕法尼亚大学的芭芭拉·梅勒斯及同事 2015 年发表于《心理科学展望》的研究发现,超级预测者——该领域最顶尖的人——对 300 天后事件的判断可靠度,相当于普通预测者对 60 天后事件的判断。梅勒斯担任科学顾问的预测研究所今年 7 月的一项分析表明,AI 系统如今已在不断演进的一组预测问题上与超级预测者持平。

本段数据:超级预测者 300 天 ≈ 普通预测者 60 天;AI 达到持平的结论发布于 2026 年 7 月。

要点:给出标尺。没有这个基线,"AI 赢了比赛"就不构成有意义的判断——作者先量化人类最好水平,再宣布追平。


6、驱动力与"额外的技巧"

A lot of this improvement has been driven by the same thing that is driving AI's progress in other domains: a handful of companies spending huge sums to scale up LLMs.

One is to combine LLMs from different firms to investigate different parts of a forecasting question, assemble the necessary data and debate among themselves the correct response.

这种进步很大程度上由与其他领域相同的因素驱动:少数公司斥巨资扩大 LLM 规模。然而,最好的系统是由初创公司和钻研者用一些额外技巧构建的。其中之一是组合不同公司的 LLM,分别研究预测问题的不同部分、汇集必要数据,并在彼此之间辩论正确答案。

要点:机制层——通用能力来自巨头,但优势来自工程组合。这里点出了一个反直觉的结构:护城河不在模型本身,而在编排方式。


7、两种人类做不到的技巧

Mantic... gives its AI esoteric datasets to which the publicly available AI models... do not have easy access, because they are behind paywalls.

And the startups can do something that no human forecaster can manage: wipe their bot's memory.

英国初创公司 Mantic 为其 AI 提供一些深奥的数据集——OpenAI、Anthropic 等公司的公开 AI 模型不易获取,因为它们处在付费墙之后。而 Preseen 则试图给新闻评论员的过往战绩打分,以决定把哪些人纳入其预测。

这些初创公司还能做一件人类预测者做不到的事:抹掉机器人的记忆。通过给 AI 提供过去数据的快照,运营者可以观察指令与信息的变化本会如何改变对过去事件的预测。这让他们能在同样的问题上反复重跑流程、反复从错误中学习,想跑多少次就跑多少次。

要点:分层揭示竞争优势——数据(付费墙内)、信号筛选(评论员战绩)、以及最独特的一项:可回测的记忆。这是全文最有技术含量的洞见。


8、反高潮:冠军是个业余爱好者

The actual winner of the latest Metaculus Cup was a bot developed by Jeffrey Liang, a self-described polymath who lives in Texas.

He spent, by his reckoning, less than 150 hours and a couple of thousand dollars... He beat four startups that have raised more than $15m between them.

尽管如此,这些花哨设计是否就是 AI 胜过人类预测者的决定性因素,仍不清楚。最近这届 Metaculus Cup 的实际冠军,是住在德克萨斯州、自称博学者的 Jeffrey Liang 开发的机器人。据他估算,他投入了不足 150 小时和几千美元的算力与数据。他击败了四家合计融资超过 1500 万美元的初创公司,而他的机器人目前正领先于 Metaculus 举办的另一场纯 AI 竞赛——该赛事奖池为 5 万美元。

本段数据:冠军投入 <150 小时、约 数千美元;对手合计融资 >1500万∗∗;另一赛事奖池∗∗1500 万**;另一赛事奖池 **50,000。

要点:全文最有力的反转。它质疑了前几节铺陈的"技巧决定论"——个人用几百小时就能打赢千万美元团队,说明这个领域还没有形成稳定的竞争壁垒。


9、人类仍占优的地方

The human forecasters who enter Metaculus's competitions are not necessarily the world's best. And the cup, which has a four-month time horizon, does not test the ability to forecast over periods of years.

Mr Riviere, for example, says that he treats his firm's AI as if it were another professional forecaster helping him improve his predictions.

参加 Metaculus 竞赛的人类预测者未必是世界最强的。而且该杯赛仅有四个月的时间跨度,无法检验跨越数年的预测能力——而这在许多机构看来至关重要。

Mantic 的预测者 Yann Riviere 认为,在那种领域,人类判断仍有优势,尽管称职的 AI 预测者出现的时间还不长,无法严格检验这一点。而且人与机器的界限也并非泾渭分明。人类预测者早已用 AI 协助完成预测所需的大量研究。Riviere 说,他把自己公司的 AI 当作另一位专业预测者,帮助他改进预测。

本段数据:杯赛时间跨度 4 个月。

要点:作者谨慎地给结论加上限定——赛事本身不代表全部预测能力。并且用一线从业者的态度说明:现实中人机是协作关系,而非替代关系。


10、成本革命与最终的边界

A forecast from human superforecasters can cost more than $10,000 and take a week. FutureSearch asks for ten minutes and a few dollars.

If AIs continue to improve they may thus, counterintuitively, reveal how much of the future is truly unknowable and how much merely so far unknown.

随着 AI 持续进步,这种平衡很可能会转变。人类将保留在现实世界收集信息的能力,而 AI 日益增长的处理能力将使其能综合比人类多得多的输入。

与此同时,AI 让每个人都能更容易地预测未来。一次人类超级预测者的预测可能要价超过 1 万美元、耗时一周;FutureSearch 只要求十分钟和几美元——尽管其他向对冲基金和政府出售服务的初创公司无疑收费更高。

这一切都留下一个开放问题:人类与机器的预测者距离"关于未来可知之物"的理论极限还有多远。例如,由于大气固有的混沌特性,天气预报在约 15 天后就变得随机。如果 AI 持续进步,它们或许反而会——反直觉地——揭示出未来的哪些部分真正不可知、哪些只是迄今未知。

本段数据:人类超级预测 >$10,000 / 一周;FutureSearch 十分钟 / 几美元;天气预报在约 15 天后失去意义。

要点:结尾把商业意义与哲学意义并置——成本下降让预测民主化,而技术逼近极限时,反而能划清"不可知"与"尚未知"的边界。这是全文最高的一层抬升,且落在一个可检验的类比(天气预报的 15 天)上。


词汇表

表达 释义
fell to the artificial variety 失守于人工智能(fall to = 被攻占)
proving ground 试炼场、试验场
seasonal Metaculus Cup 季度性的 Metaculus 杯赛
be resolved by September 在 9 月前揭晓答案(resolve 用于问题"有结果")
scored on the distance of their prediction from the true answer 按预测与真实答案的距离计分
chump change 小钱、微不足道的数目(口语)
turn $35 into $1.94m 把 35 美元变成 194 万美元
a different breed 另一个物种、完全不同的一类
underpin 支撑、构成基础
make sense of data 理解、解读数据
Soothsayers 预言者、占卜者(古雅用词,带讽刺)
weed out the charlatans and no-hopers 剔除骗子和无望者
superforecasters 超级预测者(预测研究领域的专称)
discriminate between 区分、辨别
reached parity with 与……持平、达到同等水平
esoteric datasets 深奥冷僻的数据集
behind paywalls 在付费墙之后
track records 过往战绩、历史成绩
wipe their bot's memory 抹掉机器人的记忆(此处指可回测的状态重置)
bells and whistles 花哨的附加功能
a self-described polymath 自称博学者的人
by his reckoning 据他估算
counterintuitively 反直觉地
chaos inherent in the atmosphere 大气固有的混沌

可借用句式

1. 用排名直接宣布结果,不做铺垫

For the first time an AI won the seasonal Metaculus Cup... other AIs took the second and fifth places, leaving third and fourth for humans.

用法:For the first time X did Y + 名次并置,一句话完成事实与冲击力的双重送达。科技报道开篇的常见手法。

2. 用"竞赛奖 vs 真实收益"制造落差

The human participants shared a $5,000 cash prize... But this sum is actually less than the fifth-placer had achieved in the real world.

用法:But this sum is less than... 把两个量级并置,说明真实世界的定价远高于学术评价。写能力变现类题材很有效。

3. 先否定"专用训练",再给出通用路径

Instead of being trained specifically on data regarding the questions..., they are built using the same large language models that underpin the rest of the AI boom.

用法:Instead of X, they are built using Y 精准划出技术路线的分界。解释一类新系统时,先说它"不是什么"往往比说"是什么"更清楚。

4. 用基线数字让"追平"变得可衡量

superforecasters... could discriminate... 300 days in the future as reliably as regular forecasters could manage those 60 days away.

用法:A 能做到 X 天,相当于 B 的 Y 天 给出可比较的换算。没有基线,进步幅度就无法判断。

5. 用反高潮段落质疑前文铺陈

Even so, it is unclear whether these bells and whistles are decisive... The actual winner... spent less than 150 hours and a couple of thousand dollars... He beat four startups that have raised more than $15m.

用法:Even so, it is unclear whether... 先承认不确定,再抛出反例(个人 vs 千万美元团队)。这种自我修正会让整篇报道的可信度显著提升。

6. 结尾用可行类比界定"极限"

Weather forecasts, for example, become random about 15 days out because of the chaos inherent in the atmosphere. If AIs continue to improve they may thus... reveal how much of the future is truly unknowable.

用法:for example, ... because of ... If X continues, it may reveal ... 用一个物理上已被接受的极限类比新领域,把宏大问题落到实处。


阅读地图

节点 内容
事件 9 月 5 日 AI 首次赢得 Metaculus Cup,包揽 1、2、5 名
赛制 数百人预测 9 月前揭晓的问题,按与真值的距离计分
现实落差 竞赛奖池 5,000 美元 vs 个人七个月 $35→194 万
技术分野 不用专用训练,而是用支撑 AI 热潮的通用 LLM
人类基线 超级预测者 300 天 ≈ 普通预测者 60 天;AI 已追平
独门技巧 组合多家 LLM 互相辩论、付费墙内数据、可擦除的记忆
反高潮 冠军是个人开发者,<150 小时、数千美元击败融资 >1500 万的团队
人类优势 赛事仅 4 个月跨度,长期预测仍靠人类判断;现实中是协作
成本 人类超级预测 >$10,000/一周 vs AI 十分钟/几美元
边界 类比天气预报的 15 天极限:技术进步反能划清"不可知"与"尚未知"

原文:Artificial Intelligence Now Beats Some Of The Best Human Forecasters, The Economist, 2026-09-19, Science & Technology 栏。英文引句为各段关键原句,中文为对照翻译与要点整理。