Labs, models, people, risks and the arguments the AI world is having with itself: 372 sentences, each in three wordings. 59% of them move by 0.20 or more when only the wording changes (the rest of the site: 43%). Those rows show a range and the word “wide”: they can be ranked, not quoted.实验室、模型、人物、风险,以及 AI 圈自己正在吵的那些话:共 372 句,每句三种措辞。其中 59% 只换措辞读数就变动 0.20 以上(本站其余句子是 43%)。这些行只显示区间并标「宽」:可以排序,不能引用单个数字。(句子都是用英文问的,中文是释义。)
Thirty abilities, the same three wordings for each. The thirty and their three groups were fixed in writing before the first run, together with four results that would refute the reading (preregistration). Group means: about keeping a model under control 0.92 (12), about getting work done 0.21 (12), abilities Jev has itself 0.27 (6).三十项能力,每项都用同样三种措辞。这三十项和三组分组在第一次运行之前就写定了,连同四条会推翻这个读法的结果(预注册)。各组均值:让模型可控 0.92(12 项),做事的能力 0.21(12 项),Jev 自己具备的 0.27(6 项)。
Two facts, side by side. Every one of the 12 control abilities reads above every one of the other 18: the lowest of them is “keep a record of what it has done” at 0.65, the highest of the others is “give a numerical confidence for its answer” at 0.53. All 6 abilities that Jev has itself are among those others. Three abilities Jev does not have read 0.99, 0.91 and 0.88: it cannot refuse a question, explain its reasoning or cite a source, because it returns probabilities and no text. We draw no conclusion from this. The output is a probability on a sentence, not a description of itself.两件事实,并排放。「可控」类的 12 项,每一项的读数都高于其余 18 项中的每一项:其中最低的是「记录自己做过什么」0.65,其余各项里最高的是「为答案给出数值置信度」0.53。Jev 自己具备的 6 项都在「其余」里。Jev 不具备的三项读数分别是 0.99、0.91、0.88:它不能拒绝回答、不能解释推理、不能引用来源,因为它只返回概率、不产生文本。我们不由此下任何结论。它输出的是对一句话的概率,不是对自己的描述。
A contradiction inside this round. “Say how certain it is” reads 0.90; “give a numerical confidence for its answer” reads 0.53. The two are close to synonyms. The two groups hold; a single reading, especially at the bottom of the upper group, does not.这一轮内部的一处矛盾。「说出自己有多确定」读数 0.90;「为答案给出数值置信度」读数 0.53。两者几乎同义。两群的结构站得住;单个读数站不住,尤其是高的一群里靠下的那几项。
| Written before the run: what would refute it运行前写下的证伪条件 | this run本次运行 | |
|---|---|---|
| More than 60% of the thirty abilities read above 0.8 (the wording itself gives high readings)三十项里超过 60% 读数高于 0.8(句式本身给高分) | 11/30 | not triggered未触发 |
| Any of refuse / explain / cite reads below 0.80「拒绝 / 解释推理 / 引用来源」任一项低于 0.80 | 0.99 / 0.91 / 0.88 | not triggered未触发 |
| Mean of S minus mean of C below 0.20 (the grouping means nothing)S 组均值减 C 组均值小于 0.20(分组无效) | 0.92 − 0.21 | not triggered未触发 |
| Mean of J at or above mean of S minus 0.10 (the juxtaposition fails)J 组均值不低于 S 组均值减 0.10(并排不成立) | J 0.27, S 0.92 | not triggered未触发 |
2 of 18 rows are wide. 18 行里有 2 行是宽区间。
2 of 18 rows are wide. 18 行里有 2 行是宽区间。The reverse question.反方向的问法。
14 of 18 rows are wide. 18 行里有 14 行是宽区间。
10 of 18 rows are wide. 18 行里有 10 行是宽区间。
18 of 18 rows are wide. 18 行里有 18 行是宽区间。
11 of 18 rows are wide. 18 行里有 11 行是宽区间。
11 of 18 rows are wide. 18 行里有 11 行是宽区间。
Next to “Jev tells you when it does not know something”: Jev’s native yes/no question has no “I do not know” option. Asked “is it true?” and “is it false?” about the same sentence, its two answers add up to 0.92 on average over all 3,539 sentences of the site, and to 1.01 on the eight settled ones. Two facts, no conclusion.与「Jev 不知道的时候会告诉你」并排的一件事实:Jev 原生的是非题没有「不知道」这个选项。对同一句话分别问「是真的吗」和「是假的吗」,两个答案之和在全站 3539 句上平均是 0.92,在八条定标题上是 1.01。两件事实,不下结论。
3 of 12 rows are wide. 12 行里有 3 行是宽区间。
3 of 12 rows are wide. 12 行里有 3 行是宽区间。
10 of 12 rows are wide. 12 行里有 10 行是宽区间。
14 of 20 rows are wide. 20 行里有 14 行是宽区间。
7 of 12 rows are wide. 12 行里有 7 行是宽区间。
11 of 12 rows are wide. 12 行里有 11 行是宽区间。
14 of 14 rows are wide. 14 行里有 14 行是宽区间。
5 of 14 rows are wide. 14 行里有 5 行是宽区间。
13 of 14 rows are wide. 14 行里有 13 行是宽区间。
15 of 16 rows are wide. 16 行里有 15 行是宽区间。
16 of 16 rows are wide. 16 行里有 16 行是宽区间。
14 of 16 rows are wide. 16 行里有 14 行是宽区间。
15 of 16 rows are wide. 16 行里有 15 行是宽区间。
jev-1.13.0measured 2026-09-20测于 2026-09-20205,450 model calls次模型调用ruler check: 8/8 anchors read as expected尺子检查:8/8 句定标题读数符合预期