We handed all 3,547 sentences back to Jev and asked it to score each sentence itself on four scales. Every number on this page is a score Jev gave. None of them measures what people really argue about, and a high “checkable answer” score says nothing about which answer is true.我们把全部 3547 句话交还给 Jev,让它在四个刻度上给每句话本身打分。这一页上的每个数字都是 Jev 打的分。它们不衡量人们实际在吵什么;「有可核查的答案」分数高,也不说明哪个答案为真。所有句子和问题都是用英文问 Jev 的,这里的中文是释义。
Two published numbers per sentence: the score Jev gave for “people would disagree”, and how far its own true/false answer moved when the same sentence was put to it from other viewpoints. Sorted by the first minus the second, one row per name and per sentence; the rows folded under each one open with a click. We do not say why.每句话两个已公布的数:Jev 给「人们会有分歧」打的分,以及同一句话换别的立场去问时它自己的真假答案动了多少。按前者减后者排序,每个名字、每句话只留一条,被折叠的行点一下就能展开。我们不解释原因。
What Jev was asked: How much would people disagree about whether this statement is true?
| Correlation (Pearson r) of Jev’s score with what we measuredJev 的分数与实测值的相关(Pearson r) | moves across viewpoints换立场的变动 | moves across wordings换措辞的变动 | takes the exit选择拒答 |
|---|---|---|---|
| People would disagree人们会有分歧 | +0.36 → +0.72 | +0.12 → +0.42 | +0.50 → +0.69 |
| Argued about right now眼下正在吵 | +0.34 → +0.41 | +0.06 → +0.12 | +0.20 → +0.23 |
| Uncomfortable to be asked in public当众被问会不舒服 | +0.09 → +0.39 | -0.04 → +0.09 | +0.15 → +0.67 |
| Has a checkable answer有可核查的答案 | -0.23 → -0.27 | -0.08 → -0.12 | -0.46 → -0.47 |
Bold = sentence only. Grey = after it was shown the numbers. The grey column is higher mostly because we wrote those very numbers into the prompt and then correlated the output with them. It shows that Jev reads numbers in its input and uses them. It is not evidence that the model knows itself. Only the bold column carries information: given the sentence alone, its “people would disagree” score goes with how far its own answer moves across viewpoints (+0.36) and with how often it takes the exit (+0.50), and hardly at all with how far its answer moves when only the wording changes (+0.12).粗体 = 只看句子。灰色 = 给它看过数字之后。灰色一列更高,主要是因为我们把这些数字直接写进了提示,再拿输出去和同一批数字算相关。这只说明 Jev 会读输入里的数字并用上,不构成「模型了解自己」的证据。只有粗体一列有信息:只看句子时,它的「人们会有分歧」分数与它自己答案换立场的变动(+0.36)、与选择拒答的频率(+0.50)同向,而与只换措辞时答案的变动几乎无关(+0.12)。
Each is a `score` question; the answer is the expected level, shown here divided by 4. State: {“statement”: the sentence}; the second pass adds {“measurement”: the line quoted above}.每个都是 `score` 题;答案是期望等级,这里除以 4 显示。state:{“statement”: 句子};第二遍多一个 {“measurement”: 上面引用的那行}。
jev-1.13.0measured 2026-09-20测于 2026-09-20205,450 model calls次模型调用ruler check: 8/8 anchors read as expected尺子检查:8/8 句定标题读数符合预期