Calibration

How often it's right

Measured on 10,269 held-out paragraphs from 11 sources, and on 3,000 paragraphs from web pages on sites it never trained on. The numbers include the places it does worst.

Calibration record

Model 2026-09-19.5 · tested 2026-09-19 · 10,269 held-out paragraphs · bars 0.95, machine-ish 0.93

right when it makes a call
96.3%
of paragraphs get can't tell
90%
of human paragraphs called machine-ish
0.3%
Test points, with the requirement each had to meet
TestMeasuredRequiredResult
Calibration error (ECE)0.05 or less0.0130.05 or lesspass
Top guess right, test setbeat rules alone, 50.8%69.1%beat rules alone, 50.8%pass
Top guess right, held-out attacksbeat rules alone, 5.8%55.6%beat rules alone, 5.8%pass
Separates machine from human (AUROC)0.845
Machine text, top guess machine-ish72.8%
Human text, top guess machine-ish16.6%
Web pages from sites kept out of training, top guess machine-ish20.2%
Largest probability change from 8-bit weights0.010

The three figures above are at the shipped bars. The table's top guesses are before them, when the model has to answer every paragraph.

Web check

On pages it never saw

3,000 paragraphs from 2019 web pages, on sites kept out of training. People wrote all of them, before chatbots, so any machine-ish or mixed call here is a mistake.

called machine-ish or mixed
0.7%
called human-ish
2.3%
can't tell
97.1%

Sharper reading

What the language model adds

Sharper reading is off until you turn it on. Then a small language model, SmolLM2-135M, also reads each paragraph on your device, and a model trained with its numbers makes the call. Same held-out test, same web check, same kind of bars.

The standard model against sharper reading, at the shipped bars
MeasureStandardSharper
Says can't tell89.9%77.2%
Right when it makes a call96.3%96.4%
Human text called machine-ish0.3%0.6%
Web check, called machine-ish or mixed0.7%0.9%
Web check, called human-ish2.3%13.4%

Bars

Why it says can't tell so often

The model always has a top guess. It only shows one when the guess clears the bar for that answer, which is 0.93 for machine-ish and 0.95 for the other two. The bars are set so its calls are right 96% of the time on held-out text. Drag either one to see what a lower bar would cost.

0%25%50%75%100%0.40.50.60.70.80.9Shipped
Right when it makes a callCan't tellHuman text called machine-ish
Machine-ish
0.93
Others
0.95
right when it makes a call
96%
of paragraphs get can't tell
90%
of human paragraphs called machine-ish
0.3%

Reliability

Its odds mean what they say

Group paragraphs by the probability the model gave an answer, then count how often that answer was true. A perfectly calibrated model lands on the dashed line.

005050100100
human-ish
005050100100
mixed
005050100100
machine-ish

Across is the probability given, up is how often it was true, and bigger dots hold more paragraphs. Human-ish and machine-ish stay within a few points of the line. Mixed strays furthest, by up to 10 points in its upper bins, which hold the fewest paragraphs.

Sources

Where it does well, and where it doesn't

The share of paragraphs whose top guess was right, before the bars. Human text polished by a model is the hardest case. It gets called mixed 25% of the time and machine-ish 33%, which still flags the model's hand, while 42% passes as human. Paraphrasing is the attack that hurts most.

Test sources

  • English Wikipedia

    human

    734

    89%

  • Reddit posts, 2006 to 2016

    human

    724

    88%

  • Llama 3.1 70B answers (Magpie)

    machine

    386

    87%

  • Yelp reviews

    human

    704

    87%

  • Recent models, plain prompts

    machine

    1,456

    76%

  • RAID: human documents and 11 older models

    human and machine

    3,161

    72%

  • Web pages, 2019 (C4)

    human

    801

    70%

  • ChatGPT replies to real users (WildChat)

    machine

    447

    69%

  • Recent models told to avoid the tells

    machine

    376

    67%

  • Human paragraphs polished by a model

    mixed

    1,298

    25%

  • Human openings finished by a model

    mixed

    182

    15%

Held-out attacks

  • Held-out models told to avoid the tells

    machine

    691

    67%

  • RAID model text, paraphrased

    machine

    1,240

    50%

Held-out sites

  • Web pages, 2019, from sites kept out of training

    human

    3,000

    76%

Second opinion

A hosted model did worse, so it's out

Jev is a hosted decision model from typesafe.ai. It got the same numbers the local model computes, never the text, and both made calls on the same 599 held-out paragraphs, 60 of them paraphrased machine text, and on the 2,999 paragraphs of the web check.

The local model against Jev on 599 held-out paragraphs and the 2,999 paragraphs of the web check
MeasureLocal modelJev
Says can't tell90.0%0.5%
Right when it makes a call96.7%67.4%
Right when both skip about 30%75.5%74.9%
Calibration error (ECE)0.0100.129
Human text called machine-ish0.0%21.3%
Real pages called machine-ish0.7%21.1%
Time to a callon device461 ms median

Jev almost never says can't tell, and the confidence it states is off from how often it's right by about 12.9 points on average, against 1.0 for the local model. With both skipping about 30% of paragraphs, it was 0.6 points less accurate. The 95% interval runs from 3.4 points worse to 2.3 points better.

On the 2,999 paragraphs of the web check, all written by people, it called 21% machine-ish, against 0.7% for the local model.

So it's out. Neither the extension nor this site uses it. The benchmark stays, so a new version of Jev can be measured the same way.

Data

What it learned from

Labels come from where each paragraph came from, not from annotators. Paragraphs from one source document all land in the same split, so the test never sees a version of something from training.

Paragraphs per split and label
SplithumanmachinemixedTotal
Training21,85021,3887,41650,654
Validation2,9213,0231,0066,950
Test4,4114,3781,48010,269
Held-out attacks1,9311,931
Held-out sites3,0003,000
Human
RAID's human documents (abstracts, books, news, poetry), Reddit posts from 2006 to 2016, Yelp reviews and English Wikipedia.
Machine
RAID text from 11 older models, ChatGPT replies to real users from WildChat, Llama 3.1 70B answers from Magpie, and paragraphs from nine recent models.
Mixed
Human openings finished by a model, and human paragraphs polished by one.
Held-out attacks
RAID's paraphrased model text, and text from four models told to avoid the tells. None of it was used in training.