Blog

Introducing Slop Meter

A Chrome extension and a site that mark writing paragraph by paragraph: human-ish, machine-ish, mixed, or can't tell. It runs on your device, and it would rather say nothing than accuse a person.

Zaid Mukaddam ·

In today's fast-paced digital landscape, it's important to note that artificial intelligence plays a pivotal role in shaping the future of work. Moreover, this transformative technology fosters innovation, streamlines workflows, and empowers teams to unlock their full potential. Ultimately, the future looks bright.

  • stock vocabulary · r-001
  • abstract metaphor nouns · r-006
  • significance inflation · r-031
0255075100HumanMachine% machine-shaped

98%

machine-ish

98% sure it's machine-ish.

A paragraph written to be obvious, read by the Standard model. The marked words are the ones that set its rules off.

I read a lot of text that a model wrote, and most days I only find out halfway through. It is the rhythm that gives it away: the tidy three points, the sentence that tells you what the next one is going to say, the last line about the future. So I built Slop Meter. It notices the same things I do and says so in the margin.

It is a Chrome extension and a site. The extension puts a dot beside each paragraph on the pages you read. The site does the same for anything you paste, or any document you drop on it. Both do the scoring on your device, and neither one sends your text anywhere.

Four answers

Every paragraph gets one of four answers: human-ish, machine-ish, mixed, or can't tell. Mixed means a person and a model both had a hand in it, which is how a draft looks after someone asks a chatbot to tidy it up. The answers end in -ish because this is a judgement about style. It can't tell you who wrote something, and it doesn't pretend to.

  • 2550750HumanMachine100

    human-ish

  • 2550750HumanMachine100

    mixed

  • 2550750HumanMachine100

    can't tell

  • 2550750HumanMachine100

    machine-ish

The needle shows how machine-shaped a paragraph is. The answer is a separate decision: a needle at 68 is a lean, and the meter calls it can't tell.

Behind each answer are odds, and I've tried to keep them honest. When the meter reports 90%, it is right about nine times in ten on text it has never seen. Hover a dot and a card opens with those odds and the three rules that pushed hardest. The words that set each rule off are marked in the paragraph, so you can argue with it.

A blog post with a small dot in the margin beside each paragraph. One dot is orange, and beside it a card reads machine-ish, 98%, pushed by significance inflation, stock vocabulary and lexical diversity.
The extension on a test page. Grey dots are can't tell, orange is machine-ish. The card names the rules behind the orange one and marks the words that set them off.

Why it mostly says can't tell

On 29,221 paragraphs kept out of training, it gives an answer on 10% of them. On the rest it says can't tell. That is on purpose.

A detector can be wrong in two directions, and they don't cost the same. Missing a machine paragraph takes nothing from you. Calling a person's writing machine-made is an accusation, and an accusation from a tool tends to stick. So the meter speaks only when its top guess clears a bar, and two rules set the bars. Its calls have to be right at least 96 times in 100. And it may call at most 1 human paragraph in 1,000 machine-ish. A model that can't hold both rules doesn't ship.

When it does speak, it is right 98% of the time, and it calls 0.17% of human paragraphs machine-ish. The harder check is the open web: 24,000 paragraphs from pages published in 2019, before chatbots, on sites the model never trained on. People wrote every one of them, so any machine-ish call there is a mistake. It makes that mistake on 0.26% of them.

of paragraphs get an answer
10%
of those answers are right
98%
of 24,000 human web paragraphs called machine-ish
0.26%
Held-out paragraphs and the 2019 web check, at the bars the standard model ships with.

It runs on your device

The model is 56 KB. It reads 176 numbers from each paragraph: one for each rule it can measure, and the rest plain statistics such as sentence length and how often words repeat. It runs on WebGPU where the browser has it and on the CPU where it doesn't, in well under a millisecond a paragraph. There is no server in the loop because a server would have nothing to do.

Two things do leave your device, and you have to switch both on first. Rewrite sends the one paragraph you picked to a server, and a model redoes it without the tells. Hit You're wrong and it sends the 176 numbers plus your answer. Never the text.

Two models

Standard is the model described so far. Sharper reading adds a small language model, SmolLM2 at 135 million parameters, which reads the same paragraph on your machine and reports how surprising it found each word. Machine text is text a language model finds unsurprising. Eight extra numbers from that reading let the meter answer on 25% of paragraphs instead of 10%, and it is right about as often. The price is a 125 MB download, once. The models page sets the two side by side.

MeasureStandardSharper
Says can't tell89.8%75.2%
Right when it makes a call97.7%97.3%
Human web pages called machine-ish0.26%0.24%
Downloadnone125 MB, once
Both models on the same held-out paragraphs and the same web check.

How it's built

Standard never sees your words as words. A paragraph goes in and 176 numbers come out. There is a score for each of the 36 rules a program can measure, the rate of 107 function words like the, of and which, and a few dozen counts: punctuation, how long the sentences run, whether you used a contraction. Measuring takes about a tenth of a millisecond.

Those numbers go into a network with three layers, 176 → 224 → 64 → 3. That is 54,243 weights, and I store them as 8-bit integers, which is how the whole thing fits in 56 KB. Rounding them moves no probability by more than 0.02. The network itself is one page of TypeScript, and a WGSL shader does the same sums on WebGPU.

  1. 1 · Measure

    176

    numbers from each paragraph: 36 rule scores, the rate of 107 function words, and 33 counts of punctuation, sentence length and habits like contractions.

  2. 2 · Network

    176 → 224 → 64 → 3

    Three layers and 54,243 weights, stored as 8-bit integers.

  3. 3 · Odds

    A calibration step turns the three outputs into odds that match how often it is right.

  4. 4 · Verdict

    97%

    It says machine-ish only above this, and human-ish or mixed above 95%. Below that, can't tell.

Standard, end to end. Sharper is the same picture with eight more numbers going in.

It learned from 148,291 paragraphs, and a retrain takes a few minutes on my laptop. I fit five networks from different starting points and ship one: whichever catches the most text from current models while flagging no more than 1 person's paragraph in 1,000. A last step bends the raw outputs until a 90% really is right nine times in ten.

Sharper keeps all of that and widens the input by eight, to 184 → 224 → 64 → 3. It is three of those networks, not one. I train eight, keep the three that catch the most on text they never trained on, and average what they say. The eight extra numbers come from SmolLM2 reading the first 256 tokens: how likely it found each word, how open the choice was, where the word ranked, how much that varied. Two of them are the statistics from the Fast-DetectGPT and Binoculars papers. For every word the language model puts out 49,152 numbers, so a shader boils them down on the GPU and 16 bytes a word come back. Before I wrote that shader, Sharper held about 2 GB in a Safari tab. Now it holds under 600 MB.

The rulebook

Every mark traces back to rules you can read. The rulebook has 38 of them, from stock vocabulary to the paragraph that ends by restating itself, and each comes with an example and the plain version. I use it as a style guide more often than as a detector. Most of the tells are habits of weak writing, and the models learned them from us.

Try it

Paste something and watch the needle, or add it to Chrome. The code, the rulebook, the training scripts and the dated reports are open source under AGPL-3.0.