Introducing Slop Meter
A Chrome extension and a site that mark writing paragraph by paragraph: human-ish, machine-ish, mixed, or can't tell. It runs on your device, and it would rather say nothing than accuse a person.
Zaid Mukaddam ·
In today's fast-paced digital landscape, it's important to note that artificial intelligence plays a pivotal role in shaping the future of work. Moreover, this transformative technology fosters innovation, streamlines workflows, and empowers teams to unlock their full potential. Ultimately, the future looks bright.
- stock vocabulary · r-001
- abstract metaphor nouns · r-006
- significance inflation · r-031
98%
machine-ish
98% sure it's machine-ish.
I read a lot of text that a model wrote, and most days I only find out halfway through. It is the rhythm that gives it away: the tidy three points, the sentence that tells you what the next one is going to say, the last line about the future. So I built Slop Meter. It notices the same things I do and says so in the margin.
It is a Chrome extension and a site. The extension puts a dot beside each paragraph on the pages you read. The site does the same for anything you paste, or any document you drop on it. Both do the scoring on your device, and neither one sends your text anywhere.
Four answers
Every paragraph gets one of four answers: human-ish, machine-ish, mixed, or can't tell. Mixed means a person and a model both had a hand in it, which is how a draft looks after someone asks a chatbot to tidy it up. The answers end in -ish because this is a judgement about style. It can't tell you who wrote something, and it doesn't pretend to.
human-ish
mixed
can't tell
machine-ish
Behind each answer are odds, and I've tried to keep them honest. When the meter reports 90%, it is right about nine times in ten on text it has never seen. Hover a dot and a card opens with those odds and the three rules that pushed hardest. The words that set each rule off are marked in the paragraph, so you can argue with it.

Why it mostly says can't tell
On 29,221 paragraphs kept out of training, it gives an answer on 10% of them. On the rest it says can't tell. That is on purpose.
A detector can be wrong in two directions, and they don't cost the same. Missing a machine paragraph takes nothing from you. Calling a person's writing machine-made is an accusation, and an accusation from a tool tends to stick. So the meter speaks only when its top guess clears a bar, and two rules set the bars. Its calls have to be right at least 96 times in 100. And it may call at most 1 human paragraph in 1,000 machine-ish. A model that can't hold both rules doesn't ship.
When it does speak, it is right 98% of the time, and it calls 0.17% of human paragraphs machine-ish. The harder check is the open web: 24,000 paragraphs from pages published in 2019, before chatbots, on sites the model never trained on. People wrote every one of them, so any machine-ish call there is a mistake. It makes that mistake on 0.26% of them.
- of paragraphs get an answer
- 10%
- of those answers are right
- 98%
- of 24,000 human web paragraphs called machine-ish
- 0.26%
It runs on your device
The model is 56 KB. It reads 176 numbers from each paragraph: one for each rule it can measure, and the rest plain statistics such as sentence length and how often words repeat. It runs on WebGPU where the browser has it and on the CPU where it doesn't, in well under a millisecond a paragraph. There is no server in the loop because a server would have nothing to do.
Two things do leave your device, and you have to switch both on first. Rewrite sends the one paragraph you picked to a server, and a model redoes it without the tells. Hit You're wrong and it sends the 176 numbers plus your answer. Never the text.
Two models
Standard is the model described so far. Sharper reading adds a small language model, SmolLM2 at 135 million parameters, which reads the same paragraph on your machine and reports how surprising it found each word. Machine text is text a language model finds unsurprising. Eight extra numbers from that reading let the meter answer on 25% of paragraphs instead of 10%, and it is right about as often. The price is a 125 MB download, once. The models page sets the two side by side.
| Measure | Standard | Sharper |
|---|---|---|
| Says can't tell | 89.8% | 75.2% |
| Right when it makes a call | 97.7% | 97.3% |
| Human web pages called machine-ish | 0.26% | 0.24% |
| Download | none | 125 MB, once |
How it's built
Standard never sees your words as words. A paragraph goes in and 176 numbers come out. There is a score for each of the 36 rules a program can measure, the rate of 107 function words like the, of and which, and a few dozen counts: punctuation, how long the sentences run, whether you used a contraction. Measuring takes about a tenth of a millisecond.
Those numbers go into a network with three layers, 176 → 224 → 64 → 3. That is 54,243 weights, and I store them as 8-bit integers, which is how the whole thing fits in 56 KB. Rounding them moves no probability by more than 0.02. The network itself is one page of TypeScript, and a WGSL shader does the same sums on WebGPU.
1 · Measure
176
numbers from each paragraph: 36 rule scores, the rate of 107 function words, and 33 counts of punctuation, sentence length and habits like contractions.
2 · Network
176 → 224 → 64 → 3
Three layers and 54,243 weights, stored as 8-bit integers.
3 · Odds
A calibration step turns the three outputs into odds that match how often it is right.
4 · Verdict
97%
It says machine-ish only above this, and human-ish or mixed above 95%. Below that, can't tell.
It learned from 148,291 paragraphs, and a retrain takes a few minutes on my laptop. I fit five networks from different starting points and ship one: whichever catches the most text from current models while flagging no more than 1 person's paragraph in 1,000. A last step bends the raw outputs until a 90% really is right nine times in ten.
Sharper keeps all of that and widens the input by eight, to 184 → 224 → 64 → 3. It is three of those networks, not one. I train eight, keep the three that catch the most on text they never trained on, and average what they say. The eight extra numbers come from SmolLM2 reading the first 256 tokens: how likely it found each word, how open the choice was, where the word ranked, how much that varied. Two of them are the statistics from the Fast-DetectGPT and Binoculars papers. For every word the language model puts out 49,152 numbers, so a shader boils them down on the GPU and 16 bytes a word come back. Before I wrote that shader, Sharper held about 2 GB in a Safari tab. Now it holds under 600 MB.
The rulebook
Every mark traces back to rules you can read. The rulebook has 38 of them, from stock vocabulary to the paragraph that ends by restating itself, and each comes with an example and the plain version. I use it as a style guide more often than as a detector. Most of the tells are habits of weak writing, and the models learned them from us.
Try it
Paste something and watch the needle, or add it to Chrome. The code, the rulebook, the training scripts and the dated reports are open source under AGPL-3.0.