Why annotate
Somebody has to sit down, read a real post, and say what it is. Every filter, every report, every takedown starts there.
The basics
Three steps. That's the whole idea.
Somebody says something hateful online.
A trained researcher decides exactly what it is.
That judgement is what software can finally learn from.
How it works
The second pair of eyes is the point. It is the only place a mistake gets caught before everything built on top of it copies the mistake too.
A real post, on a real platform, in front of real people.
A trained researcher reads it — with the whole thread still around it.
A second person checks the call. Nothing skips this step.
The human checkIt joins our library of labelled examples.
Why it never stops
A filter built last January is already missing things by June. Keeping up means going back to people, over and over.
It never stops
1 · Labelled examples
Thousands of judgements made by people.
2 · Detection software
Software learns the patterns those people spotted.
3 · Moderation tools
Platforms use it to find hate at scale.
4 · Hate moves on
Hate changes its words. The software starts missing things.
Thousands of judgements made by people.
Software learns the patterns those people spotted.
Platforms use it to find hate at scale.
Hate changes its words. The software starts missing things.
Finding what the software now misses, and having a person judge it, is the part most projects cannot afford to keep repeating. It is exactly the part we make cheap.
What goes wrong
None of these can be fixed later on. Each one traces straight back to how the examples were made.
Strip a post from its thread and all the software sees is vocabulary. It never learns that the same sentence was a quote, a joke, or someone pushing back.
Quoting hate to condemn it gets you flagged. The real thing slips through.
The most durable hate online is built to be deniable — numbers, nicknames, in-jokes. None of it is a slur, so none of it trips a filter.
The charts show a drop. The hate just changed words.
If a platform can't say why something came down — by whose definition, judged against what — the decision cannot be defended to anyone.
Appeals get won on process, whatever was actually said.
Blunt filters catch people discussing hate along with people spreading it — those documenting abuse, or arguing about where the line sits.
Communities under attack get moderated more than their attackers.
Faster and cheaper
The old way spends most of its money putting back the context it threw away. We never throw it away.
The missing steps aren't faster. They stop existing.
Scrape-then-label
With the extension
Find candidate content
Bulk scrape or buy an API firehose; store everything
Researcher browses the platform normally
Rebuild context
Re-join replies, parents, and media across tables — often impossible after the fact
Captured automatically at the moment of labelling
Export and hand off
Ship a CSV to a labelling vendor or platform; wait weeks
No export step exists
Label
Annotator reads a decontextualised string and guesses
Annotator reads the post as any user would
Quality control
Post-hoc agreement scoring; disputes are unresolvable without context
Lead approves or rejects in a review inbox, context attached
Deletion drift
Posts vanish between scrape and label; rows become unverifiable
Evidence is archived at capture time
3 of 6 stages disappear entirely. Not optimised — removed. The work of rebuilding context, exporting, and chasing deleted posts is never done in the first place.
The scale
Teaching software has become an industry of its own — because careful human judgement has never been in shorter supply.
$6.3B
→ $17.1B by 2030
What the world spends teaching software, mostly by paying people to label things.
Grand View Research$11.6B
→ $26.1B by 2031
What platforms spend deciding what stays up and what comes down.
Mordor Intelligence60–80%
Getting the examples ready eats most of the schedule — before any software is built.
CleverX$0.10–0.50
per label
The going rate for simple labelling. Expert work on coded hate costs far more.
BasicAIThese are outside estimates published in 2025–2026. Research firms measure this market differently and their totals vary a lot, so treat them as a rough sense of scale rather than exact figures.
People still make every call. We took away the busywork around them, so an expert's hour buys much more than it used to.