Research · FAQ
The questions we hear most often: what AddressHate does, how the AI works, how we measure accuracy, how to work with us, and why our approach is different from tools that only flag individual posts.
The essentials, for a first-time visitor.
AddressHate is a research organization that studies hate online. We build our own AI classifiers and train them with social scientists who label the data by hand, so our tools can catch the coded and implicit hate that plain keyword searches walk right past. From there, we follow how that hate moves and grows.
This dashboard is where we open that work up to the public. It's a place to see what's being said right now, how those ideas are spreading, and why it all matters, whether you're a researcher, a policymaker, or a platform trying to get ahead of a problem.
Right now we're building tools to track antisemitism, anti-Black racism, and anti-queer hate.
Over time, we want to help researchers, policymakers, and organizations make sense of every form of prejudice. The methods we're building aren't tied to any one kind of hate, so they carry over as we grow.
Most hate trackers just count slurs. We're more interested in what sits underneath them: the beliefs, the stories, and the coded language that give hate its meaning and keep it moving.
The hate that does the most damage is usually the quiet kind, the kind that never says the ugly word out loud. Watching the wider conversation lets us see which ideas are catching on, and how they change shape, before they reach the mainstream. That way we can follow a whole movement instead of reacting to one post at a time.
They're really the same idea seen at three levels: the belief, the story, and the words.
The subcategory is the belief underneath, and it's old. Antisemitism runs on a pretty fixed set of these: that Jews control the banks, that they're loyal to each other before their own country, that the Holocaust gets exaggerated. Ideas like these have been around for centuries and barely change. What changes is the occasion someone finds to reach for them. (These recurring beliefs are what our taxonomy calls subcategories.)
A narrative is one of those beliefs aimed at something specific that just happened, like an election, a war, or a violent attack. The belief is already there; the event just gives someone a reason to reach for it and offer it as the explanation.
Rhetoric is the wording itself, the form the belief takes once it becomes an actual post: a meme, a joke, an emoji, a question that isn't really a question.
You can see all three in a single line. After a major attack makes the news, someone posts "worth asking who really benefits from this." The belief doing the work underneath is that Jews engineer world events for their own profit. Pointed at this attack, it becomes a hint that the whole thing was staged by someone powerful. And the phrasing, vague and faux-curious, lets the writer plant that idea without ever committing to it. There's no slur and nothing you could pin down as a claim, but it still adds up to a complete antisemitic argument.
Most tools only read the wording, so a post like that slips straight past them. Ours is built to recognize the belief underneath it, which is how it catches the same idea whether it shows up as a slur, a meme, or a polite question.
Coded hate hides in dog whistles, euphemisms, memes, and hints instead of open slurs. That's exactly why it slides past keyword filters and ordinary moderation.
We built our AI to look for that hidden layer. It picks up hateful content from the accounts, phrases, and topics a slur search would never flag, and that's where a lot of today's real harm actually lives.
We collect content that's already public across the major platforms. Right now that's YouTube and X.
We only look at public conversation. We don't go near private accounts, direct messages, or encrypted spaces.
All of it, and it all comes out of the same pipeline. Underneath everything is the structured data: every item we analyze gets tagged by the kind of hate it carries and the specific belief it draws on.
From there we build the trend and narrative views, this dashboard, spike alerts partners can watch, and longer published pieces like the Digital Hate Review. The goal is always to give people something they can act on, not just another pile of flagged posts.
A look at the methods behind the numbers, for researchers, journalists, and platform partners.
We build our models in-house and train them on data our own social scientists and researchers label by hand.
Instead of checking text against a list of slurs, the models learn the accounts, phrases, topics, and rhetorical moves that hate tends to travel with. That's what lets them catch coded and implicit content at scale, the kind a simple word match would miss, and keep up as the volume grows.
Knowing a post is hateful barely tells you what to do about it. What actually helps is knowing which belief it's spreading and how far that belief has already gotten. With that, you can hand a platform more than a takedown list. You can show them where an idea started, where it's headed, and where there's a real chance to stop it.
And these narratives aren't random. They tend to begin in the same corners, the smaller forums and communities where a belief feels at home before it goes anywhere. From there, a handful of accounts do most of the work of carrying it into bigger, more mainstream spaces. They also get harder along the way: a joke turns into a conspiracy theory, and the conspiracy theory eventually becomes a reason to act.
Watching that happen close to real time gives you a lot of openings. A platform can catch a belief slipping out of the fringe and into ordinary comment sections and step in before it reaches millions of people. An advocacy group can answer it with counter-messaging while it's still just an idea and not yet a settled worldview. Someone in policy can notice the same pattern climbing across several platforms in the weeks before a big event, and get ready instead of scrambling afterward.
None of that is possible if all you can see is that a post was offensive. It works because we read the belief sitting underneath.
Our classifiers are built on a structured taxonomy of hate that our experts developed, organized by broad category and then by the specific beliefs within each kind of hate.
Grounding the models in clear, documented categories instead of a loose keyword list keeps our labels consistent over time and from one analyst to the next, lets outside experts check our work, and makes it possible to compare one kind of hate with another.
They're at the center of everything we do. Our social scientists and researchers write the definitions, label the training data, settle the hard calls, and keep an eye on the model's output for drift.
The AI takes their judgment and applies it across millions of items. It doesn't stand in for them, and every accuracy claim we make traces back to their work.
Text is where we're strongest today, and we're working on images, memes, and video too, since so much coded hate now lives in pictures rather than words.
Both. We dig through large historical archives to set baselines and see the long-run trends, and the pipeline keeps pulling in and classifying new content as it appears, so the dashboard stays close to the current conversation.
How we measure ourselves. Figures shown across the dashboard are illustrative; final numbers will reflect live performance at launch.
We test the model against a gold-standard set our expert annotators have labeled, and we report precision, recall, and how often it agrees with expert judgment, not one tidy number.
Because coded hate depends so much on context, the bar we care about most is agreement with our experts. And we're honest about where the model does well and where it still needs a person to weigh in.
Expert agreement is simply how often the model's call matches what our trained annotators agree on. Those annotators are social scientists and subject-matter researchers who know each kind of hate.
We treat their consensus as the truth because coded hate is a judgment call, and it's one that general-purpose labelers tend to get wrong.
We run the same expert-labeled test set through the general-purpose models and compare them with ours side by side.
Off-the-shelf systems are tuned to catch obvious hate and to avoid false alarms, so they routinely miss coded language, dog whistles, and anything that turns on context, which is the exact material ours is built for. It's not that those models are broken. It's that hate detection is a small side feature for them and the whole job for us.
Context and intent are written right into our taxonomy and annotation guidelines, and our experts handle the tricky cases by hand.
We're trying to tell hate apart from criticism, reporting, counter-speech, and satire, including blunt political speech, so that we're measuring hate and not policing opinions. When a case is genuinely a toss-up, we'd rather send it to a person than flag it too quickly.
We work only with content that's already public. We don't touch private accounts, direct messages, or end-to-end-encrypted platforms, partly on principle and partly because our job is to make sense of public conversation, not to watch individual people.
We retrain as we add more expert-labeled data and as new coded terms and slang show up, and they show up fast, so the models have to keep pace. If accuracy starts slipping on our gold-standard set, that triggers an update too.
Who it's for, and how to work with us.
Researchers, policymakers, platforms, journalists, and civil-society groups, anyone who needs solid, structured intelligence on online hate. Really it's for people who have to understand a problem before they can act on it, and who need numbers they can stand behind.
We're working toward giving vetted partners direct API access to our classifications. If that sounds like a fit for your organization, get in touch.
We partner with academic researchers and can open up our data and classifications for approved projects. Reach out and we'll talk through the scope and terms.
Yes. We help reporters cover online hate with data, historical context, and expert analysis, so a story can rest on measured trends instead of a few screenshots.
A spike is where the work starts, not where it ends. From there, partners can dig into the narratives and the specific content behind the surge, brief the people who need to know, shape moderation or policy decisions, and then check whether anything they did actually moved the numbers.
The value is in the why and the what-next, not the alert on its own.
Our pipeline looks at public content across whole platforms rather than individual reports, so we don't have a public submission tool right now. If your organization has a specific research need, reach out to us directly.
Who we are, and how we fit alongside the other people working on this.
It depends a bit on who's asking, but it really comes down to three things.
Next to advocacy and monitoring groups: we measure the whole conversation with a rigorous, taxonomy-grounded, AI-scaled approach, instead of putting out curated incident reports.
Next to for-profit moderation vendors: we're driven by research and open about our methods. We're trying to understand hate, not just delete flagged text for a client.
Next to general-purpose AI companies: our models are built for this one job and trained by experts, so they catch the coded and implicit hate that broad systems are set up to miss.
A lot of hate tracking still runs off a list of banned words and phrases. You match posts against the list, and whatever matches gets counted. It's simple and fast, but it only ever finds the hate that says the quiet part out loud.
The content that does the most damage rarely uses a slur at all. It leans on coded language, a knowing reference, a joke, a leading question, and a word list walks straight past it. Our classifier is built to read the belief underneath a post, so it catches the coded and implicit material a dictionary can't see, and it keeps working as the language shifts.
Many of the landmark findings in this field come from a single study: a researcher hand-collects a few hundred posts, writes them up, and publishes once. Those studies matter, and they've moved platform policy. But each one is a snapshot — true for the week it was gathered, and easy to wave off as too small a sample.
We run the same kind of analysis continuously, across the whole corpus, so a finding isn't a one-time headline. It's a number you can watch move week to week, and when something spikes we see it as it happens instead of months later in a report.
General-purpose models like ChatGPT are genuinely useful, and some groups now point them at this problem. But you're renting someone else's model: it's tuned for broad tasks, it can change under you without notice, and running it across millions of posts gets expensive fast — enough that studies have had to shrink their data to afford it.
We build our own classifier and train it on data our experts label by hand. So we can run the full corpus without a per-post bill, measure exactly where it's strong and weak, and improve the parts that matter most for coded hate, which is exactly where off-the-shelf models are weakest. Every batch of expert labels makes our detection better, and that capability stays ours.
We work closely with academic researchers and institutions. Their expertise shapes our taxonomy and annotation standards and keeps our work rooted in real scholarship on hate and extremism.
The Digital Hate Review is the analysis we publish from time to time on where online hate is heading. It draws on the same data and classifications behind this dashboard, and turns all that measurement into a story that partners, the press, and the public can actually follow.