the contempt atlas

Methodology & honesty note

dataset 0.3.0-wave2 · generated 1970-01-01T00:00:00Z · FIXTURE BUILD (illustrative, not findings)

What this measures — and what it doesn't

The corpus measures contempt expressed in text, not contempt actually felt. The “you overestimate” framing leans on the affective-polarization / perception-gap literature (More in Common's Perception Gap and kin), not on this corpus alone. We pair the two and surface n and confidence intervals wherever a specific number appears. We never report a single unified “America” number — every figure is per-corpus.

The metric: Attribution Asymmetry Index (AAI)

For any contempt term, real-world usage splits into DIRECT (aiming it at a target), ATTRIB (“they call us flyover country”), CLAIM (self-claiming with pride), plus neutral MENTION/META. The AAI is that distribution over MENTION/META-excluded observations. The ATTRIB↔DIRECT boundary is metric-critical and validated hardest.

How the current build was made

For each term we pulled a bounded real sample of Reddit comments (via the PullPush API, a Pushshift successor), kept only matches of the in-scope lexicon, and took a sentence-bounded context window around each. Author handles are dropped at ingest; we keep the subreddit as community context, never identity. Each window was then classified by an LLM for stance and for the group it targets and the group the commenter writes from — both chosen from a closed set of groups, or “unclear.” The observation-level data lives in a local database; the site serves only aggregates from it.

This is real text but an unvalidated classifier: the gold-set protocol that would certify the attribute-versus-aim boundary has not been run yet, so every number here is suggestive, not a finding. Several terms are sampled below the display floor and are marked accordingly.

What the data shows so far

Pooled per term, contempt words look mostly aimed — which cuts against the debunk. But split by direction, the structure flips: contempt that crosses between groups is overwhelmingly aimed, while the attributed and reclaimed usage is something groups say among themselves. The fuller argument, with the live numbers, is in the flagship essay and on the map. We present this straight, including where it contradicts what we set out to show.

Corpora & their biases

Caveats on this build

Reproducibility

The lexicon, stance rubric, few-shot examples, gold set, aggregates, and the pipeline commit are all published. Scope is class / region / lifestyle / political contempt only — racial and ethnic dehumanization slurs are out of scope and never enumerated. Author handles are stripped at ingest; we keep community, never identity.

Downloads

← Back to the perception-gap test