How the tool works, what the scores mean, and answers to common criticisms.
A tool that measures where written content sits on five independent dimensions of political-economic ideology. It does not measure bias (which implies a correct center), quality, or accuracy. It measures position — and shows you the evidence behind every score.
Ownership (who owns things), Power Concentration (how much authority the state exercises over those subject to it), Traditionalism (does the piece treat social hierarchies as constructed or natural), Governance Scope (national or transnational framing), and Democratic Accountability (is governance accountable to the governed). Each is scored independently from -4 to +4.
Because "left" and "right" collapse five different questions into one word. Reagan and Pinochet agreed on ownership but disagreed completely on state power. Sanders and Stalin agreed on ownership but disagreed completely on state power. One axis cannot capture that. Five can.
-4 is one theoretical extreme, +4 is the other, 0 is genuinely contested. Each level is defined by a governance reality profile — a description of what the world looks like at that position. The definitions come from political philosophy, not from current media norms.
No. Bias implies a correct center. We do not claim one exists. We measure position on defined axes. An article at +2 on ownership is not "biased right" — it is at a specific, documented position on a specific dimension, with evidence you can examine.
Through structured analytical lenses that examine the article at multiple levels — from surface-level framing down to what is missing entirely. Each lens produces observations mapped to specific evidence in the text. Those observations are then processed through scoring models to produce a score on each axis.
Five independent methods for converting observations into axis scores: Tripwire (which detectors fired), Elimination (which positions are impossible), Accumulation (weighing the evidence), Fingerprint (matching against known governance realities), and Debate (arguing both sides). Each shows its work differently. Users can run multiple models and compare.
Because different articles have different signal profiles. An article with heavy explicit political content benefits from a different approach than a movie review with mostly implicit framing. The models also serve as cross-checks — when multiple models converge on the same score, confidence is high.
Yes. A large language model reads the article through structured analytical lenses and produces observations. But the LLM does not produce the final scores — the scoring models do. The AI observes. The models score. This separation means the observations can be re-scored if the methodology changes, without re-reading the article.
Yes. Every score comes with an evidence chain: the score, the rubric level it matched, paraphrased exhibits from the article, and the observations that produced them. If you disagree with a score, the chain tells you exactly where your disagreement sits.
Every score has an evidence chain you can follow back to the article text. The tool does not claim objectivity. It claims auditability — every link in the chain is available for you to examine. If you disagree, you can trace the chain to find the specific observation or rubric match where your interpretation diverges.
The AI extracts observations. The scoring models process them. Both can introduce variance. The mitigation is the same as for human analysts: specify the rubric clearly enough that the choice of analyst (or model) matters less than the rubric specificity. Inter-rater reliability testing validates this.
Real-world calibration points — specific policy actions whose positions have been determined from governance records. The ACA, Glass-Steagall, Attlee's NHS, PATCO — each is a verifiable data point on the spectrum. They're what make "+1 on ownership" mean the same thing across every analysis. Without them, scores would be incommensurate. Every anchor is sourced from primary documents — legislation, executive orders, budget data — not from summaries or impressions.
We do not select — we filter. For each presidential administration since 1900, we score every piece of legislation that created or eliminated a durable institutional change lasting 10 or more years. An agency, a program, a legal framework, a right. If it meets that threshold, it goes in regardless of where we think it will score. The criterion is durability, not ideology.
International anchors fill calibration regions that no US administration reaches — for example, the NHS calibrates what -3 on healthcare ownership actually looks like in governance, because no US president has governed there.
The library covers every major policy domain across US presidents since FDR, plus key international governance actions for spectrum calibration. Each anchor is a specific policy action scored from primary source documents — legislation, executive orders, court opinions — and independently verified. The library grows continuously as we add administrations and domains.
No. All scores are unvetted drafts. They become more confident through the challenge process — if someone challenges a placement with evidence and the placement survives, it earns verified status. Scores can also be revised when new evidence emerges. Both the original and revised scores are preserved with full documentation.
The anchor is revised. Every analysis that referenced it is flagged. Users can recalibrate against the updated library. Both the original and recalibrated scores are preserved. Nothing is ever silently updated. The version history is permanent.
You are right. This tool does not measure everything. It measures one thing — how power is distributed or consolidated — across five dimensions, in written media coverage. That is a defined scope, not a theory of everything.
A piece of music criticism has aesthetic dimensions the tool does not touch. A sports column has drama and narrative it does not score. Not everything is about power, and the tool does not claim it is. When an article does not engage these questions, the tool says so. Low salience is a valid finding. "No signal" is a valid output.
That said, the pieces that seem furthest from power — the tech profile, the lifestyle piece, the education feature — are often where the most embedded assumptions live. They do not think of themselves as political. That is exactly what makes their assumptions worth surfacing.
The tool does not investigate topics. It analyzes how articles about any topic are framed — whose perspective organizes the story, what is assumed without being argued, what alternatives are absent, and where the framing sits on five defined axes.
Climate and energy policy engage primarily the ownership axis (who owns energy infrastructure, who profits from fossil fuels), the public/private boundary (public utilities vs. private energy companies), and democratic accountability (did the public choose this energy policy or did industry lobbying determine it).
The tool examines whether a climate article treats fossil fuel companies as "the energy sector" without naming public utilities as an alternative. Whether it frames transition costs as a burden on consumers without asking who profited from the status quo. Whether it proposes market solutions without considering public investment or community-owned renewables. Each of these is a measurable observation on a defined axis.
Racism is not absent from the framework. It is built into the Traditionalism axis, which measures whether a piece treats social hierarchies — including racial ones — as constructed and changeable or as natural and fixed. It also surfaces across every other axis when racial dynamics are structurally relevant.
The tool examines whether a piece about crime treats criminal behavior as individual moral failure without naming structural relationships between criminalization and race/class. Whether a piece about education ignores funding formulas tied to property values that perpetuate segregation. Whether vocabulary carries racial coding. It surfaces racism not as a cultural attitude but as a structural feature of specific policy domains, measurable in how articles frame those domains.
The framework produces surprising findings on every side. Nixon governed to the left of Obama on healthcare, regulation, and labor. Carter, a Democrat, started the deregulation era Reagan gets credit for. Sanders, "the radical socialist," is center-left social democracy — not anything close to actual socialism. Reagan, the "free market" champion, was practically protectionist on trade. The tool flags left-coded framing in Jacobin with the same rigor it flags right-coded framing in the WSJ.
The ownership axis is defined by a standard — worker ownership on one end, pure private ownership on the other — that predates American partisan politics by a century. If someone disagrees with a placement, they can challenge the axis definition, the specific evidence, or the anchor calibration. Each of those has a documented answer.
We agree. That is why we do not use one number. Every article gets five independent axis scores — a piece can be left on ownership and right on state power. Each axis also tracks how centrally the piece engages with that dimension, and whether the position is argued explicitly or assumed implicitly. The score card shows five scores. The evidence chain shows the reasoning. No single number anywhere.
Yes. We are building a formal challenge submission system where anyone can contest a placement with specific evidence. If you cite real governance records or scholarship that argues for a different placement, it triggers a review. If your evidence is stronger, the score changes. Until the formal system is live, submit challenges by email to challenge@projectoverton.org with the anchor name, the axis, your proposed score, and your evidence.
A free account is browse-only. You can read all analyzed articles on the homepage, access all published reports, and explore the full methodology documentation. You cannot run your own analyses on the free tier. Analyzing new articles requires a Pro subscription.
Pro ($15/month) is the full product. It includes 100 analyses per month, 4 scoring models (Elimination, Tripwire, Accumulation, Fingerprint), full evidence chains on every analysis, the complete Anchor Library with 80 or more verified governance records, evidence files with primary source links, and exports. Three-band profiles (speech, platform, governance) are coming soon and will be included with Pro when available.
Platform ($50/month) is for researchers, journalists, and analysts who want to build and analyze their own datasets using the same tools the project uses internally. It includes 500 analyses per month, batch article upload, all 10 analysis types including Tilt/Shift and corpus-level analysis, the Debate scoring model, a query interface across scored datasets, custom evidence files, and API access. Pro users consume individual analyses. Platform users produce datasets at scale and then run analytical tools on those datasets.
Enterprise is for institutions that need multi-seat access, custom measurement axes, custom analytical weighting, or institutional workspaces. It is not a separate tier on the pricing page. Contact us for details.