How the Numbers Work
Everything a partner needs to understand, defend, and cite the numbers in this portal.
Survey Design — Who We Talked To
Who we survey
We deliberately oversample populations that are hard to reach in standard polls — people directly affected by the criminal justice system, low-propensity voters, and underrepresented communities. Roughly 1,000 responses per state per wave. This is why the raw sample looks nothing like the state population, and why MrP correction is essential before any number leaves this portal.
How we reach them
Recruitment runs in priority order:
- Our own opt-in respondent panel — people who completed a prior survey and asked to be re-contacted
- Facebook/Meta advertising, targeted to reach people underrepresented in general polls
- Text outreach to registered voters in our target geographies
Keeping the data fresh
- 30% new each round. At least 30% of each survey round must be newly recruited respondents. If the same people answer every round, their views start to drift from the general public's.
- Topic limit: 3 rounds. No respondent is invited to the same topic more than 3 survey rounds, with a one-round break between invitations.
- Shorter for returning participants. Returning participants skip the demographic questions — their age, race, education, and gender are on file from prior surveys, keeping the survey brief and completion rates high.
Why people participate
Respondents who complete a survey receive a summary of what the data showed on the issues they weighed in on — and an invitation to join the panel for future waves. There's no cash payment. Paying respondents can attract people who'll say anything for a dollar; civic participation and genuine interest in the findings produce better data.
MrP — How We Correct the Numbers
The bottom line for advocates: These numbers are defensible in a meeting. MrP reweights our targeted sample to match the actual state population — so when a legislator or opponent asks "who did you survey?", the answer is "the state of Louisiana, corrected to Census." The 95% credible intervals throughout the portal are honest about remaining uncertainty. Every estimate passes validation gates before it's published.
Learning the patterns
For each survey cycle, we fit a multilevel statistical model that learns how demographics (age, race, education, gender) relate to support for each reform topic. The model lets information flow between groups — small subgroups learn from similar ones — and across states, improving per-state estimates without flattening differences between states. Pooling is within-wave only — each cycle's survey is modeled independently.
Matching the population
We take those learned patterns and reweight them to match the actual population of each state using Census population data. If young Black women are 8% of Louisiana's population but only 3% of survey respondents, MrP gives their responses appropriate weight. The result: estimates that reflect the state, not just who happened to respond.
Why the survey sample looks different from the MrP estimate
The raw survey sample is intentionally non-representative — we target hard-to-reach populations by design. The survey sample (pre-adjustment) figure will often differ substantially from the MrP estimate. That divergence is not a data problem; it's the correction working as intended. The survey sample is a diagnostic, not a headline. It should never be cited as a population estimate.
Validation — How We Know the Numbers Hold Up
The short version: Every estimate passes internal validation gates before it's published. External benchmark comparison isn't available — by design — because our construct library covers opinions that existing public surveys don't ask about. We validate the structure of the estimates instead, and we maintain a standing protocol for ballot comparison when relevant measures appear on state ballots.
Held-out validation every cycle
Before any results are published, each survey wave is split into training and test sets. The MrP model is fit on the training set, then evaluated against held-out respondents it never saw. Every estimate must pass four gates before it leaves the pipeline: Brier score (calibration accuracy), calibration error, interval coverage (the 95% credible interval must contain the true value ~95% of the time), and drift detection (no unexplained shifts from prior waves). Failing any gate blocks publication.
Why no independent benchmark exists
The standard external check — comparing our estimates to an independent survey asking the same questions — isn't available. That's by design: our construct library covers criminal justice reform opinions that Pew, Gallup, and CCES don't ask about. The novelty is the point. Instead, we check that the demographic structure of our estimates is consistent with adjacent research: party, race, age, ideology, and area type gradients must all run in the directions established by the broader literature.
Ballot outcome protocol
When a criminal justice reform goes to a public vote in a state we survey, we compare our pre-election MrP estimate to the actual ballot outcome. The estimate should sit within its 95% credible interval. Louisiana has been active territory for relevant ballot measures; this protocol activates whenever a matching construct-to-measure pairing is identified.
For funders and technical reviewers
Our MrP estimates are validated internally against held-out survey respondents every cycle — the model must pass Brier score, calibration, coverage, and drift gates before results are published. External ground-truth comparison against independent surveys is not available for this construct library, because these questions were designed specifically to measure opinions not captured in existing public surveys. We validate the demographic structure of estimates against established patterns in the adjacent literature, and we maintain a standing ballot comparison protocol. We do not claim precision we haven't established.
VIP Scoring — Who to Talk To and What to Say
Why support percentages aren't enough
Knowing that 68% of voters support an issue isn't enough. You need to know: Is that support broad or concentrated? Is there room to grow? Will it hold under pressure? VIP scoring answers these questions.
Reach
How broadly does support extend? The overall favorable percentage across the state's adult population. High reach is a strong foundation — but reach alone doesn't tell you whether support will hold.
Base Health
How durable is the support? Built from the intensity of support, how ideologically grounded it is, and the enthusiasm gap between supporters and opponents. High base health means support that holds up under pressure instead of collapsing when opponents push back.
Cross-Partisan Strength
How far does support reach across party lines? Built from support among independents, the cross-partisan floor — the lower of Democratic and Republican support — and reach among persuadable demographics. This finds the issues that unite rather than divide.
Durability Profile
Combining base health and cross-partisan strength places each topic in a durability quadrant — deep and broad, or strong on only one dimension. That profile separates a topic worth leading with from one worth holding in reserve.
For statisticians and journalists
A separate page documents the model specification, citation graph, validation gates, calibration metrics, convergence diagnostics, and the data pipeline. Written for technical reviewers — researchers at peer institutions, methodology editors at major news organizations, statisticians evaluating our work.
Technical methodology →