LOO Cross-State Validation
Full results of the leave-one-out validation across 28 CJ policy questions × 6 states, in two rounds: Round 1 (June 2026, respondent-demographics frame) and Round 2 (July 2026, Census-only frame — the exact configuration used for a state we have never surveyed). Companion to the Technical Methodology external validation section. Source: Cross-State Estimation Validation Memo (July 2026).
1. Overview — two rounds
Production-mode leave-one-out validation (4 chains × 1,000 tune + 1,000 draw) across all 28 CJ policy questions confirmed present in all 6 fielded states. Each state is held out in turn; the model is fit on the remaining five and used to predict the held-out state's MrP estimate. That estimate is compared to the direct per-state MrP (from local respondents) and its 95% credible interval. The validation ran twice, differing only in how the held-out state's poststratification frame was built:
- Round 1 (June 2026) — respondent frame. The held-out state's frame is built from its own respondent demographics. This isolates the model-transfer question — can five states' data predict a sixth? — while using the best available description of who lives there.
- Round 2 (July 2026) — Census (ACS) frame. The held-out state's frame is built entirely from U.S. Census ACS data, exactly as it would be for a state we had never fielded in. At prediction time the model knows nothing about the held-out state except what the Census publishes about its demographic composition — an end-to-end rehearsal of the unsurveyed-state configuration. Round 2 is the run cited on the About and Mission pages (89.4% coverage, 2.1pp MAE, 160 comparisons).
Round 1 run date: 2026-06-03 / 2026-06-04 · Round 2 run date: July 2026 · Comparison set: up to 168 per round (28 QIDs × 6 states) · K-AI-19 clean holdout fix + K-AI-20 binarizer fix applied to both.
Round 1 — respondent frame (168 of 168 comparisons)
| Metric | Result | Threshold | Pass? |
|---|---|---|---|
| LOO-CI empirical coverage | 87.5% (147/168) | ≥80% | ✓✓ |
| MAE vs direct MrP | 2.38pp | ≤5pp | ✓✓ |
| Max error (any QID/state) | 7.8pp | <8pp | ✓ |
| Questions exceeding 8pp band | 0 of 28 | 0 | ✓✓ |
Round 2 — Census (ACS) frame, the unsurveyed-state configuration (160 of 168 comparisons)
160 of 168 possible comparisons produced estimates; 8 fits were discarded by the per-prediction convergence gates (minimum bulk and tail ESS) and are excluded rather than imputed — the discarded fits were the life-without-parole item in three folds, and one fold each for a bail-process item, a domestic-violence item, a parole item, a second bail item, and a fines item. Coverage and error are therefore computed conditional on a converged fit — the same conditioning production applies, since an unconverged fit publishes nothing.
| Metric | Result | Threshold | Pass? |
|---|---|---|---|
| LOO-CI empirical coverage | 89.4% (143/160) | ≥80% | ✓✓ |
| MAE vs direct MrP | 2.1pp | ≤5pp | ✓✓ |
| Max error (any QID/state) | 9.6pp (one comparison) | — | see §4 |
| Comparisons beyond 8pp band | 1 of 160 (0.6%) | — | see §4 |
The unsurveyed-state configuration performs at parity with — on these runs, marginally better than — the respondent-frame baseline (89.4% vs 87.5% coverage; 2.1pp vs 2.38pp MAE). This is the expected result if the ACS frame and the respondent frame are two adequate measurements of the same demographic composition, and it is the direct evidence that estimates for an unfielded state hold at the stated precision, for this question domain, within the stated caveats (§5).
2. Results by state
Round 1 — respondent frame
| State | LOO-CI coverage | MAE (pp) | Max error (pp) | Note |
|---|---|---|---|---|
| North Carolina | 28/28 (100%) | 1.37 | 3.5 | Perfect |
| Virginia | 27/28 (96%) | 2.55 | 5.2 | |
| Massachusetts | 26/28 (93%) | 1.71 | 5.5 | |
| New Jersey | 26/28 (93%) | 2.62 | 5.1 | |
| Louisiana | 24/28 (86%) | 2.94 | 6.3 | First clean holdout run |
| Oklahoma | 16/28 (57%) | 3.10 | 7.8 | Consistent underperformer — see §5 |
Round 1, excluding Oklahoma: 131/140 = 93.6% LOO-CI coverage, MAE 2.24pp.
Round 2 — Census (ACS) frame
| State | N | MAE (pp) | LOO-CI coverage | Max error (pp) |
|---|---|---|---|---|
| Virginia | 28 | 1.4 | 28/28 (100%) | 4.6 |
| Massachusetts | 26 | 1.5 | 26/26 (100%) | 3.6 |
| North Carolina | 26 | 1.5 | 26/26 (100%) | 3.8 |
| New Jersey | 26 | 1.6 | 26/26 (100%) | 2.6 |
| Louisiana | 27 | 2.4 | 26/27 (96%) | 5.2 |
| Oklahoma | 27 | 4.1 | 11/27 (41%) | 9.6 |
Round 2, excluding Oklahoma: 132/133 = 99.2% LOO-CI coverage, MAE 1.7pp. Per-state N varies because 8 fits were discarded by the convergence gates (see §1).
For reference, Round 2 predictions were also scored against each state's raw (unmodeled) survey percentages: mean absolute error 3.5pp. The model's predictions sit closer to the poststratified direct estimates than to raw rates, as expected — both are estimates of the same population quantity.
3. Results by question (Round 1 detail)
Per-question detail below is from the Round 1 run. Round 2 per-comparison data (328 rows across both rounds) is retained and available on request; it is not broken out per-question here.
| Question ID | LOO-CI | MAE (pp) | Max (pp) | Note |
|---|---|---|---|---|
| CJ-BAIL1 | 6/6 | 1.42 | 2.2 | |
| CJ-BAIL2 | 6/6 | 2.03 | 3.5 | |
| CJ-BAIL3 | 4/6 | 2.50 | 5.8 | LA + OK — see §4 |
| CJ-BAIL4 | 5/6 | 2.98 | 6.3 | |
| CJ-CAND1 | 5/6 | 3.03 | 7.8 | OK miss |
| CJ-CLEMENCY1 | 6/6 | 1.25 | 2.9 | |
| CJ-CONDITIONS1 | 6/6 | 2.22 | 4.7 | |
| CJ-DETER1 | 5/6 | 2.78 | 5.8 | |
| CJ-DETER2 | 6/6 | 2.95 | 5.2 | |
| CJ-DP1 | 6/6 | 3.15 | 4.8 | |
| CJ-DV1 | 5/6 | 2.88 | 4.5 | |
| CJ-DV2 | 5/6 | 2.13 | 5.4 | |
| CJ-DV3 | 5/6 | 2.17 | 5.5 | |
| CJ-DV5 | 6/6 | 1.88 | 3.5 | |
| CJ-FAMILY1 | 6/6 | 2.07 | 3.1 | |
| CJ-FINES1 | 5/6 | 2.50 | 5.3 | |
| CJ-FINES2 | 5/6 | 2.05 | 5.1 | |
| CJ-FINES3 | 5/6 | 2.20 | 3.5 | Fixed K-AI-20 |
| CJ-FINES4 | 5/6 | 2.28 | 4.5 | |
| CJ-JUV1 | 4/6 | 3.48 | 5.1 | NJ + VA — see §4 |
| CJ-LWOP1 | 5/6 | 1.92 | 4.3 | |
| CJ-MAND1 | 5/6 | 3.28 | 5.4 | |
| CJ-PAROLE1 | 5/6 | 2.35 | 3.9 | |
| CJ-PLEA1 | 5/6 | 1.85 | 4.1 | |
| CJ-PROMISE1 | 6/6 | 2.77 | 4.6 | |
| CJ-PROP1 | 5/6 | 1.87 | 4.9 | |
| CJ-PROS1 | 5/6 | 2.27 | 5.5 | |
| CJ-REVIEW1 | 5/6 | 2.47 | 3.8 |
4. Flagged items — root-cause analysis
CJ-BAIL3 — 4/6, genuine state policy-context effect
"When deciding whether someone who has been arrested but not yet convicted should be released before trial, judges can follow fixed bail schedules or use their discretion to consider individual circumstances."
Misses: Louisiana and Oklahoma — in opposite directions. Louisiana respondents favor judicial discretion at 57% (raw); Oklahoma at 72% (raw). Both are genuine state-level effects shaped by each state's bail reform history and legal culture, not by demographic composition. The cross-state model anchors on the ~63% modal rate and cannot infer either outlier. See §5 for the honest limitation.
CJ-JUV1 — 4/6, two-cluster state structure
"Some people believe young people should be treated the same as adults in the justice system. Others believe young people deserve different consideration because they are still developing."
Misses: New Jersey and Virginia, both by under 1pp of CI width. The six states split into two clusters: Louisiana, Massachusetts, and Oklahoma show raw support of 62–67%; North Carolina, New Jersey, and Virginia cluster at ~59%. NC passes because its LOO confidence interval is 1pp wider — all three lower-cluster states have essentially identical underlying support. The cross-state model anchors on the higher cluster when those states dominate the training pool. See §5.
CJ-CAND1 × Oklahoma — Round 2's one out-of-band comparison
Round 2's single comparison beyond the 8pp band: the candidate-preference ballot-test item in the Oklahoma fold — direct 68.9%, predicted 59.3%, error 9.6pp. The same state × item pair produced Round 1's largest error (7.8pp). The consistent, two-round pattern indicates that candidate-preference responses are associated with state-specific political context that demographic composition does not carry. This item now carries a standing lower-confidence caveat in the production system, joining CJ-BAIL3 and CJ-JUV1 (flagged in Round 1 on the same grounds). A single 0.6% band exceedance, on the known-atypical state, on an item class already understood to be context-driven, is within the validated envelope given the caveat mechanism — not a silent pass.
CJ-FINES3 — resolved (K-AI-20)
Initially flagged at 0/5 with 10.5pp max error. Root cause: a binarizer substring bug — "agree" in "disagree" evaluated to True, coding unfavorable responses as favorable and inflating predictions by ~8–10pp. Fixed with a word-boundary regex match. Post-fix result: 5/6, 2.20pp MAE, 3.5pp max. The validation process identified its own coding error; the corrected results are presented throughout this document.
5. Honest limitations
The model works when opinion tracks demographics. For most of the 28 validated questions, support for criminal justice reform is primarily explained by age, race, education, and sex. In those cases, the cross-state model generalizes well — LOO errors of 1–3pp, with 87.5% empirical coverage in Round 1 and 89.4% in Round 2. Four of six states validate at 93–100% coverage in Round 1 and 100% in Round 2.
Policy-context questions are harder. When state-level policy culture or regional clustering drives opinion beyond what demographics explain, cross-state inference is less reliable. For these questions, the direct per-state MrP from local respondents is the reliable estimate; cross-state projections carry wider uncertainty. Two questions illustrate the limit: bail discretion (CJ-BAIL3), where Louisiana and Oklahoma diverge by 15pp on a binary question in ways traceable to state bail reform history; and juvenile justice differentiation (CJ-JUV1), where three states cluster at lower support (~57–59% direct MrP) and three cluster higher (62–67% raw) without a clear demographic predictor of which cluster a state belongs to.
Oklahoma, stated plainly. Oklahoma is the persistent weak fold in both rounds — 57% LOO-CI coverage in Round 1, 41% in Round 2 against a nominal 95% — and in Round 2 it accounts for essentially all shortfall. Its point estimates remain acceptable (Round 2 MAE 4.1pp; 26 of 27 within the 8pp band), but for a demographically atypical state the model's credible intervals are too narrow: point estimates remain useful, while the stated interval understates true uncertainty. Excluding Oklahoma, Round 2 coverage is 132/133 (99.2%) with MAE 1.7pp. Production consequences: (a) every modeled estimate is labeled as modeled and carries its interval; (b) this limitation is stated here explicitly rather than presenting the 89.4% aggregate as uniform; (c) expanding the training pool as new states are fielded directly addresses the cause — each new state both retires its own modeled estimates and improves everyone else's. Oklahoma direct MrP estimates (from Oklahoma respondents) are reliable; cross-state projections for Oklahoma carry higher uncertainty and are presented with that caveat.
6. Corrections identified during this run
K-AI-19 — Leaky holdout (2026-06-03). The original 6-QID LOO memo's Louisiana fold was contaminated: Louisiana appeared in training via non-canonical surveys (LA-CJ-2025-001, LA-OMN-2024-001) while only LA-CJ-2025-002 was excluded. Fixed: all surveys mapping to the held-out state are now excluded from training. The original 6-QID memo's 6/6 result for CJ-CAND1 was produced under the contaminated holdout (self-prediction); the corrected result is 5/6.
K-AI-20 — Binarizer word-boundary bug (2026-06-03). The cross-state model's likert binarizer used a naive substring check: favorable_side in response. For questions with favorable_side="agree", this matched "agree" inside "disagree" and "strongly disagree", coding unfavorable responses as favorable. CJ-FINES3 was the only affected question among the 28 validated QIDs. Fixed with a word-boundary regex. Demonstrates that the validation process can identify its own coding errors.