M Validation
Leave-one-out cross-state validation — two rounds, June & July 2026
On this page: Overview By state By question Flagged items Honest limitations Corrections found

1. Overview — two rounds

Production-mode leave-one-out validation (4 chains × 1,000 tune + 1,000 draw) across all 28 CJ policy questions confirmed present in all 6 fielded states. Each state is held out in turn; the model is fit on the remaining five and used to predict the held-out state's MrP estimate. That estimate is compared to the direct per-state MrP (from local respondents) and its 95% credible interval. The validation ran twice, differing only in how the held-out state's poststratification frame was built:

Round 1 run date: 2026-06-03 / 2026-06-04  ·  Round 2 run date: July 2026  ·  Comparison set: up to 168 per round (28 QIDs × 6 states)  ·  K-AI-19 clean holdout fix + K-AI-20 binarizer fix applied to both.

Round 1 — respondent frame (168 of 168 comparisons)

Metric Result Threshold Pass?
LOO-CI empirical coverage 87.5% (147/168) ≥80% ✓✓
MAE vs direct MrP 2.38pp ≤5pp ✓✓
Max error (any QID/state) 7.8pp <8pp
Questions exceeding 8pp band 0 of 28 0 ✓✓

Round 2 — Census (ACS) frame, the unsurveyed-state configuration (160 of 168 comparisons)

160 of 168 possible comparisons produced estimates; 8 fits were discarded by the per-prediction convergence gates (minimum bulk and tail ESS) and are excluded rather than imputed — the discarded fits were the life-without-parole item in three folds, and one fold each for a bail-process item, a domestic-violence item, a parole item, a second bail item, and a fines item. Coverage and error are therefore computed conditional on a converged fit — the same conditioning production applies, since an unconverged fit publishes nothing.

Metric Result Threshold Pass?
LOO-CI empirical coverage 89.4% (143/160) ≥80% ✓✓
MAE vs direct MrP 2.1pp ≤5pp ✓✓
Max error (any QID/state) 9.6pp (one comparison) see §4
Comparisons beyond 8pp band 1 of 160 (0.6%) see §4

The unsurveyed-state configuration performs at parity with — on these runs, marginally better than — the respondent-frame baseline (89.4% vs 87.5% coverage; 2.1pp vs 2.38pp MAE). This is the expected result if the ACS frame and the respondent frame are two adequate measurements of the same demographic composition, and it is the direct evidence that estimates for an unfielded state hold at the stated precision, for this question domain, within the stated caveats (§5).

2. Results by state

Round 1 — respondent frame

State LOO-CI coverage MAE (pp) Max error (pp) Note
North Carolina 28/28 (100%) 1.37 3.5 Perfect
Virginia 27/28 (96%) 2.55 5.2
Massachusetts 26/28 (93%) 1.71 5.5
New Jersey 26/28 (93%) 2.62 5.1
Louisiana 24/28 (86%) 2.94 6.3 First clean holdout run
Oklahoma 16/28 (57%) 3.10 7.8 Consistent underperformer — see §5

Round 1, excluding Oklahoma: 131/140 = 93.6% LOO-CI coverage, MAE 2.24pp.

Round 2 — Census (ACS) frame

State N MAE (pp) LOO-CI coverage Max error (pp)
Virginia 28 1.4 28/28 (100%) 4.6
Massachusetts 26 1.5 26/26 (100%) 3.6
North Carolina 26 1.5 26/26 (100%) 3.8
New Jersey 26 1.6 26/26 (100%) 2.6
Louisiana 27 2.4 26/27 (96%) 5.2
Oklahoma 27 4.1 11/27 (41%) 9.6

Round 2, excluding Oklahoma: 132/133 = 99.2% LOO-CI coverage, MAE 1.7pp. Per-state N varies because 8 fits were discarded by the convergence gates (see §1).

For reference, Round 2 predictions were also scored against each state's raw (unmodeled) survey percentages: mean absolute error 3.5pp. The model's predictions sit closer to the poststratified direct estimates than to raw rates, as expected — both are estimates of the same population quantity.

3. Results by question (Round 1 detail)

Per-question detail below is from the Round 1 run. Round 2 per-comparison data (328 rows across both rounds) is retained and available on request; it is not broken out per-question here.

Question ID LOO-CI MAE (pp) Max (pp) Note
CJ-BAIL1 6/6 1.42 2.2
CJ-BAIL2 6/6 2.03 3.5
CJ-BAIL3 4/6 2.50 5.8 LA + OK — see §4
CJ-BAIL4 5/6 2.98 6.3
CJ-CAND1 5/6 3.03 7.8 OK miss
CJ-CLEMENCY1 6/6 1.25 2.9
CJ-CONDITIONS1 6/6 2.22 4.7
CJ-DETER1 5/6 2.78 5.8
CJ-DETER2 6/6 2.95 5.2
CJ-DP1 6/6 3.15 4.8
CJ-DV1 5/6 2.88 4.5
CJ-DV2 5/6 2.13 5.4
CJ-DV3 5/6 2.17 5.5
CJ-DV5 6/6 1.88 3.5
CJ-FAMILY1 6/6 2.07 3.1
CJ-FINES1 5/6 2.50 5.3
CJ-FINES2 5/6 2.05 5.1
CJ-FINES3 5/6 2.20 3.5 Fixed K-AI-20
CJ-FINES4 5/6 2.28 4.5
CJ-JUV1 4/6 3.48 5.1 NJ + VA — see §4
CJ-LWOP1 5/6 1.92 4.3
CJ-MAND1 5/6 3.28 5.4
CJ-PAROLE1 5/6 2.35 3.9
CJ-PLEA1 5/6 1.85 4.1
CJ-PROMISE1 6/6 2.77 4.6
CJ-PROP1 5/6 1.87 4.9
CJ-PROS1 5/6 2.27 5.5
CJ-REVIEW1 5/6 2.47 3.8

4. Flagged items — root-cause analysis

CJ-BAIL3 — 4/6, genuine state policy-context effect

"When deciding whether someone who has been arrested but not yet convicted should be released before trial, judges can follow fixed bail schedules or use their discretion to consider individual circumstances."

Misses: Louisiana and Oklahoma — in opposite directions. Louisiana respondents favor judicial discretion at 57% (raw); Oklahoma at 72% (raw). Both are genuine state-level effects shaped by each state's bail reform history and legal culture, not by demographic composition. The cross-state model anchors on the ~63% modal rate and cannot infer either outlier. See §5 for the honest limitation.

CJ-JUV1 — 4/6, two-cluster state structure

"Some people believe young people should be treated the same as adults in the justice system. Others believe young people deserve different consideration because they are still developing."

Misses: New Jersey and Virginia, both by under 1pp of CI width. The six states split into two clusters: Louisiana, Massachusetts, and Oklahoma show raw support of 62–67%; North Carolina, New Jersey, and Virginia cluster at ~59%. NC passes because its LOO confidence interval is 1pp wider — all three lower-cluster states have essentially identical underlying support. The cross-state model anchors on the higher cluster when those states dominate the training pool. See §5.

CJ-CAND1 × Oklahoma — Round 2's one out-of-band comparison

Round 2's single comparison beyond the 8pp band: the candidate-preference ballot-test item in the Oklahoma fold — direct 68.9%, predicted 59.3%, error 9.6pp. The same state × item pair produced Round 1's largest error (7.8pp). The consistent, two-round pattern indicates that candidate-preference responses are associated with state-specific political context that demographic composition does not carry. This item now carries a standing lower-confidence caveat in the production system, joining CJ-BAIL3 and CJ-JUV1 (flagged in Round 1 on the same grounds). A single 0.6% band exceedance, on the known-atypical state, on an item class already understood to be context-driven, is within the validated envelope given the caveat mechanism — not a silent pass.

CJ-FINES3 — resolved (K-AI-20)

Initially flagged at 0/5 with 10.5pp max error. Root cause: a binarizer substring bug — "agree" in "disagree" evaluated to True, coding unfavorable responses as favorable and inflating predictions by ~8–10pp. Fixed with a word-boundary regex match. Post-fix result: 5/6, 2.20pp MAE, 3.5pp max. The validation process identified its own coding error; the corrected results are presented throughout this document.

5. Honest limitations

The model works when opinion tracks demographics. For most of the 28 validated questions, support for criminal justice reform is primarily explained by age, race, education, and sex. In those cases, the cross-state model generalizes well — LOO errors of 1–3pp, with 87.5% empirical coverage in Round 1 and 89.4% in Round 2. Four of six states validate at 93–100% coverage in Round 1 and 100% in Round 2.

Policy-context questions are harder. When state-level policy culture or regional clustering drives opinion beyond what demographics explain, cross-state inference is less reliable. For these questions, the direct per-state MrP from local respondents is the reliable estimate; cross-state projections carry wider uncertainty. Two questions illustrate the limit: bail discretion (CJ-BAIL3), where Louisiana and Oklahoma diverge by 15pp on a binary question in ways traceable to state bail reform history; and juvenile justice differentiation (CJ-JUV1), where three states cluster at lower support (~57–59% direct MrP) and three cluster higher (62–67% raw) without a clear demographic predictor of which cluster a state belongs to.

Oklahoma, stated plainly. Oklahoma is the persistent weak fold in both rounds — 57% LOO-CI coverage in Round 1, 41% in Round 2 against a nominal 95% — and in Round 2 it accounts for essentially all shortfall. Its point estimates remain acceptable (Round 2 MAE 4.1pp; 26 of 27 within the 8pp band), but for a demographically atypical state the model's credible intervals are too narrow: point estimates remain useful, while the stated interval understates true uncertainty. Excluding Oklahoma, Round 2 coverage is 132/133 (99.2%) with MAE 1.7pp. Production consequences: (a) every modeled estimate is labeled as modeled and carries its interval; (b) this limitation is stated here explicitly rather than presenting the 89.4% aggregate as uniform; (c) expanding the training pool as new states are fielded directly addresses the cause — each new state both retires its own modeled estimates and improves everyone else's. Oklahoma direct MrP estimates (from Oklahoma respondents) are reliable; cross-state projections for Oklahoma carry higher uncertainty and are presented with that caveat.

6. Corrections identified during this run

K-AI-19 — Leaky holdout (2026-06-03). The original 6-QID LOO memo's Louisiana fold was contaminated: Louisiana appeared in training via non-canonical surveys (LA-CJ-2025-001, LA-OMN-2024-001) while only LA-CJ-2025-002 was excluded. Fixed: all surveys mapping to the held-out state are now excluded from training. The original 6-QID memo's 6/6 result for CJ-CAND1 was produced under the contaminated holdout (self-prediction); the corrected result is 5/6.

K-AI-20 — Binarizer word-boundary bug (2026-06-03). The cross-state model's likert binarizer used a naive substring check: favorable_side in response. For questions with favorable_side="agree", this matched "agree" inside "disagree" and "strongly disagree", coding unfavorable responses as favorable. CJ-FINES3 was the only affected question among the 28 validated QIDs. Fixed with a word-boundary regex. Demonstrates that the validation process can identify its own coding errors.

← Technical methodology Technical Methodology Plain-language →