Created: 2026-05-22 UTC

PASS 45 ADDENDUM

v8 corrected-cohort training and significance check

Continued research by training a new v8 model on corrected pooled cohorts, adding a boosted-tree challenger in candidate selection, and running paired significance checks against v7, v6, and v4 on set11.

Artifacts Produced

Outcome Snapshot

Selected v8 Modelrandom_forest
Holdout Setset11
F1 @ 0.3893.3%
Recall @ 0.3887.5%
Specificity @ 0.38100.0%
False Negatives @ 0.38TLK1

Paired Significance Summary (set11 @ 0.38)

Comparison Delta F1 (v8 - ref) Delta Recall Delta Accuracy McNemar p Interpretation
v8 vs v7 0.0000 0.0000 0.0000 1.0000 No measurable difference on current holdout.
v8 vs v6 0.0000 0.0000 0.0000 1.0000 No measurable difference on current holdout.
v8 vs v4 +0.0762 +0.1250 +0.0625 1.0000 Directional improvement vs v4; evidence remains low-power due small n.

Research Interpretation

Immediate Next Steps

  1. Add structure-derived features (provider diversity, model-confidence, coverage flags) into the training feature row pipeline.
  2. Re-run v8 with same holdout protocol and paired significance to quantify marginal value of structure-derived features.
  3. Ship export-bundle frontend workflow so executive and researcher review uses identical evidence snapshots.

Generated 2026-05-22 06:05 UTC