v8 has 186 balanced training rows and reproducible reports.
Enough for research training. Not enough for clinical biomarker claims.
The program has moved from source discovery into reproducible biomarker/gene-prioritization modeling. The latest local evidence includes a v8 corrected-cohort training run, disjoint set11 holdout checks, paired model comparisons, and a clear next-feature roadmap.
Labels remain proxy panel/control labels, not patient outcome labels.
Plain retraining matched v7/v6; structure-derived signal is the next useful test.
Can we train a biomarker model now?
Yes, for research triage: the current dataset supports training and comparing candidate-ranking models for NDD-like biomarker/gene prioritization. No, for clinical deployment: we still need patient-level labels, larger independent cohorts, calibration evidence, and regulatory-grade validation.
Corrected all10 training with set11 as a fully disjoint holdout.
v8 numbers that matter
93 positive and 93 control.
set11 is fully gene-disjoint.
3-fold mean for selected model.
TN 8 / FP 0 / FN 1 / TP 7.
One false negative: TLK1.
No false positives on set11.
What the system is doing
What has happened across the research passes
Validated source availability, API surfaces, PTEN evidence, pathway context, and build readiness.
Loaded ETL outputs into Postgres, Neo4j, Qdrant, MinIO, and added a FastAPI service.
Built feature matrices, trained v1-v4, tuned thresholds, added OT keyword features, and created hybrid rescue policy.
Audited control purity, corrected cohorts, added disjoint stress sets, and trained v5-v7 candidates.
Expanded source pulls and validated Docker molecular/3D toolchains including Blender, OpenUSD, OpenMM, MDAnalysis, MDTraj, Biotite, and PDBFixer.
Closed Q1-Q16, launched v8 corrected-cohort training, and compared v8 against v7/v6/v4 on set11.
Latest comparative interpretation
| Comparison | Result at set11 threshold 0.38 | Interpretation |
|---|---|---|
| v8 vs v7 | Delta F1 0.0000, delta recall 0.0000, McNemar p 1.0000. | No classification lift on current holdout. |
| v8 vs v6 | Delta F1 0.0000, delta recall 0.0000, McNemar p 1.0000. | No classification lift on current holdout. |
| v8 vs v4 | Delta F1 +0.0762, delta recall +0.1250, delta accuracy +0.0625. | Directional improvement vs older baseline, but low power. |
| TLK1 boundary | v8 score 0.3367, below the 0.38 positive threshold. | Still the main residual set11 miss. |
What we are doing now
Keep v8 as a research candidate, add structure-derived features, preserve v7/v8 paired comparisons.
Continue zone-A boundary validation and collect larger fully disjoint clean cohorts.
Promote OpenAlex and Europe PMC pulls into guarded nightly canaries with budget/rate controls.
Bind model outputs to structures, include computed models with pLDDT/PAE/provenance gates.
Keep public UI read-only, add export bundles, integrate Mol* for model-linked structure drilldowns.
Add lightweight explainability per pass and keep clinical disclaimers explicit.
What still blocks stronger claims
Panel/control labels are useful for research triage but not a substitute for patient outcome labels.
set11 is clean and disjoint, but n=16 is not enough for high-confidence statistical claims.
TLK1 remains below threshold, showing feature gaps for sparse/hard positives.
Larger independent cohorts and calibration reports are needed before deployment-grade claims.