Created: 2026-05-22 UTC

Training Readiness

Enough for research training. Not enough for clinical biomarker claims.

The program has moved from source discovery into reproducible biomarker/gene-prioritization modeling. The latest local evidence includes a v8 corrected-cohort training run, disjoint set11 holdout checks, paired model comparisons, and a clear next-feature roadmap.

Research model trainingReady

v8 has 186 balanced training rows and reproducible reports.

Clinical validationNot ready

Labels remain proxy panel/control labels, not patient outcome labels.

Current actionFeature expansion

Plain retraining matched v7/v6; structure-derived signal is the next useful test.

Short answer

Can we train a biomarker model now?

Yes, for research triage: the current dataset supports training and comparing candidate-ranking models for NDD-like biomarker/gene prioritization. No, for clinical deployment: we still need patient-level labels, larger independent cohorts, calibration evidence, and regulatory-grade validation.

The correct wording is: research-grade biomarker prioritization is ready for model iteration; clinical-grade biomarker prediction is not validated yet.
Latest model candidate v8

Corrected all10 training with set11 as a fully disjoint holdout.

random forest selected boosted-tree challenger tested 0.38 threshold anchor TLK1 residual miss
Evidence snapshot

v8 numbers that matter

Training rows186

93 positive and 93 control.

Train-holdout overlap0

set11 is fully gene-disjoint.

CV ROC AUC0.9977

3-fold mean for selected model.

Holdout F1 @ 0.380.9333

TN 8 / FP 0 / FN 1 / TP 7.

Holdout recall0.8750

One false negative: TLK1.

Holdout specificity1.0000

No false positives on set11.

Pipeline view

What the system is doing

SourcesOpen Targets, ClinVar, GTEx, PubMed, OpenAlex, Europe PMC, UniProt, RCSB, AlphaFold.
Features22 tabular evidence features today; structure-derived features planned next.
CohortsCorrected all10 training, blocked contaminated controls, fully disjoint set11 holdout.
Modelsv4 baseline, v6/v7 augmented candidates, v8 corrected candidate.
SOURCE --> FEATURE --> COHORT --> MODEL --> HOLDOUT --> REVIEW
Policy0.38 deploy anchor, manual review band, rescue rules kept transparent.
ValidationLOSO checks, paired significance, disjoint holdouts, contamination audits.
FrontendRead-only executive console, model registry, pass explorer, structure atlas.
NextExports, Mol* integration, guarded source canaries, explainability.
Timeline

What has happened across the research passes

1-8
Research foundation
Validated source availability, API surfaces, PTEN evidence, pathway context, and build readiness.
9-11
Data platform
Loaded ETL outputs into Postgres, Neo4j, Qdrant, MinIO, and added a FastAPI service.
12-23
Model baseline
Built feature matrices, trained v1-v4, tuned thresholds, added OT keyword features, and created hybrid rescue policy.
24-34
Cohort hardening
Audited control purity, corrected cohorts, added disjoint stress sets, and trained v5-v7 candidates.
35-43
Source, structure, runtime
Expanded source pulls and validated Docker molecular/3D toolchains including Blender, OpenUSD, OpenMM, MDAnalysis, MDTraj, Biotite, and PDBFixer.
44-45
Direction and v8
Closed Q1-Q16, launched v8 corrected-cohort training, and compared v8 against v7/v6/v4 on set11.
Model readout

Latest comparative interpretation

ComparisonResult at set11 threshold 0.38Interpretation
v8 vs v7Delta F1 0.0000, delta recall 0.0000, McNemar p 1.0000.No classification lift on current holdout.
v8 vs v6Delta F1 0.0000, delta recall 0.0000, McNemar p 1.0000.No classification lift on current holdout.
v8 vs v4Delta F1 +0.0762, delta recall +0.1250, delta accuracy +0.0625.Directional improvement vs older baseline, but low power.
TLK1 boundaryv8 score 0.3367, below the 0.38 positive threshold.Still the main residual set11 miss.
Current workstreams

What we are doing now

Model

Keep v8 as a research candidate, add structure-derived features, preserve v7/v8 paired comparisons.

Cohorts

Continue zone-A boundary validation and collect larger fully disjoint clean cohorts.

Sources

Promote OpenAlex and Europe PMC pulls into guarded nightly canaries with budget/rate controls.

Structure

Bind model outputs to structures, include computed models with pLDDT/PAE/provenance gates.

Frontend

Keep public UI read-only, add export bundles, integrate Mol* for model-linked structure drilldowns.

Governance

Add lightweight explainability per pass and keep clinical disclaimers explicit.

Risk register

What still blocks stronger claims

Proxy labels

Panel/control labels are useful for research triage but not a substitute for patient outcome labels.

Small holdouts

set11 is clean and disjoint, but n=16 is not enough for high-confidence statistical claims.

Boundary positives

TLK1 remains below threshold, showing feature gaps for sparse/hard positives.

External validation

Larger independent cohorts and calibration reports are needed before deployment-grade claims.

Review files

Primary files behind this summary