Benchmark center
Four benchmark questions, four independent routes.
The same-context audit, external aggregate transfer, G0–G7 generalization contract and VCC 2026 readiness audit are separated so their denominators and claim boundaries cannot drift into one another.
Current model result: The model slightly improves average prediction error inside the HepG2 dataset, but it does not reliably recover response direction and does not transfer to an external perturbation dataset.
Unified benchmark matrix
Executed results are separated from future tasks
Open the exact benchmark execution matrix
| Task | Evaluation unit | Current data | Executed | Current result |
|---|---|---|---|---|
| G0 · same-context random holdout | Held-out HepG2 targets | Available | Yes | Small RMSE gain; strict gate 0/16 |
| G1 · unseen perturbation | 55 external target aggregates | Diagnostic only | Partial | No directional-transfer support |
| G2–G5 · family, donor, state, laboratory | Fully held-out context | Absent or incompatible | No | Not assessed |
| G6 · unseen disease line | Independent DMD model | No perturbation outcomes | No | Not assessed |
| G7 · unseen modality / combinations | KO, OE or observed doubles | Calibration references only | No | Not assessed |
Core HepG2 benchmark
Repeated-fold error, raw direction, residual structure and the 16 frozen strict outcomes.
Open core benchmark → 02 · External aggregateExternal-transfer diagnostic
Formula-aligned aggregate metrics on 55 Frangieh targets; no directional-transfer support.
Open external transfer → 03 · OOD matrixG0–G7 generalization ladder
Unseen perturbation, family, donor, state, laboratory, disease and modality are reported independently.
Open generalization ladder → 04 · Future-task auditVCC 2026 readiness
Infrastructure gap matrix, official-source watch and task-binding locks before launch.
Open VCC readiness →