# NMD-VCell v2.1 primary-model card addendum

Model identifier: `safe_external_v1`  
Status: primary same-HepG2 audit model  
Release role: bounded L2 internal-generalization evidence; not a therapeutic virtual-cell model

## Intended use

The model tests whether response-independent DMD-prior and curated priority features contain a reproducible signal for target-level HepG2 perturbation-response summaries. It is designed for evidence auditing, baseline comparison and experiment triage. It must not be used to predict muscle/DMD response, mechanism, efficacy, safety or clinical benefit.

## Data and inferential unit

- Source task: processed HepG2 Perturb-seq object derived from GSE264667.
- Eligible denominator: 2,160 perturbation targets with at least 20 cells.
- Response space: 2,000 highly variable response genes.
- Inferential unit: perturbation target; cells estimate target-level pseudobulk deltas and are not treated as independent biological replicates.

## Cross-fitting design

- Five non-overlapping outer folds of 432 targets.
- Every eligible target receives exactly one primary out-of-fold prediction.
- Fold assignment is balanced on prespecified cell-support and external-feature-availability strata.
- Imputation, scaling and ridge-penalty selection occur within each outer training fold.
- Candidate ridge penalties: 0.1, 1, 10, 100 and 1000, selected by training-only inner cross-validation.

The deterministic 20260717 realization is the frozen primary run. Its uncertainty is conditional on this fold construction; repeated balanced-fold stability remains a required sensitivity analysis.

## Features

- 32 included response-independent DMD-prior and priority-v1 features.
- DepMap variables are excluded.
- Dataset-wide AnnData outcome summaries are excluded.
- The downloadable feature-provenance table records all 63 audited candidate fields, inclusion state, source group, exclusion reason and held-out-response leakage risk.

## Endpoints and controls

Primary endpoint families are:

1. RMSE difference versus zero delta.
2. RMSE difference versus the outer-training mean response.
3. Raw-cosine difference versus the outer-training mean response.
4. Perturbation-specific residual cosine after subtracting the training common response.

Three negative controls preserve the fold-specific selected ridge penalty:

- feature-row shuffle;
- response-row shuffle;
- outcome-coordinate permutation.

Paired target-level bootstrap intervals use 10,000 replicates; paired sign-flip tests use 50,000 permutations; Holm correction is applied within model and endpoint family.

## Primary results

| Quantity | Estimate |
|---|---:|
| Mean model RMSE | 0.117485 |
| Mean zero-delta RMSE | 0.123799 |
| Mean train-mean RMSE | 0.117965 |
| Relative RMSE improvement vs zero | 5.10% |
| Relative RMSE improvement vs train mean | 0.41% |
| Raw-cosine difference vs train mean | 0.001637 |
| Raw-cosine bootstrap 95% CI | 0.000395 to 0.002806 |
| Raw-cosine Holm-adjusted P | 0.00356 |
| Perturbation-specific residual cosine | 0.04544 |
| Residual-cosine bootstrap 95% CI | 0.02958 to 0.06122 |

The result is statistically detectable but practically modest. It supports a small full-OOF same-dataset signal, not strong gene-specific directional prediction.

## Reliability and coverage boundaries

- Primary reliability tertiles contain 720 targets each.
- Reliability strongly tracks RMSE versus zero and residual cosine.
- Reliability does not significantly moderate RMSE improvement versus train mean (Holm P=0.239) or raw-cosine gain (Holm P=0.353).
- scGPT coverage is 232/2,160; GEARS coverage is 2,143/2,160.
- GEARS coverage is enriched for audited DMD-prior and DepMap-availability fields; native coverage is therefore not assumed exchangeable.

## Known limitations

1. One processed HepG2 dataset; no independent biological replication or prospective lockbox.
2. The raw-direction result changed inferential status between an initial unbalanced implementation and the corrected balanced implementation.
3. Target-level bootstrap intervals do not by themselves include fold-assignment/training variability.
4. Guide-pair reliability is exploratory because `sgID_AB` does not identify individual guide effects and only 133 targets are non-missing.
5. Historical Ridge/GEARS/scGPT comparisons use overlapping repeated splits and remain secondary records.
6. The model is not validated in human muscle or DMD perturbation transcriptomes.

## Reproducibility objects

The v2.1 public-resource candidate exposes unique target-level OOF scores, outer-fold assignments, inner-CV records, model summaries, primary tests, negative-control gates, feature provenance, reliability assignments/effects, coverage-selection tables, independent verification results and figure source data through SQLite, TSV and JSON. Full response-vector NPZ objects remain server-side analysis artifacts and are not required to reproduce the reported aggregate effects.
