# VCC 2026 benchmark-readiness audit

Assessed: 2026-07-29

## Verified trigger

Arc Institute has announced that the second Virtual Cell Challenge will launch on 20 August 2026 with a new prediction problem, a wider range of metrics and a USD 100,000 grand prize. The cell type, perturbation modality, training data, hidden-test design and scoring formula have not yet been disclosed.

## Current readiness

- Available: 9/26.
- Partial: 7/26.
- Missing: 10/26.
- Executed extension: formula-aligned aggregate-delta metrics on 55 external Frangieh targets; source result remains NO_DIRECTIONAL_TRANSFER_SUPPORT.
- Remaining ceiling: no donor-, dataset-, disease-stage- or muscle-perturbation truth compatible with the future task.

| layer | check | status | current evidence | next action |
| --- | --- | --- | --- | --- |
| Baseline layer | Zero-change / control-mean baseline | AVAILABLE | Repeated-fold RMSE comparison against zero response is frozen. | Retain as a mandatory comparator in every future task. |
| Baseline layer | Outer-training response mean | AVAILABLE | Leakage-safe train-mean comparator is evaluated in every outer fold. | Keep fold-specific fitting and reporting. |
| Baseline layer | Regularized linear model | AVAILABLE | The current safe external model is a strongly shrunk linear/ridge estimator. | Preserve as the interpretable model floor. |
| Baseline layer | Target-level pseudo-bulk | AVAILABLE | Cells are aggregated to target-level response objects before inference. | Publish the exact aggregation recipe when source-level raw objects are available. |
| Baseline layer | Nearest-neighbour perturbation transfer | MISSING | No frozen nearest-neighbour comparator is registered. | Add training-only perturbation similarity and a leakage-safe neighbour baseline. |
| Baseline layer | Public foundation-model comparator | MISSING | No public model output has been imported under the same split and metric contract. | Evaluate an openly reproducible model only after exact input and output parity is established. |
| Hidden generalization layer | Same-context unseen perturbation | AVAILABLE | Twenty deterministic repeated balanced five-fold realizations hold out entire HepG2 perturbation targets. | Keep this as the lowest generalization tier, not as disease transfer. |
| Hidden generalization layer | Leave-one-donor-out | MISSING | The current direct perturbation substrate does not expose multiple donors. | Acquire donor-resolved muscle/DMD perturbation data and lock donor-disjoint splits. |
| Hidden generalization layer | Leave-one-dataset-out | MISSING | External muscle datasets have different roles and do not form a harmonized perturbation response panel. | Create accession-disjoint training and test objects after modality harmonization. |
| Hidden generalization layer | Leave-one-disease-stage-out | MISSING | No stage-resolved perturbation outcome matrix is available. | Define stage labels and hold out complete disease stages, never random cells. |
| Hidden generalization layer | Cross-cell-type / context transfer | PARTIAL | A 55-target Frangieh external aggregate-delta diagnostic is executed and concludes NO_DIRECTIONAL_TRANSFER_SUPPORT; it is not muscle or DMD truth. | Generate or import matched muscle-context perturbation outcomes. |
| Hidden generalization layer | Perturbation-conditioned time course | PARTIAL | GSE52529 supplies unperturbed 0/24/48/72-hour myogenic reference states only. | Sample the same candidate perturbations across matched differentiation timepoints. |
| Multi-metric layer | Global expression error | AVAILABLE | Repeated-fold RMSE is frozen; MAE-delta is now executed on the 55-target external diagnostic beside identical zero and train-mean comparators. | Rebind MAE only after the official 2026 expression scale and prediction unit are known. |
| Multi-metric layer | Perturbation direction recovery | AVAILABLE | Raw/residual cosine plus external delta Pearson and delta Spearman endpoints are reported; the external directional result remains negative. | Retain the current negative result: raw-direction support is mixed. |
| Multi-metric layer | DES / DEG recovery | MISSING | No VCC-compatible differential-expression score is computed. | Add up/down DEG precision-recall and threshold sensitivity using cell-eval-compatible outputs. |
| Multi-metric layer | PDS / perturbation discrimination | PARTIAL | Raw L1 retrieval, a scale sweep and truth-norm-matched sensitivity are executed on 55 aggregate external targets; this is not an official AnnData or challenge run. | Re-execute with official 2026 AnnData inputs and preserve raw plus scale-audit outputs. |
| Multi-metric layer | Single-cell distribution distance | MISSING | The current model predicts target-level aggregates, not cell distributions. | Gate Wasserstein/MMD or cell-eval distribution metrics on genuine cell-level predictions. |
| Multi-metric layer | Pathway / program recovery | PARTIAL | Frozen pathway panels are descriptive overlaps, not prediction-versus-truth metrics. | Score prespecified program changes against held-out observed responses. |
| Multi-metric layer | Cell-state proportion recovery | MISSING | No generative cell-distribution output or matched proportion truth is available. | Evaluate only when model outputs represent cell-level distributions. |
| Multi-metric layer | Uncertainty and abstention | AVAILABLE | Hierarchical bootstrap intervals, strict gates and an abstention boundary are frozen. | Extend uncertainty to donor, dataset and disease-stage components. |
| Disease closure layer | DMD observational context | AVAILABLE | Four frozen DMD evidence channels expose observed direction and source agreement. | Use as context qualification, never as perturbation truth. |
| Disease closure layer | Human-myoblast perturbation phenotype | PARTIAL | GSE293514 provides bounded KO/fusion evidence for 9 of 21 candidates. | Expand candidate coverage and measure multiple muscle functions. |
| Disease closure layer | DMD-correction transcriptome | PARTIAL | GSE272233 is an orthogonal correction reference; candidates were not perturbed. | Use as a disease-state benchmark, not as candidate causal evidence. |
| Disease closure layer | Prospective prediction registry | PARTIAL | Schemas and an append-only outcome board exist, but contain zero registered outcomes. | Register predictions before assays and bind them to immutable outcome definitions. |
| Disease closure layer | Regeneration, fibrosis, inflammation and vascular endpoints | MISSING | No unified candidate perturbation outcome panel measures these disease functions. | Prespecify one functional primary endpoint and bounded secondary endpoints per study. |
| Disease closure layer | Motor function / drug or clinical response | MISSING | No validated link from current model outputs to clinical outcomes exists. | Keep clinical utility locked until independent prospective validation. |

## Executed aggregate-delta metric layer

MAE-delta, delta Pearson, delta Spearman, effect-size Spearman and raw L1-PDS now run on the frozen 55-target external diagnostic beside the ridge, training-response-mean and zero-change predictions. A raw-scale sweep and truth-norm-matched PDS sensitivity are also frozen. The latter uses observed outcome norms and is therefore diagnostic-only.

- Machine-readable results: `api/v1.1/vcc_metric_diagnostic.json`.
- Per-target table: `downloads/vcc_metric_diagnostic_by_target.tsv`.
- Reproducible runner: `downloads/build_vcc_metric_diagnostic.py`.
- Locked: DES/DEG recovery, DEG AUPRC and cell-distribution metrics because the current object contains aggregate deltas, not single-cell predictions.

## Source-verification matrix

| source | type | verification | grade |
| --- | --- | --- | --- |
| [Arc Institute 2026 Virtual Cell Challenge announcement](https://www.linkedin.com/posts/arc-institute-org_start-assembling-your-team-because-the-2026-activity-7477802662269186048-SXyD) | official_organizer_announcement | VERIFIED_PRIMARY | A_authoritative_event_source |
| [Virtual Cell Challenge 2025 Wrap-Up: Winners and Reflections](https://arcinstitute.org/news/virtual-cell-challenge-2025-wrap-up) | official_organizer_postmortem | VERIFIED_PRIMARY | A_authoritative_benchmark_source |
| [Behind the Data of the Virtual Cell Challenge](https://arcinstitute.org/news/behind-the-data-virtual-cell-challenge) | official_organizer_metric_definition | VERIFIED_PRIMARY | A_authoritative_benchmark_source |
| [Arc Virtual Cell Challenge dataset README](https://github.com/ArcInstitute/arc-virtual-cell-atlas/blob/main/virtual-cell-challenge/README.md) | official_data_repository | VERIFIED_PRIMARY | A_reproducible_repository |
| [cell-eval perturbation model evaluation suite](https://github.com/ArcInstitute/cell-eval) | official_open_source_software | VERIFIED_PRIMARY | A_reproducible_repository |
| [PRiMeFlow: Capturing Complex Expression Heterogeneity in Perturbation Response Modelling](https://arxiv.org/abs/2604.13986) | preprint | VERIFIED_EXISTENCE_PREPRINT | B_preprint_not_peer_reviewed |
| [Lingshu-Cell: A generative cellular world model for transcriptome modeling toward virtual cells](https://arxiv.org/abs/2603.25240) | preprint | VERIFIED_EXISTENCE_PREPRINT | B_preprint_not_peer_reviewed |
| [Benchmarking virtual cell models for in-the-wild perturbation response](https://arxiv.org/abs/2604.27646) | preprint | VERIFIED_EXISTENCE_PREPRINT | B_preprint_not_peer_reviewed |

Official organizer pages and repositories are treated as authoritative for event rules, released data and software behavior. The 2026 model and in-the-wild benchmark papers are preprints: their existence and author-reported findings are verified, but they are not treated as peer-reviewed validation.

## Implementation order

1. Freeze prediction/truth schemas and leakage-safe split contracts.
2. Add nearest-neighbour and open-model comparators beside zero, train-mean and ridge baselines.
3. Rebind the implemented MAE/PDS/delta engine to the official 2026 schema; keep DES/DEG and cell-distribution metrics locked until compatible inputs exist.
4. Add donor-, dataset- and disease-stage-disjoint tasks when matching perturbation truth becomes available.
5. Register candidate predictions prospectively and bind them to muscle/DMD functional outcomes.

## Claim boundary

This readiness matrix is an implementation audit and competition-preparation plan. It does not claim VCC 2026 task compatibility before the rules are released, nor does it convert same-HepG2 benchmark support into muscle, DMD, treatment or clinical validity.

AI-assisted research tools were used to locate and verify public sources; implementation decisions and claim boundaries remain explicitly auditable in the JSON contract.

