# Focused peer review of the v2.1 evidence-aligned candidate

Review date: 2026-07-17  
Review scope: scientific claim–evidence alignment, statistical interpretation, figure hierarchy, reproducibility and submission readiness.  
Files reviewed: v0.9-v2.1 manuscript, six-figure main package, thirteen-figure supplementary package, figure captions, release lock, package audit and claim–figure audit.

## Overall assessment

**Decision: Major revision before submission.**

The v2.1 revision is a substantial scientific improvement. The primary benchmark now uses all 2,160 eligible targets exactly once out of fold, removes DepMap and dataset-wide outcome summaries from the primary feature set, separates historical analyses from the current estimate, and reports effect magnitude together with uncertainty and negative controls. The revised Figure 1, Figure 3 and Supplementary Figure S13 communicate the central evidence boundary much more honestly than the prior package.

The manuscript is nevertheless not ready to submit. The main reason is no longer figure quality. The manuscript's new central analyses are not yet bound into the locked public database/archive, and the statistical evidence still conditions on one corrected balanced fold assignment even though the raw-direction inference changed across implementations. The historical comparator figure also remains in the main narrative without a v2.1-aligned comparator rerun.

## Scores

| Criterion | Score / 10 | Assessment |
|---|---:|---|
| Originality and resource concept | 8.5 | The evidence-architecture framing is distinctive and appropriate for a database/resource article. |
| Methodological rigor | 7.5 | Unique-target cross-fitting and leakage control are strong upgrades; fold-assignment uncertainty remains incompletely characterized. |
| Evidence sufficiency | 7.5 | The internal HepG2 evidence supports the bounded L1–L2 claim, but the new evidence is not yet part of the released resource authority. |
| Argument coherence | 8.5 | Primary, historical and unsupported claims are now clearly separated. |
| Writing and figure communication | 8.5 | The new figure hierarchy is readable, quantitative and appropriately cautious. |
| **Mean scientific score** | **8.1** | Strong author-review candidate; not yet a locked submission package. |

## Critical issues

None detected in the revised claim wording or in the numerical transcription of the audited primary results. The automated claim–figure audit passes 13/13 checks, and the final PDF visual audit found no remaining clipped or cross-panel labels.

## Major issues

### 1. The v2.1 evidence is not yet the released database authority

The manuscript and figures now describe a v0.9/v2.1 analysis with five balanced 432-target folds, the `safe_external_v1` model and Supplementary Figure S13. The current `RELEASE_LOCK.json`, however, still binds manuscript v0.6, resource v0.5.0 and the v10.2 figure package with twelve supplementary figures. Therefore a reader cannot yet reproduce the manuscript's central v2.1 result from the DOI-intended archive or the locked public database.

Required revision:

1. Add the v2.1 fold assignments, 2,160 target-level OOF predictions, primary metric tests, negative-control results, reliability tables, coverage-selection tables, feature-provenance records and model card to the public SQLite/TSV/JSON resource.
2. Rebuild and independently verify the resource archive and website from those objects.
3. Regenerate checksums, the release manifest, data dictionary and `RELEASE_LOCK.json` so they bind the v0.9 manuscript and the 6-main/13-supplementary figure package.
4. Mint the DOI only after this new authority passes local and browser acceptance tests.

### 2. Inference is conditional on one corrected balanced fold assignment

The balanced five-fold design gives complete target coverage, but the confidence intervals and randomization P values condition on a single deterministic fold assignment and its five trained models. Supplementary Figure S13b shows that raw-direction support changed between the first unbalanced implementation and the corrected balanced implementation. This makes split-construction stability part of the inferential question, not merely a software note.

Required revision:

1. Record the outcome-blind reason, chronology and pre-model variables used for the balancing correction.
2. Repeat balanced, non-overlapping five-fold cross-fitting over a prespecified seed set, with every target predicted once per repeat.
3. Report the distribution of RMSE and raw-cosine effects across repeats, and use an interval or sensitivity summary that reflects fold/training variability in addition to target-level sampling uncertainty.
4. Retain the present deterministic run as the frozen primary realization, but avoid treating its conditional P value as the only measure of robustness.

This does not erase the current modest point estimate; it determines how stable that estimate is to a design choice already shown to affect inference.

### 3. Main Figure 4 is a historical comparator analysis

Figure 4 is explicitly labelled historical in the caption, which is an important correction. However, its position in the main figure sequence can still be read as a current leaderboard, although it uses overlapping historical splits and a historical ridge model rather than the leakage-controlled v2.1 primary model.

Required revision: either rerun same-target comparator analyses against `safe_external_v1` under the v2.1 fold architecture, or move the historical Figure 4 and its detailed comparisons to the supplement. If a main comparator figure is retained, denominator coverage and selection should be visually inseparable from the performance estimates.

## Minor and administrative issues

1. Replace `MAINTAINER_CONTACT_PENDING`, `DOI_PENDING` and all author-controlled authorship, funding, conflict, acknowledgement and licence fields before submission.
2. Remove the author-review version note from the submitted manuscript while retaining it in the internal revision record.
3. Complete normal and private-session browser QA, HTTPS/response-header checks and an independent-user task test.
4. Keep exact software/checkpoint identifiers and the finite-randomization P-value convention visible in the final methods or model cards.

## Resolved in this review cycle

- Rebuilt Figure 1 around the 5 × 432 unique-target primary design and separated the 787 reused / 514 never-tested historical record.
- Rebuilt Figure 3 around full-OOF effect sizes, negative controls, reliability strata and an explicit practical-effect boundary.
- Added Supplementary Figure S13 for target accounting, implementation sensitivity, external-model coverage, coverage selection and reliability definitions.
- Corrected S13 PDF clipping, cross-panel label intrusion and P-value/interval overlap.
- Updated the manuscript from Sites v14 / UI v1.2 to the release-lock fact of Sites v15 / UI v1.3.
- Generated 17-page main and 14-page supplementary PDFs with 6 and 13 embedded figures; all package and claim–figure audits pass.

## Recommended order of work

1. Integrate v2.1 evidence objects into the public resource and rebuild the release authority.
2. Run repeated balanced unique-target cross-fitting to quantify fold-assignment stability.
3. Resolve the status of historical Main Figure 4.
4. Freeze author metadata, browser QA and DOI only after the scientific/resource lock is final.

The paper's defensible contribution is now clear: an auditable evidence architecture that exposes modest internal signal, weak-baseline effects, feature provenance, coverage selection and missing transfer evidence. Submission should wait until the downloadable resource proves the same result as the manuscript.
