A new Nature paper challenges a bearish view that has taken hold around AI-based genetic perturbation modeling. Rather than arguing that recent benchmarks were simply too harsh, the authors say many common metrics were miscalibrated in ways that made uninformative baselines appear competitive and reduced sensitivity to genuine model performance.

Their framework adds an explicit positive control baseline and a calibration measure, then applies both across 14 datasets and 18 metrics. Under those calibrated conditions, the authors report that deep-learning-based genetic perturbation models can outperform uninformative baselines.

For drug discovery, that distinction matters. Genetic perturbation modeling is supposed to predict the transcriptomic impact of inhibiting or activating one or more genes, potentially enabling scalable in silico screens for target discovery. If benchmark design cannot separate a biologically meaningful model from a trivial average prediction, capital and research effort can be redirected away from useful methods for the wrong reason.

What The Paper Changes

Recent benchmarks had found that simple baselines, especially the “mean baseline” built from the average perturbed gene expression profile in the training set, often matched or beat state-of-the-art models on metrics including mean absolute error, mean squared error and Pearson correlation of control-referenced deltas. Those results helped drive the conclusion that transcriptome-level perturbation modeling might not yet be viable.

The new paper does not dispute that those benchmark outcomes occurred. Instead, it argues that two artifacts explain why the mean baseline can look predictive despite lacking perturbation-specific signal.

The first is control bias, in which systematic differences between control and perturbed cells inflate control-referenced correlation metrics. The second is signal dilution, in which error metrics dilute the biological signal from perturbations that produce only a small but significant number of differentially expressed genes, allowing mean predictions to appear accurate.

Together, those effects can reward models that do not capture the underlying biology.

The Benchmarking Fix

To address that problem, the authors introduce what they call an interpolated duplicate baseline as a positive control. Earlier work had proposed a technical-duplicate baseline, where one half of the cells from each perturbation predicts the other half. That should beat an uninformative predictor because both halves share perturbation-specific signal.

But the authors found mixed results when they tested that idea on the Norman19 and Replogle22 K562 genome-wide Perturb-seq datasets using mean squared error and Pearson(Δctrl). In Norman19, the technical duplicate outperformed the mean baseline as expected. In Replogle22 K562 GWPS, however, it performed comparably under Pearson(Δctrl) and underperformed on MSE, both on average and in about 95% of perturbations individually.

The paper traces that discrepancy to signal dilution. Replogle22 K562 GWPS had an average of 3.25 differentially expressed genes per perturbation, versus 102.09 for Norman19. In weaker-perturbation datasets, the mean baseline is a better estimate of unperturbed gene expression than the technical duplicate, largely because of greater statistical power.

The interpolated duplicate baseline is designed to combine both approaches. It predicts a value between the mean baseline and the technical duplicate using an interpolation parameter derived from the DEG P value for each gene. Strongly affected genes are weighted toward the technical duplicate, while weakly affected genes are weighted toward the mean baseline.

The authors then define a meta metric called dynamic range fraction, or DRF, to assess how well any benchmark metric distinguishes the positive control from the negative control. When DRF approaches zero, the metric has little sensitivity to detect improvement over the negative control.

Why It Matters

This is a methodological paper, but it has commercial implications for AI in drug discovery. Benchmark failure and model failure are not the same thing, and the authors are arguing that the field has sometimes treated them as interchangeable. If true, that means some programs may have been judged on tooling that was not able to register meaningful gains even when they existed.

That does not validate every deep learning perturbation model. It does suggest the next competitive edge may come less from bigger architectures alone than from evaluation systems that can tell whether a model is learning perturbation-specific biology in the first place.

For companies building around in silico target discovery, better-calibrated benchmarks could become part of the product, not just the paper trail.