Why Does the True Impact of AI in Drug Discovery Lag?

AI in drug discovery keeps making headlines, but clinical success rates barely move. A new Nature Reviews Perspective explains why, and what pharma should do next.

AI in drug discovery has produced striking headlines, from protein-folding breakthroughs to AI-designed molecules reaching clinical trials. Yet a substantial new Perspective in Nature Reviews Drug Discovery argues that the technology’s clinical impact remains disappointingly limited. 

Image
Pharmatica image showing dense green data nodes that converge on one illuminated point, symbolising AI in drug discovery.

Why AI in Drug Discovery Still Struggles to Reach the Clinic

Written by a multidisciplinary team spanning academia and industry, the Nature Reviews paper gives pharma leaders a sober framework for separating genuine progress from hype, and for understanding where AI-driven models still fall short of changing outcomes at the bedside.

A useful starting point, the authors argue, is a distinction that much of the industry blurs: A ligand is not a drug.

A ligand only needs to bind its target in a biochemical assay. A drug must also clear a demanding set of preclinical and clinical hurdles, including safety, tolerability, and dosing across a diverse patient population.

This difference between ligands and actual drug compounds is why bioactive ligands logged in databases such as ChEMBL and PubChem outnumber approved drugs by roughly three-fold.

Most AI-first drug discovery companies concentrate their models on the preclinical stages, where labelled data are easiest to find.

The Perspective’s economic modelling shows this focus is misdirected because raising success rates in the first efficacy trial, typically phase II, would deliver the largest net savings across the entire development pathway, ahead of speeding up preclinical work or cutting hit-discovery costs.

Preclinical AI, in the authors’ words, resembles looking for keys where the light is, rather than in the place they’re most likely to be.

The paper is blunt in stating that reviews of AI-first drug discovery companies show that the large majority of their projects still sit in preclinical stages, with only a handful reaching phase III trials.

This gap between technological capability and clinical translation is the throughline of this across AI in drug discovery.

Image
Pharmatica image of a small glowing molecule that sits beside a larger fully formed green molecule, representing the gap between a ligand and a drug.

The Data Problem Behind Every AI Model in Drug Discovery

Good data, not algorithms, is the true challenge, and two very different types of data cause two very different problems.

Biological data resist clean labels

Biological readouts are conditional on dozens, sometimes thousands, of variables, including cell line, dose, assay timepoint, patient genetics, even smoking status or patient diet.

The same compound, at the same dose, can succeed in one patient and fail in another for reasons that have nothing to do with the drug itself.

This is called epistemic opacity where researchers often do not understand the mechanism connecting an experimental input to its output well enough to assign it a reliable label.

The example used in the paper involves correlation between an in vitro liver-toxicity assay and the actual clinical drug-induced liver injury that were both found to be low, even after correcting for human drug exposure.

In this example, a common early screening measure, thermal shift, showed only weak association with the enzymatic inhibition data it is meant to predict.

Feeding proxy data like this into a model does not make the model wrong; it makes the model’s answer only as useful as the proxy behind it.

Chemical space is too vast and too local to map

The chemical space is estimated to contain up to 10⁶⁰ small molecules, according to the paper, far larger than the roughly 10²⁴ stars in the observable universe.

Any training dataset can only illuminate a tiny, biased corner of that space.

Chemical space is also highly local. Adding a functional group to one molecular scaffold can behave completely differently from adding the same group to another.

The Perspective illustrates this with four commonly used ADME datasets covering solubility, permeability, blood-brain-barrier penetration, and clearance. Only 0.1% of compounds, and 0.6% of chemical scaffolds, appeared in all four datasets.

Models trained on each property therefore cover different, barely overlapping applicability domains, which complicates any attempt at balancing multiple compound properties at once.

Image
Pharmatica infographic showing how AI model performance translates into value in drug discovery.

 

Key data challenges the paper identifies:

  • Biological readouts shift with confounders such as cell line, dose, and patient genetics, undermining consistent labelling.
  • Chemical datasets sample only a small, biased fraction of a chemical space that may contain up to 10⁶⁰ molecules.
  • So-called external test sets are often drawn from the same pool of data available when the model was built, inflating apparent accuracy.
  • Explainable AI models can surface spurious correlations, such as flagging sugar rings as a cause of bitterness rather than a coincidental marker of it.
  • Benchmark gains do not reliably translate into better real-world decisions, a pattern the paper links to Goodhart's law.

The assay becomes part of the AI system.

If the underlying biological measurement is poorly connected to clinical reality, a sophisticated model simply learns from the wrong signal.

Model Validation Is Not the Same as Process Validation

Distinction between validating a model and validating the process that model serves is one of the most important concepts to understand about AI in drug discovery.

predictive model does not operate in a vacuum but sits between the project context that generates its training data and the follow-up decisions its predictions are meant to inform.

Optimising a generic metric, such as area under the concentration-time curve, without regard for how a model will actually be used can be actively misleading.

The paper sorts real-world use cases into three settings, summarised below:

Setting

Goal

Metric that matters

Example use case

Selection

Prioritise the most promising candidates from a large pool

Precision among top-ranked compounds

Hit finding in a large screening library

Deselection

Filter out unsuitable candidates before they progress

Recall among the undesired class

Flagging compounds with high clearance or toxicity risk

Quantification

Predict an accurate value for one fixed compound

Point-based error, such as root mean squared error

Estimating human dose from preclinical data

 

Two models with near-identical overall accuracy can behave very differently depending on which setting they are deployed in. One hypothetical model can outperform the other twofold in an early selection setting, then underperform it threefold in a later deselection setting.

A single headline accuracy figure, reported without specifying the setting, therefore tells a pharma buyer very little about whether a model will actually help.

AI should increasingly sit inside design–make–test–analyse (DMTA) workflows, where predictions lead to experiments and experimental results feed the next decision. Validation then becomes part of the operational process rather than a one-off technical exercise.

This complements Pharmatica’s analysis of integrated DMTA workflows, where the value of a faster discovery cycle depends on getting better information back into design.

The same principle applies to predictive modelling where better prediction only matters when it changes what scientists do next.

Image
Pharmatica infographic comparing selection, deselection, and quantification settings for evaluating AI models in drug discovery.

Where AI in Drug Discovery Has Delivered, and Where It Has Not

The Perspective is not uniformly sceptical. It highlights protein structure prediction, particularly AlphaFold and related methods, as a genuine success story built on around fifty years of structural data deposited in the Protein Data Bank. AlphaFold is reported to have attracted more than 3 million users across 190 countries.

The paper is quick to add a caveat, though: Protein-folding accuracy does not equal drug-discovery success.

When AlphaFold-generated structures are used for virtual screening of compound libraries, several independent studies found performance comparable to older, cheaper homology models.

Docking against a folded protein and finding an actual, developable drug candidate remain two different problems.

A similar pattern shows up in generative small-molecule design.

The technology has demonstrably improved the search for novel chemotypes and higher hit rates in early-stage screening. Its contribution, however, is still concentrated in ligand discovery rather than full drug design, because current models cannot reliably predict properties such as human bioavailability that only become apparent much later in development.

Where AI-designed drugs have reached mid-stage trials, the authors note that attributing that success to AI alone is difficult, since the lead-optimisation steps linking an AI-generated hit to a final candidate are often undisclosed.

Areas where the paper found the most credible progress:

  • Protein structure and antibody design, built on decades of curated Protein Data Bank data.
  • Synthesis-route prediction, supported by publicly available planning tools.
  • Generative chemistry for hit finding, particularly in widening the diversity of chemotypes searched.
  • High-content imaging and single-cell approaches, which move preclinical assays closer to human biology.

The Highest-Value AI Sits Closer to Clinical Success

As discussed, there is a hugely important mismatch between where AI is currently concentrated and where improvements could have the greatest economic impact.

For small-molecule discovery, much AI activity remains focused on preclinical tasks such as hit finding. These applications can reduce screening effort and accelerate optimisation.

However, as discussed, modelling suggests that improving clinical success rates, particularly around Phase II efficacy, can have a much larger effect on the cost of successfully bringing a medicine to market.

This does not make early discovery AI unimportant.

Better target selection, compound design, pharmacokinetics, toxicity prediction, and patient selection can all influence downstream outcomes. The problem is that these systems are often evaluated using proxy measures rather than the eventual clinical question.

The authors propose a more useful way to think about AI value:

  • Does the model improve a real R&D decision?
  • Does that decision influence a clinically relevant outcome?
  • Can the model work across the data and projects where it will actually be deployed?

Model accuracy alone is not enough.

Building a More Credible Path to Clinical Impact for AI in Drug Discovery

To see actual clinical impact for AI in drug discovery, efforts need to be sequenced correctly.

First, organisations need to identify where the gap between current proxy data and true clinical predictivity is largest, then generate data specifically to close it, rather than modelling whatever data happens to be already available.

Secondally, only once that foundation is in place should teams invest in scaling and productionising the AI models themselves.

Importantly, tighter integration is needed between preclinical and clinical data, which remain siloed by rigid data protection policies inside many organisations.

Learning from real clinical outcomes, including negative results, is essential to correcting the confirmation bias that can build up within ongoing pipelines.

Industry-spanning data consortia, such as those referenced in the Innovative Medicines Initiative programmes, offer one route to the scale this requires, provided that experimental design and harmonised assay conditions are treated as a first-class concern from the outset, rather than pooled retrospectively.

Regulatory momentum adds a further push. Initiatives such as the U.S. Food and Drug Administration (FDA) Modernization Act 2.0, which formally opened the door to non-animal, human-relevant testing methods, encouraging earlier adoption of predictive biological models such as induced pluripotent stem cell systems.

Get the data and the validation framework right first, and the technology should follow.

At Pharmatica, we track the gap between AI's technical promise and its clinical reality across every stage of the Discovery Loop. Our Insights connect the data, the validation science, and the regulatory context that decision-makers need before committing budget to the next AI platform.

Pharmatica analysis focuses on the point where computational capability meets biological reality, helping leaders distinguish genuine progress in AI drug discovery from technology that has yet to demonstrate translational value.

Pharmatica: Insight. Connection. Impact.

Image
Pharmatica image representing AI-enabled design that makes the test-analyse loop for pharmaceutical drug discovery.

Frequently Asked Questions

What does AI in drug discovery actually mean today?

AI in drug discovery today covers a wide range of applications, from protein structure prediction to generative chemistry and clinical trial design. The Nature Perspective stresses that most current AI applications focus on preclinical ligand discovery, which is not the same as discovering a clinically viable drug.

Why hasn't AI increased clinical success rates for new drugs?

AI in drug discovery success rates depend on data that reflects true clinical outcomes, and most AI models are trained on only proxy laboratory measures instead. This mismatch, not a lack of computing power, is the main barrier to translation.

What is the difference between a ligand and a drug?

A ligand only needs to bind its biological target in an assay. A drug must also demonstrate acceptable efficacy and safety in humans, which is why known ligands outnumber approved drugs by roughly three orders of magnitude.

Where has AI made the most credible progress in drug discovery?

Protein structure prediction and antibody design show the strongest results, built on decades of curated structural data. Generative chemistry has also improved hit finding, though its impact remains concentrated in early-stage ligand design.

How can pharma companies validate AI models more effectively?

The 2026 Nature paper recommends matching performance metrics to the model's actual use case, whether that is selecting candidates, filtering out risk, or quantifying a fixed value, rather than relying on generic accuracy scores alone.

Did you enjoy the content?

Why not support Nicole Dale by giving this content a like

Comments (0)

Enlarged image