Insights

Why better data, not more data, will determine the future of effective DMTA-led innovation

AI and automation are driving unprecedented data proliferation across the Design, Make, Test, and Analyse (DMTA) cycle, as richer analytical readouts yield more data per experiment even where compound throughput is capped. Make and Test remain difficult to scale, yet preclinical productivity now lags at Analyse, where extracting reliable, interoperable insights is a significant constraint. To realise the full predictive potential of modern technologies, innovation teams must build decision-ready ecosystems, shifting strategic focus from data quantity to the speed of actionable learning.

Insights

Why better data, not more data, will determine the future of effective DMTA-led innovation

AI and automation are driving unprecedented data proliferation across the Design, Make, Test, and Analyse (DMTA) cycle, as richer analytical readouts yield more data per experiment even where compound throughput is capped. Make and Test remain difficult to scale, yet preclinical productivity now lags at Analyse, where extracting reliable, interoperable insights is a significant constraint. To realise the full predictive potential of modern technologies, innovation teams must build decision-ready ecosystems, shifting strategic focus from data quantity to the speed of actionable learning.

Artificial intelligence (AI) is expanding what is possible to envision in drug discovery by increasing the accessible chemical space. Tools can now uncover molecular candidates for a defined target at a velocity no one could have anticipated just a few years ago.

However, research and development productivity has not matched the experimental throughput enabled by advancements in automation. Costs per approved drug have continued to climb even as computational power has multiplied, as faster candidate generation is only valuable if the rest of the DMTA cycle can keep pace with the number of incoming compounds, and it increasingly cannot. Generative gains accumulate at the design step, then stall behind empirical workflows that govern how quickly idea generation becomes trustworthy evidence.

Reconciling the abundance of data into insights that discovery teams can act on remains a largely manual task, and it is where progress loses momentum. Public databases hold millions of characterised compounds, yet the results that discriminate between candidates, proprietary assay data tied to specific programmes, remain undisclosed and scattered across incompatible systems. Even the data an organisation owns tends to be heterogeneous, produced by different assays and platforms, in formats never assembled for comparison.  

As AI accelerates the possibilities for hypothesis testing, competitive advantage will increasingly depend on the quality, consistency and usability of experimental data rather than its volume. AI has shifted the challenge in drug discovery from generating ideas to having the infrastructure for robust empirical analysis.

The coming data explosion in DMTA

We are witnessing an explosion in the volume of data flowing through DMTA workflows. AI-driven molecular design produces candidates in silico faster than any method before; automated synthesis and high-throughput, miniaturised screening increase the number of compounds physically made and assayed; and organoid and organ-on-chip systems, high-content imaging, and multi-omics platforms multiply the data each assay records. Where a screening campaign once produced a single potency value per compound, it can now generate dozens of simultaneous readouts across biochemical, cellular and phenotypic layers.

The true clinical promise of a molecule only emerges when these multidimensional readouts are combined, as isolated metrics cannot guarantee viability. However, these diverse datasets rarely share a location or a format. Integration, therefore, means reconciling records that are both redundant and unique rather than drawing from a single ordered source. Because much of this critical knowledge remains untabulated and siloed, it accumulates faster than any group of experts can process.

Consequently, as experimental workflows scale to match computational design, constraints move downstream. High-throughput systems and automated platforms can now physically produce and screen thousands of AI-generated structures in a fraction of the usual time. However, processing the resulting readouts—translating overlapping layers of raw results into clear evidence—still takes days or weeks. This mismatch illustrates the predominant bottleneck in drug discovery: the challenge is no longer generating data; it is how efficiently an organisation can extract actionable insight from it.

Why more data does not automatically improve decisions

Discovery organisations already hold years of historical assay data, but its practical usage faces familiar challenges.

Experimental variability

Over a typical seven-to-ten-year drug discovery program, DMTA cycles are run repeatedly but seldom under identical conditions. Operators change, sites vary, and cellular assays introduce biological variability. Year one and year seven datasets may be nominally identical yet statistically distinct, thereby limiting assurance of comparability.

Poor reproducibility

Changing reagent lots or minor inconsistencies in protocol execution introduce discrete batch effects. Pooling data without accounting for batch-specific effects means researchers are no longer comparing compound potency alone but rather contrasting one enzyme batch to another.

Missing metadata

A result is only as useful as its context. A potency value not linked to its assay conditions cannot be trusted, and reconstructing that workflow months later is labour-intensive. Without standardised annotations to bridge these gaps, automated systems cannot contextualise outcomes, leaving predictive models to train on isolated numbers rather than unified biological insights.

Data fragmentation

Chemistry, biology, automation, and analytical outputs are frequently isolated in separate, function-specific systems. Models trained on these disconnected, inconsistently represented datasets tend to apply silo biases rather than underlying chemical principles.

In modern drug discovery, data quality is an independent, measurable variable. Most organisations already possess volume. What they lack is interoperability. The capacity to pool, compare, and safely reuse results for sound decision-making. The consequence is innovation teams that are data-rich but insight-poor.

AI is only as good as the data it learns from

AI models cannot know which parts of their training data to distrust. They learn purely from the labels they receive. When a label carries batch effects, operator variability, or shifting assay conditions, the model cannot distinguish this noise from genuine biological signal. For instance, two results recorded as “IC50 = 50 nM” might originate from distinct reagent batches a year apart. The model treats both as reliable ground truth, fitting the variation as though it were meaningful. This is the impact of a noisy dataset: it often looks perfectly clean until a model fails to reproduce its predictions in the real world.

Experimental consistency contains this noise, yet its strength often erodes over multi-year discovery programmes. If the raw data behind a years-old potency value cannot be reproduced, doubt spreads to every result generated by that workflow. When models are trained on unstructured historical data, they inherit this exposure in full, lacking the strict quality control enforced at the bench level.

Additionally, a dataset can be perfectly accurate yet still skew outcomes due to model bias. Databases heavily favour compounds that successfully modulated a target, as inactivity is rarely recorded. A model trained on this biased record learns an incomplete picture of the chemical space.  

Deeper training data limitations originate from data sparsity around complex structure-activity relationships. Molecular machine learning models default to the assumption that structurally similar compounds behave similarly. While generally true, this breaks down at activity cliffs, where closely related analogues sharply differ in potency. As datasets rarely contain the dense, systematic sampling needed to map these sudden drop-offs, a model will confidently fail exactly where these exceptions occur, simply because it lacks the granular data to learn otherwise.

Data fragmentation compounds these issues. When chemistry, biology, and analytical results are isolated, what reads as a chemical sign may actually be an artefact of the specific system that produced it. Consequently, in drug discovery applications, improving data quality can create far more value than simply increasing dataset size. The defining challenge ahead is less about developing better algorithms and more about engineering the unified data foundations needed to yield reproducible biological insights.

Building a data-centric DMTA workflow

Translating predictive potential into clinical success demands a rigorous approach to the discovery process. While automation’s traditional role has been to add volume—particularly in high-throughput screening—its strategic future lies in stripping out operator-to-operator and site-to-site variability. Without baseline consistency, advanced predictive algorithms remain inherently vulnerable to flawed inputs.

Reproducible biological systems are equally vital. Emerging technologies, such as industrialised organoid production, are engineering consistency into naturally variable biological systems. Extending the logic of automation to complex cellular and tissue models generates highly consistent data. Consistency is fundamental to seamlessly combining biological readouts with chemical structures or text-derived insights, enriching predictive models with multifaceted perspectives that isolated datasets cannot supply.

Integrated workflows across “Make” and “Test” are also necessary to break down the departmental silos that disrupt DMTA cycles. Whether in biochemical screening, DMPK, or formulation, the underlying design-make-test-analyse logic is structurally identical. A shared infrastructure prevents functions from defaulting to disconnected tools. By resolving cross-database inconsistencies, including utilising controlled vocabularies and computing Canonical SMILES for compounds, organisations ensure datasets can be directly joined and reused, rather than constantly curated by solo teams.

The ultimate practical enabler that ties standardised execution, reproducible biology, and integrated workflows together is a unified data architecture: data lineage, full traceability, time-stamped workflows, instrument connectivity, and laboratory orchestration. These elements bind Design, Make, Test, and Analyse together rather than treating them as separate record-keeping exercises. What’s required is a mindset shift: treating data as engineered infrastructure, an absolute prerequisite for successful innovation, rather than a mere by-product of the experimental cycle.

Towards the self-improving discovery engine

The next major evolution in drug discovery transforms DMTA cycles from distinct, sequential steps into a seamless, continuous learning loop. With the integration of these phases, innovation workflows shift toward active learning, in which AI models increasingly serve as investigative partners. Each cycle directly refines the underlying models, experimental design, compound prioritisation, and strategic decision-making, ultimately reducing the number of physical cycles required to identify optimal drug candidates.

Accelerating this loop requires bridging fragmented data silos and reducing avoidable latency. Historically, distinct R&D phases generated isolated data formats, making integration a manual process. Fragmentation can lead to inefficiencies, such as spending months generating analogues for a molecule that already meets success criteria. An integrated system that unifies data and automatically routes compounds to their next test based on predefined criteria eliminates these delays, reserving human judgment for where it creates the most value.

The shape of experiments is also evolving. Modern multi-readout assays simultaneously capture target engagement, off-target effects, and early toxicity. While this reduces the number of separate physical tests required, it multiplies data complexity well beyond human capacity to pattern-match by eye. Here, AI can parse high-dimensional datasets during the Analyse phase to feed stronger, targeted insights directly back into the Design phase.

Automating these workflows is not a push for full discovery autonomy. The goal is an AI partnership that aligns digital velocity with real-world operational and GxP realities. Human oversight remains firmly anchored at key milestones to manage the iterative nature of science, while the system handles repetitive routing in between. The ultimate objective is not generating more experiments. It is accelerating learning and bringing revolutionary drugs that fulfil societal unmet needs to market.

Conclusion

In the years to come, advancing DMTA productivity will rely less on sheer experimental volume and more on how rapidly discovery teams can convert empirical outcomes into actionable learning. Leading organisations in this next phase of drug discovery will not be those who simply aggregate the largest datasets, but those who architect decision-ready data: reproducible, seamlessly connected, and properly contextualised to support robust scientific breakthroughs.

In the current innovation environment, data quality transcends back-office informatics to become a core strategic capability, directly determining the true value extracted from AI, arguably the most transformative technology of our generation. In the next era of drug discovery, competitive advantage will come not from producing more data, but from generating data that enables better decisions.

Talk to us about your next project

Last Updated
July 17, 2026

You might also like

Talk to us about your next project

We help clients with all stages of their most complex and challenging technology and product development projects.



If you're considering the next steps along your innovation journey, why not get in touch?

Thank you! Your submission has been received!
Oops! Something went wrong while submitting the form.

Form unavailable due to browser restrictions.

Your current browser or privacy settings may prevent this form from appearing. Please enable third-party scripts or submit your details through [email protected]

Get the latest from TTP

Join our community to get the latest news and updates on our work at TTP.

You will occasionally receive expert insights from across our areas of focus and hear directly from our engineers and scientists on the newest developments in the field.

Get the latest from TTP

Join our community to get the latest news and updates on our work at TTP.

Form unavailable due to browser restrictions.

Your current browser or privacy settings may prevent this form from appearing. Please enable third-party scripts or submit your details through [email protected]

Want to work 
at TTP?

Find open positions and contact us to learn more.

Overlay title

Lorem ipsum dolor sit amet, consectetur adipiscing elit. Suspendisse varius enim in eros elementum tristique.