Home Artificial intelligence How Much Can Machine Learning Be Used in HPLC Method Development?
Artificial intelligence

How Much Can Machine Learning Be Used in HPLC Method Development?

Share


Hello.

This is Pharmer.

In this article, I will organize how much machine learning can be used in HPLC method development from the perspective of the method ‘understanding’ required by ICH Q14. In the previous five articles, up to ‘Digitization of Toxicologic Pathology and AI,’ I covered how much AI can be used on the non-clinical research side. From this point on, I will proceed with five articles covering AI on the analytical and quality testing side. As the first of these, I will take up the method development stage.

To state the main point first, machine learning is effective as a tool for streamlining retention time prediction and the search for separation conditions, and for narrowing down which conditions should be verified by experiment. However, the fact that a prediction is correct does not prove that the analytical method is fit for purpose. ICH Q14 requires that the understanding gained during development be documented in a way that can be explained in terms of the Analytical Target Profile (ATP), parameter ranges, and control strategy. Machine learning is a means to deepen that understanding quickly; it does not replace experimental confirmation, validation, or routine control.

1. What kind of guideline is ICH Q14?

ICH Q14 ‘Analytical Procedure Development’ is a guideline that provides a framework for developing and maintaining analytical procedures used for quality evaluation of drug substances and drug products based on science and risk. On November 1, 2023, it reached Step 4 on the same day as ICH Q2(R2), the revised version of analytical procedure validation. Subsequently, the U.S. FDA published it as final guidance in March 2024, and it came into effect in Europe in June 2024. In Japan, it was notified on October 9, 2025, as ‘Guideline on Analytical Procedure Development’ (Yakuyakushin-hatsu 1009 No. 2). As of October 2026, it is at the implementation stage (Step 5) in Japan, the U.S., and Europe.

The relationship between Q14 and Q2(R2) is a division of roles where Q14 covers ‘how to develop and maintain’ and Q2(R2) covers ‘how to demonstrate that it is fit for purpose.’ The basics of validation were covered in a previous article, ‘Complete Guide to HPLC Method Validation for Drug Development: From Basics to Practice Based on ICH Q2(R1).’ Q14 is positioned to organize the development process that precedes that validation into a form that can be explained in regulatory submissions and quality systems.

2. Minimal Approach and Enhanced Approach

Q14 divides the development process into a minimal approach (traditional) and an enhanced approach. Even with the minimal approach, it is necessary to identify the quality attributes to be tested, select the technology, evaluate specificity, accuracy, precision, and robustness, and document the analytical procedure including the control strategy. Q14 explicitly states that the minimal approach remains a valid method.

In the enhanced approach, the following elements are incorporated as needed:

  • ATP (Analytical Target Profile): A description of the purpose of the analytical procedure, the quality attributes to be measured, the performance characteristics, and their acceptance criteria. It serves as the foundation for technology selection and is maintained throughout the lifecycle.

  • Risk assessment and use of existing knowledge: Identifying parameters that may affect the performance of the analytical procedure and determining the priority for investigation through experiments.

  • Univariate/multivariate experiments and modeling: Investigating parameter ranges and interactions.

  • Setting a control strategy: In addition to set points, defining PAR (Proven Acceptable Range) or MODR (Method Operable Design Region).

MODR is the range within which the analytical procedure is shown to be fit for purpose as a combination of two or more parameters. Changes within the approved PAR or MODR are considered not to require notification to regulatory authorities. The advantage of the enhanced approach is that the understanding gained during development leads to flexibility in post-approval changes.

3. Division of Roles between Design of Experiments and Machine Learning

Q14 cites multivariate experiments using Design of Experiments (DoE) as a means to investigate parameter ranges and interactions. Because DoE can estimate the effects of multiple factors with a small number of experiments and express the results as equations such as response surfaces, it is a method that makes it easy to explain which factors are effective and to what extent.

Machine learning is positioned to complement DoE before and after, rather than competing with it.

  • Before DoE: Predicting retention times from past analytical data or public data to estimate the factors and levels that should be investigated.

  • Situations where sequential exploration is performed instead of DoE: Proposing the next conditions to try while incorporating experimental results, and approaching the target separation more quickly.

  • After DoE: Use the model of the obtained response to computationally estimate separation under numerous conditions within the range.

On the other hand, what should ultimately be presented as the basis for set values and MODR is data that systematically verifies ranges and interactions. It is easier to organize by thinking of machine learning for streamlining exploration and planned experiments for substantiating ranges.

4. What can be done with retention time prediction and Bayesian optimization?

Data that forms the foundation for retention time prediction is also being accumulated through public initiatives. In 2019, Domingo-Almenara et al. published a retention time dataset (METLIN SMRT) of approximately 80,000 compounds measured by reversed-phase HPLC and reported a median relative error of 4.6% using deep learning-based prediction. The primary purpose of this research was to support metabolite identification, and they also noted that the predictions are specific to the chromatography conditions used for measurement and that it is necessary to check for bias when projecting to other conditions.

Regarding the exploration of conditions, Boelrijk et al. reported a method in 2023 that uses Bayesian optimization with a Gaussian process model to automatically optimize gradient conditions by linking directly with the instrument. They compared single-objective and multi-objective optimization for two types of dye mixtures with different numbers of components (over 18 and over 80), and while they noted that it is particularly suitable when modeling retention is difficult, they also pointed out that there may be a limit to the number of parameters that can be optimized simultaneously.

What can be understood from these reports is that machine learning is strong at “reducing the number of conditions to test.” If applied to pharmaceutical analytical method development, it is expected to reduce experimental rework in areas such as initial screening of columns and mobile phases, and exploration of separation conditions for closely eluting impurity peaks.

5. How to handle extrapolation outside the range of training data

Machine learning predictions have the property of being more stable for conditions and compounds similar to the data used for training, and errors tend to increase as they move further away. In HPLC, the following situations correspond to this.

  • Compound range: Structures scarce in the training data, such as novel skeletons, highly ionic compounds, and stereoisomers.

  • Condition range: Columns, mobile phase pH, temperature, and instruments different from the training data.

  • Response range: Even if retention time can be predicted, peak shape, overlap, and detection sensitivity are often outside the scope of prediction.

Annex A of Q14 shows an example of chiral HPLC where the robustness of parameters such as temperature is evaluated computationally using a retention time model created from screening data during development, and the results are verified by experimentally measuring the resolution at the center point and at conditions where the retention time of the main peak is at its minimum and maximum. The approach of using predictions to identify candidates and then experimentally confirming the edge conditions where errors are likely to occur is a useful reference even when using machine learning.

Note that the chapter on multivariate analytical methods in Q14 states that while it primarily targets multivariate models using latent variables, machine learning such as neural networks can also use similar principles, but it does not cover them in detail. Validation of analytical methods that calculate measured values themselves using machine learning will be organized again in a later installment of this series.

6. Prediction is for streamlining development; validation and routine management are separate

Even if conditions can be determined quickly with machine learning, the subsequent procedures do not change. Q14 states that while data such as robustness obtained during development can be used as part of validation, the set values used in routine operation must be supported by validation data. It also indicates that the analytical method control strategy should be defined before validation and confirmed after completion, and that the control strategy includes parameters that require control and system suitability tests (SST). The accuracy of a prediction model is not a substitute for these.

With that in mind, the following items should be kept as development records.

  • Model used for prediction and range of training data: Which compound groups and which column/conditions were used to create it.

  • Correspondence between prediction and actual measurement: Which conditions were selected by prediction, which were verified by experiment, and what the degree of deviation was.

  • Decision-making process: Conditions adopted based on prediction, conditions not adopted, and the reasons for those decisions.

  • Correspondence with ATP/risk assessment: Which items in the risk assessment the parameters narrowed down by prediction correspond to.

Q14 states that knowledge management (ICH Q10) plays a crucial role in analytical method development and its lifecycle, and that existing knowledge includes internal development experience. By documenting the results of comparing machine learning predictions with experimental data, this becomes existing knowledge for the next method development and serves as material for evaluating the impact of post-approval changes. Furthermore, Q14 explicitly states that even when using an enhanced approach, the description of the analytical method in the submission documents does not need to be simplified.

7. Summary

Machine learning in HPLC method development is an effective tool for narrowing down the conditions to be verified experimentally by predicting retention times and exploring separation conditions. On the other hand, what ICH Q14 requires is not whether a prediction is correct or incorrect, but rather documenting an understanding of how parameters affect performance in an explainable manner.

  • ICH Q14 reached Step 4 in November 2023 and is in the implementation stage in Japan, the US, and Europe as of October 2026

  • The enhanced approach consists of an ATP, risk assessment, experiments and modeling, and a control strategy including PAR and MODR

  • Machine learning is suitable for improving exploration efficiency, while planned experiments are suitable for justifying ranges

  • Predictions are prone to error outside the range of training data, so a framework that verifies edge conditions through experiments is essential

  • Predictions do not replace validation and routine management; documenting the comparison results between predictions and actual measurements is valuable for future development

Machine learning is a tool that increases the material for decision-making and records the process; the judgment and responsibility for whether an analytical method is fit for its purpose lie with the person who verifies the data. Being able to explain what was predicted and what was verified through experiments, rather than whether or not predictions were used, can be said to be the key point of analytical method development in the era of Q14.

Next time, I would like to take up the topic of whether it is acceptable to leave chromatogram integration to AI — automation of peak processing and data integrity. I hope this will help researchers grasp the position of machine learning in analytical method development.

This article is part of a series that organizes the basics of “AI x Pharmaceutical Development.” Reading this along with the previous article, “Digitization of Toxicological Pathology and AI,” and the magazine below will provide a more comprehensive understanding.



Source link

Leave a comment

Leave a Reply

Your email address will not be published. Required fields are marked *

Related Articles
Artificial intelligence

AI nonprofit will spend $10 million on journalism

The ScoopA Silicon Valley nonprofit close to the AI safety movement is...

Artificial intelligence

Trump Unveils Super Intelligence Force Amid AI Safety Debate – AndroGuider

TL;DR President Trump has unveiled a new Super Intelligence Force tasked with...

Artificial intelligence

Elon Musk Says ‘No More AI’ as He Backs Trump’s ‘Super Intelligence’ Push and Plans SpaceXSI Rebrand

Elon Musk said on 4 October 2026 that SpaceX's artificial intelligence unit...

Artificial intelligence

Scoop: A powerful new model from startup Reflection is set to shake up the AI race

A closely watched Nvidia-backed startup called Reflection is preparing to shake up...