Method for determining peak characteristics in an analytical data set
AI-driven peak integration addresses inaccuracies in chromatographic methods by adapting to complex profiles, ensuring accurate and compliant peak identification and integration, facilitating real-time release testing and automated quality control.
Patent Information
- Application Number
- JP2025536023
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2023-04-11
- Filing Date
- 2023-12-19
- Publication Date
- 2025-12-25
AI Technical Summary
Current chromatographic peak integration methods, relying on rule-based parameterization, struggle with inaccuracies due to inherent variability in chromatographic profiles, leading to erroneous peak identification and integration issues, particularly in complex profiles with coeluting or partially dissolved peaks, baseline interferences, and non-linear trends.
A computer-implemented method utilizing artificial intelligence (AI) for peak integration, involving data preprocessing, normalization, and deep learning architectures to adapt to varying chromatographic profiles, with a modular framework for model management and human-in-the-loop validation.
The AI-based method achieves accurate and reproducible peak integration across varying chromatographic profiles, reducing errors and ensuring compliance with regulatory standards, enabling real-time release testing and automated quality control.
Smart Images

Figure 2025542226000004 
Figure 2025542226000005 
Figure 2025542226000006
Abstract
Description
[Technical Field]
[0001] The present invention is directed to a method for analyzing data with a computer-implemented method to determine peak characteristics in an analyzed data set. [Background technology]
[0002] Over the past many years, the medical manufacturing sector has progressed through the development of industrial revolutions. One such industrial revolution brought about digitalization using computers and communication networks for more advanced management strategies toward higher productivity and improved product quality. Continuing modernization aims to build on the previous revolution by connecting cyberphysical systems for fully autonomous, data-driven, adaptive, and predictive control to provide advanced manufacturing. Therefore, it envisions the highest degree of digitalization, automation, virtualization, and decentralization, with significant implications for economic, environmental, and social sustainability. This quest in the medical industry requires the conversion of analog information to digital formats as "digitization," followed by enabling business and operational processes as "digitization" using information technology (IT) platforms, namely, cloud computing, artificial intelligence, big data, the Internet of Things, 3D printing, and blockchain as digital technologies. Digital transformation within the pharmaceutical and biopharmaceutical industry has the potential to bring about a fundamental change in value creation delivering higher quality products with faster time to market, cost-effectiveness, operational efficiency, flexibility and agility, enhanced safety and improved quality regulatory compliance, and reduced waste for product recalls and consistent performance.
[0003] The pharmaceutical and biopharmaceutical industries rely on technology platforms that are established with a collective knowledge base and proven to be compliant with regulatory compliance for the production of safe and effective products for their intended purpose. Because digital innovations using digital technologies are relatively new and involve a higher degree of complexity, there is a lack of experience with new digital technologies, leading to uncertainty in regulatory oversight regarding their adoption in therapeutic product manufacturing processes. The U.S. Food and Drug Administration (FDA) has worked hard over the past several years to accelerate the adoption of advanced manufacturing technologies by establishing research and regulatory programs to update regulatory processes and guidance documents and promote the publication of research on advanced manufacturing processes that can provide regulatory evidence for quality, safety, and efficacy.
[0004] The healthcare industry is exploring a number of emerging digital technologies to bring about digital innovations, including those powered by artificial intelligence (AI) and big data, across the entire healthcare value chain. AI is transforming the operational processes of medical manufacturing in the areas of drug discovery and development, visual inspection of packaging, condition-based maintenance of pharmaceutical manufacturing systems, improved data integrity and quality assurance, and predictive production processes, as well as many more emerging areas.
[0005] Within biomanufacturing, quality control (QC) operational processes are paramount to testing product quality to ensure safety and efficacy. The healthcare industry has established processes known as "pharmaceutical quality systems" to monitor and test product quality with various analytical techniques according to standardized and validated procedures that confirm therapeutic product specifications.
[0006] Among various analytical techniques, chromatography-based methods are primarily employed to determine the qualitative and quantitative analysis of components / attributes in therapeutic products. Product testing with chromatographic methods involves sample preparation, physical separation of components, sensor data acquisition as a function of signal change over time, chromatographic peak data processing, and data interpretation. These operations are performed in accordance with regulatory and compliance guidelines—clear and concisely documented Current Good Manufacturing Practices (CGMPs).
[0007] The integration of chromatographic peaks generates values indicative of product quality; therefore, the integration process must be scientifically sound, justified, and comprehensively documented to ensure the reliability of analytical results. Manual and automatic peak integration operations are available in commercially available software to detect the start and end of a peak, bridge the baseline between the start and end points, and calculate the peak's area and height. Waters Empower software offers two options for automatically integrating peaks: a) the traditional integration algorithm, which detects peaks by considering the rise of the baseline in the chromatographic profile; the integration method is optimized by setting values for peak width, intensity threshold, minimum area, and height; and b) the Apex track algorithm, which detects peaks by identifying the apex of the peak through the second derivative and is based on a measure of curvature, as defined by the rate of change in the slope of the chromatogram. Through the Liftoff% and Touchdown% values, the baseline is determined regardless of apex peak detection. These algorithms rely on rule-based parameterization for peak integration and are not automatically adaptable to highly heterogeneous and variable chromatographic profiles. Analytical runs have inherent variability resulting from retention time shifts, non-Gaussian peak shapes, and drifting baselines. This makes rule-based automatic peak integration difficult to reproduce and apply in complex chromatographic profiles with coeluting or partially dissolved peaks, excessive peak tailing, baseline and matrix interferences, and nonlinear trends. Therefore, the algorithms often result in erroneous peak identification, inappropriate peak cutting, and negative areas. Under these circumstances, manual integration is considered appropriate if the manual process is consistent across all test samples and is performed with appropriate scientific justification and controlled protocols. However, manual adjustment of the baseline can lead to egregious misleading practices related to skimming (reducing peak area) or boosting (increasing peak area) to achieve desired results.Chromatographic data processing has received attention in chromatographic peak integration and interpretation through multiple FDA 483 summonses and warning letters, reflecting the technological need for robust, reproducible, and intelligent algorithms that understand the inherent variability and perform chromatographic peak integration in accordance with data integrity principles and compliance standards. Summary of the Invention
[0008] The present invention provides a solution to the above-mentioned shortcomings by using a method including elements of artificial intelligence to process the integration of peaks in chromatographic profiles, such as those generated from analytical assays used in the quality testing of biotech products. The high complexity of biotech chromatographic profiles and current limitations of integration software have been reasons for inaccuracies. The present method for chromatographic peak integration opens up a solution that enables digitization within the areas of real-time release testing, automated quality control, process analytical technology, and continuous manufacturing.
[0009] In one embodiment, the present invention provides a computer-implemented method for determining peak characteristics from an analysis dataset, comprising: a. Obtaining analytical data in 1+q dimensions for p data sets obtained using the same analytical method; b. for each data set, extracting analytical data on the one-dimensional signal data; converting one-dimensional signal data into a two-dimensional array having cn channels, the two-dimensional array being in the form of a matrix; Normalizing a two-dimensional array having dn channels, wherein the normalized two-dimensional array having n channels is collected in a database; e. for a single data set, identifying the number of peaks in the analytical data, and for each peak, identifying a peak value, a starting peak value, and an ending peak value; A computer-implemented method is provided, comprising: [Brief explanation of the drawings]
[0010] Below is a brief description of the drawings.
[0011] [Figure 1A] Figure 1a shows the historical data set variability as observed for retention time and peak start / end times for method SEC. [Figure 1B] Figure 1b shows the historical data set variability as observed for retention time and peak start / end times for method IEX. [Figure 1C] Figure 1c shows the historical dataset variation as observed for retention time and peak start / end times for method GM. [Figure 2] Figure 2 shows an error plot for RT in the SEC dataset as evaluated on a randomly selected training set. [Figure 3-1] Figure 3 shows an error plot for RT in the SEC dataset as evaluated for progressive training numbers in the acquired data in chronological order. [Figure 3-2] Figure 3 shows an error plot for RT in the SEC dataset as evaluated for progressive training numbers in the acquired data in chronological order. [Figure 3-3] Figure 3 shows an error plot for RT in the SEC dataset as evaluated for progressive training numbers in the acquired data in chronological order. [Figure 4-1] FIG. 4 shows the error plot for PS / PE in the SEC dataset as evaluated for progressive numbers of training sessions in the time sequence of acquired data. [Figure 4-2] FIG. 4 shows the error plot for PS / PE in the SEC dataset as evaluated for progressive numbers of training sessions in the time sequence of acquired data. [Figure 4-3] FIG. 4 shows the error plot for PS / PE in the SEC dataset as evaluated for progressive numbers of training sessions in the time sequence of acquired data. [Figure 5-1] Figure 5 shows error plots for RT and PS / PE in the IEX and GM datasets as evaluated for progressive numbers of training sessions in the time sequence of acquired data. [Figure 5-2] Figure 5 shows error plots for RT and PS / PE in the IEX and GM datasets as evaluated for progressive numbers of training sessions in the time sequence of acquired data. [Figure 6] Figure 6 shows the high-level architecture of the application. [Figure 7] Figure 7 illustrates the operational modules within the application. [Figure 8] Figure 8 shows the potential application hosting in AWS. [Figure 9] Figure 9. Schematic of the GxP framework for AI-based chromatographic peak integration. DETAILED DESCRIPTION OF THE INVENTION
[0012] The present invention provides a solution by using a method that includes elements of artificial intelligence to process the integration of peaks in chromatographic profiles, such as those generated from analytical assays used in the quality testing of therapeutic products. The high complexity of chromatographic profiles of biological products and current limitations of integration software have been reasons for inaccuracies in the peak integration process. The present method for chromatographic peak integration opens up a solution that enables digitization within the areas of real-time release testing, automated quality control, process analytical technology, and continuous manufacturing.
[0013] Examples of chromatographic techniques include size-exclusion liquid chromatography, ion-exchange liquid chromatography, and hydrophilic interaction liquid chromatography. Furthermore, applications extend beyond these examples, as they can be applied to any data set with x- and y-axes to generate peak profiles. Size-exclusion liquid chromatography (SEC) methods are widely used as analytical techniques for separating monomers, fragments, and aggregates in monoclonal antibody products. These components are separated according to their size as they pass through a column bed made of porous particles. Larger components elute faster than smaller components due to differences in permeation across the column pores, i.e., smaller components permeate more easily than larger components. Ion-exchange liquid chromatography (IEX) methods are widely used as analytical techniques for separating protein components in monoclonal antibody products based on their charge. The principle involves Coulombic interactions between ionic functional groups on the stationary phase (column matrix), and oppositely charged analyte ions present in the components of the monoclonal antibody product. IEX can be performed in two modes based on the charge of the functional groups on the stationary phase: cation exchange retains positively charged ions, while anion exchange retains negatively charged ions. The composition of the mobile phase determines the separation and elution of proteinaceous components according to their net charge. In most cases, weak cation exchange mode is used to separate acidic, basic, and main variants in monoclonal antibody products. Component elution follows the order of acidic-main-basic, depending on the charge distribution of the components as imparted by post-translational modifications in monoclonal antibody products. Many glycan analysis methods (GM) are widely used to determine glycosylation in monoclonal antibody products. Hydrophilic interaction liquid chromatography (HILIC) coupled to a fluorescence detector is one common chromatography-based method performed on free and labeled glycans. HILIC provides separation based on the hydrophilicity of the glycans, and retention correlates with glycan size.The elution profile provides peaks of glycan components representing monoclonal antibody products. High-performance liquid chromatography (HPLC) or ultra-high-performance liquid chromatography (ULPC) systems are used to perform chromatography-based SEC, IEX, and GM methods, and signals are acquired using a detector system (ultraviolet or fluorescent). The elution of components is passed through a detector to record the signal, which is then converted into a representation of peaks over time. The peaks are then integrated by connecting the peak start and end points horizontally at the baseline and then vertical drop lines to separate the identities of closely eluting components. The integration then provides the proportion of each component and their retention times, which indicate their identity and distribution. Therefore, retention time (RT) and peak start / end times (PS / PE) are important variables for integrating peaks and calculating component distribution.
[0014] Data sets used to train computer-implemented methods can vary in the chromatographic profiles, e.g., SEC, IEX, and GM chromatographic profiles. SEC profiles are considered simpler because they contain fewer peaks. However, analytical variability can lead to difficulties in consistently integrating the profiles. The SEC data set included consistent profiles belonging to standards and stressed material profiles with variations in component distribution levels. Thus, the SEC data set represented a realistic data source with historical analytical variability, along with differences in component ratios. Similarly, IEX had a medium-to-complex level of complexity for peak integration due to the presence of many more peaks and analytical variability of the technique. The GM data set represented a representative case for the most complex profiles, due to the presence of many peaks and the highly sensitive analytical method prone to large variability. Thus, the data sets represented different chromatographic separation principles, varying levels of complexity for the number of peaks and resolution of separation, as well as differences in the sensitivity and variability of analytical methods. Observations from historical data set variability (Figure 1, as assessed from manually integrated chromatographic profiles) showed that, in terms of maximum shift, SEC had the lowest variability for RT (0.3 min) and PS / PE (1.2 min), and GM had the highest variability for RT (4.2 min) and PS / PE (4.9 min). One PS / PE data point in GM was observed at 50.5 min, which was considered an outlier and excluded. Meanwhile, IEX showed intermediate variability for RT (2.3 min) and PS / PE (1.9 min). Notably, the RT numbering for IEX began at RT2, which was due to the presence of inconsistencies for peak 1, which was very low in intensity. Therefore, it was excluded. However, peak start / end was retained to assess the overall shift.
[0015] To address the need for a minimal training dataset for an algorithm to consistently integrate chromatographic peaks, for example, an SEC dataset can be used to optimize an initial architecture capable of functioning with a smaller amount of data. Here, a deep learning architecture was developed and trained using 100 randomly selected datasets from 1550 runs. The absolute errors observed within the test set for RT and PS / PE were within 0.1 min (Figure 2). This provided strong confidence in the artificial intelligence (AI) technology for the chromatographic peak integration process, using only 100 datasets to train the model. Generally, deep learning approaches require a large amount of training data to provide acceptable accuracy in predictions. The algorithm herein can adequately capture features from less data, and we performed various checks for model overfitting by: a) tracking loss values for training and validation sets versus epochs; b) monitoring the comparability of the maximum absolute error between training and test samples; c) testing the model's predictive performance on unseen test data; and d) k-fold walk-forward cross-validation to check the model's performance on different sets of data. These checks confirmed the absence of model overfitting and validated the performance of models trained with fewer data sets. The random selection of data covered the expected overall analytical variability of the SEC method, which was incorporated into the training of the AI model. Therefore, we clearly demonstrate higher performance and its ability to adapt to accurate integration of chromatographic peaks. Therefore, superior performance compared to existing peak integration approaches is provided herein.
[0016] During the development of a new analytical chromatographic method and its peak integration process, large datasets are not initially available, and analytical variability is not fully known. Analytical variability evolves during routine use of the analytical method. Therefore, model performance is understood in the context of evolving analytical variability in chromatographic profiles. Herein, the training set was defined based on the time course of the run / data acquisition. SEC datasets were organized by the date and time of their acquisition, and a batch of 100 datasets was selected to train the AI model and test its performance on the next time-aligned set of data. This approach provided a realistic environment of a business operating process to test the performance of the AI model. The SEC datasets consisted of multiple data sets with different analytical sessions, instrumentation uses, and standard stress processing profiles (including variability in analyte distribution). The SEC AI model was trained on the initial 100 datasets organized by time series of data acquisition, and then tested on the subsequent 100 datasets. The subsequent 100 datasets were then incorporated into the AI model by retraining and tested on the other 100 datasets in time series. This process was repeated until all available datasets were utilized. Therefore, training and retraining batches of 100, 200, 300, 400, 500, and 1000 were used to upgrade the SEC AI model. The absolute error results for RT (Figure 3) showed a gradual improvement in accuracy from 0.25 min to 0.05 min as each model was trained and retrained from 100 datasets to 2000 datasets. Interestingly, after 400 datasets, some large absolute errors were observed above 0.1 min to 0.4 min, with two datasets between 0.4 min and 0.55 min. This phenomenon was caused by the presence of datasets from stress-treated profiles. It is well known that stress treatment of antibody products causes molecular changes that are reflected in the chromatographic profile.However, when the stress treatment profile was further incorporated into the retraining of the AI model, performance improved, with 96% of the subsequent 100 data sets being within 0.1 minutes, and four data sets reaching 0.2 minutes. The AI model thus learned the stress treatment profile and demonstrated its ability to accurately integrate peaks. This feature also demonstrates the advantages of AI in automating the chromatographic peak integration process. Similarly, PS / PE performance (Figure 4) showed the same improvement as the training data set increased over time, incorporating evolving analytical variability. In particular, the terminal drop line, PS / PE3, exhibited a wide scattering window (up to 0.8 minutes). PS / PE3 is the last drop line in the peak profile, where the eluting peak merges with the baseline. This region is typically prone to the phenomenon of "peak tailing." There are various causative factors across the sample, column, solvent, and instrument that lead to peak tailing. Different analytical methods also vary in peak tailing patterns and frequencies, which is a known uncertainty managed by operators in the peak integration process. Therefore, peak tailing leads to inconsistent peak integration in this region, often resulting in small shifts due to operator variability. However, in most cases, the shift contribution does not significantly affect the quantitative results. Nevertheless, the main peak tailing shift can contribute to differences, and regulatory agencies have issued guidelines for tracking peak tailing in measurements. Therefore, the larger scatter in absolute error for PS / PE3 was the result of operator-induced variability and uncertain contributions from the peak tailing phenomenon. PS / PE1&2 had absolute errors of less than 0.2 min for models trained on 100 runs. Similarly, when stressed sample runs were tested with models retrained from 400 and 500 runs, some large absolute errors were observed, up to 0.6 min, with some datasets being approximately 1 min.Therefore, it was expected for the stress treatment profile when it had a change in distribution, and this type of data was not recognized by the model at all. Furthermore, the performance of the peak integration process improved after retraining the model with this unique new data (all PS / PE less than 0.25 min). Therefore, the developed architecture was able to learn the past variability of the SEC method as given by all sources and efficiently perform the peak integration process.
[0017] To define a universal architecture for building models for simple to complex chromatographic profiles / methods, we repeatedly built, tested, and evaluated small architectural variations with 500k parameters. The input layer was kept configurable, and its shape depended on the dataset. The preprocessed input was passed through a group of convolutional layers, ReLU activation layers, and max-pooling layers, thereby extracting the most important features from the input. The convolutional layer, consisting of 32 channels or feature maps, was the first group to extract features from the input using linear operations. The output from the ReLU activation function was passed through max-pooling convolution to match and remove the most important features present. The number of neurons in the final output layer depended on the number of prediction outcomes desired from the model. As an example, an IEX profile consists of four peaks corresponding to four retention times (RT) and five peak onset / peak end times (PS / PE). Therefore, the final output layer (layer 22) contained four neurons for the model trained on RT outcomes and five neurons for the model trained on PS / PE. The Adam Optimizer with a learning rate of 1e-3 was used as the optimization function. The model was run for 500 epochs using an early stopping technique to avoid overfitting. The mean squared error was used as the loss function to train the model for peak retention time and peak onset / end time. The training loss and validation loss were tracked to check the model performance during training. Of the seven architecture variants, architecture 7 was based on transfer learning, and performance was evaluated across three datasets. Transfer learning is a method of applying learning from one task to another. This approach is generally used when initial data for training is lacking and AI performance needs to be improved. Transfer learning has rapidly expanded to numerous applications, with nearly 40 representative transfer learning approaches. Transfer learning was applied by freezing convolutional layers 2 to 13 in the initial architecture and using the previously optimized weights. Therefore, training was performed on the subsequent linear layers (16 to 22) in the later parts of the initial architecture.With this approach, features (lines, edges, slope change at the beginning of a peak, slope change at the end of a peak, edge curvature, etc.) were extracted from a pre-trained SEC model, and the findings (weights / features from the pre-trained SEC model) were reused to train a new model on the IEX and GM datasets. Evaluation of the performance of these architectures showed that the best architecture, consisting of an initial (small architecture with 500k parameters) + two additional layers (1 million parameters) + 0.3 dropout, demonstrated superior performance. The second best architecture was based on transfer learning. The two architectures were further evaluated on the IEX dataset, and the performance of the architecture (without transfer learning) was found to be slightly superior in accuracy. This overall better performance relative to the baseline architecture was due to the use of a dropout technique for regularization, which prevented overfitting and resulted in acceptable accuracy with less training data, allowing the AI model to perform well on test data from new time series acquisitions. Therefore, the best architecture was selected as the baseline architecture for further study.
[0018] The selected baseline architecture was applied to the IEX and GM datasets by changing only the corresponding input data (intensity vs. time, retention time, peak start and end times), the number of epochs, and the learning rate in the architecture. Indeed, the number of peaks and their start and end positions relative to the baseline were required to be provided as user-defined inputs. Using only these parameter changes, the baseline architecture was applied to generate models on the IEX and GM datasets. Training set samples were selected based on the time-acquired data, and testing was performed on subsequent data in the time series acquisition. For the IEX model, the results (Figure 5) showed that 95% of the test data were within 0.3 minutes of absolute error for RT, and 70% of the test data were within 0.5 minutes of absolute error for PS / PE. The remaining 30% of the PS / PE data showed a shift in all PS / PE drop lines. This clearly indicated that analytical variability arose from different sources, i.e., in this case, different laboratories, instrumentation, columns, and operators. This data demonstrates the sensitivity of IEX chromatographic runs and the degree of shift that occurs as variability occurs. Interestingly, the absolute RT error was less than 0.3 min for the test data showing a large shift in the PS / PE drop line. This is due to the integration approach in the IEX chromatographic profile, which is based on integrating peak clusters. Nevertheless, retraining of the IEX model is necessary from the 15th test data onward to improve the AI peak integration accuracy. For the GM model, a larger absolute error of 1.1 min for RT and PS / PE was observed for 82% of the test data. Here, 18% of the test data (three runs) showed very different profiles leading to absolute errors of up to 3.2 min for RT and 3.7 min for PS / PE. These are specific cases; typically, glycan mapping chromatographic profiles are known for their high sensitivity and variability. In such cases, user-defined mitigation is required. However, this observation is based on only very little data.By retraining the GM model with more data capture, these examples demonstrate the promise of learning these variations and applying them to future data. These results therefore amply demonstrate that this baseline architecture can be easily adapted to other new methods or profiles without requiring advanced IT expertise. The baseline architecture can be easily applied to other chromatographic profiles that are greater in number or peaks, complexity, and sensitivity to analytical variations. However, only 90 or 100 initial training data sets may not provide the desired level of accuracy. A trade-off between data and performance arises, which also increases with the complexity of the chromatographic profile and method.
[0019] AI models are built from training data and validated against test data. Model validation defines their performance in terms of their quality. Therefore, defining criteria in CBA: Model Management and Performance Monitoring is highly relevant for the routine use of AI models. The performance of machine learning models is calculated by comparing trained model predictions to actual observations. Numerous mathematical approaches exist for assessing regression model validity. Root mean square error (RMSE) is a frequently reported metric in the literature and is a function of a set of three error characteristics. Meanwhile, mean absolute error (MAE) is more explicit than RMSE and can provide a dimensional assessment for intercomparisons. The inventors have used MAE (Mean Absolute Error) as a model validation approach to evaluate the performance of trained models and their predictive ability on test data during model optimization studies and baseline architecture development. Within this application, MAE is reported as "minutes" or "%" to allow users to easily interpret it. Furthermore, mean square error (MSE) is used as a loss function for training models. For routine use, the indicators must be interpretable by operators to define and implement subsequent operations. For evaluation of the peak integration process as performed by the model, the maximum absolute error provided a more complete interpretation of performance when tested on time-series acquired test data. The RT accuracy for 90% of the test data was 0.2 min for SEC and IEX and 1 min for GM. For PS / PE, the absolute error values were 0.4 min, 0.8 min, and 1.3 min for SEC, IEX, and GM, respectively. These indicate the maximum absolute error and provide a more complete understanding for determining whether to ignore outliers or define operations. Therefore, we adopted the absolute error as a scatter plot showing the performance of each test data point. This provides better visibility into model performance to identify outliers or to inform decisions regarding retraining the AI model due to variations in the analysis method. The accuracy benchmark was a user-defined input based on scientific experience with the analysis method and its variations.During training and retraining of the AI model, accuracy was assessed using user-configured thresholds for RT and PS / PE minutes. Performance evaluation was performed at two levels: first, using ground truth data, which was then used to calculate absolute error using the operator's integration profile from Empower (extracted from the report); then, at level 2, the AI's integration profile was checked by the operator. This was required to provide a human-in-the-loop approach to guarantee the performance of such a black-box-based deep learning model. This is in line with current non-AI peak integration processes, where the final evaluation is performed by visual inspection to ensure that the peak is properly integrated and avoid any effects related to skimming or growth. Therefore, the peak integration process in our application was designed to provide a system- and human-in-the-loop approach to decision-making during the process. However, similar metrics exist for monitoring the performance of the peak integration process (in our case, relative % distribution). The performance (relative percentage value) has a dependency on PS / PE and RT as a calculation performed from the integrated peak area. Therefore, the RT and PS / PE metrics were primarily used to monitor and evaluate model performance along with visual inspection of the AI integrated peak profile. This was the approach adopted in training and validating the model on test data, such as those selected from the time-acquisition sequence. In the routine use of peak integration with an AI model, no ground truth information exists. In this case, performance monitoring of the deployed model was assessed by operator visual inspection and mitigated accordingly. We incorporated user interface functionality to enable visualization of the AI peak integration profile to the operator, allowing corrections to the drop line and baseline start and end. Therefore, we incorporated audit trail manual management of the integration process for outliers or erroneous peak integrations. Operator actions toward mitigation / correction are fully tracked by the application's audit trail component.
[0020] Guiding principles for good machine learning practice (GMLP) were jointly identified by the US Food and Drug Administration (FDA), Health Canada (HC), and the UK Medicines and Healthcare products Regulatory Agency (MHRA) to promote the use of artificial intelligence / machine learning (AI / ML)-based software as medical devices (SaMDs) to deliver safe and effective software features that improve the quality of patient care. Through a discussion paper, a proposed mechanism for a whole-of-product-lifecycle (TPLC) regulatory framework was presented to embrace the iterative improvement capabilities of AI / ML SaMDs. The concepts of GMLP, prescribed change management plans, and algorithm change protocols (ACPs) were introduced to engage in dialogue with manufacturers to represent harmonized standards and AI / ML best practices through consensus and community-driven regulatory management. Similarly, an informal network for innovation was established by the International Coalition of Medicines Regulatory Authorities (ICMRA) to adapt regulatory frameworks to emerging new technologies to promote safe and timely access to innovative medicines. Working group members within the Informal Network for Innovation comprised the Italian Medicines Agency (AIFA), the Danish Medicines Agency (DKMA), the European Medicines Agency (EMA) (Working Group Leader), the US Food and Drug Administration (FDA) (as an observer), Health Canada (HC), the Irish Therapeutic Goods Regulatory Agency (HPRA), the Swiss Therapeutic Goods Agency, and the World Health Organization (WHO). Within the ICMRA AI Working Group, horizon scanning work was undertaken to identify challenging topics relevant to regulators and stakeholders regarding the use of AI. Hypothetical case studies on AI were defined to challenge existing regulatory frameworks and develop recommendations for adoption, demonstrate the need for a risk-based approach through collaborative exchange with ICMRA, link medicinal benefits / risks to AI model management structures, and provide regulatory access to AI models and their dependent datasets to facilitate a full understanding of the validity of AI algorithms.
[0021] Following recommendations, we have defined a proposed GxP framework for our AI application, adapted from the FDA TPLC approach, the Guiding Principles for GMLP, the ICMRA AI Recommendations, and existing guidelines for data integrity under 21 CFR Part 11. Therefore, we built the application in compliance with GAMP5 and FDA 21 CFR Part 11 to ensure compliance, auditability, traceability, and data integrity. The application functionality was built using AWS physical infrastructure, virtualization, and service layers. AWS does not have specific GxP certifications for its cloud products and services. However, AWS provides commercial off-the-shelf (COTS) IT services in accordance with IT security and quality certifications, such as ISO 9001, ISO 27001, ISO 27017, and ISO 27018, and complies with NIST 800-53 under the FedRAMP compliance program. The process of the present invention complies with industry guidance provided by FDA Part 11 / CFR Title 21 by incorporating features such as LDAP-based authentication, password policies according to IdP guidelines, role-based access, audit trail and application log viewing and maintenance, and export of logs for review and distribution. Additionally, additional features pertaining to GxP requirements were incorporated such as access at a secure layer (SSL), patch management, antivirus management, encryption and backup, and disaster recovery. The application was developed in AWS following good engineering practices as recommended by the GAMP5 lifecycle approach.
[0022] The architecture of the present invention is configured in a modular configuration with a presentation layer that provides access to operators through a user interface and manages and executes functional operations. Seamless connectivity was established with a Rest API to integrate a business layer composed of logic for application management and data service management, connected to a data layer. A data science layer incorporating AI algorithms was integrated with the business layer. The application was developed in a modular format to enable independent upgrades of modules for its wide applicability and extensibility. Once access to the application is established according to authentication and policy, the first step involves the creation of a "project-session," and an identifier is assigned by the data management logic. Third-party export files representing chromatography profiles are imported into the application by the input module. For our use case, the import logic is designed to work with Waters Empower export format *.arw and *.cdf files as raw data (chromatography profiles) and *.pdf files for processing results. Additional import logic can be coded to extend the application's interoperability with other chromatography software vendors. The input module performs data import validity checks to ensure data integrity and extracts standardization information to align with the raw data management logic by assigning a unique identifier to each imported dataset. The input module also links raw data and third-party processing results to be used as ground truth for training models. The model management module is used to create new models or retrain existing models. It requires the selection of an available baseline architecture and corresponding training data, along with only the parameters "Number of Epochs" and "Learning Rate" in the architecture.A universal architecture is built into the application to train deep learning models that can be applied to any type of chromatographic profile with varying complexity resulting from different analytical methods. Flexibility is provided to configure models for the number of peaks (n) and baseline. Thus, even users with non-developer IT skills can build models from scratch. Model Management also includes a model performance dashboard, which allows operators to connect humans in the loop to review and validate trained models. This feature provides granular information related to peak integral plots, data variability within the training set, model accuracy on test data, and model performance metrics. A high level of transparency allows for building trust and reliability in trained models. To promote data integrity within the application, the Model Management module provides traceability for the training data used to train and retrain models, along with versioning tracking via unique identifiers. These details can be reviewed at any time. The final trained model, as deployed for routine use, imports raw data files managed from the Data Management module. The deep learning model facilitates the chromatographic peak integration process by predicting retention time, peak start time, and peak end time, identifying peaks, drawing drop lines that are further processed by scripts to create baselines, and calculating the area under the peaks to provide results. A human-in-the-loop module is provided to inspect the integration profile and promote data integrity principles if required to initiate manual corrections under the tracking workflow. The results are then managed in the data management module with unique identifiers.
[0023] Regarding the user interface (UI), as shown in Figure 8, a modular software application was built on an extensible, open-architecture system with an independent UI module built on a Java technology stack with Sprint-boot and an algorithm module built on Python. These modules were integrated with REST APIs to enable seamless exchange within internal services, and the application was hosted in the AWS cloud environment. The UE module provided user access to the application through modern browsers (e.g., Chrome, Microsoft Explorer, Firefox) built using SSL (Secure Socket Layer) technology. Access was granted through an internet-facing URL and was geofenced for locations / countries. User authentication was enabled through SAML 2.0 (Security Assertion Markup Language 2.0) and SSO (Single Sign-On) based on access management. User authorization was configured to be authenticated through an IdP (Identity Provider) and OAuth (Open Authentication). A root administrator was configured to provide access to a limited, required user base by pre-configuring their unique identification information (Id and email). The user interface within the application was built using React JS, covering functions such as file upload from a browser, UI-based validation, and display of results and reports. Data management and organization within the entire application was built on MySQL® (Amazon RDS) as the database. Upload of raw data files in *.cdf, *.arw, and *.pdf formats was built and managed with Java® / J2EE and Spring Boot as middleware. A script was configured to check the data integrity of the uploaded files before importing them into the application.Amazon Elastic Compute Cloud (Amazon EC2) was used for web application hosting, backup, and recovery. Storage of source data for application logs, model training, and predictions was implemented using AWS Simple Storage Service (S3). The developed baseline architecture was deployed in algorithm modules using Python as the programming language. This included artificial intelligence, machine learning, and neural network-based logic and services. NGINX was used for web servers and load balancing.
[0024] A GxP framework can be built around three main pillars: data management, model management, and human-in-the-loop, as shown in Figure 9. These elements can also address regulatory expectations and develop a system that can be trusted by operators. Data management dominates the model management and human-in-the-loop pillars, enabling seamless exchange between systems and humans with transparency to determine the validity and confidence in AI models. Based on data, decisions can be successfully assured, and confidence can be gained in the performance of AI models. In particular, the setup designed for a GxP framework consists of well-organized data management to track and trace the training data used to build the model. Furthermore, performance metrics associated with the model are also carefully crafted to provide granular information about model accuracy and validity. In particular, the model management dashboard in the UI provides metrics related to the inherent variability in the training data, a scatter plot of model performance on the test data, and model accuracy range. This high level of granularity and tracking allows operators and regulators to trust AI applications in the healthcare sector.
[0025] Working Example: Three data sets belonging to different chromatographic profiles from the analytical methods size exclusion chromatography (SEC), ion exchange chromatography (IEX), and glycan mapping (GM) were used in the study.
[0026] [Table 1]
[0027] Each chromatographic profile for the analytical technique (SEC, IEX, GM) was exported from Waters Empower software. The raw data from Empower and its operator-processed report were exported using the software's available export function. The exported files consisted of three file types, identified by their extensions: *.arw, which consisted of retention time and intensity information for the chromatographic run, and *.cdf, which contained metadata for the acquired run, i.e., "acquisition data," "retention time," "area," and "percent area." The file names for the *.arw and *.cdf files were identical. However, results from the operator-processed data in Empower were presented in the report as a .pdf file with a different filename. These files contained the results of the corresponding chromatographic profile and consisted of information for each peak observed in the profile, such as "sample name," "acquisition data," "retention time," "peak onset and end times," "slope," "area," and "% area." The raw chromatographic data were represented by the *.arw and *.cdf files, and the processed results were represented by the .pdf file. The differences in naming were due to different sets of records as stored within the Empower software. Therefore, an automated script based on text extraction and natural language processing (NLP) was used to align the raw data with its corresponding results by using a common information variable, "obtained data," that was present in both the *.cdf and °.pdf files. Furthermore, the script was used to check the presence of necessary information in the result data °.pdf files to ensure that no missing peaks were present and that the information about the results was consistent across the dataset. Furthermore, each *.arw file was checked for the number of data points per second and the end time. Files with inconsistent data were removed. Therefore, the exported data files were aligned and standardized as input files for the deep learning model construction. The input files were scaled (between 0 and 1) for more robust results.
[0028] Each chromatographic run consisted of time series data aligned to its intensity values, recorded at a frequency of 1 second. These values were converted to a square matrix to feed into the deep learning model. All values were normalized using the mean and standard deviation of the training set for easier gradient calculation in backpropagation of the deep learning model. This allowed for faster convergence of the loss function and rapid training of the neural network model. The resulting values from the °.pdf file were used as ground truth to train the deep learning model.
[0029] Deep learning approaches were explored to develop an artificial intelligence-based algorithm. The model was developed using Pytorch along with other libraries such as numpy, scikit-learn, pandas, matplotlib, etc. A convolutional neural network (CNN) model was constructed to predict the retention time, peak start time, and peak end time of the chromatographic profile. A baseline was constructed by connecting the predicted peak start time and predicted peak end time using a straight line. The equation of the line was generated using two points (predicted peak start and predicted peak end).
[0030] The CNN model was constructed in a 22-layer architecture consisting of concatenated individual blocks as convolutional layers and rectified linear units (ReLUs) as the activation function. The ReLU activation function was chosen to prevent the vanishing gradient problem and learn from nonlinear relationships between layers. The ReLU activation function, defined as R(z) = max(0, z), allows irrelevant features to be removed by zeroing them out. ReLUs were interspersed with dropout and max pooling before the fully connected layers, and a dense output layer. The overall architecture configuration consisted of one flattening layer, one dropout layer at the end to remove overfitting, three dense layers with reduced numbers of neurons (512, 256, 128, etc.), four convolutional layers with an increasing number of channels in each layer, and one output layer with a fixed number of outputs. Four max pooling layers after each convolutional layer were used to reduce the number of parameters to be trained. The CNN model was optimized over several iterations by varying hyperparameters such as the learning rate, loss function, activation function, number of layers, number of epochs, batch size, and dropout layers to optimize the loss value, i.e., accuracy as measured by the error between the predicted value and the ground truth. The operator integral profile from the exported Empower dataset result file was considered as the ground truth. The input data features were extracted in the CNN model through convolution operations by passing a weight matrix with a kernel filter of size 3. Four convolution blocks were connected with different numbers of channels to extract deeper features. Each convolution layer was followed by a ReLU activation function and a max-pooling layer, which worked in tandem to extract only the features that provided the best output, which were then trained by backpropagation through adjusting the kernel weights. Backpropagation was performed iteratively using an adam optimizer with a learning rate of 1e-3 to update the kernel weights by using the gradient of the weights relative to the output and its loss value to learn patterns in the data and improve accuracy.A max-pooling kernel was convolved in the channel to provide the maximum value from the convolutional domain. The convolutional blocks were followed by three blocks of fully connected or dense layers for deeper learning with feature mixing and matching as learned in the convolutional blocks. A dropout layer was included between the convolutional blocks and the dense layers, with a dropout value between 0 and 1 to randomly drop a portion of the neurons in the fully connected layers to avoid overfitting to the training data. Because the model was not deep, batch normalization was not included in the model architecture, and good results were obtained without it.
[0031] Model training was performed using two data organization approaches: random data selection was used to train the model with a fixed seed, with a data split ratio of 80% training, 10% validation, and 10% test. The second approach consisted of using acquisition-time-aligned data selection with a rolling window of progressive batches of training and retraining samples. The training set was progressively increased to 100, 200, 300, 400, 1000, etc., while the test set was fixed at 100 samples in subsequent data acquisition series with the most recent samples. The performance of the trained model was evaluated from the mean absolute error.
[0032] Several architecture types were explored by iteratively building, testing, and optimizing hyperparameters, as shown in Table 2. The performance of different architectures was evaluated from individual models prepared from each of the three datasets (Table 1). Furthermore, separate models were developed using these architecture types to detect peak retention times and peak onset / end times using 684,164 trainable parameters in the model. Transfer learning was explored by freezing the convolutional layers in the baseline architecture and using the weights of the previous trained model. Training was performed within the linear layer present after the baseline architecture, which has a total of 296,288 trainable parameters in the model (Table 2).
[0033] [Table 2-1] [Table 2-2]
Claims
1. 1. A computer-implemented method for analyzing data to determine peak characteristics for a sample analysis dataset, comprising: a. processing one or more analytical data sets having at least an x-axis and a y-axis to identify peak characteristics; b. storing the peak characteristics in a database; c. generating a trained model for an analytical method based on the stored peak characteristics; d. applying the trained model to the sample analysis dataset to obtain predicted retention times, peak start times, and peak end times; e. Managing the raw data files, model training data, and results by unique identifiers in a data management module; f. Managing model training and performance data with unique identifiers in a model management module; g. Outputting the peak characteristics in the form of tabular values and visual representations in the form of graphs or profiles; h. Providing a means to inspect the integration profile and provide user modifications under a tracking workflow; i. Identifying the peaks and drawing drop lines; j) outputting the peak characteristic; 12. A computer-implemented method comprising:
2. Step b) is a. acquiring analytical data in 1+q dimensions for p data sets acquired using the same analytical method; b. for each data set, extracting said analytical data in one-dimensional signal data; c. converting the one-dimensional signal data into a two-dimensional array having n channels, the two-dimensional array being in the form of a matrix; d. normalizing the two-dimensional array having n channels, wherein the normalized two-dimensional array having n channels is collected in a database; e. for a single data set, identifying the number of peaks in the analytical data, and for each peak, identifying a peak value, a start peak value, and an end peak value; 2. The computer-implemented method of claim 1 , comprising: p=1, p=2, q=3, p=4, q=5, p=6, q=7, p=8, q=9, p=10, q=11, p=12, q=13, p=14, q=15, p=16, q=17, p=18, q=19, p=19, q=19, p=1
3. The method of claim 1 or 2, wherein the identified peaks and peak values, start peak values, and end peak values are imported into a database.
4. The method of claim 1 , wherein the trained model is stored in a database.
5. The trained model further comprises: a. applying steps a) to e) of claim 2 to one or more additional analytical data sets obtained using the same analytical method; b) combining the value for the number of peaks, the peak value, the start peak value, and the end peak value in the additional analysis data set with the value for the number of peaks, the peak value, the start peak value, and the end peak value in the analysis data set obtained in step e); c. generating a trained model from the combined results; The method of claim 2 , wherein the training includes:
6. 1. A computer-implemented method for determining peak characteristics of analytical data obtained for a sample in a sample data set, comprising: a. acquiring analytical data in 1+q dimensions for p data sets acquired for the sample using the same analytical method as the sample data set; b. for each data set, extracting said analytical data in one-dimensional signal data; c. converting the one-dimensional signal data into a two-dimensional array having n channels, the two-dimensional array being in the form of a matrix; d. normalizing the two-dimensional array having n channels, wherein the normalized two-dimensional array having n channels is collected in a database; e. Identifying (for each data set) the number of peaks in the analytical data, and for each peak, identifying a peak value, a start peak value, and an end peak value; f. Generating a trained model for the analysis method according to steps d and e of claim 2; g. Extracting the analytical data from one-dimensional signal data for the sample data set; h. converting the one-dimensional signal data from a signal data set into a two-dimensional array having n channels; i. normalizing the two-dimensional array having n channels of the sample set; j. applying the trained model to the normalized two-dimensional array of the sample set; k. for each peak in the sample data set, identifying the peak value, a start peak value, and an end peak value; 12. A computer-implemented method comprising:
7. The method of claim 1 , wherein the one or more analytical datasets and the sample analytical dataset are obtained by a chromatographic analytical method.
8. The method of claim 1 , wherein the analytical data is represented by a graphical representation.
9. The method of claim 8 , wherein the graphical representation is a chromatogram or a peak profile.
10. The method of any of claims 2 to 9, wherein the two-dimensional array is in the form of a square matrix.
11. 11. The method according to claim 6, wherein the features (values for the number of peaks, peak value, start peak value, and end peak value) obtained in step k) of claim 7 are compared with the features obtained in step e) of claim 7, and a similarity score is assigned for each obtained feature, and the sample data set is accepted when the similarity score for each feature is within a predetermined threshold range, or rejected when the similarity score for each feature is outside the predetermined threshold range.
12. 12. The method of claim 11, wherein the sample is an analytical sample for a batch release.
13. The method according to any of claims 6 to 12, wherein the identified features (values for number of peaks, peak value, start peak value and end peak value) are represented graphically.
14. 1. A method for determining and quantifying the presence of an analyte in a sample, the method comprising: a. obtaining a sample containing at least one analyte; b. subjecting the sample to an analytical method to obtain a sample data set; c. determining peak characteristics of the sample data set in accordance with any one of claims 6 to 13; d. Correlating one or more peaks with the one or more analytes in the sample; e. for each correlated peak, identifying said drop line above a threshold from the baseline or baseline value, the peak start value, and the peak end value; f. Calculating the area for each correlated peak; g. determining and quantifying the presence of each analyte in the sample as correlated with each correlated peak and the area under the correlated peak; A method comprising:
15. A computer system implementing the method of claims 1 to 14, comprising: a. a presentation layer that provides operator access to a user interface and manages and executes functional operations; b. A business tier with logic related to application management and data service management, connected to the data science tier; c. a data science layer that incorporates algorithms that generate trained models, integrated with the business layer; A computer system comprising:
16. 16. The method of claim 15, wherein the computer system is in a modular format that allows for independent upgrades of modules.
17. 17. The computer system of claim 15 or 16, further comprising an input module that imports third party export files representing chromatographic profiles and processing reports and performs data import validity checks to verify the integrity of the data.
18. an input module that imports third-party export files representing chromatographic profiles and processing reports and performs data import validation checks to verify the integrity of said data; b. A model management module that creates or retrains models, selects baseline architectures and corresponding training data, and provides a model performance dashboard; c. A means for an operator to review and validate the trained model; d. A data management module that manages raw data files, model training data, and results by unique identifier; 17. The computer system of claim 16, further comprising:
19. 19. A computer-implemented method for analyzing a sample analysis dataset, said method using a computer system according to claims 15 to 18, comprising: a. accessing said computer system through a presentation layer; b. creating a project session and assigning an identifier through data management logic; c. Importing, via an input module, third-party export files representing chromatographic profiles and processing reports; d. Creating or retraining models through a model management module; e. deploying the trained model to ingest raw data files and perform peak integration; f. Inspecting the integration profile and initiating manual corrections; g. Managing raw data files, model training data, and results by unique identifiers through a data management module; 12. A computer-implemented method comprising:
20. 1. A computer-implemented method for determining predetermined sample characteristics, comprising: determining peak characteristics for a sample analysis dataset according to the method of any one of claims 1 to 14 and 19; b. comparing the peak characteristics of step a) with peak characteristics of a reference standard analytical data set; c. determining the predetermined sample characteristics for the sample analysis dataset; Including, ii) a value determining the concentration or value of an analyte; iii) the presence of an analyte in a sample for diagnosing a medical condition; iv) a value determining the purity of a sample material; v) peak characteristics for monitoring and controlling a manufacturing process; vi) peak characteristics for real-time release testing, automated quality control, process analysis, and digitization in continuous manufacturing; and vii) a variation predictor for enhancing the accuracy of the trained model or processing operation.