Drug API Prediction Method Based on DWT-CARS-MC-PLS

By performing dual-band transformation on drug infrared spectral data and combining it with CARS and Monte Carlo sampling to construct the DWT-CARS-MC-PLS model, the problems of slow speed, low accuracy, and high professional dependence in drug API analysis in existing technologies are solved, enabling rapid and accurate drug quality assessment.

CN115440315BActive Publication Date: 2025-10-31BEIJING UNIV OF TECH
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202211086843.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-09-07
Publication Date
2025-10-31
Estimated Expiration
2042-09-07

AI Technical Summary

Technical Problem

Existing methods for analyzing drug APIs suffer from problems such as slow speed, dependence on data distribution, limited model performance, high barriers to entry for professionals, and lack of adaptability in feature selection. Traditional chemical analysis methods are inefficient and highly destructive to samples, while machine learning methods in infrared spectroscopy analysis suffer from limited model optimization space and insufficient generalization ability.

Method used

A dual-band transform is used to map the one-dimensional spectral data of drugs to a two-dimensional space. The competitive adaptive reweighted sampling algorithm (CARS) is used to screen important wavelengths. Multiple PLSR sub-models are constructed through Monte Carlo sampling. These sub-models are combined using an ensemble learning method to form the DWT-CARS-MC-PLS model. The final result is the average result of the sub-models.

Benefits of technology

It improves the speed and accuracy of drug API analysis, reduces reliance on professionals, enhances the model's adaptability and generalization ability, and achieves more efficient and accurate drug quality assessment.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115440315B_ABST
    Figure CN115440315B_ABST
Patent Text Reader

Abstract

This invention proposes a drug API prediction method based on DWT-CARS-MC-PLS. The method generally consists of three stages: 1. Dual-band transformation (including four methods: differential coefficient DI, ratio coefficient RI, normalized differential coefficient NDI, and integrated two-dimensional correlation spectrum i2DCOS) can solve the problem of limited information in one-dimensional spectra, which is not conducive to modeling. 2. A competitive weighted adaptive sampling strategy is adopted to extract important wavelength features based on the fitting of the partial least squares regression (PLS) model, reducing the subjective error of human feature selection. 3. Using the idea of ​​ensemble learning, T Monte Carlo iterations are used to establish T PLS sub-models, and the mean of the prediction results of these sub-models is used as the prediction result of the entire model. The effectiveness of the proposed method is demonstrated through five repeated experiments with different partitions on a publicly available tablet dataset and comparison with commonly used machine learning methods in drug quantitative analysis.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] Infrared spectroscopy has been widely used in the quantitative detection of pharmaceuticals. Against this background, this invention combines dual-band transformation and competitive adaptive sampling algorithm, and incorporates the idea of ​​ensemble learning to design a drug API prediction method based on DWT-CARS-MC-PLS. Background Technology

[0002] API analysis of pharmaceuticals has always been a key focus and challenge in the field of chemical analysis. Therefore, rapid and accurate quantitative analysis methods are of paramount importance. Precise analytical methods can effectively help pharmaceutical evaluation personnel accurately assess drug quality. Traditional chemical quantitative analysis methods employ high-intensity liquid chromatography (HPLC), which provides highly accurate results, but is slow, time-consuming, and inefficient, and requires destructive sample processing. Infrared spectroscopy, as a rapid and non-destructive analytical method, has been widely applied in the quantitative analysis of pharmaceuticals.

[0003] While infrared spectroscopy offers a more efficient and environmentally friendly analytical method, it often requires specialized personnel to perform spectral analysis, making it a relatively demanding field. A common solution is to employ machine learning methods for quantitative analysis of pharmaceuticals. Currently, commonly used traditional machine learning algorithms include Partial Least Squares Regression (PLSR), Support Vector Machine (SVM), and Ridge Regression (Ridge). These methods have clear objectives, simple model structures, and are relatively fast, largely meeting the needs. However, these methods still have many limitations, such as heavy reliance on data distribution and a limited optimization space that severely restricts model performance.

[0004] In recent years, with the development of deep learning, many researchers have begun to use Convolutional Neural Networks (CNNs) for quantitative analysis of infrared spectral data of pharmaceuticals. Initially, CNNs were efficient models for image processing, but now CNN networks for different dimensions have emerged. For infrared spectral data, researchers can use one-dimensional CNNs (1DCNNs) for target prediction. A CNN mainly consists of convolutional layers, batch normalization layers, pooling layers, activation layers, and dense layers. First, the convolutional, batch normalization, and pooling layers extract spectral features, and then the dense layers, combined with activation functions, complete the prediction. Some researchers also use stacked autoencoders (SAEs) for quantitative analysis of pharmaceutical infrared spectra. An SAE consists of two parts: an encoder and a decoder. The network structure formed by these two parts is called an autoencoder. The SAE extracts the final encoded features through the autoencoder, and then a dense network is used to perform prediction. Unlike CNNs, which calculate gradients, SAEs calculate gradients and update parameters using the error between the decoder and the original input. Therefore, SAEs are an unsupervised feature extraction method. Although this method lacks directional information, experimental results show that it performs comparably to CNNs. SAEs have a wide range of applications in the field of infrared spectroscopy.

[0005] Besides using different machine learning models, feature selection methods are also widely used in the quantitative analysis of pharmaceutical infrared spectroscopy. Commonly used feature selection methods include Principal Component Analysis (PCA), Random Forest (RF), correlation coefficient method, standard deviation selection method (or variance selection method), univariate feature selection (UFS), variable importance in projection (VIP), and interval partial least squares (iPLS). These methods are simple to implement but lack adaptability and generalization ability. Adaptive algorithms have improved this problem. In recent years, more and more scholars have begun to use genetic algorithms (GA), successive projection algorithms (SPA), and competitive adaptive reweighted sampling (CARS). Among them, the competitive adaptive reweighted sampling algorithm has been frequently used in the quantitative analysis of infrared spectroscopy in recent years, and it is an adaptive method with strong transferability and high efficiency.

[0006] This invention proposes a drug API prediction method based on DWT-CARS-MC-PLS. First, a dual-band transform is performed on the one-dimensional spectral data of the drug to map its features to a two-dimensional space. Then, the CARS algorithm is used to screen important wavelengths. Finally, Monte Carlo sampling is used to form different drug sample subsets, and different PLSR sub-models are established. All sub-models are combined to form a DWT-CARS-MC-PLS model. The final prediction result is the average of all sub-models. The model structure is as follows: Figure 1 As shown, the detailed execution steps of the method are as follows:

[0007] The method consists of two phases: data preparation and model building.

[0008] In the data preparation phase, prepare a drug dataset of known API size m×p and a drug dataset of size n×p to be tested, where m and n represent the number of drugs and p represents the number of features of the sample.

[0009] Once the data is prepared, the model building phase begins, with the following specific steps:

[0010] Step 1: Map the one-dimensional feature space of the drug to a two-dimensional space using a two-band transform. The two-band transform includes four forms: difference coefficients (DI), ratio difference coefficients (RI), normalized difference coefficients (NDI), and integrated two-dimensional correlation spectrum (i2DCOS). One of these transforms is selected to complete the feature mapping. The transformation formulas for DI, RI, and NDI are as follows:

[0011] DI(i,j)=R i -R j (1)

[0012] RI(i,j)=R i / R j (2)

[0013] NDI(i,j)=(R i -R j ) / (R i +R j (3)

[0014] Among them, R i R j This represents the i-th and j-th features of the drug data (the absorbance values ​​of the drug at the i-th and j-th wavelengths).

[0015] Two-dimensional correlation spectroscopy is a versatile tool that can extract useful information from a series of spectra collected from a sample under certain physical or chemical stimuli. Cross-correlation analysis of spectral changes caused by perturbation can yield synchronous and asynchronous correlation spectra. Generally, synchronous correlation spectra represent the relative direction of frequency band intensity changes, while asynchronous correlation spectra represent the sequence of frequency band intensity changes. By extending spectral information in the second dimension, 2DCOS can effectively distinguish overlapping spectra and improve apparent spectral resolution.

[0016] The calculation process for the discrete form of the two-dimensional correlation spectrum is as follows:

[0017] For a perturbation t with equal intervals of m, the dynamic spectral intensity at feature v is represented by a column vector y:

[0018]

[0019] Therefore, the synchronous correlation spectrum The expression for the asynchronous correlation spectrum ψ at features i and j can be calculated using the following formula:

[0020]

[0021]

[0022] N is the Hilbert-Noda matrix, and its elements are defined as follows:

[0023]

[0024] The combined two-dimensional correlation spectrum (i2DCOS) can be obtained by multiplying the synchronous correlation spectrum and the asynchronous correlation spectrum. The expression for the combined two-dimensional correlation spectrum at features i and j is as follows:

[0025]

[0026] For practical applications, a reference spectrum r can be used to simulate equally spaced perturbations, and the calculation of the combined two-dimensional correlation spectrum can be simplified to the following form:

[0027]

[0028]

[0029]

[0030] Where t represents the spectrum to be measured. ψ represents the two-dimensional synchronous correlation spectrum, I represents the two-dimensional asynchronous correlation spectrum, and I represents the comprehensive two-dimensional correlation spectrum.

[0031] Step 2: Use CARS to select important infrared spectral features. The execution process of CARS is as follows, and the program flowchart is shown below. Figure 2 :

[0032] Step 1: Perform the i-th (i <= N) Monte Carlo sampling, select a sample subset of drugs, establish a PLSR sub-model, and conduct repeated experiments through the maximum principal component constraint to select the optimal number of principal components.

[0033] Step 2: Calculate the regression coefficient b of the PLS sub-model for each current wavelength feature. j Calculate the importance score w for each feature. j =|b j `| / sum(|bj|)` sorts the data in descending order. It then calculates the feature retention ratio `r`. i =αe -ki ,in

[0034] Step 3: Use p×r i Retain the features with high importance while filtering out features with low weight, and calculate the root mean square error (RMSECV) of the cross-validation of the sub-model under the current feature subset. If i <= N, return to Step 1.

[0035] Step 4: When i > N, select the wavelength feature corresponding to the minimum RMSECV in N Monte Carlo iterations, which is the final result of feature selection.

[0036] Step 3: Construct T PLS sub-models through T Monte Carlo iterations. Each sub-model uses 80% of the samples in the entire training set to form a training subset, and the sampling of the training set is independent of each other in each iteration. Combining the idea of ​​ensemble learning, the T sub-models are combined to form a high-precision and stable DWT-CARS-MC-PLS model. The final prediction result is the average result of each sub-model. Attached Figure Description

[0037] Appendix Figure 1 Overall architecture of the .DWT-CARS-MC-PLS model

[0038] Appendix Figure 2 .CARS Program Flowchart Detailed Implementation

[0039] To demonstrate the effectiveness of the method in this invention, a public dataset of tablet-based drugs was selected for experimentation. Details of the public dataset are as follows:

[0040] The “Tablet” dataset. Near-infrared transmission spectra of active pharmaceutical ingredients (APIs) were first publicly disclosed in a 2002 paper by Dyrby et al., and are available as open source at http: / / www.models.life.ku.dk / plates. This tablet dataset contains 310 samples, with measurements ranging from 7000 to 10500 cm⁻¹. -1 The resolution is 16cm. -1 Each sample contains a total of 404 features. The Kennard-Stone (KS) algorithm was used to divide the dataset into a calibration set with 248 samples and a validation set with 62 samples. The purpose of the analysis was to predict the active pharmaceutical ingredients (APIs) in the drug substance, and the API content (%, w / w) in the dataset was determined by high performance liquid chromatography.

[0041] The specific implementation steps of the drug API prediction method based on DWT-CARS-MC-PLS are as follows:

[0042] Step 1: Map the one-dimensional feature space of the drug to a two-dimensional space using a two-band transform. The two-band transform includes four forms: difference coefficients (DI), ratio difference coefficients (RI), normalized difference coefficients (NDI), and integrated two-dimensional correlation spectrum (i2DCOS). One of these transforms is selected to complete the feature mapping. The transformation formulas for DI, RI, and NDI are as follows:

[0043] DI(i,j)=R i -R j

[0044] RI(i,j)=R i / R j

[0045] NDI(i,j)=(R i -R j ) / (R i +R j )

[0046] Among them, R i R j Represents the i-th and j-th features of the drug data (the absorbance values ​​of the drug at the i-th and j-th wavelengths, i, j <= 404 for dataset A).

[0047]

[0048]

[0049]

[0050] Where t represents the spectrum to be measured. ψ represents the two-dimensional synchronous correlation spectrum, I represents the two-dimensional asynchronous correlation spectrum, and ψ represents the composite two-dimensional correlation spectrum. After dual-band transformation, the feature size of the drug data changes from 1×404 to 404×404.

[0051] Step 2: Use CARS to select important infrared spectral features. The execution process of CARS is as follows, and the program flowchart is shown below. Figure 2 :

[0052] Step 1: Perform the i-th (i <= 50) Monte Carlo sampling, select a sample subset of drugs, build a PLSR sub-model, and perform cross-validation within the maximum principal component constraint to select the optimal number of principal components.

[0053] Step 2: Calculate the regression coefficient b of the PLS sub-model for each current wavelength feature. j Calculate the importance score w for each feature. j =|b j | / sum(|b j Sort in descending order. Calculate the feature retention ratio r. i =αe -ki ,in

[0054] Step 3: Use p×r i Retain the features with high importance while filtering out features with low weight, and calculate the root mean square error (RMSECV) of the cross-validation of the sub-model under the current feature subset. If i <= 50, return to Step 1.

[0055] Step 4: When i > 50, select the wavelength feature corresponding to the minimum RMSECV in 50 Monte Carlo iterations, which is the final result of feature selection.

[0056] Step 3: Construct 500 PLS sub-models through 500 Monte Carlo iterations. Each sub-model uses 80% of the samples in the entire training set to form a training subset, and the sampling of the training set is independent of each other in each iteration. Combining the idea of ​​ensemble learning, T sub-models are combined to form a high-precision and stable model. The final prediction result is the average result of each sub-model.

[0057] To demonstrate the effectiveness of the method, the proposed method was compared with several commonly used models such as PLSR, SVM, CNN, and SAE on the tablet dataset using the same preprocessing and five splits. The experimental results (mean and variance are recorded) are shown in Table 1. The experiments revealed that the i2DCOS transform achieves the best results on the tablet dataset, and adding an ensemble learning strategy further improves the model's accuracy and stability. Table 1: Comparison of experiments on the tablet dataset (%, w / w)

[0058]

Claims

1. The drug API prediction method based on DWT-CARS-MC-PLS mainly includes two stages: data preparation and model building. The method consists of two phases: data preparation and model building.

1. In the data preparation stage, prepare a drug dataset with known APIs of size m×p and a drug dataset to be tested of size n×p, where m and n represent the number of drugs and p represents the number of features of the sample; 2. After the data preparation is complete, proceed to the model building stage. The specific steps are as follows: Step 1: Map the one-dimensional feature space of the drug to a two-dimensional space using a two-band transform. The two-band transform includes four forms: difference coefficients (DI), ratio difference coefficients (RI), normalized difference coefficients (NDI), and integrated two-dimensional correlation spectrum (i2DCOS). Select one of these transforms to complete the feature mapping. The transformation formulas for DI, RI, and NDI are as follows: DI(i,j)=R i -R j (1) RI(i,j)=R i / R j (2) AND(i,j)=(R i -R j ) / (R i +R j ) (3) Among them, R i R j Represents the i-th and j-th features of the drug data (the absorbance values ​​of the drug at the i-th and j-th wavelengths); Two-dimensional correlation spectroscopy is a versatile tool that can extract useful information from a series of spectra collected from a sample under certain physical or chemical stimuli. Cross-correlation analysis of spectral changes caused by perturbation can yield synchronous and asynchronous correlation spectra. Generally, synchronous correlation spectra represent the relative direction of frequency band intensity changes, while asynchronous correlation spectra represent the sequence of frequency band intensity changes. By extending spectral information in the second dimension, 2DCOS can effectively distinguish overlapping spectra and improve apparent spectral resolution. The calculation process for the discrete form of the two-dimensional correlation spectrum is as follows: For a perturbation t with equal intervals of m, the dynamic spectral intensity of the spectral feature at feature v is represented by a column vector y: Therefore, the synchronous correlation spectrum The expressions for the asynchronous correlation spectrum ψ at features i and j can be calculated using the following formula: N is the Hilbert-Noda matrix, and its elements are defined as follows: The combined two-dimensional correlation spectrum (i2DCOS) can be obtained by multiplying the synchronous correlation spectrum and the asynchronous correlation spectrum. The expression for the combined two-dimensional correlation spectrum at wavelength features i and j is as follows: For practical applications, a reference spectrum r can be used to simulate equally spaced perturbations, and the calculation of the combined two-dimensional correlation spectrum can be simplified to the following form: Where t represents the spectrum to be measured. ψ represents the two-dimensional synchronous correlation spectrum, I represents the two-dimensional asynchronous correlation spectrum, and I represents the comprehensive two-dimensional correlation spectrum. Step 2: Use CARS to select important infrared spectral wavelength features. The execution process of CARS is as follows, and the program flowchart is shown in Figure 2: Step 1: Perform the i-th (i <= N) Monte Carlo sampling, select a sample subset of drugs, build a PLSR sub-model, and perform cross-validation within the maximum principal component constraint to select the optimal number of principal components; Step 2: Calculate the regression coefficient b of the PLS sub-model for each current wavelength feature. j Calculate the importance score w for each feature. j =|b j `| / sum(|bj|)` sorts the data in descending order; calculates the feature retention ratio `r`. i =αe -ki ,in This ensures that the input space contains all features in the first iteration, and only two wavelength features in the last iteration. Step 3: Use p×r i Retain the features with high importance, while filtering out the features with low weight, and calculate the root mean square error (RMSECV) of the cross-validation of the sub-model under the current feature subset; if i <= N, then return to Step 1; Step 4: When i > N, select the wavelength feature corresponding to the minimum RMSECV in N Monte Carlo iterations, which is the final result of feature selection; Step 3: Construct T PLS sub-models through T Monte Carlo iterations. Each sub-model uses 80% of the samples in the entire training set to form a training subset, and the sampling of the training set is independent of each other in each iteration. Combining the idea of ​​ensemble learning, the T sub-models are combined to form a high-precision and stable DWT-CARS-MC-PLS model. The final prediction result is the average result of each sub-model.

Citation Information

Patent Citations

  • Method for drug storage data processing by correlation coefficients

    CN106979934A

  • Tomato gray mold degree identification method and device

    CN113902673A