Drug effect prediction model training method and system based on machine learning

Through the machine learning-based efficacy prediction model training method, Lasso regression and response surface methodology are used to determine the key process condition combinations, construct feature data sets and train efficacy prediction models, which solves the problem of complex nonlinear relationships in the efficacy prediction of traditional Chinese medicine and achieves accurate efficacy prediction and optimization.

CN120784008APending Publication Date: 2025-10-14WUHAN BOKE GUOTAI INFORMATION TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510901698.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-01
Publication Date
2025-10-14

AI Technical Summary

Technical Problem

It is difficult to effectively capture the complex nonlinear relationship between the processing technology and the efficacy of traditional Chinese medicine in the prediction of its efficacy. Traditional methods are unable to extract key features from massive data, resulting in insufficient prediction accuracy and poor repeatability.

Method used

A machine learning-based pharmacodynamic prediction model training method constructs an initial data set by extracting process parameters and chemical composition data from historical preparation records. Lasso regression and response surface methodology are used to determine the key process condition combinations, construct a feature data set, and train the pharmacodynamic prediction model. Error analysis is then performed to optimize pharmacodynamic prediction.

Benefits of technology

Effectively explore the intrinsic relationship between Chinese medicine processing parameters and efficacy changes, establish an accurate efficacy prediction model, and improve prediction accuracy and repeatability.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120784008A_ABST
    Figure CN120784008A_ABST
Patent Text Reader

Abstract

The invention provides a drug effect prediction model training method and system based on machine learning, and relates to the field of drug effect prediction.The method comprises the steps that process parameter data, chemical component data and chemical component change data are extracted from historical processing records, and an initial data set is constructed; and analyzing relevance among the process parameter data, the chemical component data and the chemical component change data based on the initial data set, and determining a key process condition combination. And constructing a feature data set based on the key process condition combination, training a drug effect prediction model through the training set to obtain a first drug effect prediction model, and predicting the test set through the first drug effect prediction model to output preliminary drug effect change prediction data. And performing error analysis on the preliminary drug effect change prediction data, and determining a parameter adjustment direction of the first drug effect prediction model. And determining a second drug effect prediction model based on the parameter adjustment direction of the first drug effect prediction model, wherein the second drug effect prediction model is used for outputting optimized drug effect change prediction data.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The embodiments of the present application relate to the field of drug efficacy prediction, and specifically to a drug efficacy prediction model training method and system based on machine learning. Background Art

[0002] As an important component of traditional Chinese medicine, Traditional Chinese Medicine (TCM) has been widely recognized for its efficacy and application value. However, the complexity of TCM makes its efficacy prediction and process optimization a significant challenge. The efficacy of TCM depends not only on the synergistic effects of its multi-component chemical components but is also significantly influenced by the processing technology (such as temperature, time, and treatment methods). Changes in these process parameters often lead to nonlinear changes in the drug components, which in turn affect their pharmacological activity. Therefore, establishing accurate efficacy prediction models and deeply exploring the relationship between processing parameters and efficacy changes are of great significance to promoting the modernization, standardization, and scientific development of TCM.

[0003] Currently, the prediction of TCM efficacy primarily relies on manual experience or traditional statistical analysis methods, such as regression analysis and principal component analysis. These methods struggle to effectively capture the complex, nonlinear relationship between preparation techniques and efficacy, and are unable to extract key features from massive amounts of data, resulting in insufficient prediction accuracy and limited generalization capabilities. Furthermore, the diverse composition of TCM ingredients and the dynamic relationship between process parameters and efficacy further complicate traditional modeling approaches, limiting the scientific nature and reproducibility of efficacy predictions. Summary of the Invention

[0004] In view of the above problems, the embodiments of the present application provide a method and system for training a drug efficacy prediction model based on machine learning, which overcomes or at least partially solves the above problem of difficulty in establishing an accurate drug efficacy prediction model and deeply explores the relationship between preparation process parameters and drug efficacy changes.

[0005] The first aspect of the embodiment of the present application provides a method for training a drug efficacy prediction model based on machine learning, comprising: extracting process parameter data, chemical composition data and chemical composition change data from historical processing records to construct an initial data set. Based on the initial data set, the correlation between the process parameter data, chemical composition data and chemical composition change data is analyzed to determine a key process condition combination, wherein the key process condition combination includes key process parameters and chemical composition change rates. Based on the key process condition combination, a feature data set is constructed, the feature data set is divided into a training set and a test set, the drug efficacy prediction model is trained with the training set to obtain a first drug efficacy prediction model, and the test set is predicted with the first drug efficacy prediction model to output preliminary drug efficacy change prediction data. Error analysis is performed on the preliminary drug efficacy change prediction data to determine the parameter adjustment direction of the first drug efficacy prediction model. Based on the parameter adjustment direction of the first drug efficacy prediction model, a second drug efficacy prediction model is determined, and the second drug efficacy prediction model is used to output optimized drug efficacy change prediction data.

[0006] In this embodiment, an initial data set is constructed by extracting process parameter data and chemical composition change data, a feature data set is constructed based on a combination of key process conditions, the relationship between process parameters and efficacy changes is learned using a first efficacy prediction model, and error analysis is performed on the preliminary efficacy change prediction data, and a second efficacy prediction model is determined to output optimized efficacy change prediction data. In this way, the intrinsic relationship between the Chinese medicine preparation process parameters and efficacy changes can be effectively explored, and an accurate efficacy prediction model can be established.

[0007] In an optional manner, before analyzing the correlation between process parameter data, chemical composition data and chemical composition change data based on the initial data set and determining a set of key process parameters that are highly correlated with the chemical composition change data, the method also includes: performing data cleaning and standardization on the initial data set to generate a standardized data set.

[0008] In an optional manner, before analyzing the correlation between process parameter data, chemical composition data and chemical composition change data based on the initial data set and determining a set of key process parameters that are highly correlated with the chemical composition change data, the method also includes: performing data cleaning and standardization on the initial data set to generate a standardized data set.

[0009] In one optional approach, the initial dataset is cleaned and normalized to generate a standardized dataset, including: filling missing values ​​in the initial dataset using an interpolation method to obtain a first dataset; removing abnormal data in the first dataset to obtain a second dataset; normalizing the second dataset to obtain a third dataset; and performing consistency checking and adjustments on the third dataset to obtain a fourth dataset.

[0010] In an optional method, the correlation between process parameter data, chemical composition data and chemical composition change data is analyzed based on the initial data set to determine the key process condition combination, including: screening key parameters through Lasso regression, building a quadratic regression model through response surface methodology to determine extreme points, and determining the key process condition combination through genetic algorithm based on key parameters and extreme points.

[0011] In an optional manner, performing error analysis on the preliminary drug efficacy change prediction data to determine the parameter adjustment direction of the first drug efficacy prediction model includes:

[0012] Calculate the deviation of the preliminary efficacy change prediction data minus the actual efficacy change data, and compare the deviation with the preset error threshold. If the deviation exceeds the error threshold, mark the data point as "needs adjustment". If the deviation is positive, reduce the value of the key process parameter in the first efficacy prediction model; otherwise, increase the value of the key process parameter in the first efficacy prediction model.

[0013] In an optional manner, after determining the second pharmacodynamic prediction model based on the parameter adjustment direction of the first pharmacodynamic prediction model, the method further includes: comparing the optimized pharmacodynamic change prediction data with the pharmacodynamic change data under actual process conditions to generate a suitability analysis result.

[0014] In an optional manner, after comparing the optimized efficacy change prediction data with the efficacy change data under actual process conditions to generate a suitability analysis result, the method further includes:

[0015] Based on the results of the applicability analysis, a report on the correlation between process parameters and drug efficacy changes is generated and displayed.

[0016] In an optional manner, the association report is presented through a three-dimensional scatter plot and a heat map.

[0017] In an optional method, the optimized efficacy change prediction data and the efficacy change data under actual process conditions are compared to generate a suitability analysis result, including: calculating the root mean square error between the optimized efficacy change prediction data and the efficacy change data under actual process conditions; if the root mean square error is higher than a preset threshold, performing a feature importance analysis to quantify the weight of each process parameter; dividing each process parameter into intervals, and calculating the root mean square error of each interval; marking the process parameter interval with the highest prediction accuracy based on the root mean square error of each interval, and determining the suitability analysis result based on the process parameter interval with the highest prediction accuracy.

[0018] The second aspect of the embodiments of the present application provides a machine learning-based pharmacodynamic prediction model training system, comprising: a construction module for extracting process parameter data, chemical composition data, and chemical composition change data from historical processing records to construct an initial data set. A first determination module for analyzing the correlation between the process parameter data, chemical composition data, and chemical composition change data based on the initial data set to determine a key process condition combination, wherein the key process condition combination includes key process parameters and chemical composition change rates. A training module for constructing a feature data set based on the key process condition combination, dividing the feature data set into a training set and a test set, training the pharmacodynamic prediction model with the training set to obtain a first pharmacodynamic prediction model, and predicting the test set with the first pharmacodynamic prediction model to output preliminary pharmacodynamic change prediction data. An adjustment module for performing error analysis on the preliminary pharmacodynamic change prediction data to determine the parameter adjustment direction of the first pharmacodynamic prediction model. A second determination module for determining a second pharmacodynamic prediction model based on the parameter adjustment direction of the first pharmacodynamic prediction model, the second pharmacodynamic prediction model being used to output optimized pharmacodynamic change prediction data.

[0019] The above description is only an overview of the technical solutions of the embodiments of the present application. In order to more clearly understand the technical means of the embodiments of the present application, they can be implemented in accordance with the contents of the specification. In order to make the above and other purposes, features and advantages of the embodiments of the present application more obvious and easy to understand, the specific implementation methods of the present application are listed below. BRIEF DESCRIPTION OF THE DRAWINGS

[0020] In order to more clearly illustrate the technical solutions of the embodiments of the present application, the following is a brief introduction to the drawings required for use in the description of the embodiments. Obviously, the drawings described below are some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without any creative work.

[0021] Figure 1 A flowchart of a method for training a drug efficacy prediction model based on machine learning is provided in some embodiments of the present application.

[0022] Figure 2 Another flowchart of a method for training a drug efficacy prediction model based on machine learning provided in some embodiments of the present application.

[0023] Figure 3 Another embodiment of the present application provides a drug efficacy prediction model training system based on machine learning. DETAILED DESCRIPTION

[0024] To make the purpose, technical solutions, and advantages of the embodiments of this application more clear, the technical solutions in the embodiments of this application will be clearly and completely described below in conjunction with the accompanying drawings in the embodiments of this application. Obviously, the described embodiments are part of the embodiments of this application, not all of the embodiments. Based on the embodiments in this application, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of this application.

[0025] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by those skilled in the art to which this application belongs; the terms used in the specification of the application herein are only for the purpose of describing specific embodiments and are not intended to limit this application.

[0026] The terms "comprises", "comprising" and "having" and any variations thereof in the specification, claims and drawings of this application are intended to cover but not exclude other contents. The word "a" or "an" does not exclude the presence of a plurality.

[0027] References to "embodiments" herein mean that a particular feature, structure, or characteristic described in connection with the embodiments may be included in at least one embodiment of the present application. The appearance of the phrase "embodiment" in various places in the specification does not necessarily refer to the same embodiment, nor does it necessarily refer to independent or alternative embodiments that are mutually exclusive of other embodiments. It is understood, both explicitly and implicitly, by those skilled in the art that the embodiments described herein may be combined with other embodiments.

[0028] In addition, the terms "first", "second", etc. in the description and claims of this application or the above-mentioned drawings are used to distinguish different objects, rather than to describe a specific order, and may explicitly or implicitly include one or more such features.

[0029] In the description of this application, unless otherwise specified, "plurality" means more than two (including two), and similarly, "multiple groups" means more than two (including two).

[0030] In the description of this application, it should be noted that, unless otherwise expressly specified or limited, the terms "mounted," "connected," and "connected" should be understood broadly. For example, "connected" or "connected" in a mechanical structure can refer to a physical connection. For example, a physical connection can be a fixed connection, such as a fixed connection via a fixing member, such as a screw, bolt, or other fixing member. A physical connection can also be a detachable connection, such as a mutual snap-fit ​​connection. A physical connection can also be an integral connection, such as a connection formed by welding, bonding, or integral molding. "Connected" or "connected" in a circuit structure can refer not only to a physical connection but also to an electrical connection or a signal connection. For example, it can be a direct connection, i.e., a physical connection, or an indirect connection through at least one intermediate element, as long as the circuit is interconnected. It can also refer to internal communication between two elements. A signal connection can refer to a signal connection through a circuit or a signal connection through a media medium, such as radio waves. Those skilled in the art will understand the specific meanings of the above terms in this application.

[0031] This application embodiment provides a method for training a drug efficacy prediction model based on machine learning. Figure 1 , Figure 1 A flowchart of a method for training a drug efficacy prediction model based on machine learning is provided in some embodiments of the present application.

[0032] like Figure 1 As shown, the drug efficacy prediction model training method based on machine learning provided in the embodiment of the present application includes the following steps 101 to 105:

[0033] Step 101: extract process parameter data, chemical composition data, and chemical composition change data from historical processing records to construct an initial data set.

[0034] Specifically, you can use a data crawler to retrieve records containing process parameter data such as temperature and time from the database. Assume that the records are in CSV format and contain the fields "Temperature (°C)", "Time (minutes)", and "Batch Number". Use Python's pandas library to read the CSV file and filter records containing temperature (e.g., 100°C to 300°C) and time (e.g., 5 to 60 minutes). For example, a record with "Temperature 150°C, Time 30 minutes" can be read using df = pd.read_csv('records.csv') and df = df.dropna().

[0035] Assuming that a record contains "Component A content (%)" and "Component B content (%)," the "batch number" can be used to associate with process parameter data. For example, the SQL query SELECT * FROM composition WHERE batch_id IN (SELECT batch_id FROM process) can be used to form a data set containing fields such as temperature, time, component A, and component B. Moreover, it is necessary to cover a variety of process conditions. For example, the temperature can be 100°C, 150°C, 200°C, 250°C, and 300°C; the time can be 5 minutes, 10 minutes, 30 minutes, and 60 minutes; the component A content can range from 5% to 20%, and so on.

[0036] In practical applications, to ensure the comprehensiveness of the data, statistical analysis can be used to check the distribution of the data. For example, the mean and standard deviation of temperature and time can be calculated, and outliers can be detected through box plots. Assume that the mean temperature is 180.5°C and the standard deviation is 20.3°C; the mean time is 25.7 minutes and the standard deviation is 10.2 minutes. If a record with a temperature of 350°C is found, it is considered that it deviates too much from the standard deviation and is removed. In addition, the k-means clustering algorithm can be used with k=5 to cluster the temperature and time process parameters and divide the process parameter intervals. For example, cluster 1 is the center, the temperature is 120°C, and the time is 10 minutes. Cluster 2 is the center, the temperature is 200°C, and the time is 40 minutes, to ensure coverage of diverse process conditions.

[0037] For example, the constructed initial data set can be saved in CSV format, containing 1000 records, covering temperature 100-300°C, time 5-60 minutes, component A 5-20%, and component B 10-30%.

[0038] In some embodiments, the Pearson correlation coefficient can be used to calculate the correlation between temperature and component A. For example, the correlation coefficient r=0.75 and the significance level p<0.05 indicate that increasing temperature may promote the increase of component A.

[0039] Step 102 : analyzing the correlation between the process parameter data, the chemical composition data, and the chemical composition change data based on the initial data set to determine a key process condition combination, wherein the key process condition combination includes key process parameters and chemical composition change rates.

[0040] Specifically, based on the initial data set, the correlation between process parameter data, chemical composition data and chemical composition change data is analyzed to determine the key process condition combination, including: screening key parameters through Lasso regression; constructing a quadratic regression model through response surface methodology to determine extreme points; and determining the key process condition combination through genetic algorithm based on key parameters and extreme points.

[0041] Exemplarily, feature extraction can be performed by principal component analysis first, assuming that the initial data set contains 10 parameters such as temperature (℃), time (minute), pressure (MPa), flow (L / min), etc. The principal component analysis algorithm can reduce the data to 2 principal components, retaining 80% of the variance, to obtain a principal component loading matrix, in which the temperature and time loadings are 0.65 and 0.55 respectively, and the loadings of other parameters such as pressure are only 0.2, indicating that temperature and time have a greater contribution to chemical composition change. Then, Lasso regression is used to screen key parameters, and the regularization parameter α is set to 0.01. The screening results show that the regression coefficients of temperature, time and flow are 0.8, 0.6 and 0.3 respectively, and the coefficients of the remaining parameters are close to 0, confirming that these three parameters are significantly related to chemical composition change. To analyze the effects of temperature and time, a quadratic regression model can be constructed using response surface methodology, setting the temperature range to 200-400℃ and the time range to 10-60 minutes. The fitted equation is:

[0042] Y = 2.5 + 0.3T + 0.2t - 0.01T 2 -0.005t 2 + 0.002Tt.

[0043] where Y is the chemical composition change rate, T is the temperature, and t is the time. By solving the partial derivative, the extreme point temperature is found to be 300℃ and the time is 35 minutes, and Y reaches a maximum value of 3.2%. To determine the key process condition combination, genetic algorithm optimization is used, with a population size of 100 and 50 iterations, the objective function is to maximize Y, and the constraint conditions are temperature 200-400℃, time 10-60 minutes, and flow 5-20 L / min. The final key process condition combination is temperature 298℃, time 36 minutes, and flow 12 L / min, and Y is 3.25%. The above process can realize principal component analysis and Lasso regression through the sklearn library of Python, the response surface method can use Design-Expert software, and the genetic algorithm is based on the DEAP framework, which can ensure the reproducibility of data results.

[0044] In step 103, a feature data set is constructed based on the key process condition combination, the feature data set is divided into a training set and a test set, a first drug efficacy prediction model is trained by the training set, and preliminary drug efficacy change prediction data is output by predicting the test set through the first drug efficacy prediction model.

[0045] Assume the key process condition combination is: temperature 298°C, time 36 minutes, flow rate 12 L / min, and Y is 3.25%. To analyze the interaction effects between chemical composition changes and process parameters, a polynomial feature generation method (PolynomialFeatures, degree = 2) is used to generate an extended feature matrix containing interaction terms, such as the interaction term between temperature and Y, and the interaction term between temperature, time, and Y. This is used to construct a feature dataset.

[0046] Dividing the feature dataset into a training set and a test set may include splitting the feature dataset into an 80% training set and a 20% test set.

[0047] For example, the first efficacy prediction model can use the random forest algorithm as a machine learning model, using the Random Forest Regressor module in the Scikit-learn library, setting the number of trees to 100 and the maximum depth to 10. The model performance is evaluated through 5-fold cross-validation, and the average mean square error is 2.3, indicating that the model prediction accuracy is high. Subsequently, the feature data set is divided into an 80% training set and a 20% test set. The model is fitted on the training set to learn the nonlinear relationship between temperature, pressure, and reaction time and the drug activity value. For example, it is found that the activity value reaches a peak of 85% when the temperature is around 50 degrees Celsius, and the activity value decreases by about 10% after the pressure exceeds 4.0 MPa. Finally, the trained model is used to predict the test set to generate preliminary efficacy change prediction data.

[0048] Step 104 : performing error analysis on the preliminary drug efficacy change prediction data to determine the parameter adjustment direction of the first drug efficacy prediction model.

[0049] Specifically, an error analysis is performed on the preliminary pharmacodynamic change prediction data to determine the parameter adjustment direction of the first pharmacodynamic prediction model, including: calculating the deviation of the preliminary pharmacodynamic change prediction data minus the actual pharmacodynamic change data, comparing the deviation with a preset error threshold, if the deviation exceeds the error threshold, marking the data point as "needs adjustment", if the deviation is a positive value, reducing the value of the key process parameter in the first pharmacodynamic prediction model, otherwise, increasing the value of the key process parameter in the first pharmacodynamic prediction model.

[0050] For example, the preliminary predicted efficacy change for a drug is 15.2%, while the actual efficacy change data shows a change of 14.5%. Assuming the key process parameters for this drug's key process conditions are temperature 298°C, time 36 minutes, and flow rate 12 L / min, a computer algorithm is used to calculate the deviation. The deviation calculation formula is: Deviation = Preliminary Efficacy Change Prediction - Actual Efficacy Change, i.e., 15.2 - 14.5 = 0.7%. This deviation is then compared with a preset error threshold. Assuming the preset error threshold is set to 1.0%, the program determines that 0.7% < 1.0%, concluding that the deviation meets the requirements. This result is recorded in the database for subsequent analysis. If the deviation exceeds the preset error threshold, for example, if the deviation is 1.5%, the data point is marked as "needs adjustment." Assuming the deviation is positive, that is, the preliminary efficacy change prediction data is higher than the actual efficacy change data, the preset rules determine that the first efficacy prediction model may overestimate the efficacy change, and automatically generate adjustment suggestions: reduce the values ​​of key process parameters in the first efficacy prediction model, for example, reduce the temperature from 298°C to 297°C, reduce the time from 36 minutes to 35 minutes, and reduce the flow rate from 12 L / min to 11.9 L / min.

[0051] Step 105 : determining a second drug efficacy prediction model based on the parameter adjustment direction of the first drug efficacy prediction model, and the second drug efficacy prediction model is used to output optimized drug efficacy change prediction data.

[0052] For example, assuming that the first pharmacodynamic prediction model is XGBoost, the learning rate is set to 0.05, the maximum tree depth is 6, and the initial parameters can be based on empirical values ​​to capture the nonlinear characteristics of pharmacodynamic changes. The first pharmacodynamic prediction model is input with the half-maximal inhibitory concentration IC50 values ​​of 1000 compounds in nanomolar (nM), with a numerical range of 0.1-100, a molecular weight range of 100 to 500, and a lipophilicity logP of -2 to 5. Use MinMaxScaler to normalize the features to the [0,1] interval to ensure the convergence stability of the first pharmacodynamic prediction model. Then, 5-fold cross-validation is used to evaluate the model performance, and the loss function is the mean square error. After the initial training, the mean square error of the test set was 0.15. By analyzing the importance of the features, it was found that logP contributed the most to the pharmacodynamic prediction. For example, the logP weight was 0.35. Based on this, XGBoost parameters were adjusted, increasing the regularization coefficient L2 to 0.1 to prevent overfitting, and the subsampling ratio was adjusted from 0.8 to 0.6 to reduce the noise sensitivity of the first efficacy prediction model. After five iterative training cycles, the mean squared error (MSE) dropped to 0.08, indicating improved prediction accuracy. Further SHAP value analysis revealed that compounds with high logP values ​​exhibited significant prediction bias, presumably due to uneven data distribution. Therefore, the SMOTE algorithm was introduced to balance the dataset, increasing the number of samples for compounds with high logP values ​​to 300. After further training, the MSE stabilized at 0.06, and the mean absolute error of the predicted IC50 value dropped to 0.04 μM, indicating improved accuracy.

[0053] In this embodiment, an initial data set is constructed by extracting process parameter data and chemical composition change data, a feature data set is constructed based on a combination of key process conditions, the relationship between process parameters and efficacy changes is learned using a first efficacy prediction model, and error analysis is performed on the preliminary efficacy change prediction data, and a second efficacy prediction model is determined to output optimized efficacy change prediction data. In this way, the intrinsic relationship between the process parameters of Chinese medicine preparation and efficacy changes can be effectively explored.

[0054] In some embodiments, before analyzing the correlation between process parameter data, chemical composition data, and chemical composition change data based on the initial data set and determining a set of key process parameters that are highly correlated with the chemical composition change data, the method also includes: performing data cleaning and standardization on the initial data set to generate a standardized data set.

[0055] Specifically, the initial dataset is cleaned and normalized to generate a standardized dataset, including: filling missing values ​​in the initial dataset using interpolation to obtain a first dataset; removing abnormal data in the first dataset to obtain a second dataset; normalizing the second dataset to obtain a third dataset; and performing consistency checking and adjustments on the third dataset to obtain a fourth dataset.

[0056] In this embodiment, a preset statistical rule is used to detect the location of missing values. If missing values ​​are detected, they are filled in by interpolation to obtain a filled first data set. For the first data set, the numerical distribution characteristics are obtained. If the values ​​are found to exceed the preset threshold range, they are marked as abnormal data and eliminated to obtain a cleaned second data set. Based on the cleaned second data set, a standardization operation is used to adjust the data scale, and the data is normalized by calculating the mean and standard deviation to obtain a standardized third data set. By performing a consistency check on the third data set, if the check result shows that the data distribution is uneven, the abnormal distribution part is adjusted twice to obtain a fourth data set with enhanced consistency.

[0057] For example, assume that the initial dataset contains three columns: temperature (°C), pressure (MPa), and flow rate (L / min). Some data is missing and contains outliers. First, examine the initial dataset and assume that the temperature column has 5% missing values, the pressure column has 3% missing values, and the flow column has no missing values. Linear interpolation is used to fill missing values. For example, for the temperature series [25.0, NaN, 27.0, 28.0], linear interpolation is used to calculate the NaN value as (25.0 + 27.0) / 2 = 26.0. Similar processing is repeated for all missing values ​​to ensure data integrity. Next, detect outliers. Using the 3σ rule based on standard deviation, the mean of the temperature column is calculated as μ = 26.5, with a standard deviation σ = 1.2. Values ​​outside the range [μ - 3σ, μ + 3σ] = [22.9, 30.1] are removed. For example, values ​​32.0 are removed, resulting in a 2% percentage of removed data. Repeat this process for pressure and flow, assuming that the pressure has 1% outliers and the flow has no outliers. After these exclusions, the number of rows in the dataset was reduced from 1000 to 980. To ensure consistency, the data was normalized using the Z-score normalization formula: z = (x - μ) / σ. For example, a temperature value of 27.0 is normalized to (27.0 - 26.5) / 1.2 ≈ 0.4167. Pressure and flow rates were processed similarly, resulting in a normalized dataset with a mean of 0 and a standard deviation of 1. Finally, a fourth dataset containing 980 rows of normalized data was generated.

[0058] It is worth noting that in this embodiment, step 102 analyzes the correlation between the process parameter data, chemical composition data and chemical composition change data based on the initial data set to determine the key process condition combination, which can be understood as analyzing the correlation between the process parameter data, chemical composition data and chemical composition change data based on the fourth data set to determine the key process condition combination.

[0059] refer to Figure 2 , Figure 2Another flowchart of a machine learning-based drug efficacy prediction model training method provided for some embodiments of the present application. In this embodiment, after step 105, steps 106 and 107 can also be included.

[0060] Step 106: Compare the optimized drug efficacy change prediction data with the drug efficacy change data under actual process conditions to generate applicability analysis results.

[0061] Specifically, comparing the optimized drug efficacy change prediction data with the drug efficacy change data under actual process conditions to generate applicability analysis results includes: calculating the root mean square error of the optimized drug efficacy change prediction data and the drug efficacy change data under actual process conditions. If the root mean square error is higher than a preset threshold, perform feature importance analysis to quantify the weight of each process parameter. Divide each process parameter by interval, and calculate the root mean square error of each interval. Based on the root mean square error of each interval, mark the process parameter interval with the highest prediction accuracy, and determine the applicability analysis results based on the process parameter interval with the highest prediction accuracy.

[0062] For example, when the temperature is 50°C and the time is 2 hours, the optimized drug efficacy change prediction data is 85%, while the drug efficacy change data measured under the same parameters under actual process conditions is 82%. Calculate the root mean square error (RMSE) of the optimized drug efficacy change prediction data and the drug efficacy change data measured under the same parameters under actual process conditions, the formula is:

[0063] Where y 1,a is the a-th optimized drug efficacy change prediction data, y 2,aFor the same parameter measured under the a-th actual process condition, n is the total number of data sets. Taking 100 data sets as an example, the RMSE is calculated to be 2.5%, and if this root mean square error is higher than the preset threshold, feature importance analysis can be used. Through the feature_importance_ method of the random forest model, the importance of temperature is 0.65 and the importance of time is 0.35, indicating that temperature contributes more to the prediction of drug efficacy changes. Further, based on stratified analysis, temperature is divided into four intervals: 30-40°C, 40-50°C, 50-60°C, and 60-70°C. The RMSE of each interval is calculated. For example, the RMSE of the 50-60°C interval is 2.1%, while the RMSE of the 30-40°C interval is 3.2%, indicating that the second drug efficacy prediction model is more accurate in the higher temperature interval. Similarly, time stratification analysis shows that the 2-3 hour interval has the lowest RMSE (2.0%). Finally, the suitability analysis results are generated, and the relationship between temperature, time, and RMSE is plotted using the matplotlib library. Statistical tests (such as t-test, p-value < 0.05) confirm that temperature has a significant impact on prediction accuracy. The suitability analysis results are: the prediction accuracy is highest under the conditions of 50-60°C and 2-3 hours, and it is recommended to prioritize adjusting the temperature parameter during process optimization. To ensure business relevance, the suitability analysis results can also be connected to the process control system, and the recommended parameters (such as temperature 55°C, time 2.5 hours) can be transmitted to the production equipment through the API to optimize drug efficacy stability.

[0064] Step 107: Based on the suitability analysis results, generate a report on the correlation between process parameters and drug efficacy changes and display it.

[0065] In some embodiments, the correlation report can be displayed through a three-dimensional scatter plot and a heat map.

[0066] Suppose the suitability analysis results obtained are: the temperature range is 50 to 80 degrees Celsius, the stirring speed range is 200 to 500 revolutions per minute, and the drug efficacy indicator is the release rate of active ingredients (percentage), ranging from 60% to 95%. Using a multiple regression analysis algorithm, a relationship model between temperature, stirring speed, and release rate is constructed, with the formula: release rate = 35.2 + 0.65 * temperature + 0.08 * stirring speed. The determination coefficient R 2 in the multiple regression analysis algorithm for evaluating the goodness of fit of the model is 0.85, indicating that this relationship model has strong explanatory power. The analysis process shows that when the temperature increases from 60 to 70 degrees Celsius, the release rate increases by about 6.5%, and when the stirring speed increases from 300 to 400 revolutions per minute, the release rate increases by about 8%, indicating that both have a significant positive impact on drug efficacy.

[0067] Next, data visualization tools were used to generate three-dimensional scatter plots and heat maps, plotted using Python's Matplotlib library. The horizontal axis represents temperature, and the vertical axis represents stirring speed. Lighter and darker colors indicate higher and lower release rates. The figure clearly shows that the release rate peaks at approximately 85% at a temperature of 70°C and a stirring speed of 400 rpm. Finally, the Apriori algorithm was used to identify parameter combination rules and generate a report correlating process parameters with changes in efficacy. The confidence level for the release rate exceeding 80% for temperatures between 65 and 75°C and stirring speeds between 350 and 450 rpm was 0.9, with a support level of 0.7. This report correlating process parameters with changes in efficacy indicates that the aforementioned parameter ranges are within the recommended range for process optimization. This method forms a complete logical chain from data modeling to visualization to rule mining. This correlation report can be directly used for process optimization and guide subsequent adjustments to production parameters.

[0068] Another embodiment of the present application further provides a drug efficacy prediction model training system based on machine learning, Figure 3 Another embodiment of the present application provides a drug efficacy prediction model training system based on machine learning, referring to Figure 3 A drug efficacy prediction model training system 3 based on machine learning includes: a construction module 31, which is used to extract process parameter data, chemical component data, and chemical component change data from historical processing records to construct an initial data set. A first determination module 32, which is used to analyze the correlation between the process parameter data, chemical component data, and chemical component change data based on the initial data set to determine a key process condition combination, wherein the key process condition combination includes key process parameters and chemical component change rates. A training module 33, which is used to construct a feature data set based on the key process condition combination, divide the feature data set into a training set and a test set, train the drug efficacy prediction model with the training set to obtain a first drug efficacy prediction model, and predict the test set with the first drug efficacy prediction model to output preliminary drug efficacy change prediction data. An adjustment module 34, which is used to perform error analysis on the preliminary drug efficacy change prediction data and determine the parameter adjustment direction of the first drug efficacy prediction model. A second determination module 35, which is used to determine a second drug efficacy prediction model based on the parameter adjustment direction of the first drug efficacy prediction model, and the second drug efficacy prediction model is used to output optimized drug efficacy change prediction data.

[0069] The machine learning-based efficacy prediction model training system provided in the embodiments of the present application can effectively explore the intrinsic relationship between Chinese medicine processing parameters and efficacy changes, and establish an accurate efficacy prediction model.

[0070] Those skilled in the art will appreciate that, although some embodiments herein include certain features included in other embodiments, combinations of features from different embodiments are intended to be within the scope of this application and to form different embodiments. For example, in the claims, any one of the claimed embodiments may be used in any combination.

[0071] As described above, the above embodiments are only used to illustrate the technical solutions of the present application, rather than to limit them. Although the present application has been described in detail with reference to the above embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the above embodiments, or make equivalent replacements for some of the technical features therein. However, these modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the embodiments of the present application.

Claims

1. A method for training a drug efficacy prediction model based on machine learning, characterized in that: The method comprises: Extract process parameter data, chemical composition data, and chemical composition change data from historical processing records to construct an initial data set; Analyzing the correlation between the process parameter data, the chemical composition data, and the chemical composition change data based on the initial data set to determine a key process condition combination, wherein the key process condition combination includes a key process parameter and a chemical composition change rate; constructing a feature data set based on the key process condition combination, dividing the feature data set into a training set and a test set, training a pharmacodynamic prediction model with the training set to obtain a first pharmacodynamic prediction model, and predicting the test set with the first pharmacodynamic prediction model to output preliminary pharmacodynamic change prediction data; performing error analysis on the preliminary drug efficacy change prediction data to determine a parameter adjustment direction of the first drug efficacy prediction model; A second drug efficacy prediction model is determined based on the parameter adjustment direction of the first drug efficacy prediction model, and the second drug efficacy prediction model is used to output optimized drug efficacy change prediction data.

2. The method according to claim 1, characterized in that Before analyzing the correlation between the process parameter data, the chemical composition data, and the chemical composition change data based on the initial data set to determine a set of key process parameters that are highly correlated with the chemical composition change data, the method further includes: The initial data set is cleaned and standardized to generate a standardized data set.

3. The method according to claim 2, characterized in that The step of performing data cleaning and standardization on the initial data set to generate a standardized data set includes: Filling missing values ​​in the initial data set by an interpolation method to obtain a first data set; performing a removal operation on abnormal data in the first data set to obtain a second data set; performing normalization processing on the second data set to obtain a third data set; The third data set is subjected to consistency check and adjustment to obtain a fourth data set.

4. The method according to claim 1, wherein The analyzing the correlation between the process parameter data, the chemical composition data, and the chemical composition change data based on the initial data set to determine a key process condition combination includes: Screen key parameters through Lasso regression; The quadratic regression model was constructed by response surface methodology to determine the extreme points; A combination of key process conditions is determined based on the key parameters and the extreme points through a genetic algorithm.

5. The method according to claim 1, wherein The performing error analysis on the preliminary drug efficacy change prediction data to determine the parameter adjustment direction of the first drug efficacy prediction model includes: Calculate the deviation of the preliminary pharmacodynamic change prediction data minus the actual pharmacodynamic change data, compare the deviation with a preset error threshold, and if the deviation exceeds the error threshold, mark the data point as "needs adjustment." If the deviation is positive, reduce the value of the key process parameter in the first pharmacodynamic prediction model; otherwise, increase the value of the key process parameter in the first pharmacodynamic prediction model.

6. The method according to claim 1, characterized in that After determining the second drug efficacy prediction model based on the parameter adjustment direction of the first drug efficacy prediction model, the method further includes: Compare the optimized efficacy change prediction data with the efficacy change data under actual process conditions to generate applicability analysis results.

7. The method according to claim 6, characterized in that After comparing the optimized drug efficacy change prediction data with the drug efficacy change data under actual process conditions to generate a suitability analysis result, the method further includes: Based on the applicability analysis results, a report on the correlation between process parameters and drug efficacy changes is generated and displayed.

8. The method according to claim 7, characterized in that The correlation report is presented through a three-dimensional scatter plot and a heat map.

9. The method according to claim 6, characterized in that The comparing the optimized efficacy change prediction data with the efficacy change data under actual process conditions to generate a suitability analysis result includes: Calculating the root mean square error between the optimized efficacy change prediction data and the efficacy change data under actual process conditions; If the root mean square error is higher than a preset threshold, a feature importance analysis is performed to quantify the weight of each process parameter; Divide each process parameter into intervals and calculate the root mean square error of each interval; The process parameter interval with the highest prediction accuracy is marked based on the root mean square error of each interval, and the applicability analysis result is determined based on the process parameter interval with the highest prediction accuracy.

10. A drug efficacy prediction model training system based on machine learning, characterized in that: The system comprises: A construction module is used to extract process parameter data, chemical composition data, and chemical composition change data from historical processing records to construct an initial data set; a first determination module, configured to analyze the correlation between the process parameter data, the chemical composition data, and the chemical composition change data based on the initial data set to determine a key process condition combination, wherein the key process condition combination includes a key process parameter and a chemical composition change rate; a training module, configured to construct a feature data set based on the key process condition combination, divide the feature data set into a training set and a test set, train a pharmacodynamic prediction model using the training set to obtain a first pharmacodynamic prediction model, and predict the test set using the first pharmacodynamic prediction model to output preliminary pharmacodynamic change prediction data; an adjustment module, configured to perform error analysis on the preliminary drug efficacy change prediction data and determine a parameter adjustment direction of the first drug efficacy prediction model; The second determination module is used to determine a second drug efficacy prediction model based on the parameter adjustment direction of the first drug efficacy prediction model, and the second drug efficacy prediction model is used to output optimized drug efficacy change prediction data.