A prediction method for drug solubilization performance based on excipient classification

A drug solubility performance prediction model was established through machine learning methods, which solved the problem of large differences in drug solubility under excipient conditions, achieved accurate prediction of drug solubility and reduced costs.

CN116631533BActive Publication Date: 2025-09-26CHINA PHARM UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202310624617.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-05-30
Publication Date
2025-09-26
Estimated Expiration
2043-05-30

AI Technical Summary

Technical Problem

Existing technologies lack convenient and accurate methods to predict the solubilization performance of drugs under different excipient conditions, resulting in large differences in drug solubility enhancement effects, time-consuming experiments and high costs.

Method used

A machine learning method was used to collect and process the drug dissolution percentage and dissolution efficiency data under excipient conditions, and a quantitative prediction model was established. The random forest algorithm was used to divide the data set and calculate the correlation and feature importance to predict the solubility performance of the drug under different excipient conditions.

Benefits of technology

It improves the prediction accuracy of drug solubilization performance, reduces the cost and risk of new drug design and development, and provides guidance for improving drug solubility under different excipient conditions.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116631533B_ABST
    Figure CN116631533B_ABST
Patent Text Reader

Abstract

The present invention discloses a method for predicting drug solubilization performance based on excipient classification, comprising the following steps: data collection and processing, component division, correlation and feature importance calculation, model establishment and prediction. The present invention predicts the drug solubility percentage and dissolution efficiency respectively under the condition that different excipients are used as main solubilizers for the drug, and analyzes the specific influence of different feature variables on the two target variables, thereby improving the accuracy of drug solubilization performance prediction, reducing the cost and risks faced in the design and development of new drugs, and guiding the research and invention of preparing inclusion compounds of existing drugs through complexation technology to improve drug solubility, thereby improving the accuracy of drug solubility percentage and dissolution efficiency prediction, and reducing the cost and risks faced in the design and development of new drugs.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to a method for predicting drug solubilization performance, in particular to a method for predicting drug solubilization performance based on excipient classification. Background Art

[0002] Most drugs have poor water solubility and low solubility, which leads to low bioavailability, weak physiological compatibility, and poor patient compliance, which seriously hinders and limits the clinical treatment and application of most drugs. Researchers have conducted a lot of relevant research to overcome the series of problems caused by low drug solubility and have made breakthrough progress. At present, various technologies for improving the solubility of low-soluble drugs include physical modification of drugs, chemical modification, such as particle size reduction, crystal engineering, salt formation, solid dispersion technology, surfactants and complexation technology (International Scholarly Research Notices, 2012.195727, (10)). Complexation technology is a good strategy and method to improve the solubility and bioavailability of drugs with poor water solubility or insoluble in water. Compared with the raw material drug, the drug solubility of the inclusion complex is greatly improved.

[0003] However, due to the wide variety of drugs, the complex properties of different excipients, and the diverse and non-fixed operating conditions. Even for the same inclusion complex, if the operating conditions are different, the excipients in the inclusion complex are flexible and diverse and have different properties, which will still lead to huge differences in the solubilization effect, and the experiment is time-consuming and costly. A large number of relevant literatures have experimentally determined the percentage of drug dissolution and dissolution efficiency at the same time. Among them, the percentage of drug dissolution (DP) is defined as the amount of drug dissolved into the total amount of inclusion complex within a given time. Dissolution efficiency (DE) is defined as the area under the dissolution curve to a certain time (t), expressed as the percentage of the rectangular area of ​​the drug dissolved at a specific time (t) to the rectangular area of ​​100% dissolution of the drug.

[0004]

[0005] Q is the dissolution percentage and t is the corresponding dissolution time.

[0006] At present, relevant literature reports have been published on the prediction of the in vitro and in vivo solubility of specific drugs under different environments by combining different machine learning methods (decision trees, random forests and artificial neural networks, etc.). For example, berberine hydrochloride (BBR) was used as a model drug for lipid drug delivery system (LBDDS) to predict the inhibition efficiency of drug-phospholipid complex delivery in vivo and in vitro by machine learning methods (International Scholarly Research Notices, 2012 (10), 195727). At present, multiple machine learning methods have been successfully used to predict the solubility of 12 anticancer drugs in supercritical carbon dioxide (Pharmaceutics, 2022, 14 (8): 1632), verifying that the use of machine learning methods to predict drug dissolution is feasible and effective. The percentage of drug solubility and dissolution efficiency are equally important in reflecting the drug solubilization effect, but traditional research mainly focuses on the solubility of drugs in different solvents (including but not limited to water or organic solvents). There is no convenient and accurate research method for the improvement of drug solubility after solubilization with different excipients. Summary of the Invention

[0007] Purpose of the invention: The purpose of the present invention is to provide a convenient and accurate prediction method for drug solubilization performance based on excipient classification.

[0008] Technical solution: A method for predicting drug solubilization performance based on excipient classification of the present invention comprises the following steps:

[0009] (1) Data collection and processing: Data values ​​of drug dissolution percentage and dissolution efficiency under conditions where different excipients were used as the main solubilizer were collected and processed, and data values ​​where multiple excipients jointly improved drug solubility were removed;

[0010] (2) Classification: The drug dissolution percentage and dissolution efficiency data values ​​obtained in step (1) are divided into a characteristic variable data set of drug composition or complex operating conditions, and the data set is divided into a training set and a test set according to a ratio of 6 to 8:2 to 4;

[0011] (3) Correlation and feature importance calculation: For the characteristic variables under the conditions where different excipients are used as the main solubilizers of the drug, the correlation with the drug dissolution percentage and dissolution efficiency and the feature importance ranking are performed respectively;

[0012] (4) Model establishment: Using machine learning algorithms, quantitative prediction models for drug dissolution percentage and dissolution efficiency were established when excipients were used as the main solubilizers of the drug.

[0013] (5) Prediction: It is used to substitute the characteristic variables of various drug compositions or complex operating conditions under different excipients as the main solubilizers of the drug into the quantitative prediction model of drug dissolution percentage and dissolution efficiency for prediction.

[0014] Preferably, in step (1), the excipients include polyethylene glycol PEG, polyvinyl pyrrolidone PVP, hydroxypropyl methylcellulose HPMC or cyclodextrin as the main solubilizer of the drug to predict the drug dissolution percentage and dissolution efficiency.

[0015] Preferably, step (1) specifically includes the following steps:

[0016] (11) The data values ​​of drug composition and complex operating conditions collected under different excipients as the main solubilizer of the drug were converted to the same units;

[0017] (12) Verify and validate the unified data to ensure that the data is valid and authentic;

[0018] (13) All the above data values ​​are checked one by one under the conditions where different excipients are used as the main solubilizers of the drug, and the data values ​​in which multiple excipients jointly improve the solubility of the drug are removed; if the characteristic variable value of a specific drug in the collected data values ​​exceeds the range of the vast majority of data values ​​and is too high or too low, the abnormal data value is selected for removal.

[0019] Preferably, in step (2), the data set is divided into a training set and a test set in a ratio of 7:3.

[0020] Preferably, in step (2), the drug composition includes drug content DC, drug molecular weight M or molar ratio of excipient to drug Molar Ratio.

[0021] Preferably, in step (2), the complex operating conditions include drug concentration C, pressure P, pH value, temperature T or dissolution time t.

[0022] Preferably, the step (3) is specifically to use Python software to calculate the correlation coefficient r and p hypothesis test value between the eight characteristic variables and the DP and DE data sets under the conditions of different excipients as the main solubilizers of the drug, verify the correlation, and calculate their characteristic importance.

[0023]

[0024] in, and represents the average of variables x and y, and r is the correlation coefficient between the two input variables. When its value is between [-1, 1], it is relatively easy to determine the linear correlation between the data. Furthermore, when the correlation coefficient is between [0, 1], the two variables are positively correlated, and when it is between [-1, 0], the correlation is negative. This means that the linear relationship between any two different variables increases as their absolute values ​​increase.

[0025]

[0026] Where N is the sample size; p is the p-value for the hypothesis test between any two variables, obtained using a double-truncated distribution with N-2 degrees of freedom. The p-value for the hypothesis test only reflects whether there is a statistically significant difference between the two variables. Generally, a p-value < 0.05 indicates a statistically significant difference, a p-value < 0.01 indicates a statistically significant difference, and a p-value < 0.001 indicates an extremely significant difference.

[0027] Preferably, in step (4), the machine learning method includes random forest, decision tree or K-fold cross validation.

[0028] Preferably, in step (4), the parameters used to evaluate the predictive ability of the quantitative prediction model are: decision coefficient R 2 And root mean square error RMSE; Among them, RMSE is an error index, the smaller the value, the smaller the prediction error and the better the model; R 2 It is a correlation index. The closer it is to 1, the better the model fit is.

[0029] R 2 : Provides information about the goodness of fit of the model. In regression, it is a statistical measure of the degree of approximation between the regression prediction and the actual data points. It describes the correlation trend between the actual value and the predicted value, not a direct description of the prediction error. When the data deviates greatly from the distribution, only R is analyzed. 2 This may lead to model evaluation errors. RMSE, also known as the standard deviation of the prediction error, can be used to quantify model quality. The prediction error is directly described, focusing more on unfavorable predictions and having high sensitivity:

[0030]

[0031] RMSE: Error indicator used to quantify model quality,

[0032] Y i exp is the actual value, Y i pred is the predicted value, is the average of the actual values, and N represents the number of compounds.

[0033] Preferably, step (4) is specifically based on the method for predicting the percentage of drug solubility and dissolution efficiency under the conditions of different excipients as the main solubilizers of the drug provided by the present invention, and at the same time, the characteristic variables: drug content, drug molecular weight, molar ratio of excipient to drug, drug concentration, pressure, pH value, dissolution time or temperature and the target variable drug solubility percentage and dissolution efficiency data set are calculated with the mean, median and variance, and the characteristics affecting the drug solubility percentage and dissolution efficiency are ranked and analyzed in importance according to the prediction results.

[0034] Principle of the invention: The present invention simultaneously conducts modeling and prediction for drug solubility percentage and dissolution efficiency under the conditions where different excipients are used as the main solubilizers of the drug, thereby improving the accuracy of drug solubilization performance prediction, reducing the cost and risks of new drug design and development, guiding the research and invention of preparing inclusion compounds of existing drugs through complexation technology to improve drug solubility, and increasing the applicability of machine learning prediction methods.

[0035] Unlike traditional studies that focus on the different solubilities of a drug in different solvents, the present invention focuses on the problem of accurately predicting the performance of excipient-solubilized drugs, particularly a method and system for predicting the drug solubility percentage and dissolution efficiency based on drug composition (drug content, drug molecular weight, and molar ratio of excipient to drug) and complex operating conditions (drug concentration, pressure, pH value, temperature, and dissolution time) under the condition that different excipients serve as the main solubilizers of the drug.

[0036] Beneficial effects: Compared with the existing technology, the present invention has the following significant advantages: using machine learning methods to predict the drug dissolution percentage and dissolution efficiency of different excipients under complex solubilizer conditions based on the different compositions of multiple drugs and complex operating conditions, and analyzing the specific effects of different characteristic variables on the two target variables, so as to improve the accuracy of drug dissolution percentage and dissolution efficiency prediction and reduce the cost and risks of new drug design and development. BRIEF DESCRIPTION OF THE DRAWINGS

[0037] Figure 1 Flowchart of the method for predicting drug solubilization performance based on excipient classification of the present invention;

[0038] Figure 2 Statistical analysis of the data values ​​of eight characteristic variables and two target variables corresponding to various drug compositions and complex operating conditions with PEG as the main drug solubilizer;

[0039] Figure 3 The results of ranking the importance of the characteristics of drug dissolution percentage and dissolution efficiency when PEG is used as the main solubilizer for the drug are shown;

[0040] Figure 4The model prediction analysis results of drug dissolution percentage and dissolution efficiency using PEG as the main solubilizer were obtained using Python software;

[0041] Figure 5 The corresponding distribution states and relative error calculation results of the actual and predicted values ​​of drug dissolution percentage and dissolution efficiency obtained when PEG was used as the main solubilizer for the drug;

[0042] Figure 6 The important feature analysis results of the drug dissolution percentage using PEG as the main solubilizer obtained by Python software;

[0043] Figure 7 These are the important characteristic analysis results of drug dissolution efficiency when PEG is used as the main solubilizer, obtained using Python software. DETAILED DESCRIPTION

[0044] The technical solution of the present invention will be further described below with reference to the accompanying drawings and embodiments.

[0045] Example 1

[0046] This example provides a method for predicting drug dissolution percentage and dissolution efficiency under different drug compositions and complex operating conditions based on PEG as the main solubilizer for the drug. Figure 1 As shown, including:

[0047] Step 1: Data collection and processing;

[0048] (1) By consulting literature and databases, we collected data on different drug compositions: drug content (DC), drug molecular weight (M) and molar ratio (PEG / Drug), and complex operating conditions: drug concentration (C), pressure (P), pH value, temperature (T) and dissolution time (t), drug dissolution percentage (DP) and dissolution efficiency (DE), as well as the corresponding drug names.

[0049] (2) Preprocess the collected data and convert different characteristic variables in the literature and database into the same unit.

[0050] (3) Data values ​​containing excipients other than PEG as the main solubilizer for drugs were removed. Finally, according to statistics, a total of 1489 drug dissolution percentage and dissolution efficiency data values ​​containing 33 drugs (Table 1) were collected when PEG was used as the main solubilizer for drugs.

[0051] Table 1. Drugs included when PEG was used as the primary solubilizer

[0052]

[0053] (4) The mean, median, upper quartile and lower quartile of the eight characteristic variables and two target variables (DP and DE) collected under the condition of PEG as the main solubilizer for different drug compositions and complex operating conditions were calculated, as shown in Figure 4. Figure 2 As shown, the drug content DC is in the range of 2.75-87%, with a relative concentration of 25%; in addition, the average values ​​of the target variables DP and DE are 56.91% and 47.69%, respectively.

[0054] Step 2, level division;

[0055] According to the drug dissolution percentage and dissolution efficiency data values ​​obtained in step (1), the data are divided into characteristic variable data sets of drug composition or complex operating conditions. The random forest (RF) method in the machine learning algorithm is used to divide the data set obtained under the condition that PEG is the main solubilizer of the drug into a training set and a test set with a ratio of 7:3.

[0056] Step 3, calculate the correlation and feature importance;

[0057] (1) Calculate correlation;

[0058] Python software was used to calculate the p hypothesis test values ​​between drug composition (DC, M and Molarratio), operating conditions (C, P, pH, T and t) and target variables (DP and DE) under the condition that PEG was the main solubilizer, and Pearson plots were drawn to present the correlation between them.

[0059] (2) Calculate feature importance;

[0060] Python software was used to calculate the importance of eight characteristic variables of drug composition (DC, M and Molarratio) and complex operating conditions (C, P, pH, T and t) to DP and DE respectively under the condition of PEG as the main solubilizer of the drug and sort them in descending order.

[0061] Origin software was used to draw the order of the importance of different feature variables on DP and DE based on the above calculation results.

[0062] like Figure 3 As shown in the figure, when PEG is used as the main solubilizer, the molar ratio (Molar Ratio) and drug content (DC) have a greater impact on the percentage of drug dissolved, among which the drug content has a more significant impact on the dissolution efficiency than the molar ratio. In addition, it can be seen that the drug composition has a greater impact on the dissolution percentage than the operating conditions, which is contrary to the results of the dissolution efficiency.

[0063] Step 4, build the model;

[0064] A quantitative prediction model for drug dissolution percentage and dissolution efficiency was constructed. When predicting the drug dissolution percentage and dissolution efficiency of different drugs, the specific excipient to be used as the main solubilizer for the drug was first determined based on the drug properties (in this example, PEG was determined as the main solubilizer for the drug). Then, eight data sets with different compositions and complex operating conditions for multiple drugs were substituted into the established quantitative prediction model for drug dissolution percentage and dissolution efficiency for fitting and prediction.

[0065] like Figure 4 As shown, when PEG is the main solubilizer for drugs, the R 2 They are 0.86 and 0.95 respectively; RMSE are 0.13 and 0.08 respectively; which proves that the established quantitative prediction model is acceptable.

[0066] Step 5, prediction;

[0067] (1) Analysis of actual and predicted values;

[0068] like Figure 5 Figures 5a and 5b illustrate the accuracy of predicting drug dissolution percentage and dissolution efficiency using the random forest model, based on visualization of the actual and predicted values ​​randomly sampled in step 3. The prediction results show that, with PEG as the primary solubilizer, the predicted values ​​are largely consistent with the actual values, demonstrating the applicability of the random forest algorithm for predicting drug dissolution percentage and dissolution efficiency, with excellent performance.

[0069] (2) Relative error analysis;

[0070] like Figure 5 5c and 5d are based on Figure 5 The calculation results of the relative errors of drug dissolution percentage and dissolution efficiency were obtained by comparing the actual values ​​and predicted values ​​in step 5a and 5b. The results show that under the condition of PEG as the main solubilizer, the relative errors of most actual and predicted values ​​of drug dissolution percentage are within 10%, and only a few relative errors exceed 10%. The relative errors of most dissolution efficiency values ​​are within 5%, which is consistent with the R 2 This is consistent with the RMSE calculation results, further demonstrating that the prediction results of drug dissolution percentage and dissolution efficiency have certain universality and stability.

[0071] (3) Analysis of important feature results;

[0072] like Figure 6 6a-6d and Figure 7 7a-7d in the figure, based on Figure 3 The first four important features for drug dissolution percentage and dissolution efficiency were analyzed in detail one by one. When analyzing the impact of one feature variable on drug dissolution percentage and dissolution efficiency, the other feature variables were average values. Among them, E(DP) and E(DE) are the average values ​​of drug dissolution percentage and dissolution efficiency, respectively. Figure 6 and 7 As shown in the vertical axis, when the PDP value (thick line) of the dissolution percentage and dissolution efficiency is higher than E(DP) and E(DE), the corresponding characteristic variables are favorable conditions, otherwise they are unfavorable conditions.

[0073] The results showed that when PEG was used as the main solubilizer for drugs, the molar ratio, drug content, drug concentration and dissolution time were more important for the percentage of drug dissolution than other characteristic variables; drug content, dissolution time, molar ratio and drug molecular weight were more important for dissolution efficiency. Taking dissolution time as an example: when the dissolution time was between 0 and 100 minutes, the percentage of drug dissolution continued to increase ( Figure 6 6d in the figure); the dissolution efficiency increased sharply at 55 minutes and remained unchanged for 120 minutes, and increased sharply again at 175 minutes and did not change any more ( Figure 7 7d).

Claims

1. A method for predicting drug solubilization performance based on excipient classification, characterized in that: The prediction method comprises the following steps: (1) Data collection and processing: Data values ​​of drug dissolution percentage and dissolution efficiency under conditions where different excipients were used as the main solubilizer were collected and processed, and data values ​​where multiple excipients jointly improved drug solubility were removed; (2) Classification: Based on the drug dissolution percentage and dissolution efficiency data values ​​obtained in step (1), the data sets are divided into characteristic variable data sets of drug composition or complex operating conditions, and the data sets are divided into training sets and test sets according to the ratio of 6 to 8:2 to 4; (3) Correlation and feature importance calculation: Correlation calculation and feature importance ranking were performed on the feature variables under the conditions where different excipients were used as the main solubilizers for the drug, respectively, with the drug dissolution percentage and dissolution efficiency. The feature variables were drug content DC, drug molecular weight M, excipient to drug molar ratio Molar Ratio, drug concentration C, pressure P, pH value, temperature T, or dissolution time t; (4) Model establishment: Using machine learning algorithms, quantitative prediction models for drug dissolution percentage and dissolution efficiency were established when excipients were used as the main solubilizers. (5) Prediction: It is used to substitute the characteristic variables of multiple drug compositions or complex operating conditions under different excipients as the main solubilizers for the quantitative prediction model of drug dissolution percentage and dissolution efficiency for prediction. The characteristic variables of the complex operating conditions refer to drug concentration C, pressure P, pH value, temperature T or dissolution time t.

2. The prediction method according to claim 1, characterized in that In step (1), the drug dissolution percentage and dissolution efficiency are predicted under the condition that the excipient includes polyethylene glycol, polyvinyl pyrrolidone, hydroxypropyl methylcellulose or cyclodextrin as a drug solubilizer.

3. The prediction method according to claim 1, wherein: Step (1) specifically includes the following steps: (11) The data values ​​of drug composition and complex operating conditions collected under different excipients as the main solubilizer of the drug were converted to the same units; (12) Verify and validate the unified data to ensure that the data is valid and authentic; (13) All the above data values ​​are checked one by one under the conditions where different excipients are used as the main solubilizers of the drug, and the data values ​​in which multiple excipients jointly improve the solubility of the drug are removed; if the characteristic variable value of a specific drug in the collected data values ​​exceeds the range of the vast majority of data values ​​and is too high or too low, the abnormal data value is selected for removal.

4. The prediction method according to claim 1, wherein: In step (2), the data set is divided into a training set and a test set in a ratio of 7:

3.

5. The prediction method according to claim 1, wherein: In step (2), the drug composition includes drug content DC, drug molecular weight M or molar ratio of excipient to drug Molar Ratio.

6. The prediction method according to claim 1, characterized in that In step (2), the complex operating conditions include drug concentration C, pressure P, pH value, temperature T or dissolution time t.

7. The prediction method according to claim 1, wherein: Specifically, step (3) is to use Python software to calculate the correlation coefficients r and p hypothesis test values ​​between the eight characteristic variables and the DP and DE data sets under the conditions where different excipients are used as the main solubilizers of the drug, verify the correlation, and calculate their characteristic importance; in, and represents the average value of variables x and y, r is the correlation coefficient between the two input variables, and its value is between [-1, 1]. When the correlation coefficient value is between [0, 1], the two variables are positively correlated, and between [-1, 0], they are negatively correlated; Where N is the sample size; p is the p hypothesis test value between any two variables, which is obtained through a double-truncated distribution with N-2 degrees of freedom. The p hypothesis test value reflects whether there is a significant statistical difference between the two. p<0.05 indicates a statistical difference, p<0.01 indicates a significant statistical difference, and p<0.001 indicates an extremely significant statistical difference.

8. The prediction method according to claim 1, wherein: In step (4), the machine learning method includes random forest, decision tree or K-fold cross validation.

9. The prediction method according to claim 1, characterized in that In step (5), the parameters used to evaluate the predictive ability of the quantitative prediction model are: decision coefficient R 2 and root mean square error RMSE; where R 2 It is a correlation index that provides information about the goodness of fit of the model; RMSE: Error indicator used to quantify model quality, is the actual value, Y i pred is the predicted value, is the average of the actual values, and N represents the number of compounds.

10. The prediction method according to claim 1, wherein: Step (5) is specifically, based on the prediction method of drug solubility percentage and dissolution efficiency under the conditions of different excipients as the main solubilizers of the drug, the characteristic variables: drug content, drug molecular weight, molar ratio of excipient to drug, drug concentration, pressure, pH value, dissolution time or temperature and the target variable drug solubility percentage and dissolution efficiency data set are calculated, and the characteristics affecting the drug solubility percentage and dissolution efficiency are ranked and analyzed in terms of importance according to the prediction results.

Citation Information

Patent Citations

  • Solubility prediction model of compound molecules and application

    CN114334022A

  • KR20220065378A