Adult patient vancomycin clearance rate prediction scheme design combined with Stacking model
By combining the Stacking model with various machine learning and deep learning models, the instability and inaccuracy of vancomycin pharmacokinetic prediction in existing technologies have been resolved, achieving higher accuracy and stable clearance rate prediction, and supporting personalized treatment.
Patent Information
- Application Number
- CN202511082359.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-08-04
- Publication Date
- 2025-11-14
AI Technical Summary
Existing machine learning models for vancomycin pharmacokinetic prediction are unstable and inaccurate, lack medical mechanism explanation, and are difficult to meet the requirements of personalized treatment.
By combining the Stacking model with population pharmacokinetics (PPK), integrating various machine learning and deep learning models such as Random Forest, CNN, RNN, XGBoost, and LightGBM, and improving prediction accuracy through Bayesian optimization algorithms, a vancomycin clearance prediction scheme for adult patients was constructed.
It improves the accuracy and stability of vancomycin clearance prediction, provides a more reliable basis for medication, supports individualized treatment, reduces model error, and enhances the scientific interpretability of predictions.
Smart Images

Figure CN120954560A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of vancomycin pharmacokinetic technology, and in particular to the design of a vancomycin clearance prediction scheme for adult patients using a Stacking model. Background Technology
[0002] Vancomycin is a glycopeptide antibiotic discovered in the 1950s. It works by specifically binding to the D-alanyl-D-alanine (D-Ala-D-Ala) dipeptide at the terminal end of the bacterial cell wall peptidoglycan precursor, blocking the catalytic activity of transpeptidase and transglycosylase, thereby strongly inhibiting the synthesis of the cell wall of Gram-positive bacteria and ultimately leading to bacterial death. Due to its unique antibacterial mechanism and irreplaceable role in treating severe Gram-positive bacterial infections, vancomycin has long been considered one of the important "last-line" antibiotics in clinical practice. However, due to vancomycin's narrow therapeutic window and significant inter-individual pharmacokinetic variability, accurately predicting its drug clearance rate is crucial for ensuring therapeutic efficacy and avoiding toxic side effects.
[0003] In recent years, with the advancement of computer technology, research hotspots have emerged continuously in the field of pharmacokinetics. In the prediction of pharmacokinetic parameters, traditional population pharmacokinetic (PPK) models have achieved certain results. (Zhang Ting) [1] Studies by researchers have shown that machine learning models such as XGBoost significantly improve the accuracy of vancomycin AUC prediction, surpassing traditional pharmacokinetic models. Similarly, the application of machine learning in personalized medicine has gradually gained widespread attention, but methods for accurately predicting pharmacokinetic parameters still need further improvement to meet the requirements of individualized clinical treatment.
[0004] However, despite the promising applications of machine learning in the field of pharmacokinetics, there are still certain limitations to using machine learning models alone. Their pharmacokinetic predictions are unstable and inaccurate, and many traditional machine learning models lack scientific explanation of medical mechanisms, which hinders their widespread application.
[0005] Unlike single machine learning models, this invention adopts an ensemble learning strategy, constructing a stacking model based on multiple machine learning and deep learning models, and combining it with a PPK model to establish a more accurate prediction framework, thereby providing a more reliable scientific basis for drug use. Summary of the Invention
[0006] In view of the shortcomings of the prior art described above, the purpose of this invention is to provide a design for a vancomycin clearance prediction scheme for adult patients that incorporates a Stacking model, in order to solve the problems existing in the prior art.
[0007] To achieve the above and other related objectives, this invention provides a design for a vancomycin clearance prediction scheme for adult patients that incorporates a Stacking model. The scheme is characterized by combining the Stacking model with population pharmacokinetics (PPK) principles to improve the accuracy of vancomycin clearance prediction in adult patients. The design of this vancomycin clearance prediction scheme for adult patients that incorporates a Stacking model consists of four parts.
[0008] The first part involves using a screened simulated dataset containing data on the sex, age, weight, serum creatinine (Scr) levels, and clearance rates based on population pharmacokinetic (PPK) principles and Gibbs sampling from 1500 adult patients. Defective samples (such as incomplete data) were removed during the data screening process. Data from 500 patients was used to compare the performance of the PPK model with the model ultimately proposed in this invention, while the remaining 1000 patients' data were used to train the aforementioned machine learning model, the Stacking model.
[0009] The second part involves using machine learning and deep learning training, combined with patient information (i.e. predictive factors), to construct various predictive clearance rate models.
[0010] The third part is: based on five machine learning models and deep learning models, namely Random Forest, CNN, RNN, XGBoost and LightGBM, the Stacking model is integrated and the Bayesian optimization algorithm is used to improve the prediction accuracy of the model.
[0011] The fourth part involves selecting 500 patients from the simulated dataset as an external dataset, and comparing the final model, which combines the Stacking model and the PPK model proposed in this invention, with the PPK model.
[0012] As described above, the proposed scheme for predicting vancomycin clearance in adult patients, combining a Stacking model, demonstrates improved prediction accuracy and stability compared to single models or traditional methods, based on simulation data validation. In the future, with the continuous accumulation of data and algorithm optimization, this model is expected to be more widely applied in practice, further providing a reference for the development of precision medicine and personalized treatment. Simultaneously, this invention also provides a reference for pharmacokinetic prediction studies of other drugs. Attached Figure Description
[0013] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the accompanying drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are merely some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without any creative effort.
[0014] Figure 1 This is a block diagram illustrating the design of a vancomycin clearance prediction scheme for adult patients using a Stacking model, as an example of the present invention.
[0015] Figure 2 This is a block diagram illustrating the construction of a Stacking model for predicting vancomycin clearance in adult patients, as exemplified by an example of the present invention.
[0016] Figure 3 This is a comparison chart of the average residuals between the final model and the PPK model of a vancomycin clearance prediction scheme for adult patients that combines the Stacking model, as an example of the present invention. Detailed Implementation
[0017] The present invention will now be described in detail with reference to specific embodiments. These embodiments will help those skilled in the art to further understand the present invention, but do not limit the invention in any way. It should be noted that those skilled in the art can make various modifications and improvements without departing from the concept of the present invention. These all fall within the scope of protection of the present invention.
[0018] refer to Figure 1 This invention proposes a scheme for predicting vancomycin clearance in adult patients using a Stacking model, comprising the following steps:
[0019] Data preprocessing: This embodiment of the invention uses a screened simulated clinical dataset containing data on the gender, age, weight, and serum creatinine (Scr) levels of 1500 adult patients. A vancomycin clearance prediction model for adult patients is derived using PPK. [2] Using Gibbs sampling, a reference clearance rate was obtained for 1500 patients. The formula for calculating the clearance rate (CL) for adult patients is as follows:
[0020]
[0021] Where CLcr represents creatinine clearance rate and Age represents age. and This indicates variability among patients, which can be obtained through sampling. , The formula for calculating creatinine clearance (CLcr) is as follows.
[0022]
[0023]
[0024] Wherein, gender indicates sex; weight indicates body weight in kg; and Scr indicates serum creatinine in umol / L.
[0025] During the data screening process, defective samples (such as incomplete data) were removed. Data from 500 patients were used to compare the performance of the PPK model with the final model combining the Stacking model and the PPK model proposed in this patent. Data from the remaining 1,000 patients were used to train the Stacking model.
[0026] Creating clearance rate models: Machine learning and deep learning training are performed using patient information (i.e., predictors) to build various predictive clearance rate models.
[0027] Building a Stacking Model: Based on five machine learning and deep learning models—Random Forest, CNN, RNN, XGBoost, and LightGBM—the Stacking model is integrated and Bayesian parameter optimization is used to improve the model's predictive performance.
[0028] Model comparison: 500 patients were selected from the simulated dataset as an external dataset, and the final model combining the Stacking model and PPK model proposed in this invention was compared with the PPK model.
[0029] Stacking is an ensemble strategy that leverages predictions generated by multiple machine learning algorithms as new features, feeding them into a meta-learner for further learning. Through training the meta-learner, these initial predictions can be optimally combined to generate a more accurate set of final predictions. Throughout this process, Stacking fully utilizes the strengths of different learners and intelligently integrates these predictions via the meta-learner to improve the overall model performance.
[0030] refer to Figure 2 This paper demonstrates the architecture of the Stacking model designed in this invention. It selects Convolutional Neural Network (CNN), Lightweight Gradient Boosting (LightGBM), Recurrent Neural Network (RNN), Extreme Gradient Boosting (XGBoost), and Random Forest as the base models, and selects XGBoost as the meta-learner.
[0031] Choosing these five models as the base models ensures diversity in the base models to capture different data features. XGBoost has non-linear modeling capabilities and can handle missing data. LightGBM is a gradient boosting framework based on decision trees that captures non-linear interaction features between data. Random Forest reduces the risk of overfitting by integrating multiple decision trees and is suitable for handling noise in high-dimensional data. Although CNNs are generally used to process image or time series data, their convolutional layers can extract local spatial features, which is suitable for capturing local correlations between patient physiological characteristics. RNNs, through their recursive structure, can capture the potential dependencies between different physiological indicators in the data.
[0032] By integrating multiple model types, the Stacking model can comprehensively capture linear, nonlinear, and local features in the data, thereby significantly improving the robustness of predictions. Therefore, the Stacking model was chosen to be combined with the PPK model, and linear regression integrated these complementary predictions, resulting in a final model with high prediction accuracy.
[0033] This invention selects four statistical evaluation indicators as the evaluation indicators for the model. R² is a statistic of the goodness of fit of the regression model, and its value is between 0 and 1. The closer R² is to 1, the better the model fits the data. The formula for calculating R² is:
[0034]
[0035] Where SSE represents the residual sum of squares and SST represents the total sum of squares.
[0036] Mean squared error (MSE) is a statistical measure of the difference between actual and predicted values; it is the average of the squares of the differences between each predicted and actual value. Root mean squared error (RMSE) is the square root of MSE, restoring the units of the error to the units of the original data. Mean absolute error (MAE) is similar to MSE, but uses the absolute value of the error instead of its square, making it sensitive to outliers. The smaller the MSE, RMSE, and MAE, the better the model's predictive performance. The formulas for calculating MSE, RMSE, and MAE are:
[0037]
[0038]
[0039]
[0040] Where n is the number of patients, This represents the simulated reference clearance rate for the i-th patient. This represents the predicted clearance rate for the i-th patient.
[0041] To further illustrate that the Stacking model performs best across key evaluation metrics, the table below compares performance indicators: the Stacking model has the lowest mean squared error (MSE=0.5588), root mean square error (RMSE=0.7475), and mean absolute error (MAE), while its coefficient of determination (R²=0.9036) is significantly better than other models (such as Random Forest, XGBoost, LightGBM, RNN, and CNN). This indicates that the Stacking model has the smallest deviation between predicted and actual values and the strongest explanatory power for data variability.
[0042] Model Mean Squared Error (MSE) Root Mean Square Error (RMSE) Coefficient of determination (R²) Mean Absolute Error (MAE) Random Forest 1.2257 1.1071 0.7924 0.9004 XGBoost 1.1068 1.0520 0.8066 0.7891 LightGBM 2.4031 1.5502 0.5930 1.2731 RNN 2.7160 1.6480 0.5400 1.3072 CNN 3.7313 1.9317 0.3572 1.6074 Stacking 0.5588 0.7475 0.9036 0.5932
[0043] refer to Figure 3 Based on simulation data, the average residual distribution (horizontal axis: sample size; vertical axis: average residual value) of the vancomycin clearance rate predicted by the population pharmacokinetic (PPK) model and the final model proposed in this invention was compared. The residuals of the PPK model (light gray line) fluctuated significantly within the range of [-1.5, 1.5], especially showing obvious systematic bias in the sample size range of 225 to 300. In contrast, the residual fluctuation range of the hybrid model of this invention (dark gray line) was significantly reduced to [-1, 1]. The analysis shows that the relative error of the final model proposed in this patent is reduced by 56% compared with the PPK model, and the width of its average residual distribution range is also correspondingly reduced.
[0044] In conclusion, based on the validation results using simulation data, the model proposed in this invention exhibits relatively stable performance in predicting vancomycin clearance rates in adult patients. Compared to single models or traditional methods, the model demonstrates improved prediction accuracy and stability. This achievement provides clinical pharmacists with medication guidance, thereby enhancing the safety and effectiveness of patient treatment. Furthermore, the methodological framework established by this model also provides a reference framework for research on the prediction of pharmacokinetic parameters of other drugs.
[0045] The above embodiments should be understood as illustrative only and not as limiting the scope of protection of the present invention. After reading the description of the present invention, those skilled in the art can make various alterations or modifications to the present invention, and these equivalent changes and modifications also fall within the scope defined by the claims of the present invention. References
[0046] Zhang T, Chen Yuheng, Ding Lanping, et al. Prediction of vancomycin AUC24 h using population pharmacokinetic model and XGBoost algorithm [J]. Chinese Journal of Hospital Pharmacy, 2025, 45(08):859-865.
[0047] Gao Yucheng, Jiao Zheng, Huang Hong, et al. Development of a vancomycin personalized dosing decision support system [J]. Acta Pharmaceutica Sinica, 2018, 53(01):104-110.
Claims
1. A scheme for predicting vancomycin clearance in adult patients using a stacking model is designed, combining the application effects of a population pharmacokinetic (PPK) model and a machine learning model in predicting drug pharmacokinetic parameters. The PPK model, based on classical pharmacokinetic theory, can describe the behavior of vancomycin in vivo. The machine learning model, through learning from data, can automatically capture nonlinear relationships and complex patterns in the data, exhibiting strong adaptability and predictive capabilities.
2. The design of a vancomycin clearance prediction scheme for adult patients combined with a Stacking model as described in claim 1, characterized in that: The simulated data included gender, age, weight, serum creatinine (Scr) of 1500 adult patients, as well as clearance rate data based on population pharmacokinetics (PPK) principles and Gibbs sampling. Data from 1000 randomly selected patients were used to train the Stacking model, while the remaining 500 patients served as an external dataset for comparing the performance of the final model combining the Stacking and PPK models proposed in this invention with the PPK model.
3. The design of a vancomycin clearance prediction scheme for adult patients combined with a Stacking model as described in claim 2, characterized in that: Patient information (i.e. predictors) was used for machine learning and deep learning training to build various models predicting vancomycin clearance rates.
4. The design of a vancomycin clearance prediction scheme for adult patients combined with a Stacking model as described in claim 3, characterized in that: Based on five machine learning models—Random Forest, CNN, RNN, XGBoost, and LightGBM—and a deep learning model, the Stacking model is integrated and Bayesian parameter optimization is used to further improve the model's predictive performance.
5. The design of a vancomycin clearance prediction scheme for adult patients combined with a Stacking model as described in claim 4, characterized in that: 500 patients were selected from the simulated dataset as samples for comparing the performance of the final model combining the Stacking model and the PPK model proposed in this invention with that of the PPK model.