Method for predicting and optimizing treatment effect of bacterial infection diseases

By using SHAP algorithm and Boruta algorithm to select features and build an integrated machine learning model, the problems of inaccurate prediction of treatment effects of bacterial infection diseases in the prior art are solved, and higher prediction accuracy and better treatment plan optimization effects are achieved.

CN120089273APending Publication Date: 2025-06-03SHANGHAI PUDONG HOSPITAL
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510200904.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-24
Publication Date
2025-06-03

AI Technical Summary

Technical Problem

The prior art cannot predict the effects of multiple therapeutic effects of bacterial infection diseases at the same time, especially the effects of single-use and combined drugs, and there are problems such as insufficient input characteristics, unused time series information, insufficient model interpretation, lack of standardized evaluation systems, data quality and scale restrictions, and lack of practical applications.

Method used

The core feature set is calculated by using SHAP algorithm and Boruta algorithm, and the optimal feature combination is selected from it, different optimal basic models are constructed, and integrated machine learning models are then built using the integrated algorithm to integrate the patient's multi-dimensional information and time series data to improve the accuracy and interpretability of predictions.

Benefits of technology

It improves the prediction accuracy of the treatment effect of bacterial infection diseases, provides scientific basis for optimization of anti-infection treatment, reduces antibacterial drug exposure and medical costs, reduces adverse reactions in drug treatment, and improves the generalization ability and explanatory nature of the model.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120089273A_ABST
    Figure CN120089273A_ABST
Patent Text Reader

Abstract

The invention provides a method for predicting and optimizing the treatment effect of bacterial infection diseases. The method comprises the following steps: 1) collecting clinical information of a patient to obtain original data; 2) preprocessing the original data to obtain an original data set; 3) dividing the original data set to obtain a training set and a test set; step 4) based on the training set, using an SHAP algorithm and a Boruta algorithm to calculate a core feature set; 5) selecting an optimal feature combination from the core feature set to construct different optimal basic models, and 6) using an integration algorithm to construct an integrated machine learning model based on the basic models, and the method can solve the problem that multiple treatment effects of bacterial infection diseases cannot be predicted at the same time in the prior art.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to a method for predicting and optimizing the treatment effect of bacterial infectious diseases, belonging to the technical field of machine learning model prediction. Background Art

[0002] Bacterial infectious diseases are diseases that can be prevented and treated. However, currently, bacterial infectious diseases have become one of the main causes of death and disability globally. With the development of medical technology, the prediction and optimization of the treatment effect of bacterial infectious diseases have mainly gone through the following stages: 1) The first stage is qualitative prediction based on clinical experience. Doctors mainly rely on personal experience and case accumulation, combined with patients' symptoms, signs, and basic examination results, to make empirical judgments on the treatment effect. This method is too dependent on doctors' personal experience, and the prediction results are highly subjective. 2) The second stage is the introduction of statistical methods. By collecting a large amount of case data, establishing a statistical model, and analyzing the key factors affecting the treatment effect, a preliminary quantitative prediction of the treatment effect is made. Although this method has a certain degree of scientificity, it fails to fully consider individual differences. 3) The third stage is the adoption of machine learning technology. Using artificial intelligence algorithms to analyze multi-dimensional information such as patients' clinical data, laboratory test results, imaging data, and even bacterial DNA, a prediction model is established. This method can better discover the complex relationships between data and improve the prediction accuracy.

[0003] However, there are also some problems in the existing technologies: 1) The input features of the prediction model are not comprehensive enough. Currently, the prediction models often only consider some clinical indicators and fail to fully integrate important biological parameters such as patients' gene information, immune status, and microbiome characteristics, which affects the prediction accuracy. 2) It fails to effectively combine time series information. Most prediction models only analyze static data at a certain time point, ignoring the dynamic change process of disease development and making it difficult to accurately grasp the evolution trend of the treatment effect. 3) The interpretability of the model is insufficient. Many machine learning models have the "black box" problem and it is difficult to explain the specific prediction basis, which affects the trust of doctors and patients in the prediction results. 4) There is a lack of a standardized evaluation system. The evaluation indicators adopted by different studies are not unified, making it difficult to objectively compare and evaluate different prediction methods. 5) Limited by data quality and scale. It is difficult to obtain high-quality clinical data, and the data volume is often insufficient to support the training of complex models, which affects the generalization ability of the prediction model. 6) Lack of practical applications in specific scenarios. The construction of existing related models is often limited to paper discussions and lacks actual implementation and deployment.

[0004] For example, acute exacerbation of chronic obstructive pulmonary disease (AECOPD) is a preventable and treatable disease, but it has currently become one of the major causes of death and disability globally; with the aging of the population and the continuous existence of risk factor exposures, the disease burden of AECOPD is expected to further increase; the mortality related to AECOPD is also increasing significantly year by year; studies have shown that bacterial infection is an important cause of AECOPD, accounting for about 50% of cases, especially more common in patients with moderate to severe AECOPD; therefore, timely implementation of effective and appropriate antibacterial treatment is the key to the clinical management strategy for AECOPD complicated with bacterial lower respiratory tract infection.

[0005] Currently, major clinical practice guidelines unanimously recommend amoxicillin-clavulanic acid, cephalosporins, respiratory quinolones, and piperacillin-tazobactam as the first-line treatment regimens for AECOPD complicated with lower respiratory tract infection; however, due to the lack of robust clinical trial data and the limitations of traditional pathogen detection methods, the selection of antibiotics still largely relies on experience; but this empirical treatment model has the following deficiencies: 1) lack of personalized treatment guidance for specific patients, leading to an increased risk of treatment failure; 2) unnecessary exposure to antibacterial drugs may result in the emergence of drug-resistant strains; 3) it is difficult for traditional statistical methods to comprehensively consider multiple clinical variables for predicting treatment effects.

[0006] In recent years, the application of machine learning technology in the medical field has gradually emerged, and it has unique advantages in dealing with complex data relationships and predicting treatment effects; however, there are few existing studies on multi-task learning models for the treatment of bacterial infection diseases, and there is a lack of clinical decision support tools combined with network applications; therefore, there is an urgent need to develop a technical method that integrates data integration, treatment effect prediction, and optimization plan recommendation to provide support for clinical precision medicine. Summary of the Invention

[0007] The present invention proposes a method for predicting and optimizing the treatment effect of bacterial infection diseases, aiming to solve the problem that the existing technology cannot simultaneously predict multiple treatment effects (especially single and combined drug use) of bacterial infection diseases.

[0008] The technical solution of the present invention: A method for predicting and optimizing the treatment effect of bacterial infection diseases, the method includes: Step 1) Collect patient clinical information to obtain original data; Step 2) Preprocess the original data to obtain an original data set; Step 3) Divide the original data set to obtain a training set and a test set; Step 4) Based on the training set, use the SHAP algorithm and the Boruta algorithm to calculate the core feature set; Step 5) Select the optimal feature combination from the core feature set and construct different optimal basic models; Step 6) Based on the basic models, use an ensemble algorithm to construct an ensemble machine learning model.

[0009] Furthermore, the original data obtained by collecting the clinical information of patients in Step 1) includes original feature variables and target variables; the target variable is the clinical outcome, and the target variable is divided into two situations: treatment success and treatment failure; the original feature variables include the general conditions of patients, co-existing disease conditions, co-administered medication conditions, and examination and test indicators; the general conditions of patients include patient gender, patient age, and patient weight; the co-existing disease conditions include co-existing diabetes, co-existing tumors, and co-existing cardiovascular diseases; the co-administered medication conditions include whether other anti-infective drugs are combined; the examination and test indicators include white blood cell count, C-reactive protein, and procalcitonin; the data types of the original feature variables include continuous variables and binary variables.

[0010] Furthermore, the original data set obtained by preprocessing the original data in Step 2) specifically includes: Step 2-1) Write English code names for the names of each original feature variable and the target variable to facilitate subsequent data processing; Step 2-2) Correct the obviously incorrect data in the original data; Step 2-3) Verify the outliers in the original data; Step 2-4) Impute the missing values in the original data.

[0011] Furthermore, in Step 3) of dividing the original data set into a training set and a test set, it specifically includes: randomly dividing the original data set obtained after preprocessing into a training set and a test set according to a ratio of 7:3 or 8:2.

[0012] Furthermore, in Step 4), the SHAP algorithm and the Boruta algorithm are used to calculate the core feature set, specifically including selecting the collection of the top N feature variables ranked by feature importance from the original feature variables according to the SHAP algorithm and the Boruta algorithm as the core feature set, and N is preferably ≥15.

[0013] Furthermore, in Step 5) of constructing different optimal basic models, it specifically includes: Step 5-1) Algorithm selection: The algorithms used to construct the basic models include: random forest algorithm, support vector machine algorithm, logistic regression algorithm, XGBoost algorithm, naive Bayes algorithm, and gradient boosting decision tree algorithm; each algorithm constructs a basic model respectively; Step 5-2) Determination of training subset: Using the core feature set in Step 4) as the initial feature set to be examined, extract a training subset from the training set that only contains the features of the feature set to be examined and the target variable; Step 5-3) Data balancing: If the training subset is unbalanced data, that is, there are significant differences in the proportions of each category in the target variable of the data set, then resample the training subset to obtain the resampled training subset; Step 5-4) Hyperparameter search: Based on the resampled training subset, use the algorithm in Step 5-1) to perform model fitting respectively for hyperparameter search; The methods of hyperparameter search include any one or several of grid search, random search, and Bayesian optimization; Step 5-5) Cross-validation and manual adjustment of hyperparameters: Using ROC-AUC as the evaluation index, based on the hyperparameter search results and the resampled training subset, use the algorithm in Step 5-1) to perform model fitting again for cross-validation; According to the performance of the ROC-AUC of the model, perform manual parameter tuning on the basis of the hyperparameters determined in Step 5-4), and then repeat cross-validation until the model obtains the optimal ROC-AUC performance, thereby obtaining optimized hyperparameters; Step 5-6) Fitting of the basic model: Based on the algorithm in Step 5-1), the resampled training subset in Step 5-3), and the optimal hyperparameters determined in Step 5-5), perform model fitting again to obtain the final basic model; Step 5-7) Evaluation of the basic model: Based on the test set, comprehensively evaluate the final basic model obtained in Step 5-6) using the core indicators; The core indicators for model verification include accuracy, precision, ROC-AUC, recall rate, F1 score, and log loss; Step 5-8) Obtain the optimal basic model of a certain algorithm; Step 5-9) Obtain the optimal basic models of all algorithms.

[0014] Furthermore, the feature set to be examined is a dynamic feature set, which is adjusted dynamically according to the final model performance; Standardize the continuous variables in the training subset to obtain the standardized training subset, and at the same time obtain the standardizer, which are used as the data set and standardization basis for subsequent processing respectively.

[0015] Further, the step 5-8) of obtaining the optimal basic model of a certain algorithm specifically includes: Based on a certain algorithm in step 5-1), starting from step 5-2), each time a feature is randomly removed from the to-be-examined feature set in step 5-2) to form a new to-be-examined feature set, and then the loop from step 5-2) to step 5-7) is executed once; if the comprehensive performance of the model obtained after the loop is improved, then confirm to remove feature A from the to-be-examined feature set, and randomly reduce another feature, and re-enter the loop from step 5-2); if the performance of the model decreases after the loop, then retain feature A and randomly remove another feature, and re-enter the loop from step 5-2); after several such loops, under a certain optimal feature set, satisfactory cross-validation and verification results based on the test set are obtained, so as to obtain the optimal basic model of a certain algorithm; the step 5-9) of obtaining the optimal basic models of all algorithms specifically includes: Determine an algorithm, and then loop and execute steps 5-2) to 5-8), so as to obtain the optimal basic models of all algorithms in step 5-1); the optimal basic model generated by each algorithm corresponds to a unique optimal feature set, that is, the input features of the optimal basic model.

[0016] Further, the construction of the integrated machine learning model in step 6) specifically includes: Step 6-1) Compare the comprehensive performance of the optimal basic models generated by different algorithms in step 5), and select the optimal basic models of several algorithms with the best comprehensive performance; Step 6-2) Use the voting method or stacking integration method to fit several of the best-performing basic models to obtain an integrated learning model; Step 6-3) After obtaining the integrated learning model, use the cross-validation method to evaluate the performance of the integrated learning model, and test the generalization ability of the integrated learning model on an independent test set; Step 6-4) By adjusting the integration method, integrated model parameters, and the number of basic models, continuously iterate the process of steps 6-1) to 6-4); after each iteration, re-perform cross-validation and verification based on the test set until satisfactory cross-validation results and test set results are obtained; finally, obtain the integrated machine learning model.

[0017] Further, the prediction and optimization method for the treatment effect of a bacterial infection disease is characterized by further including: Step 7) Implementation of prediction; the implementation of step 7) of prediction specifically includes: First, find the collection of the input features of the optimal basic model selected in step 6) as the input features of the integrated machine learning model; collect the clinical information corresponding to the input features of the integrated machine learning model, and after being processed by a normalizer, input it into the integrated machine learning model, and then the prediction result of the treatment effect of the corresponding bacterial infection disease can be obtained.

[0018] Advantages of the present invention: 1) The present invention can help clinicians quickly identify the optimal treatment plan, thereby improving the treatment success rate and reducing a series of subsequent problems caused by treatment failure; the present invention not only improves the prediction accuracy of the treatment effect, but also provides a scientific basis for the optimization of anti-infection treatment, which is of great significance for reducing antimicrobial exposure, reducing medical costs, and reducing adverse drug reactions, etc.; 2) The method of single-task learning modeling and multi-task prediction deployment of the present invention effectively improves the prediction performance and generalization ability of the model, and is applicable to a wider range of clinical scenarios; the present invention further solves the problems of insufficient interpretability of the prediction model and the inability of the prediction model to be implemented; 3) By accurately predicting the effect of each treatment plan, the present invention can avoid the overuse of antimicrobial drugs in empirical combination drug use; reduce antimicrobial exposure; reduce adverse drug reactions related to drug treatment; improve the success rate of anti-infection treatment through efficacy prediction, and reduce the average length of hospital stay, meeting the current clinical requirements for antimicrobial management and rational use; 4) The present invention uses single models such as random forest algorithm, support vector machine algorithm, and logistic regression algorithm to construct an ensemble learning system, and combines the advantages of different models through a stacking classifier or a voting machine, significantly improving the prediction accuracy and robustness; the model performance has been verified by various indicators such as ROC-AUC (area under the receiver operating characteristic curve), accuracy, recall rate, and F1 score. Description of the Drawings

[0019] Att Figure 1 is a schematic overall flow chart of the method of Embodiment 1 of the present invention.

[0020] Att Figure 2 is a SHAP force plot for model interpretability of Embodiment 1 of the present invention. Detailed Embodiments

[0021] A method for predicting and optimizing the treatment effect of bacterial infection diseases, the method comprising: Step 1) Collect patient clinical information to obtain raw data; Step 2) Preprocess the raw data to obtain a raw data set; Step 3) Divide the raw data set to obtain a training set and a test set; Step 4) Based on the training set, use the SHAP algorithm and the Boruta algorithm to calculate the core feature set; Step 5) Select the optimal feature combination from the core feature set to construct different optimal basic models; Step 6) Based on the basic models, use an ensemble algorithm to construct an ensemble machine learning model.

[0022] A method for predicting and optimizing the treatment effect of bacterial infectious diseases, the method further includes: Step 7) Implementation of the prediction.

[0023] The original data obtained by collecting the clinical information of the patient in step 1) includes original feature variables and target variables; the target variable is the clinical outcome, and the target variable is divided into two situations: treatment success and treatment failure, which is a binary variable; the original feature variables include the general situation of the patient, co-existing diseases, co-administered medications, examination and test indicators; the general situation of the patient includes the patient's sex (Sex), age (Age), weight (Weight), etc.; the co-existing diseases include co-existing diabetes (DM), co-existing tumors (CA), co-existing cardiovascular diseases (CVD), etc.; the co-administered medications include whether other anti-infective drugs are combined (Ery), etc.; the examination and test indicators include white blood cell count (WBC), C-reactive protein (CRP), procalcitonin (PCT), etc.; the data types of the original feature variables include continuous variables (such as: WBC, etc.) and binary variables (such as: co-existing DM, etc.).

[0024] The original data set obtained by preprocessing the original data in step 2) specifically includes: Step 2-1) Write English codes for the names of each original feature variable and the target variable to facilitate subsequent data processing; Step 2-2) Correct the obviously incorrect data in the original data; Step 2-3) Verify the outliers in the original data; Step 2-4) Impute the missing values in the original data.

[0025] The step 3) of dividing the original data set into a training set and a test set specifically includes: randomly dividing the original data set obtained after preprocessing into a training set and a test set according to a ratio of 7:3 or 8:2; the random division method preferably uses random stratified sampling.

[0026] In step 4), using the SHAP algorithm and the Boruta algorithm based on the training set, calculate the core feature set, specifically including calculating the relevant indicators of feature importance using the SHAP algorithm and the Boruta algorithm respectively, and sorting the importance of the features according to the level of the indicators (the higher the feature index, the higher its importance), so as to perform feature selection. Specifically: when performing feature selection, select the collection of the top N feature variables with the highest feature importance from the original feature variables according to the SHAP algorithm and the Boruta algorithm as the core feature set, and N is preferably ≥15.

[0027] In step 5), constructing different optimal basic models specifically includes: Step 5-1) Algorithm Selection: The algorithms used to build the basic model include: Random Forest algorithm (RF), Support Vector Machine algorithm (SVM), Logistic Regression algorithm (LR), XGBoost algorithm, Naive Bayes algorithm (GNB), Gradient Boosting Decision Tree algorithm (GBDT), etc.; Each algorithm corresponds to building a basic model respectively. For example, use the RF algorithm to build an RF model; Step 5-2) Determination of Training Subset: Use the core feature set in Step 4) as the initial feature set to be examined, and extract a training subset from the training set that only contains the features of the feature set to be examined and the target variable; The feature set to be examined is a dynamic feature set, which is adjusted dynamically according to the final model performance; Standardize the continuous variables in the training subset to obtain the standardized training subset, and at the same time obtain the standardizer, which are used as the data set for subsequent processing and the standardization basis respectively; The standardization methods mainly include Z-score standardization, min-max normalization, and mean normalization, etc.; Step 5-3) Data Balancing: If the training subset is unbalanced data, that is, there are significant differences in the proportions of various types in the target variable of the data set, then resample the training subset to obtain the resampled training subset; The resampling methods include oversampling and / or undersampling, etc.; Step 5-4) Hyperparameter Search: Based on the resampled training subset, use the algorithms in Step 5-1) to perform model fitting respectively and conduct hyperparameter search; The hyperparameter search methods include any one or several of grid search, random search, and Bayesian optimization; Step 5-5) Cross-validation and Manual Adjustment of Hyperparameters: Using ROC-AUC as the evaluation index, based on the hyperparameter search results and the resampled training subset, use the algorithms in Step 5-1) to perform model fitting again and conduct cross-validation; The preferred cross-validation methods are StratifiedKFold (10-fold) method and / or Bootstrap (1000 iterations) method; In this step, according to the performance of the ROC-AUC of the model, manual parameter tuning can be performed on the basis of the hyperparameters determined in Step 5-4), and then cross-validation is repeated until the model obtains the optimal ROC-AUC performance, so as to obtain the optimized hyperparameters; Step 5-6) Fitting of Basic Model: Based on the algorithms in Step 5-1), the resampled training subset in Step 5-3), and the optimal hyperparameters determined in Step 5-5), perform model fitting again to obtain the final basic model; Step 5-7) Basic model evaluation: Based on the test set, comprehensively evaluate the final basic model obtained in Step 5-6) using core metrics; the core metrics for model verification include accuracy, precision, ROC-AUC, recovery rate, F1 score (the harmonic mean of precision and recall), and LogLoss (logarithmic loss); Step 5-8) Obtaining the optimal basic model for a certain algorithm: Based on a certain algorithm in Step 5-1) (such as RF), starting from Step 5-2), each time randomly remove one feature (denoted as feature A) from the set of features to be investigated in Step 5-2) to form a new set of features to be investigated, and then loop once from Step 5-2) to Step 5-7); if the comprehensive performance of the model obtained after the loop improves, confirm the removal of feature A from the set of features to be investigated, and randomly reduce another feature, and re-enter the loop from Step 5-2); if the performance of the model deteriorates after the loop, retain feature A and randomly remove another feature, and re-enter the loop from Step 5-2); after several such loops, under a certain optimal feature set, obtain satisfactory cross-validation and verification results based on the test set, so as to obtain the optimal basic model for a certain algorithm (such as RF); the satisfactory verification results refer to that based on the core metrics, in the above-mentioned several loops, both the cross-validation and the verification results based on the test set perform optimally; Step 5-9) Obtaining the optimal basic models for all algorithms: Determine an algorithm and then loop through Steps 5-2) to 5-8), so as to obtain the optimal basic models for all algorithms in Step 5-1); each optimal basic model generated by an algorithm corresponds to a unique optimal feature set, that is, the input features of this optimal basic model.

[0028] The construction of the integrated machine learning model in Step 6) specifically includes: Step 6-1) Compare the comprehensive performance of the optimal basic models generated by different algorithms in Step 5), and select the optimal basic models of several (preferably 2 or more) algorithms with the best comprehensive performance; Step 6-2) Use the voting method (VotingClassifier) or the stacking method (StackingClassifier) to fit several (preferably 2 or more) best-performing basic models to obtain an integrated learning model; the voting method (VotingClassifier) uses hard voting (Majority Voting) or soft voting (Soft Voting) for integration; for the stacking method (StackingClassifier), when performing the prediction function, a meta-learner needs to be constructed, and the prediction results of the basic models based on the training set are used as the input of the meta-learner to train the meta-learner; After obtaining the integrated learning model in step 6-3), the cross-validation method is used to evaluate the performance of the integrated learning model, and the generalization ability of the integrated learning model is tested on an independent test set; Step 6-4) By adjusting the integration method (voting method or stacking method), the parameters of the integrated model, the number of base models, etc., continuously iterate the processes of the above steps 6-1) to 6-4); after each iteration, re-perform cross-validation and validation based on the test set until satisfactory cross-validation results and test set results are obtained; finally, an integrated machine learning model is obtained.

[0029] In the present invention, each time a new model is generated, cross-validation and validation based on the test set are required. All cross-validations are based on the optimal training subset of the model; all test sets are the same test set, which is an independent test set.

[0030] The implementation of the prediction in step 7): The optimal base models selected in step 6) may have different input features. First, find the union of the input features of these optimal base models as the input features of the integrated machine learning model; collect the clinical information corresponding to the input features of the integrated machine learning model, and after being processed by the normalizer in step 5-2), input it into the integrated machine learning model to obtain the prediction results of the treatment effects of the corresponding bacterial infection diseases; the prediction results include the prediction results of monotherapy and / or combination therapy.

[0031] The present invention combines the clinical characteristics of patients and treatment plan variables, and uses the integrated machine learning model obtained in step 6) to achieve accurate prediction of the treatment effects of the corresponding treatment plans for bacterial infection diseases; for example: the present invention can achieve accurate prediction of the treatment effects of piperacillin sodium tazobactam monotherapy and combined with erythromycin for acute exacerbation of chronic obstructive pulmonary disease complicated with bacterial lower respiratory tract infection, and provide a reference for the optimization of anti-infection plans. Example 1

[0032] The following further illustrates the specific implementation manner of the method for predicting and optimizing the treatment effects of a bacterial infection disease according to the present invention by taking the acute exacerbation of chronic obstructive pulmonary disease (AECOPD) complicated with bacterial lower respiratory tract infection (LRTIs) as an example in combination with the attached drawings, which specifically includes the following steps: Step 1) Original data collection; As shown in the attached Figure 1 figure, collect the data of all eligible AECOPD patients complicated with bacterial LRTIs from January 1, 2021 to July 31, 2023 according to the established inclusion and exclusion criteria from the inpatient doctor station system and the patient 360 system, a total of 658 cases; among them, 489 cases received monotherapy with piperacillin sodium tazobactam for injection (TZP), and 169 cases received combined treatment with TZP and erythromycin lactobionate for injection; Step 2) Data preprocessing: The collected raw data includes 1 target variable (Result) and 24 feature variables. The feature variables include: ① Continuous variables: Age, Weight, Estimated Glomerular Filtration Rate (EGFR), Blood Urea Nitrogen (BUN), Serum Albumin (ALB), White Blood Cell Count (WBC), D-dimer (DD), C-reactive Protein (CRP), Procalcitonin (PCT), Neutrophil-to-Lymphocyte Ratio (NLR), and Lymphocyte Count (LYM), a total of 11 variables; ② Categorical variables (all binary variables): Body Temperature (Temp), Gender (Sex), Respiratory Failure (RF), Recent Hospitalization History (RH), Diabetes Mellitus (DM), Active Tumor (CA), Bronchiectasis (BE), Interstitial Lung Disease (ILD), Cerebrovascular Disease (CVD), Heart Failure (HF), Outpatient Treatment History before Admission (OPTH), TZP Dose Type (TZPD), and Treatment Regimen (i.e., whether combined with Erythromycin Lactobionate for injection, Ery), a total of 13 variables. Perform preprocessing such as outlier confirmation and missing value imputation on the 24 feature variables in the above raw data to obtain the original dataset. Step 3) Dataset division: Randomly sample the preprocessed original dataset according to a ratio of 7:3 and divide it into a training set and a test set. Step 4) Feature selection: Based on the training set in Step 3), calculate the feature importance indicators of the 24 feature variables using the SHAP algorithm and the Boruta algorithm respectively. Sort the importance of the features according to the above indicators from high to low, and find the union of the top 18 features ranked by the two algorithms. Finally, determine that a total of 16 features, namely CRP, ALB, WBC, DD, Age, PCT, BUN, NLR, LYM, Weight, EGFR, TZPD, CA, DM, OPTH, and Ery, are the core feature set. Step 5) Basic model construction: Use RF, SVM, LR, and GBDT to construct basic models respectively. The construction of the basic model, taking RF as an example, specifically includes the following steps: Step 5-1) Determination of the training subset: Use the core feature set in Step 4) as the feature set to be investigated, and combine it with the training set in Step 3) to obtain a training subset containing 16 feature variables and 1 target variable. Perform Z-score standardization on the continuous variables (CRP, ALB, WBC, DD, Age, PCT, BUN, NLR, LYM, Weight, and EGFR) in this training subset to obtain the standardized training subset and the standardizer. Step 5-2) Data balancing: Undersample the training subset finally obtained in Step 5-1) to obtain the optimal balanced subset. Step 5-3) Hyperparameter grid search: Based on the balanced optimal subset and the RF algorithm, perform grid search for hyperparameters to obtain the preliminary hyperparameters of the RF model; Step 5-4) Cross-validation and manual tuning: Using ROC-AUC as the evaluation metric, based on the preliminary hyperparameters of the RF model, the optimal subset in Step 5-2), and the RF algorithm, perform cross-validation using Bootstrap (1000 iterations); in this step, manual tuning can be performed according to the cross-validation results of ROC-AUC, and then cross-validation is performed again until satisfactory ROC-AUC results are obtained, thereby obtaining the optimized hyperparameters of the RF model; Step 5-5) RF basic model fitting: Based on the optimal subset in Step 5-2), the optimized hyperparameters in Step 5-4), and the RF algorithm, perform model fitting to finally obtain the RF basic model corresponding to the feature set to be investigated in Step 5-1); Step 5-6) RF basic model evaluation: Based on the test set in Step 3), evaluate the RF basic model in Step 5-5) using accuracy, precision, recall, ROC-AUC, F1 score, and Log loss; Step 5-7) Obtain the optimal RF basic model: The above steps start from 5-1), each time randomly removing one feature (denoted as feature A) from the feature set to be investigated in Step 5-1) to form a new feature set to be investigated, and then looping once from Step 5-1) to Step 5-6); if the comprehensive performance of the model obtained after the loop improves, confirm the removal of feature A from the feature set to be investigated, and randomly reduce another feature, and re-enter the loop from Step 5-1); if the performance of the model decreases after the loop, retain feature A and randomly remove another feature, and re-enter the loop from Step 5-1); after several such loops, under the optimal feature set, satisfactory cross-validation and verification results based on the test set are obtained, thereby obtaining the optimal basic model of RF; the features included in the optimal feature set are: CRP, ALB, WBC, DD, Age, PCT, BUN, NLR, LYM, Weight, EGFR, TZPD, CA, OPTH, and Ery, a total of 15 features; Step 5-8) Obtain the optimal basic models of all algorithms: Determine an algorithm and then loop through Steps 5-1) to 5-7), so as to obtain the optimal basic models of RF, SVM, LR, and GBDT respectively; the optimal basic model of each algorithm corresponds to a unique optimal feature set, that is, the input features of this optimal basic model; Step 6) Ensemble learning model construction: Select the three optimal basic models of RF, SVM, and LR in Step 5-8) for ensemble; the main steps for constructing the ensemble learning model include: Step 6-1) Determine the input features of the ensemble learning model: The optimal features of the three basic models of RF, SVM, and LR determined in Step 5) are all: CRP, ALB, WBC, DD, Age, PCT, BUN, NLR, LYM, Weight, EGFR, TZPD, CA, OPTH, and Ery; these 15 features are used as the input features of the ensemble learning model; Step 6-2) Determine the adjustable parameters: 1) The adjustable parameters of the ensemble learning model mainly include the type of ensemble (voting method or stacking method); 2) When using the voting method, whether it is soft voting or hard voting; 3) Whether to select two or three of the three basic models of RF, SVM, and LR; 4) When selecting two of the aforementioned three models, determine which two to select; 5) How to allocate the weights of different basic models; Step 6-3) Construct the ensemble learning model: Based on the adjustable parameters in Step 6-1), combine different model ensemble strategies and fit different ensemble learning models; Step 6-4) Validate the ensemble learning model: Based on the test set, use the BootStrap (1000 iterations) cross-validation method and validation based on the test set to evaluate the performance of the ensemble models under different parameter combination strategies; Step 6-5) Obtain the optimal ensemble learning model: Based on accuracy, precision, recall rate, ROC-AUC, F1 score, and Log loss, finally determine that the ensemble learning model based on the voting method (soft voting) with the three models of RF, SVM, and LR as the basic models (weight ratio 1:1:1) is the optimal ensemble learning model; Step 7) Model deployment and prediction: Deploy the trained ensemble learning model in Step as a Web application. Users input 14 features of the patient: CRP, ALB, WBC, DD, Age, PCT, BUN, NLR, LYM, Weight, EGFR, TZPD, CA, OPTH. Then the application can predict the treatment effects of two treatment regimens: single-agent TZP (feature Ery = 0) and combination therapy (feature Ery = 1) at the same time; for the 15th feature (Ery), the web backend automatically adds two cases of Ery = 0 and Ery = 1 respectively.

[0033] Effect of this embodiment:

[0034] 1. Performance verification of the ensemble learning model: On the test set, the ensemble learning model shows excellent prediction performance: 1) ROC-AUC: 0.71; 2) Accuracy: 0.69; 3) Recall rate: 0.72; 4) F1 score: 0.47.

[0035] 2. Model interpretability verification: As shown in the appendix Figure 2The SHAP force plot shown below demonstrates the model's explanation for a single prediction case: 1) Case of treatment failure (Figure a in Figure 2 ): The predicted probability of this case is 0.84, indicating a high probability of treatment failure. The SHAP force plot shows that the main factors promoting treatment failure (dark part) are: relatively high NLR value (%: 0.93), WBC (11.27×10 9 / L), and DD (1.72 μg / ml) values; the factors inhibiting treatment failure (light part) are: relatively low CRP (0.5 mg / l) value; factors such as age (74 years old) and outpatient treatment history (OPTH = 0) also have an impact on the prediction result.

[0036] 2) Case of treatment success (Figure b in Figure 2 ): The predicted probability of this case is 0.23, indicating a high probability of treatment success. The SHAP force plot shows that the main factors promoting treatment success (dark part) are: ideal ALB (42.9 g / L), relatively low CRP (6.31 mg / l), and PCT (0.18 ng / ml) values; indicators such as age (52 years old) and EGFR (103.15 ml / min) are at a favorable level; although underlying diseases such as bronchiectasis (BE = 1) exist, their impact is limited.

[0037] 3. Clinical application value: The method of the present invention can accurately identify the expected effects of different treatment regimens, such as: the situation where single-drug treatment fails while combination treatment is successful, and the situation where both treatment regimens are successful, etc.

[0038] It can be seen from the above embodiments that the present invention can not only realize the prediction of personalized treatment regimens for AECOPD complicated with bacterial LRTIs, but also facilitate the actual implementation of the prediction model in the form of a Web application, providing a reliable reference basis for clinical decision-making. This method can be extended and applied to the prediction and optimization of the treatment effects of other bacterial infectious diseases.

Claims

1. A method for predicting and optimizing the therapeutic effect of bacterial infection diseases, characterized in that include: Step 1) Collect the patient's clinical information to obtain the original data; Step 2) Preprocess the original data to obtain the original data set; Step 3) Divide the original data set into training set and test set; Step 4) Based on the training set, use the SHAP algorithm and the Boruta algorithm to calculate the core feature set; Step 5) Select the optimal feature combination from the core feature set and build different optimal basic models; Step 6) Based on the basic model, use the integrated algorithm to build an integrated machine learning model.

2. The method for predicting and optimizing the treatment effect of bacterial infection diseases according to claim 1, characterized in that The original data obtained by collecting the clinical information of the patient in step 1) includes original characteristic variables and target variables; the target variable is the clinical outcome, and the target variable is divided into two situations: treatment success and treatment failure; the original characteristic variables include the patient's general condition, concomitant diseases, concomitant medication, and inspection and testing indicators; the patient's general condition includes the patient's gender, age, and weight; the concomitant diseases include concomitant diabetes, concomitant tumors, and concomitant cardiovascular diseases; the concomitant medication includes whether other anti-infective drugs are used in combination; the inspection and testing indicators include white blood cell count, C-reactive protein, and procalcitonin; The data types of original feature variables include continuous variables and binary variables.

3. The method for predicting and optimizing the treatment effect of bacterial infection diseases according to claim 1, characterized in that The step 2) preprocesses the original data to obtain the original data set, specifically including: Step 2-1) Write the names of each original feature variable and target variable in English to facilitate subsequent data processing; Step 2-2) Correct the obviously erroneous data in the original data; Step 2-3) Verify outliers in the original data; Steps 2-4) Interpolate missing values ​​in the original data.

4. The method for predicting and optimizing the treatment effect of bacterial infection diseases according to claim 1, characterized in that The step 3) divides the original data set into a training set and a test set, specifically comprising: randomly dividing the original data set obtained after preprocessing into a training set and a test set in a ratio of 7:3 or 8:

2.

5. The method for predicting and optimizing the treatment effect of bacterial infection diseases according to claim 1, characterized in that In the step 4), the SHAP algorithm and the Boruta algorithm are used to calculate the core feature set, which specifically includes selecting the set of feature variables ranked top N in terms of feature importance from the original feature variables according to the SHAP algorithm and the Boruta algorithm as the core feature set, where N≥15.

6. The method for predicting and optimizing the treatment effect of bacterial infection diseases according to claim 1, characterized in that In step 5), different optimal basic models are constructed, specifically including: Step 5-1) Algorithm selection: The algorithms used to build the basic model include: random forest algorithm, support vector machine algorithm, logistic regression algorithm, XGBoost algorithm, naive Bayes algorithm, gradient boosting decision tree algorithm; each algorithm corresponds to building a basic model; Step 5-2) Determine the training subset: Take the core feature set in step 4) as the initial feature set to be examined, and extract a training subset containing only the features of the feature set to be examined and the target variable from the training set; Step 5-3) Data balance: If the training subset is unbalanced data, that is, the proportion of each category in the target variable of the data set is significantly different, then the training subset is resampled to obtain the resampled training subset; Step 5-4) Hyperparameter search: Based on the resampled training subset, use the algorithm in step 5-1) to perform model fitting and hyperparameter search respectively; the hyperparameter search method includes any one or more of grid search, random search, and Bayesian optimization; Step 5-5) Cross-validation and manual adjustment of hyperparameters: Using ROC-AUC as the evaluation indicator, based on the hyperparameter search results and the resampled training subset, use the algorithm in step 5-1) to fit the model again and perform cross-validation; Based on the ROC-AUC performance of the model, manually adjust the hyperparameters determined in step 5-4), and then repeat cross-validation until the model achieves the best ROC-AUC performance, thereby obtaining the optimized hyperparameters; Step 5-6) Basic model fitting: Based on the algorithm in step 5-1), the training subset after resampling in step 5-3), and the optimal hyperparameters determined in step 5-5), the model is fitted again to obtain the final basic model; Step 5-7) Basic model evaluation: Based on the test set, the final basic model obtained in step 5-6) is comprehensively evaluated using core indicators; the core indicators of the model verification include accuracy, precision, ROC-AUC, recovery rate, F1 score and logarithmic loss; Steps 5-8) Obtain the optimal basic model of a certain algorithm; Steps 5-9) Obtain the optimal basic model for all algorithms.

7. A method for predicting and optimizing the treatment effect of bacterial infection diseases according to claim 6, characterized in that The feature set to be examined is a dynamic feature set, which is dynamically adjusted according to the final model performance; the continuous variables in the training subset are standardized to obtain a standardized training subset, and a normalizer is obtained at the same time, which are used as a data set and a standardization basis for subsequent processing respectively.

8. The method for predicting and optimizing the treatment effect of bacterial infection diseases according to claim 6, characterized in that The step 5-8) obtains the optimal basic model of a certain algorithm, specifically including: based on a certain algorithm in step 5-1), the above steps start from 5-2), randomly remove a feature from the feature set to be examined in step 5-2) each time to form a new feature set to be examined, and then cycle from step 5-2) to step 5-7) once; if the comprehensive performance of the model obtained after the cycle is improved, confirm to remove feature A from the feature set to be examined, and randomly reduce another feature, and re-enter the cycle from step 5-2); if the performance of the model decreases after the cycle, retain feature A, and then The machine removes another feature and re-enters the loop from step 5-2); after several cycles, under a certain optimal feature set, satisfactory cross-validation and verification results based on the test set are obtained, thereby obtaining the optimal basic model of a certain algorithm; the step 5-9) obtains the optimal basic model of all algorithms, specifically including: determining an algorithm, and then cyclically executing steps 5-2) to 5-8), so as to obtain the optimal basic model of all algorithms in step 5-1); the optimal basic model generated by each algorithm corresponds to a unique optimal feature set, that is, the input feature of the optimal basic model.

9. The method for predicting and optimizing the treatment effect of bacterial infection diseases according to claim 1, characterized in that The step 6) of constructing an integrated machine learning model specifically includes: Step 6-1) Compare the comprehensive performance of the optimal basic models generated by different algorithms in step 5), and select the optimal basic models of several algorithms with the best comprehensive performance; Step 6-2) Use voting ensemble or stacking ensemble to fit several basic models with the best performance to obtain an ensemble learning model; Step 6-3) After obtaining the ensemble learning model, the cross-validation method is used to evaluate the performance of the ensemble learning model, and the generalization ability of the ensemble learning model is tested on an independent test set; Step 6-4) Continuously iterate the above steps 6-1) to 6-4) by adjusting the integration method, integration model parameters, and the number of basic models; after each iteration, re-perform cross-validation and validation based on the test set until satisfactory cross-validation results and test set results are obtained; finally, an integrated machine learning model is obtained.

10. A method for predicting and optimizing the therapeutic effect of bacterial infection diseases according to any one of claims 1 to 9, characterized in that It also includes: step 7) implementation of prediction; the implementation of step 7) prediction specifically includes: first, obtaining the set of optimal basic model input features selected in step 6) as the input features of the integrated machine learning model; collecting clinical information corresponding to the input features of the integrated machine learning model, and after processing by the normalizer, inputting the clinical information into the integrated machine learning model, so as to obtain the prediction results of the treatment effect of the corresponding bacterial infection disease.