Method for predicting drug resistance of IVIG in KD based on interpretable machine learning model
By optimizing feature selection using the random forest ranger algorithm and SHAP values, combined with novel biomarkers, a highly interpretable IVIG resistance prediction model was constructed. This model addresses the existing issues of insufficient nonlinear relationship capture, variable screening, and data processing in IVIG resistance prediction, enabling efficient and real-time clinical applications and reducing the risk of coronary artery complications in patients with Kawasaki disease.
Patent Information
- Application Number
- CN202510935514.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-08
- Publication Date
- 2025-10-17
AI Technical Summary
Existing technologies for predicting IVIG resistance in Kawasaki disease have problems such as difficulty in capturing nonlinear relationships, reliance on subjective expert experience in variable screening, insufficient data processing capabilities, information loss, and insufficient interpretability. They also lack real-time visualization tools and fail to fully utilize new biomarkers.
The random forest ranger algorithm was combined with the SHAP value and recursive feature elimination algorithm for feature selection, retaining the continuous variable form and building a highly interpretable machine learning model. A web-based prediction platform based on the R language was developed, introducing new biomarkers such as venous oxygen content and lactate to adapt to large-scale data analysis.
It improves the accuracy and interpretability of IVIG resistance prediction and provides a real-time, visual clinical decision support tool to help identify high-risk patients and optimize treatment strategies, thereby reducing the incidence of coronary artery complications.
Smart Images

Figure CN120809207A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The application provides an IVIG resistance prediction method based on an interpretable machine learning model in KD, and belongs to the technical field of biological information processing. BACKGROUND
[0002] Kawasaki Disease (KD) is an acute febrile systemic vasculitis in childhood, mainly involving small and medium-sized arteries, especially the coronary arteries, which is the leading cause of acquired heart disease in children in developed countries. Early standardized treatment can significantly reduce the risk of cardiovascular complications, and the standard treatment regimen of intravenous immunoglobulin (IVIG) 2g / kg has become the internationally recognized preferred strategy. However, about 10% to 20% of children have poor response to the first IVIG treatment, and develop persistent or recurrent high fever, which is clinically referred to as "IVIG resistance (IVIGR)".
[0003] Studies have shown that IVIG resistance is highly related to the occurrence of coronary artery lesions (CALs). Early identification of high-risk patients and timely administration of intensive treatment (such as combined glucocorticoids or immunosuppressants) can help improve the prognosis. Therefore, it is of great significance to establish an accurate and reliable IVIG resistance prediction model for optimizing clinical treatment strategies.
[0004] At present, several prediction scoring systems have been proposed, such as Egami, Kobayashi and Sano scores, which are all based on logistic regression method. Due to its few variables, simple calculation and strong interpretability, it has certain application value in clinical practice. However, these traditional models still have the following limitations: the artificial division of continuous variables into two categories may introduce bias, it is difficult to capture the complex nonlinear relationship between variables, and the performance is limited when dealing with large-scale data.
[0005] In recent years, the development of artificial intelligence, especially machine learning (ML) technology, has provided a new way for disease risk prediction. ML can handle large-scale, multi-dimensional and high-complexity clinical data, mine potential feature associations and improve model prediction performance. However, due to its "black box" characteristics, ML models still face the problem of insufficient interpretability in clinical application. SUMMARY
[0006] The present application proposes an IVIG resistance prediction method based on an interpretable machine learning model in KD and an online prediction system to solve the following technical problems:
[0007] 1. Enhance the modeling ability of non-linear and complex feature relationships
[0008] Traditional linear models such as logistic regression have difficulty effectively capturing the non-linear relationships and high-order interaction effects between clinical variables, limiting the predictive accuracy of the model. This invention uses non-linear ensemble learning algorithms such as Random Forest with ranger algorithm (RF ranger) to strengthen the modeling ability of complex variable relationships, thereby improving the prediction performance.
[0009] 2. Realize objective and interpretable feature selection
[0010] Traditional variable screening relies heavily on expert subjective experience, which has the risk of missing important features and bias. This invention combines SHAP (Shapley Additive Explanations) value ranking and Recursive Feature Elimination (RFE) algorithm to build a scientific and transparent variable screening process, ensuring the objectivity and interpretability of feature selection.
[0011] 3. Improve the processing ability of large-scale clinical data
[0012] With the development of Electronic Medical Record (EMR) systems, the dimension and complexity of clinical data are increasing, and traditional models are difficult to fully integrate and utilize high-dimensional information. This invention automatically extracts and processes 66 clinical and laboratory variables, fully utilizes multi-dimensional data features, and adapts to the analysis needs of large-scale high-dimensional data.
[0013] 4. Avoid information loss caused by variable binary classification processing
[0014] Traditional models usually perform binary classification on continuous variables based on thresholds, resulting in information loss and classification bias.
[0015] This invention retains the original continuous variable form, combines SHAP value to analyze the actual contribution of each variable to the prediction result, and improves the accuracy and stability of the model.
[0016] 5. Improve the clinical practicability of prediction tools
[0017] Existing models are mostly in the research stage, lacking visual, real-time and easy-to-operate clinical prediction tools.
[0018] The application develops a web-based prediction platform based on R language, and clinical doctors can input individual characteristics of patients to obtain real-time individualized IVIG drug resistance risk prediction results, thereby significantly improving the usability and practical value of the model in clinical application.
[0019] 6. Expanding variable selection to mine potential biomarkers
[0020] Existing studies are mostly limited to traditional laboratory indicators, ignoring the role of new potential biomarkers.
[0021] The application introduces new indicators such as central venous oxygen content (CvO2) and lactic acid (LAC), and evaluates their predictive importance through the SHAP method, enriching the biological interpretation basis of the prediction model.
[0022] 7. Adapt to regional population characteristics and realize model localization
[0023] In view of the clinical characteristics of KD patients, the application develops and optimizes an IVIG drug resistance prediction model suitable for Suzhou area, thereby improving the prediction accuracy of the model.
[0024] To solve the above problems, the application is based on the electronic medical record data of Kawasaki disease children, uses 13 kinds of machine learning algorithms to construct an IVIG drug resistance prediction model, and screens key variables through the SHAP method to realize the explainability optimization of the model. Finally, an IVIG drug resistance prediction model with accuracy and explainability is constructed, and an online visualization tool is developed for real-time access and decision support of clinical doctors, helping to realize precision medicine.
[0025] SHAP (SHapley Additive exPlanations) value as a model explanation method based on game theory can quantify the specific contribution of each feature to the model prediction result, effectively improve the transparency and clinical understandability of the ML model, and promote its practical application in the medical field.
[0026] The specific technical scheme is:
[0027] The KD IVIG drug resistance prediction method based on the explainable machine learning model comprises the following steps:
[0028] S1. Data acquisition
[0029] Automatic data extraction is performed based on the electronic medical record system, and the accuracy of 20% samples is verified by double researchers.
[0030] S2. Data processing
[0031] Variables with missing values over 50% were removed, and the rest were filled with multiple imputation;
[0032] Categorical variables were ordered, and continuous variables were directly included in modeling;
[0033] Through Spearman correlation analysis, when the correlation coefficient is greater than 0.6, exclude highly correlated redundant variables;
[0034] Finally, 66 variables were retained for modeling;
[0035] The 66 variables include: gender, coronary artery lesions, incomplete Kawasaki disease, conjunctival hyperemia, oral mucosal changes, hand and foot edema, perionychium desquamation, rash, cervical lymph node enlargement, perianal desquamation, number of clinical features, time to diagnosis, age of onset, white blood cell count, neutrophil percentage, lymphocyte percentage, monocyte percentage, eosinophil percentage, basophil percentage, red blood cell count, hemoglobin, platelet count, platelet hematocrit, urine white blood cell count, urine red blood cell count, urine bacterial count, lipase, serum albumin, serum globulin, alanine aminotransferase, aspartate aminotransferase, alkaline phosphatase, gamma glutamyl transpeptidase, direct bilirubin, indirect bilirubin, blood urea nitrogen, creatinine, uric acid, lactate dehydrogenase, creatine kinase, total cholesterol, triglycerides, C-reactive protein, complement C3, complement C4, immunoglobulin A, immunoglobulin G, immunoglobulin M, prealbumin, pH, carbon dioxide partial pressure, oxygen partial pressure, blood glucose, venous oxygen content, oxygen saturation, sodium ion, potassium ion, calcium ion, chloride ion, anion gap, lactic acid, creatine kinase isoenzyme MB, cardiac troponin T, magnesium ion, glycocholic acid, fever duration.
[0036] The data set is divided into training set (70%) and test set (30%);
[0037] S3. Model construction
[0038] Construct the performance of 13 prediction models for comparison;
[0039] Cross-validation to assess model generalization ability;
[0040] Adjust the model parameter optimization model;
[0041] Select the RF ranger model with the best performance and clinical interpretability as the final model;
[0042] S4. Model explanation module:
[0043] Introduce SHAP method for global and local visual explanation of the model;
[0044] Combine SHAP value ranking and recursive feature elimination strategy, compare the area under the ROC curve AUC under different variable numbers, and finally determine the optimal subset containing 10 key features: coronary artery lesion CALs, C-reactive protein CRP, fever duration, venous oxygen content CvO2, oral mucosa change, direct bilirubin, lactic acid, gender, neutrophil percentage NEU%, and rash;
[0045] S5. Webpage deployment
[0046] An R language Shiny platform is used to build a web-based real-time calculator to realize model deployment;
[0047] The user inputs 10 key variable values to automatically generate the IVIG resistance risk probability and prompts the high-risk threshold; if the risk probability exceeds the threshold, the patient belongs to the high-risk group.
[0048] Further, S1. Data collection covers demographic characteristics, clinical symptoms, laboratory test indicators, and imaging data.
[0049] S3. The machine learning algorithm used in model construction includes random forest RF, gradient boosting machine GBM, support vector machine SVM, logistic regression LR, elastic net regularized generalized linear model GLMnet, K-nearest neighbor KNN, decision tree DT, random forest based on ranger algorithm RF ranger, least absolute shrinkage and selection operator LASSO, naive Bayes NB, adaptive boosting Adaboost, neural network NNET, and support vector machine based on kernlab algorithm SVM kernlab.
[0050] In S5. Webpage deployment, the high-risk threshold is based on the Youden index, and the threshold is 11.6%.
[0051] Based on the large-scale electronic medical record data of Kawasaki disease patients, the machine learning model with good interpretability is constructed to predict IVIG resistance, and the online webpage calculation tool is built to input the key indicators in real time for the clinician, and automatically output the IVIG resistance risk probability of the patient. The tool sets 11.6% as the optimal prediction threshold, which can help doctors identify high-risk individuals, adjust treatment strategies in time, and reduce the occurrence of coronary artery complications. BRIEF DESCRIPTION OF DRAWINGS
[0052] Figure 1 is a flowchart of the present application;
[0053] Figure 2 is a ROC curve of 13 machine learning models of the embodiments;
[0054] Figure 3 is a top 10 important features of the RF ranger model based on SHAP values of the embodiments;
[0055] Figure 4 is a top 10 important features of the GBM model based on SHAP values of the embodiments;
[0056] Figure 5 is a top 10 important features of the LASSO model based on SHAP values of the embodiments;
[0057] Figure 6 is a top 10 important features of the LR model based on SHAP values of the embodiments;
[0058] Figure 7 is a top 10 important features of the GLMnet model based on SHAP values of the embodiments;
[0059] Figure 8 is one of the model local explanations using SHAP values of the embodiments;
[0060] Figure 9 is two of the model local explanations using SHAP values of the embodiments;
[0061] Figure 10 is three of the model local explanations using SHAP values of the embodiments;
[0062] Figure 11 is four of the model local explanations using SHAP values of the embodiments;
[0063] Figure 12 is the AUC values after incorporating different number of features of the embodiments;
[0064] Figure 13 is one of the feature scatter plots of the embodiments;
[0065] Figure 14 is two of the feature scatter plots of the embodiments;
[0066] Figure 15 is three of the feature scatter plots of the embodiments;
[0067] Figure 16 is four of the feature scatter plots of the embodiments;
[0068] Figure 17 Figure 5 is a scatter plot of characteristics for an embodiment;
[0069] Figure 18 Figure 6 is a scatter plot of characteristics for an embodiment;
[0070] Figure 19 Figure 7 is a scatter plot of characteristics for an embodiment;
[0071] Figure 20 Figure 8 is a scatter plot of characteristics for an embodiment;
[0072] Figure 21 Figure 9 is a scatter plot of characteristics for an embodiment;
[0073] Figure 22 Figure 10 is a scatter plot of characteristics for an embodiment;
[0074] Figure 23 Figure 11 is an embodiment web-based IVIGR risk prediction calculator;
[0075] Figure 24 Figure 12 is an embodiment web-based IVIGR risk prediction probability pie chart;
[0076] Figure 25 Figure 13 is an embodiment distribution of predicted probabilities of IVIG resistance for two groups of patients;
[0077] Figure 26 Figure 14 is an embodiment Spearman variable correlation heat map. DETAILED DESCRIPTION
[0078] This embodiment is based on the electronic medical record data of 2,371 children with Kawasaki disease in a certain region, and 13 kinds of machine learning algorithms are used to construct IVIG resistance prediction model, and the SHAP method is used to screen key variables to realize the interpretability optimization of the model. Finally, a set of IVIG resistance prediction model with accuracy and interpretability is constructed, and developed as an online visualization tool for real-time access and decision support for clinicians, helping to realize precision medicine. As shown in FIG. 1, the embodiment specifically comprises the following steps: Figure 1
[0079] S1. Data collection
[0080] Automatic data extraction based on electronic medical record system, covering demographic characteristics, clinical symptoms, laboratory test indicators, imaging data, etc. And through double researcher verification 20% sample accuracy.
[0081] S2. Data processing
[0082] Variables with more than 50% missing values are removed, and the rest are filled with missing values by multiple imputation method;
[0083] Categorical variables were ordered and continuous variables were directly included in the modeling;
[0084] Highly correlated redundant variables were excluded by Spearman correlation analysis (correlation coefficient > 0.6);
[0085] Finally, 66 variables were retained for modeling.
[0086] The dataset was divided into training set (70%) and test set (30%);
[0087] S3. Model construction
[0088] 13 kinds of mainstream machine learning algorithms were used to build prediction models, including random forest (RF), gradient boosting machine (GBM), support vector machine (SVM), logistic regression (LR), etc. Cross-validation was used to evaluate the generalization ability of the model. Adjusting model parameters can optimize the model.
[0089] The performance of 13 machine learning models in Kawasaki disease intravenous immunoglobulin resistance is shown in Table 1.
[0090] Table 1 Performance of 13 machine learning models in Kawasaki disease intravenous immunoglobulin resistance
[0091] Model AUC Sensitivity Specificity Positive predictive value Negative predictive value RF 0.754 0.973 0.488 0.879 0.520 GBM 0.777 0.984 0.499 0.881 0.674 SVM 0.753 0.897 0.508 0.898 0.504 GLMnet 0.768 0.754 0.582 0.925 0.459 KNN 0.556 0.903 0.466 0.868 0.291 DT 0.508 0.984 0.433 0.874 0.431 RF ranger 0.775 0.943 0.531 0.893 0.575 SVM kernlab 0.768 0.753 0.571 0.923 0.454 LASSO 0.770 0.759 0.582 0.925 0.462 NB 0.687 0.948 0.509 0.891 0.573 LR 0.769 0.746 0.571 0.922 0.449 AdaBoost 0.743 0.971 0.443 0.885 0.619 NNET 0.499 0.021 0.989 0.929 0.329
[0092] The AUC values of different models are shown in Figure 2 .
[0093] The RF ranger model with the best performance and good clinical interpretability was selected as the final model. The performance of RF ranger model based on different feature numbers is shown in Table 2.
[0094] Table 2 Performance of RF ranger model based on different feature numbers
[0095]
[0096]
[0097]
[0098] S4. Model explanation module:
[0099] The SHapley Additive exPlanation (SHAP) method was introduced to visualize the global and local explanation of the model;
[0100] Figures 3 to 7 The top 10 important features of RF ranger, GBM, LASSO, LR and GLMnet models based on SHAP values are shown in Table 3.
[0101] Figures 8 to 11 Model local explanation using SHAP values for the example.
[0102] Combining SHAP value ranking and recursive feature elimination strategy, the area under the ROC curve (AUC) of the RF ranger model with different numbers of variables was compared as shown in FIG. 2, and finally the optimal subset containing 10 key features was determined: coronary artery lesions (CALs), C-reactive protein (CRP), fever duration, venous oxygen content (CvO2), oral mucosa changes, direct bilirubin, lactic acid, gender, neutrophil percentage (NEU%), and rash. Figure 12
[0103] Figures 13 to 22 Feature scatter plot for the example, with the y-axis being the SHAP value and the x-axis being the actual feature value. Feature values with SHAP threshold > 0 will push the outcome towards IVIGR.
[0104] Comparison of IVIG resistance prediction scoring systems as shown in FIG. 3.
[0105] Table 3 Comparison of IVIG resistance prediction scoring systems
[0106]
[0107]
[0108] S5. Webpage deployment
[0109] A web-based real-time calculator was constructed using the R language Shiny platform to realize model deployment.
[0110] Users can input the values of the 10 key variables to automatically generate the IVIG resistance risk probability and prompt the high-risk threshold (based on the Youden index, the threshold is 11.6%). If the risk probability exceeds the threshold, the patient belongs to the high-risk group, and the clinician is recommended to implement a more aggressive and personalized stratified management plan for the patient, thereby achieving early intervention, reducing the potential complication risk of IVIG resistance, and improving treatment effect and patient prognosis.
[0111] Figure 23 and Figure 24 Web-based IVIGR risk prediction calculator. Figure 25 Probability distribution of IVIG resistance for the two groups of patients. By inputting the actual 10 feature values of the patient, the probability of IVIGR occurrence is automatically calculated.
[0112] Figure 26 Spearman variable correlation heatmap. After excluding variables with missing values >50%, multiple imputation was performed for missing values of the remaining variables. Spearman correlation analysis was performed for all variables, and variables with a correlation coefficient >0.6 and low correlation with IVIGR were excluded. The final heatmap showed the correlation between the remaining 66 variables.
[0113] The technical key points of the present application are:
[0114] 1. Machine learning model construction technology based on large-scale electronic medical record data
[0115] Using the clinical data of a large number of Kawasaki disease patients in a certain area, a machine learning model with good interpretability was developed to effectively predict the risk of IVIG resistance.
[0116] 2. Design and implementation of online web computing tools
[0117] A convenient online computing platform was built to support clinicians to input patient data in real time, automatically output the probability of IVIG resistance risk, and set the optimal threshold (11.6%) to assist clinical decision-making.
[0118] 3. High-risk patient identification and clinical guidance
[0119] Through the prediction model of the present application, doctors can automatically obtain the IVIG resistance risk probability of patients based on the input of key variables, and use the high-risk threshold of 11.6% (determined based on Youden index) as the judgment standard to identify patients belonging to high risk in a timely manner. When the risk probability of a patient exceeds this threshold, the doctor should pay attention to the fact that the patient may have a higher risk of IVIG resistance. Based on this, doctors can tailor individualized treatment plans, such as optimizing treatment strategies, to achieve precision medicine. This not only improves the scientificity and effectiveness of clinical decision-making, but also significantly reduces the incidence of serious complications such as coronary artery lesions caused by IVIG resistance. Through the setting of risk threshold and timely identification, high-risk patients can receive more proactive and personalized management, greatly improving patient prognosis and treatment effectiveness.
Claims
1. A method for predicting IVIG resistance in KD based on an interpretable machine learning model, characterized in that: The following steps are involved: S1. Data Collection Automated data extraction was performed based on the electronic medical record system, and 20% sample accuracy was verified by two researchers; S2. Data Processing Variables with more than 50% missing values were eliminated, and the remaining missing values were filled using multiple imputation; Categorical variables were coded ordinal, and continuous variables were directly included in the modeling; Through Spearman correlation analysis, highly correlated redundant variables were excluded when the correlation coefficient was >0.6; Finally, 66 variables were retained for modeling; The dataset is divided into 70% as training set and 30% as test set; S3. Model construction Construct and compare the performance of multiple prediction models; Cross-validation evaluates the generalization ability of the model; Adjust model parameters to optimize the model; The RF ranger model with the best performance and clinical interpretability was selected as the final model; S4. Model interpretation module: The SHAP method is introduced to provide global and local visual explanations of the model; Combining SHAP value ranking and recursive feature elimination strategy, the area under the receiver operating characteristic (ROC) curve (AUC) was compared under different numbers of variables, and the optimal subset of 10 key features was finally determined: coronary artery lesions (CALs), C-reactive protein (CRP), fever duration, venous oxygen content (CvO2), oral mucosal changes, direct bilirubin, lactate, gender, neutrophil percentage (NEU%), and rash; S5. Web page deployment Use the R language Shiny platform to build a real-time calculator on the Web to implement model deployment; By inputting the values of 10 key variables, the user can automatically generate the IVIG resistance risk probability and prompt the high-risk threshold; If the risk probability exceeds the threshold, the patient belongs to the high-risk group.
2. The method for predicting IVIG resistance in KD based on an interpretable machine learning model according to claim 1, wherein S1. Data collection covers demographic characteristics, clinical symptoms, laboratory test indicators, and imaging data.
3. The method for predicting IVIG resistance in KD based on an interpretable machine learning model according to claim 1, wherein: S3. The machine learning algorithms used in model construction include random forest (RF), gradient boosting machine (GBM), support vector machine (SVM), logistic regression (LR), generalized linear model (GLMnet) with elastic net regularization, K-nearest neighbor (KNN), decision tree (DT), random forest (RF) based on the ranger algorithm, least absolute shrinkage and selection operator (LASSO), naive Bayes (NB), adaptive boosting (Adaboost), neural network (NNET), and support vector machine (SVM) based on the kernlab algorithm (kernlab).
4. The method for predicting IVIG resistance in KD based on an interpretable machine learning model according to claim 1, wherein: S5. In web page deployment, the high risk threshold is based on the Youden index, and the threshold is 11.6%.
Citation Information
Cited By
Prediction method, system and platform for gamma globulin non-reactive Kawasaki disease
CN122291106A