Prediction method for giant coronary tumor in Kawasaki disease based on interpretable machine learning
By constructing a prediction model for giant coronary artery aneurysms in Kawasaki disease based on interpretable machine learning, the limitations of linear assumptions and the subjectivity of feature selection in existing models are resolved. This enhances the transparency and generalization ability of the model, provides a convenient real-time prediction tool, and enables early identification and stratified management of high-risk children.
Patent Information
- Application Number
- CN202510927400.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-07
- Publication Date
- 2025-10-28
AI Technical Summary
Existing prediction models for giant coronary artery aneurysms in Kawasaki disease suffer from limitations such as linear assumptions, reliance on subjective experience for feature selection, insufficient generalization ability, loss of information from continuous variables, and a lack of convenient real-time prediction tools, resulting in limited predictive performance and insufficient practical application.
We constructed a predictive model based on interpretable machine learning, used the SHAP method for variable interpretation, combined it with recursive feature elimination to screen key variables, and developed an online predictive tool to support doctors in real-time input of key indicators for individualized risk assessment.
It improves model transparency, reduces human bias, preserves continuous variable information, enhances model generalization ability, provides convenient real-time clinical decision support, and enables early identification and stratified management of high-risk children.
Smart Images

Figure CN120853879A_ABST
Abstract
Description
Technical Field
[0001] This invention provides a method for predicting giant coronary artery aneurysms in Kawasaki disease based on interpretable machine learning, belonging to the field of bioinformatics technology. Background Technology
[0002] Kawasaki disease (KD) is an acute febrile systemic vasculitis that primarily affects children under 5 years old, mainly involving small and medium-sized arteries, especially the coronary arteries. It is currently the most common cause of acquired heart disease in children in developed countries, and without timely treatment, the incidence of coronary artery lesions (CALs) can be as high as 25%. With continuous improvements in the diagnosis and treatment of KD, the risk of CALs has significantly decreased, and most coronary artery dilatations can resolve during the acute phase. However, moderate or giant coronary artery aneurysms (CAA) often persist, significantly increasing the risk of serious cardiovascular events in the long term, such as myocardial infarction, vascular rupture, and even sudden death. Therefore, early identification of children at high risk for moderate or giant CAA and timely initiation of intensive interventions (such as hormone therapy) are crucial for improving prognosis.
[0003] Currently, several studies have developed predictive models for CAA, but most focus on overall risk assessment for all types of CAA. Although some studies have preliminarily identified risk factors associated with moderate to giant CAA, predictive tools for moderate to severe CAA are still lacking. There is an urgent need to develop predictive models that can enable early identification of high-risk children with moderate to giant CAA in clinical practice, in order to achieve stratified management of these children.
[0004] With the development of artificial intelligence technology, especially the widespread application of machine learning (ML) methods in medical research, a new path has been provided for establishing accurate, stable, and efficient predictive models for medium and large coronary artery aneurysms. ML models can process multi-dimensional, large-scale clinical data with complex correlation structures, demonstrating significant advantages in early disease prediction and risk stratification.
[0005] Two studies have already constructed predictive models for giant coronary artery aneurysms in Kawasaki disease. Jiang et al. explored the predictive factors for giant CAA in KD patients and established a risk assessment model applicable to the Chinese population (S.Jiang,M.Li,K.Xu,Y.Xie,P.Liang,C.Liu,Q.Su,B.Li,Predictive factors of medium-giant coronaryartery aneurysms in Kawasaki disease,Pediatr Res 95(1)(2024)267-274.). The study employed a retrospective cohort analysis (2015-2020, 1331 patients) and a prospective validation (2021, 193 patients). Multivariate logistic regression (LR) analysis revealed that male sex, older age, longer duration of fever, IVIG resistance, elevated platelet count, and decreased albumin were independent predictors of medium to large CAA. The model performance was evaluated using ROC curves and the Hosmer-Lemeshow test, and the model stability was validated by 10-fold internal cross-validation.
[0006] Zhao et al. conducted a prospective cohort analysis, including 1856 children with Kawasaki disease (KD), and systematically evaluated the predictive value of dynamic changes in inflammatory markers before and after intravenous immunoglobulin (IVIG) treatment for medium-giant coronary artery aneurysms (CAA) (L. Zhao, J. Wu, X. Liu, K. Zhou, Y. Hua, S. Shao, C. Wang, Risk factors for predicting medium-giant coronary artery aneurysms in Kawasaki disease, Immunologic research 73(1)(2025)52.). Multivariate logistic regression analysis revealed that prolonged fever duration, IVIG resistance, cardiac enlargement, and elevated pre-treatment white blood cell count (≥12.05 × 10⁻⁶) were associated with the development of CAA. 9 / L), hypoalbuminemia (≤37.25g / L), and decreased neutrophil percentage change (ΔN≤30.2%) were independent risk factors for predicting medium-sized CAA, and their predictive power was assessed by AUC.
[0007] Current published studies on CAA prediction for large and medium-sized enterprises employ logistic regression statistical methods, which cannot fully capture the complex nonlinear relationships between variables. Furthermore, the vast majority of models lack online real-time prediction tools. Specifically, they suffer from the following technical drawbacks:
[0008] 1. Limitations of linear assumptions: Traditional predictive models (such as logistic regression) are usually based on the assumption of linear relationships, which makes it difficult to effectively capture the complex nonlinear associations between clinical variables.
[0009] 2. Feature selection relies on subjective experience: Traditional methods rely heavily on clinical experience in variable selection, which can easily introduce human bias and thus miss important indicators with potential predictive value.
[0010] 3. Insufficient generalization ability: Due to differences in population and region, the predictive performance of existing models may be limited in Suzhou and Fuzhou.
[0011] 4. Loss of information in continuous variables: Logistic regression models often dichotomize continuous variables, leading to loss of variable information, which may reduce model accuracy and increase the risk of bias.
[0012] 5. Insufficient clinical deployment: The lack of convenient real-time prediction tools limits their practical application in clinical decision-making.
[0013] 6. Insufficient consideration of correlation between variables: Some models do not analyze the potential multicollinearity among the included variables, which may affect the accuracy of the model. Summary of the Invention
[0014] This invention aims to solve several key technical problems in current risk prediction models for moderate and giant coronary artery aneurysms in Kawasaki disease, specifically including:
[0015] (1) Lack of predictive tools: There is currently a lack of dedicated risk prediction models for moderate and above CAA, and a lack of clinically available early identification tools, which limits the timely intervention and management of high-risk children.
[0016] (2) Limited model capabilities: Existing research is mainly based on linear models such as logistic regression, which are difficult to capture the complex nonlinear relationships between variables and thus have limited predictive performance; the binary classification of continuous variables loses some information and affects predictive effectiveness.
[0017] (3) Feature selection is highly subjective: Traditional models rely on human experience in the variable selection process, which can easily introduce bias and may miss potential key variables.
[0018] (4) Insufficient generalization ability: The existing model is limited in use in Suzhou and Fuzhou.
[0019] (5) Insufficient clinical deployment: Existing models lack convenient online real-time prediction tools, making it difficult to apply them quickly in clinical practice and limiting their practical value.
[0020] To this end, this invention constructs a high-performance and interpretable prediction model for medium-to-giant coronary artery aneurysms based on large-scale, multi-center electronic medical record data. It combines the SHAP (SHapley Additive exPlanations) method for variable interpretation to improve model transparency. Furthermore, it is deployed through a visual web tool, supporting doctors to input key indicators in real time and quickly output individualized risk assessment results, thereby enabling early identification and stratified intervention for high-risk children in clinical practice.
[0021] This invention aims to achieve the following objectives:
[0022] (1) Enhance model transparency and clinical interpretability;
[0023] (2) Optimize feature selection to reduce human bias;
[0024] (3) Use unbiased EMRs data to screen and discover potential novel predictive molecules;
[0025] (4) Preserve complete information about continuous variables;
[0026] (5) Provide convenient real-time clinical decision support tools;
[0027] (6) Develop a predictive model applicable to medium and large coronary artery aneurysms.
[0028] A method for predicting giant coronary artery aneurysms in Kawasaki disease based on interpretable machine learning includes the following steps:
[0029] S1. Data source: A retrospective cohort study was conducted to collect cases of hospitalized KD children and exclude cases that did not meet the requirements. The exclusion criteria were as follows: (1) no echocardiography was performed; (2) no initial IVIG treatment was received; (3) IVIG treatment was received before admission; (4) glucocorticoids were used before or at the same time as IVIG treatment; (5) recurrent Kawasaki disease.
[0030] S2. Coronary artery disease assessment and classification;
[0031] Echocardiographic results were collected from patients using standardized PHILIPS and GE equipment to ensure consistent image quality. The diameters of the left main coronary artery (LMCA), left anterior descending artery (LAD), left circumflex artery (LCX), and right coronary artery (RCA) were measured, and the Dallaire Z-score was calculated based on body surface area. According to the AHA guidelines, the severity of coronary arteries (CALs) is graded by Z-score as follows:
[0032] (1) No involvement: Z<2;
[0033] (2) Simple expansion: Z≥2 and <2.5;
[0034] (3) Small aneurysms: Z ≥ 2.5 and < 5;
[0035] (4) Medium-sized aneurysm: Z≥5 and <10, and absolute diameter <8mm;
[0036] (5) Large or giant aneurysms: Z≥10, or absolute diameter≥8mm.
[0037] The maximum Z-score (Zmax) refers to the highest Z-score measured in any one of the four coronary arteries.
[0038] S3. Variable preprocessing;
[0039] (1) Variables were extracted from the electronic medical record system EMRs, covering basic demographic characteristics, clinical manifestations, routine laboratory indicators and echocardiographic parameters;
[0040] (2) Variables with a missing rate of more than 50% were removed, and the remaining missing values were processed by multiple imputation.
[0041] (3) Spearman correlation analysis was used to remove highly correlated variables with a correlation coefficient > 0.6 to avoid the influence of multicollinearity on the modeling results.
[0042] S4. Machine learning model building;
[0043] (1) The medical records were randomly divided into training set and test set in a 7:3 ratio.
[0044] (2) Construct and compare the performance of 11 machine learning algorithms: Random Forest (RF), Gradient Boosting Machine (GBM), Support Vector Machine (SVM), Generalized Linear Model with Elastic Network Regularization (GLMnet), K Nearest Neighbors (KNN), Decision Tree (DT), Random Forest (RF) based on Ranger algorithm, Support Vector Machine (SVM) based on Kernlab algorithm (Kernlab), Naive Bayes (NB), Adaptive Boosting Algorithm (Adaboost), and Neural Network (NNET).
[0045] (3) The model performance evaluation indicators include: area under the curve (AUC), sensitivity, specificity, positive predictive value (PPV), negative predictive value (NPV), accuracy, and F1 score.
[0046] (4) Use cross-validation to evaluate the robustness of the model;
[0047] (5) Models with AUC, Sensitivity, Specificity, Accuracy and F1 all above 0.7 are selected to proceed to the next stage of analysis.
[0048] (6) Select the high-performance Random Forest (RF) model based on the Ranger algorithm and the Support Vector Machine (SVM) model based on the Kernlab algorithm.
[0049] S5. Variable selection;
[0050] (1) The importance of variables is ranked based on the SHAP method, and the impact of different combinations of variable numbers on model performance is evaluated by combining the recursive feature elimination (RFE) method.
[0051] (2) The SVM kernlab model with better performance was selected, and the optimal combination of variables was determined by the DeLong test. Finally, seven core predictors that were highly correlated with medium and large coronary artery aneurysms were identified: C-reactive protein (CRP), eosinophil percentage (EOS%), monocyte percentage (MONO%), neutrophil percentage (NEU%), rash, triglycerides (TG), and time of diagnosis.
[0052] S6. Model Explanation;
[0053] (1) Global and local interpretation of the model: At the global level, the influence of each variable on the model output is analyzed based on the SHAP value; at the local level, an individualized prediction interpretation diagram is output to help doctors understand the specific influence of each variable on the prediction results of a certain child.
[0054] (2) Use other medical records for external independent verification.
[0055] S7. Determine the optimal intervention threshold;
[0056] The Youden index was used to determine the optimal prediction threshold, which served as the boundary for identifying high-risk individuals.
[0057] S8. Develop online forecasting tools;
[0058] An interactive web-based tool is deployed using the Shiny framework in R language. Clinicians can input the above 7 core predictive factors, and the system will automatically output the probability of the patient developing moderate / giant CAA.
[0059] This invention establishes a risk prediction model and online tool for giant coronary artery aneurysms in Kawasaki disease based on an interpretable machine learning model. The optimal-performing SVM kernlab model was selected through multi-model comparison, and combined with the SHAP method and recursive feature elimination, a set of key predictive molecules was identified: CRP, EOS%, MONO%, NEU%, rash, triglycerides, and time to diagnosis. The model was validated using multi-center data to improve its applicability. A risk threshold of 50.2% was calculated, which helps in the early identification and stratified management of high-risk children. A web-based tool was developed to suggest early intervention for patients with a risk higher than 50.2%, improving the prognosis of these children. Attached Figure Description
[0060] Figure 1 is a flow chart of the present invention;
[0061] Figure 2 ROC curves for 11 machine learning models in the example;
[0062] Figure 3 This example demonstrates the ranking of the top 10 key features of an SVM kemlab model based on SHAP values.
[0063] Figure 4 As an example, the top 10 key features of the RF ranger model are ranked based on SHAP values.
[0064] Figure 5 ROC curves for different numbers of variables in the example;
[0065] Figure 6 Sensitivity of different variable number models for the example;
[0066] Figure 7 Model specificity for different numbers of variables in the example;
[0067] Figure 8 This is a global interpretation of the SVM kernlab model using SHAP values in the example.
[0068] Figure 9 This is a scatter plot of the diagnosis time in the example.
[0069] Figure 10 This is a scatter plot of the percentage of mononuclear cells in the example.
[0070] Figure 11 Scatter plot of the rash in the example;
[0071] Figure 12 This is a scatter plot of the percentage of eosinophils in the example.
[0072] Figure 13 The C-reactive protein scatter plot is shown in the example.
[0073] Figure 14 The scatter plot of triglycerides in the example;
[0074] Figure 15 This is a scatter plot of the percentage of neutrophils in the example.
[0075] Figure 16 For example, patient 1 was locally explained using the SHAP value model;
[0076] Figure 17For example, patient 2 was partially explained using the SHAP value model;
[0077] Figure 18 For example, patient 3 was locally explained using the SHAP value model;
[0078] Figure 19 For example, patient 4 was locally explained using the SHAP value model;
[0079] Figure 20 External verification of the model for the embodiment;
[0080] Figure 21 The webpage prediction tool used in this example;
[0081] Figure 22 The following is a probability distribution diagram of giant coronary artery aneurysms in two groups of patients in the example;
[0082] Figure 23 The example is a Spearman variable correlation heatmap. Detailed Implementation
[0083] The specific technical solutions of the present invention will be described with reference to the embodiments.
[0084] like Figure 1 As shown, the method for predicting giant coronary artery aneurysms in Kawasaki disease based on interpretable machine learning includes the following steps:
[0085] S1. Data Source:
[0086] The data in this embodiment comes from the electronic medical record system of the Children's Hospital Affiliated to Soochow University. A retrospective cohort study method was used to collect a total of 2,765 cases of children with KD who were hospitalized between March 2019 and June 2024.
[0087] Exclusion criteria were as follows: (1) no echocardiography was performed; (2) no initial IVIG treatment was received; (3) IVIG treatment was received before admission; (4) glucocorticoids were used before or concurrently with IVIG treatment; (5) recurrent Kawasaki disease. A total of 2,334 cases were ultimately included. In addition, 443 children with KD from Fujian Provincial Hospital were included as an external validation cohort.
[0088] S2. Coronary artery lesion assessment and classification:
[0089] All patients underwent echocardiography performed by two pediatric cardiovascular ultrasound physicians with over five years of experience, and the results were further confirmed by a cardiologist. Standardized PHILIPS and GE equipment were used to ensure consistent image quality. The diameters of the left main coronary artery (LMCA), left anterior descending artery (LAD), left circumflex artery (LCX), and right coronary artery (RCA) were measured, and the Dallaire Z-score was calculated based on body surface area. According to the AHA guidelines, the severity of coronary arteries (CALs) is graded by Z-score as follows:
[0090] (1) No involvement: Z<2;
[0091] (2) Simple expansion: Z≥2 and <2.5;
[0092] (3) Small aneurysms: Z ≥ 2.5 and < 5;
[0093] (4) Medium-sized aneurysm: Z≥5 and <10, and absolute diameter <8mm;
[0094] (5) Large or giant aneurysms: Z≥10, or absolute diameter≥8mm.
[0095] The maximum Z-score (Zmax) refers to the highest Z-score measured in any one of the four coronary arteries.
[0096] S3. Variable preprocessing:
[0097] (1) Variables were extracted through electronic medical records (EMRs) systems, covering basic demographic characteristics, clinical manifestations, routine laboratory indicators and echocardiographic parameters;
[0098] (2) Variables with a missing rate of more than 50% were removed, and the remaining missing values were processed by multiple imputation.
[0099] (3) Spearman correlation analysis was used to remove highly correlated variables with a correlation coefficient > 0.6 to avoid the influence of multicollinearity on the modeling results. Figure 23 Example: Spearman variable correlation heatmap;
[0100] S4. Machine Learning Model Building:
[0101] (1) 2334 patients in Suzhou were randomly divided into training set and test set in a 7:3 ratio.
[0102] (2) Construct and compare the performance of 11 machine learning algorithms: Random Forest (RF), Gradient Boosting Machine (GBM), Support Vector Machine (SVM), Generalized Linear Model with Elastic Net Regularization (GLMnet), K-Nearest Neighbor (KNN), Decision Tree (DT), Random Forest with Ranger Algorithm (RFranger), Support Vector Machine with Kernlab Algorithm (SVM kernlab), and Naive Bayes. Bayes (NB), Adaptive Boosting (Adaboost), and Neural Network (NNET). Figure 2 ROC curves for 11 machine learning models used in the examples.
[0103] (3) Model performance evaluation metrics include: Area Under the Receiver Operating Characteristic Curve (AUC), Sensitivity, Specificity, Positive Predictive Value (PPV), Negative Predictive Value (NPV), Accuracy, and F1 Score. The performance of the 11 machine learning models in predicting medium-to-giant coronary artery aneurysms is shown in Table 1.
[0104] Table 1. Performance of 11 machine learning models in predicting medium-to-giant coronary artery aneurysms.
[0105] Model AUC Sensitivity Specificity Negative predictive value Positive predictive value Accuracy F1 value RF 0.781 0.839 0.688 0.991 0.091 0.835 0.909 GBM 0.773 0.628 0.813 0.993 0.049 0.632 0.770 SVM 0.760 0.946 0.313 0.983 0.119 0.931 0.964 GLMnet 0.767 0.792 0.688 0.991 0.072 0.790 0.880 KNN 0.608 0.170 0.938 0.991 0.026 0.187 0.290 DT 0.528 0.057 0.999 0.999 0.024 0.079 0.108 RF ranger 0.781 0.770 0.750 0.992 0.071 0.770 0.867 SVM kernlab 0.774 0.799 0.750 0.993 0.081 0.798 0.886 NB 0.727 0.818 0.688 0.991 0.081 0.815 0.897 Adaboost 0.767 0.641 0.813 0.993 0.050 0.645 0.779 NNET 0.627 0.102 0.938 0.986 0.024 0.122 0.186
[0106] (4) Use cross-validation to evaluate the robustness of the model;
[0107] (5) Models with AUC, Sensitivity, Specificity, Accuracy and F1 all above 0.7 are selected to proceed to the next stage of analysis.
[0108] (6) Select the high-performance Random Forest (RF) model based on the Ranger algorithm and the Support Vector Machine (SVM) model based on the Kernlab algorithm.
[0109] S5. Variable Selection:
[0110] (1) The importance of variables is ranked based on the SHAP (Shapley Additive Explanations) method, and the impact of different combinations of variable numbers on model performance is evaluated by combining the recursive feature elimination (RFE) method. Figure 3 This example demonstrates the ranking of the top 10 key features of an SVM kemlab model based on SHAP values. Figure 4 As an example, the top 10 important features of the RF ranger model are ranked based on SHAP values. Table 2 shows the performance of the SVM Kernlab model with different numbers of features. Table 3 shows the performance of the RF ranger model with different numbers of features. Figure 5 ROC curves for different numbers of variables in the example; Figure 6 Sensitivity of different variable number models for the example; Figure 7 The model specificity for different numbers of variables in the example.
[0111] Table 2 shows the performance of the SVM kernlab model based on different numbers of features.
[0112]
[0113]
[0114] Table 3 shows the performance of the RFranger model based on different numbers of features.
[0115]
[0116]
[0117] (2) The SVM kernlab model with better performance was selected, and the optimal combination of variables was determined by the DeLong test. Finally, seven core predictors that were highly correlated with medium and large coronary artery aneurysms were identified: C-reactive protein (CRP), eosinophil percentage (EOS%), monocyte percentage (MONO%), neutrophil percentage (NEU%), rash, triglyceride (TG), and time to diagnosis.
[0118] S6. Model Explanation:
[0119] (1) Global and local interpretation of the model: At the global level, the influence of each variable on the model output is analyzed based on the SHAP value; at the local level, an individualized prediction interpretation diagram is output to help doctors understand the specific influence of each variable on the prediction results of a certain child. Figure 8 For a global interpretation of the SVM kernlab model using SHAP values, Figures 9-15 This is a scatter plot of the characteristics of each variable. Figures 16-19 This provides a local interpretation of the model using SHAP values.
[0120] (2) External independent validation was conducted using 443 children from Fujian Provincial Hospital, such as... Figure 20 .
[0121] S7. Determine the optimal intervention threshold:
[0122] The optimal prediction threshold of 50.2% was determined using the Youden index and used as the identification line for high-risk individuals.
[0123] S8. Develop online forecasting tools:
[0124] like Figure 21 An interactive web tool based on the Shiny framework in R language is deployed. Clinicians can input the above 7 core indicators, and the system will automatically output the probability of the patient developing moderate / giant CAA. If the predicted value exceeds 50.2%, early intervention is recommended. Figure 22 The following is a probability distribution diagram of giant coronary artery aneurysms in two groups of patients in the example.
Claims
1. A method for predicting giant coronary artery aneurysms in Kawasaki disease based on interpretable machine learning, characterized in that, Includes the following steps: S1. Data source: A retrospective cohort study was conducted to collect cases of hospitalized children with koiloma syndrome (KD), excluding cases that did not meet the requirements. S2. Coronary artery disease assessment and classification; Collect patients' echocardiographic examination results using standardized PHILIPS and GE equipment to ensure consistent image quality; measure the internal diameters of the left main coronary artery (LMCA), left anterior descending artery (LAD), left circumflex artery (LCX), and right coronary artery (RCA), and calculate the Dallaire Z score based on body surface area; S3. Variable preprocessing; (1) Variables were extracted from the electronic medical record system EMRs, covering basic demographic characteristics, clinical manifestations, routine laboratory indicators and echocardiographic parameters; (2) Variables with a missing rate of more than 50% were removed, and the remaining missing values were processed by multiple imputation. (3) Spearman correlation analysis was used to remove highly correlated variables with a correlation coefficient > 0.6 to avoid the influence of multicollinearity on the modeling results; S4. Machine learning model building; (1) The medical records were randomly divided into a training set and a test set at a ratio of 7:3; (2) Construct and compare the performance of various machine learning algorithms; (3) Model performance evaluation metrics include: area under the curve (AUC), sensitivity, specificity, positive predictive value (PPV), negative predictive value (NPV), accuracy, and F1 score; (4) Use cross-validation to evaluate the robustness of the model; (5) Models with AUC, Sensitivity, Specificity, Accuracy and F1 all above 0.7 are selected to proceed to the next stage of analysis; (6) Select the high-performance Random Forest (RFranger) based on the Ranger algorithm and the Support Vector Machine (SVMkernlab) based on the Kernlab algorithm; S5. Variable selection; (1) The importance of variables is ranked based on the SHAP method, and the impact of different combinations of variable numbers on model performance is evaluated by combining the recursive feature elimination (RFE) method. (2) Select the better-performing SVM kernlab model and determine the optimal variable combination through the DeLong test, and finally determine the core predictive factors that are highly correlated with medium and large coronary artery aneurysms; S6. Model Explanation; (1) Perform global and local interpretations of the model; (2) External independent verification using other medical records; S7. Determine the optimal intervention threshold; The Youden index was used to determine the optimal prediction threshold, which served as the identification boundary for high-risk individuals. S8. Develop online forecasting tools; An interactive web-based tool is deployed using the Shiny framework in R language. Clinicians can input core predictive factors, and the system will automatically output the probability of the patient developing moderate / giant CAA.
2. The method for predicting giant coronary artery aneurysms in Kawasaki disease based on interpretable machine learning according to claim 1, characterized in that, The exclusion criteria for S1 are as follows: (1) no echocardiography was performed; (2) no initial IVIG treatment was received; (3) IVIG treatment was received before admission; (4) glucocorticoids were used before or during IVIG treatment; (5) recurrent Kawasaki disease.
3. The method for predicting giant coronary artery aneurysms in Kawasaki disease based on interpretable machine learning according to claim 1, characterized in that, According to the AHA guidelines, the severity of CALs in S2 is graded by Z-score as follows: (1) No involvement: Z<2; (2) Simple expansion: Z≥2 and <2.5; (3) Small aneurysms: Z ≥ 2.5 and < 5; (4) Medium-sized aneurysm: Z≥5 and <10, and absolute diameter <8mm; (5) Large or giant aneurysms: Z≥10, or absolute diameter≥8mm; The maximum Z-score (Zmax) refers to the highest Z-score measured in any one of the four coronary arteries.
4. The method for predicting giant coronary artery aneurysms in Kawasaki disease based on interpretable machine learning according to claim 1, characterized in that, Machine learning models in S4 include: Random Forest, Gradient Boosting Machine, Support Vector Machine, Generalized Linear Model with Elastic Network Regularization, K-Nearest Neighbors Algorithm, Decision Tree, Random Forest based on Ranger Algorithm, Support Vector Machine based on Kernlab Algorithm, Naive Bayes, Adaptive Boosting Algorithm, and Neural Network.
5. The method for predicting giant coronary artery aneurysms in Kawasaki disease based on interpretable machine learning according to claim 1, characterized in that, The core predictive factors in S5 include: C-reactive protein (CRP), eosinophil percentage (EOS%), monocyte percentage (MONO%), neutrophil percentage (NEU%), rash, triglycerides (TG), and time to diagnosis.
6. The method for predicting giant coronary artery aneurysms in Kawasaki disease based on interpretable machine learning according to claim 1, characterized in that, In S6, the global and local interpretations of the model are as follows: at the global level, the influence of each variable on the model output is analyzed based on the SHAP value; at the local level, an individualized prediction interpretation graph is output to help doctors understand the specific impact of each variable on the prediction results of a particular child.
Citation Information
Cited By
Whole blood donation adverse reaction risk prediction method and system
CN121054268A
A method and system for predicting the risk of adverse reactions of whole blood donation
CN121054268B