Muscle fatty degeneration risk assessment method based on machine learning

Through the AFIE algorithm and IS-SHAP method, the problem of low feature screening accuracy was solved, the prediction accuracy and interpretation transparency of the myosteatosis risk assessment model were improved, and more accurate feature screening and SHAP value interpretation were achieved.

CN120600307APending Publication Date: 2025-09-05THE FIRST PEOPLES HOSPITAL OF CHANGZHOU
View PDF 0 Cites 2 Cited by

Patent Information

Application Number
CN202510695735.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-28
Publication Date
2025-09-05

AI Technical Summary

Technical Problem

When the existing technology uses the Boruta algorithm and Lasso regression to screen relevant features in data samples, the feature screening accuracy is low, resulting in inaccurate prediction results of the muscle fatty degeneration risk assessment model, affecting the accuracy of SHAP value interpretation.

Method used

The adaptive feature importance evaluation (AFIE) algorithm is adopted, combined with the Boruta algorithm and Lasso regression. The feature importance score is calculated iteratively, the interaction between features is considered, and the feature selection threshold is dynamically adjusted. At the same time, the interaction-sensitive SHAP value calculation method (IS-SHAP) is introduced to expand the traditional SHAP framework and capture the interaction effect between features.

Benefits of technology

The accuracy of feature screening and the predictive performance of the model are improved, the global and local interpretation transparency of the model is enhanced, the stability and consistency of feature importance ranking are significantly improved, and feature interaction effects that are ignored by traditional methods are identified.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120600307A_ABST
    Figure CN120600307A_ABST
Patent Text Reader

Abstract

The invention provides a muscle fatty degeneration risk assessment method based on machine learning, and relates to the technical field of disease risk assessment. The method comprises the following steps: acquiring a data sample of an HD patient; carrying out feature screening on related features in the data sample according to a Boruta algorithm and Lasso regression based on interaction importance among the features, and generating an important feature set; based on the important feature set, constructing and training a plurality of machine learning prediction models by using the corresponding data samples, and determining an optimal model; inputting data of an actual HD patient into the optimal model to generate a prediction result; performing global explanation and local explanation on the prediction result based on the SHAP value; and generating a risk assessment result according to the global explanation and the local explanation of the prediction result. By adopting the method, the features can be accurately screened, and the prediction result of the model can be explained by accurately utilizing the SHAP value.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the technical field of disease risk assessment, and in particular to a method for assessing the risk of myosteatosis based on machine learning. Background Art

[0002] End-stage renal disease (ESRD), the terminal stage of various chronic kidney diseases, has been recognized as a major global public health issue. In 2019, the estimated global prevalence of chronic kidney disease was 13.4%, with approximately 4.902 million to 7.083 million patients entering the ESRD stage and relying on hemodialysis (HD) to prolong their lives.

[0003] Myosteatosis, a manifestation of abnormal distribution of adipose tissue between and within muscle cells, reflects poor muscle quality and has been considered a strong predictor of poor disease prognosis. Metabolic factors such as inflammation, obesity, insulin resistance, and dyslipidemia increase HD patients' susceptibility to myosteatosis, leading to worse clinical outcomes.

[0004] Although muscle biopsy is the gold standard for diagnosing myosteatosis, its invasiveness, clinical applicability, and cost-effectiveness limit its implementation in large clinical populations. Age, obesity, diabetes, and cardiovascular disease are currently recognized risk factors for myosteatosis. However, other potential and effective clinical indicators for assessing myosteatosis remain to be explored.

[0005] Electronic medical records not only contain rich and comprehensive demographic, anthropometric, laboratory, and imaging data, but can also be quickly and accurately accessed, making them an excellent tool for clinical disease screening. However, synergistic effects inevitably exist between clinical data, and complex nonlinear relationships may exist between characteristic factors and diseases, posing a significant challenge to the application of traditional linear statistical methods.

[0006] Machine learning technology is not only highly inclusive of data and offers high prediction accuracy, but it can also provide feature importance rankings and Shapley Additive Explanation (SHAP) values ​​to directly interpret the model. Currently, when using machine learning to predict the risk of myosteatosis, the existing technology typically uses the Boruta algorithm and Lasso regression to screen relevant features in the data sample during the feature selection phase and directly takes the intersection to obtain important features. The relevant data is then input into a pre-set prediction model for prediction, and the prediction results are finally interpreted using SHAP values.

[0007] In the process of using the above method, the inventors discovered the following problem: when using the Boruta algorithm and Lasso regression to screen the relevant features in the data sample and directly take the intersection to obtain the important features, the feature screening has the problem of low accuracy, which makes the model's prediction results inaccurate and also affects the accuracy of the subsequent SHAP value interpretation. Summary of the Invention

[0008] Based on this, it is necessary to provide a risk assessment method for myosteatosis based on machine learning to address the above technical issues, which can accurately screen features and accurately use SHAP values ​​to interpret the prediction results of the model.

[0009] The present application provides a method for assessing the risk of myosteatosis based on machine learning, which is characterized by comprising:

[0010] Obtain data samples from HD patients;

[0011] Based on the interaction importance between features, the relevant features in the data sample are screened using the Boruta algorithm and Lasso regression to generate an important feature set;

[0012] Based on the important feature set, use the corresponding data samples to build and train multiple machine learning prediction models and determine the optimal model;

[0013] The data of actual HD patients were input into the optimal model to generate prediction results;

[0014] Perform global and local interpretations of prediction results based on SHAP values;

[0015] Generate risk assessment results based on global and local interpretations of the prediction results.

[0016] In one embodiment, the data sample includes demographic characteristics, anthropometric characteristics, clinical characteristics, comorbidities, concomitant medications, laboratory tests and chest CT images; wherein, demographic characteristics include age and gender; anthropometric characteristics include height, weight and body mass index; clinical characteristics include dialysis age, smoking history, and drinking history; comorbidities include diabetes, hypertension, and cardiovascular disease; concomitant medications include lipid-lowering drugs, iron supplements, erythropoietin and compound α-keto acids; laboratory tests include white blood cell count, hemoglobin, platelet count, serum albumin, neutrophil count, lymphocyte count, monocyte count, fasting blood glucose, total cholesterol, triglycerides, low-density lipoprotein and high-density lipoprotein; and the chest CT image must include the first lumbar transverse process level.

[0017] In one embodiment, feature screening of relevant features in a data sample is performed based on the interaction importance between features using the Boruta algorithm and Lasso regression, including:

[0018] Based on the Boruta algorithm, the initial feature importance set is obtained according to the relevant features in the data sample, and based on the Lasso regression, the sparse feature set is obtained according to the relevant features in the data sample;

[0019] Calculate the interaction importance between features based on the preset feature interaction importance matrix;

[0020] Calculate the comprehensive score of features based on the interaction importance between features;

[0021] When the comprehensive score is greater than the threshold, the corresponding feature is selected into the important feature set.

[0022] In one embodiment, the calculation formula of the feature interaction importance matrix is ​​as follows:

[0023]

[0024] Among them, I jk is the interaction importance between feature j and feature k; f(x i ) is the model for sample x i The prediction function of It is the second-order partial derivative of the prediction function with respect to feature j and feature k, which is used to measure the strength of the interaction between the two features.

[0025] In one embodiment, the threshold is a dynamic threshold, and the formula is as follows:

[0026] θ=μ S +α·σ S

[0027] Among them, θ is the dynamic threshold; μ S is the mean of the comprehensive scores of all features; σ S is the standard deviation of the comprehensive score of all features; α is an adjustment parameter with a default value of 0.8.

[0028] In one embodiment, the plurality of machine learning prediction models include naive Bayes, nearest neighbor, decision tree, random forest, support vector machine, extreme gradient boosting, and logistic regression.

[0029] In one embodiment, the prediction results are globally and locally interpreted based on the SHAP value, including:

[0030] Calculate the total SHAP value of each feature and the interaction SHAP value between features;

[0031] Generate the main SHAP value of each feature based on the total SHAP value and the interaction SHAP value;

[0032] The relative strength of the interaction between features is evaluated based on the interaction SHAP value and the main SHAP value;

[0033] The corresponding visualization graphs are constructed using the total SHAP value, interactive SHAP value, main SHAP value and the relative strength of the interaction between features, and the visualization graphs are used to provide global and local explanations for the prediction results.

[0034] In one embodiment, the formula for calculating the interaction SHAP value between features is as follows:

[0035] Δ ij (S) = f(x S∪{i,j} )-f(x S∪{i} )-f(x S∪{j} )+f(x S )

[0036]

[0037] Among them, Δ i j(S) is the interactive difference between feature i and feature j in feature subset S; f(x S ∪{i,j}) is the model prediction value corresponding to the input containing feature subset S, feature i and feature j; f(x S ∪{i}) is the model prediction value corresponding to the input containing the feature subset S and feature i; f(x S ∪{j}) is the model prediction value corresponding to the input containing feature subset S and feature j; f(x S ) is the model prediction value corresponding to the input containing only the feature subset S; ∪ is the set union operator; φ ij (f,x) is the SHAP value of the interaction between feature i and feature j; f is the prediction function of the machine learning model; x is the input sample to be explained; S is the feature subset that does not contain feature i and feature j; N is the set of all features; i and j are feature indices; |S| is the size of feature subset S, that is, the number of features contained; |N| is the total number of features; |S|! is the factorial of |S|; (|N|-|S|-2)! is the factorial of (|N|-|S|-2); (|N|-1)! is the factorial of (|N|-1); Δ ij (S) is the interaction difference between feature i and feature j under feature subset S.

[0038] In one embodiment, the total SHAP value of each feature is divided into main effects and interaction effects; the formula for generating the main SHAP value of each feature based on the interaction SHAP value is as follows:

[0039]

[0040] in, is the total SHAP value of feature i; is the main effect SHAP value of feature i; is the average distribution coefficient of the interaction effect; φ ij (f,x) is the interaction SHAP value of feature i and feature j.

[0041] In one embodiment, the formula for evaluating the relative strength of the interaction between features based on the interaction SHAP value and the main SHAP value is as follows:

[0042]

[0043] Among them, R ij It is the relative strength of the interaction between features, and its value range is [0,1]. The larger the value, the stronger the interaction effect and the weaker the main effect.

[0044] This application adopts the above-mentioned myosteatosis risk assessment method based on machine learning, which has the following beneficial effects:

[0045] 1. This paper introduces an adaptive feature importance evaluation algorithm. The core idea of ​​this algorithm is to calculate the feature importance score in an iterative manner, consider the interaction between features, and dynamically adjust the feature selection threshold according to the data distribution. It combines the advantages of the traditional Boruta algorithm and Lasso regression, and enhances the accuracy of feature screening through feature interaction weights and dynamic threshold adjustment mechanism.

[0046] 2. An interaction-sensitive SHAP value calculation method is proposed, which expands the traditional SHAP framework to better capture the interaction effects between features, provide more accurate global and local explanations of the model, and enhance model transparency and credibility. BRIEF DESCRIPTION OF THE DRAWINGS

[0047] Figure 1 Flowchart of a method for assessing risk of myosteatosis based on machine learning in one embodiment;

[0048] Figure 2 A Myosteatosis diagram based on L1 level measurement in one embodiment;

[0049] Figure 3 A feature screening diagram based on a training set in one embodiment;

[0050] Figure 4 AUC comparison chart of the AFIE algorithm and the single-use method in one embodiment;

[0051] Figure 5 A comparison chart of the number of screening features in one embodiment;

[0052] Figure 6 A comparison chart of key performance indicators of the AFIE algorithm and other methods in one embodiment;

[0053] Figure 7 A comparison of the areas under the curves of seven different prediction models in one embodiment;

[0054] Figure 8 is a global explanation graph based on the SHAP method in one embodiment;

[0055] Figure 9 is a local explanation diagram based on the SHAP method in one embodiment;

[0056] Figure 10 A comparison diagram of IS-SHAP and traditional SHAP in one embodiment;

[0057] Figure 11 A visualization diagram of the IS-SHAP interaction effect in one embodiment;

[0058] Figure 12 A schematic diagram of a computational tool for Myosteatosis risk stratification on the WeB side in one embodiment. DETAILED DESCRIPTION

[0059] In order to make the purpose, technical solutions and advantages of this application more clear, the following further describes this application in detail with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain this application and are not intended to limit this application.

[0060] Reference Figure 1 , the present application provides a method for assessing the risk of myosteatosis based on machine learning, characterized by including;

[0061] S100, obtaining data samples of HD patients;

[0062] In one embodiment, the data sample includes demographic characteristics, anthropometric characteristics, clinical characteristics, comorbidities, concomitant medications, laboratory tests and chest CT images; wherein, demographic characteristics include age and gender; anthropometric characteristics include height, weight and body mass index; clinical characteristics include dialysis age, smoking history, and drinking history; comorbidities include diabetes, hypertension, and cardiovascular disease; concomitant medications include lipid-lowering drugs, iron supplements, erythropoietin and compound α-keto acids; laboratory tests include white blood cell count, hemoglobin, platelet count, serum albumin, neutrophil count, lymphocyte count, monocyte count, fasting blood glucose, total cholesterol, triglycerides, low-density lipoprotein and high-density lipoprotein; and the chest CT image must include the first lumbar transverse process level.

[0063] In addition, the following derived indicators were calculated based on the above data samples, including the neutrophil-lymphocyte ratio (NLR), platelet-lymphocyte ratio (PLR), lymphocyte-monocyte ratio (LMR), and glucose-lymphocyte ratio (GLR). NLR is the ratio of neutrophil count to lymphocyte count; PLR is the ratio of platelet count to lymphocyte count; LMR is the ratio of lymphocyte count to monocyte count; and GLR is the ratio of fasting blood glucose to lymphocyte count.

[0064] Finally, after the data samples of HD patients were obtained, the data from one of the independent centers were used as the external validation set, and the data from the remaining centers were randomly divided into the training set and internal validation set in a ratio of 8:2.

[0065] Preferably, data collection uses a multicenter stratified sampling strategy to ensure balanced data distribution across centers and avoid data bias. For missing data, multiple imputation methods are used to construct multiple complete data sets and combine the results to reduce estimation bias and improve the accuracy of statistical inference.

[0066] S200, based on the interaction importance between features, the relevant features in the data sample are screened according to the Boruta algorithm and Lasso regression to generate an important feature set.

[0067] Reference Figures 2 to 6 In one embodiment, step S200 includes:

[0068] Step S210: obtaining an initial feature importance set based on relevant features in the data sample based on the Boruta algorithm, and obtaining a sparse feature set based on relevant features in the data sample based on Lasso regression.

[0069] Based on the Boruta algorithm, by creating "shadow features" (random arrangement of original features) and building a random forest model, the importance scores of real features and shadow features are compared to identify significantly important features.

[0070] Lasso regression is used to obtain a sparse feature set. By introducing the L1 regularization term and compressing the regression coefficient, feature selection is achieved. The specific formula is as follows:

[0071]

[0072] Where N is the number of samples; y i is the target variable of the i-th sample; x ij is the jth eigenvalue of the i-th sample; β0 is the intercept term; β j is the coefficient of the jth feature; λ is the regularization parameter; j is the feature index, and p is the total number of features.

[0073] Step S220: Calculate the interaction importance between features based on a preset feature interaction importance matrix.

[0074] Specifically, the feature interaction importance matrix I introduced in this application is calculated as follows:

[0075]

[0076] Among them, I jk represents the interaction importance between feature j and feature k; f(x i ) represents the model's response to sample x i The prediction function of Represents the second-order partial derivative of the prediction function with respect to feature j and feature k; it is used to measure the strength of the interaction between two features.

[0077] Step S230: Calculate the comprehensive score of the features based on the interaction importance between the features.

[0078] Specifically, the formula for calculating the comprehensive score of the feature is as follows:

[0079]

[0080] Among them, S j is the comprehensive score of feature j; B j is the importance score obtained by the Boruta algorithm; L j is the absolute value of the characteristic coefficient of Lasso regression; W B 、W L and W I are the Boruta score, Lasso coefficient, and interaction importance weights, respectively. Their initial values ​​are set to 0.4, 0.4, and 0.2, respectively, and can be adjusted according to data characteristics.

[0081] Step S240: When the comprehensive score is greater than the threshold, the corresponding feature is selected into the important feature set.

[0082] In one embodiment, the threshold is a dynamic threshold, and the formula is as follows:

[0083] θ=μ S +α·σ S

[0084] Among them, θ is the dynamic threshold; μ S is the mean of the comprehensive scores of all features; σ S is the standard deviation of the comprehensive score of all features; α is an adjustment parameter with a default value of 0.8.

[0085] This paper introduces the Adaptive Feature Importance Evaluation (AFIE) algorithm. The core idea of ​​this algorithm is to calculate the feature importance score in an iterative manner, consider the interaction between features, and dynamically adjust the feature selection threshold according to the data distribution. It combines the advantages of the traditional Boruta algorithm and Lasso regression, and enhances the accuracy of feature screening through feature interaction weights and dynamic threshold adjustment mechanism.

[0086] By using the AFIE algorithm of the present invention, 7 key features were screened out from the initial 24 features (age, gender, BMI, dialysis age, smoking history, drinking history, hypertension, diabetes, cardiovascular disease, lipid-regulating drugs, iron, erythropoietin, compound α-keto acid, white blood cells, hemoglobin, serum albumin, NLR, PLR, LMR, GLR, total cholesterol, triglycerides, low-density lipoprotein and high-density lipoprotein): age, body mass index, diabetes, cardiovascular disease, serum albumin, glucose-lymphocyte ratio and high-density lipoprotein. Experiments have shown that compared with using the Boruta algorithm or Lasso regression alone, the feature set screened by the AFIE algorithm has significant improvements in predictive performance and interpretability, with the model AUC increased by 3.5% and the number of features reduced by 30%. Figures 9 to 11 The AFIE algorithm achieved an AUC of 0.800 on the external validation set, a 3.5% improvement over the simple intersection method (from 0.765 to 0.800), and 4.6% and 3.8% improvements over the Boruta algorithm and Lasso regression used alone, respectively. The AFIE algorithm ultimately selected seven key features, a 30% reduction compared to the 10 features of the simple intersection method, a 53% reduction compared to the 15 features of the Boruta algorithm, and a 36% reduction compared to the 11 features of the Lasso regression method. The AFIE algorithm outperformed other methods in multiple metrics, including accuracy, sensitivity, specificity, F1 score, and Kappa coefficient, demonstrating comprehensive performance advantages.

[0087] S300, based on the important feature set, uses the corresponding data samples to build and train multiple machine learning prediction models and determine the optimal model.

[0088] Reference Figure 7 Based on the important feature sets and corresponding training sets, we built and trained a variety of machine learning prediction models, including Naive Bayes, Nearest Neighbor, Decision Tree, Random Forest, Support Vector Machine, Extreme Gradient Boosting, and Logistic Regression. For each model, we optimized its parameters on the training set, and then evaluated its performance on internal and external validation sets. Evaluation metrics included area under the curve, accuracy, sensitivity, specificity, positive predictive value, negative predictive value, F1 score, and Kappa score.

[0089] In this application, after multi-dimensional comparison, the XGBoost model was found to have the best overall performance on both the internal and external validation sets. In particular, on the external validation set, the XGBoost model achieved an AUC of 0.789, an accuracy of 0.789, a sensitivity of 0.863, and a specificity of 0.776, all exceeding those of other models. Therefore, the extreme gradient boosting model was determined to be the optimal prediction model.

[0090] Preferably, the extreme gradient boosting model uses the following parameter settings: maximum tree depth of 5, learning rate of 0.1, subsample ratio of 0.8, feature sampling ratio of 0.7, L1 regularization coefficient of 0.01, L2 regularization coefficient of 0.1, and minimum number of child node samples of 20. These parameters are determined by Bayesian optimization method to ensure the generalization ability of the model while avoiding overfitting.

[0091] S400 , inputting data of actual HD patients into the optimal model to generate prediction results.

[0092] S500, performing global and local interpretations on the prediction results based on the SHAP value.

[0093] Reference Figures 8 to 11 ,This paper proposes an interaction-sensitive SHAP value calculation method, which expands the traditional SHAP framework, better captures the interaction effects between features, and provides a,more accurate model interpretation.

[0094] Specifically, the total SHAP value of each feature and the interaction SHAP value between features are calculated; the main SHAP value of each feature is generated based on the total SHAP value and the interaction SHAP value; the relative strength of the interaction between features is evaluated based on the interaction SHAP value and the main SHAP value; the corresponding visualization graph is constructed using the total SHAP value, interaction SHAP value, main SHAP value and the relative strength of the interaction between features, and the visualization graph is used to provide global and local explanations for the prediction results.

[0095] like Figure 12 Compared with traditional SHAP, IS-SHAP has the following significant advantages: (1) Interpretation consistency is improved by 15%: The consistency of IS-SHAP's interpretation results on different data subsets is significantly improved; (2) Feature ranking stability is improved by 21%: The stability of feature importance ranking in different scenarios is greatly improved; (3) Interaction effect recognition rate is improved by 53%: IS-SHAP can identify feature interaction effects that traditional SHAP ignores, such as the important interactive relationship between age and cardiovascular disease, serum albumin and glucose lymphocyte ratio; (4) Figure 6As shown in Figure 3, IS-SHAP decomposes feature contributions into main effects and interaction effects, making the model interpretation more accurate and comprehensive; (5) Provides an interaction strength matrix: It clearly quantifies the interaction strength between each feature pair and finds that age×CVD (0.38) and ALB×GLR (0.32) have strong interaction effects; (6) Provides richer visualization methods: It provides interaction-sensitive SHAP waterfall charts, feature interaction strength matrices and other visualization methods to make the model interpretation more intuitive.

[0096] The IS-SHAP method revealed several key clinical findings, including that the increased risk of myosteatosis in elderly patients with concurrent cardiovascular disease is significantly greater than the sum of the independent effects of the two factors, and that high albumin levels can partially mitigate the risk associated with elevated GLR. These findings provide new insights into the pathological mechanisms of myosteatosis.

[0097] Specifically, the traditional SHAP method is based on the concept of Shapley value in cooperative game theory to calculate the contribution of each feature to the model output. The total SHAP value is calculated as follows:

[0098]

[0099] Among them, φ i (f,x) is the total SHAP value of feature i; N is the set of all features; S is the feature subset that does not contain feature i; f(x S∪{i} ) represents the predicted value of subset S and feature i; f(x S ) indicates that only the predicted values ​​of subset S are included.

[0100] The IS-SHAP method introduces two key innovations:

[0101] (1) Explicit modeling of interaction effects:

[0102]

[0103] Among them, φ ij (f,x) is the SHAP value of the interaction between features i and j; f is the prediction function of the machine learning model; x is the input sample to be explained; S is the feature subset that does not contain features i and j; N is the set of all features; i and j are feature indices; |S| is the size of feature subset S, that is, the number of features contained; |N| is the total number of features; |S|! is the factorial of |S|; (|N| - |S| - 2)! is the factorial of (|N| - |S| - 2); (|N| - 1)! is the factorial of (|N| - 1).

[0104] Among them, Δ ij (S) is the interaction difference, which is calculated as follows:

[0105] Δ ij (S) = f(x S∪{i,j} )-f(x S∪{i} )-f(x s∪{j} )+f(x s )

[0106] Among them, Δ i j(S) is the interactive difference between feature i and feature j in feature subset S; f(x S ∪{i,j}) is the model prediction value corresponding to the input containing feature subset S, feature i and feature j; f(x S ∪{i}) is the model prediction value corresponding to the input containing the feature subset S and feature i; f(x S ∪{j}) is the model prediction value corresponding to the input containing feature subset S and feature j; f(x S ) is the model prediction value corresponding to the input containing only the feature subset S; ∪ is the set union operator.

[0107] (2) Multi-level attribution decomposition: IS-SHAP decomposes the feature contribution into two parts: main effect and interaction effect:

[0108]

[0109] in, is the total SHAP value of feature i, is the main effect SHAP value, φ ij (f,x) is the SHAP value of the interaction with feature j. Ensure that the interaction effect is evenly distributed between the two features involved in the interaction.

[0110] Using the IS-SHAP method, three visualization methods are provided:

[0111] (1) Globally explained SHAP summary bar chart: shows the average contribution of each feature to the model and distinguishes the proportion of main effects and interaction effects.

[0112] (2) SHAP summary dot plot: displays the eigenvalues ​​of each sample and their SHAP value distribution, distinguishing high and low eigenvalues ​​by color (yellow indicates high values, purple indicates low values).

[0113] (3) Local interpretation SHAP waterfall plot: For individual patients, it shows the specific contribution of each feature to the prediction results, including main effects and main interaction effects.

[0114] Using the IS-SHAP method, significant interactions were found between age and cardiovascular disease, as well as serum albumin and glucose-lymphocyte ratio. This finding provides a new perspective for clinical understanding of risk factors for myosteatosis. Experiments showed that compared with traditional SHAP, IS-SHAP improved the consistency of interpretation by 15% and the stability of feature importance ranking by 21%.

[0115] S600: Generate risk assessment results based on the global interpretation and local interpretation of the prediction results.

[0116] Based on the output probability of the optimal prediction model, patients are divided into three levels: low risk (predicted probability <0.3), medium risk (0.3≤predicted probability <0.7) and high risk (predicted probability ≥0.7), providing a reference for clinicians and guiding subsequent intervention and treatment plans.

[0117] Preferably, refer to Figure 12 This method also includes deploying the optimal prediction model based on a web framework, providing a simple data input interface and result display interface for easy use by clinicians. The web application integrates data input, prediction calculation, result display, and interpretation visualization functions to achieve a one-stop myosteatosis risk assessment.

[0118] This application also provides an operational example. Specifically, in the actual operation process, the acquisition and processing of HD patient-related data samples are as follows:

[0119] First, patients who were on maintenance hemodialysis (dialysis duration ≥ 3 months, hemodialysis frequency: 3 times / week) at four tertiary-level A hospitals were independently screened. Inclusion criteria: (1) regular hemodialysis due to end-stage renal disease; (2) urea clearance index (Kt / V) > 1.2; (3) aged between 18 and 80 years; (4) chest non-enhanced multi-slice CT scan images. Exclusion criteria: (1) chest CT without L1 slice images; (2) acute heart failure or myocardial infarction; (3) malignant tumors; (4) liver failure and respiratory failure; (5) inflammatory bowel disease; (6) Alzheimer's disease; (7) serious missing clinical and laboratory data.

[0120] Then, clinical data of hemodialysis patients were collected, including the above-mentioned demographic characteristics, anthropometric characteristics, clinical characteristics, comorbidities, concomitant medications, and laboratory tests, and the corresponding derived indicators were calculated.

[0121] At the same time, if Figure 2As shown, the degree of myosteatosis was measured using chest CT images. The mean muscle density (SMD; Hounsfield unit, HU) of total skeletal muscle was measured using chest CT L1 images to quantify skeletal muscle mass. Myosteatosis was defined as SMD ≤ 35.91 HU (males) or ≤ 30.64 HU (females), and normal muscle mass (non-myosteatosis) was defined as SMD > 35.91 HU (males) or > 30.64 HU (females).

[0122] CT equipment parameters were uniformly set: tube voltage 120 kV, tube current 100 mAs, pitch 0.8, and slice thickness 5 mm. Regions of interest at the L1 level included the external oblique, quadratus lumborum, internal oblique, erector spinae, transverse abdominis, and rectus abdominis muscles, and quantification was performed using 3DSlicer software.

[0123] Based on the collected HD patient-related data samples, the AFIE algorithm proposed in this application is used to perform feature screening, specifically including:

[0124] First, the Boruta algorithm is used to assess the initial feature importance. This algorithm is implemented through the following steps: creating "shadow features" by randomly permuting the original features; building a random forest model that includes the original features and the shadow features; calculating the importance scores of all features; determining the maximum importance score (MZSA) among the shadow features; labeling each original feature with a status: if the feature importance is significantly higher than the MZSA, it is marked as "confirmed"; if it is significantly lower than the MZSA, it is marked as "rejected"; and the rest are marked as "tentative." This step is repeated until the maximum number of iterations (set to 100) is reached or all features are marked as "confirmed" or "rejected."

[0125] Then, we use Lasso regression for sparse feature selection. The steps are as follows: set family = "binomial" and use 10-fold cross-validation to determine the optimal regularization parameter λ; use the optimal λ value to build a Lasso model and extract features corresponding to non-zero coefficients; calculate the feature importance score, which is the absolute value of the feature coefficient.

[0126] Next, the feature interaction importance evaluation is performed as follows:

[0127] Construct a random forest model f(x). For each pair of features (j, k), calculate the interaction importance by permuting the feature values. The specific formula is as follows:

[0128]

[0129] in, Represents the sample after feature j and feature k are randomly arranged at the same time.

[0130] To improve computational efficiency, the H-statistic is used to estimate the importance of interaction:

[0131]

[0132] Where V is the total variance of the model; V j and V k are the main effect variances of feature j and feature k respectively; V jk is the common variance contribution of feature j and feature k.

[0133] Based on the three aforementioned evaluation methods, the comprehensive score of each feature is calculated:

[0134]

[0135] Among them, B, L, and I represent the Boruta score, Lasso coefficient, and interaction importance set of all features, respectively. The numerator is normalized so that the scores from different sources are weighted and combined on the same scale.

[0136] Finally, the dynamic threshold is determined to select the final important features. The steps include:

[0137] Calculate the distribution characteristics of the comprehensive score of all features (mean μ S and standard deviation σ S ); dynamically determine the threshold value based on the data characteristics; select the features with scores greater than the threshold value as the final important feature set. The threshold value is calculated as follows:

[0138] θ=μ S +α·σ S ·(1+β·CV S )

[0139] Among them, CV S is the coefficient of variation, equal to σ S / μ S , β is a tuning parameter with a default value of 0.2, which is used to adjust the threshold according to the degree of dispersion of the score distribution.

[0140] In this example, the AFIE algorithm ultimately screened out seven key features: age, BMI, diabetes, CVD, ALB, GLR, and HDL. Compared to using the Boruta algorithm or Lasso regression alone, the AFIE algorithm significantly improved the stability of feature selection, increasing consistency across different data subsets by 21.5%. It also improved predictive performance, with the model's AUC increasing by 3.5%, while also reducing the number of features from 11 in traditional methods to just 7.

[0141] After feature screening is completed, a preset model is constructed and trained based on the screened important feature set and the corresponding training set.

[0142] Specifically, this embodiment constructs and comprehensively evaluates seven machine models, including Naive Bayes, Nearest Neighbor, Decision Tree, Random Forest, Support Vector Machine, Extreme Gradient Boosting, and Logistic Regression. Among them, Naive Bayes assumes independence between features and calculates posterior probabilities based on Bayes' theorem; Nearest Neighbor classifies based on distance metrics in the sample feature space; Decision Tree recursively partitions data by building a tree structure; Random Forest integrates multiple decision trees and determines the classification result through voting; Support Vector Machine searches for the optimal hyperplane to separate different categories; Extreme Gradient Boosting is an ensemble learning algorithm based on gradient boosting decision trees; and Logistic Regression models the nonlinear relationship between independent and dependent variables.

[0143] For each model, parameters were optimized on the training set, and performance was evaluated on internal and external validation sets. Evaluation metrics included area under the curve (AUC), accuracy, sensitivity, specificity, positive predictive value (PPV), negative predictive value (NPV), F1 score, and Kappa score. AUC reflects the model's ability to distinguish between different classes; accuracy reflects the proportion of correctly predicted samples; sensitivity reflects the true positive rate (TRR), the ability to correctly identify positive samples; specificity reflects the true negative rate (TRR), the ability to correctly identify negative samples; PPV reflects the proportion of true positives among positive predictions; NPV reflects the proportion of true negatives among negative predictions; F1 score reflects the harmonic mean of precision and recall; and Kappa score reflects the consistency of the model, taking into account the probability of random classification.

[0144] A multi-dimensional comparison revealed that the extreme gradient boosting model performed best overall on both the internal and external validation sets. In particular, on the external validation set, the extreme gradient boosting model achieved an AUC of 0.789, accuracy of 0.789, sensitivity of 0.863, and specificity of 0.776, all exceeding those of the other models.

[0145] In addition, to ensure the reliability of the evaluation results, five-fold and ten-fold cross-validation were performed on the validation set to further verify the stability of the model performance.

[0146] This embodiment details the implementation of the interaction-sensitive SHAP value calculation (IS-SHAP) method. Figures 8 to 11 As shown, the IS-SHAP method provides global and local explanations of the model and enhances model transparency.

[0147] The complete implementation steps of the IS-SHAP method are as follows:

[0148] First, calculate the total SHAP value of each feature: for sample x, determine the conditional expectation E[f(X)|XS=xS], where S is the feature subset; use Monte Carlo sampling to approximate the calculation, and the specific formula is as follows:

[0149]

[0150] in, is the estimated total SHAP value of feature i; M is the number of sampling times (set to 1000 in this embodiment); is the sample containing feature i in the mth sampling; is a sample that does not contain feature i.

[0151] Next, calculate the interaction SHAP value between feature pairs: define the interaction difference Δ ij (S), the specific formula is as follows:

[0152] Δ ij (S) = f(x S∪{i,j} )-f(x S∪{i} )-f(x s∪{j} )+f(x S )

[0153] Based on the interaction difference, the interaction SHAP value is calculated. The specific formula is as follows:

[0154]

[0155] In order to improve the calculation efficiency, the replacement sampling method is used for approximate calculation. The specific formula is as follows:

[0156]

[0157] Among them, Δ i j(S) is the interactive difference between feature i and feature j in feature subset S; f(x S ∪{i,j}) is the model prediction value corresponding to the input containing feature subset S, feature i and feature j; f(x S ∪{i}) is the model prediction value corresponding to the input containing the feature subset S and feature i; f(x S ∪{j}) is the model prediction value corresponding to the input containing feature subset S and feature j; f(x S ) is the model prediction value corresponding to the input containing only the feature subset S; ∪ is the set union operator; φ ij(f,x) is the SHAP value of the interaction between feature i and feature j; f is the prediction function of the machine learning model; x is the input sample to be explained; S is the feature subset that does not contain feature i and feature j; N is the set of all features; i and j are feature indices; |S| is the size of feature subset S, that is, the number of features contained; |N| is the total number of features; |S|! is the factorial of |S|; (|N|-|S|-2)! is the factorial of (|N|-|S|-2); (|N|-1)! is the factorial of (|N|-1); Δ ij (S) is the interactive difference between feature i and feature j in feature subset S; S m is the feature subset randomly sampled for the mth time.

[0158] The total SHAP value of each feature is decomposed into main effects and interaction effects. The specific formula is as follows:

[0159]

[0160] in, represents the main effect of feature i, and the calculation formula is:

[0161]

[0162] in, is the total SHAP value of feature i; is the main effect SHAP value of feature i; is the average distribution coefficient of the interaction effect, φ ij (f,x) is the interaction SHAP value of feature i and feature j.

[0163] Total SHAP value of feature i It comes from two parts: the first part is the main effect SHAP value of feature i It represents the independent impact of feature i on the model prediction results; the second part is the sum of the interactive effects of feature i and all other features. Indicates the influence of feature i on the model prediction results through interaction with other features. The role of this is to ensure that each interaction effect φ ij (f, x) is evenly distributed between the two interacting features i and x, avoiding repeated calculations of interaction effects. This ensures that the sum of the SHAP values ​​of all features equals the difference between the model's final prediction and the baseline. This decomposition adheres to the axiomatic requirements of Shapley values ​​and ensures fair and complete contribution distribution.

[0164] Evaluate the relative strength of the interaction between features. The specific formula is as follows:

[0165]

[0166] Among them, R ij The value range is [0,1]. The larger the value, the stronger the interaction effect and the weaker the main effect.

[0167] Based on the IS-SHAP value, three visualizations are implemented:

[0168] (1) SHAP summary bar chart: displays the average contribution of features and displays the proportion of main effects and interaction effects in segments through bar charts;

[0169] (2) SHAP summary dot plot: The relationship between the eigenvalue and the SHAP value is displayed in the form of a scatter plot, and the color coding indicates the size of the eigenvalue;

[0170] (3) SHAP dependency plot: visualizes the impact of feature value changes on SHAP values, including main effect curves and interaction effect intervals.

[0171] In addition, for local interpretation of individual samples, a SHAP waterfall plot is used to show the contribution of each feature, sorted by contribution size, and mark the main interaction effects.

[0172] The IS-SHAP method was used for analysis and found that: (1) there was a positive interaction effect between age and cardiovascular disease, that is, when elderly people were accompanied by cardiovascular disease, the increased risk of muscle fatty degeneration was greater than the sum of the two independent effects; (2) there was a negative interaction effect between serum albumin and glucose lymphocyte ratio, and high albumin levels could partially alleviate the risk brought by high GLR; (3) the interaction effect between body mass index and high-density lipoprotein varied in different BMI ranges, and the interaction effect was strongest when the BMI was between 23-28.

[0173] These findings provide a new perspective for clinical understanding of the pathological mechanisms of myofeeatosis and a theoretical basis for individualized intervention strategies.

[0174] This example details the risk stratification assessment method and system deployment. Based on the probability values ​​output by the optimal prediction model, patients are categorized into three risk levels: low risk, medium risk, and high risk. Low risk: predicted probability < 0.3; medium risk: 0.3 ≤ predicted probability < 0.7; and high risk: predicted probability ≥ 0.7.

[0175] To facilitate clinical application, the present invention deploys the prediction model as a Web application based on the Streamlit framework. Figure 12 As shown in the figure, the application provides a simple data input interface. Doctors only need to enter the patient's seven key characteristics (age, BMI, diabetes, CVD, ALB, GLR and HDL), and the system can automatically calculate the risk probability and risk level of myosteatosis.

[0176] The web application also provides a wealth of visualization capabilities, including: 1. Data distribution display: used to understand the overall distribution of training data; 2. Feature distribution display: shows the distribution differences of each feature in different risk groups; 3. Model performance indicators: displays the evaluation results of the model, including ROC curves, accuracy, etc.; 4. SHAP value interpretation: provides global and individual feature contribution analysis; 5. Risk assessment results: intuitively displays the predicted probability and risk level.

[0177] The deployment of this application greatly simplifies the use of predictive models, allowing clinicians to quickly and conveniently assess patients' risk of myosteatosis and provide support for clinical decision-making.

[0178] This application demonstrates the practical effects of the present invention through a clinical application case, as follows:

[0179] A hospital used the risk stratification assessment system presented in this paper to assess the risk of myosteatosis in a group of 56-year-old hemodialysis patients. The doctor collected the patient's relevant data and entered it into the system: age 56, BMI 23.10, no history of diabetes or cardiovascular disease, ALB 35.97 g / L, GLR 8.00, and HDL 0.96 mmol / L.

[0180] The system automatically processed the input data and calculated a predicted probability of 89.25% using an extreme gradient boosting model, indicating a high risk. The system also interpreted the predictions based on the IS-SHAP method, indicating that the patient's advanced age was the primary risk factor, contributing a SHAP value of approximately 0.6, followed by a relatively low ALB level. The IS-SHAP analysis also revealed an interaction effect between age and ALB, explaining an additional 10.5% of the predicted variance compared to considering only the main effects.

[0181] Based on the evaluation results, the doctor developed a targeted intervention plan for the patient, including optimizing nutrition, increasing moderate exercise, and regular follow-up visits. This plan effectively prevented the onset and progression of myosteatosis. At a follow-up visit three months later, the patient's ALB level had increased to 41.3 g / L, and reassessment indicated that his risk of myosteatosis had fallen to a moderate level.

[0182] Through six months of follow-up observation, the proportion of patients assessed as high-risk using this system who actually developed muscle fatty degeneration was 78.3%, while only 12.5% ​​of patients assessed as low-risk developed muscle fatty degeneration, verifying the predictive accuracy and clinical practical value of this system.

[0183] This case demonstrates the advantages of this invention in practical clinical applications: accurate risk prediction, providing explainable prediction basis, supporting individualized intervention decisions, and enabling rapid evaluation through a convenient web interface.

[0184] The technical features of the above embodiments can be combined arbitrarily. To make the description concise, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.

[0185] The above-described embodiments merely represent several implementation methods of the present application. While the descriptions are relatively specific and detailed, they should not be construed as limiting the scope of the present application. It should be noted that a person of ordinary skill in the art may make various modifications and improvements without departing from the spirit of the present application, and these modifications and improvements fall within the scope of protection of the present application. Therefore, the scope of protection of the present application shall be determined by the appended claims.

Claims

1. A method for assessing the risk of myosteatosis based on machine learning, characterized in that: include; Obtain data samples from HD patients; Based on the interaction importance between features, the relevant features in the data sample are screened using the Boruta algorithm and Lasso regression to generate an important feature set; Based on the important feature set, use the corresponding data samples to build and train multiple machine learning prediction models and determine the optimal model; The data of actual HD patients were input into the optimal model to generate prediction results; Perform global and local interpretations of prediction results based on SHAP values; Generate risk assessment results based on global and local interpretations of the prediction results.

2. The method according to claim 1, characterized in that The data sample includes demographic characteristics, anthropometric characteristics, clinical characteristics, comorbidities, concomitant medications, laboratory tests and chest CT images; among them, demographic characteristics include age and gender; anthropometric characteristics include height, weight and body mass index; clinical characteristics include dialysis age, smoking history, and drinking history; comorbidities include diabetes, hypertension, and cardiovascular disease; concomitant medications include lipid-lowering drugs, iron supplements, erythropoietin and compound α-keto acids; laboratory tests include white blood cell count, hemoglobin, platelet count, serum albumin, neutrophil count, lymphocyte count, monocyte count, fasting blood glucose, total cholesterol, triglycerides, low-density lipoprotein and high-density lipoprotein; chest CT images must include the first lumbar transverse process level.

3. The method according to claim 1 or 2, characterized in that Based on the importance of interaction between features, the relevant features in the data sample are screened using the Boruta algorithm and Lasso regression, including: Based on the Boruta algorithm, the initial feature importance set is obtained according to the relevant features in the data sample, and based on the Lasso regression, the sparse feature set is obtained according to the relevant features in the data sample; Calculate the interaction importance between features based on the preset feature interaction importance matrix; Calculate the comprehensive score of features based on the interaction importance between features; When the comprehensive score is greater than the threshold, the corresponding feature is selected into the important feature set.

4. The method according to claim 3, characterized in that The calculation formula of the feature interaction importance matrix is ​​as follows: Among them, I jk is the interaction importance between feature j and feature k; f(x i ) is the model for sample x i The prediction function of It is the second-order partial derivative of the prediction function with respect to feature j and feature k, which is used to measure the strength of the interaction between the two features.

5. The method according to claim 3, characterized in that The threshold is a dynamic threshold, and the formula is as follows: θ=μ S +a·s S Among them, θ is the dynamic threshold; μ S is the mean of the comprehensive scores of all features; σ S is the standard deviation of the comprehensive score of all features; α is an adjustment parameter with a default value of 0.

8.

6. The method according to claim 1 or 2, characterized in that Multiple machine learning prediction models including Naive Bayes, Nearest Neighbor, Decision Tree, Random Forest, Support Vector Machine, Extreme Gradient Boosting, and Logistic Regression.

7. The method according to claim 1 or 2, characterized in that Global and local interpretations of prediction results based on SHAP values ​​include: Calculate the total SHAP value of each feature and the interaction SHAP value between features; Generate the main SHAP value of each feature based on the total SHAP value and the interaction SHAP value; The relative strength of the interaction between features is evaluated based on the interaction SHAP value and the main SHAP value; The corresponding visualization graphs are constructed using the total SHAP value, interactive SHAP value, main SHAP value and the relative strength of the interaction between features, and the visualization graphs are used to provide global and local explanations for the prediction results.

8. The method according to claim 7, characterized in that The formula for calculating the interaction SHAP value between features is as follows: Δ ij (S)=f(x S∪{i,j} )-f(x S∪{i} )-f(x S∪{j} )+f(x S ) Among them, Δ i j(S) is the interactive difference between feature i and feature j in feature subset S; f(x S ∪{i,j}) is the model prediction value corresponding to the input containing feature subset S, feature i and feature j; f(x S ∪{i}) is the model prediction value corresponding to the input containing the feature subset S and feature i; f(x S ∪{j}) is the model prediction value corresponding to the input containing feature subset S and feature j; f(x S ) is the model prediction value corresponding to the input containing only the feature subset S; ∪ is the set union operator; φ ij (f,x) is the SHAP value of the interaction between feature i and feature j; f is the prediction function of the machine learning model; x is the input sample to be explained; S is the feature subset that does not contain feature i and feature j; N is the set of all features; i and j are feature indices; |S| is the size of feature subset S, that is, the number of features contained; |N| is the total number of features; |S|! is the factorial of |S|; (|N|-|S|-2)! is the factorial of (|N|-|S|-2); (|N|-1)! is the factorial of (|N|-1); Δ ij (S) is the interaction difference between feature i and feature j under feature subset S.

9. The method according to claim 7, characterized in that The total SHAP value of each feature is divided into main effect and interaction effect; the formula for generating the main SHAP value of each feature based on the interaction SHAP value is as follows: in, is the total SHAP value of feature i; is the main effect SHAP value of feature i; is the average distribution coefficient of the interaction effect; φ ij (f,x) is the interaction SHAP value of feature i and feature j.

10. The method according to claim 9, characterized in that The formula for evaluating the relative strength of the interaction between features based on the interaction SHAP value and the main SHAP value is as follows: Among them, R ij is the relative strength of the interaction between features, and its value range is [0,1].

Citation Information

Cited By

  • Method for predicting intestinal preparation quality of diabetic patient based on SHAP

    CN121122679A

  • Diabetes patient bowel preparation quality prediction method based on shap

    CN121122679B