CVD (Chemical Vapor Deposition) death subgroup identification method and device in combination with causal reasoning and consensus clustering

By combining causal reasoning and consensus clustering methods, identifying CVD death subgroups is solved, and the existing models lack mediated causal relationships are achieved, precise prediction and subgroup identification of CVD death risk are achieved, and strategies and causal mechanism understanding are provided to reduce mortality.

CN120260930AActive Publication Date: 2025-07-04THE FIRST AFFILIATED HOSPITAL OF XIAMEN UNIV

Patent Information

Application Number
CN202510727453.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-03
Publication Date
2025-07-04
Estimated Expiration
2045-06-03

AI Technical Summary

Technical Problem

The existing CVD death risk prediction model lacks mediated causal inference, cannot reveal potential causal relationships, and fails to effectively analyze CVD death subpopulations, making it difficult to achieve precise intervention.

Method used

Combining causal reasoning and consensus clustering methods, we collect and preprocess feature data, mediate causal reasoning, use machine learning models to predict death risk, use SHAP algorithm to select the optimal model and feature set, conduct consensus clustering to identify CVD death subgroups, and conduct statistical analysis to identify pathways.

Benefits of technology

Revealing the potential pathways for CVD death, identifying key feature variables, reducing model complexity, providing references to risk stratification and prevention strategies, and revealing the causal mechanisms between different subpopulations.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120260930A_ABST
    Figure CN120260930A_ABST
Patent Text Reader

Abstract

The invention discloses a CVD (Chemical Vapor Deposition) death subgroup identification method and device combining causal reasoning and consensus clustering. The method comprises the following steps: collecting and preprocessing original feature data; performing intermediary causal relationship reasoning on the preprocessed feature data, and taking features contained in an intermediary causal relationship reasoning result as an initial feature set; performing death risk prediction according to the input feature data by using a plurality of machine learning models, and training by using the initial feature set to obtain a plurality of death risk prediction models; calculating SHAP values of the feature variables of all the death risk prediction models by adopting an SHAP algorithm, and selecting an optimal model and an optimal feature set by utilizing the SHAP values of the feature variables; performing consensus clustering on the optimal feature set and a death risk prediction result output by the optimal model to obtain a plurality of death subgroups; and carrying out statistical analysis on the feature data and survival results of each death subgroup, and identifying a path of each death subgroup in combination with an intermediary causal relationship reasoning result.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of medical and health information technology, and particularly to a method and device for identifying CVD death subgroups by combining causal inference and consensus clustering. Background Art

[0002] As an emerging artificial intelligence technology, ML has been widely used in CVD death risk prediction in recent years and has shown significant advantages, especially in dealing with large-scale, complex, and multi-dimensional data. ML can help medical institutions identify individuals who may be at risk of CVD death at an early stage, so as to achieve timely intervention to reduce CVD-related deaths. Currently, a variety of features have been used to construct CVD death risk prediction models, including dietary data, lifestyle factors, environmental exposure data, retinal imaging, low-dose computed tomography, and traditional clinical variables such as demographic characteristics, medical history, and basic laboratory indicators.

[0003] However, existing CVD death risk prediction models are mainly targeted at specific disease populations. Although these models have achieved good performance in predicting future CVD death risks, there are still certain limitations. The key features identified by these models are usually closely related to CVD death in specific patient groups, but lack mediating causal inference and cannot reveal the potential mediating causal relationships related to CVD death. Understanding these causal relationships is crucial for fundamentally reducing CVD mortality and clarifying its underlying mechanisms. In addition, previous studies rarely conduct stratified analysis on CVD death subgroups to explore the potential mechanisms of the transition of individuals between different survival states, and this aspect is particularly critical for medical institutions to prevent the further development of diseases and implement precise interventions. Summary of the Invention

[0004] The purpose of the present invention is to solve the problems in the prior art.

[0005] The technical solution adopted by the present invention to solve its technical problems is: to provide a method for identifying CVD death subgroups by combining causal inference and consensus clustering, including the following steps:

[0006] Collect original feature data and perform preprocessing to obtain preprocessed feature data, and the feature variables in the preprocessed feature data include exposure factors and laboratory indicators;

[0007] Perform mediating causal relationship inference on the preprocessed feature data, and use the exposure factors and laboratory indicators included in the mediating causal relationship inference result as the initial feature set;

[0008] Use a number of machine learning models to perform death risk prediction based on the input feature data respectively, and then train them using the initial feature set respectively to obtain a number of death risk prediction models; use the SHAP algorithm to calculate the SHAP values of the feature variables of all death risk prediction models, and select the optimal model and the optimal feature set using the SHAP values of the feature variables of each death risk prediction model;

[0009] Perform consensus clustering on the death risk prediction results output by the optimal feature set and the optimal model to obtain a number of death subgroups;

[0010] Perform statistical analysis on the feature data and survival results of each death subgroup, and identify the pathways of each death subgroup in combination with the statistical analysis results and the results of mediating causal relationship reasoning.

[0011] Preferably, the preprocessing includes:

[0012] Eliminate feature variables with missing values exceeding 50%;

[0013] Eliminate the feature data of subjects under 20 years old;

[0014] Eliminate the feature data of subjects with missing values exceeding 50%;

[0015] Classify each subject based on the survival status of each subject at different time points.

[0016] Preferably, the classifying each subject based on the survival status of each subject at different time points includes:

[0017] At the three-year follow-up, classify the subjects into the survival group or the death group according to whether they survive;

[0018] At the five-year follow-up, classify the subjects into the survival group or the death group according to whether they survive;

[0019] At the ten-year follow-up, classify the subjects into the survival group or the death group according to whether they survive.

[0020] Preferably, the performing mediating causal relationship reasoning on the preprocessed feature data includes the following steps:

[0021] Use the K-nearest neighbor algorithm to fill in the preprocessed feature data;

[0022] Exhaustively combine the exposure factors and laboratory indicators; for each pair of exposure factors and laboratory indicators, verify them based on a number of conditional independent tests to obtain a number of P-values, and use the maximum value among them as the overall P-value;

[0023] If the overall P-value is less than 0.05, it means that the exposure factor can affect CVD death through the laboratory indicator as a mediating variable;

[0024] Among them, the several conditionally independent tests include:

[0025] The exposure factor is significantly associated with the death outcome;

[0026] After adjusting for the death outcome, the exposure factor is still significantly associated with the mediating variable; the mediating variable refers to laboratory indicators;

[0027] After adjusting for the exposure factor, the mediating variable is still significantly associated with the death outcome;

[0028] After adjusting for the mediating variable, the exposure factor is not associated with the death outcome.

[0029] Preferably, the SHAP values are calculated for the feature variables of all death risk prediction models using the SHAP algorithm, and the optimal model and the optimal feature set are selected using the SHAP values of the feature variables of each death risk prediction model. Specifically:

[0030] The SHAP values are calculated for the feature variables of all death risk prediction models using the SHAP algorithm;

[0031] For each death risk prediction model, the number of input features is gradually reduced based on the SHAP values of the feature variables, and the prediction effect is evaluated. The optimal model is selected by combining the prediction effect and the number of feature variables; the set of input feature variables of the optimal model is the optimal feature set.

[0032] Preferably, the optimal feature set includes 4 exposure factors and 6 laboratory indicators; the 4 exposure factors are age, household income-to-poverty ratio, gender, and marital status; the 6 laboratory indicators are lymphocyte percentage, neutrophil percentage, hemoglobin, hematocrit, mean corpuscular volume, and osmotic pressure; the optimal model is a logistic regression model with the optimal feature set as the input feature variables.

[0033] Preferably, the death risk prediction results output by the optimal feature set and the optimal model are subjected to consensus clustering to obtain several death subgroups, including the following steps:

[0034] The feature values of the final feature set are input into the optimal prediction model to obtain the death risk prediction results;

[0035] The feature values and the prediction results are subjected to consensus clustering to obtain three death subgroups.

[0036] Preferably, the statistical analysis of the feature data and survival results of each death subgroup includes:

[0037] Survival analysis was performed. The Kaplan-Meier survival curve was used to estimate the survival probability of each death subgroup, and the Log-rank test was used to compare the differences in survival curves between groups.

[0038] Descriptive analysis was carried out. Continuous variables were expressed as the median and quartiles, and categorical variables were expressed as frequencies and percentages.

[0039] Comparison between groups: For the comparison between two groups, the Mann-Whitney U test was used for continuous variables, and the chi-square test or Fisher's exact test was used for categorical variables. For multiple comparisons among three groups, the Bonferroni correction was used to adjust the P value to control type I error.

[0040] Among them, the continuous variables included age, household income-to-poverty ratio, lymphocyte percentage, neutrophil percentage, hemoglobin, hematocrit, mean corpuscular volume, and osmotic pressure; the categorical variables included gender and marital status.

[0041] Preferably, the pathways of each death subgroup were identified by combining the statistical analysis results and the results of mediating causal relationship reasoning. The specific pathways of the death subgroups were as follows:

[0042] Age affects the transformation between different subgroups by influencing hemoglobin, hematocrit, lymphocyte percentage, neutrophil percentage, and osmotic pressure.

[0043] Gender affects the transformation between different subgroups by influencing lymphocyte percentage, osmotic pressure, and mean corpuscular volume.

[0044] Marital status affects the transformation between different subgroups by influencing the changes in lymphocyte percentage, neutrophil percentage, osmotic pressure, and mean corpuscular volume.

[0045] Household income-to-poverty ratio affects the transformation between different subgroups by influencing the changes in hemoglobin and hematocrit.

[0046] The present invention also provides a CVD death subgroup identification device combining causal reasoning and consensus clustering, including:

[0047] A data collection module, which collects original feature data and performs preprocessing to obtain preprocessed feature data. The feature variables in the preprocessed feature data include exposure factors and laboratory indicators.

[0048] A causal reasoning module, which performs mediating causal relationship reasoning on the preprocessed feature data, and uses the exposure factors and laboratory indicators included in the results of mediating causal relationship reasoning as the initial feature set.

[0049] The model construction module uses a number of machine learning models to predict the death risk based on the input feature data respectively, and then trains them using the initial feature set respectively to obtain a number of death risk prediction models; the SHAP algorithm is used to calculate the SHAP values of the feature variables of all the death risk prediction models, and the optimal model and the optimal feature set are selected using the SHAP values of the feature variables of each death risk prediction model;

[0050] The subgroup acquisition module performs consensus clustering on the death risk prediction results output by the optimal feature set and the optimal model to obtain a number of death subgroups;

[0051] The subgroup identification module performs statistical analysis on the feature data and survival results of each death subgroup, and identifies the pathways of each death subgroup by combining the statistical analysis results and the results of mediating causal relationship reasoning.

[0052] The present invention has the following beneficial effects: First, it reveals the potential pathways of CVD death by using mediating causal inference, providing an important reference for medical institutions to reduce the CVD mortality rate and optimize prevention strategies; Second, it identifies the 10 most critical feature variables in the CVD death pathway, and these features are all common clinical indicators, effectively reducing the model complexity while ensuring the model performance, laying a good foundation for the promotion of the model in the real world; Finally, these 10 feature variables can be further used for risk stratification of the CVD death population, and reveal the potential causal mechanism of the mutual transformation between different CVD death subgroups.

[0053] The present invention will be further described in detail below in conjunction with the drawings and embodiments, but the present invention is not limited to the embodiments. Description of the Drawings

[0054] Figure 1 It is a method step diagram of a method for identifying CVD death subgroups combining causal reasoning and consensus clustering according to an embodiment of the present invention;

[0055] Figure 2 It is a schematic diagram of the data preprocessing process of a method for identifying CVD death subgroups combining causal reasoning and consensus clustering according to an embodiment of the present invention;

[0056] Figure 3 It is a visualization schematic diagram of the results of mediating causal relationship reasoning of a method for identifying CVD death subgroups combining causal reasoning and consensus clustering according to an embodiment of the present invention;

[0057] Figure 4 It is the prediction results of four death risk prediction models for 3-year follow-up of a method for identifying CVD death subgroups combining causal reasoning and consensus clustering according to an embodiment of the present invention;

[0058] Figure 5The change curve of AUC of the death risk prediction model for the CVD death subgroup identification method combining causal reasoning and consensus clustering in the embodiments of the present invention with the reduction of features;

[0059] Figure 6 The Kaplan-Meier survival curve diagram in different subgroups for the CVD death subgroup identification method combining causal reasoning and consensus clustering in the embodiments of the present invention;

[0060] Figure 7 The distribution diagram of different features in different subgroups for the CVD death subgroup identification method combining causal reasoning and consensus clustering in the embodiments of the present invention;

[0061] Figure 8 The visualization schematic diagram of the conversion network between different death subgroups for the CVD death subgroup identification method combining causal reasoning and consensus clustering in the embodiments of the present invention;

[0062] Figure 9 The structural schematic diagram of the CVD death subgroup identification device combining causal reasoning and consensus clustering in the embodiments of the present invention. Detailed implementation manners

[0063] See Figure 1 As shown, it is the method step diagram of the CVD death subgroup identification method combining causal reasoning and consensus clustering in the embodiments of the present invention, including the following steps:

[0064] S101, collect the original feature data and perform preprocessing to obtain the preprocessed feature data, and the feature variables in the preprocessed feature data include exposure factors and laboratory indicators;

[0065] S102, perform mediation causal relationship reasoning on the preprocessed feature data to obtain the exposure factors and laboratory indicators included in the mediation causal relationship reasoning result as the initial feature set;

[0066] S103, use several machine learning models to perform death risk prediction respectively according to the input feature data, and then use the initial feature set for training to obtain several death risk prediction models; use the SHAP algorithm to calculate the SHAP values of the feature variables of all death risk prediction models, and select the optimal model and the optimal feature set by using the SHAP values of the feature variables of each death risk prediction model;

[0067] S104, perform consensus clustering on the death risk prediction results output by the optimal feature set and the optimal model to obtain several death subgroups;

[0068] S105. Perform statistical analysis on the characteristic data and survival outcomes of each dead subgroup, and identify the pathways of each dead subgroup by combining the statistical analysis results and the results of mediating causal relationship inference.

[0069] Specifically, in S101, the embodiments of the present invention use the NHANES dataset, including 40,617 participants, whose characteristic variables include basic demographic information, anthropometric data, laboratory indicators, and questionnaire survey data, and preprocess the data therein. For the preprocessing, first, the characteristic variables with missing values exceeding 50% are excluded, and 20 exposure factors and 47 laboratory indicators are selected for further analysis; then, referring to Figure 2 as shown, individuals under 20 years old and participants with characteristic missing values exceeding 50% are excluded to further optimize the study population and advance the analysis in the next stage; the death status data is matched with the exposure data of NHANES, and the subjects are classified as dead or alive at different time points. Since there are differences in the follow-up durations of the subjects, the follow-up intervals are standardized, and three key time points of 3 years, 5 years, and 10 years are preset. At these 3 time points, all subjects are classified into the CVD death group or the survival group. In addition, to ensure that the research focus is on CVD-related deaths, all individuals who died of non-CVD causes are excluded.

[0070] Specifically, the 20 exposure factors include: Age, age; Gender, gender; Education Level - Adults 20+, education level (20 years old and above); Marital Status, marital status; Race, race; Body Mass Index (kg / m²), body mass index (BMI); Waist Circumference (cm), waist circumference; Minutes of sedentary activity, sedentary time (minutes); Ratio of family income to poverty, ratio of family income to poverty; Ever told you had high blood pressure, ever been diagnosed with high blood pressure; Doctor told you have diabetes, doctor diagnosed with diabetes; Ever told you had coronary heart disease, ever been diagnosed with coronary heart disease; Ever told you had a stroke, ever been diagnosed with a stroke; Had at least 12 alcohol drinks / 1 yr, drank alcohol ≥ 12 times in the past year; Vigorous work activity, vigorous work activity; Moderate work activity, moderate-intensity work activity; Walk or bicycle, walking or cycling activity; Vigorous recreational activities, vigorous recreational activities; Moderate recreational activities, moderate-intensity recreational activities; Smoked at least 100 cigarettes in life, smoked ≥ 100 cigarettes in a lifetime. The 47 laboratory indicators include: Albumin_urine (mg / L), urinary albumin; Creatinine_urine (umol / L), urinary creatinine; Direct HDL-Cholesterol (mmol / L), high-density lipoprotein cholesterol (HDL); White blood cell count (1000 cells / uL), white blood cell count; Lymphocyte percent (%), lymphocyte percentage; Monocyte percent (%), monocyte percentage; Segmented neutrophils percent (%), neutrophil percentage; Eosinophils percent (%), eosinophil percentage;Basophils percent(%), percentage of basophils; Lymphocyte number (1000 cells / uL), absolute value of lymphocytes; Monocyte number (1000 cells / uL), absolute value of monocytes; Segmented neutrophils num(1000 cell / uL), absolute value of segmented neutrophils; Eosinophils number (1000 cells / uL), absolute value of eosinophils; Basophils number (1000 cells / uL), absolute value of basophils; Red blood cellcount (million cells / uL), red blood cell count; Hemoglobin (g / dL), hemoglobin; Hematocrit(%), hematocrit; Mean cell volume (fL), mean corpuscular volume (MCV); Mean cell hemoglobin(pg), mean corpuscular hemoglobin (MCH); Mean cell hemoglobin concentration (g / dL), mean corpuscular hemoglobin concentration (MCHC); Red cell distribution width (%), red cell distribution width (RDW); Platelet count (1000 cells / uL), platelet count (PLT); Mean platelet volume (fL), mean platelet volume (MPV); Glycohemoglobin (%), glycated hemoglobin (HbA1c); Albumin (g / L), serum albumin; Alanine Aminotransferase (ALT) (U / L), alanine aminotransferase; AspartateAminotransferase (AST) (U / L), aspartate aminotransferase; Alkaline Phosphatase (ALP) (IU / L), alkaline phosphatase; Blood urea nitrogen (mmol / L), blood urea nitrogen; Total calcium (mmol / L), total calcium; Cholesterol (mmol / L), total cholesterol; Bicarbonate (mmol / L), bicarbonate (HCO3⁻); Creatinine (umol / L), serum creatinine; Gamma Glutamyl Transferase (GGT) (U / L), gamma-glutamyl transferase;Glucose_serum (mmol / L), blood glucose; Iron_refigerated (umol / L), serum iron; LactateDehydrogenase (LDH) (U / L), lactate dehydrogenase; Phosphorus (mmol / L), phosphorus; Total bilirubin(umol / L), total bilirubin; Total protein (g / L), total protein; Triglycerides (mmol / L), triglycerides; Uric acid (umol / L), uric acid; Sodium (mmol / L), sodium ion; Potassium (mmol / L), potassium ion; Chloride (mmol / L), chloride ion; Osmolality (mmol / Kg), osmolality; Globulin (g / L), globulin.;

[0071] Specifically, in S102, for the subject characteristic data with missing values, the K-nearest neighbor (KNN) algorithm is used for imputation. KNN is a non-parametric statistical learning method that can estimate missing data using the feature values of similar individuals in the dataset. After the missing value imputation is completed, the R software package 'cit' is used to explore the mediating causal relationship between exposure factors, laboratory indicators, and CVD death. This software package uses a likelihood-based hypothesis testing method to evaluate the causal mediation effect. Specifically, it examines whether the exposure factor affects CVD death through the laboratory indicator as a mediating variable and is verified based on the following four conditional independence tests: (1) The exposure factor is significantly correlated with the death outcome; (2) After adjusting for the death outcome, the exposure factor is still significantly correlated with the mediating variable (laboratory indicator); (3) After adjusting for the exposure factor, the mediating variable is still significantly correlated with the death outcome; (4) After adjusting for the mediating variable, the exposure factor is not correlated with the death outcome. If the P-values of the above four independent tests are less than 0.05, it means that its hypothesis is satisfied. For example, if the exposure factor is significantly correlated with the death outcome, the output P-value should be less than 0.05. The P-values of the four tests are comprehensively evaluated using the intersection-union test. The significance level of the overall causal inference test is set to P < 0.05, that is, the maximum value of the P-values output by the four tests respectively is the final overall P-value. If the overall P-value is less than 0.05, it means that a certain exposure factor can affect CVD death through a certain laboratory indicator as a mediating variable.

[0072] In the embodiments of the present invention, subjects are divided into a survival group and a CVD death group according to different follow-up time points. There are significant differences between the two groups in terms of exposure factors and laboratory indicators (for example, compared with the survival group, the CVD death group has a significantly higher age and a larger proportion of males). Through mediating causal inference, it is found that within the 3-year follow-up period, a total of 18 exposure factors mediated the occurrence of CVD death through 36 laboratory indicators; while within the 5-year and 10-year follow-up periods, 19 exposure factors mediated CVD death through 40 and 42 laboratory indicators respectively. These exposure factors and laboratory indicators with mediating causal relationships will be used in the next step for the construction of the prediction model. Finally, the mediating causal network is visualized using Cytoscape (version 3.10.2). The schematic diagram of the visualization effect is shown in Figure 3 as shown, where the items in the left column are exposure factors, the items in the middle column are mediating variables (using laboratory indicators as mediating variables), and the items in the right column are death outcomes; the causal relationships between the items in different columns are represented by lines.

[0073] Specifically, in S103, the ML models used include Logistic Regression (LR), Random Forest (RF), Support Vector Machine (SVM), and Extreme Gradient Boosting (XGBoost); its input features are divided into categorical variables and continuous variables. All continuous variables are standardized before model training to eliminate biases caused by different measurement scales. Categorical variables are converted into a numerical format using one-hot encoding to prevent the dataset from introducing ordinal relationships or implicit distance assumptions. After feature preprocessing, the dataset is randomly divided into a training set and a validation set in a ratio of 7:3, and the model is trained only on the training set. To determine the optimal hyperparameters and prevent overfitting, the embodiments of the present invention combine grid search cross-validation and manual tuning for parameter optimization. The hyperparameters to be adjusted are as follows:

[0074] Logistic Regression (LR): C, max_iter, penalty, solver;

[0075] Random Forest (RF): max_depth, min_samples_leaf, n_estimators;

[0076] Support Vector Machine (SVM): C, gamma, kernel;

[0077] Extreme Gradient Boosting (XGBoost): colsample_bytree, gamma, learning_rate, max_depth, n_estimators, subsample.

[0078] All models adopted 5-fold cross-validation and used the area under the receiver operating characteristic (ROC) curve (AUC) as the main evaluation index to select the optimal model. To address the class imbalance problem, the present invention adjusted the class_weight parameter in LR, RF, and SVM, and adjusted the scale_pos_weight parameter in XGBoost. All models were trained using the best estimator and validated on the validation set. The model performance evaluation metrics included sensitivity (Sn), specificity (Sp), positive predictive value (PPV), negative predictive value (NPV), F1 score, Matthews correlation coefficient (MCC), and accuracy (Acc). Meanwhile, the AUC was also used to comprehensively evaluate the model performance. In addition, to further evaluate the robustness of the model, the bootstrap method was adopted in the validation set to calculate all performance evaluation metrics and determine their 95% confidence intervals (CIs).

[0079] In the embodiments of the present invention, to evaluate the performance of exposure factors and laboratory indicators with mediating causal relationships in the CVD death pathway in constructing a CVD death prediction model, the exposure factors and laboratory indicators were used separately or in combination, and four traditional ML methods were used for modeling. The results showed that the model constructed with the exposure factors and laboratory indicators jointly as feature variables had better predictive performance than the models using only exposure factors or laboratory indicators. Among them, the LR model performed best in predicting CVD death at 3 years, 5 years, and 10 years, with AUC values of 0.914 (95% CI: 0.879–0.941), 0.923 (95% CI: 0.895–0.946), and 0.908 (95% CI: 0.886–0.929), respectively. Taking CVD death at 3 years as an example, the death risk prediction results of the four models are shown in Figure 4 shown.

[0080] Due to the "black box" nature of ML, it is difficult to explain the contribution of each feature to the model prediction. Therefore, the SHapley Additive exPlanations (SHAP) algorithm was introduced to explain the output of the machine learning model. The SHAP algorithm assigns a SHAP value to each feature to measure the impact of the feature on the prediction model. Based on the SHAP values, the importance of all features was ranked in different machine learning models, and then models with reduced feature numbers were constructed according to these rankings to confirm the optimal machine learning model and the appropriate number of features. The AUC was used as the main index to evaluate the model performance. In addition, the model complexity, training cost, and interpretability were also comprehensively considered to finally select the best model architecture.

[0081] Specifically, the SHAP algorithm was used to uniformly interpret the four ML models to ensure the comparability of the results. By comparing the feature SHAP values of the four ML models, it was found that although there were slight differences in the feature rankings between different algorithms, there was a considerable overlap among the top 20 features, and these features maintained a highly consistent ranking trend in the 3-year, 5-year, and 10-year CVD death prediction models. Notably, age ranked first in all models, which is consistent with the traditional understanding of CVD death risk factors. Subsequently, the number of features in the model was gradually reduced according to the SHAP value ranking to leverage the advantage of ML in feature selection, identify the most critical CVD death driving factors, while maintaining relatively good model performance and reducing model complexity; the trend of the model's AUC with the reduction of features is shown in Figure 5 As shown, the top 61, 50, 40, 30, 20, 10, and 5 features with the highest SHAP values were selected to evaluate the prediction effects respectively; as can be seen from the figure, in the CVD death prediction at different follow-up time points, the AUC of the LR and XGBoost models decreased the least when the number of features decreased; when determining the final number of features, it was found that when the number of features of the LR model decreased from 20 to 5, the AUC decreased significantly. To balance model performance and complexity, it was decided to select the LR model with 10 features as the optimal model. The 10 features include age, household income-to-poverty ratio, lymphocyte percentage, neutrophil percentage, hemoglobin, hematocrit, mean corpuscular volume, osmotic pressure, gender, and marital status; its AUCs for predicting 3-year, 5-year, and 10-year CVD death were 0.884 (95% CI: 0.847–0.917), 0.896 (95% CI: 0.865–0.924), and 0.879 (95% CI: 0.855–0.902) respectively, and other performance evaluation indicators of the model also performed well, as shown in Table 1. Considering factors such as low model complexity, low training cost, and strong interpretability, the LR model was finally selected as the optimal model.

[0082] Table 1 - Performance evaluation results of the LR model at different follow-up time points:

[0083]

[0084] Specifically, in S104, consensus clustering is used to identify CVD death subgroups; the input data for consensus clustering includes the features of the above-mentioned optimal model and the predicted values of individual CVD death probabilities. In the embodiments of the present invention, all individuals who died of CVD during the 10-year follow-up period are screened out and consensus clustering analysis is performed. The input data for clustering includes all the features of the optimal model and the predicted probabilities of CVD death generated by the model. In the embodiments of the present invention, the 'ConsensusClusterPlus' R package is used for consensus clustering, and the parameter settings are as follows: maxK = 6, reps = 1000, pItem = 0.8, pFeature = 1, clusterAlg = "km", distance = "euclidean". The determination of the number of clusters is based on two factors: the average pairwise agreement matrix (consensus matrix) within the consensus clustering and the relative change area (deltaplot) under the cumulative distribution function (CDF) curve. Finally, the optimal number of subgroups is determined to be 3, and this decision is based on the following two points: (1) When K = 3 and K = 4, the separation effect of the consensus matrix is the clearest; (2) When K = 4, the increase in the area under the CDF curve is smaller than that when K = 3. Therefore, K = 3 is the most suitable choice. According to the clustering results, the CVD death population is divided into three subgroups.

[0085] Specifically, in S105, survival analysis is performed on the subgroups. The purpose of survival analysis is to compare the differences in CVD death risks among individuals in different groups during the follow-up period. In the embodiments of the present invention, the Kaplan-Meier survival curve is used to estimate the survival probabilities of each group; the Log-rank test is used to compare the differences in survival curves between groups. For example, the CVD death risk of a certain subgroup is significantly higher than that of other groups (Log-rank test, P < 0.05), indicating a survival difference. Continuous variables are expressed as the median (25th and 75th percentiles), and categorical variables are expressed as frequencies (percentages).

[0086] To further explore the differences in characteristics among the three subgroups, statistical tests were performed on all 10 features among different subgroups, including descriptive analysis and comparison between groups. The purpose is to summarize and describe the basic characteristics of the research sample and compare the differences in variables between different groups, including: continuous variables are expressed as the median and quartiles, and categorical variables are expressed as frequencies and percentages; for the comparison between two groups, the Mann-Whitney U test is used for continuous variables; the chi-square test (χ²) or Fisher's exact test is used for categorical variables; for multiple comparisons among three groups, the P value is adjusted using the Bonferroni correction to control type I error. For example, the median of a certain experimental group on variable X is significantly higher than that of the control group, P < 0.05, indicating a statistical difference.

[0087] All statistical tests were two-sided tests, and P < 0.05 was considered statistically significant. Survival analysis was performed using R (version 4.4.2), and the remaining statistical analyses were performed using Python (version 3.11.5). In the embodiments of the present invention, Kaplan-Meier survival curves of different CVD death subgroups were obtained. As shown in Figure 6 Figure, the survival of Subgroup 1 was the best, followed by Subgroup 2, while the survival of Subgroup 3 was the worst. The statistical distribution results of different characteristics in different death subgroups were obtained. As shown in Figure 7 and Table 2, where Figure 7 the statistical graphs in were represented by box plots or stacked bar graphs. ns indicated no significance, ** indicated p < 0.01, *** indicated p < 0.001. The horizontal axes of the statistical graphs all represented subgroups, and the vertical axes represented different characteristic variables. It could be seen that except for the ratio of family income to poverty, there were differences of varying degrees in the remaining characteristics among the three groups.

[0088] Table 2 - Results of characteristic statistical analysis of participants in different CVD death subgroups:

[0089]

[0090] Finally, these statistical analysis results were integrated with the previous mediation causal inference results to construct a transformation network between different CVD death subgroups and visualize it to enhance the understanding of the potential mediation causal relationships driving subgroup transformation. The integration process was to change the red nodes in Figure 3 to different subgroups, streamline the characteristics to the above 10 important characteristics selected, and the changing trends of the red arrows and blue arrows were drawn according to the results of the differential characteristics. For example, if the age of Subgroup 2 (Subgroup1) was higher than that of Subgroup 1 (Subgroup1), a red arrow was added next to the age. If the age of Subgroup 3 (Subgroup3) was higher than that of 2, two red arrows were added. If it decreased, a blue arrow was added, and so on. The visualization effect of the embodiments of the present invention is shown in Figure 8As shown in the figure, the network contains 4 exposure factors, 6 laboratory indicators, and 3 identified CVD death subgroups. The single red arrow in the figure indicates an increase in value, two arrows indicate a further increase in value, and the blue arrow indicates a decrease in value. The results show that age can mediate the transformation between different subgroups by affecting hemoglobin, hematocrit, lymphocyte percentage, neutrophil percentage, and osmotic pressure; gender affects subgroup transformation by influencing lymphocyte percentage, osmotic pressure, and mean corpuscular volume; marital status affects the transformation between subgroups by mediating changes in lymphocyte percentage, neutrophil percentage, osmotic pressure, and mean corpuscular volume; while household income and poverty ratio mainly affect the transformation between subgroups by mediating changes in hemoglobin and hematocrit.

[0091] Specifically, refer to Figure 9 As shown in the figure, it is a schematic structural diagram of a CVD death subgroup identification device combining causal reasoning and consensus clustering according to an embodiment of the present invention, including:

[0092] A data collection module 901, which collects original feature data and performs preprocessing to obtain preprocessed feature data. The feature variables in the preprocessed feature data include exposure factors and laboratory indicators;

[0093] A causal reasoning module 902, which performs mediating causal relationship reasoning on the preprocessed feature data to obtain an initial feature set with the exposure factors and laboratory indicators included in the mediating causal relationship reasoning result;

[0094] A model construction module 903, which uses several machine learning models to perform death risk prediction according to the input feature data respectively, and then trains them using the initial feature set respectively to obtain several death risk prediction models; the SHAP value is calculated for the feature variables of all death risk prediction models using the SHAP algorithm, and the optimal model and the optimal feature set are selected using the SHAP values of the feature variables of each death risk prediction model;

[0095] A subgroup obtaining module 904, which performs consensus clustering on the death risk prediction results output by the optimal feature set and the optimal model to obtain several death subgroups;

[0096] A subgroup identification module 905, which performs statistical analysis on the feature data and survival results of each death subgroup, and combines the statistical analysis results and the mediating causal relationship reasoning results to identify the pathways of each death subgroup.

[0097] When this embodiment works, first, the potential pathway of CVD death is revealed by using mediation causal inference, providing an important reference for medical institutions to reduce the CVD mortality rate and optimize prevention strategies; secondly, combined with ML technology, the 10 most critical feature variables in the CVD death pathway are identified. These features are all common clinical indicators, effectively reducing the model complexity while ensuring the model performance and laying a good foundation for the promotion of the model in the real world; finally, the 10 clinical variables can be further used for the risk stratification of the CVD death population and reveal the potential causal mechanism of the mutual transformation between different CVD death subgroups.

[0098] The above are only the preferred embodiments of the present invention and are not intended to limit the present invention. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present invention shall be included within the protection scope of the present invention.

Claims

1. A method for identifying CVD death subgroups by combining causal inference and consensus clustering, characterized in that, It includes the following steps: Collect the original feature data and perform preprocessing to obtain preprocessed feature data. The feature variables in the preprocessed feature data include exposure factors and laboratory indicators. Perform mediation causal relationship reasoning on the preprocessed feature data, and use the exposure factors and laboratory indicators included in the mediation causal relationship reasoning results as the initial feature set. Use several machine learning models to predict the death risk according to the input feature data respectively, and then train them using the initial feature set respectively to obtain several death risk prediction models; use the SHAP algorithm to calculate the SHAP values of the feature variables of all death risk prediction models, and select the optimal model and the optimal feature set using the SHAP values of the feature variables of each death risk prediction model. Perform consensus clustering on the death risk prediction results output by the optimal feature set and the optimal model to obtain several death subgroups. Perform statistical analysis on the feature data and survival results of each death subgroup, and identify the pathways of each death subgroup in combination with the statistical analysis results and the mediation causal relationship reasoning results.

2. The method for identifying CVD death subgroups by combining causal inference and consensus clustering according to claim 1, wherein The preprocessing includes: Eliminate the feature variables with missing values exceeding 50%. Eliminate the feature data of the subjects under 20 years old. Eliminate the feature data of the subjects with missing values exceeding 50%. Classify each subject based on the survival status of each subject at different time points.

3. The method for identifying CVD death subgroups by combining causal inference and consensus clustering according to claim 2, wherein The classification of each subject based on the survival status of each subject at different time points includes: At the three-year follow-up, classify the subjects as the survival group or the death group according to whether they are alive. At the five-year follow-up, classify the subjects as the survival group or the death group according to whether they are alive. At the ten-year follow-up, classify the subjects as the survival group or the death group according to whether they are alive.

4. The method for identifying CVD death subgroups by combining causal inference and consensus clustering according to claim 1, wherein The mediation causal relationship reasoning on the preprocessed feature data includes the following steps: Use the K-nearest neighbor algorithm to fill in the preprocessed feature data. Perform an exhaustive combination of exposure factors and laboratory indicators; for each pair of exposure factors and laboratory indicators, verify them based on several conditional independence tests to obtain several P values, and use the maximum value among them as the overall P value. If the overall P value is less than 0.05, it means that the exposure factor can affect CVD death through the laboratory indicator as a mediating variable. Among them, the several conditional independence tests include: The exposure factor is significantly correlated with the death outcome. After adjusting the death outcome, the exposure factor is still significantly correlated with the mediating variable; the mediating variable refers to the laboratory indicator. After adjusting the exposure factor, the mediating variable is still significantly correlated with the death outcome. After adjusting the mediating variable, the exposure factor is not related to the death outcome.

5. The method for identifying CVD death subgroups by combining causal inference and consensus clustering according to claim 1, wherein Use the SHAP algorithm to calculate the SHAP values of the feature variables of all death risk prediction models, and select the optimal model and the optimal feature set using the SHAP values of the feature variables of each death risk prediction model. Specifically: Use the SHAP algorithm to calculate the SHAP values of the feature variables of all death risk prediction models. For each death risk prediction model, gradually reduce the number of input features based on the SHAP values of the feature variables, evaluate the prediction effect, and select the optimal model by combining the prediction effect and the number of feature variables; the set of input feature variables of the optimal model is the optimal feature set.

6. The method for identifying CVD death subgroups by combining causal inference and consensus clustering according to claim 1, wherein The optimal feature set includes 4 exposure factors and 6 laboratory indicators; the 4 exposure factors are age, household income-to-poverty ratio, gender, and marital status; the 6 laboratory indicators are lymphocyte percentage, neutrophil percentage, hemoglobin, hematocrit, mean corpuscular volume, and osmotic pressure; the optimal model is a logistic regression model with the optimal feature set as the input feature variables.

7. The method for identifying CVD death subgroups by combining causal inference and consensus clustering according to claim 1, wherein Performing consensus clustering on the death risk prediction results output by the optimal feature set and the optimal model to obtain several death subpopulations, including the following steps: Input the feature values of the final feature set into the optimal prediction model to obtain the death risk prediction results; Perform consensus clustering on the feature values and the prediction results to obtain three death subpopulations.

8. The method for identifying CVD death subgroups by combining causal inference and consensus clustering according to claim 1, characterized in that, Performing statistical analysis on the feature data and survival results of each death subpopulation, including: Survival analysis, using the Kaplan-Meier survival curve to estimate the survival probability of each death subpopulation, and using the Log-rank test to compare the differences in survival curves between groups; Descriptive analysis, representing continuous variables by median and quartiles, and representing categorical variables by frequency and percentage; Inter-group comparison, for the comparison between two groups, the Mann-Whitney U test is used for continuous variables, and the chi-square test or Fisher's exact test is used for categorical variables; for multiple comparisons between three groups, the Bonferroni correction is used to adjust the P value to control type I error; Among them, the continuous variables include age, household income-to-poverty ratio, lymphocyte percentage, neutrophil percentage, hemoglobin, hematocrit, mean corpuscular volume, and osmotic pressure; the categorical variables include gender and marital status.

9. The method for identifying CVD death subgroups by combining causal inference and consensus clustering according to claim 1, wherein Combining the statistical analysis results and the mediation causal relationship inference results to identify the pathways of each death subpopulation, and the pathways of the death subpopulation are specifically: Age affects the transformation between different subpopulations by influencing hemoglobin, hematocrit, lymphocyte percentage, neutrophil percentage, and osmotic pressure; Gender affects the transformation between different subpopulations by influencing lymphocyte percentage, osmotic pressure, and mean corpuscular volume; Marital status affects the transformation between different subpopulations by influencing the changes in lymphocyte percentage, neutrophil percentage, osmotic pressure, and mean corpuscular volume; The household income-to-poverty ratio affects the transformation between different subpopulations by influencing the changes in hemoglobin and hematocrit.

10. A CVD death subgroup identification device combining causal reasoning and consensus clustering, characterized in that, Including: A data collection module that collects the original feature data and performs preprocessing to obtain preprocessed feature data, and the feature variables in the preprocessed feature data include exposure factors and laboratory indicators; A causal inference module that performs mediation causal relationship inference on the preprocessed feature data, and uses the exposure factors and laboratory indicators included in the mediation causal relationship inference results as the initial feature set; The model construction module uses several machine learning models to respectively predict the death risk based on the input feature data, and then trains them respectively using the initial feature set to obtain several death risk prediction models; the SHAP algorithm is used to calculate the SHAP values of the feature variables of all death risk prediction models, and the optimal model and the optimal feature set are selected using the SHAP values of the feature variables of each death risk prediction model; The subpopulation acquisition module performs consensus clustering on the death risk prediction results output by the optimal feature set and the optimal model to obtain several death subpopulations; The subpopulation identification module performs statistical analysis on the feature data and survival results of each death subpopulation, and identifies the pathways of each death subpopulation in combination with the statistical analysis results and the results of mediating causal relationship reasoning.

Citation Information

Patent Citations

  • Medical risk factor analysis method for discovering senile chronic disease based on causality

    CN116597999A

  • Causal intermediary analysis method, system and device for transverse data and medium

    CN117393053A

  • Clustering method and device for people dead due to cardiovascular and cerebrovascular diseases based on DNN and consensus clustering

    CN118965051A

Cited By

  • Method for identifying ICU patient death risk and safety index range

    CN121306572A

  • A method for identifying mortality risk and safety indicator ranges in ICU patients

    CN121306572B