Cardiovascular health assessment method, electronic equipment and medium

Weighted evaluation and multi-model prediction of cardiovascular health indicators through COX proportional hazard regression and machine learning algorithms solve the problem of inaccurate existing cardiovascular health assessment methods, and achieve higher evaluation accuracy and personalized health risk prediction.

CN119964797APending Publication Date: 2025-05-09XIANGYA HOSPITAL CENT SOUTH UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510034601.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-09
Publication Date
2025-05-09

AI Technical Summary

Technical Problem

The existing cardiovascular health assessment methods are not accurate enough and fail to fully consider the impact of multiple health indicators on cardiovascular health in different populations.

Method used

The weighted evaluation of cardiovascular health indicators is carried out using COX proportional hazard regression method and machine learning algorithms (such as extreme gradient enhancement, gradient enhancement decision tree, and random forest) to establish multiple prediction models to improve the accuracy of scores.

Benefits of technology

Through weighted evaluation and multi-model prediction, the accuracy of cardiovascular health assessment is improved and the individual's cardiovascular health status and health risks can be more accurately reflected.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure BDA0005235191040000051
    Figure BDA0005235191040000051
  • Figure BDA0005235191040000061
    Figure BDA0005235191040000061
  • Figure BDA0005235191040000062
    Figure BDA0005235191040000062
Patent Text Reader

Abstract

The invention provides a cardiovascular health assessment method, electronic equipment and a medium. The method comprises the following steps: acquiring scores of cardiovascular health indexes of different people; taking the score of each cardiovascular health index as an independent variable, taking age, education degree, annual income, Deprivation index, gender and ethnic as covariants, and obtaining the weight of the score of each cardiovascular health index through a COX proportional risk regression method; and carrying out weighted average on the score of each cardiovascular health index by using the weight to obtain a cardiovascular health score. According to the method, multiple important health indexes influencing cardiovascular health are added, the cardiovascular health score is evaluated through weighting, the influence degree of different health indexes of different people on cardiovascular health is considered, and the accuracy of cardiovascular health evaluation is improved. The cardiovascular health score is used for predicting different health risks, a new perspective and data support are provided for theoretical research in related fields, development of the physical health promotion field is promoted, and health aging is promoted.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of health assessment, and in particular relates to a cardiovascular health assessment method, electronic equipment and medium. Background Art

[0002] As the population ages, the incidence of cardiovascular disease (CVD) increases, leading to higher disability and mortality rates, which has a huge impact on the public health system. In June 2022, the American Heart Association (AHA) released the latest cardiovascular health (CVH) assessment guidelines, namely Life's Essential 8 TM (LE8), for anyone aged 2 years and over.

[0003] Life's Essential 8 for Cardiovascular Health TM The content is divided into two main areas - health behaviors and health factors. Health behaviors include diet, physical activity, smoking and sleep. Health factors are body mass index, cholesterol level, blood sugar and blood pressure. The American Heart Association's Life's Essential 8 model assesses cardiovascular health through a combination of health behaviors and health factors. Life's Essential 8 TM Each component of the My Life Check tool is assessed by a built-in scoring system ranging from 0 to 100. The overall cardiovascular health score, which ranges from 0 to 100, is the average of the scores of the eight health indicators. A total score below 50 indicates "poor" cardiovascular health, 50-79 is considered "moderate" cardiovascular health, and scores of 80 and above indicate "high" cardiovascular health. A large number of studies have found that a high level of cardiovascular health, that is, a high cardiovascular health score, is not only associated with a lower risk of cardiovascular disease, but also helps reduce the incidence of a variety of diseases such as cancer, diabetes and dementia, as well as the risk of death, and prolongs life.

[0004] In addition, more and more studies have shown that mental health, drinking, and long periods of sitting are also important predictors of cardiovascular health. Among them, drinking is particularly closely related to all-cause mortality; social and structural determinants, as well as mental health and well-being, are key basic factors for individuals or communities to maintain and improve cardiovascular health. However, none of them have yet been included in the evaluation system, the evaluation indicators are imperfect, and the current cardiovascular health score is the average of multiple health indicator scores, which affects the accuracy of cardiovascular health assessment. Summary of the invention

[0005] The purpose of the present invention is to provide a cardiovascular health assessment method, electronic equipment and medium to improve the accuracy of cardiovascular health assessment in view of the deficiencies in the prior art.

[0006] In order to achieve the above object, the technical solution adopted by the present invention is:

[0007] A cardiovascular health assessment method comprises the following steps:

[0008] Obtaining scores of cardiovascular health indicators of different populations, wherein the cardiovascular health indicators include diet, exercise, smoking, sleep, body mass index, blood lipids, blood sugar, blood pressure, and drinking, and the cardiovascular health indicators also include mental health or sedentary life, wherein the different populations include people of different ages, different education levels, different annual incomes, different Townsend deprivation indexes, different genders, different nationalities, and different diseases;

[0009] The scores of each cardiovascular health index were used as independent variables, and age, education level, annual income, Townsend deprivation index, gender, and ethnicity were used as covariates. The weights of the scores of each cardiovascular health index were obtained by the COX proportional hazards regression method;

[0010] The scores of each cardiovascular health index are weighted averaged according to the weights of the scores of each cardiovascular health index to obtain the cardiovascular health score;

[0011] The expression of cardiovascular health score is as follows:

[0012] S=x1β1+…+x i β i +…+x m β m

[0013] Among them, S is the cardiovascular health score, x i is the score of the i-th cardiovascular health index, β i is the weight of the score of the ith cardiovascular health indicator, and m is the number of cardiovascular health indicators.

[0014] The present invention adds a number of important health indicators that affect cardiovascular health, and through weighted assessment of cardiovascular health scores, takes into account the degree of influence of different health indicators of different populations on cardiovascular health, thereby improving the accuracy of cardiovascular health assessment.

[0015] Furthermore, the scores of the cardiovascular health indicators of different populations are weighted averaged according to the weights of the scores of each cardiovascular health indicator to obtain the cardiovascular health scores of different populations;

[0016] Multiple prediction models for cardiovascular health scores were established using extreme gradient boosting, gradient boosted decision trees, least absolute shrinkage and selection operators, and random forests;

[0017] The scores of cardiovascular health indicators of different populations are used as the input of the prediction model to train the prediction model;

[0018] The areas under the receiver operating characteristic curves of multiple prediction models were compared, and the prediction model with the largest area under the receiver operating characteristic curve was used as the final prediction model for cardiovascular health score.

[0019] Furthermore, the accuracy, calibration slope and Brier score were used to evaluate the prediction model.

[0020] Furthermore, the SHAP algorithm was used to obtain the contribution ranking of each cardiovascular health index score to the final prediction model output.

[0021] The contribution of different cardiovascular health indicators to model prediction can be understood, and cardiovascular health indicators can be screened.

[0022] Furthermore, based on the cardiovascular health score, the relationship between the cardiovascular health score and various health conditions is evaluated.

[0023] Furthermore, if the cardiovascular health indicator includes mental health, the health status includes life expectancy, all-cause mortality, cardiovascular disease, and cancer-specific morbidity and mortality.

[0024] Further, if the cardiovascular health indicator includes sitting, the health status includes life expectancy, all-cause mortality, stroke and dementia-specific morbidity.

[0025] Cardiovascular health scores are used to predict different health status and risks, provide new perspectives and data support for theoretical research in related fields, advance theoretical development in the field of physical health promotion, and promote healthy aging.

[0026] Based on the same inventive concept, the present invention also provides an electronic device, including:

[0027] one or more processors;

[0028] A memory having one or more programs stored thereon, which, when executed by the one or more processors, enables the one or more processors to implement the steps of the cardiovascular health assessment method.

[0029] Based on the same inventive concept, the present invention also provides a computer-readable storage medium storing a computer program, which implements the steps of the cardiovascular health assessment method when executed by a processor.

[0030] Compared with the prior art, the present invention has the following beneficial effects:

[0031] The present invention adds multiple important health indicators that affect cardiovascular health, and through weighted evaluation of cardiovascular health scores, takes into account the impact of different health indicators of different populations on cardiovascular health, thereby improving the accuracy of cardiovascular health assessment. It will help to more accurately reflect the cardiovascular health status of individuals and measure the physical health level.

[0032] Cardiovascular health scores are used to predict different health risks, provide new perspectives and data support for theoretical research in related fields, advance theoretical development in the field of physical health promotion, and promote healthy aging. BRIEF DESCRIPTION OF THE DRAWINGS

[0033] Figure 1 Schematic diagram of cardiovascular health indicators including mental health;

[0034] Figure 2 A schematic diagram of cardiovascular health indicators including sedentary behavior;

[0035] Figure 3 The nonlinear relationship between cardiovascular health score and all-cause mortality, cardiovascular disease, and cancer-specific morbidity and mortality was investigated;

[0036] Figure 4 The nonlinear relationship between cardiovascular health score and all-cause mortality, stroke, and dementia incidence was investigated;

[0037] Figure 5 Comparison of the areas under the receiver operating characteristic curves of multiple prediction models for cardiovascular health indicators including mental health;

[0038] Figure 6 Comparison of the areas under the receiver operating characteristic curves of multiple prediction models for cardiovascular health indicators including sedentary behavior;

[0039] Figure 7 A ranking chart of the contribution of each cardiovascular health indicator score to cardiovascular health indicators including mental health;

[0040] Figure 8 A ranking chart of the contribution of cardiovascular health indicators including sedentary cardiovascular health indicators;

[0041] Fig. 9 A comparison chart of the prediction model of cardiovascular health indicators including mental health and the LE8 model;

[0042] Fig.10 Comparison chart of the prediction model for cardiovascular health indicators including sedentary time and the LE8 model;

[0043] Fig.11The relationship between the cardiovascular health score, which includes psychological health, and all-cause mortality, cardiovascular disease, and cancer-specific morbidity and mortality was plotted for cardiovascular health indicators;

[0044] Fig.12 A graph showing the relationship between cardiovascular health scores and life expectancy for cardiovascular health indicators including mental health;

[0045] Fig.13 To plot the relationship between cardiovascular health scores for cardiovascular health indicators including sedentary time and all-cause mortality, stroke, and dementia-specific incidence;

[0046] Fig.14 A graph showing the relationship between cardiovascular health scores and life expectancy for cardiovascular health indicators including sedentary behavior;

[0047] Fig.15 Schematic representation of the method for estimating life expectancy for multistate survival models. DETAILED DESCRIPTION

[0048] The present invention will be described in detail below in conjunction with the embodiments. It should be noted that the embodiments and features in the embodiments of the present invention can be combined with each other without conflict.

[0049] Example

[0050] The cardiovascular health assessment method of this embodiment includes the following steps:

[0051] Through surveys, the parameters of cardiovascular health indicators of different populations, such as different ages, different education levels, different annual incomes, different Townsend deprivation indexes, different genders, different nationalities, and different diseases, were obtained. The scores of each cardiovascular health indicator were obtained according to Table 2 (Reference: Lloyd-Jones DM, Allen NB, Anderson CAM, Black T, Brewer LC, Foraker RE, Grandner MA, Lavretsky H, Perak AM, Sharma G, Rosamond W; American Heart Association. Life's Essential 8: Updating and Enhancing the American Heart Association's Construct of Cardiovascular Health: A Presidential Advisory From the American Heart Association. Circulation. 2022 Jun 29:101161CIR0000000000001078.doi:10.1161 / CIR.0000000000001078.), cardiovascular health indicators include diet, exercise, smoking, sleep, body mass index, blood lipids, blood sugar, blood pressure, drinking, sedentary, mental health, etc. Figure 1 , Figure 2 .

[0052] The Townsend Deprivation Index (TDI) provides a comprehensive way to assess the degree of poor material living conditions by taking into account unemployment, car ownership, homelessness, and household overcrowding. A higher Townsend Deprivation Index value means that more people in the area or community are troubled by these problems, indicating a higher level of material deprivation and insufficient resources. By analyzing and comparing various variables between different regions or community faults and the overall economic situation, the Townsend Deprivation Index can be used to determine whether a particular group enjoys fairness and welfare protection.

[0053] Diet quality is assessed in the most recent dietary recommendations for cardiovascular health, which consider adequate intake of fruits, vegetables, whole grains, fish, shellfish, dairy products, and vegetable oils, and reduced intake of refined grains, processed meats, unprocessed meats, and sugar-sweetened beverages (Table 1).

[0054] Table 1 UK Biobank diet scoring details

[0055]

[0056]

[0057] Table 2 Scores of cardiovascular health indicators

[0058]

[0059]

[0060]

[0061] The scores of each cardiovascular health indicator were used as independent variables, all-cause mortality was used as the dependent variable, and age, sex at birth, race and ethnicity, Townsend deprivation index, education level and annual household income were used as covariates. The weights of the scores of each cardiovascular health indicator were obtained by the COX proportional hazard regression method. The scores of each cardiovascular health indicator of different populations were weighted averaged according to the weights of the scores of each cardiovascular health indicator to obtain the cardiovascular health scores of different populations.

[0062] S=x1β1+…+x i β i +…+x m β m

[0063] Among them, S is the cardiovascular health score, x i is the score of the i-th cardiovascular health index, β i is the weight of the score of the ith cardiovascular health indicator, and m is the number of cardiovascular health indicators.

[0064] In some implementations of this embodiment, the cardiovascular health score is predicted and evaluated through scores of diet, physical activity, smoking, sleep, body mass index, cholesterol level, blood sugar, blood pressure, drinking, mental health, etc.

[0065] In some other implementations of this embodiment, the cardiovascular health score is predicted and evaluated through scores of diet, physical activity, smoking, sleep, body mass index, cholesterol level, blood sugar, blood pressure, drinking, sitting, etc.

[0066] The COX proportional hazards regression model is a semiparametric regression model, which is mainly used in survival analysis, especially multivariate survival analysis, to explore the impact of multiple factors on survival time.

[0067] In COX regression, the process of adjusting covariates usually involves multiple steps, including selecting appropriate covariates, building models, estimating parameters, and interpreting results. Effectively adjusting covariates in COX regression can more accurately assess the impact of independent variables on dependent variables.

[0068] 1. Choosing Appropriate Covariates

[0069] First, the purpose and background of the study need to be clarified, and the covariates that need to be corrected should be determined based on the research questions. Covariates can be continuous numerical variables, discrete categorical variables, or level variables. These covariates are usually factors that affect the dependent variables (survival time and survival status) and need to be considered in the correction process.

[0070] 2. Constructing the COX regression model

[0071] Data preparation: Prepare the survival time variable, survival status variable, focal factor (i.e. the independent variable of the main study) and covariates that need to be adjusted.

[0072] Model construction: In statistical software (such as SPSS, SAS, etc.), select the COX regression analysis method.

[0073] Feed the survival time variables into the time frame and the survival state variables into the state frame.

[0074] Enter the focal factor and the covariates that need to be adjusted into the Covariates box.

[0075] 3. Parameter Estimation

[0076] After building the model, parameter estimation is required. This usually involves the following steps:

[0077] Select the estimation method: According to the data characteristics and research needs, select the appropriate estimation method, such as maximum likelihood estimation.

[0078] Perform stepwise regression: You can use stepwise regression to screen variables to find the covariates that contribute most to the model.

[0079] Output results: Statistical software will output a series of results, including model parameter estimates, standard errors, confidence intervals, P values, etc.

[0080] The P value refers to the probability of an event that a statistical summary (such as the difference between the mean values ​​of two groups of samples) is the same as or even greater than the actual observed data in a probability model. The P value is the probability that the null hypothesis is true or more serious. If the P value is smaller than the selected significance level (0.05 or 0.01), the null hypothesis will be rejected and unacceptable, but this does not directly indicate that the original hypothesis is correct. The P value is a random variable that follows a normal distribution. In actual use, there is uncertainty due to various factors such as samples, and the results produced may be controversial.

[0081] IV. Interpretation of Results

[0082] After obtaining the parameter estimation results, it is necessary to interpret the results. This usually involves the following aspects:

[0083] Overall test: Check whether the model as a whole is statistically significant, usually through a likelihood ratio test or a score test.

[0084] Variable screening: Based on indicators such as P value and confidence interval, determine which covariates have significant contributions to the model.

[0085] Relative risk (HR): Explain the relative risk (HR) of each covariate, that is, the degree of change in survival risk when the covariate changes by one unit.

[0086] Survival curve: Survival curves at different covariate levels can be drawn to intuitively show the impact of covariates on survival time.

[0087] The following is the process of obtaining weights using the COX proportional hazards regression method:

[0088] Data preparation: First, you need to collect and organize relevant data, including survival time and event outcome variables (such as death, disease occurrence, etc.), which usually come from medical research or other types of follow-up studies. And independent variables that may affect survival time (such as diet, exercise, smoking, sleep, body mass index, blood lipids, blood sugar, blood pressure, drinking, sedentary or mental health).

[0089] Model construction: Use the COX regression function in statistical software (such as R, SAS, SPSS, etc.) to build a COX proportional hazard regression model based on the collected data. In R language, the coxph function can be used to fit the model, where Surv (time, status) represents the survival time and event outcome variables, and the right side is the independent variable of the model.

[0090] h(t)=h0(t)exp(β1x' 1+ β2x' 2+…+ β p x' p )

[0091] where h(t) is a function with independent variables x'1, x'2, ... x' p The risk function of an individual at time t; x'1, x'2, ... x' p is the independent variable; β'1, β'2, ... β' p is the partial regression coefficient of the independent variable; h0(t) is x'1=x'2=…=x' p = 0, the risk function at time t is called the baseline risk function. The COX model does not make any assumptions about the content of the first factor h0(t), and the second factor has a parametric model. All COX models are actually semi-parametric models.

[0092] Parameter estimation: The parameters in the COX model include the partial regression coefficient of the independent variable and the baseline risk function. The partial regression coefficient reflects the influence of the independent variable on the risk function and can be estimated by the maximum likelihood estimation method. The baseline risk function is the risk function when all independent variables are 0. It usually does not need to be estimated directly, but is obtained indirectly through methods such as residuals.

[0093] The likelihood probability is as follows:

[0094]

[0095] Among them, X i is the i-th instance, representing the independent variable vector; β′ is the partial regression coefficient vector, T i and T j Indicates the time when the event occurred. represents multiplying the probabilities of all instances where the event of interest occurred.

[0096] Select β′ when L(β′) reaches the maximum value, which is the weight, that is, argmax β′ {L(β′)}.

[0097] Restricted cubic spline fitting was performed on cardiovascular health scores and all-cause mortality, cardiovascular disease, cancer-specific morbidity and mortality, stroke and dementia incidence, etc. to explore the nonlinear relationship between cardiovascular health scores and them. According to the results, cardiovascular health scores were divided into three levels: low level CVH (cardiovascular health score less than 60 points), moderate level CVH (cardiovascular health score 60-79 points) and high level CVH (cardiovascular health score 80 points and above).

[0098] Restricted cubic splines mainly provide clues for determining the cutoff values ​​for dividing cardiovascular health scores into low, medium and high levels.

[0099] Restricted Cubic Splines (RCS) is a nonlinear regression method used to capture the complex relationship between independent variables X' and dependent variables Y' in a regression model. The following is a detailed introduction to restricted cubic splines:

[0100] Knots: Knots are specific locations of the spline function within the range of the independent variable X', usually selected based on quantiles or other criteria. These nodes divide the range of the independent variable X' into several intervals, and a polynomial is used to fit each interval.

[0101] Spline Function: A spline function is a piecewise polynomial function that fits the data with a low-order polynomial (usually a cubic polynomial) on each segment. At each node, the spline function is connected, but has continuous derivatives at the connection points to ensure smooth transitions.

[0102] Restricted: In order to avoid unreasonable extrapolation results of the spline function outside the boundaries of the data range, RCS imposes restrictions on the parts outside the nodes, usually by setting the slope of the external spline to zero. This restrictive constraint makes the fitting behavior of the spline at the endpoints subject to certain restrictions to maintain the smoothness and stability of the curve.

[0103] RCS essentially selects the location and number of nodes and fits the spline function RCS(X′) so that the continuous variable X′ presents a smooth curve in the entire value range.

[0104] After RCS transforms the independent variable X' through the spline function RCS(X'), it selects the appropriate link function according to the distribution type of the dependent variable Y', and then fits the model g(Y') = constant term + RCS(X') + other independent variables, where g is the link function. The spline function RCS(X') includes a linear term X' and K-2 cubic terms (S), that is, RCS(X') = β"0X'+β"1S1+…+β" (k-2) S (k-2) , the complete formula is shown below. Ci(x′) is the cubic component falling in the i-th node, and g is the link function.

[0105]

[0106] The location of the node has little effect on the fitting of the spline function and is generally selected according to the percentile of the continuous variable. The number of nodes has a greater impact on the spline function. The number of nodes determines the shape of the curve. When the number of nodes is 2, the fitted curve is a straight line. Studies have shown that the spline function fits better when the number of nodes is 3 to 5, and K = 4 is generally recommended. Harrell (2010) pointed out that "for many data sets, K = 4 provides adequate fit of the model and is a good compromise between flexibility and position loss caused by overfitting small samples" (Harrell, FE (2010) Regression Modeling Strategies: with applications to linear models, logistic regression, and survival analysis. Springer-Verlag New York, Inc. New York, USA.). By comparing the AIC values ​​corresponding to different numbers of nodes (K = 3, 4, 5), the node with the smallest AIC value (K = 4) is selected to fit the model, such as Figure 3 , Figure 4 .

[0107] In order to select the optimal machine learning algorithm for cardiovascular health assessment model, four machine learning methods were evaluated: extreme gradient boosting (XGBT), gradient boosting decision tree (GBDT), least absolute shrinkage and selection operator (Lasso), random forest (RandomForest). Each algorithm has unique advantages in feature selection, interpretability and handling complex relationships in data.

[0108] Multiple prediction models of cardiovascular health scores were established using gradient boosting decision tree, gradient boosting decision tree, least absolute shrinkage and selection operator, and random forest. The appropriate time cutoff was selected to divide the external validation set. For the data collected before this time cutoff, the scores of various cardiovascular health indicators of different populations and 70% of the cardiovascular health scores of different populations were randomly selected as training sets, and the remaining 30% were test sets. The training set data was used for prediction model training, and the test set was used for internal validation.

[0109] A network search strategy (Grid Search) is used to obtain the best hyperparameter configuration for each machine learning algorithm. Grid search is a systematic hyperparameter tuning method that finds the best hyperparameter combination by exhaustively searching a predefined hyperparameter space. Specifically, grid search lists all possible hyperparameter combinations, and then trains and evaluates the model for each combination; sets model performance evaluation criteria, using accuracy, calibration slope, and Brier score as evaluation criteria. Finally, the combination that performs best on the validation set is selected. Suppose there are two hyperparameters α and γ, each with three possible values. Grid search tries all possible (α, γ) combinations. In this way, it is guaranteed to find the optimal combination within a given hyperparameter space.

[0110] Accuracy is one of the commonly used indicators to evaluate the performance of classification models, which indicates the proportion of samples correctly predicted by the classification model to all samples.

[0111]

[0112] Among them, TP: True positive examples, the number of samples that are actually positive and predicted to be positive; TN: True negative examples, the number of samples that are actually negative and predicted to be negative; FP: False positive examples, the number of samples that are actually negative but predicted to be positive; FN: False negative examples, the number of samples that are actually positive but predicted to be negative; n: the number of samples.

[0113] Calibration slope, also known as the slope of the calibration curve, refers to the slope of a straight line used in statistics to describe the relationship between predicted probabilities and actual outcomes. In machine learning and statistics, a calibration curve is a graphical tool used to represent the relationship between predicted probabilities and actual probabilities. The closer the calibration slope is to 1, the better the consistency between the model's predictions and actual observations.

[0114] Calculation method: The calibration slope is usually obtained through linear regression analysis, and its calculation formula is slope = covariance (predicted value, actual value) / variance (predicted value). In actual calculation, it is necessary to collect the predicted probability of the model and the corresponding actual results, and then use these data points to fit a best fit line (calibration curve), and the slope of this line is the calibration slope.

[0115] The Brier score is usually defined as the average squared error between the predicted probability and the actual label. For a binary classification problem, given a sample point x, its true label is y, and the predicted probability is p(y=1|x), the Brier score is BrierScore=∑(yp(y=1|x)) 2The smaller the value of the Brier score, the more accurate the model is. Generally speaking, a Brier score of 0 means that the model predictions are 100% correct. Brier scores between 0-0.1 are considered excellent, while those between 0.1 and 0.25 are generally considered good. If the Brier score is higher than 0.25, it indicates that the model's predictions are not very accurate and need to be improved.

[0116] Comparison of the areas under the receiver operating characteristic curves of prediction models for multiple cardiovascular health scores, Figure 5 Comparison of the areas under the receiver operating characteristic curves of multiple prediction models for cardiovascular health indicators including mental health. Figure 6 The area under the receiver operating characteristic curve of multiple prediction models for cardiovascular health indicators including sedentary life is compared. The prediction model with the largest area under the receiver operating characteristic curve is used as the final prediction model for cardiovascular health score.

[0117] The receiver operating characteristic (ROC) curve is a classification performance indicator. The ROC curve is also called the sensitivity curve. The closer the ROC curve is to the diagonal line, the lower the accuracy of the model.

[0118] The X-axis of the receiver operating characteristic curve (1-specificity): True Positive Rate (TPR), also known as Sensitivity, is calculated as: TPR = TP / (TP+FN), where TP is the number of true positives and FN is the number of false negatives. It represents the proportion of positive samples that are correctly classified as positive.

[0119] The Y axis of the receiver operating characteristic curve (sensitivity): False Positive Rate (FPR), also known as 1-Specificity, is calculated as: FPR = FP / (FP+TN), where FP is the number of false positives and TN is the number of true negatives. It indicates the proportion of negative samples that are misclassified as positive samples.

[0120] Draw the ROC curve: Take the FPR and TPR values ​​corresponding to each threshold as coordinate points, draw these points in the coordinate system, and connect these points to form the ROC curve. Usually, the ROC curve is a curve from the lower left corner to the upper right corner.

[0121] The area under the receiver operating characteristic curve (AUC) is defined as the area under the ROC curve (ROC integral), which represents the area between the ROC curve and the y=x line of the model under all possible classification thresholds. Randomly select a positive sample and a negative sample, and the probability that the value of the positive sample is higher than that of the negative sample is the AUC value. The larger the AUC value (area), the better the model performance.

[0122] Calculate AUC: Calculate the area under the ROC curve, that is, the area under the curve between the true positive rate and the false positive rate. The AUC value ranges from 0 to 1. AUC equals 1, that is, all patients in the test subjects are positive, and all non-patients are negative. AUC equals 0.5, which means that the performance of the model is the same as random guessing, that is, it has no classification ability. This usually means that the output of the model is random or its performance is very poor. AUC between 0.5-1 indicates that the model has a certain classification ability. The closer the AUC is to 1, the better the model performance.

[0123] The scores of each cardiovascular health index are input into the final prediction model of the cardiovascular health score to obtain the cardiovascular health score.

[0124] Based on the final prediction model of cardiovascular health score, the SHAP local interpretation algorithm is used to obtain the contribution ranking of different population parameters and scores of various cardiovascular health indicators to the final prediction model output (model prediction value), such as Figure 7 , Figure 8 shown. Figure 7 Among them, age, smoking score and income are the most predictive features in the machine learning model, followed by sex, blood pressure score, alcohol score, glucose score, physical activity score, Townsend deprivation index (TDI), sleep score, BMI score, mental health score, blood lipid score, healthy diet score and education level. By obtaining the contribution ranking of different population parameters and scores of various cardiovascular health indicators to model prediction, we can understand the contribution of different population parameters and different cardiovascular health indicators to model prediction, and can screen cardiovascular health indicators.

[0125] The SHAP (SHapley Additive exPlanations) algorithm is a method for explaining machine learning models. The mathematical principle of SHAP is based on the Shapley value in game theory, which is used to solve the problem of allocating the contribution of participants to the total revenue in a cooperative game. In the field of machine learning, SHAP regards the machine learning model as a cooperative game and each feature as a cooperative participant.

[0126] The SHAP algorithm treats the contribution of each feature value to the model output as a "fair" distribution, ensuring that the contribution of each feature value is its due share. The core idea of ​​the SHAP algorithm is to decompose the predicted value of the model output into the sum of the contributions of each feature. For a given prediction, it calculates the contribution of each feature value to the prediction result by considering all permutations and combinations of feature values. This process is based on the following two principles:

[0127] Fairness: The contribution of each feature value is based on its actual impact on the model output, ensuring that the contribution of each feature value is fair.

[0128] Local independence: When calculating the contribution of an eigenvalue, it is assumed that other eigenvalues ​​are independent, which simplifies the calculation process.

[0129] The advantages of the SHAP algorithm include:

[0130] Fairness: Ensuring that the contribution of each feature value is fair helps understand the decision-making process of the model.

[0131] Model-independence: It can be used to explain any machine learning model, including deep learning models.

[0132] Easy to understand: SHAP values ​​provide an intuitive way to understand the impact of features on prediction results.

[0133] The SHAP method is as follows:

[0134] First, for each prediction sample, the average influence estimate of all features (i.e., the mean of the feature for all samples) is subtracted from the model prediction value to obtain the marginal contribution of each feature to the prediction;

[0135] Then, the Shapley value of each feature is calculated based on the marginal contribution of each feature and the number of occurrences of the feature;

[0136] Finally, the Shapley values ​​of each feature are added together to obtain the SHAP value of the sample.

[0137] The expression for the Shapley value is as follows:

[0138]

[0139] Among them, φ i represents the Shapley value of the score i of the cardiovascular health index, M is the number of scores of the cardiovascular health index, N is the score set of the cardiovascular health index, S is the score subset of the cardiovascular health index, and f(.) represents the output of the model.

[0140] like Fig. 9 , Fig.10 ,Comparing the cardiovascular health prediction model (LE10) with the LE8 method, the results of the ROC curve, clinical decision curve and calibration curve of the cardiovascular health prediction model of this embodiment are better than those of the LE8 method. Fig. 9 , Fig.10 In the figure, from top to bottom are the ROC curve, clinical decision curve, and calibration curve.

[0141] The clinical decision curve is drawn as follows:

[0142] a. Determine the prediction model and decision threshold: First, you need to determine the range of prediction models and decision thresholds to be evaluated. The decision threshold refers to the boundary used to convert the predicted probability into a class label.

[0143] b. Calculate the true positive rate and false positive rate: For each decision threshold, calculate the true positive rate (TPR) and false positive rate (FPR). The true positive rate indicates the proportion of all positive samples that are correctly classified as positive, and the false positive rate indicates the proportion of all negative samples that are incorrectly classified as positive.

[0144] c. Calculate the net benefit: The net benefit is the difference between the true positive rate and the false positive rate, or the performance measure of the classifier at a specific threshold. By calculating the net benefit at different decision thresholds, a net benefit curve can be drawn.

[0145] d. Draw a decision curve: Draw a decision curve with the sensitivity and specificity of the prediction model as the horizontal axis and the benefit (or net benefit) as the vertical axis. The decision curve shows the benefits of using the prediction model to make decisions at different decision thresholds.

[0146] e. Comparing different models: The shape of the decision curve can reflect the performance changes of the prediction model at different decision thresholds. Generally speaking, the steeper the curve, the more significant the performance change of the model at a specific threshold. The position of the decision curve can reflect the overall performance of the model. If one decision curve is closer to the top relative to another curve, it usually means that the model has higher benefits or better decision performance at different decision thresholds.

[0147] Based on the cardiovascular health score, the relationship between the cardiovascular health score and various health statuses was explored. If the cardiovascular health index includes mental health, the health status includes life expectancy, all-cause mortality, cardiovascular disease, cancer-specific morbidity and mortality, such as Fig.11 , Fig.12 If cardiovascular health indicators include sedentary behavior, health status includes life expectancy, all-cause mortality, stroke, and dementia-specific incidence, such as Fig.13 , Fig.14 The higher the cardiovascular health score, the lower the risk of all-cause mortality and the longer the life expectancy.

[0148] In order to evaluate the relationship between CVH levels and life expectancy and calculate life expectancy under different health states, a population-based multi-state life table method (MSLT) was applied, which is based on the theory of continuous-time multi-state survival models (including absorbent death states). Multi-state survival models are used to describe, understand and predict health-related processes over time. By establishing a multi-state life table, the Markov chain of individuals in different states is systematically constructed and described. Markov chain is a mathematical statistical model that assumes that the state of a system at a certain moment in the future is only related to the current state, but not to the past state. In order to establish the model, a finite number of health states are defined, each of which represents a specific situation or attribute of the system or individual at a specific moment, and distinguishes the potential transitions between these states. Next, the transition probabilities between states are analyzed. These probabilities reflect the probability of an individual transitioning from the current state to other states. In the Markov chain model, the transition probabilities between states constitute a transition probability matrix, which describes in detail the laws and trends of individuals transitioning between different states. With this transition probability matrix, the transition between states can be predicted. Specifically, the life expectancy of reaching a target state after a series of state transitions from an initial state is calculated. The life expectancy in a specific state is defined as the expected number of remaining years in that state, conditioned on the current age.

[0149] For the elderly population, the three-state disease-death model is defined by the health state, the disease state, and the death state. For individuals of a certain age, two remaining life expectancies are distinguished, namely the expected remaining time in the health state and the expected remaining time in the disease state. The sum of these two expected remaining times is the total remaining life expectancy.

[0150] Fig.15The method of estimating life expectancy for a multi-state survival model is shown. The model has three states: disease-free (state 1), disease development (state 2), and death (state 3). The possible transition directions are (1) from the absence of cardiovascular disease and cancer to the development of cardiovascular disease or cancer; (2) from the absence of cardiovascular disease and cancer to death without the development of cardiovascular disease or cancer; (3) from the development of cardiovascular disease or cancer to death from all causes. In this model, transitions from state 2 back to state 1 are not allowed, and only the first entry into a state is considered.

[0151] The present invention adds multiple important health indicators that affect cardiovascular health, uses machine learning methods to evaluate cardiovascular health, takes into account the impact of different health indicators of different populations on cardiovascular health, and improves the accuracy of cardiovascular health evaluation. It will help to more accurately reflect the cardiovascular health status of individuals and measure their physical health level.

[0152] Cardiovascular fitness scores are used to predict different health outcomes to promote healthy aging.

[0153] By updating and improving the cardiovascular health assessment model, a more comprehensive and accurate cardiovascular health assessment tool is provided, the theoretical basis for disease prevention and non-medical intervention is improved, and the concept of integration of sports and medicine is deepened.

[0154] By exploring the relationship between cardiovascular health levels and diseases, the correlation between cardiovascular health and specific diseases has been clarified. The disease-free life expectancy of people with different cardiovascular health levels has been predicted, providing new perspectives and data support for theoretical research in related fields and promoting theoretical development in the field of physical health promotion.

[0155] Another embodiment of the present invention provides an electronic device, including:

[0156] one or more processors;

[0157] A memory stores one or more programs, which, when executed by one or more processors, enable the one or more processors to implement the steps of the cardiovascular health assessment method.

[0158] In some implementations, the memory may be a high-speed random access memory (RAM), and may also include a non-volatile memory, such as at least one disk memory.

[0159] In some other implementations, the processor may be a central processing unit (CPU), a digital signal processor (DSP), or other general-purpose processors of various types, which are not limited herein.

[0160] Another embodiment of the present invention provides a computer-readable storage medium storing a computer program, which implements the steps of the cardiovascular health assessment method when executed by a processor.

[0161] The contents explained in the above embodiments should be understood as these embodiments are only used to more clearly illustrate the present invention, and are not used to limit the scope of the present invention. After reading the present invention, various equivalent forms of modifications to the present invention by those skilled in the art all fall within the scope defined by the claims attached to this application.

Claims

1. A cardiovascular health assessment method, characterized in that: The process includes: Obtaining scores of cardiovascular health indicators of different populations, wherein the cardiovascular health indicators include diet, exercise, smoking, sleep, body mass index, blood lipids, blood sugar, blood pressure, and drinking, and the cardiovascular health indicators also include mental health or sedentary life, wherein the different populations include people of different ages, different education levels, different annual incomes, different Townsend deprivation indexes, different genders, different nationalities, and different diseases; The scores of each cardiovascular health index were used as independent variables, and age, education level, annual income, Townsend deprivation index, gender, and ethnicity were used as covariates. The weights of the scores of each cardiovascular health index were obtained by the COX proportional hazards regression method; The scores of each cardiovascular health index are weighted averaged according to the weights of the scores of each cardiovascular health index to obtain the cardiovascular health score; The expression of cardiovascular health score is as follows: S=x1β1+…+x i b i +…+x m b m Among them, S is the cardiovascular health score, x i is the score of the i-th cardiovascular health index, β i is the weight of the score of the ith cardiovascular health indicator, and m is the number of cardiovascular health indicators.

2. The cardiovascular health assessment method according to claim 1, characterized in that: The scores of cardiovascular health indicators of different populations were weighted averaged according to the weights of the scores of each cardiovascular health indicator to obtain the cardiovascular health scores of different populations; Multiple prediction models for cardiovascular health scores were established using extreme gradient boosting, gradient boosted decision trees, least absolute shrinkage and selection operators, and random forests; The scores of cardiovascular health indicators of different populations are used as the input of the prediction model to train the prediction model; The areas under the receiver operating characteristic curves of multiple prediction models were compared, and the prediction model with the largest area under the receiver operating characteristic curve was used as the final prediction model for cardiovascular health score.

3. The cardiovascular health assessment method according to claim 2, characterized in that: The prediction model was evaluated using accuracy, calibration slope and Brier score.

4. The cardiovascular health assessment method according to claim 2, characterized in that: The SHAP algorithm was used to obtain the ranking of the contribution of each cardiovascular health index score to the final prediction model output.

5. The cardiovascular health assessment method according to claim 1, characterized in that: Based on the cardiovascular health score, the relationship between the cardiovascular health score and various health status was evaluated.

6. The cardiovascular health assessment method according to claim 5, characterized in that: If the cardiovascular health indicator includes mental health, the health status includes life expectancy, all-cause mortality, cardiovascular disease, and cancer-specific morbidity and mortality.

7. The cardiovascular health assessment method according to claim 5, characterized in that: If the cardiovascular health indicator includes sitting, the health status includes life expectancy, all-cause mortality, stroke, and dementia-specific incidence.

8. An electronic device, characterized in that: include: one or more processors; A memory having one or more programs stored thereon, which, when the one or more programs are executed by the one or more processors, enables the one or more processors to implement the steps of the method according to any one of claims 1 to 7.

9. A computer-readable storage medium, characterized in that: The computer program is stored therein, and when the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 7 are implemented.