CVD risk assessment tool based on wearable device data and machine learning algorithm
By combining wearable devices and XGBoost algorithm CVD risk assessment tools, the real-time and personalized problems of vascular disease risk assessment in existing technology are solved, dynamic assessment and timely early warning of individual cardiovascular health status are achieved, and the accuracy of risk prediction and personalized management capabilities of cardiovascular disease are improved.
Patent Information
- Application Number
- CN202510251231.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-04
- Publication Date
- 2025-07-04
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
The existing cardiovascular disease risk assessment methods mainly rely on static clinical data, which are difficult to reflect the dynamic changes in individual health status, cannot achieve real-time disease risk monitoring, and have limitations when applied outside the hospital, especially in telemedicine and individualized interventions.
Combining wearable device data and XGBoost machine learning algorithms, a cardiovascular disease risk assessment tool is built, traditional risk factors and health measurement indicators are obtained through the data acquisition module, and features are screened using LASSO, random forests and Logistic regression models, XGBoost model is constructed to predict individual CVD risks, and feature contribution is evaluated through SHAP values to provide personalized risk warnings.
It has achieved a dynamic assessment of individual cardiovascular health status, provided timely and accurate risk warnings, improved the accuracy and applicability of cardiovascular disease risk prediction, promoted the timeliness of individual health management and public health awareness, and promoted the integrated development of medical services.
Smart Images

Figure CN120260902A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the fields of public health and preventive medicine, specifically relates to the field of cardiovascular disease risk assessment, and particularly relates to a CVD risk assessment tool based on wearable device data and machine learning algorithms. Background Art
[0002] With the aging of the population and the change of lifestyle, cardiovascular disease (CVD) has become an important disease affecting public health. Its high incidence and mortality make effective early risk assessment particularly urgent.
[0003] At present, the research on the prediction of cardiovascular disease risk has been gradually developing. The earliest cardiovascular disease risk prediction model originated from the Framingham Heart Study (FHS) in the United States. Subsequently, many 10-year cardiovascular disease risk prediction models have been developed in European and American countries, such as the Systematic Coronary Risk Estimation (SCORE) model in Europe, the QRISK model in the United Kingdom, and the Pooled Cohort Equations (PCE) model in the United States. The prediction tools have been continuously optimized for specific populations or cardiovascular outcomes. In China, researchers have established multiple cardiovascular disease risk prediction models applicable to Chinese adults, including the cardiovascular disease incidence risk model based on the "Chinese Multi-provincial Cohort Study (CMCS)" by our team (Wang Wei, Zhao Dong, Liu Jing, et al. Prospective study on risk factors and incidence risk prediction models of cardiovascular diseases in Chinese people aged 35-64 years [J]. Chinese Journal of Cardiology, 2003, 31(12): 902-908) and the coronary heart disease prediction model (Liu J, Hong Y, Sr DRB, et al. Predictive value for the Chinese population of the Framingham CHD risk assessment tool compared with the Chinese Multi-Provincial Cohort Study [J]. JAMA, 2004, 291(21): 2591-2599), the ischemic cardiovascular disease incidence risk model based on the "USA–People’s Republic of China Collaborative Study of Cardiovascular and Cardiopulmonary Epidemiology (USA-PRC)" ("Research group on the comprehensive risk assessment and intervention program for coronary heart disease and stroke". Development and research on the assessment method and simple assessment tool for the incidence risk of ischemic cardiovascular diseases in Chinese people [J]. Chinese Journal of Cardiology, 2003, 31(12): 893-901), and the atherosclerotic cardiovascular disease prediction model based on the "Prediction for ASCVD Risk in China (China-PAR)" (Yang X, Li J, Hu D, et al.Predicting the 10-year risks of atherosclerotic cardiovascular disease in Chinese population: the China-PAR Project (Prediction for ASCVD Risk in China)[J]. Circulation, 2016, 134(19): 1430-1440), and the cardiovascular event prediction model based on the China Kadoorie Biobank (CKB) (Yang S, Han Y, Yu C, et al. Development of a model to predict 10-year risk of ischemic and hemorrhagic stroke and ischemic heart disease using the China Kadoorie Biobank[J]. Neurology, 2022, 98(23): e2307-2317). These studies provide a scientific basis for the prevention, control and management of cardiovascular diseases in China.
[0004] Traditional CVD risk assessment methods mainly rely on static clinical data (such as age, blood pressure, cholesterol, etc.). These data can provide certain disease risk information, but it is difficult to reflect the changes and early signs of an individual's health status.
[0005] Chinese Patent CN 110428901A (Title: Stroke Incidence Risk Prediction System and Application) is based on the China-PAR cohort study and provides a prediction system for evaluating the stroke incidence risk of individuals, especially suitable for Chinese adults. It can accurately evaluate the 10-year and lifetime stroke incidence risks of individuals. This model can identify high-risk individuals, but some indicators rely on in-hospital laboratory tests, making it difficult to achieve real-time disease risk monitoring. At the same time, there are also certain challenges in the promotion of out-of-hospital disease prevention and control and individual compliance.
[0006] Chinese Patent CN 112331362A (Title: A Method for Predicting the Incidence Risk of Cardiovascular Diseases (CVD)) is a CVD 10-year incidence risk assessment formula and a simple risk stratification tool that can simultaneously evaluate the total CVD incidence risk including ASCVD and hemorrhagic stroke. However, its evaluation is based on laboratory indicators, and there are certain limitations in long-term health management and out-of-hospital application of individualized interventions.
[0007] Chinese patent CN 113764105A (name: A method for predicting cardiovascular data in the middle-aged and elderly) combines pathological information with a machine learning algorithm to determine the risk of cardiovascular disease in the middle-aged and elderly, and builds a Spark big data platform based on memory-based elastic distributed data sets to protect the health of the elderly and reduce the workload of community doctors. However, this patent relies on clinical information and is difficult to implement in the general population outside the hospital. In addition, this patent also cannot achieve dynamic monitoring of cardiovascular risks in the general population, and its application in long-term health management and individualized intervention is limited.
[0008] The prediction factors used in the prediction model constructed in Chinese patent CN 114783606 A (name: A method for predicting the risk of cardiovascular disease that is easy to promote and apply) are all information routinely collected from the health records of Chinese residents. The prediction model is constructed based on an ultra-large-scale CKB natural population cohort to meet the needs of simplicity and universality when applied in different regions of China. However, this patent uses a more traditional prediction model and does not involve new prediction models such as machine learning. In addition, this patent does not combine wearable device data and cannot realize dynamic monitoring of cardiovascular risks in the general population.
[0009] With the rapid development of mobile health technology, the use of wearable devices related to cardiovascular diseases has become more and more popular. These devices can collect health data such as steps, heart rate, sleep quality, etc. in real time, giving risk assessment models the ability to dynamically monitor, thereby providing individuals with more accurate and timely health management recommendations.
[0010] At the same time, with the rapid advancement of data science, machine learning has gradually been applied to cardiovascular disease risk prediction. Machine learning methods can extract complex patterns from massive multi-dimensional data, improving the accuracy and flexibility of the model.
[0011] Chinese patent CN 117877732 A (name: A dynamic risk assessment method and system based on smart wearable devices) conducts dynamic cardiovascular risk assessment and prediction based on user multimodal big data and provides personalized health behavior intervention suggestions. It constructs a cardiovascular risk assessment and prediction model to assess and predict the user's risk probability and risk level of cardiovascular and cerebrovascular diseases in the next 3 months, 1 year, 3 years, 5 years, and 10 years. However, the assessment of risk probability does not involve a machine learning model, and does not reflect the contribution of individual characteristics to the risk obtained by the model.
[0012] The above models mainly rely on static data, making it difficult to capture the dynamic changes in an individual's health status, restricting the real-time monitoring of cardiovascular risks, and making it difficult to provide more accurate and timely risk assessments. In addition, these models rely on laboratory test data, which is complex to obtain, has a time delay, and is difficult to continuously monitor in daily life, especially in the aspects of telemedicine or out-of-hospital individual management. Finally, since some data of these models need to be sourced from hospital visits, it is difficult to achieve personalized health guidance, and their application is limited in long-term out-of-hospital health management and individualized intervention. From the perspective of compliance, it is also difficult to provide individuals with a continuous and effective health management plan.
[0013] There is an urgent need to develop new cardiovascular disease risk assessment tools. Based on traditional risk factors, these tools should fully integrate health measurement indicators from wearable devices, provide more accurate and personalized cardiovascular risk assessments using dynamic health data, and capture potential risk signals to provide a more efficient solution for long-term health management. Summary of the Invention
[0014] To solve the above technical problems, the present invention provides a CVD risk assessment tool that combines wearable device data and machine learning algorithms. It can real-time monitor an individual's multi-dimensional health data (such as heart rate, number of steps, sleep duration, etc.), build a model based on the XGBoost machine learning algorithm, and accurately predict the risk probability of cardiovascular disease. It can dynamically evaluate an individual's cardiovascular health status, provide personalized and timely risk warnings, and is applicable to scenarios such as daily health management and remote monitoring, providing a new solution for the early intervention and personalized medicine of cardiovascular diseases.
[0015] On the one hand, the present invention provides a cardiovascular disease risk assessment system based on wearable device data, including:
[0016] A data collection module that obtains traditional cardiovascular risk factors through a cardiovascular risk assessment scale and obtains monitoring data through a wearable device;
[0017] A feature screening module that determines the range of candidate predictors, and then extracts features and screens predictors for the population wearing wearable devices by jointly using the Least Absolute Shrinkage and Selection Operator (LASSO), Random Forest (RF), and Logistic regression models;
[0018] A model construction module that builds a model based on the XGBoost machine learning algorithm;
[0019] A risk diagnosis module that outputs the probability of an individual having a medium to high CVD risk within 10 years based on the feature data obtained in real-time from wearable device data.
[0020] In a preferred embodiment, the cardiovascular disease risk assessment system of the present invention further includes: a feature importance evaluation module that evaluates the contribution of features using SHAP values. For a given model f(x) and feature set N = {1, 2, …, n}, the SHAP value of feature i is defined as:
[0021]
[0022] where S represents the subset of features excluding feature i, f(S) is the output of the model trained only based on S, and φ i represents the SHAP value of feature i.
[0023] In a preferred embodiment, the feature importance evaluation module of the cardiovascular disease risk assessment system of the present invention visually reflects the cumulative contribution of individual features to the model output using the waterfall plot of XGBoost.
[0024] In a preferred embodiment, the data collection module of the cardiovascular disease risk assessment system of the present invention also collects demographic data, including age and gender.
[0025] In a preferred embodiment, in the feature screening module of the cardiovascular disease risk assessment system of the present invention, the extracted features include the number of steps, calorie consumption, duration of moderate-to-high-intensity exercise, sleep duration, maximum stress value, average stress value, maximum heart rate, and resting heart rate.
[0026] On the other hand, the present invention provides a method for assessing the risk of cardiovascular disease based on wearable device data, including the following steps:
[0027] S1) Data collection: Obtain traditional cardiovascular risk factors through a cardiovascular risk assessment scale and obtain monitoring data through a wearable device;
[0028] S2) Feature screening: Determine the range of candidate predictors, and then extract features and screen predictors for the population wearing the wearable device by jointly using the LASSO, RF, and Logistic regression models;
[0029] S3) Model construction: Based on the XGBoost machine learning algorithm;
[0030] S4) Risk diagnosis: Based on the feature data obtained in real time from the wearable device data, output the probability that the 10-year CVD risk of the evaluated individual is medium to high risk.
[0031] In a preferred embodiment, the method for assessing the risk of cardiovascular disease of the present invention further includes: S5) Feature importance evaluation: Evaluate the contribution of features using SHAP values. For a given model f(x) and feature set N = {1, 2, …, n}, the SHAP value of feature i is defined as:
[0032]
[0033] Among them, S represents the feature subset that does not include feature i, f(S) is the output of the model trained only based on S, and φ i represents the SHAP value of feature i.
[0034] In a preferred embodiment, in the S5) feature importance evaluation of the cardiovascular disease risk assessment method of the present invention, the cumulative contribution of individual features to the model output is visually reflected by using the waterfall chart of XGBoost.
[0035] In a preferred embodiment, in the S1) data collection of the cardiovascular disease risk assessment method of the present invention, demographic data including age and gender are also collected.
[0036] In a preferred embodiment, in the S2) feature screening of the cardiovascular disease risk assessment method of the present invention, the extracted features include the number of steps, calorie consumption, duration of moderate-to-high-intensity exercise, sleep duration, maximum stress value, average stress value, maximum heart rate, and resting heart rate.
[0037] The present invention has the following beneficial effects:
[0038] 1) Higher accuracy: The present invention combines the influence of the measurement indexes of wearable devices on the occurrence risk of cardiovascular diseases, and based on the XGBoost machine learning algorithm with the best comprehensive prediction performance, the prediction of the probability of medium-to-high risk of individual CVD risk has higher accuracy.
[0039] 2) Improve the risk warning ability: The present invention continuously and dynamically monitors the health characteristic data of individuals based on wearable devices, captures the change trends of key indexes such as stress, heart rate, sleep, and exercise in real time, and combines the constructed risk prediction model, so as to realize accurate cardiovascular disease risk prediction and early warning, with the advantages of strong real-time performance, convenient data collection, continuous measurement, easy popularization, etc., and improves the timeliness and effectiveness of health management.
[0040] 3) Improve the personalization ability: The present invention constructs an individual-specific cardiovascular disease risk prediction model through personalized health data accumulation and machine learning modeling, can more accurately reflect the real risk level of individuals, can dynamically evaluate the cardiovascular health status of individuals, provide personalized and timely risk warnings, is applicable to scenarios such as daily health management and remote monitoring, and improves the accuracy and applicability of cardiovascular disease risk prediction.
[0041] 4) Improve public health awareness: The present invention not only provides an accurate method for assessing the risk of cardiovascular diseases for individuals, but also can compare the contributions of different characteristics in predicting risks, solving the "black box" problem in machine learning for cardiovascular risk prediction. Through personalized risk assessment and health data feedback, it encourages individuals to pay more active attention to cardiovascular health and cultivate good living habits, such as reasonable sleep, regular exercise, and stress reduction management, thereby effectively reducing the risk of cardiovascular diseases.
[0042] 5) Improve the efficiency of health management: The present invention can help individuals monitor their own health status at any time through wearable devices, combine with an intelligent early warning system for active intervention, reduce the dependence on chronic disease management, improve the ability of self-health management, and at the same time promote the development of the hierarchical diagnosis and treatment model, improving the overall efficiency of health management.
[0043] 6) Promote the integrated development of mobile health technologies and medical services: Relying on intelligent wearable devices, big data analysis, and telemedicine technologies, this method promotes the deep integration of digital health and traditional medical models. Medical institutions can provide accurate interventions and personalized management based on real-time health data, improve the efficiency of disease prevention and chronic disease management, and at the same time promote medical technology innovation and accelerate the construction and development of the intelligent medical system. Brief Description of the Drawings
[0044] Figure 1 The receiver operating characteristic curves of Model 1 and Model 2 fitted by the XGBoost model in the training and validation sets for Embodiment 1 of the present invention.
[0045] Figure 2 The contribution degree of each feature in Model 1 fitted by the XGBoost model in the training set for Embodiment 2 of the present invention.
[0046] Figure 3 The contributions of each feature in the XGBoost model for Embodiment 2 of the present invention. A is the swarm plot of the SHAP values of each feature, and B is the waterfall plot of one example. Detailed Embodiments
[0047] The technical solutions of the present invention will be clearly and completely described below in conjunction with specific embodiments. The embodiments given are only for better explaining the present invention, rather than limiting the scope of the present invention.
[0048] Provide a cardiovascular disease risk assessment system based on wearable device data, including:
[0049] A data acquisition module, which obtains traditional cardiovascular risk factors through a cardiovascular risk assessment scale and obtains monitoring data through wearable devices;
[0050] The feature screening module determines the range of candidate predictors, and then extracts features and screens predictors for the population wearing wearable devices by jointly using the LASSO, RF, and Logistic models;
[0051] The model construction module constructs a model based on the XGBoost machine learning algorithm;
[0052] The risk diagnosis module outputs the probability of an individual having a medium to high CVD risk within 10 years based on the feature data obtained in real time from the wearable device data.
[0053] In the present invention, the primary outcome is the 10-year CVD risk (medium / high risk or low risk) predicted according to the "Cardiovascular Risk Assessment Scale". This scale is designed based on the cardiovascular disease primary prevention risk assessment process for Chinese adults in the "Primary Version of the Chinese Guidelines for Cardiovascular Disease Primary Prevention", and assesses the risk of cardiovascular disease in the next 10 years based on individual cardiovascular disease risk factors, with high accuracy and authority.
[0054] In the present invention, monitoring data is obtained through a wearable device, and preferably demographic information including age and gender is also collected. The wearable device can measure exercise parameters and medical parameters, such as the number of steps, calorie consumption, duration of moderate to high-intensity exercise, sleep duration, maximum pressure value, average pressure value, maximum heart rate, and resting heart rate. The wearable device can be designed to be worn on the patient's wrist, such as a wristband, watch, or bracelet.
[0055] In a preferred embodiment, the features extracted by the feature screening module from the wearable device include the number of steps, calorie consumption, duration of moderate to high-intensity exercise, sleep duration, maximum pressure value, average pressure value, maximum heart rate, and resting heart rate.
[0056] The inventors studied the prediction performance of various models in CVD risk assessment and found that the XGBoost model has the best comprehensive prediction performance. Therefore, the present invention is based on the XGBoost model. XGBoost is a machine learning algorithm based on the gradient boosting framework, which gradually optimizes the model performance by iteratively constructing multiple weak learners (such as decision trees). Its goal is to minimize the objective function, which includes a loss function and a regularization term, to balance the model's fitting ability and complexity. The output of the model is represented by the weighted sum of multiple trees:
[0057]
[0058] where, is the predicted value of the i-th sample, x i is the feature vector, f k is the k-th decision tree, is the function space of the tree model.
[0059] Through second-order Taylor expansion, the optimization objective is approximated as:
[0060]
[0061] where g i and h i are the first-order derivative and second-order derivative of the loss function respectively, and Ω(f t ) is the regularization term, which is used to control the model complexity.
[0062] Select prediction factors based on feature screening to construct the XGBoost model, and use grid search and cross-validation to select the optimal parameters to finally fit the model.
[0063] In a preferred embodiment, the cardiovascular disease risk assessment system of the present invention further includes: a feature importance evaluation module, which uses SHAP values to evaluate the contribution of features.
[0064] SHAP values are based on the Shapley values in cooperative game theory. By calculating the marginal contribution of features to the model output, the importance of each feature and its impact on the prediction result are quantified. Its core idea is to perform weighted averaging on all possible feature combinations to ensure the fairness and consistency of the interpretation.
[0065] For a given model f(x) and feature set N = {1, 2, …, n}, the SHAP value of feature i is defined as:
[0066]
[0067] where S represents the feature subset that does not include feature i, f(S) is the output of the model trained only based on S, and φ i represents the SHAP value of feature i. SHAP values have additivity, that is where f0 is the baseline output of the model. Thus, the contribution degree of each feature of an individual can be obtained, which is convenient for feedback on which features have a greater impact on the individual and can be reasonably controlled. The swarm plot of the SHAP values of each feature can be used to reflect the contribution of each feature. Preferably, the waterfall plot of XGBoost is used for visualization to reflect the cumulative contribution of individual features to the model output.
[0068] Example 1
[0069] This example provides a model that integrates the health index dimensions of wearable devices and can predict the probability of an individual having a medium to high total CVD risk within the next 10 years. It includes:
[0070] 1. Dataset preparation
[0071] Our team selected the "Cardiovascular Risk Assessment Scale" as the baseline characteristics, including demographic data and traditional cardiovascular risk factors, and obtained information such as exercise, sleep, stress, heart rate, and blood pressure through wearable devices.
[0072] The data was divided into a training set and a test set at a ratio of 7:3, and the baseline characteristics are as follows:
[0073]
[0074] In this example, the main outcome: the 10-year CVD risk (medium / high risk / low risk) predicted by the "Cardiovascular Risk Assessment Scale". This scale is designed according to the cardiovascular disease primary prevention risk assessment process for Chinese adults in the "Primary Care Edition of the Chinese Guidelines for Cardiovascular Disease Primary Prevention", and based on individual cardiovascular disease risk factors, it assesses the risk of developing cardiovascular disease in the next 10 years, with high accuracy and authority.
[0075] The "Cardiovascular Risk Assessment Scale" is as follows:
[0076] Question 1 / 10 Your gender (single choice)
[0077] · Male
[0078] · Female
[0079] Question 2 / 10 Your year of birth:
[0080] Question 3 / 10 Your height (cm):
[0081] Question 4 / 10 Your weight (kg):
[0082] Question 5 / 10 Do you drink alcohol? (single choice)
[0083] · Frequently (≥ once a week)
[0084] · Occasionally
[0085] · Never
[0086] Question 6 / 10 Do you smoke? (single choice)
[0087] · Frequently (≥ once a week)
[0088] · Occasionally
[0089] · Never
[0090] · Have quit smoking
[0091] Question 7 / 10 Your current blood pressure situation (single choice)
[0092] · Systolic blood pressure < 130 mmHg and diastolic blood pressure < 80 mmHg
[0093] · Systolic blood pressure 130 - 139 mmHg, and / or diastolic blood pressure 80 - 89 mmHg
[0094] · Diagnosed with hypertension and taking medication
[0095] · Diagnosed with hypertension but not taking medication
[0096] · Uncertain
[0097] Question 8 / 10. What is your current blood glucose situation (single choice)
[0098] · Normal
[0099] · Diagnosed with diabetes and taking medication
[0100] · Diagnosed with diabetes but not taking medication
[0101] · Uncertain
[0102] Question 9 / 10. What is your current blood lipid situation (single choice)
[0103] · Normal
[0104] · Diagnosed with dyslipidemia and taking medication
[0105] · Diagnosed with dyslipidemia but not taking medication
[0106] · Uncertain
[0107] Question 10 / 10. Do you have any of the following diseases (multiple choices)
[0108] · Coronary heart disease
[0109] · Stroke
[0110] · Chronic kidney disease
[0111] · Heart failure
[0112] · Atrial fibrillation
[0113] · Malignant tumor
[0114] · None of the above diseases
[0115] 2. Feature screening
[0116] (1) Predictors
[0117] Demographic data: Gender, age (years);
[0118] Wristband data: Steps (steps / day), distance (km / day), calories (kcal / day), moderate - to - high - intensity exercise (minutes / day), sleep duration (hours / day), maximum stress, minimum stress, average stress, maximum heart rate (beats / min), minimum heart rate (beats / min), resting heart rate (beats / min).
[0119] (2) Combine the population feature extraction method to screen for predictive factors
[0120] By jointly using the LASSO, RF, and Logistic regression models, extract features and screen for predictive factors in the population wearing wearable devices.
[0121] A. LASSO
[0122] The diagnostic model of LASSO controls the model complexity and performs feature selection by introducing a penalty term to limit the magnitude of the regression coefficients. Its objective function is:
[0123]
[0124] where N is the number of study subjects, p is the number of predictive factors, β0 is the constant term, β is the regression coefficient matrix, and t is the penalty term, which controls the model complexity to avoid overfitting. LASSO can achieve variable screening by shrinking the regression coefficients of some predictive factors to zero.
[0125] B. Random Forest
[0126] Random Forest is an ensemble learning method that makes predictions by constructing multiple decision trees and performing weighted voting. Its main process includes the following steps:
[0127] 1. Data sampling: Use the bootstrap method to randomly sample from the original data to generate multiple subsets (in-bag data), and the remaining approximately 37% of the data is used as out-of-bag data.
[0128] 2. Feature selection and tree construction: At each node of each decision tree, randomly select some features for splitting and generate a decision tree without pruning.
[0129] 3. Weighted voting: Vote on or take the weighted average of the prediction results of all decision trees to determine the final output result.
[0130] The total prediction probability of the RF model can be expressed as:
[0131]
[0132] where B is the total number of decision trees in the random forest, and f b (x) is the output of the b-th decision tree.
[0133] The importance index (variable importance) of RF can be estimated by calculating the change in prediction error of the out-of-bag data, and its formula is:
[0134]
[0135] where is the prediction error of the out-of-bag data of the b-th decision tree, is the prediction error of the out-of-bag data after the j-th feature is permuted (randomly shuffled). A larger VI(j) value indicates a higher importance of this feature to the model prediction result.
[0136] Finally, 8 predictors are selected to construct the basic model (Model 1) of the wearable device, including the number of steps, calorie consumption, duration of moderate-to-high-intensity exercise, sleep duration, maximum stress value, average stress value, maximum heart rate, and resting heart rate. On this basis, age and gender are added to construct the extended model (Model 2).
[0137] 3. Model Construction
[0138] Based on the XGBoost machine learning algorithm, multiple weak learners (such as decision trees) are iteratively constructed to gradually optimize the model performance. Its goal is to minimize the objective function, which includes the loss function and the regularization term, to balance the fitting ability and complexity of the model. The output of the model is represented by the weighted sum of multiple trees:
[0139]
[0140] where, is the predicted value of the i-th sample, x i is the feature vector, f k is the k-th decision tree, is the function space of the tree model.
[0141] Through the second-order Taylor expansion, the optimization objective is approximated as:
[0142]
[0143] where, g i and h i are the first derivative and the second derivative of the loss function respectively, and Ω(f t ) is the regularization term, which is used to control the model complexity.
[0144] Select the predictors based on feature screening to construct the XGBoost model, and use grid search and cross-validation to select the optimal parameters and finally fit the model.
[0145] The model is internally validated by dividing the training set and the test set to test whether the model is overfitting. The features included in Model 1 are the number of steps, calorie consumption, duration of moderate-to-high-intensity exercise, sleep duration, maximum stress value, average stress value, maximum heart rate, and resting heart rate. Model 2 includes age and gender on the basis of Model 1.
[0146] The predictive performance of the base model (Model 1) and the model with age and gender added (Model 2) was evaluated in the training set. The area under the curve (AUC) was used to measure the discrimination ability of the classifier, the calibration slope was used to evaluate the consistency between the predicted probability and the actual result, the accuracy was used to measure the overall correctness of the prediction result, the recall was used to evaluate the ability to identify positive examples, the specificity was used to measure the discrimination ability of negative examples, and the Brier score was used to comprehensively evaluate the accuracy and reliability of the predicted probability.
[0147] Before calculating the calibration slope, accuracy, recall, and specificity, Platt transformation was performed to convert the model output values (such as the scores of XGBoost) into probability values. Its core is to fit the relationship between the decision value and the label through Logistic regression, satisfying the following formula:
[0148]
[0149] where A and B are parameters optimized by maximum likelihood estimation to provide a more intuitive probability interpretation and an output convenient for threshold setting.
[0150] The specific evaluation results are as follows:
[0151]
[0152] Using only the data obtained from the wearable device, the AUC of the model has reached 0.751, and the 95% CI is (0.735 - 0.766); after adding age and gender on this basis, the AUC is further increased to 0.841, and the 95% CI is (0.829 - 0.853), indicating that the XGBoost model has good performance in predicting the probability of medium to high risk of 10-year cardiovascular events. The receiver operating characteristic curves of Model 1 and Model 2 fitted by the XGBoost model in the training and validation sets are shown in the appendix Figure 1 .
[0153] Example 2
[0154] Based on the model constructed in Example 1, the importance of each feature was evaluated.
[0155] In Model 1, the average |SHAP value| of each feature was calculated to compare the contribution of each feature to the model and judge the importance. The contribution degree of each feature of Model 1 fitted by the XGBoost model in the training set is shown in the appendix Figure 2 . It can be seen that the maximum heart rate, calories, and resting heart rate show relatively high importance.
[0156] The waterfall plot of XGBoost is used to visualize the cumulative contribution of individual features to the model output. As Figure 3As shown, starting from the baseline value E[f(x)] of the model (the mean of the predicted values of all samples), by gradually adding the SHAP values of each feature, it shows how the features drive the predicted value to change from the baseline value E[f(x)] to the final model output f(x). A positive SHAP value indicates a positive contribution of the feature to the predicted value, while a negative SHAP value indicates an inhibitory effect. The waterfall plot intuitively reflects the step-by-step impact of the features on the model's decision-making, facilitating feedback on which features have a greater impact on the individual and can be reasonably controlled.
[0157] It can be seen from the evaluation results of the embodiments of the present invention that the present invention has good performance in predicting the probability of medium to high risk of 10-year cardiovascular events. It can not only evaluate the probability of medium to high risk of 10-year occurrence risk of individual CVD, but also obtain the contribution degree of each feature of the individual, facilitating feedback on which features have a greater impact on the individual and can be reasonably controlled.
Claims
1. A cardiovascular disease risk assessment system based on wearable device data, comprising: A data acquisition module that obtains traditional cardiovascular risk factors through a cardiovascular risk assessment scale and obtains monitoring data through a wearable device; A feature screening module that determines the range of candidate predictors, and then extracts features and screens predictors for the population wearing wearable devices by jointly using the Least Absolute Shrinkage and Selection Operator (LASSO), Random Forest (RF), and Logistic regression model; A model construction module that constructs a model based on the XGBoost machine learning algorithm; A risk diagnosis module that outputs the probability that the evaluated individual has a medium to high risk of CVD within 10 years based on the feature data obtained in real time from the wearable device data.
2. The cardiovascular disease risk assessment system based on wearable device data according to claim 1, characterized in that It further includes: A feature importance evaluation module that evaluates the contribution of features using SHAP values. For a given model f(x) and feature set N = {1, 2, …, n}, the SHAP value of feature i is defined as: where S represents the subset of features that does not include feature i, f(S) is the output of the model trained only based on S, and φ i represents the SHAP value of feature i.
3. The cardiovascular disease risk assessment system based on wearable device data according to claim 2, characterized in that: The feature importance evaluation module visually reflects the cumulative contribution of individual features to the model output using the waterfall plot of XGBoost.
4. The cardiovascular disease risk assessment system based on wearable device data according to any one of claims 1 to 3, characterized in that: The data acquisition module also collects demographic information, including age and gender.
5. The cardiovascular disease risk assessment system based on wearable device data according to any one of claims 1 to 3, characterized in that: The features extracted by the feature screening module include the number of steps, calorie consumption, duration of moderate to high-intensity exercise, sleep duration, maximum stress value, average stress value, maximum heart rate, and resting heart rate.
6. A method for cardiovascular disease risk assessment based on wearable device data, comprising the following steps: S1) Data acquisition: Obtain traditional cardiovascular risk factors through a cardiovascular risk assessment scale and obtain monitoring data through a wearable device; S2) Feature screening: Determine the range of candidate predictors, and then extract features and screen predictors for the population wearing wearable devices by jointly using the Least Absolute Shrinkage and Selection Operator (LASSO), Random Forest (RF), and Logistic regression model; S3) Model construction: Based on the XGBoost machine learning algorithm; S4) Risk diagnosis: Output the probability that the evaluated individual has a medium to high risk of CVD within 10 years based on the feature data obtained in real time from the wearable device data.
7. The method for cardiovascular disease risk assessment based on wearable device data according to claim 6, characterized in that It further includes: S5) Feature importance evaluation: Evaluate the contribution of features using SHAP values. For a given model f(x) and feature set N = {1, 2, …, n}, the SHAP value of feature i is defined as: where S represents the feature subset that does not contain feature i, f(S) is the output of the model trained only based on S, and φ i represents the SHAP value of feature i.
8. The cardiovascular disease risk assessment method based on wearable device data according to claim 7, characterized in that: In S5) Feature importance evaluation, the cumulative contribution of individual features to the model output is visually reflected using the waterfall plot of XGBoost.
9. The cardiovascular disease risk assessment method based on wearable device data according to any one of claims 6 to 8, characterized in that: In S1) Data acquisition, demographic information, including age and gender, is also collected.
10. The method for cardiovascular disease risk assessment based on wearable device data according to any one of claims 6 to 8, characterized in that: In S2) Feature screening, the features extracted include the number of steps, calorie consumption, duration of moderate to high-intensity exercise, sleep duration, maximum stress value, average stress value, maximum heart rate, and resting heart rate.
Citation Information
Patent Citations
Stroke onset risk prediction system and application
CN110428901A
Method for predicting onset risk of cardiovascular disease (CVD)
CN112331362A
Method for predicting cardiovascular data of middle-aged and elderly people
CN113764105A
Cardiovascular disease onset risk prediction method easy to popularize and apply
CN114783606A
Dynamic risk assessment method and system based on intelligent wearable device
CN117877732A