Cardiovascular disease risk assessment method and system based on machine learning model
By building a cardiovascular disease risk assessment system based on multiple machine learning models, using SHAP values to select key features and construct discriminative sequences and matrices, the problem of low accuracy and reliability of traditional cardiovascular disease risk assessment models is solved, and efficient and accurate risk assessment and model scalability are achieved.
Patent Information
- Application Number
- CN202510448188.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-10
- Publication Date
- 2025-07-11
- Estimated Expiration
- 2045-04-10
AI Technical Summary
The existing cardiovascular disease risk assessment model has problems such as insufficient feature representation, weak generalization ability of the model, and low accuracy and reliability of risk assessment. Traditional models are difficult to train and are prone to local optimality, and diagnostic rules are cumbersome and have artificial limitations.
A variety of machine learning models (such as LR, GB, NNK, GNB, SVM) are used to build a cardiovascular disease risk assessment system, select key features through SHAP values, build a discrimination sequence and identification matrix, and comprehensively utilize the prediction results of different models to conduct efficient and accurate 5-categorized risk assessment.
It realizes efficient and accurate cardiovascular disease risk assessment, improves the scalability and diagnostic accuracy of the model, eliminates the interference of input parameters on classification, and improves the applicability of the model to predict results in different users.
Smart Images

Figure CN120299716A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of artificial intelligence, and particularly relates to a cardiovascular disease risk assessment method and system based on a machine learning model.
Background Art
[0002] With the gradual aggravation of the aging process of the social population, chronic diseases have become the number one killer of the global residents' health. Among them, cardiovascular diseases are major diseases that the elderly have to prevent and control. Society is facing great pressure of continuous increase in cardiovascular diseases. Cardiovascular diseases refer to diseases that affect the functions of the heart and blood vessels. Common cardiovascular diseases include hypertension, myocardial infarction, coronary heart disease, heart failure, heart disease, arrhythmia, etc. At present, there is no effective chronic disease management model. Cardiovascular diseases are also called cardio-cerebrovascular diseases. The mortality rate of cardiovascular diseases ranks first, higher than that of tumors and other diseases.
[0003] With the development of big data and artificial intelligence technologies, combined with intelligent models, and the mature application of intelligent wearable warning devices, there is also great room for improvement in research on comprehensively and accurately assessing cardiovascular disease risk factors, regularly monitoring, effectively intervening and preventing, and early clinical diagnosis and treatment. With the development of the computer field, medical intelligent diagnosis systems have emerged. The knowledge base in a medical intelligent diagnosis system is used to store the diagnostic rules of diseases. The diagnostic rules mainly describe the manifestations and characteristics of diseases and represent the causal relationship between diseases and symptoms. The early knowledge bases were completely organized according to the thinking processes of clinicians. Manually sorting out the knowledge bases based on doctors' experience is very cumbersome, time-consuming and laborious, and the generation of rules has certain human limitations, not considering enough the diversity, variability and uncertainty factors of diseases, and some special rules may be omitted.
[0004] Although methods for generating a rule base through machine learning are gradually emerging, they are somewhat lacking in terms of technological maturity and rule rationality. In practical applications, the diagnostic accuracy rate is relatively low, and there is still a large room for improvement. Cardiovascular disease risk prediction and assessment at home and abroad are basically based on predictive variables such as blood pressure, body mass index, blood sugar, total cholesterol, smoking, disease history, and family history. There are problems in the construction of traditional cardiovascular disease risk assessment models, such as insufficient feature representation, weak model generalization ability, and low accuracy and reliability of risk assessment; there are problems in the optimization of traditional cardiovascular disease risk assessment models, such as low search efficiency, being prone to falling into local optima, and parameter adjustment relying on manual experience. The cardiovascular disease risk assessment system based on machine learning is an innovative medical technology that uses machine learning algorithms to deeply mine and learn medical data, constructs a cardiovascular disease risk assessment model, and can achieve accurate prediction and assessment of cardiovascular disease risks. Currently, in many studies, neural networks or even deep neural networks are used for prediction and assessment. A neural network usually consists of multiple layers, including an input layer, a hidden layer, and an output layer, and each layer is composed of multiple neurons. The structure of traditional machine learning models is relatively simple. For example, a linear model only has simple parameters, and a decision tree has a branching structure, but it does not have a multi-layer structure like a neural network. The complexity of a neural network is higher than that of traditional machine learning models, and its training difficulty is very high. It is basically difficult to achieve a valuable training goal, while the use of machine learning models is relatively simple, the training speed is faster, and a lot of training experience has been accumulated. How to make full use of the advantages of machine learning models to achieve the intelligent upgrade of risk assessment is an issue to be solved. Based on the above problems, the present invention makes full use of diagnostic disease indicators and multiple machine learning models, and fully utilizes the prediction results and their differential characteristics of different machine learning models to achieve efficient and accurate 5-class risk assessment.
Summary of the Invention
[0005] To solve the above problems in the prior art, the present invention proposes a method and system for cardiovascular disease risk assessment based on machine learning models. The method includes:
[0006] Step S1: Construct I machine learning models as prediction models, and select N user features as alternative input features;
[0007] Step S2: Select N1 user features from the alternative input features as the first input features of the prediction models;
[0008] Step S3: Train and validate the prediction models based on the N1 user features of the samples, and set evaluation parameters to evaluate the prediction models;
[0009] Step S4: constructing a discrimination sequence associated with each two classifications based on the evaluation parameters; the discrimination sequence includes a prediction model identifier, and the recognition ability of distinguishing between the two classifications is ranked according to the prediction model;
[0010] Step S5: Obtain I classification results R corresponding to I prediction models for N1 user features of the user to be evaluated i =(r i,k );where: r i,k is the kth element of the classification result of the i-th prediction model; determine the maximum element value max(r i,k ), if the maximum values of the elements in the classification results of I prediction models all point to the same classification, then the mean of the I classification results is taken as the evaluation result; if M1 prediction models greater than or equal to PER1 point to the same classification k1, determine the classification pointed to by the maximum value of the elements in the classification results of the prediction models that do not point to the same classification, which is called classification k2, obtain the recognition sequence corresponding to the classification group (k1, k2), perform recognition based on the recognition sequence, and take the recognition result as the evaluation result.
[0011] Furthermore, I=5, and the 5 machine learning models are 5 different machine learning models.
[0012] Furthermore, the five machine learning models are LR, GB, NNK, GNB, and SVM machine learning models.
[0013] Furthermore, N=29.
[0014] Furthermore, N1=20.
[0015] Furthermore, the N1 user features are arranged in sequence to form an N-dimensional vector as the first input feature of the machine learning model.
[0016] Further, PER1=50%.
[0017] Furthermore, the identification is performed based on the identification sequence, specifically: when the maximum element value of the M1 prediction results points to the same classification k1, the classification result of the I-M1 prediction model sorted last in the identification sequence is deleted; if all the classification results that have not been deleted point to the same classification, the mean of all the classification results that have not been deleted is used as the evaluation result; otherwise, set I=M1 and return to step S5.
[0018] Furthermore, before training the prediction model, the sample data used for training is preprocessed, and all samples containing missing values are removed.
[0019] A cardiovascular disease risk assessment system based on a machine learning model, and the cardiovascular disease risk assessment system based on the machine learning model is used to implement the cardiovascular disease risk assessment method based on the machine learning model.
[0020] The beneficial effects of the present invention include:
[0021] (1) Making full use of diagnostic disease indicators and multiple machine learning models, and fully utilizing the prediction results and their differential characteristics of different machine learning models to achieve efficient and accurate 5-class risk assessment.
[0022] (2) For the prediction results of different users, using 2 classifications to construct an identification sequence and an identification matrix to select multiple prediction models, which can make full use of the classification results of different prediction models; when the classification results are inconsistent but the differences between different prediction models are not obvious, by constructing an identification matrix, it is possible to make full use of the evaluation of all dimensions on the prediction models, find the prediction model with the most obvious differential performance to participate in the prediction evaluation, and comprehensively use the learning results of multiple prediction models; by hierarchically using diagnostic disease indicators to eliminate the classification confusion caused by input parameters to the prediction models, multiple prediction models can reach rapid consistency under basically the same conditions; at the same time, the scalability of this method is also very strong.
Description of the Drawings
[0023] The drawings described herein are used to provide a further understanding of the present invention, form a part of this application, but do not constitute an improper limitation to the present invention. In the drawings:
[0024] Figure 1 It is a schematic diagram of the cardiovascular disease risk assessment method based on the machine learning model provided by the present invention.
[0025] Figure 2 is a schematic diagram of the prediction ROC curves of five machine learning models provided by the present invention. Among them: Figure 2(A) shows patients with coronary atherosclerotic heart disease and healthy people; Figure 2(B): patients with unstable angina and healthy people; Figure 2(C): patients with heart failure and healthy people; Figure 2(D): the ROC curves of five machine learning models in predicting patients with heart failure and other cardiovascular diseases.
[0026] Figure 3 It is an assessment schematic diagram of predicting coronary atherosclerotic heart disease and healthy people provided by the present invention.
[0027] Figure 4 It is an assessment schematic diagram of predicting unstable angina and healthy people provided by the present invention.
[0028] Figure 5 It is an assessment schematic diagram of predicting heart failure and healthy people provided by the present invention.
[0029] Figure 6 It is a schematic diagram for evaluating the prediction of heart failure and other cardiovascular diseases provided by the present invention.
Specific Embodiments
[0030] The present invention will be described in detail below in conjunction with the accompanying drawings and specific embodiments, where the illustrative embodiments and descriptions are only used to explain the present invention, but not to limit the present invention.
[0031] The present invention proposes a method and system for cardiovascular disease risk assessment based on a machine learning model. As shown in the attached Figure 1 figure, the method includes the following steps:
[0032] Step S1: Construct I machine learning models as prediction models, and select N user features as alternative input features;
[0033] Preferably: I = 5, and the 5 machine learning models are 5 different machine learning models, namely LR, GB, NNK, GNB, and SVM machine learning models; the N user features are features closely related to blood cells, blood lipids, and blood glucose;
[0034] Preferably: N = 29; the 29 user features include gender (Sex), age (Age), absolute basophil count (BAS#), basophil percentage (BAS%), absolute eosinophil count (EOS#), eosinophil percentage (EOS%), hematocrit (HCT), hemoglobin (HGB), absolute lymphocyte count (LYM#), lymphocyte percentage (LYM%), mean corpuscular hemoglobin content (MCH), mean corpuscular hemoglobin concentration (MCHC), mean corpuscular volume (MCV), absolute monocyte count (MON#), monocyte percentage (MON%), mean platelet volume (MPV), absolute neutrophil count (NEU#), neutrophil percentage (NEU%), plateletcrit (PCT), platelet distribution width (PDW), platelet count (PLT), red blood cell count (RBC), red blood cell distribution width (RDW), white blood cell count (WBC), glucose (GLU), high-density lipoprotein cholesterol (HDL-C), low-density lipoprotein cholesterol (LDL-C), total cholesterol (TC), triglyceride (TG);
[0035] Step S2: Select N1 user features from the alternative input features as the first input features of the prediction model; specifically: assign a SHAP value to each user feature, and the SHAP value is used to indicate the influence of the user feature on the prediction model; select N1 user features from the alternative input features as the first input features of the prediction model based on the magnitude of the SHAP value;
[0036] Preferably: N1 user features are arranged in sequence to form an N-dimensional vector as the first input feature;
[0037] Preferably: the classification results of the prediction model include classifications 1-5, which are coronary atherosclerotic heart disease, unstable angina, heart failure, other cardiovascular diseases, and health; the corresponding setting of the output of the prediction model is a 5-tuple vector, each element in the output corresponds to the above 5 classification results, the sum of all elements in the output vector is 1, and the value of each element indicates the confidence of the prediction model for its corresponding classification result; for example: (0.1, 0.2, 0.4, 0, 0) indicates that the confidence for coronary atherosclerotic heart disease, heart failure, and unstable angina is 0.1, 0.2, and 0.4, respectively, and the confidence for unstable angina is the highest; but in fact, the expression of 0.4 cannot bring reliable information to the evaluation, and the reference value is limited;
[0038] Preferably: the SHAP value of the user feature is calculated by the "shap python" package (version 0.41.0), and for 5 prediction models, N1=20 is set, and the top 20 key user features obtained by the SHAP algorithm are counted as the first input feature; the "Counter" function is used in the algorithm to count the feature word frequency, the word frequency is calculated by flattening the data, and the importance score is constructed based on the feature importance ranking, and the comprehensive score is generated by combining the word frequency and importance to identify the key features;
[0039] Alternatively: determine the first N1 key user features as the first input features based on user feedback; the user can select the key user features based on expert experience;
[0040] Step S3: training and verifying the prediction model based on the N1 user features of the sample, and setting one or more evaluation parameters to evaluate the prediction model; specifically: training and verifying the prediction model based on the N1 user features of the sample, and determining the evaluation parameters based on the area under the ROC curve AUC to evaluate the performance of each prediction model;
[0041] Preferably: before training the prediction model, preprocess the sample data used for training and remove all samples with missing values;
[0042] Preferably: for each prediction model, the data of healthy individuals and disease patients are randomly divided into a training set and a validation set, and the data of all cases are randomly divided into a training set (70%) for model development and a test set (30%) for performance evaluation using the "Sample" function in the "Pandas" package; wherein: the training set is used to train the model, and the test set is used to evaluate the effect of the model;
[0043] Preferably, the evaluation parameters include the area under the ROC curve AUC, accuracy Acc, sensitivity Sn, specificity Sp, positive predictive value PPV, negative predictive value NPV, Matthews correlation coefficient MCC, and / or F1 score. The model performance is evaluated based on these 8 evaluation parameters; the following formulas (1)-(7) are used to determine the evaluation parameters; where: TP, TN, FP, and FN represent true positive, true negative, false positive, and false negative, respectively;
[0044] ACC = (TP + TN) / (TP + TN + FP + FN) (1);
[0045] Sn = TP / (TP + FN) (2);
[0046] Sp = TN / (TN + FP) (3);
[0047] PPV = TP / (TP + FP) (4);
[0048] NPV = TN / (TN + FN) (5);
[0049]
[0050] F1 = 2TP / (2TP + FN + FP) (7);
[0051] The above 8 evaluation parameters are used in different scenarios. For example, Acc accuracy is effective when the data is balanced, but may be misleading in unbalanced cases; AUC is more stable, especially in binary classification problems. Sensitivity and specificity are for different classes respectively, while PPV and NPV are the probabilities of inferring the true class from the prediction results; MCC takes into account all four elements of the confusion matrix, so it is more meaningful in the case of class imbalance; F1 is suitable when precision and recall are equally important; therefore, for the prediction results of different users, using 2 classifications to construct classification groups and selecting multiple prediction models based on AUC can make full use of the classification results of different prediction models;
[0052] Step S4: Construct a discrimination sequence associated with every 2 classifications based on the evaluation parameters; the discrimination sequence contains the prediction model identifier, and is sorted according to the discrimination ability of the prediction model to distinguish between these 2 classifications;
[0053] The discrimination sequence; specifically: constructing the discrimination sequence based on AUC; after completing the training of the 5 prediction models, the evaluation data obtained is shown in Figure 2-6; it can be seen from the figure that when predicting patients with coronary atherosclerotic heart disease and healthy people, the GB model shows good comprehensive performance, with an AUC of 0.97 (95% CI: 0.93-0.99); when differentiating patients with unstable angina and healthy people, the GNB model shows the best comprehensive performance, with an AUC of 0.94 (95% CI: 0.89-0.99); in differentiating patients with heart failure and healthy people, the LR and GNB models have good comprehensive performance, with corresponding AUCs of 0.99 (95% CI: 0.96-1.00) and 0.99 (95% CI: 0.97-1.00); similarly, in the construction of the model for differentiating heart failure from other cardiovascular diseases, the GB model shows quite excellent comprehensive performance, with an AUC of 0.81 (95% CI: 0.69-0.92); the AUCs of the LR, GNB, NN, and SVM models are 0.76 (95% CI: 0.59-0.91), 0.76 (95% CI: 0.61-0.90), 0.76 (95% CI: 0.58-0.90), and 0.76 (95% CI: 0.59-0.89) in turn; then the discrimination sequence for the classification group (3,4) is (GB(0.81), LR(0.76), GNB(0.76), NN(0.76), SVM(0.76)); the construction of the discrimination sequence for other classification groups is similar;
[0054] Alternatively: the specific step S4 is: constructing a discrimination matrix for every 2 classifications based on all evaluation parameters; the element value m in the discrimination matrix i,j indicates the parameter value of the jth evaluation parameter of the prediction model i; when the classification results are inconsistent but the differences between different prediction models are not obvious, by constructing the discrimination matrix, it is possible to make full use of the evaluation of all dimensions of the prediction models, find the prediction model with the most obvious differential performance to participate in the prediction evaluation, and comprehensively utilize the learning results of multiple prediction models; at the same time, the scalability of this method is also very strong;
[0055] Step S5: Obtain I classification results R corresponding to I prediction models for N1 user characteristics of the user to be evaluated i =(r i,k ); where: r i,k is the kth element of the classification result of the ith prediction model; determine the maximum element value max(r i,k)For the classification pointed to, if the maximum values of the elements in the classification results of I prediction models all point to the same classification, then the mean of the I classification results is used as the evaluation result; if M1 prediction models greater than or equal to PER1 point to the same classification k1, determine the classification pointed to by the maximum value of the elements in the classification results of the prediction models that do not point to the same classification, which is called classification k2, obtain the identification sequence corresponding to the classification group (k1, k2), perform identification based on the identification sequence, and use the identification result as the evaluation result; if at most 1 prediction result points to the same classification (all classification results point to different classifications), then feedback is performed;
[0056] Preferably: In the initial state, set I = 5, and input the user characteristics to be evaluated into 5 prediction models to obtain classification results for each prediction model;
[0057] Preferably: PER1 is a preset percentage, for example: PER1 = 50%;
[0058] Alternatively: If the maximum values of the elements in the classification results of I prediction models all point to the same classification, then the classification result corresponding to the prediction model with the largest AUC value is used as the evaluation result;
[0059] Preferably: When there are multiple classification k2, select the classification with the largest number of times pointed to by the maximum value as classification k2; when multiple classifications have the same number of times pointed to by the maximum value, randomly select one as classification k2;
[0060] The said feedback is specifically: Determine that the intelligent evaluation is invalid; and prompt that there may be an abnormality in user feature collection;
[0061] The said identification based on the identification sequence is specifically: At this time, the maximum element values of M1 prediction results point to the same classification k1, and delete the classification results of the I - M1 prediction models with the last sorting in the identification sequence; if all the remaining classification results point to the same classification, then use the mean of all the remaining classification results as the evaluation result; otherwise, set I = M1 and return to step S5;
[0062] Alternatively: If all the remaining classification results point to the same classification, then use the classification result corresponding to the prediction model with the largest AUC value among them as the evaluation result;
[0063] Alternatively: If all the remaining classification results do not point to the same classification, then use the classification result of the one with the highest F1 score when identifying the classification group (k1, k2) among the remaining prediction models as the evaluation result;
[0064] Replaceable: If all the undropped classification results do not point to the same classification, when identifying for the classification group (k1, k2) in the undropped prediction models, obtain the classification results corresponding to the one with the highest F1 score and the one with the highest MCC coefficient; among the classification results, r i,k1 and r i,k2 Take the classification result with the largest difference as the evaluation result;
[0065] Corresponding to the replaceable manner of step S4, replaceable: The specific step SI is as follows: Obtain I classification results R i =(r i,k ) corresponding to I prediction models for the user characteristics to be evaluated; Determine the classification pointed to by the maximum element value max(r i,k ) in each of the I classification results. If the maximum element values in the classification results of more than or equal to I0% of the prediction models point to the same classification, take the mean of the I classification results as the evaluation result; If the maximum values of all the classification results point to different classifications, give feedback; Otherwise, determine the 2 classifications to which the maximum element value of the prediction result points the most, and call them k1 and k2 respectively. Obtain the identification matrix corresponding to the classification group (k1, k2), and perform identification based on the identification matrix, and take the identification result as the evaluation result;
[0066] The identification based on the identification matrix is specifically as follows: Determine the prediction model with the largest difference in evaluation ability based on the evaluation parameters, and call it the difference prediction model. Insert the shadow classification result based on this difference prediction model, and determine the identification result based on the classification result and the shadow classification result; Specifically, it includes the following steps:
[0067] Step S5A1: Take the mean of the I classification results as the first shadow result;
[0068] Preferably: Set the initial value of I to 5; That is to say, the present invention only exemplifies the case of 5 prediction models, and the number of prediction models can be extended according to needs, and it is preferably an odd number;
[0069] Step S5A2: For any row i in the identification matrix, determine the two elements with the largest difference between any two column elements m i,j ; Determine the row i where the two elements with the largest difference in all rows are located and its corresponding prediction model identifier, and call it the difference prediction model; Delete the columns where the two elements are located in the identification matrix; In this way, the same row and its corresponding prediction model identifier will not be selected again due to repeated evaluation factor reasons;
[0070] Step S5A3: Assign a first weight to the classification result of the difference prediction model, and assign a second weight to other classification results and the first shadow result; Wherein: The first weight is greater than the second weight;
[0071] Preferably: the first weight is equal to 2 times the second weight;
[0072] Step S5A4: Based on the first weight and the second weight, obtain the weighted value of the first shadow result and other classification results as the second shadow result;
[0073] Step S5A5: Determine whether the maximum value of elements in more than 50% of the I + 2 classification results points to the same classification. If so, take the mean of the I + 2 classification results as the evaluation result; otherwise, set I = I + 2 and return to Step S5A1;
[0074] Alternatively: Step S5 is specifically: obtain I first classification results R1 corresponding to I prediction models for the user characteristics to be evaluated i =(r1 i,k ); if the maximum value of elements in more than 50% of the first classification results points to the same classification, take the mean of the I first classification results as the evaluation result; otherwise, proceed to the next step S6;
[0075] Step S6: Obtain I second classification results R2 corresponding to I prediction models for N2 user characteristics of the user to be evaluated i =(r2 i,k ); if the maximum value of elements in more than 50% of the second classification results of the prediction models points to the same classification, take the mean of the I second classification results as the evaluation result; otherwise, proceed to Step S7;
[0076] Alternatively: if the maximum value of elements in more than 50% of the second classification results of the prediction models points to the same classification, take the mean of the I first classification results and the I second classification results as the evaluation result;
[0077] Preferably: select N2 user characteristics from N1 user characteristics based on SHAP values; the N2 user characteristics are the N2 more important user characteristics among the N1 user characteristic values; obtain the N2 user characteristics of the user to be evaluated to construct the input of the prediction model; for the part of the input that is not the N2 user characteristics, fill it with default values; input the input data filled with default values into the prediction models respectively to obtain I second classification results corresponding to the N2 user characteristics of I prediction models; that is to say, when there are differences in the prediction models, simply processing the classification results through evaluation parameters still cannot bypass the internal logic of the prediction models, and some of the N2 user characteristics are factors causing the differences and are not helpful for discovering the correct classification results. Especially when there are differences between the prediction models, too many user characteristics will cause "trouble" to the models. Therefore, evaluate by eliminating the disturbing feature factors.
[0078] Step S7: Combine the first classification result and the second classification result to form the classification result. If the maximum value of the elements in the second classification results of more than 50% of the prediction models points to the same classification, then use the mean value of the 10 classification results as the evaluation result; otherwise, give feedback.
[0079] Based on the same inventive concept, the present invention also provides a cardiovascular disease risk assessment system based on a machine learning model, and the system is used to complete the above-mentioned cardiovascular disease risk assessment method based on a machine learning model.
[0080] Based on the same inventive concept, the present invention also provides a cardiovascular disease risk assessment server based on a machine learning model, and the server is used to complete the above-mentioned cardiovascular disease risk assessment method based on a machine learning model.
[0081] A computer program (also known as a program, software, software application, script, or code) can be written in any form of programming language, including assembly or interpreted languages, declarative or procedural languages, and it can be deployed in any form, including as a stand-alone program or as a module, component, subroutine, object, or other unit suitable for use in a computing environment. A computer program may or may not correspond to a file in a file system. The program can be stored as part of a file that holds other programs or data (such as one or more scripts in a markup language document), in a single file dedicated to the program, or in multiple cooperating files (such as files that store one or more modules, subroutines, or code portions). A computer program can be deployed to execute on one computer or on multiple computers located at one site or distributed across multiple sites and interconnected by a communication network.
[0082] Those skilled in the art should understand that the embodiments of the present invention can be provided as a method, a system, or a computer program product. Therefore, the present invention can take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present invention can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.
[0083] The present invention is described with reference to the flowcharts and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of the present application. It should be understood that each flow and / or block in the flowchart and / or block diagram, and the combination of flows and / or blocks in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing devices to generate a machine, such that the instructions executed by the processor of the computer or other programmable data processing devices produce a means for implementing the functions specified in one flow Figure 1 one flow or multiple flows and / or blocks Figure 1 or in one block or multiple blocks.
[0084] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing device to work in a specific manner, such that the instructions stored in the computer-readable memory produce a manufactured article including an instruction means that implements the functions specified in one flow Figure 1 one flow or multiple flows and / or blocks Figure 1 or in one block or multiple blocks.
[0085] These computer program instructions can also be loaded onto a computer or other programmable data processing device, such that a series of operation steps are executed on the computer or other programmable device to generate a computer-implemented process, so that the instructions executed on the computer or other programmable device provide steps for implementing the functions specified in one flow Figure 1 one flow or multiple flows and / or blocks Figure 1 or in one block or multiple blocks.
[0086] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and not to limit them. Although the present invention has been described in detail with reference to the above embodiments, those of ordinary skill in the art should understand that: modifications or equivalent replacements can still be made to the specific embodiments of the present invention, and any modification or equivalent replacement that does not depart from the spirit and scope of the present invention should be covered by the protection scope of the claims of the present invention.
Claims
1. A method for assessing the risk of cardiovascular diseases based on a machine learning model, characterized in that, The method comprises: Step S1: construct I machine learning model as a prediction model, and select N user features as candidate input features; Step S2: Select N1 user features from the candidate input features as the first input features of the prediction model; Step S3: training and verifying the prediction model based on the N1 user features of the sample, and setting evaluation parameters to evaluate the prediction model; Step S4: constructing a discrimination sequence associated with each two classifications based on the evaluation parameters; the discrimination sequence includes a prediction model identifier, and the recognition ability of distinguishing between the two classifications is ranked according to the prediction model; Step S5: Obtain I classification results R corresponding to I prediction models for N1 user features of the user to be evaluated i =(r i,k ); where: r i,k is the k-th element of the classification result of the i-th prediction model; determine the classification pointed to by the maximum element value max(r i,k ) in each of the I classification results. If the maximum element values in the classification results of the I prediction models all point to the same classification, take the mean of the I classification results as the evaluation result. If M1 prediction models greater than or equal to PER1 point to the same classification k1, determine the classification pointed to by the maximum element value in the classification results of the prediction models that do not point to the same classification, which is called classification k2, obtain the identification sequence corresponding to the classification group (k1, k2), perform identification based on the identification sequence, and take the identification result as the evaluation result.
2. The cardiovascular disease risk assessment method based on a machine learning model according to claim 1, wherein I=5, 5 machine learning models are 5 different machine learning models.
3. The method for cardiovascular disease risk assessment based on a machine learning model according to claim 2, wherein The five machine learning models are LR, GB, NNK, GNB, and SVM machine learning models.
4. The cardiovascular disease risk assessment method based on a machine learning model according to claim 3, wherein N=29。 5. The cardiovascular disease risk assessment method based on a machine learning model according to claim 4, wherein N1=20。 6. The method for cardiovascular disease risk assessment based on a machine learning model according to claim 5, characterized in that, Arrange the N1 user features in sequence to form an N-dimensional vector as the first input feature of the machine learning model.
7. The cardiovascular disease risk assessment method based on a machine learning model according to claim 6, characterized in that, PER1=50%.
8. A method for cardiovascular disease risk assessment based on a machine learning model according to claim 7, characterized in that, The identification is performed based on the identification sequence, specifically: when the maximum element value of the M1 prediction results points to the same classification k1, the classification result of the I-M1 prediction model sorted last in the identification sequence is deleted; if all the classification results that have not been deleted point to the same classification, the mean of all the classification results that have not been deleted is used as the evaluation result; otherwise, set I=M1 and return to step S5.
9. The cardiovascular disease risk assessment method based on a machine learning model according to claim 8, characterized in that, Before training the prediction model, the sample data used for training is preprocessed and all samples with missing values are removed.
10. A cardiovascular disease risk assessment system based on a machine learning model, characterized in that, The cardiovascular disease risk assessment system based on a machine learning model is used to implement the cardiovascular disease risk assessment method based on a machine learning model as described in any one of claims 1 to 9.
Citation Information
Patent Citations
Cardiovascular disease non-planned re-hospitalization risk prediction method
CN110347837A
Multispectral riverway remote sensing monitoring method based on semi-supervised learning
CN112084843A
Alzheimer disease risk assessment method and device
CN114418966A
Chronic disease risk assessment method and system based on multi-model algorithm
CN115602325A
Risk prediction model for severe fever with thrombocytopenia syndrome as well as construction method and application of risk prediction model
CN119092124A