VTE risk prediction method and system based on gradient increasing
By using a gradient incremental algorithm to screen VTE-related features and calculate risk thresholds, a machine learning model was established. This solved the problems of one-sided feature selection and insufficient interpretability in existing technologies, and enabled accurate prediction of VTE risk and improved clinical acceptance.
Patent Information
- Application Number
- CN202511021908.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-24
- Publication Date
- 2025-11-11
- Estimated Expiration
- 2045-07-24
AI Technical Summary
Existing machine learning-based VTE risk prediction methods suffer from biased feature selection and insufficient model interpretability, resulting in low clinical acceptance.
A gradient-incremental VTE risk prediction method was adopted. Highly correlated features were screened by Pearson correlation coefficient, and risk thresholds were calculated by combining Youden index method, ROC curve and kernel density estimation method. A machine learning model was established using gradient-incremental algorithm to automatically learn feature weights and perform clinical consistency verification.
It achieves accurate prediction of VTE risk, improves model interpretability and clinical acceptability, and provides intelligent prediction results that combine accuracy and clinical applicability.
Smart Images

Figure CN120932877A_ABST
Abstract
Description
Technical Field
[0001] This application belongs to the field of smart medical big data disease risk prediction technology, specifically, it relates to a VTE risk prediction method and system based on gradient increment. Background Technology
[0002] Venous thromboembolism (VTE), including deep vein thrombosis (DVT) and pulmonary embolism (PE), is the third leading cause of vascular death worldwide, frequently occurring in surgical patients, long-term bedridden individuals, and those with specific lifestyles. Its early symptoms are often insidious, and without timely intervention, it can lead to fatal complications such as pulmonary embolism. Therefore, accurate risk prediction is crucial for clinical decision-making. Currently, widely used risk assessment scales in clinical evaluation primarily rely on subjective ratings of patient symptoms and medical history by healthcare professionals, assigning weights and calculating risk levels through pre-defined rules. While these methods are simple to operate, the assessment factors are often designed based on expert experience, making it difficult to cover complex clinical variables; furthermore, the lack of data-driven optimization in weight allocation leads to significant differences in assessment results among different medical institutions, exhibiting clear limitations.
[0003] The limitations of traditional assessment methods have spurred research into VTE prediction based on machine learning. While existing machine learning models can handle nonlinear relationships and integrate multidimensional data, most studies rely solely on linear statistical methods such as the Pearson correlation coefficient to screen features, potentially overlooking key factors that are nonlinearly associated with VTE, resulting in biased feature selection. Furthermore, risk thresholds for numerical features such as D-dimer are often calculated independently based on data distribution, without incorporating diagnostic criteria from clinical guidelines. Consequently, their predictions are disconnected from actual treatment procedures, leading to insufficient medical interpretability and low clinical acceptance of existing machine learning models. Summary of the Invention
[0004] To address the issues of limited feature selection and insufficient model interpretability in existing machine learning-based VTE risk prediction methods, which lead to low clinical acceptance, this application provides a gradient-based VTE risk prediction method.
[0005] In one approach, a gradient-incremental VTE risk prediction method comprises the following steps: S1. Collect sample feature data groups by exporting or accessing APIs through the hospital's internal system; S2. Perform irrelevant feature removal, missing value processing, variable type conversion, and normalization on the collected sample feature data; S3. The processed sample characteristic data were analyzed for VTE disease correlation using Pearson correlation coefficient. S4. Calculate the risk threshold for numerical sample features, and compare the calculation of the risk threshold for numerical sample features using the Youden index method, ROC curve and kernel density estimation method. S5. Use the gradient increment algorithm to build a machine learning model and train the model; S6. Test the predictive performance of the model and conduct clinical consistency verification; S7. Input the user's individual sample feature data and use the gradient incremental model to predict the user's VTE risk.
[0006] Specifically, in step S2, the basic sample feature data related to VTE prediction obtained after deleting irrelevant features includes 15 items: age, plasma D-dimer, white blood cells, platelets, hemoglobin, neutrophil percentage, hypertension, coronary heart disease, tumor, gender, history of surgery, diabetes, history of blood transfusion, smoking history, and history of comorbidities. Among them, age, white blood cell count, hemoglobin, platelet count, percentage of neutrophils, and plasma D-diglycol are numerical sample characteristics.
[0007] Furthermore, in step S3, after calculating the Pearson correlation coefficient of each basic sample feature data, eight highly correlated sample features related to VTE disease are obtained: age, plasma D-dimer, platelets, hemoglobin, coronary heart disease, tumor, gender, and past comorbidities.
[0008] Furthermore, in step S4, the risk thresholds for each numerical sample characteristic obtained by the Youden index method are as follows: age 80 years, white blood cell count 8.07×10⁹ / L, hemoglobin 78g / L, platelet count 307×10⁹ / L, neutrophil percentage 72%, and plasma D-dimer 284ng / mL DDU.
[0009] The risk thresholds for each numerical sample characteristic obtained from the ROC curve were as follows: age 80 years, white blood cell count 6.14×10⁹ / L, hemoglobin 126g / L, platelet count 222×10⁹ / L, neutrophil percentage 68.3%, and plasma D-dimer 256ng / mL DDU.
[0010] The risk thresholds for each numerical sample characteristic obtained by the nuclear density estimation method are as follows: age 80 years, white blood cell count 8.69×10⁹ / L, hemoglobin 131.1g / L, platelet count 197×10⁹ / L, neutrophil percentage 69.8%, and plasma D-dimer 965ng / mL DDU.
[0011] Furthermore, the risk thresholds and accuracy of numerical sample characteristics obtained by the Youden index method, ROC curve, and kernel density estimation method were compared, and the comprehensive risk thresholds were obtained by combining the curve fluctuations of the kernel density curve in the corresponding numerical region: age 65 years, white blood cell count 8.69×10⁹ / L, hemoglobin 126g / L, platelet count 197×10⁹ / L, neutrophil percentage 72%, and plasma D-dimer 965ng / mL DDU.
[0012] Furthermore, in step S5, two different VTE risk prediction models are established using XGBoost or LightGBM algorithms, respectively, with 15 basic sample feature data and 8 highly correlated sample feature data as inputs. The initial prediction value is 0.5, and the initial disease probability of all test subjects is assumed to be 50%. The prediction of VTE disease probability of test subjects is gradually adjusted and accumulated by combining the sample feature data of test subjects.
[0013] In one approach, the model training process is as follows: Initialize the model, set the initial prediction score to 0.5 as the initial prediction value (i.e., the initial probability value) for all samples, and specify that a decision tree should be used as the base learner; After training begins, based on the parameter settings, instead of column sampling, all features are used for iteration. In each iteration, the gradient residual between the current model prediction and the true label is calculated. The new tree fits these residuals by minimizing the objective function, which includes information about the first derivative-gradient and the second derivative-Hessian matrix. Tree node splitting is strictly controlled, and splitting only occurs when the loss reduction from splitting is at least 4.5. The sum of the sample weights of child nodes is limited to 1 to prevent overfitting. Leaf node weights are smoothed using the algorithm's default L2 regularization. After the tree is built, the contribution of the new tree is scaled according to a learning rate of 0.1 to prevent overfitting, and the extreme changes in the weights of the leaf nodes are further constrained to enhance the stability of the model. Repeat the iteration until training stops when the performance on the validation set no longer improves. Finally, the prediction results of all trees are weighted and combined to form the final model.
[0014] Furthermore, in step S6, the predictive performance of the model is evaluated using four evaluation metrics: ROC curve, AUC value, Matthews correlation coefficient, and Kappa coefficient.
[0015] Furthermore, to achieve the above objectives, this application also provides a gradient-incremental VTE risk prediction system. The system is used to implement the aforementioned gradient-incremental VTE risk prediction method and includes: The data acquisition module is used to obtain the basic sample feature data set for VTE prediction through the hospital's internal system or API. The data preprocessing module is used to perform irrelevant feature removal, missing value handling, variable type conversion, and normalization. The sample feature analysis module includes a correlation analysis unit that quantifies the correlation of VTE disease with basic sample feature data based on Pearson correlation coefficient, and a risk threshold calculation unit that calculates thresholds for numerical sample features. The model building and evaluation module takes the output data from the sample feature analysis module as input, builds a machine learning model based on the gradient increment algorithm, and evaluates the model's performance on the test set. The risk prediction module receives the individual characteristic data of the test subject and calls the gradient incremental model to calculate the probability of VTE risk.
[0016] Furthermore, the data acquisition module is also equipped with an anonymization processing unit and adopts an encrypted transmission mechanism; when collecting basic sample feature data, direct identification data involving patient personal identity information is deleted or hashed, and quasi-identifiers indirectly associated with identity features are desensitized.
[0017] One or more technical solutions provided in this application have at least the following technical effects or advantages: This application combines Pearson correlation coefficient and gradient-incremental model feature importance ranking to perform dual screening of feature data related to VTE disease, while capturing both linear and nonlinear correlation features. It utilizes the Youden index method, ROC curves, and kernel density estimation method, combined with clinical guideline standards, to calculate risk thresholds, achieving dynamic calibration of numerical features and avoiding thresholds driven by single data that are detached from clinical practice. Based on a machine learning model constructed using the gradient-incremental algorithm, it automatically learns feature weights and adapts to data distribution. Through feature importance analysis, it transforms the model's decision logic into interpretable medical evidence, assisting physicians in decision-making and providing intelligent predictive results for early VTE screening that are both accurate and clinically practical. Attached Figure Description
[0018] To more clearly illustrate the technical solutions of the embodiments of this application, the drawings used in the description of the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0019] Figure 1 This is a flowchart of a gradient-increment-based VTE risk prediction method in one embodiment of this application; Figure 2 This is a heatmap of correlation analysis in one embodiment of this application; Figure 3 This is a density curve of age in one embodiment of this application; Figure 4 This is a white blood cell density curve in one embodiment of this application; Figure 5 This is a density curve of hemoglobin in one embodiment of this application; Figure 6 This is a platelet density curve in one embodiment of this application; Figure 7 This is a density curve of the percentage of neutrophils in one embodiment of this application; Figure 8 This is a density curve of plasma D-dimer in one embodiment of this application; Figure 9 This is a ROC curve of the model prediction performance in one embodiment of this application, with 15 basic sample features as input. Figure 10 This is a ROC curve of the model prediction performance in one embodiment of this application, with eight highly correlated sample features as input. Detailed Implementation
[0020] To make the objectives, technical solutions, and advantages of the embodiments of this application clearer, the technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, not all embodiments. The modules of the embodiments of this application described and shown in the accompanying drawings can be arranged and designed in various different configurations. Therefore, the following detailed description of the embodiments of this application provided in the accompanying drawings is not intended to limit the scope of the claimed application, but merely to illustrate selected embodiments of this application. All other embodiments obtained by those skilled in the art based on the embodiments of this application without creative effort are within the scope of protection of this application.
[0021] Unless otherwise defined, all technical and scientific terms used in this application have the same meaning as commonly understood by one of ordinary skill in the art to which this application pertains. The terminology used herein is for the purpose of describing particular embodiments only and is not intended to be limiting of the application.
[0022] To address the issues of limited feature selection and insufficient model interpretability in existing machine learning-based VTE risk prediction methods, leading to low clinical acceptance, this application provides a gradient-based VTE risk prediction method. Please refer to [link to relevant documentation]. Figure 1 The method includes the following steps: S1. Collect sample feature data groups by exporting or accessing APIs through the hospital's internal system.
[0023] In step S1, the hospital's internal system or access API is used to obtain a group of original sample feature data containing medical element information from the hospital, providing source data for subsequent correlation analysis and prediction model establishment.
[0024] S2. Perform irrelevant feature removal, missing value processing, variable type conversion, and normalization on the collected sample feature data.
[0025] In step S2, the raw sample feature data obtained from the hospital contains many sample features irrelevant to VTE prediction, such as department and patient number. Before performing correlation analysis, the raw sample feature data is preprocessed by deleting irrelevant sample feature items, resulting in basic sample feature data containing only 15 items relevant to VTE prediction: age, plasma D-dimer, white blood cell count, platelet count, hemoglobin, neutrophil percentage, hypertension, coronary heart disease, tumor, gender, previous surgical history, diabetes, blood transfusion history, smoking history, and past comorbidities. Among these, age, white blood cell count, hemoglobin, platelet count, neutrophil percentage, and plasma D-dimer are numerical sample features.
[0026] In step S2, there are still some missing values in the original sample feature data that are not completely random. To avoid model prediction errors caused by these missing values, it is necessary to process them. Common methods for processing missing values include direct deletion, mean imputation, and mode imputation. To improve the accuracy of model prediction, this application chooses mean imputation to process the missing values. For example, the missing values in the column of sample feature "coronary heart disease" are filled using the mean of this column, 1.59396.
[0027] Furthermore, for the gender, past comorbidities, past surgical history, blood transfusion history and smoking history in the basic sample feature data, variable type conversion is required. This application uses binary variables for conversion, and the specific conversion results are shown in Table 1.
[0028] Table 1 Results of Basic Sample Feature Variable Type Conversion Furthermore, numerical sample feature data such as age and neutrophil percentage are normalized, transforming the feature values to the specific interval [0, 1] to eliminate the influence of different dimensions between features. This makes the numerical sample feature data numerically comparable, which helps improve the performance and convergence speed of the machine learning model. Specifically, a min-max normalization method is used to linearly map the values of each numerical sample feature data to the range [0, 1], preserving the original distribution characteristics of each numerical sample feature data.
[0029] S3. The processed sample feature data were analyzed for VTE disease correlation using Pearson correlation coefficient.
[0030] In step S3, correlation is used to reflect the degree of association between each sample feature and VTE incidence. In machine learning, it can be used to select features that have a significant impact on the target variable. Highly correlated sample features may contribute significantly to the model's predictive ability, while low-correlation sample features can be considered for removal to improve model performance and training efficiency. Specifically, correlation is usually measured by the correlation coefficient. This application uses the sample Pearson correlation coefficient to measure the correlation of basic sample features. The Pearson correlation coefficient is a commonly used statistical indicator used to measure the degree of linear association between two continuous variables. The correlation heatmap of 15 basic sample features was obtained by calculating the Pearson correlation coefficient matrix, as shown below. Figure 2 As shown.
[0031] Furthermore, the weights of corresponding sample features are measured based on their contribution to reducing the impurity of the decision tree. Each sample feature generates a reduction in impurity at each node of the decision tree. This reduction is weighted, and then the reductions over all nodes are summed. Finally, normalization is performed to obtain the weights of each basic sample feature. Specifically, the correlation and weights between each of the 15 basic sample features and VTE prediction are analyzed and calculated, as shown in Table 2.
[0032] Table 2 Summary of Correlation and Weights of Basic Sample Features Among them, items marked with "*" indicate sample characteristics that are highly correlated with the outcome of whether or not the patient is a VTE patient. The judgment criteria are a correlation coefficient > 0.05, that is, eight sample characteristics with high correlation to VTE are obtained: age, plasma D-dimer, platelets, hemoglobin, coronary heart disease, tumor, gender, and past comorbidities.
[0033] S4. Calculate the risk threshold of numerical sample features. Use the Youden index method, ROC curve and kernel density estimation method to compare and calculate the risk threshold of numerical sample features.
[0034] In step S4, the risk threshold is a numerical boundary that distinguishes between normal and abnormal. For numerical sample features, determining the risk threshold can provide a reference for doctors and machine learning models to judge the patient's disease risk. This application uses the Youden index method, ROC curve, and kernel density estimation method to calculate the risk thresholds of six numerical sample features, and comprehensively analyzes and selects the most suitable risk threshold.
[0035] The Youden index, also known as the accuracy index, is an important indicator in medicine and statistics used to evaluate the accuracy of binary diagnostic tests. It measures the test's ability to distinguish between "sick" and "non-sick" groups by comprehensively considering the sensitivity and specificity of the diagnostic test. The formula is: Youden Index = Sensitivity + Specificity - 1. The ROC curve, on the other hand, is an important tool in medical statistics and machine learning for evaluating the performance of binary classification models. It visualizes the relationship between the true positive rate (sensitivity) and false positive rate (specificity) of a model at different thresholds, helping researchers select the optimal decision threshold and compare the diagnostic capabilities of different models. The ROC curve plots the false positive rate (FPR) on the x-axis and the true positive rate (TPR) on the y-axis. The point in the upper left corner represents the ideal classification state: FPR = 0 – no false positives, all negative samples are correctly classified; TPR = 1 – all positive samples are correctly classified. In scenarios involving determining whether someone is sick, this means accurately identifying all sick individuals while avoiding misdiagnosing healthy individuals as sick.
[0036] Specifically, firstly, ROC curves for each basic sample feature are plotted. In the ROC curves, sensitivity and specificity change with the threshold. By calculating the Youden index under different thresholds, the threshold corresponding to the maximum Youden index is found. This threshold can be used as a better risk threshold. After obtaining the risk threshold, their accuracy is verified based on the original sample data. The results obtained by the Youden index method are shown in Table 3.
[0037] Table 3. Risk threshold results using the Youden index method Furthermore, since it is difficult to reach the ideal point at the top left corner of the ROC curve, the threshold closest to this ideal point is chosen as a better cutoff value, i.e., the risk threshold. The logic behind this cutoff value selection is to achieve a balance between maximizing the correct identification rate (TPR) for patients and minimizing the false positive rate (FPR) for healthy individuals. Intuitively, the closer a sample is to the top left corner of the ROC curve, the better its overall performance in distinguishing between the diseased and healthy populations. The risk threshold results obtained from the ROC curve are shown in Table 4.
[0038] Table 4. Risk threshold results for ROC curves Furthermore, kernel density estimation is a nonparametric statistical method used to estimate the probability density function of a random variable based on sample data. It constructs a continuous density curve by smoothly weighting the sample points using a "kernel function".
[0039] For details, please refer to Figures 3 to 8The kernel density estimation method was used to calculate the probability density gradient of sample feature values under different risk states, and the intersection of the density curves or the point where the density difference changes significantly was used as the risk threshold. This method is relatively more objective, as it can use the distribution information of the data to determine the threshold, and the kernel density curve is drawn based on real sample feature data, which has high reference value. The risk thresholds obtained by the kernel density estimation method are shown in Table 5.
[0040] Table 5. Risk threshold results using kernel density estimation method Furthermore, by Figure 3 The density curves of age show that the difference in curve density changes significantly between the ages of 60 and 65, and the curves intersect at the age of 80. Based on this, the accuracy rates for ages of 60, 65, and 80 were calculated, and the results are shown in Table 6.
[0041] Table 6. Accuracy calculation results for ages 60, 65, and 80. Furthermore, considering the accuracy of each threshold and the screening range, the optimal comprehensive risk threshold for age sample characteristics was finally determined to be 65 years old. Similarly, the optimal comprehensive risk thresholds for other numerical sample characteristics were obtained through analysis and comparison, as shown in Table 7.
[0042] Table 7 Comprehensive Risk Thresholds for Numerical Sample Features S5. Use the gradient increment algorithm to build a machine learning model and train the model.
[0043] In step S5, after data preprocessing, two machine learning models are established using 15 basic sample feature data and 8 highly correlated sample feature data as inputs. This application employs the gradient ascending algorithm for model building. The core principle of the gradient ascending model is to iteratively train multiple weak learners and combine their predictions into a strong learner through accumulation, thereby improving overall prediction performance. It is primarily based on the gradient ascent algorithm. In machine learning, for a differentiable objective function, the gradient ascent algorithm iteratively updates parameters along the gradient direction of the objective function, continuously increasing the value of the objective function. Therefore, the gradient ascending model can effectively utilize the gradient information of the objective function to find the optimal parameter values. Compared with methods such as random search, it can find the parameter values that maximize the objective function more quickly and accurately, thus improving model performance. Furthermore, for complex objective functions, especially in high-dimensional parameter spaces, the gradient ascending algorithm can gradually optimize the objective function by continuously updating parameters along the gradient direction, demonstrating good performance in machine learning and prediction.
[0044] For example, the gradient increasing model is built using the XGBoost algorithm, with the main parameters set as follows: (1) base_score=0.5, the initial predicted score for all samples. (2) booster='gbtree', select gradient boosting decision tree as the type of base learner. (3) colsample_bylevel=1, the sampling ratio (level) of columns during tree construction. (4) colsample_bynode=1, the sampling ratio of columns (nodes) during tree construction. (5) gamma=4.5, the minimum loss reduction required for tree node splitting. (6) learning_rate=0.1, learning rate, controls the step size of each iteration. (7) max_delta_step=3, limits the maximum step size for weight updates. (8) min_child_weight=1, the minimum sum of weights required for child nodes. (9) objective='binary:logistic', sets the objective function type to logistic regression for binary classification problems. The core of the model is to gradually correct errors using multiple small decision trees. In the initial stage, it is assumed that everyone has a 50% probability of developing VTE. The mission of each tree is to correct the error of the previous prediction. For example, if the first tree finds that patients with "plasma D-dimer > risk threshold" are more likely to develop the disease, it will increase the predicted probability of that patient from 0.5 to 0.6. The second tree may find that patients with "age > 60 years and smoking" have a higher risk, and further adjust the predicted value. The final prediction result is the sum of the adjusted values of all trees, which is then converted into a probability between 0 and 1 through a logistic function, that is, the probability of the test subject developing VTE.
[0045] Furthermore, the model training process is as follows: Initialize the model, set the initial prediction score to 0.5 as the initial prediction value (i.e., the initial probability value) for all samples, and specify that a decision tree should be used as the base learner; After training begins, based on the parameter settings, instead of column sampling, all features are used for iteration. In each iteration, the gradient residual between the current model prediction and the true label is calculated. The new tree fits these residuals by minimizing the objective function, which includes information about the first derivative-gradient and the second derivative-Hessian matrix. Tree node splitting is strictly controlled, and splitting only occurs when the loss reduction from splitting is at least 4.5. The sum of the sample weights of child nodes is limited to 1 to prevent overfitting. Leaf node weights are smoothed using the algorithm's default L2 regularization. After the tree is built, the contribution of the new tree is scaled according to a learning rate of 0.1 to prevent overfitting, and the extreme changes in the weights of the leaf nodes are further constrained to enhance the stability of the model. Repeat the iteration until training stops when the performance on the validation set no longer improves. Finally, the prediction results of all trees are weighted and combined to form the final model.
[0046] S6. Test the predictive performance of the model and conduct clinical consistency verification.
[0047] In step S6, to evaluate the prediction accuracy of the prediction model, evaluation metrics are introduced to analyze the accuracy and performance of the prediction model. In this embodiment, the prediction error is generally the difference between the model's predicted value and the actual value of the dataset. By evaluating the prediction results, it is reasonable to determine whether the prediction model has high accuracy and reliability. This application uses four evaluation metrics—ROC curve, AUC value, Matthews correlation coefficient, and Kappa coefficient—to evaluate the model's prediction results.
[0048] Among them, the AUC value is an important indicator used in machine learning and statistics to evaluate the performance of binary classification models. It is the area under the receiver operating characteristic curve, that is, the area under the ROC curve. In essence, it is the probability that the model will predict the positive sample as positive when a positive sample and a negative sample are randomly selected. The value ranges from 0.5 to 1. The closer the AUC value is to 1, the better the classification performance of the classifier.
[0049] The Matthews correlation coefficient is also a statistical metric used to evaluate the performance of binary classification models, especially suitable for imbalanced scenarios. It comprehensively considers four types of results: true positive (TP), true negative (TN), false positive (FP), and false negative (FN), providing a single numerical value that can summarize the classification quality and more robustly reflect the classification accuracy of the model.
[0050] The Kappa coefficient is a statistical indicator used to measure the consistency of classification results. It considers the proportion of consistent classification results while also correcting for the influence of random consistency, thus more objectively evaluating the reliability of the classifier, evaluator, or model. Its value ranges from -1 to 1, where 1 represents perfect classification (the predicted result is completely consistent with the actual classification); 0 represents classification accuracy equal to random guessing; and a negative value indicates classification accuracy lower than random guessing.
[0051] For details, please refer to Figure 9 and Figure 10 The evaluation results of the prediction performance of the two gradient incremental models with 15 basic sample feature data and 8 highly correlated sample feature data as input are shown in Table 8.
[0052] Table 8. Evaluation results of model prediction performance As shown in Table 8, both prediction models achieve good prediction results and have high accuracy. Based on the Matthews correlation coefficient and Kappa coefficient, the models exhibit good classification ability and high consistency with the actual labels. Figure 9 and Figure 10 It can be seen that both prediction models perform well under different thresholds. The prediction model with 15 basic sample features as input has better performance and can be regarded as the preferred model group.
[0053] S7. Input the user's individual sample feature data and use the gradient incremental model to predict the user's VTE risk.
[0054] In step S7, the individual sample feature data of the test users are received using the evaluated gradient incremental model. The entered data is then converted and normalized. The model is substituted into the processed data to perform the calculation process, and then provides personalized reports to the test users, such as the probability of VTE, risk level, deviation of key features, and ranking of feature importance, to support the doctor's clinical decision-making.
[0055] Therefore, this application combines Pearson correlation coefficient and gradient-incremental model feature importance ranking to perform dual screening of feature data related to VTE disease, while capturing both linear and nonlinear correlation features. It utilizes the Youden index method, ROC curves, and kernel density estimation method, combined with clinical guideline standards, to calculate risk thresholds, achieving dynamic calibration of numerical features and avoiding thresholds driven by single data that are detached from clinical practice. Based on a machine learning model constructed using the gradient-incremental algorithm, it automatically learns feature weights and adapts to data distribution. Through feature importance analysis, it transforms the model's decision logic into interpretable medical evidence, assisting physicians in decision-making and providing intelligent prediction results for early VTE screening that are both accurate and clinically practical.
[0056] As one implementation, this application also provides a gradient-incremental VTE risk prediction system, which is used to implement a gradient-incremental VTE risk prediction method in the foregoing embodiments. The system includes: The data acquisition module is used to obtain basic sample characteristic data for VTE prediction through the hospital's internal system or API, including patient age, plasma D-dimer, white blood cell count, platelet count, hemoglobin, neutrophil percentage, hypertension, coronary heart disease, tumor, gender, surgical history, diabetes, blood transfusion history, smoking history, and comorbidity information.
[0057] Furthermore, the data acquisition module is also equipped with an anonymization processing unit and adopts an encrypted transmission mechanism; when collecting basic sample feature data, direct identifier data involving patients' personal identity information, such as patient name, ID number, and medical insurance number, are deleted or hashed, and quasi-identifiers that are indirectly related to identity features, such as residential address and consultation timestamp, are desensitized to prevent the risk of leakage of patients' privacy data in the collection, transmission, and storage stages.
[0058] The data preprocessing module performs irrelevant feature removal, missing value handling, variable type conversion, and normalization. Before correlation analysis, the original sample feature data is preprocessed by removing irrelevant features to obtain basic sample feature data containing only 15 items relevant to VTE prediction: age, plasma D-dimer, white blood cell count, platelet count, hemoglobin, neutrophil percentage, hypertension, coronary heart disease, tumor, gender, past surgical history, diabetes, transfusion history, smoking history, and past comorbidities. Simultaneously, mean imputation is performed on numerical features, mode imputation is performed on categorical features, and normalization is applied to eliminate the influence of different dimensions between features.
[0059] The sample feature analysis module includes a correlation analysis unit that quantifies the correlation between basic sample feature data and VTE based on Pearson correlation coefficient, thereby screening a subset of highly correlated features, and a risk threshold calculation unit that calculates thresholds for numerical sample features. It combines Youden's index method, ROC curve analysis and kernel density estimation method, and integrates clinical diagnostic guidelines to generate dynamic risk thresholds.
[0060] The model building and evaluation module includes a model building unit and a model evaluation unit. The model building unit takes the output data from the sample feature analysis module as input and uses either XGBoost or LightGBM algorithms to build a gradient-increasing model. For binary classification tasks, log loss is used. The optimization objective is to minimize the deviation between the predicted probability and the true label. The residual learning logic involves calculating the residual of the current prediction result in each iteration, adding a new decision tree to fit the residual, and adjusting the weights of the split nodes and leaves of the tree through gradient descent to gradually approximate the true distribution. The model evaluation unit uses four evaluation metrics—ROC curve, AUC value, Matthews correlation coefficient, and Kappa coefficient—to evaluate the model's performance on the test set.
[0061] The risk prediction module is equipped with an interactive input unit and an interpretable report generation unit. The interactive input unit receives the individual characteristic data of the test subject and performs variable type conversion and normalization on the data. After the risk prediction module calls the gradient incremental model to calculate the test subject's VTE risk probability, the interpretable report generation unit outputs the VTE prevalence probability, feature importance ranking, local interpretability, and Sapley value analysis results to assist doctors in making clinical decisions.
[0062] In summary, this application effectively addresses the prediction bias and insufficient interpretability issues caused by the one-sided feature selection and threshold settings deviating from clinical practice in existing machine learning models by constructing a VTE risk prediction system that deeply integrates data-driven approaches with clinical knowledge. The system automatically captures nonlinear correlations between features based on a gradient incremental algorithm, combined with a dynamic risk threshold generation mechanism, overcoming the limitations of traditional linear screening methods and ensuring the accurate identification of key risk factors such as the interaction effect between plasma D-dimer and surgical history. Simultaneously, through visualized decision-making paths and clinical consistency verification, the model's prediction logic is transformed into medical evidence understandable to physicians, improving the credibility and clinical acceptance of the prediction results. Furthermore, the system supports dual-mode input and privacy-compliant processing, balancing prediction accuracy with practical feasibility. Through dynamic iterative optimization, it adapts to diverse medical scenarios, providing efficient, transparent, and compliant intelligent decision support for early VTE warning.
[0063] The above description of the disclosed embodiments enables those skilled in the art to make or use this application. Various modifications to these embodiments will be readily apparent to those skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of this application. Therefore, this application is not to be limited to the embodiments shown herein, but is to be accorded the widest scope consistent with the principles and novel features disclosed herein.
Claims
1. A gradient-incremental VTE risk prediction method, characterized in that, The method includes the following steps: S1. Collect sample feature data groups by exporting or accessing APIs through the hospital's internal system; S2. Perform irrelevant feature removal, missing value processing, variable type conversion, and normalization on the collected sample feature data; S3. The processed sample characteristic data were analyzed for VTE disease correlation using Pearson correlation coefficient. S4. Calculate the risk threshold for numerical sample features, and compare the calculation of the risk threshold for numerical sample features using the Youden index method, ROC curve and kernel density estimation method. S5. Use the gradient increment algorithm to build a machine learning model and train the model; S6. Test the predictive performance of the model and conduct clinical consistency verification; S7. Input the user's individual sample feature data and use the gradient incremental model to predict the user's VTE risk.
2. The VTE risk prediction method based on gradient increment according to claim 1, characterized in that, In step S2, after deleting irrelevant features, the basic sample feature data related to VTE prediction includes 15 items: age, plasma D-dimer, white blood cells, platelets, hemoglobin, neutrophil percentage, hypertension, coronary heart disease, tumor, gender, history of surgery, diabetes, history of blood transfusion, smoking history, and history of comorbidities. Among them, age, white blood cell count, hemoglobin, platelet count, percentage of neutrophils, and plasma D-diglycol are numerical sample characteristics.
3. The VTE risk prediction method based on gradient increment according to claim 1, characterized in that, In step S3, after calculating the Pearson correlation coefficient of each basic sample feature data, eight highly correlated sample features related to VTE disease are obtained: age, plasma D-dimer, platelets, hemoglobin, coronary heart disease, tumor, gender, and past comorbidities.
4. The VTE risk prediction method based on gradient increment according to claim 1, characterized in that, In step S4, The risk thresholds for each numerical sample characteristic obtained by the Youden index method are as follows: age 80 years, white blood cell count 8.07 × 10⁻⁶. 9 / L, hemoglobin 78g / L, platelets 307×10 9 / L, neutrophil percentage 72%, plasma D-dimer 284 ng / mL LDDU; The risk thresholds for each numerical sample characteristic obtained from the ROC curves are as follows: age 80 years, white blood cell count 6.14 × 10⁻⁶. 9 / L, hemoglobin 126g / L, platelets 222×10 9 / L, neutrophil percentage 68.3%, plasma D-dimer 256 ng / mL DDU; The risk thresholds for each numerical sample feature obtained by the nuclear density estimation method are as follows: age 80 years, white blood cell count 8.69 × 10⁻⁶. 9 / L, hemoglobin 131.1g / L, platelets 197×10 9 / L, neutrophil percentage 69.8%, plasma D-dimer 965ng / mL DDU.
5. The VTE risk prediction method based on gradient increment according to claim 4, characterized in that, By comparing the risk thresholds and accuracy rates of numerical sample characteristics obtained using the Youden index method, the ROC curve, and the kernel density estimation method, and by combining the curve fluctuations of the kernel density curve in the corresponding numerical region, a comprehensive risk threshold is obtained: age 65 years, white blood cell count 8.69 × 10⁻⁶. 9 / L, hemoglobin 126g / L, platelets 197×10 9 / L, neutrophil percentage 72%, plasma D-dimer 965ng / mL DDU.
6. The VTE risk prediction method based on gradient increment according to claim 3, characterized in that, In step S5, two different VTE risk prediction models are established using 15 basic sample feature data and 8 highly correlated sample feature data as inputs, respectively, and the XGBoost or LightGBM algorithm is used. The initial prediction value is 0.5, and the initial disease probability of all test subjects is assumed to be 50%. The prediction of VTE disease probability of test subjects is gradually adjusted and accumulated by combining the sample feature data of test subjects.
7. The VTE risk prediction method based on gradient increment according to claim 6, characterized in that, The training process for the model is as follows: Initialize the model, set the initial prediction score to 0.5 as the initial prediction value (i.e., the initial probability value) for all samples, and specify that a decision tree should be used as the base learner; After training begins, based on the parameter settings, instead of column sampling, all features are used for iteration. In each iteration, the gradient residual between the current model prediction and the true label is calculated. The new tree fits these residuals by minimizing the objective function, which includes information about the first derivative-gradient and the second derivative-Hessian matrix. Tree node splitting is strictly controlled, and splitting only occurs when the loss reduction from splitting is at least 4.
5. The sum of the sample weights of child nodes is limited to 1 to prevent overfitting. Leaf node weights are smoothed using the algorithm's default L2 regularization. After the tree is built, the contribution of the new tree is scaled according to a learning rate of 0.1 to prevent overfitting, and the extreme changes in the weights of the leaf nodes are further constrained to enhance the stability of the model. Repeat the iteration until training stops when the performance on the validation set no longer improves. Finally, the prediction results of all trees are weighted and combined to form the final model.
8. The VTE risk prediction method based on gradient increment according to claim 1, characterized in that, In step S6, the predictive performance of the model is evaluated using four evaluation metrics: ROC curve, AUC value, Matthews correlation coefficient, and Kappa coefficient.
9. A VTE risk prediction system based on gradient increment, characterized in that, The system is used to implement the gradient-incremental VTE risk prediction method according to any one of claims 1 to 8, the system comprising: The data acquisition module is used to acquire the basic sample feature data set for VTE prediction through the hospital's internal system or API; The data preprocessing module is used to perform irrelevant feature removal, missing value processing, variable type conversion, and normalization processing. The sample feature analysis module includes a correlation analysis unit that quantifies the VTE disease correlation of basic sample feature data based on Pearson correlation coefficient, and a risk threshold calculation unit that performs threshold calculation for numerical sample features. The model building and evaluation module uses the output data of the sample feature analysis module as input to build a machine learning model based on the gradient increment algorithm and evaluates the performance of the model on the test set. The risk prediction module receives the individual characteristic data of the test subject and calls the gradient incremental model to calculate the VTE risk probability.
10. The VTE risk prediction system based on gradient increment according to claim 9, characterized in that, The data acquisition module is also equipped with an anonymization processing unit and adopts an encrypted transmission mechanism; when collecting basic sample feature data, direct identification data involving patient personal identity information is deleted or hashed, and quasi-identifiers that are indirectly associated with identity features are desensitized.
Citation Information
Patent Citations
Risk behavior information prediction method and system based on gradient boosting decision tree
CN116502742A
Prediction and early warning method and system for medical instrument related pressure damage
CN119423712A
Apparatus and method for training an artificial intelligence-supported diagnostic assessment tool
US12333413B1