Prediction method, model and system for perioperative cerebral apoplexy

By establishing a perioperative database and using machine learning methods to build a prediction model, the problem of predicting stroke risk after craniocerebral surgery is solved, efficient and reliable prediction results are achieved, and the gap in the existing technology is filled.

CN120072292APending Publication Date: 2025-05-30THE THIRD AFFILIATED HOSPITAL OF SUN YAT SEN UNIV
View PDF 0 Cites 4 Cited by

Patent Information

Application Number
CN202510107590.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-23
Publication Date
2025-05-30

AI Technical Summary

Technical Problem

The prior art is difficult to accurately predict the risk of stroke after craniocerebral surgery, and there is a lack of effective prediction methods and models.

Method used

By establishing a perioperative database, combining machine learning methods to build a prediction model, using preoperative, intraoperative and postoperative clinical sample data for preprocessing and deletion, selecting multiple machine learning algorithms for model training, and finding the best parameter combination through grid search and K-fold cross-validation, and finally selecting the best performance prediction model.

Benefits of technology

It has achieved scientific and reliable predictions of postoperative stroke risk in patients with craniocerebral surgery, with high specificity, good model performance and high sensitivity, filling the gap in the existing technology that lacks effective prediction methods and models.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120072292A_ABST
    Figure CN120072292A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of risk prediction, in particular to a prediction method, model and system for perioperative cerebral apoplexy. The prediction model construction method for postoperative cerebral apoplexy comprises the following steps: collecting clinical sample data of a patient who has accepted a selected craniocerebral operation, the clinical sample data including preoperative, intraoperative and postoperative data; randomly dividing the collected clinical sample data into a training set and a verification set according to a proportion; preprocessing and deleting variables in the training set to obtain a first training set containing a plurality of variables; performing model training on the plurality of different machine learning models by using the first training set to obtain a plurality of candidate models; and using the verification set to predict and compare the performance of the plurality of candidate models, and selecting a prediction model with the best performance as a final prediction model. According to the method, the blank that in the prior art, an effective prediction method and a prediction model for postoperative cerebral apoplexy of a patient subjected to a selective craniocerebral operation are lacked is filled up.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of risk prediction, and particularly to a prediction method, model and system for perioperative stroke. Background Art

[0002] Stroke is a disease caused by various reasons leading to cerebrovascular problems and resulting in focal or overall brain tissue damage. Stroke is one of the most serious complications after craniocerebral surgery, which can significantly affect the patient's rehabilitation process and quality of life. Accurately evaluating the possibility of a patient having a stroke before surgery, thus formulating personalized surgical and nursing plans and optimizing perioperative management, helps reduce the potential risk of stroke. Therefore, it is extremely important to identify high-risk patients as early as possible. However, the current clinical routine scoring scales cannot predict the risk of stroke in postoperative patients, and a tool that can predict the risk of postoperative stroke in patients through preoperative and intraoperative data is needed to actively intervene in high-risk patients at an early stage. Summary of the Invention

[0003] The purpose of the present invention is to establish a prediction model for postoperative stroke in intracranial surgery by combining a perioperative database with machine learning methods to predict the risk of postoperative stroke in patients undergoing elective craniocerebral surgery. To achieve the purpose of the present invention, the following technical solutions are adopted:

[0004] The first aspect of the present invention proposes a method for constructing a prediction model for postoperative stroke, and the method includes the following steps:

[0005] Collect clinical sample data of patients who have undergone elective craniocerebral surgery, where the clinical sample data covers preoperative, intraoperative, and postoperative data;

[0006] Randomly divide the collected clinical sample data into a training set and a validation set according to a ratio;

[0007] Preprocess and select variables in the training set to obtain a first training set containing several variables;

[0008] Use the first training set to train several different machine learning models to obtain several candidate models;

[0009] Use the validation set to predict and compare the performance of several candidate models, and select the candidate model with the best performance as the final prediction model.

[0010] A further improvement lies in that the specific method for preprocessing and selecting variables in the training set to obtain a first training set containing several variables includes:

[0011] Exclude cases with incomplete data;

[0012] Perform mapping processing on categorical variables and map them to binary values;

[0013] The standard deviation standardization method was used to standardize the data for continuous variables;

[0014] Use the mean to fill in the missing values ​​of continuous variables and the mode to fill in the missing values ​​of categorical variables to ensure data integrity;

[0015] Delete variables with missing values ​​exceeding the preset value;

[0016] Pearson correlation coefficients were calculated between variables, and highly correlated variables were removed;

[0017] The data were transformed by YeoJohnson to make the distribution closer to normality;

[0018] Through centering and normalization transformation, the scales of different features are made consistent;

[0019] The LASSO sparse variables in the training set are screened out through LASSO regression to obtain the first training set containing several variables.

[0020] A further improvement is that the specific method of using the first training set to perform model training on several different machine learning models to obtain several candidate models includes:

[0021] Selecting several different machine learning algorithms to build several basic models, wherein the several different machine learning algorithms include at least one or more of a logistic regression algorithm, a decision tree algorithm, a Gaussian Bayes algorithm, a random forest algorithm, a gradient boosting algorithm, an adaptive boosting algorithm, a neighbor algorithm, an extreme gradient boosting algorithm, and a guided clustering algorithm;

[0022] Use grid search combined with K-fold cross validation to find the best parameter combination for each base model;

[0023] The first training set and the best parameter combination are used to train the corresponding basic model to obtain several candidate models.

[0024] A further improvement is that the specific method of using the grid search method combined with K-fold cross validation to find the best parameter combination for each base model includes:

[0025] Set the parameters and parameter value lists that need to be grid searched for each basic model, and perform cross-combinations;

[0026] The first training set is randomly divided into K equal subsets, and each equal subset is used in turn as the test set. The remaining K-1 equal subsets are used as the second training set to train and test each basic model under a specific parameter combination. Each basic model generates K evaluation indicators under each parameter combination, and the model score of each basic model under different parameter combinations is obtained according to the evaluation indicators.

[0027] Determine the best parameter combination for each base model according to the model score.

[0028] A further improvement is that the specific method for obtaining the model scores of each base model under different parameter combinations according to the evaluation index is: taking the average of K evaluation indexes as the model score of the corresponding base model under the corresponding parameter combination.

[0029] A further improvement is that when using the validation set to predict and compare the performance of several candidate models, the comparison indexes include at least one or more of the area under the ROC curve, accuracy, sensitivity, specificity, and F1 score.

[0030] A further improvement is that the method further includes the following steps:

[0031] Use permutation importance or partial dependence plots to evaluate the impact of each feature on the model prediction for the final prediction model, and make charts to visually display and help understand which factors have a significant impact on the occurrence of postoperative stroke.

[0032] A further improvement is that the method further includes the following steps:

[0033] Deploy the final prediction model to the clinical environment for real-time prediction of the postoperative stroke risk of new patients;

[0034] Continuously monitor the performance of the final prediction model in actual applications, collect new data and update the model regularly to ensure its long-term effectiveness and accuracy.

[0035] The second aspect of the present invention proposes a prediction model construction system for postoperative stroke, including:

[0036] A data acquisition module for collecting clinical sample data of patients who have undergone elective craniocerebral surgery, wherein the clinical sample data covers pre-operative, intra-operative, and post-operative data;

[0037] A data division module for randomly dividing the collected clinical sample data into a training set and a validation set according to a ratio;

[0038] A preprocessing and screening module for preprocessing and screening the variables in the training set to obtain a first training set containing several variables;

[0039] A first processing module for training several different machine learning models using the first training set to obtain several candidate models;

[0040] A second processing module for predicting and comparing the performance of several candidate models using the validation set, and selecting the best-performing prediction model as the final prediction model.

[0041] In the third aspect of the present invention, a prediction system for postoperative stroke is proposed. The system includes an input device, a processor, and a computer-readable medium. The computer-readable medium stores a plurality of instructions. The input device is used to obtain the measured values of relevant detection indexes of the patient undergoing craniocerebral surgery to be measured. The processor is connected to the input device and is used to process the data obtained by the input device and output the predicted value of the stroke risk. The instructions direct the input device and the processor to execute the method for predicting stroke after craniocerebral surgery. The method includes the following steps: 1) Obtain the measured values of 5 indexes, namely, the preoperative plasma albumin level, ASA classification, preoperative blood routine hemoglobin level, plasma albumin / globulin ratio, and total bilirubin level, of the patient undergoing craniocerebral surgery. 2) Standardize the measured values of the 5 indexes in step 1), load the trained logistic regression model, and input the result parameters of the standardized 5 indexes into the trained logistic regression model for calculation to obtain the predicted value of the stroke risk.

[0042] Advantages of the present invention:

[0043] The present invention provides a scientific, reliable, highly specific, good model performance, and highly sensitive model for predicting early postoperative stroke in patients undergoing craniocerebral surgery, as well as a construction method and system, filling the gap in the prior art of lacking effective prediction methods and prediction models for postoperative stroke in elective craniocerebral surgery patients.

[0044] The present invention predicts postoperative stroke in patients undergoing craniocerebral surgery by incorporating preoperative and intraoperative variables, and is expected to predict postoperative stroke in patients undergoing craniocerebral surgery in future clinical applications, which is helpful for early decision-making on stroke in clinical work. Brief Description of the Drawings

[0045] Figure 1 It is a flowchart of a method for constructing a prediction model for perioperative stroke in an embodiment of the present invention;

[0046] Figure 2 It is an overall flowchart of the construction of the prediction model in an embodiment of the present invention;

[0047] Figure 3 It is an ROC curve graph of each candidate model for postoperative stroke;

[0048] Figure 4 It is a fan-shaped graph of the feature contribution degree of the LR model for postoperative stroke. Detailed Embodiments

[0049] To enable those skilled in the art to better understand the solution of the present invention, the technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without making creative efforts shall fall within the protection scope of the present invention.

[0050] It should be noted that the terms "first", "second", etc. in the specification and claims of the present invention and the above-mentioned drawings are used to distinguish similar objects, and do not necessarily need to be used to describe a specific order or sequence. It should be understood that such data can be interchanged under appropriate circumstances so that the embodiments of the present invention described herein can be implemented in an order other than those illustrated or described herein. In addition, the terms "comprising" and "having" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product or device comprising a series of steps or units does not necessarily have to be limited to those steps or units clearly listed, but may include other steps or units not clearly listed or inherent to these processes, methods, products or devices.

[0051] The following is an explanation of the terms such as the algorithm model and performance indicators involved in the present invention:

[0052] Logistic Regression (LR): Logistic regression is a generalized linear model used for binary or multi-class classification problems. It predicts the class label by estimating the probability of an event occurring. Logistic regression assumes a linear relationship between features and log-odds.

[0053] Decision Tree (DT): A decision tree is a tree-structured model that gradually divides data into different subsets through a series of conditional judgments (nodes) and finally reaches the leaf nodes for classification or regression prediction. Each internal node represents a test on a feature, each branch represents a test result, and each leaf node represents a class or value.

[0054] Gaussian Naive Bayes (GNB): Gaussian Naive Bayes is a classification algorithm based on Bayes' theorem. It assumes conditional independence between features and assumes that features follow a Gaussian distribution. It predicts the class by calculating the posterior probability.

[0055] Random Forest (RF): Random Forest is an ensemble learning method that improves the performance of the model by building multiple decision trees and aggregating their prediction results. Each tree uses a randomly selected subset of features and a randomly sampled subset of data during training.

[0056] Gradient Boosting (GB): Gradient Boosting is an iterative ensemble learning method that corrects the errors of the previous model by gradually adding new models. Each new model focuses on reducing the residual of the previous model, usually using a decision tree as the base learner.

[0057] Adaptive Boosting (AdaBoost): Adaptive boosting is an iterative ensemble learning method that gradually adjusts the model to pay more attention to misclassified samples by assigning different weights to training samples. After each iteration, misclassified samples are given higher weights so that subsequent models pay more attention to these samples.

[0058] K-Nearest Neighbors (KNN): K-Nearest Neighbors is an instance-based learning method that calculates the distance between the sample to be predicted and all samples in the training set, selects the K nearest samples (neighbors), and performs classification or regression prediction based on the majority class or average of these neighbors.

[0059] Extreme Gradient Boosting (XGBoost): XGBoost is an optimized version of gradient boosting that introduces regularization terms to prevent overfitting and improves training speed through parallelization and cache optimization. It also supports sparse data processing and automatic feature selection.

[0060] Bootstrap Aggregating (Bagging): Bootstrap Aggregating is an ensemble learning method that randomly selects multiple subsets from the training set (with replacement sampling), trains multiple base learners separately, and finally aggregates their prediction results by voting or averaging.

[0061] AUC (area under the ROC curve): a comprehensive indicator to measure the distinguishing ability of the model, ranging from 0.5 to 1.0. The higher the value, the stronger the distinguishing ability of the model.

[0062] Accuracy: The proportion of correct predictions, but it may not perform well on imbalanced datasets.

[0063] Sensitivity: Also known as recall, it refers to the proportion of true positives that are correctly identified.

[0064] Specificity: It refers to the proportion of true negatives that are correctly identified.

[0065] F1 score: It is the harmonic mean of precision and recall, and is particularly suitable for classification problems with imbalanced classes.

[0066] Grid Search is an exhaustive search method that traverses all possible combinations in the specified parameter space to find the best parameter settings.

[0067] K-fold Cross-Validation is to divide the training set into K mutually exclusive subsets (or "folds"), and then alternately use each subset as the validation set, and the remaining K-1 subsets as the training set for K times of training and validation.

[0068] Please refer to the appendix Figure 1 - Appendix Figure 4 In the first aspect of the embodiments of the present invention, a prediction method, model, and system for perioperative stroke are proposed. Generally speaking, it realizes the prediction of postoperative stroke by constructing a prediction model for postoperative stroke. Among them, in this embodiment, as Figure 1 shown, a method for constructing a prediction model for postoperative stroke includes the following steps:

[0069] Step S1: Collect clinical sample data of patients who have undergone elective craniotomy. Among them, the clinical sample data covers pre-operative, intra-operative, and post-operative data to ensure a comprehensive reflection of the patient's health status and surgical process.

[0070] Specifically, clinical sample data (raw data) can be obtained from the perioperative database of the hospital. Comprehensive data collection helps improve the accuracy and reliability of the model because it can capture various factors affecting stroke, including but not limited to medical history, physiological indicators, surgical type, etc.

[0071] Step S2: Randomly divide the collected clinical sample data into a training set and a validation set according to a certain proportion.

[0072] Specifically, in this embodiment, to obtain reliable evaluation and avoid overfitting, the collected clinical sample data is randomly divided into a training set (80%) and a validation set (20%) for subsequent model training and performance evaluation. By using an independent validation set, the generalization ability of the model, that is, the performance of the model on unseen data, can be objectively evaluated, thereby avoiding overfitting and ensuring that the model has good external validity. Of course, those skilled in the art can also adjust the ratio according to actual needs, such as a training set (70%) and a validation set (30%). Patients' postoperative stroke can be diagnosed based on imaging changes and discharge diagnosis.

[0073] Step S3: Preprocess and select variables in the training set to obtain a first training set containing several variables.

[0074] It can be understood that by preprocessing and selecting training set variables, noise can be effectively reduced, redundant features can be eliminated, the stability and prediction accuracy of the model can be improved, the interpretability and generalization ability of the model can be enhanced, and a solid foundation can be laid for constructing an efficient perioperative stroke prediction model.

[0075] Specifically, in this embodiment, 34 features (variables) of the randomly assigned training set and validation set are included for preprocessing and selection. The included features are: gender, preoperative ICU admission, drug allergy, renal function, liver function, five coagulation items, preoperative hypertension, preoperative diabetes, ASA, age, BMI, neutrophil percentage NEUT, total white blood cell count WBC, hemoglobin concentration HGB, platelet count PLT, erythrocyte sedimentation rate ESR, troponin I cTnI, activated partial thromboplastin time APTT, prothrombin standardized ratio PT, fibrinogen concentration Fib, calcium ion concentration, sodium ion concentration, potassium ion concentration, glucose GLU, albumin ALB, alanine aminotransferase ALT, aspartate aminotransferase AST, total bilirubin TBILI, blood urea nitrogen BUN, estimated glomerular filtration rate eGFR, albumin / globulin A / G, creatine phosphokinase CK, myoglobin, C-reactive protein CRP. Finally, 16 variables are left after screening.

[0076] Step S4: Use the first training set to train several different machine learning models to obtain several candidate models.

[0077] It can be understood that different types of models may be good at capturing different types of data patterns. By comparing multiple models, the most suitable model for solving a specific problem can be selected, thereby improving the prediction performance.

[0078] Step S5: Use the validation set to predict and compare the performance of several candidate models, and select the candidate model with the best performance as the final prediction model.

[0079] It is understandable that the final prediction model should be the one that can exhibit the best comprehensive performance on the validation set. This method ensures that the selected model not only performs well on the training data but also has reliable prediction ability on new data, which is crucial for stroke prediction in practical applications.

[0080] Through the above method steps, the present invention aims to construct an effective and reliable perioperative stroke prediction model for realizing the prediction of perioperative stroke, providing a scientific basis for medical decision-making, helping doctors take preventive measures in advance, and reducing the risk of stroke occurrence.

[0081] In a preferred solution of this embodiment, in step S3, the specific method for preprocessing and screening variables in the training set to obtain the first training set containing several variables includes:

[0082] Step S31: Exclude cases with incomplete data.

[0083] It is understandable that excluding cases with incomplete data can ensure data quality and reduce the deviation caused by insufficient information.

[0084] Step S32: Perform mapping processing on categorical variables and map them to binary values.

[0085] It is understandable that converting non-numerical features (binary categorical variables) to binary values (such as "yes" / "no", "true" / "false") facilitates the processing of machine learning algorithms. For example, map "yes", "true", "male" to 1, and map "no", "false", "female" to 0.

[0086] Step S33: Perform data standardization processing on continuous variables using the standard deviation normalization method.

[0087] Specifically, the calculation formula for standard deviation normalization is x' = (x - mean) / std, where x' is the value after standardization, x is the original value, mean is the mean value, and std is the standard deviation. Through data standardization processing, the influence of features with different scales on the model can be made consistent, making the model training more stable.

[0088] Step S34: Use the mean value to fill in the missing values of continuous variables and use the mode to fill in the missing values of categorical variables to ensure data integrity.

[0089] It is understandable that by using the mean value to fill in the missing values of continuous variables and using the mode (the value with the highest occurrence frequency) to fill in the missing values of categorical variables, data integrity can be ensured, and as many samples as possible can be retained without affecting the model to reduce the impact of data missing on the model.

[0090] Step S35: Delete variables with missing values exceeding a preset value.

[0091] Specifically, in this embodiment, the preset value is 30%. It can be understood that if the missing ratio of a variable exceeds 30%, it is considered that the information it provides is insufficient to support effective prediction, so it is removed from the analysis. Specifically, in this embodiment, after the processing of this step, 8 variables including BMI, neutrophil percentage NEUT, erythrocyte sedimentation rate ESR, cardiac troponin I cTnI, estimated glomerular filtration rate eGFR, creatine phosphokinase CK, myoglobin, and C-reactive protein CRP are removed.

[0092] Step S36: Calculate the Pearson correlation coefficient between variables and remove highly correlated variables.

[0093] It can be understood that too strong a correlation will affect the performance of the trained machine learning model. When there is a high correlation (Pearson correlation coefficient > 0.7) between two or more variables, they may provide redundant information. In this case, one variable with stronger prediction performance will be retained, and other related variables will be removed to reduce the problem of multicollinearity. In this embodiment, after the processing of this step, one variable with poor prediction performance and a correlation greater than 0.7, aspartate aminotransferase AST, is deleted.

[0094] Step S37: Perform YeoJohnson transformation on the data to make the distribution closer to normal.

[0095] It can be understood that YeoJohnson transformation is a normalization transformation method that can be applied to both non-negative and negative values, and it can help improve the distribution characteristics of variables.

[0096] Step S38: Through centering and normalization transformation, make the scales of different features consistent.

[0097] It can be understood that through centering and normalization transformation, the variable scales are further adjusted to ensure that all features change within the same range, which is beneficial to the convergence of optimization algorithms such as gradient descent.

[0098] Step S39: Screen out the variables that are LASSO sparse in the training set through LASSO regression to obtain a first training set containing several variables.

[0099] It is understandable that LASSO (Least Absolute Shrinkage and Selection Operator) is a linear regression method. The characteristic of LASSO regression is to perform variable screening and complexity adjustment while fitting a generalized linear model. It shrinks the coefficients of unimportant features by introducing a penalty term (L1 regularization), and even compresses some coefficients to zero, thereby achieving automated feature selection.

[0100] After the processing of steps S31 - S39, there are finally 16 variables left for modeling: gender, preoperative ICU admission, renal function, liver function, preoperative diabetes mellitus, ASA, age, total white blood cell count WBC, hemoglobin concentration HGB, prothrombin standardized ratio PT, fibrinogen concentration Fib, potassium ion concentration, albumin ALB, alanine aminotransferase ALT, total bilirubin TBILI, albumin / globulin A / G.

[0101] In a preferred solution of this embodiment, in step S4, the specific method for training a number of different machine learning models using the first training set to obtain a number of candidate models includes:

[0102] Step S41: Select several different machine learning algorithms to construct a number of basic models.

[0103] Step S42: Use the grid search method combined with K - fold cross - validation to find the best parameter combination for each basic model.

[0104] Step S43: Use the first training set and the best parameter combination to train the corresponding basic models to obtain a number of candidate models.

[0105] It is understandable that by selecting multiple types of machine learning algorithms to construct basic models, different algorithms have their own advantages and disadvantages and are suitable for different types of data and problems. By comparison, the most suitable model for the current dataset and prediction task can be found. By using K - fold cross - validation to find the best parameter combination for each basic model, the limited data can be utilized more fully and a more robust estimate of the model performance can be provided. By combining grid search with K - fold cross - validation, the parameter space can be systematically explored while maintaining the data utilization rate, so as to find the best parameter combination for each basic model. After determining the best parameter combination for each basic model, the entire first training set and these best parameters are used to retrain the model. This step ensures that the model is optimized on all available training data, rather than just based on the subset during the cross - validation process.

[0106] Specifically, in this embodiment, in step S42, the specific method of using the grid search method combined with K-fold cross-validation to find the best parameter combination for each base model includes:

[0107] Step S421: Set the parameters and parameter value lists that each base model needs to perform grid search, perform cross-combinations, and generate all possible parameter combinations.

[0108] Step S422: Randomly divide the first training set into K equal subsets, and take turns using each equal subset as the test set, and the remaining K-1 equal subsets as the second training set to train and test each base model under a specific parameter combination. Each base model generates K evaluation indicators under each parameter combination, and the model scores of each base model under different parameter combinations are obtained according to the evaluation indicators.

[0109] Simply put, for each round (k = 1, 2,..., K), take turns selecting a subset as the test set, merge the remaining K-1 subsets into the second training set, use the second training set to train the base model, and evaluate it with the test set, and record the evaluation indicators of this round (such as accuracy, F1 score, AUC-ROC, etc.). In this embodiment, the value of K is preferably 5.

[0110] Step S423: Judge the best parameter combination of each base model according to the model score.

[0111] Specifically, in step S422, the specific method of obtaining the model scores of each base model under different parameter combinations according to the evaluation indicators is: taking the average value of the K evaluation indicators as the model score of the corresponding base model under the corresponding parameter combination.

[0112] It can be understood that taking the average value of the K evaluation indicators as the model score of the corresponding base model under the corresponding parameter combination, the model score reflects the overall level of the model performance under this parameter combination. At the same time, the K-fold cross-validation reduces the accidental error caused by a single division and improves the reliability of the scoring.

[0113] In this embodiment, the several different machine learning algorithms in step S41 include at least one or more of logistic regression algorithm, decision tree algorithm, Gaussian Bayesian algorithm, random forest algorithm, gradient boosting algorithm, adaptive boosting algorithm, nearest neighbor algorithm, extreme gradient boosting, and agglomerative algorithm. Preferably, in this embodiment, the above 9 machine learning algorithms are included at the same time, and a total of 9 candidate models are generated in step S43.

[0114] Specifically, in step S5, when predicting and comparing the performance of several candidate models using a validation set, the comparison metrics include at least one or more of the area under the ROC curve (AUC), accuracy, sensitivity, specificity, and F1 score.

[0115] It can be understood that by testing 9 candidate models with a validation set, the performance of each candidate model can be obtained, and the best-performing prediction model can be selected as the final prediction model.

[0116] It can be understood that the ROC curve is a curve plotted with the true positive rate as the ordinate and the false positive rate as the abscissa. The AUC is defined as the area enclosed by the ROC receiver operating characteristic curve and the coordinate axes, and is obtained by summing the areas of each part under the ROC curve. As a prediction metric, the AUC reflects the overall (average) performance, so the accuracy is high. At the same time, the sensitivity, specificity, and accuracy of the four models are also evaluated: Sensitivity (also known as the true positive rate) refers to the proportion of samples that are actually positive and are judged to be positive (the ability to correctly judge actual diseased cases as diseased, that is, the probability that a patient is judged to be positive); Specificity (also known as the true negative rate) refers to the proportion of samples that are actually negative and are judged to be negative (the ability to correctly judge actual non-diseased cases, that is, the proportion of test results that are negative); Accuracy, also known as efficiency, is expressed as the percentage of the sum of the number of true positives and true negatives to the number of subjects tested. The larger all evaluation metrics are, the better the result is, and the larger the value threshold (0, 1) is, the better the prediction performance of the model is.

[0117] In this embodiment, after testing, the performance parameters of 9 candidate models are shown in Table 1 below:

[0118]

[0119] Table 1

[0120] From the data in Table 1, it can be seen that among the 9 candidate models, the logistic regression (LR) model has the best overall performance in predicting postoperative stroke in patients undergoing craniotomy, with an AUC of 0.741, an accuracy of 66.8%, and a sensitivity of 65.0%.

[0121] In a preferred solution of this embodiment, the method further includes the following steps:

[0122] Step S6: Use permutation importance or partial dependence plots to evaluate the impact of each feature on the model prediction for the final prediction model, and make charts to visually display and help understand which factors have a significant impact on the occurrence of postoperative stroke.

[0123] It should be understood that permutation importance is a method for measuring the contribution of features to the prediction performance of a model. It evaluates the importance of a feature by randomly shuffling the values of a single feature and then measuring the degree of decline in model performance (such as evaluation metrics like accuracy, AUC, etc.). If the model performance significantly declines, it indicates that the feature is very important for the prediction result; otherwise, it indicates that the influence of the feature is relatively small. Partial dependence plots are used to show the relationship between one or two features and the model prediction, while ignoring the influence of other features. It can help understand how features individually or interactively affect the prediction result.

[0124] Specifically, in this embodiment, finally, the top 5 important indicators for risk assessment of postoperative stroke in patients undergoing craniotomy in the early stage are obtained, such as Figure 4 shown, the top five in the important features of the LR model are "preoperative plasma albumin level", "ASA grade", "preoperative hemoglobin level in blood routine", "plasma albumin / globulin ratio", and "total bilirubin level", enabling anesthesiologists to perform early intervention during the operation.

[0125] In a preferred solution of this embodiment, the method further includes the following steps:

[0126] Step S7: Deploy the final prediction model into the clinical environment for real-time prediction of the postoperative stroke risk of new patients.

[0127] Step S8: Continuously monitor the performance of the final prediction model in actual application, collect new data and update the model regularly to ensure its long-term effectiveness and accuracy.

[0128] A second aspect of the embodiment of the present invention proposes a prediction model construction system for postoperative stroke, which is used to execute a prediction model construction method for postoperative stroke proposed in the first aspect of the embodiment, including:

[0129] A data acquisition module, which is used to collect clinical sample data of patients who have undergone elective craniotomy, wherein the clinical sample data covers preoperative, intraoperative, and postoperative data.

[0130] A data division module, which is used to randomly divide the collected clinical sample data into a training set and a validation set according to a ratio.

[0131] A preprocessing and screening module, which is used to preprocess and screen the variables in the training set to obtain a first training set containing several variables.

[0132] A first processing module, which is used to train several different machine learning models using the first training set to obtain several candidate models.

[0133] A second processing module, configured to use a validation set to predict and compare the performance of a number of candidate models, and select the candidate model with the best performance prediction as the final prediction model.

[0134] In a third aspect of the embodiments of the present invention, a prediction system for postoperative stroke is proposed. The system includes an input device, a processor, and a computer-readable medium. The computer-readable medium stores a plurality of instructions. The input device is configured to obtain measurement values of relevant detection indexes of a patient undergoing a craniocerebral operation to be measured. The processor is connected to the input device, and the processor is configured to process the data obtained by the input device and output a prediction value of the stroke risk. The instructions direct the input device and the processor to execute a method for predicting stroke after craniocerebral surgery. The method includes the following steps: 1) Obtain measurement values of five indexes, namely, the preoperative plasma albumin level, ASA grade, preoperative hemoglobin level in blood routine, plasma albumin / globulin ratio, and total bilirubin level, of a patient undergoing a craniocerebral operation; 2) Standardize the measurement values of the five indexes in step 1), load a trained logistic regression model, and input the result parameters of the standardized five indexes into the trained logistic regression model for calculation to obtain a prediction value of the stroke risk. In addition, the training method of the logistic regression model adopts that described in the first aspect of the embodiments of the present invention.

[0135] The present invention provides a scientific, reliable, highly specific, good model performance, and high-sensitivity model that can predict the early stage of postoperative stroke in patients undergoing craniocerebral surgery, as well as a construction method and system, filling the blank in the prior art of lacking effective prediction methods and prediction models for postoperative stroke in elective craniocerebral surgery patients.

[0136] The present invention predicts postoperative stroke in patients undergoing craniocerebral surgery by incorporating preoperative and intraoperative variables, and is expected to predict postoperative stroke in patients undergoing craniocerebral surgery in future clinical applications, contributing to the early decision-making of stroke in clinical work.

[0137] The above-described embodiments are only used to illustrate the technical solutions of the present application, and are not intended to limit the same. Although the present application has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that: they can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements on some of the technical features; and these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present application, and should all be included in the protection scope of the present application.

Claims

1. A method for constructing a prediction model for postoperative stroke, characterized in that: The method comprises the following steps: Collect clinical sample data of patients who have undergone elective cranial surgery, wherein the clinical sample data covers preoperative, intraoperative and postoperative data; The collected clinical sample data were randomly divided into training set and validation set according to the proportion; Preprocess and delete the variables in the training set to obtain a first training set containing several variables; Using the first training set to perform model training on several different machine learning models to obtain several candidate models; The validation set is used to compare the performance of several candidate models, and the prediction model with the best performance is selected as the final prediction model.

2. A method for constructing a prediction model for postoperative stroke according to claim 1, characterized in that: The specific method of preprocessing and deleting the variables in the training set to obtain the first training set containing several variables includes: Cases with incomplete data were excluded; Map the categorical variables into binary values; The standard deviation standardization method was used to standardize the data for continuous variables; Use the mean to fill in the missing values ​​of continuous variables and the mode to fill in the missing values ​​of categorical variables to ensure data integrity; Delete variables with missing values ​​exceeding the preset value; Pearson correlation coefficients were calculated between variables, and highly correlated variables were removed; The data were transformed by YeoJohnson to make the distribution closer to normality; Through centering and normalization transformation, the scales of different features are made consistent; The LASSO sparse variables in the training set are screened out through LASSO regression to obtain the first training set containing several variables.

3. The method for constructing a prediction model for postoperative stroke according to claim 1, characterized in that: The specific method of using the first training set to perform model training on several different machine learning models to obtain several candidate models includes: Selecting several different machine learning algorithms to build several basic models, wherein the several different machine learning algorithms include at least one or more of a logistic regression algorithm, a decision tree algorithm, a Gaussian Bayes algorithm, a random forest algorithm, a gradient boosting algorithm, an adaptive boosting algorithm, a neighbor algorithm, an extreme gradient boosting algorithm, and a guided clustering algorithm; Use grid search combined with K-fold cross validation to find the best parameter combination for each base model; The first training set and the best parameter combination are used to train the corresponding basic model to obtain several candidate models.

4. A method for constructing a prediction model for postoperative stroke according to claim 3, characterized in that: The specific method of using grid search combined with K-fold cross validation to find the best parameter combination for each base model includes: Set the parameters and parameter value lists that need to be grid searched for each basic model, and perform cross-combinations; The first training set is randomly divided into K equal subsets, and each equal subset is used in turn as the test set. The remaining K-1 equal subsets are used as the second training set to train and test each basic model under a specific parameter combination. Each basic model generates K evaluation indicators under each parameter combination, and the model score of each basic model under different parameter combinations is obtained according to the evaluation indicators. The best parameter combination for each base model is determined based on the model score.

5. A method for constructing a prediction model for postoperative stroke according to claim 4, characterized in that: The specific method for obtaining the model score of each basic model under different parameter combinations according to the evaluation indicators is: taking the average value of K evaluation indicators as the model score of the corresponding basic model under the corresponding parameter combination.

6. The method for constructing a prediction model for postoperative stroke according to claim 1, characterized in that: When using the validation set to compare the performance of several candidate models, the comparison indicators include at least one or more of the area under the ROC curve, accuracy, sensitivity, specificity and F1 score.

7. The method for constructing a prediction model for postoperative stroke according to claim 1, characterized in that: The method further comprises the following steps: The final prediction model was evaluated using permutation importance or partial dependence plots to assess the impact of each feature on model prediction, and a graph was made to visually display the factors that significantly affected the occurrence of postoperative stroke.

8. The method for constructing a prediction model for postoperative stroke according to claim 1, characterized in that: The method further comprises the following steps: The final prediction model will be deployed in a clinical setting for real-time prediction of postoperative stroke risk in new patients; Continuously monitor the performance of the final prediction model in actual applications, collect new data and update the model regularly to ensure its long-term effectiveness and accuracy.

9. A prediction model construction system for postoperative stroke, characterized in that: include: A data acquisition module, used to collect clinical sample data of patients who have undergone elective cranial surgery, wherein the clinical sample data includes preoperative, intraoperative and postoperative data; A data partitioning module is used to randomly divide the collected clinical sample data into a training set and a validation set according to a certain proportion; A preprocessing and filtering module is used to preprocess and filter the variables in the training set to obtain a first training set containing several variables; A first processing module is used to perform model training on a plurality of different machine learning models using a first training set to obtain a plurality of candidate models; The second processing module is used to use the verification set to predict and compare the performance of several candidate models, and select the prediction model with the best performance as the final prediction model.

10. A prediction system for postoperative stroke, characterized in that: The system includes an input device, a processor and a computer-readable medium, wherein the computer-readable medium stores a plurality of instructions, wherein the input device is used to obtain the measured values ​​of relevant detection indicators of the patient undergoing craniocerebral surgery; the processor is connected to the input device, and the processor is used to process the data obtained by the input device and output the predicted value of stroke risk; the instructions instruct the input device and the processor to execute a method for predicting stroke after craniocerebral surgery; the method includes the following steps: 1) obtaining the measured values ​​of five indicators of the patient undergoing craniocerebral surgery, namely, preoperative plasma albumin level, ASA grade, preoperative hemoglobin level in routine blood tests, plasma albumin / globulin ratio, and total bilirubin level; 2) standardizing the measured values ​​of the five indicators in step 1), loading a trained logistic regression model, and inputting the result parameters of the five standardized indicators into the trained logistic regression model for calculation to obtain the predicted value of stroke risk.

Citation Information

Cited By

  • Postoperative cerebral apoplexy analysis method, system and device based on multi-modal image analysis

    CN120727293A

  • Postoperative stroke analysis method, system and device based on multi-modal image analysis

    CN120727293B

  • Prediction method and system for postoperative early-stage bad results of craniopharyngeal tubuloma patient

    CN120895266A

  • Method and system for optimizing diabetic complication prediction model

    CN121768690A