A method, device, equipment and medium for predicting the risk value of postoperative pulmonary complications
By combining the prediction models of XGBoost tree, LSTM and nmODE modules, the problem of insufficient modeling ability of neural networks to model heterogeneous tables in the prior art is solved, and a more accurate prediction of the risk of postoperative pulmonary complications is achieved, providing better auxiliary support for clinical practice.
Patent Information
- Application Number
- CN202311645628.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-12-04
- Publication Date
- 2025-06-20
- Estimated Expiration
- 2043-12-04
AI Technical Summary
The neural network of the existing postoperative pulmonary complication prediction model lacks the ability to model heterogeneous tables, affects the model performance, and the data source is incomplete and cannot fully reflect the risk of postoperative complications.
The XGBoost tree module, LSTM module and nmODE module were used to construct a predictive model of postoperative complications in lungs. The importance of feature is extracted through the XGBoost tree, and the LSTM analyzes feature dependence, and the memory and learning ability of nmODE are used to make the final prediction.
The model's ability to model heterogeneous tabular data is improved, and the accuracy of predicting postoperative complication risks is enhanced, providing clinicians with an important reference basis.
Smart Images

Figure CN117637167B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of artificial intelligence, relates to the field of postoperative risk value prediction, and particularly relates to a method, device, equipment and medium for predicting the risk value of postoperative pulmonary complications. Background Technique
[0002] Postoperative pulmonary complications refer to the occurrence of abnormal pulmonary conditions and dysfunction after surgery, which have a negative impact on the health and recovery of patients, may lead to cardiovascular and immune system problems, prolong the hospital stay and increase medical costs. With the development of the population and the progress of perioperative medicine, the number of patients with multiple diseases in the thoracic surgery department has been increasing. However, the incidence of postoperative pulmonary complications (PPCs) in thoracic surgery is still as high as 30.0% to 50.0%, which has a significant impact on the prognosis of patients. In order to be able to detect the risk factors of postoperative pulmonary complications at an early stage, identify high-risk groups, and take corresponding intervention measures to prevent and reduce the occurrence of postoperative pulmonary complications, a risk prediction model for postoperative pulmonary complications has emerged. By using the prediction model, precise and stratified management of patients can be achieved, the occurrence of postoperative pulmonary complications can be effectively prevented and reduced, and the treatment effect and recovery quality of surgical patients can be improved. The existing methods for predicting postoperative pulmonary complications mainly use statistical methods to establish a risk prediction model, using methods such as Logistic regression and Cox regression to analyze independent influencing factors and establish a model based on the regression coefficients.
[0003] The invention patent application with the application number 202110967700.8 discloses a system for constructing a risk prediction model for persistent air leakage after lung cancer resection, which includes: a data collection module, connected to the central control module, for collecting thoracic surgery cases and their related data from various hospitals, and the case-related data includes multiple clinical and physiological index data; a data processing module, connected to the central control module, for processing the collected corresponding cases and their related data; a central control module, connected to the data collection module, the data processing module, the classification and extraction module, the screening module, the feature extraction module, the model construction module, and the evaluation module, for processing data and using a single-chip microcomputer or a controller to control the normal operation of each module; a model construction module, connected to the central control module, for constructing a PAL risk prediction model after lung cancer resection based on the processed data and the feature extraction results; the construction of the PAL risk prediction model after lung cancer resection based on the processed data and the feature extraction results includes: variable screening according to the results of multicollinearity test, feature extraction results, and single-factor and multi-factor logistic regression screening results; drawing a Nomogram, drawing a characteristic curve, and determining the classification critical value according to the Youden index; and dividing the processed case data into a training set and an internal validation set according to a ratio of 2:1; constructing a PAL risk prediction model after lung cancer resection using ANN and RF; and training the constructed model using the training set; performing internal validation on the trained model based on random splitting of samples and cross-validation of the internal validation set; using other central data sets as external validation sets for external validation of the model; an evaluation module, connected to the central control module, for evaluating the model effect by calculating the discrimination and calibration; the evaluation module evaluating the model effect by calculating the discrimination and calibration includes: using C-index, accuracy, sensitivity, specificity, positive likelihood ratio, negative likelihood ratio, positive predictive value, and negative predictive value to describe the discrimination; and quantitatively evaluating the calibration of the model by drawing a calibration curve, Hosmer-Lemeshow goodness-of-fit test, and calculating the Brier score.
[0004] In the prior art, the neural network prediction of postoperative complications of the lungs requires complex manual analysis. The reason is that the data of postoperative complications of the lungs are usually in the form of tables, which contain a mixture of discrete and continuous variables, and the neural network lacks the ability to model heterogeneous tabular data, thus affecting the model performance; in addition, most of the existing prediction models only incorporate local surgical data and do not comprehensively consider the influence of seasonal factors on the same day, resulting in incomplete data sources for the prediction models and being unable to fully reflect the occurrence risk of postoperative complications of the lungs. Summary of the Invention
[0005] The object of the present invention is to provide a method, device, equipment and medium for predicting the risk value of postoperative pulmonary complications in order to solve the problems that the neural network in the existing prediction model lacks the ability to model heterogeneous tables, thus affecting the model performance, and the data source of the prediction model is not comprehensive enough to fully reflect the risk of postoperative pulmonary complications.
[0006] In order to achieve the above object, the present invention specifically adopts the following technical solutions:
[0007] A method for predicting the risk value of postoperative pulmonary complications, comprising the following steps:
[0008] Step S1, obtaining sample data;
[0009] Obtain the postoperative pulmonary sample data, and calibrate whether complications occur in the postoperative pulmonary sample data to form label data;
[0010] Step S2, constructing a risk value prediction model;
[0011] Construct a risk value prediction model, which includes an XGBoost tree module, an LSTM module and an nmODE module;
[0012] Step S3, training the risk value prediction model;
[0013] Use the sample data in step S1 to train the risk value prediction model constructed in step S2;
[0014] Step S4, real-time risk value prediction;
[0015] Obtain the real-time postoperative data and input it into the risk value prediction model, and the risk value prediction model outputs the predicted risk value.
[0016] Further, in step S1, the postoperative pulmonary sample data includes age group, gender number, patient's surgery date, daily air pressure, air pressure higher than the average, daily temperature, temperature higher than the average, daily humidity, humidity higher than the average, sunshine, sunshine higher than the average value, AQI, PM2.5, PM10, SO2, CO, NO2, O3_8h, AQI classification, 2.5 classification, 10 classification, SO2 higher than the average value, CO higher than the average value, NO2 higher than the dangerous value, O3 higher than the dangerous value, monthandseason, season number, length of hospital stay, hypertension, cardiovascular disease, COPD emphysema, diabetes, tumor history, other combined diseases, smoking history, whether still smoking, pulmonary function, MVV condition, diffusion function condition, BMI classification, open chest or minimally invasive, operation minutes, resection range number, number of indwelling roots.
[0017] Furthermore, the SMOTE algorithm is used to preprocess the sample data after lung surgery. The specific SMOTE algorithm is as follows:
[0018] Based on the K nearest neighbor samples of each sample, randomly select N neighboring points, and generate a threshold between 0 and 1 through interpolation to synthesize data; the input is the data set X, and the output is the sampled data set, with a size of (N / 100)*T;
[0019] Among them, T represents the number of samples to be processed, N represents the sampling ratio, and K represents the number of nearest neighbor samples.
[0020] Furthermore, in step S2, the XGBoost tree module outputs the weight corresponding to the sample according to the sample data. Specifically:
[0021] Step S2-1-1 constructs an XGBoost tree module including n decision trees, which is used to fit the sample data and calculate the split gain of each feature in each decision tree;
[0022] Step S2-1-2 obtains the score of each feature by analyzing the split gain of each feature in each decision tree;
[0023] Step S2-1-3 discards the samples with scores lower than the threshold, retains the samples with scores higher than the threshold, and outputs the weights of the samples.
[0024] Furthermore, in step S2, the LSTM module includes multiple LSTM sub-modules A using forward propagation and multiple LSTM sub-modules A` using backward propagation. The sample data and the weight of the sample output by the XGBoost tree module are used as the inputs of the LSTM sub-module A and the LSTM sub-module A`. After the outputs of the LSTM sub-module A and the LSTM sub-module A` are jump-connected with the corresponding sample data, they are used as the output of the LSTM module.
[0025] Furthermore, both the LSTM sub-module A and the LSTM sub-module A` are added with a peephole structure and use a coupled forget mechanism.
[0026] Furthermore, in step S2, the neural ordinary differential equation of the nmODE module is:
[0027]
[0028] Among them, represents the input of the nmODE module, represents the output of the data x in the LSTM module, represents the state of the nmODE module at time t.
[0029] A risk value prediction system for postoperative pulmonary complications, comprising:
[0030] A sample data acquisition module, configured to acquire postoperative pulmonary sample data and calibrate whether complications occur in the postoperative pulmonary sample data to form labeled data;
[0031] A risk value prediction model construction module, configured to construct a risk value prediction model, where the risk value prediction model includes an XGBoost tree module, an LSTM module, and an nmODE module;
[0032] A risk value prediction model training module, configured to train the risk value prediction model constructed by the risk value prediction model construction module using the sample data of the sample data acquisition module;
[0033] A risk value real-time prediction module, configured to acquire real-time postoperative data and input it into the risk value prediction model, and the risk value prediction model outputs the predicted risk value.
[0034] A computer device, comprising a memory and a processor, where the memory stores a computer program, and when the computer program is executed by the processor, the processor executes the steps of the above method.
[0035] A computer-readable storage medium, storing a computer program, and when the computer program is executed by a processor, the processor executes the steps of the above method.
[0036] The beneficial effects of the present invention are as follows:
[0037] 1. In the present invention, the XGBoost tree structure is used to encode data and extract the importance of features. Then, the improved LSTM is used to analyze the relevant feature dependencies in the data and extract the processed features. Finally, the memory and learning capabilities of the neural ordinary differential equation are utilized to analyze the extracted features and give the final prediction result. This method can solve the problems that the neural network in the existing prediction model lacks the ability to model heterogeneous tables, thus affecting the model performance, and the data source of the prediction model is not comprehensive enough to fully reflect the risk of postoperative pulmonary complications. It can accurately predict the risk value of postoperative pulmonary complications and provide important reference for clinicians.
[0038] 2. The present invention proposes a new neural network model, which combines the tree model and the neural network model, and improves the LSTM network and the nmODE module to predict postoperative complications. By comparing with other models on the dataset, it is proved that this model has excellent performance and can be used as a preoperative evaluation tool to guide the treatment of patients during the perioperative period, more effectively prevent the occurrence of postoperative pulmonary complications, and provide help for doctors.
[0039] 3. The trained model of the present invention has the ability of rapid detection and prediction, supports batch operations, and can improve the speed with the expansion of the device. This model can save the human and material resources required for primary screening, enabling doctors to concentrate their work on higher-level diagnosis and the design of treatment plans. In addition, the model can automatically process the backlogged data that has not been fully analyzed without human intervention. BRIEF DESCRIPTION OF THE DRAWINGS
[0040] Figure 1 is a schematic flowchart of the present invention;
[0041] Figure 2 is a schematic structural diagram of the risk value prediction model in the present invention;
[0042] Figure 3 is a schematic structural diagram of LSTM sub-module A and LSTM sub-module A' in the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0043] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention. Apparently, the described embodiments are some, but not all, of the embodiments of the present invention.
[0044] Therefore, based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.
[0045] Embodiment 1
[0046] This embodiment provides a method for predicting the risk value of postoperative pulmonary complications. First, surgical environment and season data are incorporated into the dataset, and a new network model is used to predict the risk value of postoperative pulmonary complications, which can provide more accurate and reliable prediction results and better assist clinical practice. As Figure 1 shown, the specific steps are as follows:
[0047] Step S1, obtain sample data;
[0048] Obtain the postoperative pulmonary sample data and calibrate whether complications occur in the postoperative pulmonary sample data to form label data.
[0049] The postoperative pulmonary sample data comes from West China Hospital, and a total of 5,000 cases of sample data are obtained.
[0050] The sample data after lung surgery contains 31 features that lead to complications, specifically including age group, gender code, patient surgery date, daily air pressure, air pressure above the mean, daily temperature, temperature above the mean, daily humidity, humidity above the mean, sunshine, sunshine above the average, AQI, PM2.5, PM10, SO2, CO, NO2, O3_8h, AQI classification, 2.5 classification, 10 classification, SO2 above the average, CO above the average, NO2 above the danger value, O3 above the danger value, month and season, season code, length of hospital stay, hypertension, cardiovascular disease, COPD emphysema, diabetes, tumor history, other co-morbidities, smoking history, whether still smoking, lung function, MVV, diffusion function, BMI classification, open chest or minimally invasive, surgical minutes, resection range code, number of indwelling roots.
[0051] In addition, due to the serious problem of sample imbalance in the sample data, the number of negative samples accounts for the vast majority, and untreated may lead to overfitting. Therefore, the sample data after lung surgery is preprocessed, and the SMOTE algorithm is used to preprocess the sample data after lung surgery. The specific steps of the SMOTE algorithm are as follows:
[0052] Based on the K nearest neighbor samples of each sample, randomly select N neighboring points, and generate a threshold between 0 and 1 through interpolation to synthesize data; the input is the data set X, and the output is the sampled data set, with a size of (N / 100)*T;
[0053] Among them, T represents the number of samples to be processed, N represents the sampling ratio, and K represents the number of nearest neighbor samples.
[0054] By using a combination method of oversampling and undersampling, better classification performance can be obtained in the ROC space than simple oversampling. This experiment uses a combination method of oversampling and undersampling. The ratio of positive and negative samples in the initial data is 4234:766. After sampling and undersampling, the ratio of positive and negative samples is adjusted to 2:1. At the same time, the data set is divided into a training set and a test set according to a ratio of 0.8:0.2.
[0055] Step S2, construct a risk value prediction model;
[0056] Construct a risk value prediction model, which includes an XGBoost tree module, an LSTM module, and an nmODE module. Specifically as Figure 2 shown.
[0057] The XGBoost tree module outputs the weights corresponding to the samples according to the sample data, specifically as:
[0058] Step S2-1-1 constructs an XGBoost tree module including n decision trees, which is used to fit the sample data and calculate the split gain of each feature in each decision tree.
[0059] Step S2-1-2 obtains the score of each feature by analyzing the split gain of each feature in each decision tree.
[0060] Step S2-1-3 discards the samples with scores lower than the threshold, retains the samples with scores higher than the threshold, and outputs the weights of the samples.
[0061] The LSTM module includes multiple LSTM sub-modules A using forward propagation and multiple LSTM sub-modules A' using backward propagation. The sample data and the weights of the sample output by the XGBoost tree module are used as the inputs of the LSTM sub-modules A and LSTM sub-modules A'. The outputs of the LSTM sub-modules A and LSTM sub-modules A' are jump-connected with the corresponding sample data and then used as the output of the LSTM module.
[0062] Both the LSTM sub-module A and the LSTM sub-module A' are added with a peephole structure and use a coupled forget mechanism, as specifically Figure 3 shown. A peephole structure is introduced in the LSTM sub-module. This structure allows the gate layer to receive the input of the cell state simultaneously, enabling the model to better utilize historical information. By introducing the information of the cell state into the gate layer, the model can better decide which information needs to be forgotten and which information needs to be added. In the LSTM sub-module, a coupled forget mechanism is also adopted. Different from the traditional LSTM, here the information to be forgotten and added is determined simultaneously instead of operating separately, which can improve the model's ability to model long-term dependencies.
[0063] The neural ordinary differential equation of the nmODE module is:
[0064]
[0065] where represents the input of the nmODE module, represents the output of the data x in the LSTM module, represents the state of the nmODE module at time t.
[0066] Step S3 trains the risk value prediction model.
[0067] Use the sample data in Step S1 to train the risk value prediction model constructed in Step S2.
[0068] The training method of the risk value prediction model and the loss function used in the training can be implemented using existing technologies. During the training process, the Adam optimizer and the BCELoss loss function were used, and the learning rate was set to 0.0001. The model was trained for 200 rounds.
[0069] It should be noted that different parameter selections during the training process may affect the performance of the model, and adjustments and optimizations can be made according to the actual situation. In this embodiment, metrics such as Accuracy, Precision, Roc, Recall, and F1-score were used to evaluate the model. The higher the values of these metrics, the better the performance of the model.
[0070] Step S4, real-time risk value prediction;
[0071] Obtain real-time postoperative data and input it into the risk value prediction model, and the risk value prediction model outputs the predicted risk value.
[0072] Embodiment 2
[0073] This embodiment provides a risk value prediction system for postoperative pulmonary complications, specifically including:
[0074] A sample data acquisition module, which is used to acquire postoperative pulmonary sample data and calibrate whether complications occur in the postoperative pulmonary sample data to form labeled data.
[0075] The postoperative pulmonary sample data comes from West China Hospital, and a total of 5000 cases of sample data were acquired.
[0076] The postoperative pulmonary sample data includes 31 features that cause complications, specifically including age group, gender code, patient's surgery date, daily air pressure, air pressure higher than the mean, daily temperature, temperature higher than the mean, daily humidity, humidity higher than the mean, sunshine, sunshine higher than the average, AQI, PM2.5, PM10, SO2, CO, NO2, O3_8h, AQI classification, 2.5 classification, 10 classification, SO2 higher than the average, CO higher than the average, NO2 higher than the dangerous value, O3 higher than the dangerous value, monthandseason, season code, length of hospital stay, hypertension, cardiovascular disease, COPD emphysema, diabetes, tumor history, other coexisting diseases, smoking history, whether still smoking, pulmonary function status, MVV status, diffusion function status, BMI classification, open chest or minimally invasive, operation minutes, resection range code, number of indwelling roots.
[0077] In addition, due to the serious problem of sample imbalance in the sample data, the number of negative samples accounts for the vast majority, which may lead to overfitting if not processed. Therefore, the sample data after lung surgery is preprocessed, and the SMOTE algorithm is used for data preprocessing of the sample data after lung surgery. The specific steps of the SMOTE algorithm are as follows:
[0078] Based on the K nearest neighbor samples of each sample, randomly select N neighboring points, and generate a threshold between 0 and 1 through interpolation to synthesize data; the input is the dataset X, and the output is the sampled dataset with a size of (N / 100)*T;
[0079] Among them, T represents the number of samples to be processed, N represents the sampling ratio, and K represents the number of nearest neighbor samples.
[0080] By using a combined method of oversampling and undersampling, better classification performance can be obtained in the ROC space than that of simple oversampling. This experiment adopted a combined method of oversampling and undersampling. The ratio of positive to negative samples in the initial data was 4234:766. After oversampling and undersampling, the ratio of positive to negative samples was adjusted to 2:1. At the same time, the dataset was divided into a training set and a test set according to a ratio of 0.8:0.2.
[0081] The risk value prediction model construction module is used to construct a risk value prediction model, and the risk value prediction model includes an XGBoost tree module, an LSTM module, and an nmODE module. Specifically as Figure 2 shown.
[0082] The XGBoost tree module outputs the weight corresponding to the sample according to the sample data. Specifically as follows:
[0083] Step S2-1-1: An XGBoost tree module including n decision trees is constructed to fit the sample data and calculate the split gain of each feature in each decision tree;
[0084] Step S2-1-2: By analyzing the split gain of each feature in each decision tree, the score of each feature is obtained;
[0085] Step S2-1-3: Discard the samples with scores lower than the threshold, retain the samples with scores higher than the threshold, and output the weights of the samples.
[0086] The LSTM module includes multiple LSTM sub-modules A using forward propagation and multiple LSTM sub-modules A` using backward propagation. The sample data and the weight of the sample output by the XGBoost tree module are used as the inputs of the LSTM sub-modules A and LSTM sub-modules A`. After the outputs of the LSTM sub-modules A and LSTM sub-modules A` are jump-connected with the corresponding sample data, they are used as the output of the LSTM module.
[0087] Both the LSTM sub-module A and the LSTM sub-module A' are added with peephole structures and use the coupled forget mechanism, specifically as follows Figure 3 shown. A peephole structure is introduced in the LSTM sub-module. This structure allows the gate layer to receive the input of the cell state simultaneously, enabling the model to better utilize historical information. By introducing the information of the cell state into the gate layer, the model can better decide which information needs to be forgotten and which information needs to be added. In the LSTM sub-module, a coupled forget mechanism is also adopted. Different from the traditional LSTM, here the information to be forgotten and added is determined simultaneously instead of operating separately, which can improve the model's ability to model long-term dependencies.
[0088] The neural ordinary differential equation of the nmODE module is:
[0089]
[0090] where, represents the input of the nmODE module, represents the output of the data x in the LSTM module, represents the state of the nmODE module at time t.
[0091] The risk value prediction model training module is used to train the risk value prediction model constructed by the risk value prediction model construction module using the sample data of the sample data acquisition module.
[0092] For the training method of this risk value prediction model and the loss function used in the training, existing technologies can be adopted. During the training process, the Adam optimizer and the BCELoss loss function are used, and the learning rate is set to 0.0001, and the model is trained for 200 rounds.
[0093] It should be noted that different parameter selections during the training process may affect the performance of the model, and adjustments and optimizations can be made according to the actual situation. In this embodiment, indicators such as Accuracy, Precision, Roc, Recall, and F1-score are used to evaluate the model, and the higher the values of these indicators, the better the performance of the model.
[0094] The risk value real-time prediction module is used to obtain real-time postoperative data and input it into the risk value prediction model, and the risk value prediction model outputs the predicted risk value.
[0095] Embodiment 3
[0096] A computer device includes a memory and a processor. The memory stores a computer program. When the computer program is executed by the processor, the processor is caused to execute the steps of a method for predicting the risk value of postoperative pulmonary complications.
[0097] Among them, the computer device can be a computing device such as a desktop computer, a notebook, a palm computer, and a cloud server. The computer device can perform human-computer interaction with the user through a keyboard, a mouse, a remote control, a touchpad, or a voice control device, etc.
[0098] The memory includes at least one type of readable storage medium. The readable storage medium includes flash memory, a hard disk, a multimedia card, a card-type memory (such as an SD or D-interface display memory, etc.), a random access memory (RAM), a static random access memory (SRAM), a read-only memory (ROM), an electrically erasable programmable read-only memory (EEPROM), a programmable read-only memory (PROM), a magnetic memory, a magnetic disk, an optical disk, etc. In some embodiments, the memory can be an internal storage unit of the computer device, such as the hard disk or memory of the computer device. In other embodiments, the memory can also be an external storage device of the computer device, such as a plug-in hard disk equipped on the computer device, a Smart Media Card (SMC), a Secure Digital (SD) card, a Flash Card, etc. Of course, the memory can also include both the internal storage unit and the external storage device of the computer device. In this embodiment, the memory is commonly used to store the operating system installed on the computer device and various application software, such as the program code of the method for predicting the risk value of postoperative pulmonary complications. In addition, the memory can also be used to temporarily store various data that have been output or will be output.
[0099] In some embodiments, the processor can be a Central Processing Unit (CPU), a controller, a microcontroller, a microprocessor, or other data processing chips. The processor is generally used to control the overall operation of the computer device. In this embodiment, the processor is used to run the program code stored in the memory or process data, such as running the program code of the method for predicting the risk value of postoperative pulmonary complications.
[0100] Embodiment 4
[0101] A computer-readable storage medium stores a computer program. When the computer program is executed by a processor, the processor is caused to execute the steps of a method for predicting the risk value of postoperative pulmonary complications.
[0102] Among them, the computer-readable storage medium stores an interface display program, and the interface display program can be executed by at least one processor, so that the at least one processor executes the steps of the risk value prediction method for postoperative pulmonary complications as described above.
[0103] Through the description of the above embodiments, those skilled in the art can clearly understand that the methods of the above embodiments can be implemented by means of software plus a necessary general hardware platform. Of course, they can also be implemented by hardware, but in many cases the former is a better implementation. Based on such an understanding, the technical solution of the present application, in essence, or the part that contributes to the prior art can be embodied in the form of a software product. The computer software product is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk), and includes several instructions for causing a terminal device (which can be a mobile phone, computer, server or network device, etc.) to execute the risk value prediction method for postoperative pulmonary complications described in the embodiments of the present application.
Claims
1. A method for predicting the risk value of postoperative pulmonary complications, characterized in that, It includes the following steps: Step S1, obtaining sample data; Obtaining sample data after lung surgery and calibrating whether complications occur in the sample data after lung surgery to form labeled data; Step S2, constructing a risk value prediction model; Constructing a risk value prediction model, which includes an XGBoost tree module, an LSTM module, and an nmODE module; Step S3, training the risk value prediction model; Using the sample data in Step S1 to train the risk value prediction model constructed in Step S2; Step S4, real-time prediction of risk value; Obtaining real-time postoperative data and inputting it into the risk value prediction model, and the risk value prediction model outputs the predicted risk value; In Step S2, the XGBoost tree module outputs the weight corresponding to the sample according to the sample data, specifically: Step S2-1-1, constructing an XGBoost tree module including n decision trees, which is used to fit the sample data and calculate the split gain of each feature in each decision tree; Step S2-1-2, obtaining the score of each feature by analyzing the split gain of each feature in each decision tree; Step S2-1-3, discarding the samples with scores lower than the threshold, retaining the samples with scores higher than the threshold, and outputting the weights of the samples; In Step S2, the LSTM module includes multiple LSTM sub-modules A using forward propagation and multiple LSTM sub-modules A' using backward propagation. The sample data and the weight of the sample output by the XGBoost tree module are used as the inputs of the LSTM sub-modules A and LSTM sub-modules A'. After the outputs of the LSTM sub-modules A and LSTM sub-modules A' are jump-connected with the corresponding sample data, they are used as the output of the LSTM module; Both the LSTM sub-module A and the LSTM sub-module A' are added with a peephole structure and use a coupled forget mechanism; In Step S2, the neural ordinary differential equation of the nmODE module is: Among them, represents the input of the nmODE module, represents the output of the data x in the LSTM module, represents the state of the nmODE module at time t.
2. The method for predicting the risk value of postoperative pulmonary complications according to claim 1, characterized in that, In Step S1, the sample data after lung surgery includes age group, gender number, patient's surgery date, daily air pressure, air pressure higher than the mean, daily temperature, temperature higher than the mean, daily humidity, humidity higher than the mean, sunshine, sunshine higher than the average, AQI, PM2.5, PM10, SO2, CO, NO2, O3_8h, AQI classification, 2.5 classification, 10 classification, SO2 higher than the average, CO higher than the average, NO2 higher than the dangerous value, O3 higher than the dangerous value, monthandseason, season number, length of hospital stay, hypertension, cardiovascular disease, COPD emphysema, diabetes, tumor history, other comorbidities, smoking history, whether still smoking, lung function, MVV, diffusion function, BMI classification, open chest or minimally invasive, operation minutes, resection range number, number of indwelling roots.
3. The method for predicting the risk value of postoperative pulmonary complications according to claim 2, characterized in that, Using the SMOTE algorithm to perform data preprocessing on the sample data after lung surgery, and the SMOTE algorithm is specifically: Based on the K nearest neighbor samples of each sample, randomly select N neighboring points, and generate a threshold between 0 and 1 through interpolation to synthesize data; the input is the data set X, and the output is the sampled data set, with a size of (N / 100)*T; Among them, T represents the number of samples to be processed, N represents the sampling ratio, and K represents the number of nearest neighbor samples.
4. A system for predicting the risk value of postoperative pulmonary complications, characterized in that, It includes: A sample data acquisition module, which is used to acquire the sample data after lung surgery, calibrate whether complications occur in the sample data after lung surgery, and form label data; A risk value prediction model construction module, which is used to construct a risk value prediction model. The risk value prediction model includes an XGBoost tree module, an LSTM module, and an nmODE module; A risk value prediction model training module, which is used to train the risk value prediction model constructed by the risk value prediction model construction module with the sample data of the sample data acquisition module; A risk value real-time prediction module, which is used to acquire real-time postoperative data and input it into the risk value prediction model, and the risk value prediction model outputs the predicted risk value; In the risk value prediction model construction module, the XGBoost tree module outputs the weight corresponding to the sample according to the sample data, specifically: Step S2-1-1, an XGBoost tree module including n decision trees is constructed, which is used to fit the sample data and calculate the split gain of each feature in each decision tree; Step S2-1-2, by analyzing the split gain of each feature in each decision tree, the score of each feature is obtained; Step S2-1-3, discard the samples with scores lower than the threshold, retain the samples with scores higher than the threshold, and output the weights of the samples; In the risk value prediction model construction module, the LSTM module includes multiple LSTM sub-modules A using forward propagation and multiple LSTM sub-modules A' using backward propagation. The sample data and the weight of the sample output by the XGBoost tree module are used as the inputs of the LSTM sub-module A and the LSTM sub-module A'. The outputs of the LSTM sub-module A and the LSTM sub-module A' are jump-connected with the corresponding sample data and used as the output of the LSTM module; Both the LSTM sub-module A and the LSTM sub-module A' are added with a peephole structure and use a coupled forget mechanism; In the risk value prediction model construction module, the neural ordinary differential equation of the nmODE module is: Among them, represents the input of the nmODE module, represents the output of data x in the LSTM module, represents the state of the nmODE module at time t.
5. A computer device, characterized in that: It includes a memory and a processor. The memory stores a computer program. When the computer program is executed by the processor, the processor executes the steps of the method according to any one of claims 1 to 3.
6. A computer-readable storage medium, characterized in that: Stores a computer program. When the computer program is executed by the processor, the processor executes the steps of the method according to any one of claims 1 to 3.
Citation Information
Patent Citations
A system for constructing a prediction model for the risk of persistent air leakage after lung cancer resection
CN113936804B
Malignant tumor combined venous thromboembolism risk prediction method
CN113674864A
TAVR postoperative complication risk value prediction method based on aggregation neural network
CN114242234A