Mortality Prediction Method and System Optimized Based on LightGBM
By adopting a mortality prediction method based on LightGBM optimization in the intensive care unit, using the random forest and Pearson correlation algorithm for feature selection, and optimizing model parameters through the sparrow search algorithm, the problem of inability to efficiently and accurately predict patient mortality in the existing technology is solved, and the accuracy of prediction and resource utilization efficiency are improved.
Patent Information
- Application Number
- CN202111317863.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-11-09
- Publication Date
- 2025-06-13
- Estimated Expiration
- 2041-11-09
AI Technical Summary
The mortality rate of patients in the intensive care unit cannot be predicted efficiently and accurately in the prior art, resulting in the inability to reasonably plan and use of intensive care unit and equipment.
The mortality prediction method based on LightGBM is adopted to obtain patient monitoring data and use the preset LightGBM model to predict. The model uses the random forest algorithm and the Pearson correlation algorithm to jointly select feature, and optimizes model parameters through the sparrow search algorithm.
The accuracy of patient mortality prediction is improved, the accuracy of feature selection and the optimal combination of model parameters is ensured, and the problem of low prediction efficiency and accuracy in the prior art is solved.
Smart Images

Figure CN114093503B_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of medical care technology, and in particular, to a mortality prediction method and system optimized based on LightGBM. Background Art
[0002] The intensive care unit (ICU) is a medical monitoring unit in a hospital for centrally monitoring patients and specifically treating critically ill patients. Its goal is to improve the success rate of rescue for critically ill patients and reduce the number of deaths. The difference between the ICU and an ordinary ward is that there is a central monitoring station, which can directly observe all monitoring wards, providing the best guarantee for patients fundamentally. However, according to the data, the ICU resource allocation in each hospital is very scarce, and the cost is relatively high. It is simply difficult for ordinary families to bear it for a long time. Therefore, it is particularly important to monitor the intensive care data of ICU patients and make a more accurate prediction of the patient's mortality rate to achieve the purpose of reasonably using ICU resource allocation.
[0003] In the prior art, due to the gradually increasing richness and complexity of modern medical monitoring equipment, various data in intensive care often have problems such as a large amount of data, high complexity, and highly imbalanced data. Many existing deep learning methods are not very ideal in terms of efficiency and accuracy in predicting the mortality rate of patients, and cannot accurately predict the mortality rate of patients, so they cannot be reasonably planned and used for the use of intensive care wards and equipment. Summary of the Invention
[0004] The present application provides a mortality prediction method and system optimized based on LightGBM to solve the problem in the prior art that the mortality rate of patients cannot be accurately predicted with high efficiency, resulting in the inability to reasonably plan and arrange the use of intensive care wards and equipment.
[0005] The above object of the present application is achieved by the following technical solutions:
[0006] In a first aspect, an embodiment of the present application provides a mortality prediction method optimized based on LightGBM, including:
[0007] Obtain the monitoring data of the patient to be detected;
[0008] Input the monitoring data into a preset LightGBM model to obtain a mortality prediction result of the patient to be detected; wherein, the LightGBM model is a LightGBM mortality prediction model obtained by jointly performing feature selection on a preset data set through a preset random forest algorithm and a preset Pearson correlation algorithm, and optimizing the model parameters through a preset sparrow search algorithm.
[0009] Output the mortality prediction result.
[0010] Further, the feature selection jointly performed by the random forest algorithm and the Pearson correlation algorithm based on a preset data set includes:
[0011] Determine the data set, and perform data processing on the data set to obtain the features to be selected;
[0012] Calculate the importance value of each of the features to be selected through a preset random forest algorithm;
[0013] Calculate the correlation of each of the features to be selected through a preset Pearson correlation algorithm;
[0014] Based on the importance value and the correlation of the features to be selected, obtain the mortality impact value corresponding to each of the features to be selected;
[0015] Select the features to be selected based on the mortality impact value.
[0016] Further, the determining the data set and performing data processing on the data set to obtain the features to be selected includes:
[0017] Determine the data set;
[0018] Filter, clean, and standardize the data in the data set;
[0019] Based on the results of filtering, cleaning, and standardization, determine the features to be selected.
[0020] Further, the selecting the features to be selected based on the mortality impact value includes:
[0021] Sort the features to be selected based on the mortality impact value;
[0022] Select and discard the features to be selected based on the sorting result to complete the feature selection.
[0023] Further, the data set is an intensive care medicine information set.
[0024] Further, the optimizing the model parameters through a preset sparrow search algorithm includes:
[0025] Define the LightGBM algorithm as the fitness function, and use the parameter value range therein as the activity range of each sparrow in the preset sparrow search algorithm;
[0026] Calculate the fitness value of each sparrow and sort them to obtain the optimal fitness value;
[0027] Determine the position of the optimal fitness based on the optimal fitness value;
[0028] Determine the position of the optimal fitness as the optimal parameter combination of the LightGBM algorithm.
[0029] Furthermore, determining the position of the optimal fitness based on the optimal fitness value includes:
[0030] After obtaining the optimal fitness value and the position of the optimal fitness for the first time, check whether the sparrow search algorithm has reached the maximum number of iterations;
[0031] If not, update the positions of the discoverers, joiners, and vigilantes in the sparrow search algorithm, recalculate the fitness values of each sparrow and sort them, and obtain the position of the optimal fitness until the sparrow search algorithm reaches the maximum number of iterations;
[0032] If it has reached, determine the position of the optimal fitness based on the optimal fitness value after reaching the maximum number of iterations.
[0033] In a second aspect, an embodiment of the present application provides a mortality prediction system optimized based on LightGBM, including:
[0034] An acquisition module for the monitoring data of the patient to be detected;
[0035] A calculation model module for bringing the monitoring data into a preset LightGBM model to obtain the mortality prediction result of the patient to be detected;
[0036] The calculation model module further includes:
[0037] A feature selection sub-module for jointly performing feature selection based on a preset data set through a preset random forest algorithm and a preset Pearson correlation algorithm;
[0038] A parameter optimization sub-module for optimizing the model parameters through a preset sparrow search algorithm to obtain a LightGBM mortality prediction model;
[0039] An output module for outputting the mortality prediction result.
[0040] The technical solutions provided by the embodiments of the present application may include the following beneficial effects:
[0041] In the technical solution provided by the embodiments of the present application, first, the monitoring data of the patient to be detected is obtained; then, the monitoring data is brought into the preset LightGBM model to obtain the mortality prediction result of the patient to be detected; finally, the mortality prediction result is output. Among them, the LightGBM model is a LightGBM mortality prediction model obtained by jointly performing feature selection based on a preset dataset through a preset random forest algorithm and a preset Pearson correlation algorithm, and optimizing the model parameters through a preset sparrow search algorithm. In this way, by jointly performing feature selection through the random forest algorithm and the Pearson correlation algorithm, the accuracy of feature selection and rejection is ensured, and the optimal parameter combination of the LightGBM algorithm is determined through the sparrow search algorithm, thus not only retaining the advantages of the original algorithm such as fast speed and low consumption, but also making up for the deficiency that it cannot determine the optimal parameter combination, and improving the accuracy of patient mortality prediction in multiple aspects.
[0042] It should be understood that the above general description and the following detailed description are only exemplary and explanatory, and cannot limit the present application. Brief Description of the Drawings
[0043] The drawings here are incorporated into the specification and constitute a part of this specification, showing embodiments consistent with the present application, and are used together with the specification to explain the principles of the present application.
[0044] Figure 1 It is a schematic flowchart of a mortality prediction method optimized based on LightGBM provided by an embodiment of the present application;
[0045] Figure 2 It is a schematic diagram for the construction and use of a mortality prediction model optimized based on LightGBM provided by an embodiment of the present application;
[0046] Figure 3 It is a schematic diagram of the principle of feature selection in the mortality prediction method optimized based on LightGBM provided by an embodiment of the present application;
[0047] Figure 4 It is a schematic diagram of the principle of determining the optimal parameters of the model in the mortality prediction method optimized based on LightGBM provided by an embodiment of the present application;
[0048] Figure 5 It is a schematic diagram of the RandomForest feature importance ranking in the mortality prediction method optimized based on LightGBM provided by an embodiment of the present application;
[0049] Figure 6 It is a schematic diagram of the LightGBM feature importance ranking in the mortality prediction method optimized based on LightGBM provided by an embodiment of the present application;
[0050] Figure 7 It is a schematic diagram of the Pearson correlation between the death situation of ICU patients and each feature in the mortality prediction method optimized based on LightGBM provided by the embodiment of the present application;
[0051] Figure 8 It is a schematic diagram of the AUC value under different feature selection methods in the mortality prediction method optimized based on LightGBM provided by the embodiment of the present application;
[0052] Figure 9 It is a schematic diagram of the change trend of the optimal fitness value in the mortality prediction method optimized based on LightGBM provided by the embodiment of the present application;
[0053] Figure 10 It is a comparison chart of the experimental results of different algorithms in the verification process of the mortality prediction method optimized based on LightGBM provided by the embodiment of the present application;
[0054] Figure 11 It is a comparison chart of the ROC curves of different algorithms in the verification process of the mortality prediction method optimized based on LightGBM provided by the embodiment of the present application;
[0055] Figure 12 It is a schematic flow chart of a mortality prediction system optimized based on LightGBM provided by the embodiment of the present application. Detailed implementation manners
[0056] Here, the exemplary embodiments will be described in detail, and the examples are shown in the drawings. When the following description refers to the drawings, unless otherwise indicated, the same numbers in different drawings represent the same or similar elements. The implementation manners described in the following exemplary embodiments do not represent all implementation manners consistent with the present application. On the contrary, they are only examples of devices and methods consistent with some aspects of the present application as detailed in the appended claims.
[0057] In order to solve the above problems, the present application provides a mortality prediction method optimized based on LightGBM to efficiently and accurately predict the patient mortality rate, so as to help relevant medical staff plan the use of intensive care units and equipment, thereby improving the utilization efficiency of intensive care unit and equipment resources.
[0058] Embodiment
[0059] Refer to Figure 1 , Figure 1 It is a schematic flow chart of a mortality prediction method optimized based on LightGBM provided by the embodiment of the present application, as Figure 1As shown, the method at least includes the following steps:
[0060] S101. Obtain the monitoring data of the patient to be detected.
[0061] S102. Input the monitoring data into a preset LightGBM model to obtain the mortality prediction result of the patient to be detected.
[0062] Among them, the LightGBM model is a LightGBM mortality prediction model obtained by jointly performing feature selection on a preset data set through a preset random forest algorithm and a preset Pearson correlation algorithm, and optimizing the model parameters through a preset sparrow search algorithm.
[0063] Specifically, in the mortality prediction method optimized based on LightGBM provided in the embodiments of the present application, LightGBM (Light Gradient Boosting Machine) is used as the basis of the prediction model. It should be noted that LightGBM has the advantages of fast training speed, low memory consumption, high accuracy, etc. Among them, LightGBM includes two new technologies: Gradient-based One-side Sampling and Exclusive Feature Bundling (EFB). These two new technologies are optimized from the perspectives of reducing the number of samples and reducing the feature dimension respectively to handle the related problems of a large number of data instances and a large number of sample features. In dealing with the related problem of reducing the number of samples, the Gradient-based One-side Sampling (GOSS) method is adopted. Each time the weak classifier is updated, GOSS compresses the training data set without changing the feature value distribution and without losing accuracy, reducing the computational amount. In addition, the Gradient-based One-side Sampling increases the diversity of the weak classifier, thereby improving the generalization ability of the model. In dealing with the related problem of reducing the feature dimension, the method of EFB independent feature merging is adopted. EFB constructs a weighted undirected graph, models the constructed feature set into a graph coloring problem, uses a similar greedy algorithm to obtain the result, and uses the method of dividing histograms to bundle the mutually exclusive features in the feature set.
[0064] In practical applications, after obtaining the monitoring data of the patient to be tested, the monitoring data generally has a large number of features, and it is necessary to retain the features that are more important for predicting mortality, and discard the features that have less impact on predicting mortality. In the process of feature selection, two algorithms are used in this application to select features, namely the random forest algorithm and the Pearson correlation algorithm. Among them, the random forest algorithm Random Forest (RF) is a classification and regression algorithm based on the idea of model aggregation, which measures the importance of features by calculating the average impurity attenuation obtained by all decision trees in the random forest. In this application, on the one hand, the importance of each feature variable is scored by the random forest algorithm, thereby providing a basis for feature selection.
[0065] On the other hand, feature selection is also performed using the Pearson correlation algorithm. In statistics, the Pearson correlation coefficient (PCCs) is a metric used to calculate the correlation (linear correlation) between two variables X and Y, and its value is between -1 and 1. Through the Pearson correlation calculation, the correlation of features is obtained, providing a basis for feature extraction.
[0066] Based on the use of random forest to calculate the importance of features, Pearson correlation analysis was used to calculate the correlation between each laboratory test item and the mortality situation. Since the importance and correlation of the features are in the same dimension, the feature importance and correlation of each laboratory test item were added to obtain the impact value of the features of each laboratory test item on the mortality situation. The features of the test items were ranked using the impact value, and the features with the highest ranking were retained, while the features with the lowest ranking could be appropriately discarded, thereby reducing the impact of feature data with little or no effect on the accuracy of the prediction results, reducing the amount of calculation, and alleviating system pressure.
[0067] Further, after determining the basic training model and the corresponding features, in order to make the prediction results more accurate, in the mortality prediction method based on LightGBM optimization provided in the embodiments of the present application, it further includes optimizing the model parameters through the sparrow search algorithm. In practical applications, although the LightGBM algorithm has the above excellent characteristics, since it is difficult to determine whether the parameter values of the LightGBM algorithm are the optimal values of its model, in the mortality prediction method based on LightGBM optimization provided in the embodiments of the present application, the sparrow search algorithm (SSA) is used to search for the optimal parameter combination of the LightGBM algorithm. For example, in the parameter tuning of the LightGBM algorithm, several hyperparameters such as learning_rate, n_estimators, num_leaves, min_data_in_leaf, and max_depth are selected for optimization to determine their optimal values. Then, when inputting the test set samples, the accuracy of predicting the patient's mortality is ensured in multiple aspects.
[0068] S103. Output the mortality prediction result.
[0069] Finally, the obtained mortality prediction result is output, so that relevant staff can reasonably plan and arrange the use of the intensive care unit and equipment based on the patient mortality prediction result, improving the resource utilization rate.
[0070] The embodiments of the present application provide a mortality prediction method based on LightGBM optimization, including first obtaining the monitoring data of the patient to be detected; then bringing the monitoring data into the preset LightGBM model to obtain the mortality prediction result of the patient to be detected; and finally outputting the mortality prediction result. Among them, the LightGBM model is a LightGBM mortality prediction model obtained by jointly performing feature selection through a preset random forest algorithm and a preset Pearson correlation algorithm based on a preset data set, and optimizing the model parameters through a preset sparrow search algorithm. In this way, through the joint feature selection of the random forest algorithm and the Pearson correlation algorithm, the accuracy of feature selection and rejection is ensured, and the optimal parameter combination of the LightGBM algorithm is determined through the sparrow search algorithm, thus not only retaining the advantages of the original algorithm such as fast speed and low consumption, but also making up for its inability to determine the optimal parameter combination.
[0071] Figure 2 It is a schematic diagram for the construction and use of a mortality prediction model based on LightGBM optimization provided in the embodiments of the present application, as Figure 2 shown:
[0072] In the mortality prediction method optimized based on LightGBM provided by the embodiments of the present application, it mainly includes the construction process of the mortality prediction model optimized based on LightGBM and the usage process of this model.
[0073] The construction process includes selecting a data set to obtain training samples. For example, after a series of data preprocessing processes such as sorting and screening the data of ICU patients in the MIMIC-III data set, training samples are obtained. The features are sorted based on the calculation results of random forest importance and Pearson correlation, and then combined with the LightGBM algorithm and the Sparrow Search Algorithm (SSA) to obtain the RF-PCCs (i.e., random forest and Pearson correlation algorithm) feature selection and SSA (i.e., Sparrow Search Algorithm)-LightGBM mortality prediction model provided by the embodiments of the present application.
[0074] Specifically, data is obtained from the MIMICIII data set, and the data is preprocessed, including data screening, data cleaning, and standardization processing to obtain training set samples; then through Figure 2 the process of optimizing LightGBM parameters by the SSA algorithm marked by the dashed line in the figure, the optimal fitness value and the combination corresponding to the fitness value are determined by the Sparrow Search Algorithm, i.e., the SSA algorithm in the figure, to optimize the LightGBM parameters; and through Figure 2 the feature selection process marked by the dashed line in the figure, including jointly calculating the mortality impact value of the features based on the calculation results of random forest importance and Pearson correlation, then sorting the mortality impact values of the features, and selecting or rejecting the features based on the sorting results to obtain a complete prediction model. By inputting the test set samples, the mortality prediction results corresponding to the test set samples can be obtained and output.
[0075] Figure 3 It is a schematic diagram of the principle of feature selection in a mortality prediction method optimized based on LightGBM provided by the embodiments of the present application. As Figure 3 shown, the feature selection in the mortality prediction method optimized based on LightGBM provided by the embodiments of the present application mainly includes:
[0076] First, calculate the importance of the current feature based on the random forest algorithm. Random Forest (RF) is a classification and regression algorithm proposed by Breiman (2001) based on the idea of model aggregation. The feature importance is measured by calculating the average impurity decay obtained from all decision trees in the random forest. Since the Gini index method is relatively fast and simple to calculate without using logarithms, the Gini index is selected as the evaluation index of feature importance in this application. The importance score of the feature variable (variable importance measures) is denoted as VIM, and the Gini index is denoted as GI. According to its calculation formula, in the i-th decision tree, the Gini index of node n is:
[0077]
[0078] In the formula: K represents that there are K categories at the feature node n, and P nk represents the probability that a randomly selected sample belongs to category k at node n.
[0079] In the i-th decision tree, if the node where feature j appears belongs to the set Q, then the importance of feature j at the feature node n of this decision tree is:
[0080]
[0081] In the formula: ΔVIM represents the change in the Gini index before and after node n is split, and GI l represents the Gini index of the new node after node splitting.
[0082] If there are t trees in the random forest, then the importance of feature variable j in the random forest is:
[0083]
[0084] Then, calculate the correlation between features through the Pearson correlation algorithm. In statistics, the Pearson correlation coefficient (PCCs) is a measure used to calculate the correlation (linear correlation) between two variables X and Y, and its value ranges from -1 to 1. The Pearson correlation coefficient between two variables is defined as the quotient of the covariance and the standard deviation between the two variables. Denote the Pearson correlation as P. According to the Pearson correlation calculation formula, the correlation P j,X between feature variable j and the death situation X is:
[0085]
[0086] In the formula, P j,Xis the correlation coefficient between the feature variable j and the death situation X, and σ j , σ X are the standard deviations of the feature variable j and the death situation X respectively, μ j , μ X are the expected values of the feature variable j and the death situation X respectively, and cov(j, X) is the covariance between the feature variable j and the death situation X.
[0087] In practical applications, based on the above theoretical basis, the features can be selected by designing an RF-PCCs feature selection model. When predicting the mortality rate of ICU patients, in the method provided by the embodiments of the present application, each laboratory test item of patients with more than 10,000 examinations is selected. After screening, there are a total of 56 features. Since different feature combinations will affect the prediction results of the patient's mortality rate, the Gini index of each laboratory test item of ICU patients can be calculated, and the calculation result of its Gini index is used as the importance of this test item, so as to screen out factors with stronger correlation with the mortality rate. In addition, when calculating the feature importance, there may be a situation where different test items of the same patient have the same importance. Blindly selecting or discarding features will affect the accuracy and reliability of the mortality prediction model. Therefore, the present application proposes an RF-PCCs feature selection model. On the basis of using a random forest to calculate the importance of features, Pearson correlation analysis is used to calculate the correlation between each laboratory test item and the death situation. Since the importance and correlation of features are in the same dimension, the feature importance and correlation of each laboratory test item are added to obtain the influence value of each laboratory test item on the death situation.
[0088] Feature variable X j ∈(X 1 , X 2 , X 3 ,..., X M ), where M is the total number of features. The key definitions of the RF-PCCs feature selection model proposed in this paper are as follows:
[0089] Definition: (Mortality influence value Mor_value) The sum of the importance and correlation of a certain feature variable is called the mortality influence value of this feature. Then the mortality influence value of the feature variable X j can be expressed as
[0090]
[0091] In some specific implementation processes, the construction and use process of the RF-PCCs feature selection model includes:
[0092] Step 1: Input the training sample set, the maximum number of features M, initialize the current iteration number j = 1, and define an empty dictionary Mor_dic = {};
[0093] Step 2: Calculate the feature importance j and correlation P of the current feature variable X j,X ;
[0094] Step 3: Calculate the mortality impact value Mor_value j of the current feature variable X j ;
[0095] Step 4: Add the feature variable name of the current feature variable X j and the corresponding mortality impact value Mor_value j to the dictionary Mor_dic, Mor_dic = {X j : Mor_value j};
[0096] Step 5: If j < M, the current iteration number j = j + 1, and continue to execute the process between the above Step 2 and Step 4;
[0097] Step 6: First convert the dictionary Mor_dic into a tuple, then sort it in descending order according to the mortality impact values corresponding to each feature variable in the tuple, and output the sorting result. Feature selection based on the sorting result can ensure the rationality of feature selection to the greatest extent, thereby improving the accuracy of the model prediction result.
[0098] Figure 4 is a schematic diagram of the principle for determining the optimal parameters of the model in a mortality prediction method optimized based on LightGBM provided by an embodiment of the present application. As Figure 4 shown, in the embodiment of the present application, the sparrow search algorithm is used to determine the optimal parameters of the model, and a mathematical model of SSA-LightGBM is established based on the optimal parameters to ensure the accuracy of the model prediction result.
[0099] The Sparrow Search Algorithm (SSA) is a new optimization algorithm inspired by the foraging and anti-predation behaviors of sparrows. Sparrows can be divided into two types: producers and scroungers. Producers usually have good energy reserves and are responsible for finding food in the population and providing better foraging areas and directions for the entire sparrow population. Scroungers obtain food by using producers, and the fitness value corresponding to each sparrow individual is used as an indicator to measure the level of a sparrow's energy reserve. The identities of producers and scroungers change dynamically to find better food resources, but the proportion of producers and scroungers in the entire population remains fixed. During the foraging process, scroungers can always search for the producers that provide the best food resources and then forage from the best food or forage around those producers. At the same time, some scroungers may continuously monitor the producers to seize resources in order to increase their predation rate. There are also some scroungers that are in relatively poor foraging positions in the entire population. To obtain more food resources, these scroungers may fly to other areas to forage. During the predation process, once a sparrow discovers a predator, the individual starts to emit a chirping sound as an alarm signal. When the alarm value is greater than the safety value, the producers will take the scroungers to other safe areas to forage. When realizing the danger, the sparrows at the edge of the group will quickly move to the safe area to obtain a better position, and the sparrows in the middle of the group will then move around to get closer to other sparrows to reduce their probability of being captured.
[0100] In this application, assume that d represents the dimension of the variables of the problem to be optimized, and n represents the total number of sparrows. The sparrow population X composed of n sparrows is expressed in the following form:
[0101]
[0102] Define the LightGBM algorithm as the fitness function, and use the value ranges of the parameters of the LightGBM algorithm as the activity position ranges of each sparrow. Calculate the error rate of the test set as the fitness value f of each sparrow individual. The smaller the fitness value, the higher the energy reserve of the sparrow. The position corresponding to the optimal fitness value is the optimal parameter combination of the LightGBM algorithm.
[0103] Assume that t represents the current iteration number, and iter max represents the maximum number of iterations. represents the position information of the i-th sparrow in the j-th dimension, R 2 (R 2If \(S\in[0,1)\) and \(ST\in[0.5,1]\) represent the safety value and the warning value respectively, then the position update formula for the discoverer is as follows:
[0104]
[0105] where \(\alpha\in(0,1]\) is a random number. \(Q\) represents a random number that follows a normal distribution. \(L\) represents a \(1\times d\) matrix, where each element in the matrix is all 1.
[0106] Suppose \(X\) P represents the optimal position occupied by the current discoverer, and \(X\) worst represents the worst position in the current global. Then the position update formula for the joiner is as follows:
[0107]
[0108] where \(A\) represents a \(1\times d\) matrix, where each element in the matrix can be randomly assigned a value of 1 or -1, and \(A\) + = \(A\) T (AA T ) -1 .
[0109] When , it means that the \(i\)-th joiner with low energy reserve at this time needs to obtain more energy, so it flies to other places to forage.
[0110] Suppose \(X\) best is the current global optimal position, and \(f\) i is the fitness value of the current sparrow individual. \(f\) g and \(f\) w are the current global best and worst fitness values respectively. Then the position update formula for the vigilant is as follows:
[0111]
[0112] where \(\beta\) is the step size control parameter and follows a standard normal distribution random number. \(K\in[-1,1]\) is a random number, and \(\varepsilon\) is a constant to avoid the denominator being zero.
[0113] When \(f\) i > \(f\) g , it means that the peripheral sparrows are at the edge of the population and have discovered a predator, and are extremely vulnerable to predator attacks.
[0114] When \(f\) i = \(f\) g , it means that the sparrows in the middle of the population are aware of the danger and need to move randomly in the safe area to get closer to other sparrows to reduce their risk of being preyed upon. \(K\) represents the direction of the sparrow's movement and is also the step size control parameter.
[0115] In some specific implementation processes, the mathematical model of SSA-LightGBM may include the following steps:
[0116] The first step: Set the population size pop, the maximum number of iterations MaxIter, set the warning value ST, the ratio PD of predators and joiners, the ratio SD of sparrows aware of danger, and the value ranges of the parameters of the LightGBM algorithm as shown in Tables 1 and 2;
[0117] The second step: Calculate the fitness value of each sparrow and sort them to obtain the optimal fitness value;
[0118] The third step: Obtain the corresponding optimal fitness value position combination according to the optimal fitness value;
[0119] The fourth step: Determine whether the maximum number of iterations is reached. If it is reached, terminate the operation. Otherwise, first update the positions of the discoverers, joiners, and vigilants according to Formulas 8 and 9; then repeat the process of the second to fourth steps;
[0120] The fifth step: Output the optimal fitness value and the position of the optimal fitness. The position of the optimal fitness value is the optimal parameter combination of the LightGBM algorithm.
[0121] Among them, the settings of the parameters of the above SSA-LightGBM mortality prediction model are shown in Table 1, and the value ranges of the parameters of the LightGBM algorithm are shown in Table 2:
[0122] Table 1 Settings of the parameters of the SSA-LightGBM mortality prediction model
[0123]
[0124] Table 2 Preset parameter ranges
[0125]
[0126] Next, a complete embodiment will be used to elaborate in detail on the mortality prediction method optimized based on LightGBM provided by the embodiments of the present application:
[0127] First, the experimental platform for the model use and prediction result verification in this embodiment is an Intel(R) Core(TM) i7-5500U CPU@2.40GHz, the operating system is Windows 10, the program development environment is PyCharm Community Edition 2020.2.3, and the programming language is python3.6.5.
[0128] Determination and processing of the dataset: In the study on predicting the mortality of ICU patients, in order to verify the effectiveness of the SSA-LightGBM mortality prediction model based on RF-PCCs feature selection provided in the embodiments of the present application, the dataset used is the publicly available dataset - Medical Information Mart for Intensive Care III (MIMIC-III), which is jointly supported by Beth Israel Deaconess Medical Center, the Computational Physiology Laboratory of the Massachusetts Institute of Technology, and Philips. This dataset includes demographic information, vital signs, laboratory tests, medications, and other health data related to approximately 60,000 visits to the intensive care unit and unknown identities. The MIMIC-III dataset is a relational database consisting of 26 tables. In this experiment, five tables, namely ADMISSIONS, PATIENTS, ICUSTAYS, D_LABITEMS, and LABEVENTS, were used. The descriptions of these five tables are shown in Table 3 below.
[0129] Table 3 Description of the experimental data table
[0130]
[0131] This application studies the mortality of ICU patients. A total of 61,532 hospitalization records of ICU patients were queried from the ICUSTAY table. Since a patient may enter the ICU ward multiple times, there will be multiple ICU hospitalization records. Therefore, the ICU patients' hospitalization records were screened, and only the hospitalization records of each patient's first entry into the ICU were retained. After screening, there were 46,476 hospitalization records of ICU patients. The age of each patient was calculated based on the patient's death date and birth date. It was found that there were 7,870 neonatal death patients and 1,990 death patients over 91 years old. Since there are too many missing items in laboratory tests for neonatal patients and patients with too old ages, the research objects selected in the embodiments of the present application do not include these two parts of hospitalization records. After the final screening, there were a total of 36,616 ICU hospitalization records of patients. The experimental data statistical table is shown in Table 4.
[0132] Table 4 Experimental data statistical table
[0133]
[0134] After determining the dataset, it is necessary to preprocess and standardize the data. It should be noted that in the experiment of the embodiments of the present application, five tables in the Medical Information Mart for Intensive Care III (MIMIC-III) dataset were selected, and these five tables are interconnected through SUBJECT_ID.
[0135] First, tags need to be added to the patient's death situation to complete supervised learning. The mortality prediction in the embodiments of this application is the survival status of the patient 24 hours after leaving the ICU. The method of adding tags is as follows: First, convert the death time dod of the ICU patient and the time dischtime when the patient leaves the ICU into a form measured in hours. If dod ≤ (dischtime + 24), it is marked as dead; otherwise, it is marked as alive. Among the 36,616 screened ICU patients, 4,086 patients are no longer alive, and 32,530 patients are still alive. The screened data is divided into a training set and a test set, with 70% divided into the training set and 30% divided into the test set.
[0136] Immediately afterwards, feature selection is also required. It can be seen from the D_LABITEMS table that the laboratory test items are the above-mentioned features, with a total of 753 items. According to the item ID (itemid), the measurement values corresponding to the laboratory test items done by ICU patients can be viewed in the LABEVENTS table. Since only some patients have undergone some laboratory test items, in this experiment, only the test items with the number of examined patients greater than 10,000 are selected. In addition, for the same patient, multiple different values may be produced in the same laboratory test item because the same patient has undergone the same item of examination at different time periods. Therefore, we take the mean of the multiple values produced by the same patient in the same examination item. Finally, missing values and outliers in the data set are processed. After processing, the data set has a total of 56 laboratory test items. The distribution of the number of people for each test item is shown in Table 5.
[0137] Table 5 Statistical table of the number of people for laboratory test items
[0138]
[0139] Specifically, in the study of predicting the mortality rate of ICU patients, since different features are selected, it will affect the results of predicting the patient's mortality rate. Therefore, how to select features is particularly important. First, calculate the importance values of each feature affecting ICU patients according to Random Forest and LightGBM respectively. The results of the feature importance values are as Figure 5 , Figure 6 shown. The horizontal axis is the feature name, and the vertical axis is the value calculated for feature importance.
[0140] From Figure 5It can be seen that in the calculation of feature importance using random forest, the feature importance of Lactate is the highest, which is 0.117935, and the feature importance of Nitrite is the lowest, only 0.000692. This indicates that the Lactate feature has a greater impact on the prediction result of the death situation of ICU patients. In contrast, the Nitrite feature has a very small impact on the prediction result of the death situation of ICU patients.
[0141] From Figure 6 It can be seen that in the calculation of feature importance using LightGBM, the feature importance of Sodium is the highest, which is 211, and the feature importance of Nitrite is the lowest, only 8. This indicates that the Sodium feature has a greater impact on the prediction result of the death situation of ICU patients. In contrast, the Nitrite feature has a very small impact on the prediction result of the death situation of ICU patients. The feature importance values of Anion Gap and RDW are the same, which is 151; the feature importance values of pH, Urea Nitrogen, Potassium, and PTT are the same, which is 143; the feature importance values of Magnesium and MCV are the same, which is 143; the feature importance values of Base Excess and PT are the same, which is 104; the feature importance values of Free Calcium and MCH are the same, which is 94; the feature importance values of Red BloodCells and Lactate Dehydrogenase are the same, which is 86.
[0142] Since reasonable selection and rejection of feature values will have an important impact on the accuracy of the ICU patient death prediction model proposed in this paper, therefore, in order to prevent blindly selecting and rejecting features when multiple features have the same importance result, in the solution of this embodiment, on the basis of calculating the feature importance results using random forest and LightGBM, the Pearson correlation between the death situation of ICU patients and each feature is calculated, and the results are as Figure 7 shown.
[0143] It can be seen from the Pearson correlation graph that Anion Gap, Lactate, and Urea Nitrogen have a relatively high positive correlation with the death situation of ICU patients, and Bicarbonate and Base Excess have a relatively high negative correlation with the death situation of ICU patients.
[0144] Moreover, since the feature importance and correlation are of the same dimension, in the embodiments of the present application, the feature importance and correlation of each laboratory test item are added together to obtain the final influence value of each laboratory test item on the death situation and sort them. Here, the top ten final influence value rankings of the RF-PCCs and LightGBM-PCCs feature selection models are listed separately in Tables 6 and 7 as follows.
[0145] Table 6 Top Ten Final Influence Value Rankings of the RF-PCCs Feature Selection Model
[0146]
[0147] Table 7 Top Ten Final Influence Value Rankings of the LightGBM-PCCs Feature Selection Model
[0148]
[0149] Next, the experimental results and prediction results of the above embodiments will be analyzed;
[0150] The confusion matrix is an evaluation index to measure the classification performance of a classifier, as shown in Table 8. In the embodiments of the present application, the accuracy, recall, precision, and AUC-ROC are used to evaluate the proposed SSA-LightGBM mortality prediction model based on RF-PCCs feature selection.
[0151] Table 8 Confusion Matrix
[0152]
[0153] Precision:
[0154] Recall:
[0155] Accuracy:
[0156] F1:
[0157] To ensure the selection of reasonable features when predicting the mortality of ICU patients, in the embodiments of the present application, the AUC values under different feature numbers F are compared according to the sorting results of the four feature selection methods (RF, LightGBM, RF-PCCs, LightGBM-PCCs) from high to low, and the results are as Figure 8 shown.
[0158] FromFigure 8 As shown in the figure, among the four feature selection methods of RF, LightGBM, RF-PCCs, and LightGBM-PCCs, when the number of features is 32 for the RF feature selection method, the highest AUC value is 0.7477; when the number of features is 47 for the LightGBM feature selection method, the highest AUC value is 0.7470; when the number of features is 48 for the RF-PCCs feature selection method, the highest AUC value is 0.7508; when the number of features is 24 for the LightGBM-PCCs feature selection method, the highest AUC value is 0.7437. In summary, in the embodiments of the present application, RF-PCCs is selected for feature selection, and the number F of features is 48.
[0159] In this experiment, four classifier models of SVM, XGBoost, RF, and LightGBM were selected for comparison with the model proposed in the embodiments of the present application. Among them, the change trend of the optimal fitness value of the SSA-LightGBM mortality prediction model proposed in the embodiments of the present application is as Figure 9 shown, and the obtained optimal fitness value is 0.06982.
[0160] When the optimal fitness value is 0.06982, the optimal parameter combination obtained by the SSA-optimized LightGBM algorithm is shown in Table 9:
[0161] Table 9 Optimal Parameter Combination
[0162]
[0163] Other parameter settings for the SSA algorithm to optimize the LightGBM model: the seed number seed = 33, during the iteration process, the ratio of the training data to the total number is bagging_fraction = 0.8, and the number of bagging is set to 6. The experimental results and their ROC curves are shown in Table 10, Figure 10 , and Figure 11 shown, where Figure 10 in Figure 10 shown, for Accuracy, Auc, and F1, in each group, from left to right are SVM, RF, XGBoost, LightGBM, and The Proposed:
[0164] Table 10 Performance Comparison Table of Different Algorithms
[0165]
[0166] Through Table 10 and Figure 11It can be seen that the prediction result of the SSA-LightGBM mortality prediction model proposed in the embodiment of the present application based on RF-PCCs feature selection is the best. The precision value of this mortality prediction model is 0.930, the F1 value is 0.627, and the AUC-ROC value is 0.751. In terms of precision, it is 0.045 higher than SVM, 0.048 higher than RF, 0.031 higher than XGBoost, and 0.035 higher than LightGBM; in terms of F1 value, it is 0.107 higher than SVM, 0.173 higher than RF, 0.131 higher than XGBoost, and 0.053 higher than LightGBM; in terms of AUC-ROC, it is 0.073 higher than SVM, 0.091 higher than RF, 0.124 higher than XGBoost, and 0.206 higher than LightGBM. The prediction effect of the SSA-LightGBM mortality prediction model proposed in the present application based on RF-PCCs feature selection is better than that of SVM, RF, XGBoost, and LightGBM.
[0167] The prediction model proposed in the present application not only takes into account the selection of features, but also uses the SSA algorithm to make up for the shortcoming that it is difficult to determine the optimal parameter combination of the LightGBM algorithm, improves the efficiency and accuracy of mortality prediction, and provides a new idea for ICU mortality prediction.
[0168] In addition, based on the same inventive concept, the embodiment of the present application also provides a mortality prediction system optimized based on LightGBM. Figure 12 As shown in Figure 12 is a schematic flowchart of a mortality prediction system optimized based on LightGBM provided by an embodiment of the present application. As
[0169] An acquisition module 201, configured to obtain the monitoring data of the patient to be detected;
[0170] A calculation model module 202, configured to bring the monitoring data into a preset LightGBM model to obtain the mortality prediction result of the patient to be detected;
[0171] The calculation model module 202 further includes:
[0172] A feature selection sub-module 2021, configured to jointly perform feature selection based on a preset data set through a preset random forest algorithm and a preset Pearson correlation algorithm;
[0173] The parameter optimization sub-module 2022 is used to optimize the model parameters through a preset sparrow search algorithm to obtain a LightGBM mortality prediction model;
[0174] The output module 203 is used to output the mortality prediction result.
[0175] It can be understood that the same or similar parts in the above embodiments can be referred to each other, and the content not detailed in some embodiments can be referred to the same or similar content in other embodiments.
[0176] It should be noted that in the description of the present application, the terms "first", "second", etc. are only used for descriptive purposes and cannot be understood as indicating or implying relative importance. In addition, in the description of the present application, unless otherwise specified, the meaning of "a plurality" means at least two.
[0177] Any process or method description shown in the flowchart or described in other ways herein can be understood to represent a module, segment, or part of code including one or more executable instructions for implementing a specific logical function or process. The scope of the preferred embodiments of the present application includes additional implementations, where the functions can be executed in a substantially simultaneous manner or in the reverse order according to the involved functions, rather than in the order shown or discussed, which should be understood by those skilled in the technical field of the embodiments of the present application.
[0178] It should be understood that each part of the present application can be implemented by hardware, software, firmware, or a combination thereof. In the above embodiments, multiple steps or methods can be implemented by software or firmware stored in a memory and executed by a suitable instruction execution system. For example, if implemented by hardware, as in another embodiment, any one or a combination of the following well-known technologies in the art can be used: discrete logic circuits with logic gate circuits for implementing logical functions on data signals, application-specific integrated circuits with appropriate combinational logic gate circuits, programmable gate arrays (PGAs), field-programmable gate arrays (FPGAs), etc.
[0179] Those of ordinary skill in the technical field of the present application can understand that all or part of the steps carried by the methods of the above embodiments can be completed by instructing relevant hardware through a program. The program can be stored in a computer-readable storage medium. When the program is executed, it includes one or a combination of the steps of the method embodiments.
[0180] In addition, each functional unit in various embodiments of the present application may be integrated into a processing module, or each unit may exist physically alone, or two or more units may be integrated into one module. The above-mentioned integrated module may be implemented in the form of hardware or in the form of a software functional module. When the integrated module is implemented in the form of a software functional module and sold or used as an independent product, it may also be stored in a computer-readable storage medium.
[0181] The above-mentioned storage medium may be a read-only memory, a magnetic disk, an optical disc, etc.
[0182] In the description of this specification, the description with reference to the terms "one embodiment", "some embodiments", "example", "specific example", or "some examples", etc. means that the specific features, structures, materials, or characteristics described in connection with the embodiment or example are included in at least one embodiment or example of the present application. In this specification, the schematic representation of the above terms does not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials, or characteristics described may be combined in any one or more embodiments or examples in a suitable manner.
[0183] Although the embodiments of the present application have been shown and described above, it can be understood that the above embodiments are exemplary and should not be construed as limiting the present application. Those of ordinary skill in the art can make changes, modifications, substitutions, and variations to the above embodiments within the scope of the present application.
Claims
1. A mortality prediction method optimized based on LightGBM, characterized in that, it includes: Obtain the monitoring data of the patient to be detected; Input the monitoring data into a preset LightGBM model to obtain the mortality prediction result of the patient to be detected; wherein, the LightGBM model is a LightGBM mortality prediction model obtained by jointly performing feature selection on a preset data set through a preset random forest algorithm and a preset Pearson correlation algorithm, and optimizing the model parameters through a preset sparrow search algorithm; Output the mortality prediction result; Among them, the joint feature selection through the random forest algorithm and the Pearson correlation algorithm based on the preset data set includes: Determine the data set, and perform data processing on the data set to obtain features to be selected; Through a preset random forest algorithm, calculate the importance value of each feature to be selected; Through a preset Pearson correlation algorithm, calculate the correlation of each feature to be selected; Based on the importance value and correlation of the features to be selected, obtain the mortality impact value corresponding to each feature to be selected, including: adding the importance value and correlation of the features to be selected for each laboratory test item to obtain the mortality impact value corresponding to the features to be selected for each laboratory test item; Select the features to be selected based on the mortality impact value, including: sorting the features to be selected based on the mortality impact value; making choices for the features to be selected based on the sorting result to complete feature selection.
2. The mortality prediction method optimized based on LightGBM according to claim 1, characterized in that, The determination of the data set and the data processing of the data set to obtain features to be selected includes: Determine the data set; Screen, clean and standardize the data in the data set; Based on the results of screening, cleaning and standardization, determine the features to be selected.
3. The mortality prediction method optimized based on LightGBM according to claim 1, characterized in that, The data set is an intensive care medicine information set.
4. The mortality prediction method optimized based on LightGBM according to claim 1, characterized in that, The optimization of the model parameters through the preset sparrow search algorithm includes: Define the LightGBM algorithm as the fitness function, and use the parameter value range therein as the activity range of each sparrow in the preset sparrow search algorithm; Calculate the fitness value of each sparrow and sort to obtain the optimal fitness value; Determine the position of the optimal fitness based on the optimal fitness value; Determine the optimal parameter combination of the LightGBM algorithm as the position of the optimal fitness.
5. The mortality prediction method optimized based on LightGBM according to claim 4, characterized in that, The determination of the position of the optimal fitness based on the optimal fitness value includes: After obtaining the optimal fitness value and the position of the optimal fitness for the first time, check whether the sparrow search algorithm has reached the maximum number of iterations; If not, update the positions of the discoverer, joiner, and scout in the sparrow search algorithm, recalculate the fitness values of each sparrow and sort them, and obtain the position with the optimal fitness value until the sparrow search algorithm reaches the maximum number of iterations; If it has reached, determine the position with the optimal fitness value based on the optimal fitness value after reaching the maximum number of iterations.
6. A mortality prediction system optimized based on LightGBM Characterized in that It includes: An acquisition module for the monitoring data of the patient to be detected; A calculation model module for bringing the monitoring data into a preset LightGBM model to obtain the mortality prediction result of the patient to be detected; The calculation model module further includes: A feature selection sub-module for jointly performing feature selection based on a preset dataset through a preset random forest algorithm and a preset Pearson correlation algorithm, specifically including: Determine the dataset and perform data processing on the dataset to obtain the features to be selected; Through the preset random forest algorithm, calculate the importance value of each feature to be selected respectively; Through the preset Pearson correlation algorithm, calculate the correlation of each feature to be selected respectively; Based on the importance value and correlation of the features to be selected, obtain the mortality impact value corresponding to each feature to be selected, including: adding the importance value and correlation of the features to be selected for each laboratory test item to obtain the mortality impact value corresponding to the features to be selected for each laboratory test item; Select the features to be selected based on the mortality impact value, including: sorting the features to be selected based on the mortality impact value; making decisions on the features to be selected based on the sorting result to complete feature selection; A parameter optimization sub-module for optimizing the model parameters through a preset sparrow search algorithm to obtain a LightGBM mortality prediction model; An output module for outputting the mortality prediction result.
Citation Information
Patent Citations
ICU death rate prediction method and system based on punishment integrated model
CN111370126A
Fan main bearing fault monitoring and diagnosing method based on XGBoost algorithm model
CN111380686A