Method for predicting hospitalization duration of cholelithiasis patient based on ensemble learning
Through integrated learning methods, an accurate prediction model for hospitalization time for patients with cholelithiasis is constructed, which solves the problem of differences caused by the division of time intervals in traditional prediction methods, and realizes accurate prediction of hospitalization time and resource optimization, improving medical efficiency and patient satisfaction.
Patent Information
- Application Number
- CN202510359036.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-25
- Publication Date
- 2025-07-11
AI Technical Summary
Traditional hospitalization time prediction methods often use interval division, which leads to a large difference between the actual hospitalization time and the predicted time of cholelithiasis patients, affecting the accuracy of bed turnover and treatment plan.
Using an integrated learning-based approach, data preprocessing, key feature screening and model training are carried out by constructing a benchmark data set, using Bayesian optimization strategy and stacking generalization strategy, combined with XGBoost and LightGBM base models and random forest metamodels, accurately predicting the length of hospitalization for patients with cholelithiasis.
Accurate numerical prediction of the length of hospitalization for patients with cholelithiasis is achieved, helping the hospital optimize resource allocation, improve medical resource utilization and patient satisfaction, assisting doctors to evaluate rehabilitation status, and optimize treatment plans.
Smart Images

Figure CN120299651A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to a method for predicting the length of hospital stay, and more specifically, to a method for predicting the length of hospital stay of cholelithiasis patients based on ensemble learning. Background Art
[0002] The length of hospital stay (LOS) is a key parameter index for hospital ward bed allocation and is of great significance for measuring the operation efficiency, medical level, and overall work quality of the hospital. For a long time, the problems of long patient hospital stay and low ward bed turnover rate have continuously troubled hospital managers and become urgent problems to be solved. Accurately predicting the length of a patient's hospital stay can help the hospital plan the patient's treatment and care plan in advance, reduce the waiting time caused by insufficient beds, optimize resource allocation while improving patient satisfaction. At the same time, as an indirect indicator of the severity of the disease and the treatment effect, the length of hospital stay also helps the hospital to comprehensively manage and control the patient's disease.
[0003] Traditional methods for predicting the length of hospital stay often tend to divide the length of hospital stay into several intervals, such as 1 - 5 days, 5 - 10 days, and more than 10 days, etc. This way of classification prediction simplifies the problem to a certain extent and inevitably introduces errors. Taking the cholelithiasis disease as an example, for inpatients, their treatment methods generally include surgical treatment, through laparoscopic surgery or cholecystolithotomy, and supplemented with drug treatment to promote bile secretion and control inflammation. The length of a patient's hospital stay varies due to various factors such as the surgical method and the patient's individual conditions. If the traditional simple length division is used to allocate patients, it is very likely that there will be a large difference between the actual length of the patient's hospital stay and the predicted length of hospital stay, affecting the bed turnover and the planning of treatment and care plans.
[0004] By predicting the length of hospital stay, the hospital can more accurately grasp the patient's treatment process, thereby reasonably arranging medical resources; at the same time, it also helps doctors more accurately judge the patient's recovery situation, thereby formulating a more reasonable recovery plan for the patient, reducing the patient's anxiety, and improving patient satisfaction. The present invention adopts a regression prediction method, and based on the multi-dimensional information of the patient, uses an advanced algorithm based on ensemble learning to accurately numerically predict the number of hospital days of cholelithiasis patients, which has significant advantages. Summary of the Invention
[0005] In order to overcome the above technical problems, the present invention provides a method for predicting the length of hospital stay of cholelithiasis patients based on ensemble learning.
[0006] The method for predicting the length of hospital stay of cholelithiasis patients based on ensemble learning according to the present invention is characterized by being realized through the following steps:
[0007] a). Construct a benchmark dataset; collect multi-dimensional diagnosis and treatment data of cholelithiasis patients from the hospital HIS and the electronic pathology system, and construct a benchmark dataset D for predicting the length of hospital stay.
[0008] b). Preprocess the dataset; preprocess the data in the benchmark dataset D so that the processed data meets the training requirements.
[0009] c). Screen key features; perform feature engineering on the preprocessed dataset, and then screen out the key features suitable for modeling from the dataset D.
[0010] d). Model training; input the dataset D after key feature screening into the base model for iterative training. During the training process, use the Bayesian optimization strategy to optimize the model, save the optimal base model, and use the prediction results of the base learner as new features to form a feature matrix.
[0011] e). Model integration and prediction; use the prediction results of the base learner to train the meta-learner, and use the trained meta-learner to predict the samples to be predicted to obtain the predicted value of the length of hospital stay of cholelithiasis patients.
[0012] For the method for predicting the length of hospital stay of cholelithiasis patients based on ensemble learning of the present invention, the construction of the benchmark dataset described in step a) is specifically implemented through the following steps:
[0013] a-1). Collect demographic information; collect general demographic information of each cholelithiasis patient, including gender, age, ethnicity, height, weight, ward area, cost category, medical insurance status, occupation, admission route, admission diagnosis, admission date, and discharge date, from the hospital HIS and the electronic pathology system.
[0014] a-2). Collect surgical information; collect surgical information of cholelithiasis patients, including surgical name, surgical category, surgical level, major surgery flag, anesthesia method, and wound type.
[0015] a-3). Collect test information; collect test information of cholelithiasis patients, including white blood cells, hemoglobin, potassium, sodium, N-terminal pro-brain natriuretic peptide, creatinine, albumin, red blood cells, hematocrit, prothrombin time, D-dimer, glutamic oxaloacetic transaminase, glucose, interleukin, procalcitonin, C-reactive protein, and multi-drug resistance marker. The above test information is based on the first test result in the first medical encounter, that is, for patients with multiple test results during hospitalization, the first one is taken to reduce the bias caused by treatment effects.
[0016] a-4). Collect medication information; collect medication information of cholelithiasis patients, including aspirin, clopidogrel bisulfate tablets, and compound reserpine and triamterene tablets.
[0017] Taking the multi-dimensional diagnosis and treatment data composed of the general demographic information, surgical information, test information, and medication information of each gallstone patient collected as one row and storing it in the set D to form the original benchmark dataset D for predicting the length of hospital stay, D = {(X1, X2,..., X n ) T}, where n is the number of gallstone patients collected, and X i is the multi-dimensional diagnosis and treatment data of the i-th gallstone patient, and X i = {x i1 , x i2 ,..., x im}, where m is the number of multi-dimensional diagnosis and treatment data of the patient, and x ij is the j-th multi-dimensional diagnosis and treatment data of the i-th gallstone patient.
[0018] For the method for predicting the length of hospital stay of gallstone patients based on ensemble learning of the present invention, the preprocessing of the dataset described in step b) is specifically implemented through the following steps:
[0019] b-1). Elimination of outliers; Based on statistical methods, samples with a length of hospital stay significantly exceeding the normal range are eliminated to reduce the impact of outliers on model training;
[0020] b-2). Filling of missing values; For the missing test information values in the set D, the column mean filling method is used for processing, as shown in formula (1):
[0021]
[0022] In the formula, x ij-fill is the filled value of the missing test information j of the i-th patient, n' is the number of non-missing values in the j-th column of the set D, and x ik is the corresponding non-missing value in the j-th column of the set D;
[0023] b-3). Standardization of numerical features; The numerical features in the set D are standardized according to formula (2) so that the processed numerical values are between [0, 1] to eliminate the influence of different dimensions on subsequent modeling:
[0024]
[0025] In the formula, x i ′ j is the standardized processing result of the numerical feature x ij in the set D, x imax is the maximum value of the numerical feature in the j-th column of the set D, and x imin is the minimum value of the numerical feature in the j-th column of the set D.
[0026] The method for predicting the length of hospital stay of patients with cholelithiasis based on ensemble learning, the key feature screening described in step c) is specifically implemented through the following steps:
[0027] c-1). Feature augmentation; Generate two new features from the original benchmark dataset D. First, extract the month of patient admission based on the admission time of the patient to explore the potential relationship between the admission month and the length of the patient's hospital stay; Second, calculate the BMI index of each patient based on the patient's height and weight data to explore the potential relationship between the BMI index and the length of the patient's hospital stay. The BMI index is shown in formula (3):
[0028]
[0029] In the formula, w represents the patient's weight in kg, and h represents the patient's height in m;
[0030] c-2). Non-numerical feature encoding; Adopt the label encoding method to digitally map the non-numerical features in the dataset D to obtain structured data;
[0031] c-3). Feature screening; Based on the model-based feature ranking method, use random forest to screen features, and screen the top N features with the highest contribution for modeling according to the shape value. The calculation of the shape value is shown in formula (4):
[0032] y′ i =y base +f(X i ,1)+f(X i ,2)+...+f(X i ,l) (4)
[0033] In the formula, X i represents the i-th sample of the dataset D, y′ i represents the predicted value of the i-th sample, y base represents the mean of the predicted values of all samples, f(X i ,j) represents the contribution value of the j-th feature in the i-th sample to the final predicted value y′ i ;
[0034] Finally, select the top 20 features with the highest contribution for subsequent modeling. The top 20 features with the highest contribution selected are: operation name, albumin, occupation code, operation level, wound type, blood potassium, D-dimer, N-terminal pro-brain natriuretic peptide, ward area, red blood cells, interleukin, procalcitonin, prothrombin time, BMI, creatinine, anesthesia method, age, patient weight, blood sodium, glucose.
[0035] The method for predicting the length of hospital stay of patients with cholelithiasis based on ensemble learning, the model training described in step d) is specifically implemented through the following steps:
[0036] d-1). Dataset division; divide the samples in dataset D into a training set B and a test set C according to a ratio of 7:3;
[0037] d-2). Select base models; first divide the training set B evenly into 5 subsets, i.e., B = [subset B1, subset B2, subset B3, subset B4, subset B5], select the base models XGBoost and LightGBM of the efficient gradient boosting algorithm widely used in the field of machine learning. For each base model, optimize the hyperparameters of the model using 5-fold cross-validation k-flod and Bayesian optimization;
[0038] d-3). Set the objective function; set the objective function as the mean absolute error MAE (mean absolute error, MAE) to evaluate the model performance, as shown in formula (5):
[0039]
[0040] In the formula, n represents the number of samples, y i represents the true value of the i-th sample, y i ′ represents the predicted value of the i-th sample. Before the start of model training, the initial value of MAE is set to infinity;
[0041] d-4). Define the search space; define the hyperparameter search space of the model, that is, determine the value range or candidate value set of each hyperparameter of the model;
[0042] d-5). Iterative training process and cross-validation; in each iterative training and cross-validation, select four subsets as the training subsets, and the remaining one subset as the validation subset. This operation is repeated five times to ensure that each subset has taken turns to serve as the training set and the validation set;
[0043] For each cross-validation training process, record the current MAE value, and evaluate the model performance by comparing the MAE values under different parameter combinations; after every 5 cross-validation trainings are completed, calculate the average value of MAE in these 5 training processes as the overall evaluation score of this iterative cycle to measure the generalization ability of the model under the current parameter combination;
[0044] d - 6). Model optimization and saving: Repeat the training steps in d - 5) until the model converges or reaches the preset number of iterations. If the MAE score of the current iteration cycle is lower than that of the previous iteration cycle, determine that the model in the current iteration cycle is the optimal model so far and save it to ensure the continuous optimization and performance improvement of the model during training.
[0045] For the method for predicting the length of hospital stay of cholelithiasis patients based on ensemble learning of the present invention, the specific method of model ensemble and prediction in step e) is as follows: Based on the stacking ensemble strategy, further improve the model prediction performance; Select the random forest as the meta - model, and use the prediction results of the above two base models XGBoost and LightGBM on the training set as the input of the meta - model. The meta - model further optimizes the final prediction value by learning the prediction results of the base models, reducing the bias of a single model; Through ensemble learning, while retaining the advantages of the base models, the prediction performance and generalization ability of the model are further improved; And use the trained meta - learner to predict the samples to be predicted, and obtain the predicted value of the length of hospital stay of cholelithiasis patients.
[0046] The beneficial effects of the present invention are as follows: For the method for predicting the length of hospital stay of cholelithiasis patients based on ensemble learning of the present invention, first collect the general demographic information, surgical information, test information, and medication information of cholelithiasis patients from the hospital's HIS and electronic pathology systems to construct the benchmark dataset D; Then perform pre - processing on the sample data in dataset D, including outlier removal, missing value filling, and standardization of numerical features. Then, perform key feature screening, including feature augmentation, non - numerical feature encoding, and feature selection, to screen out the N (20) features with the highest contribution. Then, divide the dataset into a training set and a test set, establish an objective function, and use the base models XGBoost and LightGBM for iterative training to obtain the optimal model; Finally, use the random forest as the meta - model to further optimize and obtain the meta - learner. The trained meta - learner can be used to accurately predict the length of hospital stay of cholelithiasis patients; Since the method for predicting the length of hospital stay of cholelithiasis patients based on ensemble learning of the present invention can accurately predict the length of hospital stay of cholelithiasis patients, using the predicted length of hospital stay of patients is beneficial to the optimization of medical resources, can help the hospital management better allocate ward beds, improve the utilization rate of medical resources and the operation efficiency of the hospital, contribute to the overall management and control of diseases such as cholelithiasis in the hospital, and also help doctors evaluate the prognosis of patients, adjust treatment plans in a timely manner, optimize the diagnosis and treatment process, and improve patient satisfaction. Description of the Drawings
[0047] Figure 1 It is a flowchart of the method for predicting the length of hospital stay of cholelithiasis patients based on ensemble learning of the present invention;
[0048] Figure 2 It is the training flow chart of the base learner in the present invention;
[0049] Figure 3 It is the shape value graph of feature screening in the embodiment of the present invention. Detailed implementation manners
[0050] The present invention will be further described below in conjunction with the accompanying drawings and embodiments.
[0051] As Figure 1 shown, the flow chart of the prediction method for the length of hospital stay of cholelithiasis patients based on ensemble learning of the present invention is given, which is realized through the following steps:
[0052] a). Construct a benchmark data set; collect multi-dimensional diagnosis and treatment data of cholelithiasis patients from the hospital HIS and the electronic pathology system, and construct a benchmark data set D for predicting the length of hospital stay;
[0053] HIS is the abbreviation of Hospital Informance System, that is, the hospital information management system. In this step, cholelithiasis patients with the main admission diagnosis of K80 are selected according to the ICD-10 coding, and the general demographic information, test information, surgical information and medication information of each patient are extracted from the HIS system and the electronic medical record system to construct a sample feature set. A total of 4942 sample data are selected for modeling, and the specific steps are as follows:
[0054] a-1). Collect demographic information; collect the general demographic information of each cholelithiasis patient, including gender, age, ethnicity, height, weight, ward area, cost category, medical insurance status, occupation, admission route, admission diagnosis, admission date, and discharge date, from the hospital HIS and the electronic pathology system;
[0055] a-2). Collect surgical information; collect the surgical information of cholelithiasis patients, including the name of the operation, type of operation, level of operation, major operation flag, anesthesia method, and type of wound;
[0056] a-3). Collect test information; collect the test information of cholelithiasis patients, including white blood cells, hemoglobin, potassium in blood, sodium in blood, N-terminal pro-brain natriuretic peptide, creatinine, albumin, red blood cells, hematocrit, prothrombin time, D-dimer, glutamic oxaloacetic transaminase, glucose, interleukin, procalcitonin, C-reactive protein, and multi-drug resistance marker. The above test information is based on the first test result in the first medical contact, that is, for those with multiple test results during the hospital stay, the first one is taken to reduce the bias caused by treatment effects;
[0057] a-4). Collect medication information; collect the medication information of cholelithiasis patients, including aspirin, clopidogrel bisulfate tablets, and compound reserpine and triamterene tablets;
[0058] Taking the multi-dimensional diagnosis and treatment data composed of the general demographic information, surgical information, test information, and medication information of each collected cholelithiasis patient as a row, storing it in the set D, and forming the original benchmark dataset D for predicting the length of hospital stay, D = {(X1, X2,..., X n ) T}, where n is the number of collected cholelithiasis patients, X i is the multi-dimensional diagnosis and treatment data of cholelithiasis patient i, and X i = {x i1 , x i2 ,..., x im}, where m is the number of multi-dimensional diagnosis and treatment data of the patient, and x ij is the j-th multi-dimensional diagnosis and treatment data of cholelithiasis patient i.
[0059] As shown in Table 1, the multi-dimensional diagnosis and treatment data of a cholelithiasis patient in the dataset D are given, including general demographic information, surgical information, test information, and medication information.
[0060] Table 1
[0061]
[0062] To illustrate the multi-dimensional diagnosis and treatment data of cholelithiasis patients in the dataset D, Tables 2 to 5 give the multi-dimensional diagnosis and treatment data of the first 5 patients and the last 5 patients in the dataset D.
[0063] Table 2
[0064]
[0065]
[0066] Table 3
[0067]
[0068]
[0069] Table 4
[0070]
[0071]
[0072] Table 5
[0073]
[0074]
[0075] Among them, Table 3, Table 4, and Table 5 are continuation tables of Table 2. It can be seen that in the dataset D, the general demographic information, surgical information, test information, and medication information of each cholelithiasis patient are stored as one row.
[0076] b). Preprocessing of the dataset; preprocess the data in the benchmark dataset D so that the processed data meets the training requirements;
[0077] The preprocessing of the dataset is specifically achieved through steps b-1) to b-3):
[0078] b-1). Removal of outliers; based on statistical methods, remove samples with a hospitalization duration significantly exceeding the normal range to reduce the impact of outliers on model training;
[0079] For example, if a patient's hospitalization duration exceeds 20 days, it is considered that their hospitalization duration significantly exceeds the normal range, and they are removed from the dataset D.
[0080] b-2). Filling of missing values; for the missing test information values in the set D, use the column mean filling method for processing, as shown in formula (1):
[0081]
[0082] In the formula, x ij-fill is the filled value of the missing test information j of patient i, n′ is the number of non-missing values in the j-th column of the set D, and x ik is the corresponding non-missing value in the j-th column of the set D;
[0083] b-3). Standardization of numerical features; standardize the numerical features in the set D according to formula (2) so that the processed numerical values are between [0,1] to eliminate the impact of different dimensions on subsequent modeling:
[0084]
[0085] In the formula, x i ′ j is the standardized processing result of the numerical feature x ij in the set D, x imax is the maximum value of the numerical feature in the j-th column of the set D, and x imin is the minimum value of the numerical feature in the j-th column of the set D.
[0086] c). Screening of key features; perform feature engineering on the preprocessed dataset, and then screen out the key features suitable for modeling from the dataset D;
[0087] The screening of key features is specifically achieved through steps c-1) to c-3):
[0088] c-1). Feature augmentation; Generate two new features from the original benchmark dataset D. First, extract the month of the patient's admission based on the admission time of the patient to explore the potential relationship between the admission month and the length of the patient's hospital stay. Second, calculate the BMI index for each patient based on the patient's height and weight data to explore the potential relationship between the BMI index and the length of the patient's hospital stay. The BMI index is shown in formula (3):
[0089]
[0090] In the formula, w represents the patient's weight in kg, and h represents the patient's height in m;
[0091] c-2). Non-numerical feature encoding; Adopt the method of label encoding to digitally map the non-numerical features in the dataset D to obtain structured data;
[0092] For example, for the numerical features of gender, ethnicity, ward, medical insurance status, occupation, admission diagnosis, surgical name, and anesthesia method in the dataset D, they are digitized using the method of label encoding.
[0093] c-3). Feature screening; Based on the feature ranking method of the model, use random forest to screen features, and screen the top N features with the highest contribution according to the shape value for modeling. The calculation of the shape value is shown in formula (4):
[0094] y′ i =y base +f(X i ,1)+f(X i ,2)+...+f(X i ,l) (4)
[0095] In the formula, X i represents the i-th sample of the dataset D, y′ i represents the predicted value of the i-th sample, y base represents the mean of the predicted values of all samples, and f(X i ,j) represents the contribution value of the j-th feature in the i-th sample to the final predicted value y′ i ;
[0096] As Figure 3 shown, the feature screening shape value diagram in the embodiment of the present invention is given. Finally, the top 20 features with the highest contribution are selected for subsequent modeling. The top 20 features with the highest contribution selected are: surgical name, albumin, occupation code, surgical level, wound type, blood potassium, D-dimer, N-terminal pro-B-type natriuretic peptide, ward, red blood cells, interleukin, procalcitonin, prothrombin time, BMI, creatinine, anesthesia method, age, patient weight, blood sodium, glucose.
[0097] d). Model training: Input the dataset D after key feature screening into the base model for iterative training. During the training process, use the Bayesian optimization strategy to optimize the model, save the optimal base model, and use the prediction results of the base learner as new features to form a feature matrix.
[0098] As Figure 2 shown, the training flow chart of the base learner in the present invention is given. The model training is specifically implemented through steps d-1) to d-6):
[0099] d-1). Dataset division: Divide the samples in the dataset D into a training set B and a test set C according to a ratio of 7:3.
[0100] d-2). Select the base model: First, evenly divide the training set B into 5 subsets, i.e., B = [subset B1, subset B2, subset B3, subset B4, subset B5]. Select the base models XGBoost and LightGBM of the efficient gradient boosting algorithm widely used in the field of machine learning. For each base model, use 5-fold cross-validation k-flod and Bayesian optimization to optimize the hyperparameters of the model.
[0101] d-3). Set the objective function: Set the objective function as the mean absolute error MAE (mean absolute error, MAE) to evaluate the model performance, as shown in formula (5):
[0102]
[0103] In the formula, n represents the number of samples, y i represents the true value of the i-th sample, and y i ' represents the predicted value of the i-th sample. Before the start of model training, the initial value of MAE is set to infinity.
[0104] d-4). Define the search space: Define the hyperparameter search space of the model, that is, determine the value range or candidate value set of each hyperparameter of the model.
[0105] d-5). Iterative training process and cross-validation: In each iterative training and cross-validation, select four subsets as the training subsets, and the remaining one subset as the validation subset. This operation is repeated five times to ensure that each subset has served as the training set and the validation set in turn.
[0106] For each cross-validation training process, record the current MAE value, and evaluate the model performance by comparing the MAE values under different parameter combinations; after every 5 cross-validation trainings are completed, calculate the average value of MAE in these 5 training processes as the overall evaluation score for this iteration cycle, which is used to measure the generalization ability of the model under the current parameter combination.
[0107] d - 6) Model optimization and saving; repeat the training steps in d - 5) until the model converges or reaches the preset number of iterations. If the MAE score of the current iteration cycle is lower than that of the previous iteration cycle, determine that the model in the current iteration cycle is the optimal model so far and save it to ensure the continuous optimization and performance improvement of the model during the training process.
[0108] Among them, it is necessary to define the hyperparameter search space of the model, that is, determine the value range or candidate value set of each hyperparameter of the model. For example, for the base model XGBoost, the hyperparameters include the learning rate set to [0.001, 0.5], the maximum depth of the tree (max_depth) set to [3, 15], and the number of trees (n_estimators) set to [100, 1000].
[0109] e) Model ensemble and prediction; use the prediction results of the base learners to train the meta-learner, and use the trained meta-learner to predict the samples to be predicted to obtain the predicted value of the hospitalization duration of cholelithiasis patients.
[0110] In this step, the specific method of model ensemble and prediction is as follows: based on the stacking generalization stacking ensemble strategy, further improve the model prediction performance; select the random forest as the meta-model, and use the prediction results of the above two base models XGBoost and LightGBM on the training set as the input of the meta-model. The meta-model further optimizes the final prediction value by learning the prediction results of the base models, reducing the bias of a single model; through ensemble learning, while retaining the advantages of the base models, the prediction performance and generalization ability of the model are further improved; and use the trained meta-learner to predict the samples to be predicted to obtain the predicted value of the hospitalization duration of cholelithiasis patients.
[0111] It can be seen that the method for predicting the hospitalization duration of cholelithiasis patients based on ensemble learning of the present invention has the following characteristics:
[0112] (1) The method for predicting the hospitalization duration of cholelithiasis patients based on ensemble learning of the present invention constructs an accurate regression prediction model by extracting multi-dimensional diagnosis and treatment information of cholelithiasis patients (including demographic information, test information, surgical information, and medication information), and can accurately predict the hospitalization duration of patients at the early stage of hospitalization. Compared with the previous duration interval prediction method, the prediction accuracy is higher.
[0113] (2) The hospitalization duration prediction method for cholelithiasis patients based on ensemble learning in the present invention significantly enhances the data representation ability and the learning efficiency of the model through carefully designed data preprocessing steps and feature engineering means. In the data preprocessing stage, the data is first comprehensively cleaned, including handling missing values, outlier detection and correction, data type conversion, etc., to ensure the accuracy of the data input into the model. This step effectively avoids model bias caused by data quality problems and improves the stability of the prediction results. In the feature engineering stage, potential features closely related to the hospitalization duration of cholelithiasis patients are deeply mined, enhancing the ability to express the non-linear relationship between features. Through feature selection, variables highly relevant to the prediction target are selected, redundant and irrelevant features are removed, simplifying the model structure and improving the generalization ability of the prediction model.
[0114] (3) The hospitalization duration prediction method for cholelithiasis patients based on ensemble learning in the present invention uses ensemble learning for modeling. During the modeling process, cross-validation and Bayesian optimization strategies are adopted, cleverly converting the model optimization problem into the minimum value problem of the objective function to achieve fine-tuning of the model.
[0115] (4) The hospitalization duration prediction method for cholelithiasis patients based on ensemble learning in the present invention is beneficial to the optimization of medical resources through the prediction of patients' hospitalization duration. It can help hospital management better allocate ward beds, improve the utilization rate of medical resources and the operation efficiency of the hospital. As an indirect indicator of disease severity and treatment effect, the hospitalization duration not only helps the hospital in the overall management and control of diseases such as cholelithiasis, but also helps doctors evaluate the prognosis of patients, adjust treatment plans in a timely manner, optimize the diagnosis and treatment process, and improve patient satisfaction.
Claims
1. A method for predicting the length of hospital stay of patients with cholelithiasis based on ensemble learning, characterized in that, It is achieved through the following steps: a). Construct a benchmark dataset; collect multi-dimensional diagnosis and treatment data of cholelithiasis patients from the hospital HIS and the electronic pathology system, and construct a benchmark dataset D for predicting the length of hospital stay; b). Preprocessing of the dataset; Preprocess the data in the benchmark dataset D so that the processed data meets the training requirements; c). Screening of key features; Perform feature engineering on the preprocessed dataset, and then screen out the key features suitable for modeling from the dataset D; d). Model training; input the dataset D after key feature screening into the base model for iterative training. During the training process, use the Bayesian optimization strategy to tune the model, save the optimal base model, and use the prediction results of the base learner as new features to form a feature matrix; e). Model integration and prediction; Use the prediction results of the base learner to train the meta-learner, and use the trained meta-learner to predict the samples to be predicted to obtain the predicted value of the length of hospital stay of cholelithiasis patients.
2. The method for predicting the length of hospital stay of cholelithiasis patients based on ensemble learning according to claim 1, wherein The construction of the benchmark dataset described in step a) is specifically achieved through the following steps: a-1). Collect demographic information; collect general demographic information of each cholelithiasis patient, including gender, age, ethnicity, height, weight, ward area, cost category, medical insurance status, occupation, admission route, admission diagnosis, admission date, and discharge date, from the hospital HIS and the electronic pathology system; a-2). Collect surgical information; collect surgical information of cholelithiasis patients, including surgical name, surgical category, surgical level, major surgery flag, anesthesia method, and wound type; a-3). Collect test information; collect test information of cholelithiasis patients, including white blood cells, hemoglobin, potassium, sodium, N-terminal pro-brain natriuretic peptide, creatinine, albumin, red blood cells, hematocrit, prothrombin time, D-dimer, glutamic oxaloacetic transaminase, glucose, interleukin, procalcitonin, C-reactive protein, and multi-drug resistance marker. The above test information is based on the first test result in the first medical contact, that is, for patients with multiple test results during hospitalization, the first one is taken to reduce the bias caused by treatment effects; a-4). Collect medication information; collect medication information of cholelithiasis patients, including aspirin, clopidogrel bisulfate tablets, and compound reserpine and triamterene tablets; Taking the multi-dimensional diagnosis and treatment data composed of the general demographic information, surgical information, test information, and medication information of each collected cholelithiasis patient as a row, storing it in the set D, and forming the original benchmark dataset D for predicting the length of hospital stay as D = {(X1, X2,..., X n ) T}, where n is the number of collected cholelithiasis patients, X i is the multi-dimensional diagnosis and treatment data of the i-th cholelithiasis patient, and X i = {x i1 , x i2 ,..., x im}, where m is the number of multi-dimensional diagnosis and treatment data of the patient, and x ij is the j-th multi-dimensional diagnosis and treatment data of the i-th cholelithiasis patient.
3. The method for predicting the length of hospital stay of cholelithiasis patients based on ensemble learning according to claim 2, wherein, The preprocessing of the dataset described in step b) is specifically achieved through the following steps: b-1). Removal of outliers; based on statistical methods, remove samples with a hospital stay significantly exceeding the normal range to reduce the impact of outliers on model training; b-2). Filling of missing values; for the missing test information values in the set D, use the column mean filling method for processing, as shown in formula (1): where \(x\) ij-fill is the filled value of the missing test information \(j\) of patient \(i\), \(n'\) is the number of non-missing values in the \(j\)-th column of the set \(D\), and \(x\) ik is the corresponding non-missing value in the \(j\)-th column of the set \(D\); b-3). Standardization of numerical features; standardize the numerical features in the set D according to formula (2) so that the processed values are between [0,1] to eliminate the impact of different dimensions on subsequent modeling: where x' ij is the normalized result of the numerical feature x ij in set D, x imax is the maximum value of the j-th column of numerical features in set D, and x imin is the minimum value of the j-th column of numerical features in set D.
4. The method for predicting the length of hospital stay of cholelithiasis patients based on ensemble learning according to claim 3, wherein, The screening of key features described in step c) is specifically achieved through the following steps: c-1). Feature augmentation: Generate two new features from the original benchmark dataset D. First, extract the admission month of the patient based on the admission time to explore the potential relationship between the admission month and the length of the patient's hospitalization. Second, calculate the BMI index for each patient based on the patient's height and weight data to explore the potential relationship between the BMI index and the length of the patient's hospitalization. The BMI index is shown in formula (3): In the formula, w represents the patient's weight in kg, and h represents the patient's height in m. c-2). Non-numerical feature encoding: Adopt the label encoding method to digitally map the non-numerical features in the dataset D to obtain structured data. c-3). Feature screening: Based on the model-based feature ranking method, use random forest to screen features, and select the top N features with the highest contribution for modeling according to the shape value. The calculation of the shape value is shown in formula (4): y' i = y base + f(X i , 1) + f(X i , 2) +... + f(X i , l) (4) Wherein, X i represents the i-th sample of the data set D, and y' i represents the predicted value of the i-th sample, and y base represents the mean of the predicted values of all samples, and f(X i , j) represents the contribution value of the j-th feature in the i-th sample to the final predicted value y' i ; Finally, select the top 20 features with the highest contribution for subsequent modeling. The top 20 features with the highest contribution selected are: surgical name, albumin, occupation code, surgical level, wound type, blood potassium, D-dimer, N-terminal pro-B-type natriuretic peptide, ward area, red blood cells, interleukin, procalcitonin, prothrombin time, BMI, creatinine, anesthesia method, age, patient weight, blood sodium, glucose.
5. The method for predicting the length of hospital stay of cholelithiasis patients based on ensemble learning according to claim 4, wherein The model training described in step d) is specifically implemented through the following steps: d-1). Dataset division: Divide the samples in the dataset D into a training set B and a test set C according to a ratio of 7:
3. d-2). Select the base model: First, evenly divide the training set B into 5 subsets, i.e., B = [subset B1, subset B2, subset B3, subset B4, subset B5]. Select the base models XGBoost and LightGBM of the efficient gradient boosting algorithm widely used in the field of machine learning. For each base model, optimize the hyperparameters of the model using 5-fold cross-validation k-flod and Bayesian optimization. d-3). Set the objective function: Set the objective function as the mean absolute error MAE (mean absolute error, MAE) to evaluate the model performance, as shown in formula (5): where n represents the number of samples, and y i represents the true value of the i-th sample, and y' i represents the predicted value of the i-th sample. Before the start of model training, the initial value of MAE is set to infinity; d-4). Define the search space: Define the hyperparameter search space of the model, that is, determine the value range or candidate value set of each hyperparameter of the model. d-5). Iterative training process and cross-validation: In each iterative training and cross-validation, select four subsets as the training subsets, and the remaining one subset as the validation subset. This operation is repeated five times to ensure that each subset has served as the training set and the validation set in turn. For each cross-validation training process, record the current MAE value, and evaluate the model performance by comparing the MAE values under different parameter combinations; after every 5 cross-validation trainings are completed, calculate the average value of the MAE in these 5 training processes as the overall evaluation score of this iterative cycle to measure the generalization ability of the model under the current parameter combination. d - 6). Model optimization and saving: Repeat the training steps in d - 5) until the model converges or reaches the preset number of iterations. If the MAE score in the current iteration cycle is lower than that in the previous iteration cycle, determine that the model in the current iteration cycle is the optimal model so far and save it to ensure the continuous optimization and performance improvement of the model during training.
6. The method for predicting the length of hospital stay of cholelithiasis patients based on ensemble learning according to claim 5, wherein The specific method for model integration and prediction in step e) is as follows: Based on the stacking integration strategy, further improve the model prediction performance; Select the random forest as the meta - model, and use the prediction results of the above two base models, XGBoost and LightGBM, on the training set as the input of the meta - model. The meta - model further optimizes the final prediction value by learning the prediction results of the base models, reducing the bias of a single model; Through ensemble learning, while retaining the advantages of the base models, further improve the prediction performance and generalization ability of the model; And use the trained meta - learner to predict the samples to be predicted to obtain the predicted value of the hospitalization duration of patients with cholelithiasis.