Model for early predicting acute kidney injury risk of senile sepsis patient after transferring into ICU (Intensive Care Unit) based on clinical variables and immune inflammation indexes
By constructing an XGBoost algorithm model based on elderly sepsis patients and screening key clinical and immune inflammatory indicators, the problem of early prediction of acute kidney injury risk in elderly sepsis patients after ICU was solved, achieving efficient and accurate risk assessment and clinical application.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- BEIJING CHAOYANG HOSPITAL CAPITAL MEDICAL UNIVERSITY
- Filing Date
- 2026-02-06
- Publication Date
- 2026-05-19
AI Technical Summary
Existing technologies struggle to accurately predict the risk of acute kidney injury in elderly sepsis patients after admission to the ICU, especially given the time lag and insufficient model generalization ability in large-scale clinical data.
Based on clinical variables and immune inflammatory markers of 627 elderly patients with sepsis, a predictive model was constructed using the XGBoost algorithm. Key indicators such as gender, sequential organ failure score, neutrophil count, procalcitonin, lactate, heart rate, complement C3, T lymphocyte ratio, and natural killer cell ratio were selected to establish an early predictive model. Interpretive analysis was provided using the SHAP method.
It achieves efficient prediction of acute kidney injury risk in elderly patients with sepsis. The model has an ROC-AUC of no less than 0.85, demonstrating good predictive efficacy and clinical applicability. Furthermore, the model is highly interpretable and easy to apply in electronic medical record systems.
Smart Images

Figure CN122067769A_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of biomedicine, specifically relating to a model for early prediction of the risk of acute kidney injury in elderly sepsis patients after admission to the ICU based on clinical variables and immune inflammatory indicators. Background Technology
[0002] Sepsis is a life-threatening multi-organ dysfunction caused by host dysregulation due to infection, and has become one of the leading causes of death in patients in the intensive care unit. Globally, the incidence and mortality of sepsis remain high, especially posing a serious threat to elderly patients. In this complex clinical syndrome, acute kidney injury is one of the most common and devastating complications. Epidemiological studies show that about 50% of sepsis patients will develop acute kidney injury, which increases in-hospital mortality by 2-3 times and significantly increases the long-term risk of patients progressing to chronic kidney disease[1].
[0003] Currently, the diagnosis of sepsis-related acute kidney injury in clinical practice mainly relies on the clinical practice guidelines published by the Kidney Disease Improvement Global Organization (KDIGO). These guidelines depend on two functional indicators: elevated serum creatinine and decreased urine output. However, these traditional biomarkers have significant limitations: First, serum creatinine levels only rise significantly when the glomerular filtration rate decreases by more than 50%, failing to reflect early renal parenchymal damage; second, creatinine levels are influenced by various factors such as age, sex, and muscle mass, resulting in lower specificity in elderly patients; furthermore, urine output monitoring is easily affected by diuretic use and volume status, limiting its reliability. These factors collectively lead to a significant time lag in diagnosis based on the KDIGO criteria, severely limiting the valuable window for early intervention.
[0004] To achieve early identification of sepsis-related acute kidney injury, researchers have developed various clinical prediction models. Existing models are primarily based on conventional clinical variables, such as sequential organ failure assessment scores, acute physiological and chronic health assessment scores, procalcitonin levels, and lactate levels. While these models can identify high-risk patients to some extent, their predictive efficacy generally faces limitations; most models have an area under the curve (AUC) between 0.75 and 0.80, which is insufficient to meet the needs of precision medicine.
[0005] In the existing technology, although a few studies have explored the association between a single immune marker and the prognosis of sepsis, most of these studies are single-center, retrospective designs with limited sample sizes and have failed to effectively integrate immune markers with traditional clinical variables. More importantly, there is a lack of predictive models specifically for the elderly population, which is precisely the highest-risk group for sepsis-related acute kidney injury [2][3].
[0006] Some studies have explored the application value of red blood cell distribution width and fibrinogen in predicting the risk of acute kidney injury in elderly sepsis patients[4]. However, the sample size of elderly sepsis patients included was small, with only 158 patients and the proportion of acute kidney injury (AKI) in the database was too high (70.9%), which may have selection bias and limited the extrapolation of the results.
[0007] In another study that established a risk prediction model for acute kidney injury in elderly patients with sepsis based on a database[5], the training set used was based on the MIMIC-Ⅳ database from abroad, and the validation set data came only from sepsis information from a hospital in China. Therefore, the generalization ability of the model in different populations and medical environments still needs further verification.
[0008] In studies on the predictive value of dynamic monitoring of plasma SOD, CysC, and KIM-1 levels for the risk of acute kidney injury in elderly patients with sepsis[6], dynamic monitoring is poorly available, has high testing costs, and takes a long time to return results, which seriously limits its application in clinical practice.
[0009] Therefore, developing a more accurate model for early prediction of the risk of acute kidney injury in elderly sepsis patients after being transferred to the ICU, based on a large number of clinical cases in China, is an urgent clinical problem to be solved.
[0010] References
[0011] [1] Wen Jingli. Research progress on early prediction of sepsis-related acute kidney injury. [J].Advances in Clinical Medicine, 2022, 12(08):8071-8076.DOI:10.12677 / acm.2022.1281162.
[0012] [2] Zhang Hu. The value of platelet / lymphocyte ratio in early prediction and prognosis of acute kidney injury in sepsis [D]. Wannan Medical College, 2021.
[0013] [3] Ao Xue, Deng Chao, Hou Yu, et al. Analysis of risk factors for acute kidney injury in patients with sepsis and construction of Nomogram prediction model [J]. Guangdong Medical Journal, 2024, 45(6):712-716.
[0014] [4] Dou Peng. Application value of red blood cell distribution width and fibrinogen in predicting the risk of acute kidney injury in elderly patients with sepsis [J]. China Medical Guide, 2025, 23(26):88-91.
[0015] [5] Zhao Jingjing, Chen Fujin, Chen Ting, Wang Jing, Sui Xiuhua, Feng Hanmin, Yao Li. Establishment of a risk prediction model for acute kidney injury in elderly patients with sepsis based on database [J]. Chinese Journal of Geriatrics, 2023, (Vol. 2).
[0016] [6] Gou Lixia1, Liu Chaochao2, Chen Na3. Predictive value of dynamic monitoring of plasma SOD, CysC and KIM-1 levels for the risk of acute kidney injury in elderly patients with sepsis [J]. Journal of Wuhan University (Medical Sciences), 2024, (No. 4). Summary of the Invention
[0017] To address the lack of a reliable method for accurately predicting the risk of acute kidney injury (AKI) in elderly sepsis patients after ICU admission based on large-scale domestic clinical data, this study included 627 elderly sepsis patients and collected their clinical variables and immune-inflammatory markers within 24 hours of ICU admission. The top 6 clinical variables and top 4 immune-inflammatory markers associated with AKI were identified. Four machine learning algorithms were used to construct predictive models. The XGBoost algorithm demonstrated the best predictive performance, followed by the logistic regression algorithm. A visually appealing nomogram was created based on the logistic regression algorithm, establishing a highly accurate early prediction model. The specific technical solution is as follows:
[0018] S1. Patient Screening: Elderly sepsis patients aged ≥65 years who required ICU treatment due to their condition were included. The patient cohort was determined based on inclusion and exclusion criteria.
[0019] A total of 627 elderly patients with sepsis were included, of whom 270 (43.1%) developed acute kidney injury within 7 days of admission to the ICU;
[0020] S2. Data Collection: The system collects clinical variables and immune inflammatory markers within 24 hours of admission to the ICU;
[0021] S3. Patient grouping: The 627 patients included in the cohort were randomly divided into two groups at a ratio of 8:2: 501 patients in the training set and 126 patients in the validation set.
[0022] S4. Data preprocessing: Data cleaning was performed using the pandas v2.3.3 and NumPy v2.3.5 libraries;
[0023] S5. Screening and Ranking of Clinical Variables: Clinical variables and immune inflammatory markers associated with the occurrence of acute kidney injury were screened. The top 6 clinical variables were: gender, Sequential Organ Failure Assessment (SOFA) score, neutrophil count (NE), procalcitonin (PCT), lactate (Lac), and heart rate (HR). The top 4 immune inflammatory markers were: complement C3 (C3), T lymphocyte percentage (T4%), natural killer cell percentage (NK%), and complement C4 (C4).
[0024] S6. Model building, training, validation and evaluation
[0025] In the training and validation sets, the top 6 clinical variables and 4 immune inflammatory markers were used to build predictive models using four machine learning algorithms. The ROC-AUC of the four machine learning algorithms, ranked from largest to smallest, are: Extreme Gradient Boosting (XGBoost), Logistic Regression, Random Forest, and Support Vector Machine.
[0026] The predictive performance of four machine learning algorithms was evaluated using ROC-AUC, PR curves, and calibration curves.
[0027] S7. Model Interpretive Analysis: The SHAP method is used to perform interpretive analysis on the XGBoost model, generating feature importance maps and individual prediction interpretation maps;
[0028] S8. Nonograph creation: Using the top 6 clinical variables and the top 4 immune inflammatory markers as independent influencing factors, determine the regression coefficients and corresponding scores of each clinical variable and immune inflammatory marker, and create a nonograph based on the independent influencing factors.
[0029] S9. The models with 6 clinical variables, 6 clinical variables plus 5 immune inflammatory markers, and 6 clinical variables plus 3 immune inflammatory markers were examined. The results showed that their ROC-AUC was lower than that of the model with 6 clinical variables plus 4 immune inflammatory markers.
[0030] Compared with existing technologies, the advantages of this invention are: it establishes an early prediction model for the risk of acute kidney injury in elderly Chinese sepsis patients transferred to the ICU based on a large sample size, with the following advantages:
[0031] 1. Good predictive performance, with an AUC of not less than 0.85;
[0032] 2. This model is based on strict patient selection criteria and has strong clinical applicability;
[0033] 3. This model provides visual interpretive analysis through SHAP diagrams, facilitating clinical understanding and application;
[0034] 4. The clinical variables and immune inflammatory indicators included in this model can be easily obtained from electronic medical record systems used in clinical practice. Attached Figure Description
[0035] Figure 1 Research flowchart;
[0036] Figure 2 , Four PR curve of a model built using a machine learning algorithm;
[0037] Figure 3 , Four DCA curves of models constructed using various machine learning algorithms;
[0038] Figure 4 , Four ROC curves of models built using various machine learning algorithms;
[0039] Figure 5 SHAP feature importance summary diagram;
[0040] Figure 6 1. Analysis results of feature importance of the XGBoost algorithm model;
[0041] Figure 7 , line chart;
[0042] Figure 8 , nomograms of predicted ROC curves and AUC on the training and validation sets;
[0043] Figure 9 , 6 ROC curves of models combining 1 clinical variable with 0, 3, and 5 immune inflammatory indicators, respectively. Detailed Implementation
[0044] The present application will now be described in further detail with reference to specific embodiments. The embodiments given are merely illustrative of the present application and are not intended to limit its scope. The embodiments provided below can serve as a guide for further improvements by those skilled in the art and do not constitute a limitation on the present application in any way.
[0045] Unless otherwise specified, the experimental methods in the following embodiments are conventional methods, performed in accordance with the techniques or conditions described in the literature in this field or in accordance with the product instructions.
[0046] Unless otherwise specified, all materials and reagents used in the following examples are commercially available.
[0047] Example 1, Sample Collection
[0048] This prospective study collected data from 1344 elderly patients with sepsis who required ICU treatment at five comprehensive tertiary hospitals in Beijing between June 2023 and October 2025. Participating institutions and their respective ethical codes are: Peking Union Medical College Hospital (ethics code: K3148, I-22PJ1104), Beijing Shijitan Hospital (ethics code: ITT2023-007-002), Beijing Jishuitan Hospital (ethics code: K2023-195-00), Beijing Hospital (ethics code: 2023BJYYEC-150-01), and Beijing Chaoyang Hospital affiliated with Capital Medical University (ethics code: 2025-ke-869).
[0049] This study has been registered at chictr.org.cn (registration number: ChICTR2300074175). Due to ethical and data protection requirements, case information from all participating centers was de-identified. During analysis and reporting, data from each center were presented in coded form, and the specific case distribution was not disclosed to protect patient privacy and ensure the objectivity of the analysis.
[0050] 1. Inclusion criteria
[0051] (1) Admitted to the ICU and diagnosed with sepsis;
[0052] (2) And the applicant's age is ≥65 years old;
[0053] (3) And the treatment time in the ICU exceeds 48 hours;
[0054] (4) And baseline renal function is normal (e.g., estimated glomerular filtration rate eGFR ≥ 60 mL / min / 1.73 m² at enrollment).
[0055] 2. Exclusion criteria
[0056] (1) Patients with a history of chronic kidney disease (CKD);
[0057] (2) Patients who have been admitted to the ICU for less than 48 hours;
[0058] (3) Or patients with autoimmune diseases, tumors, or blood diseases;
[0059] (4) Patients with a missing rate of more than 30% for key clinical variables or immune indicators.
[0060] 3. Patient screening
[0061] Based on the inclusion and exclusion criteria, 214 cases of CKD, 286 cases of tumors, 74 cases of autoimmune diseases, 54 cases of hematological diseases, and 89 cases with an ICU stay of less than 48 hours were excluded, totaling 717 cases. Finally, 627 elderly sepsis patients were included, of whom 270 (43.1%) developed acute kidney injury within 7 days of diagnosis. Sepsis-associated acute kidney injury (SA-AKI) was diagnosed according to the Sepsis-3.0 combined with KDIGO criteria.
[0062] 4. Data Collection
[0063] Clinical variables were collected from 627 elderly patients with sepsis within 24 hours of ICU admission, including patient age, sex, medical history, whether admitted via the emergency department, whether transferred due to pulmonary infection, infection focus, comprehensive organ function (expressed using the Sequential Organ Failure Scale (SOFA)), and disease severity (expressed using the Acute Physiology and Chronic Health Evaluation II (APACHE) score). II) indicates the following: basic vital signs (including body temperature, heart rate, mean arterial blood pressure, respiratory rate, blood lactate, superior vena cava oxygen saturation, arterial-venous carbon dioxide partial pressure difference, and amount of vasoactive drugs used); organ function (including left ventricular ejection fraction, oxygenation index, serum creatinine, blood urea nitrogen, alanine aminotransferase, total bilirubin, direct bilirubin, and serum amylase); collection of all subject outcome-related indicators, including ICU stay, hospital stay, ventilator use time, and 28-day mortality; collection of peripheral blood from all subjects to analyze their immune and inflammatory status, including interleukin-6, interleukin-8, interleukin-10, tumor necrosis factor-α, neutrophil count, lymphocyte count, neutrophil-to-lymphocyte ratio, high-sensitivity C-reactive protein, and procalcitonin; collection of peripheral blood to analyze immunoglobulin A, immunoglobulin G, immunoglobulin M, complement C3, complement C4, and lymphocyte subsets (including CD19). + Total B lymphocyte count, CD3 + Total number of T lymphocytes, CD4 + T lymphocytes, CD8 + T lymphocytes and CD3 - CD16 + CD56 + Natural killer (NK) cells.
[0064] In addition, peripheral blood samples were collected at 0, 24, 48, and 72 hours after admission to the ICU to determine absolute lymphocyte counts, generating longitudinal sequence measurements. Peripheral blood mononuclear cells were stained with fluorescent monoclonal antibodies and then analyzed by flow cytometry using a three-color EPICS-XL flow cytometer (Beckman Coulter, Brea, CA) to detect T cells (CD3+). + CD4 +and CD8 + T cell subsets, B cells (CD19) + ) and NK cells (CD3) - CD16 + CD56 + ).
[0065] 5. Grouping
[0066] Using a random allocation method, 627 subjects were randomly divided into two groups at an 8:2 ratio. The model training set contained 501 subjects, and the model validation set contained 126 subjects. The case screening process is detailed below. Figure 1 .
[0067] Example 2: Model Construction, Training, Validation, and Evaluation
[0068] 1. Screening of clinical variables and immune inflammatory markers
[0069] In the training set, Python 3.10 was used as the primary computing platform for the entire data analysis and model building process. During the data preprocessing stage, we used the pandas (v2.3.3) and NumPy (v2.3.5) libraries for data cleaning.
[0070] Feature selection follows the principle of "methodological priority and clinical rationality verification".
[0071] The specific process is as follows: First, univariate tests are performed on the candidate clinical variables and immune inflammatory markers on the cleaned training set. Continuous variables are tested using the Welch t-test (without assuming homogeneity of variance), and binary variables are also tested using the t-test as a robust approximation, and sorted by p-value from smallest to largest. Then, collinearity control and pruning are performed according to the following constraints:
[0072] (1) If SOFA and APACHE II coexist, only the one with a stronger univariate effect (smaller p) is retained;
[0073] (2) When white blood cells (WBC) and norepinephrine (NE) coexist, NE is preferentially preserved; based on this,
[0074] Six clinical variables were automatically added from the top-ranked univariate variables without any forced inclusion. The automatically selected variables were: gender, Sequential Organ Failure Assessment (SOFA) score, neutrophil count (NE), procalcitonin (PCT), lactate (Lac), and heart rate (HR). In the initial screening based on routine clinical indicators, we systematically evaluated the performance of different combinations of features based on the feature importance ranking generated by the XGBoost model. The results showed that the first six features constituted a "performance inflection point": they covered the core dimensions of disease pathophysiology, and the model's discriminative ability reached a stable plateau at this point. Variables ranked 7th (prothrombin time, PT) and below showed significantly lower importance scores, and their inclusion in the independent validation set did not statistically significantly improve model performance; instead, it increased model complexity and the risk of overfitting. Therefore, selecting the first six variables achieved an optimal balance between simplicity and robustness while ensuring model effectiveness.
[0075] From the top-ranked univariate immune inflammatory markers, four markers with the most substantial incremental changes were automatically selected: complement C3 (C3), T lymphocyte percentage (T4%), natural killer cell percentage (NK%), and complement C4 (C4). Immune variables were only retained as percentages (%) or quantitative measurements, and were not used in parallel with counts (#) to avoid redundancy and information leakage. The above names are the results of the actual run and were not pre-specified. The same selection logic was used when supplementing immune inflammatory markers. The top four immune markers represent key and complementary pathways such as innate immunity, specific responses, and inflammatory mediators, providing the model with new information independent of conventional clinical indicators. From a data-driven perspective, these four markers collectively brought significant performance gains. However, the fifth-ranked marker, the percentage of B lymphocytes in lymphocytes (B%), showed a sharp decline in importance, and its added predictive value was limited; selecting only the top three would have missed an important immune dimension, resulting in incomplete information. Therefore, selecting the first four immune indicators represents a precise balance between maximizing the contribution of immune information and avoiding the introduction of redundant noise.
[0076] In summary, the final "6+4" feature combination was not arbitrarily set, but determined through triple validation using cross-validation performance curves, feature importance breakpoint analysis, and clinicopathological significance. This ensures that the model fully utilizes core information from different data sources while maintaining a streamlined structure, guaranteeing its generalization ability and ease of use in future clinical practice.
[0077] The names mentioned above are the actual filtering results from this run, and are not pre-specified.
[0078] 2. Model building, training, validation and evaluation
[0079] To build an efficient prediction model, we trained and compared four mainstream classification algorithms: logistic regression, random forest, support vector machine (SVM), and extreme gradient boosting tree (XGBoost).
[0080] Building a prediction model based on the XGBoost algorithm first requires systematic data preprocessing, including handling missing values, encoding categorical variables, and feature standardization. Then, the XGBoost model is initialized, selecting the appropriate objective function and evaluation metric based on the prediction task type, and setting basic parameters such as the learning rate, maximum tree depth, and number of trees.
[0081] A predictive model was constructed using logistic regression. Gender, SOFA, neutrophil count, procalcitonin, lactate, heart rate, C3, T4%, NK%, and C4 were identified as independent influencing factors. The regression coefficients for each independent influencing factor were determined, and the logistic regression equation is as follows:
[0082] Logit(P)=4.4589+0.0328×SOFA+0.0080×PCT+0.1185×Lac+0.0380×NE+0. 0065×HR+0.9184×gender-2.8986×C3-0.0812×T4%+0.0438×C4-0.1144×NK%
[0083] In the above formula, for gender, male is 1 and female is 0.
[0084] Table 1 shows the comparison results of ROC-AUC, AP, sensitivity, and specificity predicted by the models built by the four machine learning algorithms on the training set.
[0085] Table 1. Comparison of ROC-AUC, sensitivity, and specificity of four machine learning algorithms on the training set.
[0086]
[0087] As shown in Table 1, the ROC-AUC values predicted by the models built by the four machine learning algorithms in the training set are sorted from largest to smallest: Extreme Gradient Boosting (XGBoost), Logistic Regression, Random Forest, and Support Vector Machine (SVM).
[0088] PR curve ( Figure 2The results showed that XGBoost (AP=0.818) and Logistic Regression (AP=0.812) performed better, indicating that they could better balance precision and recall in the case of imbalanced positive and negative samples; Random Forest (AP=0.780) and SVM (AP=0.764) performed slightly worse.
[0089] Further analysis of the DCA curve reveals the model's clinical applicability. Figure 3 The results showed that XGBoost had a higher net benefit than the "all-inclusive" or "no-treatment" strategies within a wider threshold range (e.g., 0.2-0.8), indicating that it can provide a better risk-benefit ratio when used for clinical decision-making and has better translational application value.
[0090] Table 2 shows the comparison results of ROC-AUC, AP, sensitivity, and specificity predicted by the models built by the four machine learning algorithms on the validation set.
[0091] Table 2. Comparison of ROC-AUC, sensitivity, and specificity of four machine learning algorithms on the validation set.
[0092]
[0093] As shown in Table 2, the ROC-AUC values predicted by the four machine learning algorithms on the validation set are sorted from largest to smallest: XGBoost, Logistic Regression, Random Forest, and SVM.
[0094] Figure 4 The ROC curves for four machine learning algorithms are shown on the training and validation sets.
[0095] From Table 1-2 and Figure 2-4 It can be seen that among the four machine learning algorithms, the Extreme Gradient Boosting (XGBoost) model has the best prediction performance, followed by the Logistic Regression model.
[0096] The XGBoost model cannot provide a simple formula because it is essentially a committee composed of hundreds or thousands of "decision trees." During prediction, each tree assigns a score based on a series of "if...then..." rules (such as "if lactate is greater than 3"), and the final result is the sum of the scores from all the trees. This process captures the complex interactions between lactate and inflammatory markers, but it also makes its decision-making logic, like the human brain's comprehensive judgment, impossible to simplify into a single formula.
[0097] In clinical applications, its practicality lies precisely in this "comprehensive judgment" capability. It can be embedded in electronic medical record systems, transforming into a tireless super assistant. When an elderly patient with sepsis is admitted, the system can instantly analyze all indicators, including their SOFA score and procalcitonin, directly outputting a quantified probability of acute kidney injury risk (e.g., "85% high risk"), providing doctors with the most direct early warning. Simultaneously, through "interpretable AI" technology, we can retrospectively analyze why the model made a high-risk judgment, for example, discovering that the specific combination of "extremely high SOFA score combined with mildly elevated lactate" triggered the alarm. Therefore, its core value is not providing manually calculable equations, but rather providing a high-precision risk radar trained on massive amounts of data, driving early intervention and accurate monitoring, ultimately becoming a powerful aid to doctors' decision-making.
[0098] Example 3: Model Interpretive Analysis
[0099] The SHAP method is used to perform interpretive analysis on the trained XGBoost model, generating a SHAP feature importance summary diagram. Figure 2 The SHAP summary plot shows the overall contribution of each variable to the prediction results. This SHAP summary plot reveals the key impact of each feature on the prediction results in the XGBoost model. T4% and NK% are the most significant negative features, and the higher the feature value (red), the lower the model output. HR (heart rate) is a significant positive feature, and the higher the heart rate, the higher the output. C3 and C4 (complement index) also have a negative impact, and the lower the feature value (purple), the lower the output. NE (neutrophils), SOFA score, PCT (procalcitonin) and Lac (lactic acid) have a positive impact, and the higher the feature value, the higher the output. The gender feature shows gender differences: males (1) positively drive the output, and females (0) negatively pull down the output. The color bars (blue→red) indicate the level of feature values, and the scatter distribution intuitively reflects the correlation between feature values and SHAP values. Overall, negative features (T4%, NK%, C3, C4) dominate model decisions. High T4% and NK% reduce risk prediction, while high HR and SOFA increase prediction probability, providing a clear basis for model interpretation.
[0100] Figure 6The results of feature importance analysis based on the XGBoost algorithm are presented, with features ranked from highest to lowest importance. Complement component C3 (g / L) is the most important predictive feature, with an importance value close to 0.2, dominating the model's decision-making. Gender is the second most important feature, with an importance value of approximately 0.127, indicating a significant impact of gender on model predictions. Immune-related indicators T4% (importance ≈ 0.115) and NK% (importance ≈ 0.110) rank third and fourth, respectively, reflecting the crucial role of cellular immune status in the model. Neutrophil percentage (NE, importance ≈ 0.095) and heart rate (HR, importance ≈ 0.078) follow closely behind, showing the contribution of inflammatory response and circulatory status to prediction. Procalcitonin (PCT, importance ≈ 0.072) and lactate (Lac, importance ≈ 0.070) were equally important as markers of infection and metabolism. The SOFA score (importance ≈ 0.060) had a relatively low impact on the degree of organ failure, while complement C4 (g / L) had the lowest importance (≈ 0.055) and contributed the least to the model. Overall, immune indicators (C3, T4%, NK%) and basal physiological parameters (NE, HR) had a more significant driving effect on model prediction, suggesting that the contribution of these features should be given special attention in relevant prediction tasks.
[0101] Example 4: Application of the XGBoost Model
[0102] Application Example 1: Suppose an elderly sepsis patient, upon admission to the ICU, has a SOFA score of 10, procalcitonin level of 5 ng / mL, lactate level of 3 mmol / L, and neutrophil count of 15 × 10⁻⁶. 9 The values of the following variables were input into the XGBoost predictive model: / L, heart rate 110 beats / min, sex male, complement C3 1.2g / L, T lymphocyte percentage 50%, complement C4 0.3g / L, and NK cell percentage 10%. The model output an acute kidney injury risk probability of 23%, indicating a low risk of SA-AKI.
[0103] Application Example 2: Suppose an elderly sepsis patient, upon admission to the ICU, has a SOFA score of 12, procalcitonin level of 50 ng / mL, lactate level of 4 mmol / L, and neutrophil count of 20 × 10⁻⁶. 9 The values of the following variables were input into the XGBoost predictive model: / L, heart rate 110 beats / min, male sex, complement C3 0.85g / L, T lymphocyte percentage 40%, complement C4 0.25g / L, and NK cell percentage 8%. The model output an acute kidney injury risk probability of 80.9%, indicating a high risk of SA-AKI, requiring immediate clinical intervention.
[0104] In clinical applications, its practicality lies precisely in this "comprehensive judgment" capability. It can be embedded in electronic medical record systems, transforming into a tireless super assistant. When an elderly patient with sepsis is admitted, the system can instantly analyze all indicators, including their SOFA score and procalcitonin, directly outputting a quantified probability of acute kidney injury risk (e.g., "85% high risk"), providing doctors with the most direct early warning. Simultaneously, through "interpretable AI" technology, we can retrospectively analyze why the model made a high-risk judgment, for example, discovering that the specific combination of "extremely high SOFA score combined with mildly elevated lactate" triggered the alarm. Therefore, its core value is not providing manually calculable equations, but rather providing a high-precision risk radar trained on massive amounts of data, driving early intervention and accurate monitoring, ultimately becoming a powerful aid to doctors' decision-making.
[0105] Example 5: Establishing a predictive model for the risk of acute kidney injury in elderly sepsis patients transferred to the ICU using nomograms.
[0106] A nomogram is based on a multi-factor model. It integrates multiple independent variables and uses line segments with scales to draw the fitted functional relationships in the multi-factor model on the same plane. It is used to express the interrelationships and relative importance of the independent variables in the prediction model.
[0107] Using the top 6 clinical variables and top 4 immune inflammatory markers associated with acute kidney injury identified in Example 2 as independent influencing factors, the regression coefficients and corresponding scores of each independent influencing factor were determined. Based on the logistic regression equation in Example 2, a nomogram was created based on the independent influencing factors. The nomogram simplifies the complex relationship between the 6 clinical variables and 4 immune inflammatory markers. (Normally shown). Figure 7 As shown.
[0108] The first line is a score scale, with a score range of 0 to 100;
[0109] The second line is the SOFA score. A SOFA of -2 corresponds to a score of 0 on the scale, and a SOFA of 24 corresponds to a score of 16.25 on the scale. The scores are divided into average intervals.
[0110] The third row is procalcitonin (PCT). A PCT score of -10 corresponds to a score of 0 on the scale, and a PCT score of 120 corresponds to a score of 18 on the scale. The intervals between these values are averaged.
[0111] The fourth row is lactic acid (Lac). A Lac of 0 corresponds to a score of 0 on the scale, and a Lac of 24 corresponds to a score of 15 on the scale. The intervals between these values are averaged.
[0112] The fifth row shows the neutrophil count (NE). An NE of 0 corresponds to a score of 0 on the scale, and an NE of 90 corresponds to a score of 53 on the scale. The intervals between these values are averaged.
[0113] The sixth line is heart rate (HR). An HR of 0 corresponds to a score of 0 on the scale, and an HR of 200 corresponds to a score of 15 on the scale. The intervals between these values are averaged.
[0114] The seventh category is gender, with females receiving 0 points and males receiving 15 points.
[0115] The eighth row is C3. C3 is 2, which corresponds to 0 points on the score scale. C3 is 0, which corresponds to 76 points on the score scale. The interval between them is divided equally.
[0116] The ninth row is T4%, where T4% is 100 points and corresponds to 0 points on the score scale, and T4% is 10 points and corresponds to 99 points on the score scale. The intervals between these two values are averaged.
[0117] The tenth row is C4. C4 is 0, which corresponds to 0 points on the score scale. C4 is 1, which corresponds to 6 points on the score scale. The intervals between them are averaged.
[0118] The eleventh row is NK%. NK% is 65, which corresponds to 0 points on the score scale. NK% is 0, which corresponds to 94 points on the score scale. The intervals between these two values are averaged.
[0119] The twelfth and thirteenth rows represent the total score and its corresponding predicted probability of acute kidney injury, with a probability range of 0.1 to 0.9.
[0120] In the nomogram, rows 2 through 11 represent the top 6 clinical variables and the top 4 immune inflammatory markers associated with acute kidney injury within 7 days. Different clinical variables and immune inflammatory markers correspond to different scores on the scale. The sum of the scores of the clinical variables and immune inflammatory markers in rows 2 through 11 is projected onto the corresponding positions in rows 12 and 13, which is the predicted probability of acute kidney injury within 7 days.
[0121] The ROC curves for the nomogram prediction of the training and validation sets are shown below. Figure 8 The AUCs were 0.856 and 0.879, respectively.
[0122] Example 6: Application of Nodal Charts
[0123] Application Example 1: Suppose an elderly sepsis patient, upon admission to the ICU, has a SOFA score of 10, procalcitonin level of 5 ng / mL, lactate level of 3 mmol / L, and neutrophil count of 15 × 10⁻⁶. 9The values of the variables were: / L, heart rate 110 beats / min, sex male, complement C3 1.2g / L, T lymphocyte percentage 50%, complement C4 0.3g / L, and NK cell percentage 10%. The total score was 206 points, and the AKI probability was 21%, which is close to the 23% of the XGBoost model.
[0124] Application Example 2: Suppose an elderly sepsis patient, upon admission to the ICU, has a SOFA score of 12, procalcitonin level of 50 ng / mL, lactate level of 4 mmol / L, and neutrophil count of 20 × 10⁻⁶. 9 The values of the variables were: / L, heart rate 110 beats / min, sex male, complement C3 0.85g / L, T lymphocyte percentage 40%, complement C4 0.25g / L, and NK cell percentage 8%. The total score was 245 points, and the AKI probability was 81%, which is close to the 80.9% of the XGBoost model.
[0125] Example 7: Predictive model for 6 clinical variables (Comparative Example 1)
[0126] A predictive model for six variables—gender, SOFA, neutrophil count, procalcitonin, lactate, and heart rate—was established using logistic regression. The logistic regression equation is as follows:
[0127] Logit(P) = -3.7375 + 0.1152 × SOFA + 0.0102 × Procalcitonin + 0.1235 × Lactate + 0.0465 × Neutrophil Count + 0.0112 × Heart Rate + 0.8819 × Sex
[0128] The prediction model was evaluated on the training set, and the results are as follows: Figure 9 As shown, the ROC-AUC is 0.717, which is less than the ROC-AUC (0.862) of the model constructed using logistic regression algorithm with 10 variables (6 clinical variables and 4 immune inflammatory indicators) in Example 2. This indicates that the predictive performance of the model with 6 clinical variables is not as good as that of the model with 10 variables (6 clinical variables and 4 immune inflammatory indicators).
[0129] Example 8: Predictive model for 6 clinical variables and 5 immune inflammatory markers (Comparative Example 2)
[0130] A logistic regression algorithm was used to establish a predictive model for eleven variables: gender, SOFA, neutrophil count, procalcitonin, lactate, heart rate, C3, T4%, NK%, C4, and B%. The logistic regression equation is as follows:
[0131] Logit(P)=5.275+0.05×SOFA+0.012×PCT+0.046×Lac+0.06×NE+0.007×H R+0.923×gender-2.689×C3-0.094×T4%+0.145×C4-0.123×NK%-0.047×B%
[0132] The prediction model was evaluated on the training set, and the results are as follows: Figure 9 As shown, the ROC-AUC was 0.863, which is only 0.001 higher than the ROC-AUC (0.862) of the model constructed using logistic regression with 10 variables (6 clinical variables and 4 immune inflammatory markers) in Example 2. Although B% may be statistically significant in univariate analysis, its contribution to the overall discriminative power of the model is limited and does not reach the generally accepted threshold for clinically or practically significant improvement. Considering the principles of model simplicity, avoiding overfitting, and the convenience of future clinical applications, we believe that the evidence for including the B% variable in the current model is insufficient and therefore it was removed from the final model.
[0133] Example 9: Predictive model for 6 clinical variables and 3 immune inflammatory markers (Comparative Example 3)
[0134] A predictive model for nine variables—gender, SOFA, neutrophil count, procalcitonin, lactate, heart rate, C3, T4%, and NK%—was established using logistic regression. The logistic regression equation is as follows:
[0135] Logit(P)=3.883+0.045×SOFA+0.01×PCT+0.042×Lac+0.043×NE+0.007×HR+1.102×gender-2.7×C3-0.08×T4%-0.104×NK%
[0136] The prediction model was evaluated on the training set, and the results are as follows: Figure 9 As shown, the ROC-AUC is 0.83, which is less than the ROC-AUC (0.862) of the model constructed using the logistic regression algorithm with 6 clinical variables and 4 immune inflammatory indicators (a total of 10 variables) in Example 2. This indicates that the model with 6 clinical variables and 3 immune inflammatory indicators (a total of 9 variables) performs worse than the model with 6 clinical variables and 4 immune inflammatory indicators (a total of 10 variables).
[0137] The above embodiments demonstrate that the extreme gradient boosting algorithm, logistic regression algorithm, and nomogram prediction model provided by the present invention have high accuracy, interpretability, and clinical applicability, and can provide an effective tool for early warning of acute kidney injury in elderly patients with sepsis.
[0138] Although the present invention has been described in detail above with general descriptions and specific embodiments, modifications or improvements can be made to it, which will be obvious to those skilled in the art. Therefore, all such modifications or improvements made without departing from the spirit of the present invention fall within the scope of protection claimed by the present invention.
Claims
1. A method for constructing an early predictive model for the risk of acute kidney injury in elderly sepsis patients transferred to the ICU, characterized in that, The steps are as follows: (1) Data collection and grouping: elderly sepsis patients were screened according to the inclusion and exclusion criteria, and clinical variables and immune inflammatory indicators of patients were collected within 24 hours of admission to the ICU. Patients were randomly assigned to the training set and the validation set. (2) Data preprocessing: Use pandas and NumPy libraries to clean the training set; (3) Screening and ranking of clinical variables and immune inflammatory markers: The screening follows the principle of "methodological priority and clinical rationality verification". The process is as follows: univariate tests are performed on the cleaned training set for candidate clinical variables and immune inflammatory markers. Welch t test is used for continuous variables, without assuming homogeneity of variance. T test is also used for binary variables as a robust approximation, and they are ranked in ascending order of p-value. Subsequently, collinearity control and trimming are performed according to the following constraints: A. If both SOFA and APACHE II exist, only the one with a stronger univariate effect, i.e., a smaller p-value, is retained. B. When WBC and NE coexist, NE is retained first; based on this, the number of variables is automatically increased to 6 from the variables that are ranked first in the single-variable sort. The top 6 univariate clinical variables associated with the occurrence of acute kidney injury were: gender, sequential organ failure score (SOFA), neutrophil count (NE), procalcitonin (PCT), lactate (Lac), and heart rate (HR). From the top-ranked univariate immune inflammatory markers, the 4 markers with the most significant increase were automatically selected: complement C3 (C3), T lymphocyte percentage (T4%), natural killer cell percentage (NK%), and complement C4 (C4). (4) Model training, validation and evaluation: The XGBoost algorithm is used for model training and validation, and the prediction performance is evaluated by ROC curve.
2. The construction method as described in claim 1, characterized in that, The clinical variables include: patient age, gender, past medical history, whether admitted through the emergency department, whether transferred due to pulmonary infection, infection focus, SOFA score, APACHE II, baseline vital signs, and organ function; the immune inflammatory markers include: interleukin-6, interleukin-8, interleukin-10, tumor necrosis factor-α, neutrophil count, lymphocyte count, neutrophil-to-lymphocyte ratio, high-sensitivity C-reactive protein, procalcitonin; immunoglobulin A, immunoglobulin G, immunoglobulin M, complement C3, complement C4, and lymphocyte subsets. The basic vital signs mentioned include body temperature, heart rate, mean arterial blood pressure, respiratory rate, blood lactate, superior vena cava oxygen saturation, arterial-venous carbon dioxide partial pressure difference, and the amount of vasoactive drugs used. The organ functions mentioned include: left ventricular ejection fraction, oxygenation index, serum creatinine, blood urea nitrogen, alanine aminotransferase, total bilirubin, direct bilirubin, and serum amylase. The lymphocyte subsets mentioned include: CD19 + Total B lymphocyte count, CD3 + Total number of T lymphocytes, CD4 + T lymphocytes, CD8 + T lymphocytes and CD3 - CD16 + CD56 + Natural killer (NK) cells.
3. The construction method as described in claim 1, characterized in that, The inclusion criteria include: patients diagnosed with sepsis upon admission to the ICU, aged ≥65 years, treated in the ICU for more than 48 hours, and with normal baseline renal function; the exclusion criteria include: patients with a history of chronic kidney disease, or patients admitted to the ICU for less than 48 hours, or patients with autoimmune diseases, tumors, hematological diseases, or patients with a missing rate of more than 30% of key clinical variables.
4. An early prediction model for the risk of acute kidney injury in elderly sepsis patients transferred to the ICU, obtained by the construction method described in any one of claims 1-3.
5. The early prediction model as described in claim 4, characterized in that, The early prediction model was subjected to interpretive analysis using the SHAP method to demonstrate the overall contribution of each variable to the prediction results.
6. A method for constructing an early predictive model for the risk of acute kidney injury in elderly sepsis patients transferred to the ICU, characterized in that, The steps are as follows: (1) Data collection and grouping: elderly sepsis patients were screened according to the inclusion and exclusion criteria, and clinical variables and immune inflammatory indicators of patients were collected within 24 hours of admission to the ICU. Patients were randomly assigned to the training set and the validation set. (2) Data preprocessing: Use pandas and NumPy libraries to clean the training set; (3) Screening and ranking of clinical variables and immune inflammatory markers: The screening follows the principle of "methodological priority and clinical rationality verification". The process is as follows: univariate tests are performed on the cleaned training set for candidate clinical variables and immune inflammatory markers. Welch t test is used for continuous variables, without assuming homogeneity of variance. T test is also used for binary variables as a robust approximation, and they are ranked in ascending order of p-value. Subsequently, collinearity control and trimming are performed according to the following constraints: A. If both SOFA and APACHE II exist, only the one with a stronger univariate effect, i.e., a smaller p-value, is retained. B. When WBC and NE coexist, NE is retained first; based on this, the number of variables is automatically increased to 6 from the variables that are ranked first in the single-variable sort. The top 6 univariate clinical variables associated with the occurrence of acute kidney injury were: gender, sequential organ failure score (SOFA), neutrophil count (NE), procalcitonin (PCT), lactate (Lac), and heart rate (HR). From the top-ranked univariate immune inflammatory markers, the 4 markers with the most significant increase were automatically selected: complement C3, T lymphocyte percentage, natural killer cell percentage, and complement C4. (4) Model training, validation and evaluation: The predictive model was constructed using the logistic regression algorithm: The top 6 clinical variables and the top 4 immune inflammatory markers were used as independent influencing factors, and the regression coefficients of each independent influencing factor were determined. The logistic regression equation is as follows: Logit(P)=4.4589+0.0328×SOFA+0.0080×PCT+0.1185×Lac+0.0380×NE+0. 0065×HR+0.9184×gender-2.8986×C3-0.0812×T4%+0.0438×C4-0.1144×NK% In the gender field, male is represented by 1 and female by 0. The model was trained and validated using a logistic regression algorithm, and its predictive performance was evaluated using the ROC curve.
7. An early prediction model for the risk of acute kidney injury in elderly sepsis patients transferred to the ICU, obtained by the construction method described in claim 6.
8. The nomogram obtained by the construction method according to claim 6, characterized in that, The aforementioned nomogram can be used to visually determine the early risk of acute kidney injury in elderly sepsis patients transferred to the ICU, as follows: (1) A nomogram was constructed using the top 6 clinical variables and the top 4 immune inflammatory markers as independent influencing factors. The top 6 clinical variables and the top 4 immune inflammatory markers were identified as independent influencing factors. The scores of the top 6 clinical variables were determined, and a nomogram was created based on the independent influencing factors according to the logistic regression equation. (2) Using the nomogram to predict the risk of acute kidney injury in elderly sepsis patients after admission to the ICU Clinical information was collected from elderly sepsis patients within 24 hours of admission to the ICU, including gender, sequential organ failure score, neutrophil count, procalcitonin, lactate, heart rate, complement C3, T lymphocyte percentage, natural killer cell percentage, and complement C4. The corresponding scores for each clinical variable and immune inflammatory marker were found in the nomogram. The total score was obtained by summing the scores of each clinical variable and immune inflammatory marker. The probability of acute kidney injury within 7 days was then calculated in the "Probability of Acute Kidney Injury" column.
9. The nodal chart as described in claim 8, characterized in that, The nomogram includes: The first line is a score scale, with a score range of 0 to 100; The second line is the SOFA score. A SOFA of -2 corresponds to a score of 0 on the scale, and a SOFA of 24 corresponds to a score of 16.25 on the scale. The scores are divided into average intervals. The third row is procalcitonin PCT. A PCT of -10 corresponds to a score of 0 on the scale, and a PCT of 120 corresponds to a score of 18 on the scale. The intervals between these values are averaged. The fourth row is lactic acid (Lac). A Lac value of 0 corresponds to a score of 0 on the scale, and a Lac value of 24 corresponds to a score of 15 on the scale. The intervals between these values are averaged. The fifth row shows the neutrophil count (NE). An NE of 0 corresponds to a score of 0 on the scale, and an NE of 90 corresponds to a score of 53 on the scale. The intervals between these values are averaged. The sixth line is the heart rate (HR). An HR of 0 corresponds to a score of 0 on the scale, and an HR of 200 corresponds to a score of 15 on the scale. The intervals between these values are averaged. The seventh line is gender, with 0 points for females and 15 points for males. The eighth row is C3. C3 is 2, which corresponds to 0 points on the score scale. C3 is 0, which corresponds to 76 points on the score scale. The interval between them is divided equally. The ninth row is T4%, where T4% is 100 points and corresponds to 0 points on the score scale, and T4% is 10 points and corresponds to 99 points on the score scale. The intervals between these values are averaged. The tenth row is C4. C4 is 0, which corresponds to 0 points on the score scale. C4 is 1, which corresponds to 6 points on the score scale. The intervals between them are averaged. The eleventh row is NK%. NK% is 65, which corresponds to 0 points on the score scale. NK% is 0, which corresponds to 94 points on the score scale. The intervals between these two values are averaged. The twelfth and thirteenth rows represent the total score and its corresponding predicted probability of acute kidney injury, with a probability range of 0.1 to 0.
9. In the nomogram, rows 2 through 11 represent the top 6 clinical variables and the top 4 immune inflammatory markers associated with acute kidney injury within 7 days. Different clinical variables and immune inflammatory markers correspond to different scores on the scale. The sum of the scores of the clinical variables and immune inflammatory markers in rows 2 through 11 is projected onto the corresponding positions in rows 12 and 13, which is the predicted probability of acute kidney injury within 7 days.