Model for early prediction of acute kidney injury risk after geriatric sepsis patient is transferred into ICU (Intensive Care Unit) based on clinical variables

By screening clinical variables of elderly sepsis patients, a prediction model based on random forest and logistic regression was constructed, which solved the problem of early prediction of the risk of acute kidney injury after elderly sepsis patients were transferred to the ICU. It provides an intuitive and visual prediction tool that is suitable for application in primary hospitals and improves the accuracy and interpretability of the prediction.

CN122067768APending Publication Date: 2026-05-19BEIJING CHAOYANG HOSPITAL CAPITAL MEDICAL UNIVERSITY
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202610173085.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-02-06
Publication Date
2026-05-19

AI Technical Summary

Technical Problem

Existing technologies are insufficient for early prediction of the risk of acute kidney injury after ICU admission in elderly sepsis patients, especially given the lack of model generalization ability in the Chinese population, and the time lag and reliability issues of traditional biomarkers.

Method used

Based on clinical variables from 627 elderly patients with sepsis, the top 6 clinical variables associated with acute kidney injury were selected. A predictive model was constructed using random forest and logistic regression algorithms, and interpretive analysis was performed using the SHAP method to create an intuitive nomogram, providing an early predictive model.

Benefits of technology

This model enables early prediction of acute kidney injury risk in elderly sepsis patients in China. It has strong clinical accessibility, rapid results, is suitable for application in primary hospitals, and is highly interpretable, helping clinicians to take targeted preventive measures.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122067768A_ABST
    Figure CN122067768A_ABST
Patent Text Reader

Abstract

The invention discloses a model for early prediction of acute kidney injury risks after geriatric sepsis patients are transferred into ICU based on clinical variables, 627 geriatric sepsis patients are included in the research, clinical variables of the geriatric sepsis patients in 24 hours in the ICU are collected, and the first 6 clinical variables related to acute kidney injury occurrence are screened out; four machine learning algorithms are adopted to construct prediction models based on the first 6 clinical variables ranked, the prediction models of the random forest and the logistic regression algorithm are good in performance, and a visual column diagram is made according to the logistic regression algorithm. Therefore, the model which is high in clinical accessibility and is used for early predicting the acute kidney injury risk of the senile sepsis patient after the senile sepsis patient is transferred into ICU is established.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of biomedicine, specifically relating to a model for early prediction of the risk of acute kidney injury in elderly sepsis patients after admission to the ICU based on clinical variables. Background Technology

[0002] Sepsis is a life-threatening multi-organ dysfunction caused by host dysregulation due to infection, and has become one of the leading causes of death in patients in the intensive care unit. Globally, the incidence and mortality of sepsis remain high, especially posing a serious threat to elderly patients. In this complex clinical syndrome, acute kidney injury is one of the most common and devastating complications. Epidemiological studies show that about 50% of sepsis patients will develop acute kidney injury, which increases in-hospital mortality by 2-3 times and significantly increases the long-term risk of patients progressing to chronic kidney disease[1].

[0003] Currently, the diagnosis of sepsis-related acute kidney injury in clinical practice mainly relies on the clinical practice guidelines published by the Kidney Disease Improvement Global Organization (KDIGO). These guidelines depend on two functional indicators: elevated serum creatinine and decreased urine output. However, these traditional biomarkers have significant limitations: First, serum creatinine levels only rise significantly when the glomerular filtration rate decreases by more than 50%, failing to reflect early renal parenchymal damage; second, creatinine levels are influenced by various factors such as age, sex, and muscle mass, resulting in lower specificity in elderly patients; furthermore, urine output monitoring is easily affected by diuretic use and volume status, limiting its reliability. These factors collectively lead to a significant time lag in diagnosis based on the KDIGO criteria, severely limiting the valuable window for early intervention.

[0004] In order to achieve early identification of sepsis-related acute kidney injury, researchers have developed a variety of clinical prediction models. Existing models are mainly based on routine clinical variables, such as sequential organ failure assessment scores, acute physiological and chronic health assessment scores, procalcitonin, lactate levels, etc. Although these models can identify high-risk patients to a certain extent, their predictive efficacy generally faces bottlenecks, such as small sample size, possible selection bias; or models based on foreign databases are not suitable for the Chinese population; or poor availability of indicators. Some studies have explored the application value of red blood cell distribution width and fibrinogen in predicting the risk of acute kidney injury in elderly sepsis patients [2], but the sample size of elderly sepsis patients included is small, with only 158 patients and the proportion of acute kidney injury (AKI) in the database is too high (70.9%), which may be due to selection bias, resulting in limited extrapolation of results. Another study that established a database-based model for predicting the risk of acute kidney injury in elderly sepsis patients[3] used a training set based on the foreign MIMIC-Ⅳ database, and the validation set data came only from sepsis information from a domestic hospital. Therefore, the generalization ability of the model in different populations and medical environments still needs further verification. In the study on the predictive value of dynamic monitoring of plasma SOD, CysC, and KIM-1 levels for the risk of acute kidney injury in elderly sepsis patients[4], dynamic monitoring has poor availability, high detection costs, and long result return time, which seriously limits its promotion and application in clinical practice.

[0005] Therefore, based on a large number of clinical cases in China, developing a clinically accessible model for early prediction of the risk of acute kidney injury in elderly sepsis patients after admission to the ICU is an urgent clinical problem to be solved.

[0006] References

[0007] [1] Wen Jingli. Research progress on early prediction of sepsis-related acute kidney injury [J]. Advances in Clinical Medicine, 2022, 12(08):8071-8076.DOI:10.12677 / acm.2022.1281162.

[0008] [2] Dou Peng. Application value of erythrocyte distribution width and fibrinogen in predicting the risk of acute kidney injury in elderly patients with sepsis [J]. China Medical Guide, 2025, 23(26):88-91.

[0009] [3] Zhao Jingjing, Chen Fujin, Chen Ting, Wang Jing, Sui Xiuhua, Feng Hanmin, Yao Li. Establishment of a risk prediction model for acute kidney injury in elderly patients with sepsis based on database [J]. Chinese Journal of Geriatrics, 2023, (Vol. 2).

[0010] [4] Gou Lixia1, Liu Chaochao2, Chen Na3. Predictive value of dynamic monitoring of plasma SOD, CysC and KIM-1 levels for the risk of acute kidney injury in elderly patients with sepsis [J]. Journal of Wuhan University (Medical Sciences), 2024, (No. 4). Summary of the Invention

[0011] To address the lack of a model for early prediction of acute kidney injury (AKI) risk in elderly sepsis patients admitted to the ICU based on readily accessible routine clinical variables and large-scale domestic clinical cases, this study included 627 elderly sepsis patients and collected their clinical variables within 24 hours of ICU admission. The top six clinical variables associated with AKI were identified, and four machine learning algorithms were used to construct predictive models based on these six variables. Random forest and logistic regression algorithms showed better predictive performance. A visually appealing nomogram was created using the logistic regression algorithm, establishing a clinically accessible early prediction model. The specific technical solution is as follows:

[0012] S1. Patient Screening: Elderly sepsis patients aged ≥65 years who required ICU treatment due to their condition were included. The patient cohort was determined based on inclusion and exclusion criteria.

[0013] A total of 627 elderly patients with sepsis were included, of whom 270 (43.1%) developed acute kidney injury within 7 days of admission to the ICU;

[0014] S2. Data Collection: The system collects clinical variables within 24 hours of admission to the ICU;

[0015] S3. Patient grouping: The 627 patients included in the cohort were randomly divided into two groups at a ratio of 8:2: 501 patients in the training set and 126 patients in the validation set.

[0016] S4. Data preprocessing: Data cleaning was performed using the pandas v2.3.3 and NumPy v2.3.5 libraries;

[0017] S5. Screening and ranking of clinical variables: Screening out clinical variables associated with the occurrence of acute kidney injury, the top 6 clinical variables ranked are: gender, sequential organ failure score (SOFA), neutrophil count (NE), procalcitonin (PCT), lactate (Lac), and heart rate (HR).

[0018] S6. Model building, training, validation and evaluation

[0019] In the training and validation sets, the top 6 clinical variables were used to build predictive models using four machine learning algorithms. The ROC-AUC of the four machine learning algorithms, ranked from largest to smallest, are: Random Forest, Logistic Regression, Support Vector Machine, and Extreme Gradient Boosting.

[0020] The predictive performance of four machine learning algorithms was evaluated using ROC-AUC, PR curves, and calibration curves.

[0021] S7. Model Interpretive Analysis: Use the SHAP method to perform interpretive analysis on the model, generating feature importance maps and individual prediction interpretation maps;

[0022] S8. Nonograph creation: Using the top 6 ranked clinical variables as independent influencing factors, determine the regression coefficients and corresponding scores of each clinical variable, and create a nonograph based on the independent influencing factors;

[0023] S9. The prediction models with 3 combinations of 5 clinical variables and 7 clinical variables were examined. The ROC-AUC of both models was lower than that of the prediction model with 6 clinical variables.

[0024] Compared with existing technologies, the advantages of this invention are: it establishes an early prediction model for the risk of acute kidney injury in elderly Chinese sepsis patients transferred to the ICU based on a large sample size, with the following advantages:

[0025] 1. High clinical accessibility

[0026] The clinical variables in the model can be easily obtained and calculated in the electronic medical record system in clinical work. It has the advantages of convenient acquisition, low cost and quick results, and is more suitable for large-scale application in clinical work in primary hospitals.

[0027] 2. The model has strong interpretability.

[0028] Clinicians can use the SHAP chart to identify the main risk factors for acute kidney injury in a patient, and then take targeted preventive measures.

[0029] 3. Intuitive and visual

[0030] Nodal plots can provide an early and intuitive indication of the probability of developing acute kidney injury. Attached Figure Description

[0031] Figure 1 Research flowchart;

[0032] Figure 2 , Four PR curve of the training set of a model of a machine learning algorithm;

[0033] Figure 3 , Four DCA curves of the training set of a model using a machine learning algorithm;

[0034] Figure 4 , FourROC curves of the training and validation sets of various machine learning algorithms;

[0035] Figure 5 SHAP feature importance summary diagram;

[0036] Figure 6 Feature importance analysis results based on the random forest algorithm;

[0037] Figure 7 , line chart;

[0038] Figure 8 , nomogram of ROC curves predicted on the training and validation sets;

[0039] Figure 9 , 3 ROC curves of a logistic regression model with 5 clinical variables on the training set;

[0040] Figure 10 , 7 ROC curves of a logistic regression model for 10 clinical variables on the training set. Detailed Implementation

[0041] The present application will now be described in further detail with reference to specific embodiments. The embodiments given are merely illustrative of the present application and are not intended to limit its scope. The embodiments provided below can serve as a guide for further improvements by those skilled in the art and do not constitute a limitation on the present application in any way.

[0042] Unless otherwise specified, the experimental methods in the following embodiments are conventional methods, performed in accordance with the techniques or conditions described in the literature in this field or in accordance with the product instructions.

[0043] Unless otherwise specified, all materials and reagents used in the following examples are commercially available.

[0044] Example 1, Sample Collection

[0045] This prospective study collected data from 1344 elderly patients with sepsis who required ICU treatment at five comprehensive tertiary hospitals in Beijing between June 2023 and October 2025. Participating institutions and their respective ethical codes are: Peking Union Medical College Hospital (ethics code: K3148, I-22PJ1104), Beijing Shijitan Hospital (ethics code: ITT2023-007-002), Beijing Jishuitan Hospital (ethics code: K2023-195-00), Beijing Hospital (ethics code: 2023BJYYEC-150-01), and Beijing Chaoyang Hospital affiliated with Capital Medical University (ethics code: 2025-ke-869).

[0046] This study has been registered at chictr.org.cn (registration number: ChICTR2300074175). Due to ethical and data protection requirements, case information from all participating centers was de-identified. During analysis and reporting, data from each center were presented in coded form, and the specific case distribution was not disclosed to protect patient privacy and ensure the objectivity of the analysis.

[0047] 1. Inclusion criteria

[0048] (1) Admitted to the ICU and diagnosed with sepsis;

[0049] (2) And the applicant's age is ≥65 years old;

[0050] (3) And the treatment time in the ICU exceeds 48 hours;

[0051] (4) And baseline renal function is normal (e.g., estimated glomerular filtration rate eGFR ≥ 60 mL / min / 1.73 m² at enrollment).

[0052] 2. Exclusion criteria

[0053] (1) Patients with a history of chronic kidney disease (CKD);

[0054] (2) Patients who have been admitted to the ICU for less than 48 hours;

[0055] (3) Or patients with autoimmune diseases, tumors, or blood diseases;

[0056] (4) Patients with a missing rate of more than 30% of key clinical variables.

[0057] 3. Patient screening

[0058] Based on the inclusion and exclusion criteria, 214 cases of CKD, 286 cases of tumors, 74 cases of autoimmune diseases, 54 cases of hematological diseases, and 89 cases with an ICU stay of less than 48 hours were excluded, totaling 717 cases. Finally, 627 elderly sepsis patients were included, of whom 270 (43.1%) developed acute kidney injury within 7 days of diagnosis. Sepsis-associated acute kidney injury (SA-AKI) was diagnosed according to the Sepsis-3.0 combined with KDIGO criteria.

[0059] 4. Data Collection

[0060] Clinical variables were collected from 627 elderly patients with sepsis admitted to the ICU within 24 hours, including patient age, sex, medical history, whether admitted through the emergency department, whether transferred due to pulmonary infection, infection focus, comprehensive organ function (expressed using the Sequential Organ Failure Scale (SOFA)), disease severity (expressed using the Acute Physiology and Chronic Health Evaluation II (APACHE II) score), baseline vital signs (including body temperature, heart rate, mean arterial blood pressure, respiratory rate, blood lactate, superior vena cava oxygen saturation, arterial-venous carbon dioxide partial pressure difference, and number of vasoactive drugs used), and various organ functions (including left ventricular ejection fraction, oxygenation index, serum creatinine, blood urea nitrogen, alanine aminotransferase, total bilirubin, direct bilirubin, and serum amylase). Outcome-related indicators for all subjects were collected, including ICU stay, length of hospital stay, duration of mechanical ventilation, and 28-day mortality.

[0061] 5. Grouping

[0062] Using a random allocation method, 627 subjects were randomly divided into two groups at an 8:2 ratio. The model training set contained 501 subjects, and the model validation set contained 126 subjects. The case screening process is detailed below. Figure 1 .

[0063] Example 2: Constructing a Prediction Model

[0064] 1. Screening of clinical variables

[0065] Feature selection follows the principle of "methodological priority and clinical rationality verification".

[0066] The specific process is as follows: In the data preprocessing stage, we use the pandas (v2.3.3) and NumPy (v2.3.5) libraries for data cleaning.

[0067] On the cleaned training sample, univariate tests were performed on the candidate clinical variables. Welch t test was used for continuous variables (without assuming homogeneity of variance). T test was also used for binary variables as a robust approximation, and they were sorted by p value from smallest to largest. Then collinearity control and pruning were performed according to the following constraints: (1) If SOFA and APACHE II coexist, only the one with stronger univariate effect (smaller p value) was retained; (2) When white blood cells (WBC) and NE coexist, NE was retained first. Based on this, the number of variables was automatically increased to 6 from the top univariate ranking, without any "forced inclusion".

[0068] Following the aforementioned principles and procedures, the top six clinical variables were selected as follows: gender, Sequential Organ Failure Assessment (SOFA) score, neutrophil count (NE), procalcitonin (PCT), lactate (Lac), and heart rate (HR). In the initial screening based on routine clinical indicators, we systematically evaluated the performance of different combinations of features based on the feature importance ranking generated by the random forest model. The results showed that the top six features collectively constituted a "performance inflection point": they encompassed the core dimensions of disease pathophysiology, and the model's discriminative ability reached a stable plateau at this point. Variables ranked seventh (prothrombin time, PT) and below showed significantly reduced importance scores, and their inclusion in the independent validation set did not statistically significantly improve model performance; instead, it increased model complexity and the risk of overfitting. Therefore, selecting the top six variables achieved an optimal balance between simplicity and robustness while ensuring model effectiveness.

[0069] The clinical variable names mentioned above are the results of the actual screening in this run, and are not pre-specified.

[0070] 2. Model building, training, validation and evaluation

[0071] In the training set, Python 3.10 was used as the main computing platform for the entire data analysis and model building process.

[0072] To build an efficient prediction model, we trained and compared four mainstream classification algorithms: logistic regression, random forest, support vector machine (SVM), and extreme gradient boosting tree (XGBoost).

[0073] The logistic regression algorithm involves taking the top 6 ranked clinical variables as independent influencing factors, determining the regression coefficients of each clinical variable, and constructing a logistic regression equation. The logistic regression equation is as follows:

[0074] Logit(P) = -3.7375 + 0.1152 × SOFA + 0.0102 × Procalcitonin + 0.1235 × Lactate + 0.0465 × Neutrophil Count + 0.0112 × Heart Rate + 0.8819 × Sex

[0075] Table 1 shows the comparison results of ROC-AUC, AP, sensitivity, and specificity of four machine learning models in predicting the risk of acute kidney injury on the training set.

[0076] Table 1. Comparison of ROC-AUC, sensitivity, and specificity of four machine learning algorithms on the training set.

[0077]

[0078] As shown in Table 1, the ROC-AUC values ​​predicted by the models built by the four machine learning algorithms in the training set are sorted from largest to smallest: Random Forest, Logistic Regression, Support Vector Machine (SVM), and Extreme Gradient Boosting (XGBoost).

[0079] PR curve ( Figure 2 The results show that, in the case of imbalanced positive and negative samples, the ability to balance precision and recall, ranked from highest to lowest, is: Logistic Regression, Random Fores, SVM, and XGBoost.

[0080] Further analysis of the DCA curve reveals the model's clinical applicability. Figure 3 The results showed that Logistic Regression and Random Forest had higher net benefits than the "all-inclusive" or "no-treatment" strategies within a wider threshold range (e.g., 0.2-0.8), indicating that they can provide a better risk-benefit ratio when used for clinical decision-making and have better translational application value.

[0081] Table 2 shows the comparison results of ROC-AUC, AP, sensitivity, and specificity predicted by the models built by the four machine learning algorithms on the validation set.

[0082] Table 2. Comparison of ROC-AUC, sensitivity, and specificity of four machine learning algorithms on the validation set.

[0083]

[0084] Figure 4 The ROC curves for four machine learning algorithms are shown on the training and validation sets.

[0085] From Table 1-2 and Figure 2-4 It can be seen that among the four machine learning algorithms, the Random Forest algorithm has the best prediction performance, followed by the Logistic Regression algorithm.

[0086] Example 3: Model Interpretive Analysis

[0087] The SHAP method was used to perform interpretive analysis on the trained random forest model, generating a SHAP feature importance summary plot. (See attached image.) Figure 5 This demonstrates the overall contribution of each variable to the prediction results. The feature importance analysis of this random forest model (training set) is presented using horizontal bar charts (…). Figure 6The data clearly demonstrates the contribution of each feature to the prediction results: gender ranks first with the highest importance value of 0.26, indicating that gender difference is the core driving factor of model decision-making; SOFA (Sequential Organ Failure Assessment Score) follows closely with an importance value of 0.20, highlighting the key influence of organ failure degree on prediction. NE (neutrophil count) and PCT (ng / mL) (procalcitonin) have similar contributions with an importance value of 0.143 and 0.138, both at a moderate to high level, reflecting the synergistic effect of infection and circulatory status; Lac (lactic acid) with an importance value of 0.130 is slightly lower than the former two, and HR (heart rate) has the lowest importance value (0.125), having the weakest impact in the model. Overall, gender and SOFA constitute the dominant predictive framework, HR contributes the least, and the remaining features provide complementary support, jointly constructing the model's predictive logic for the target variable and providing quantitative evidence for clinical decision-making.

[0088] Example 4: Application of the Random Forest Model

[0089] Application Example 1: Suppose a sepsis patient is admitted to the ICU with a SOFA score of 12, procalcitonin level of 35 ng / mL, lactate level of 3.8 mmol / L, and neutrophil count of 18 × 10⁻⁶. 9 The individual had a heart rate of 115 beats / min and was male. Inputting these clinical variables into a random forest model's prediction system, the model output a 78.6% probability of acute kidney injury, indicating a high risk of SA-AKI and requiring immediate clinical intervention.

[0090] Application Example 2: Suppose a sepsis patient is admitted to the ICU with a SOFA score of 4, procalcitonin level of 6 ng / mL, lactate level of 1.2 mmol / L, and neutrophil count of 7 × 10⁻⁶. 9 The individual's blood pressure was 85 / L, heart rate was 85 beats / min, and gender was female. These clinical variables were input into a random forest model's prediction system. The model output a 20.2% probability of acute kidney injury, indicating a low risk of SA-AKI.

[0091] Example 5: Establishing a predictive model for the risk of acute kidney injury in elderly sepsis patients transferred to the ICU using nomograms.

[0092] A nomogram is based on a multi-factor model. It integrates multiple independent variables and uses line segments with scales to draw the fitted functional relationships in the multi-factor model on the same plane. It is used to express the interrelationships and relative importance of the independent variables in the prediction model.

[0093] Using the top six clinical variables associated with acute kidney injury identified in Example 2 as independent influencing factors, the regression coefficients and corresponding scores of each clinical variable were determined. Based on the logistic regression equation in Example 2, a nomogram was created based on the independent influencing factors. The nomogram simplifies the complex relationships among the six clinical variables, as shown below. Figure 7 As shown.

[0094] The first line is a score scale, with a score range of 0 to 100;

[0095] The second line is the SOFA score. A SOFA of -2 corresponds to a score of 0 on the scale, and a SOFA of 24 corresponds to a score of 71 on the scale. The intervals between these scores are averaged.

[0096] The third row is procalcitonin (PCT). A PCT score of -10 corresponds to a score of 0 on the scale, and a PCT score of 120 corresponds to a score of 31.5 on the scale. The intervals between these values ​​are averaged.

[0097] The fourth row is lactic acid (Lac). A Lac of 0 corresponds to a score of 0 on the scale, and a Lac of 24 corresponds to a score of 70 on the scale. The intervals between these values ​​are averaged.

[0098] The fifth row is the neutrophil count (NE). An NE of 0 corresponds to a score of 0 on the scale, and an NE of 90 corresponds to a score of 100 on the scale. The intervals between these values ​​are averaged.

[0099] The sixth line is heart rate (HR). An HR of 0 corresponds to a score of 0 on the scale, and an HR of 200 corresponds to a score of 43 on the scale. The intervals between these values ​​are averaged.

[0100] The seventh category is gender, with females receiving 0 points and males receiving 23 points.

[0101] The eighth and ninth rows represent the total score and the corresponding predicted probability of acute kidney injury, with a probability range of 0.05 to 0.95.

[0102] In the nomogram, rows 2 through 7 are the top 6 clinical variables associated with acute kidney injury within 7 days. Different clinical variables correspond to different scores on the scale. The sum of the scores of the clinical variables in rows 2 through 7 is projected onto the corresponding positions in rows 8 and 9, which is the predicted probability of acute kidney injury within 7 days.

[0103] ROC curves for predicting the risk of acute kidney injury using nomograms on the training and validation sets are shown below. Figure 8 The AUCs were 0.719 and 0.720, respectively.

[0104] Example 6: Application of Nodal Charts

[0105] Application Example 1: Suppose a sepsis patient, upon admission to the ICU, has a SOFA score of 12 (corresponding to a score of 38 on the scale), procalcitonin level of 35 ng / mL (corresponding to a score of 10 on the scale), lactate level of 3.8 mmol / L (corresponding to a score of 15 on the scale), and neutrophil count of 18 × 10⁻⁶. 9 / L (corresponding to a score of 20 points), heart rate of 115 beats / min (corresponding to a score of 20 points), gender is male (corresponding to a score of 23 points), total score is 121 points, corresponding to an AKI probability of 80%, which is close to the AKI probability of 78.6% predicted by the random forest model.

[0106] Application Example 2: Suppose a sepsis patient, upon admission to the ICU, has a SOFA score of 4 (corresponding to a score of 16 on the scale), procalcitonin level of 6 ng / mL (corresponding to a score of 5 on the scale), lactate level of 1.2 mmol / L (corresponding to a score of 3 on the scale), and a neutrophil count of 7 × 10⁻⁶. 9 / L (corresponding to a score of 8 on the scale), heart rate of 85 beats / min (corresponding to a score of 12 on the scale), gender is female (corresponding to a score of 0 on the scale), total score is 44 points, corresponding to an AKI probability of 18%, which is close to the AKI probability of 20.2% predicted by the random forest model.

[0107] Example 7: Predictive models for five indicators: sex, SOFA, neutrophil count, lactate, and heart rate (Comparative Example 1)

[0108] A logistic regression algorithm was used to establish a predictive model for five indicators: sex, SOFA, neutrophil count, lactate, and heart rate (procalcitonin deficiency, PCT). The logistic regression equation is as follows:

[0109] logit(p) = -3.795 + 0.879 × sex + 0.124 × SOFA + 0.050 × neutrophil count + 0.127 × lactate + 0.012 × heart rate

[0110] The prediction model was evaluated on the training set, and the results are as follows: Figure 8 As shown, the ROC-AUC of the training set is 0.717, which is less than the ROC-AUC of the training set of the logistic regression algorithm based on 6 clinical variables in Example 2 (0.747), indicating that the predictive performance of the model with 5 clinical variables of procalcitonin deficiency is not as good as the model with 6 clinical variables.

[0111] Example 8: Predictive models for five indicators: sex, SOFA, neutrophil count, procalcitonin, and lactate (Comparative Example 2)

[0112] A logistic regression algorithm was used to establish a predictive model for five indicators (heart rate, heart rate) including sex, SOFA, neutrophil count, procalcitonin, and lactate. The logistic regression equation is as follows:

[0113] logit(p) = -2.817 + 0.878 × sex + 0.125 × SOFA + 0.048 × neutrophil count + 0.129 × lactate + 0.011 × procalcitonin

[0114] The prediction model was evaluated on the training set, and the results are as follows: Figure 8 As shown, the training set ROC-AUC is 0.715, which is less than the ROC-AUC (0.747) of the prediction training set of the logistic regression algorithm based on 6 clinical variables in Example 2. This indicates that the predictive performance of the model with 5 clinical variables of heart rate is not as good as the model with 6 clinical variables.

[0115] Example 9: Predictive models for five indicators: sex, SOFA, neutrophil count, procalcitonin, and heart rate (Comparative Example 3)

[0116] A logistic regression algorithm was used to establish a predictive model for five indicators (lactate deficiency, Lac) including gender, SOFA, neutrophil count, procalcitonin, and heart rate. The logistic regression equation is as follows:

[0117] logit(p) = -2.817 + 0.828 × sex + 0.128 × SOFA + 0.047 × neutrophil count + 0.012 × heart rate + 0.011 × procalcitonin

[0118] The prediction model was evaluated on the training set, and the results are as follows: Figure 8 As shown, the ROC-AUC is 0.720, which is less than the ROC-AUC (0.747) of the prediction training set of the logistic regression algorithm based on 6 clinical variables in Example 2, indicating that the predictive performance of the model with 5 clinical variables of lactate deficiency is not as good as that of the model with 6 clinical variables.

[0119] Example 10: Predictive models for seven indicators: sex, SOFA, neutrophil count, procalcitonin, heart rate, lactate, and PT (Comparative Example 4).

[0120] A predictive model for seven indicators—gender, SOFA, neutrophil count, procalcitonin, heart rate, lactate, and PT—was established using logistic regression. The logistic regression equation is as follows:

[0121] Logit(P) = -4.685 + 0.111 × SOFA + 0.011 × Procalcitonin + 0.102 × Lactate + 0.047 × NE + 0.011 × Heart Rate + 0.89 × Gender + 0.07 × Prothrombin Time

[0122] The prediction model was evaluated on the training set, and the results are as follows: Figure 10As shown, the ROC-AUC of the training set is 0.734, which is less than the ROC-AUC of the prediction training set of the logistic regression algorithm based on 6 clinical variables in Example 2 (0.747). This indicates that the predictive performance of the model with 7 clinical variables that increase prothrombin time is not as good as the model with 6 clinical variables.

[0123] The above embodiments demonstrate that the prediction model provided by the present invention has high accuracy, interpretability, and clinical applicability, and can provide an effective tool for early warning of acute kidney injury in elderly patients with sepsis.

[0124] Although the present invention has been described in detail above with general descriptions and specific embodiments, modifications or improvements can be made to it, which will be obvious to those skilled in the art. Therefore, all such modifications or improvements made without departing from the spirit of the present invention fall within the scope of protection claimed by the present invention.

Claims

1. A method for constructing an early predictive model for the risk of acute kidney injury in elderly sepsis patients transferred to the ICU, characterized in that, The steps are as follows: (1) Data collection and grouping: elderly sepsis patients were screened according to the inclusion and exclusion criteria, and clinical variables of patients were collected within 24 hours of admission to the ICU. Patients were randomly assigned to a training set and a validation set. (2) Data preprocessing: Data cleaning was performed using pandas and NumPy libraries; (3) Screening and ranking of clinical variables: Screening follows the principle of "methodological priority and clinical rationality verification". The process is as follows: univariate tests are performed on the cleaned training set for candidate clinical variables. Welch t test is used for continuous variables. Homogeneity of variance is not assumed. T test is also used for binary variables as a robust approximation. The variables are ranked from smallest to largest p-value. Then, collinearity control and trimming are performed according to the following constraints: A. If both SOFA and APACHE II exist, only the one with a stronger univariate effect, i.e., a smaller p-value, is retained. B. When WBC and NE coexist, NE is retained first; based on this, the number of variables is automatically increased to 6 from the variables that are ranked first in the single-variable sort. The top 6 univariate clinical variables associated with the occurrence of acute kidney injury were: gender, Sequential Organ Failure Assessment (SOFA) score, neutrophil count (NE), procalcitonin (PCT), lactate (Lac), and heart rate (HR). (4) Model training, validation and evaluation: The random forest algorithm is used to train and validate the model, and the prediction performance is evaluated by the ROC curve.

2. The construction method as described in claim 1, characterized in that, The clinical variables for patients admitted to the ICU within 24 hours include: patient age, gender, medical history, whether admitted through the emergency department, whether transferred due to pulmonary infection, infection focus, SOFA score, APACHE II, baseline vital signs, and organ function. The basic vital signs mentioned include body temperature, heart rate, mean arterial blood pressure, respiratory rate, blood lactate, superior vena cava oxygen saturation, arterial-venous carbon dioxide partial pressure difference, and the amount of vasoactive drugs used. The organ functions mentioned include: left ventricular ejection fraction, oxygenation index, serum creatinine, blood urea nitrogen, alanine aminotransferase, total bilirubin, direct bilirubin, and serum amylase.

3. The construction method as described in claim 1, characterized in that, The inclusion criteria include: patients diagnosed with sepsis upon admission to the ICU, aged ≥65 years, treated in the ICU for more than 48 hours, and with normal baseline renal function; the exclusion criteria include: patients with a history of chronic kidney disease, or patients admitted to the ICU for less than 48 hours, or patients with autoimmune diseases, tumors, hematological diseases, or patients with a missing rate of more than 30% of key clinical variables.

4. An early prediction model for the risk of acute kidney injury in elderly sepsis patients transferred to the ICU, obtained by the construction method described in any one of claims 1-3.

5. The early prediction model as described in claim 4, characterized in that, The early prediction model was subjected to interpretive analysis using the SHAP method to demonstrate the overall contribution of each clinical variable to the prediction results.

6. A method for constructing an early predictive model for acute kidney injury in elderly sepsis patients transferred to the ICU, characterized in that, (1) Data collection and grouping: elderly sepsis patients were screened according to the inclusion and exclusion criteria, and clinical variables of patients were collected within 24 hours of admission to the ICU. Patients were randomly assigned to a training set and a validation set. (2) Data preprocessing: Data cleaning was performed using pandas and NumPy libraries; (3) Screening and ranking of clinical variables: Screening follows the principle of "methodological priority and clinical rationality verification". The process is as follows: univariate tests are performed on the cleaned training set for candidate clinical variables. Welch t test is used for continuous variables. Homogeneity of variance is not assumed. T test is also used for binary variables as a robust approximation. The variables are ranked from smallest to largest p-value. Then, collinearity control and trimming are performed according to the following constraints: A. If both SOFA and APACHE II exist, only the one with a stronger univariate effect, i.e., a smaller p-value, is retained. B. When WBC and NE coexist, NE is retained first; based on this, the number of variables is automatically increased to 6 from the variables that are ranked first in the single-variable sort. The top 6 univariate clinical variables associated with the occurrence of acute kidney injury were: gender, Sequential Organ Failure Assessment (SOFA) score, neutrophil count (NE), procalcitonin (PCT), lactate (Lac), and heart rate (HR). (4) Model training, validation and evaluation Establish a predictive model using logistic regression: The top 6 ranked clinical variables are considered as independent influencing factors. The regression coefficients for each clinical variable are determined, and the logistic regression equation is constructed as follows: Logit(P) = -3.7375 + 0.1152 × SOFA + 0.0102 × Procalcitonin + 0.1235 × Lactate + 0.0465 × Neutrophil Count + 0.0112 × Heart Rate + 0.8819 × Sex In the gender section, male is represented by 1, and female by 0; The model was trained and validated using a logistic regression algorithm, and its predictive performance was evaluated using the ROC curve.

7. An early prediction model for the risk of acute kidney injury in elderly sepsis patients transferred to the ICU, obtained by the construction method described in claim 6.

8. The nomogram obtained by the construction method according to claim 6, characterized in that, The aforementioned nomogram can be used to visually determine the early risk of acute kidney injury in elderly sepsis patients transferred to the ICU, as follows: (1) Create a nomogram using the top 6 clinical variables as independent influencing factors. The top 6 clinical variables were identified as independent influencing factors, and their corresponding scores were determined. Based on the logistic regression equation, a nomogram was created based on the independent influencing factors. (2) Using the nomogram to predict the risk of acute kidney injury in elderly sepsis patients after admission to the ICU Clinical information such as gender, sequential organ failure score, neutrophil count, procalcitonin, lactate, and heart rate were collected from elderly sepsis patients within 24 hours after admission to the ICU. The corresponding scores of each clinical variable for the patient were found according to the nomogram. The total score of the patient was obtained by adding up the scores of each clinical variable. The specific probability value of the patient developing acute kidney injury within 7 days was obtained in the column of probability of acute kidney injury.

9. The nodal chart as described in claim 8, characterized in that, The nomogram includes: The first line is a score scale, with a score range of 0 to 100; The second line is the SOFA score. A SOFA of -2 corresponds to a score of 0 on the scale, and a SOFA of 24 corresponds to a score of 71 on the scale. The intervals between these scores are averaged. The third row is procalcitonin (PCT). A PCT score of -10 corresponds to a score of 0 on the scale, and a PCT score of 120 corresponds to a score of 31.5 on the scale. The intervals between these values ​​are averaged. The fourth row is lactic acid (Lac). A Lac of 0 corresponds to a score of 0 on the scale, and a Lac of 24 corresponds to a score of 70 on the scale. The intervals between these values ​​are averaged. The fifth row is the neutrophil count (NE). An NE of 0 corresponds to a score of 0 on the scale, and an NE of 90 corresponds to a score of 100 on the scale. The intervals between these values ​​are averaged. The sixth line is heart rate (HR). An HR of 0 corresponds to a score of 0 on the scale, and an HR of 200 corresponds to a score of 43 on the scale. The intervals between these values ​​are averaged. The seventh category is gender, with females receiving 0 points and males receiving 23 points. The eighth and ninth rows show the total score and its corresponding predicted probability of acute kidney injury, with a probability range of 0.05 to 0.

95. In the nomogram, rows 2 through 7 are the top 6 clinical variables associated with acute kidney injury within 7 days. Different clinical variables correspond to different scores on the scale. The sum of the scores of the clinical variables in rows 2 through 7 is projected onto the corresponding positions in rows 8 and 9, which is the predicted probability of acute kidney injury within 7 days.