A model for predicting the risk of severe SA-AKI in ICU sepsis patients based on clinical variables

By constructing a large-sample machine learning model based on the ADQI standard, key clinical variables were screened and interpreted, which solved the limitations of existing models in predicting the risk of severe SA-AKI in ICU sepsis patients, and enabled early identification of high-risk patients and provision of personalized intervention recommendations.

CN122291004APending Publication Date: 2026-06-26BEIJING CHAOYANG HOSPITAL CAPITAL MEDICAL UNIVERSITY
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
BEIJING CHAOYANG HOSPITAL CAPITAL MEDICAL UNIVERSITY
Filing Date
2026-03-24
Publication Date
2026-06-26

AI Technical Summary

Technical Problem

Existing models rely heavily on the KDIGO criteria and are based on changes in serum creatinine when predicting the risk of severe SA-AKI in ICU sepsis patients. They ignore urine output criteria, and the small sample size and limited population result in limited generalization ability and an inability to identify high-risk patients in the early stages.

Method used

Based on a large-scale prospective clinical study in China, the latest definition of ADQI was adopted as the diagnostic criteria for SA-AKI. Six clinical variables with the strongest association with severe SA-AKI were screened, and an early prediction model was constructed using various machine learning algorithms. The SHAP method was used for interpretive analysis, and a visual nomogram was established.

Benefits of technology

It achieves accurate prediction of the risk of severe SA-AKI, improves the robustness and generalization ability of the model, provides a convenient clinical application tool, and helps doctors develop personalized intervention measures.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122291004A_ABST
    Figure CN122291004A_ABST
Patent Text Reader

Abstract

This invention discloses a model for predicting the risk of severe SA-AKI in ICU sepsis patients based on clinical variables. This study is based on a recent prospective, multicenter, large-sample cohort study conducted in China. Using the latest ADQI definition as the diagnostic criteria for severe SA-AKI, 1715 ICU sepsis patients were included. Clinical variables were collected within 24 hours after diagnosis of ICU sepsis patients, and six feature variables with the strongest association with severe SA-AKI were screened. Six machine learning algorithms were used to construct an early prediction model, among which SVM and Logistic Regression algorithms showed better predictive performance. The SHAP method was introduced to generate feature importance maps and individual prediction interpretation maps, which can intuitively show the contribution of each clinical variable to the risk prediction of a specific patient and explain the reasons for the occurrence of severe SA-AKI, so as to help doctors take targeted preventive measures.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of biomedicine, specifically relating to a model for predicting the risk of severe SA-AKI in ICU sepsis patients based on clinical variables. Background Technology

[0002] Sepsis is a life-threatening organ dysfunction caused by a dysregulation of the host response to infection. According to the consensus report of the 28th Acute Disease Quality Initiative Working Group (ADQI) in 2023, sepsis-associated acute kidney injury (SA-AKI) is defined as acute kidney injury (AKI) occurring within 7 days of a sepsis diagnosis, and must meet both the sepsis-3 and Kidney Disease Improving Global Outcomes (KDIGO) diagnostic criteria for AKI.

[0003] In current clinical practice, the diagnosis of AKI is mainly based on the KDIGO Clinical Practice Guidelines for Acute Kidney Injury, which rely on two functional indicators: elevated serum creatinine and decreased urine output.

[0004] Regarding SA-AKI, although the 2023 ADQI consensus further clarified its standard definition and limited the onset time window to within 7 days after the diagnosis of sepsis, its core diagnosis still does not break through the dependence on the two traditional indicators of creatinine and urine output.

[0005] However, these traditional markers have obvious limitations: on the one hand, serum creatinine only shows a significant increase after the glomerular filtration rate decreases by about 50%, which is difficult to reflect early renal parenchymal damage; on the other hand, urine output is affected by multiple factors such as volume status, hemodynamic intervention and diuretic use, and its specificity and stability are limited. Therefore, "diagnostic" identification based on KDIGO criteria has often missed the most valuable intervention window [1].

[0006] Unlike AKI caused by other etiologies, SA-AKI often presents clinically as a severe form of AKI, namely stage 2-3 in the KDIGO staging. Unlike subclinical or mild AKI (KDIGO stage 1), which is often reversible and responds well to fluid therapy, severe AKI (KDIGO stages 2-3) often shows irreversible progression and has a limited response to interventions such as fluid therapy, which essentially reflects organic damage or even necrosis of renal tubular epithelial cells. Numerous clinical studies have shown that severe SA-AKI is significantly associated with long-term continuous deterioration of renal function and high mortality [2]. Early identification of high-risk patients with severe SA-AKI can help clinicians trigger early interventions (such as targeting specific etiologies, adjusting nephrotoxic drugs, and optimizing hemodynamics) before irreversible kidney damage occurs, thereby guiding clinical decisions, reducing kidney damage, and improving the clinical prognosis of patients.

[0007] The 2023 ADQI consensus meeting clearly stated that early identification of patients at risk of progressing to severe and / or persistent kidney injury is crucial for timely initiation of appropriate support measures.

[0008] Regardless of the severity of their condition, all SA-AKI patients should immediately adopt AKI clustering measures based on the KDIGO guidelines, namely avoiding nephrotoxic exposure, optimizing hemodynamics, strictly controlling blood glucose, and closely monitoring renal function.

[0009] For severe SA-AKI, bundled treatment measures need to be further upgraded and strengthened. The focus of treatment will be on preventing death and multiple organ failure caused by AKI. In addition, from the perspective of health economics, limited medical resources should be precisely allocated to the most needy populations [3,4]: ① High-intensity renal function monitoring: For high-risk populations of severe SA-AKI, the monitoring frequency of serum creatinine and urine output needs to be shortened, and new biomarkers need to be introduced for more accurate risk stratification in order to identify and intervene in a timely manner. ② Refined hemodynamic and fluid management: For high-risk populations of severe SA-AKI, more refined hemodynamic management needs to be implemented. Goal-oriented treatment should be carried out through non-invasive or invasive hemodynamic monitoring (such as cardiac output, stroke variability monitoring, etc.), while refined volume management (not only to avoid hypoperfusion, but also to be wary of fluid overload) should be implemented to pursue individualized renal perfusion targets and ensure that oxygen supply meets the needs of the kidneys. ③ Renal Replacement Therapy (RRT) Preparation: RRT is a core treatment for severe SA-AKI patients, especially KDIGO stage 3. When life-threatening complications occur (such as hyperkalemia, acidosis, pulmonary edema), RRT is required to replace kidney function, remove metabolic waste and excess water, and maintain homeostasis. Non-severe SA-AKI does not require RRT. For high-risk individuals with severe SA-AKI, a contingency plan should be developed in advance, RRT needs should be assessed regularly, vascular access should be established, and the timing of RRT initiation should be individualized. ④ Etiological Intervention: For high-risk individuals with severe SA-AKI, more aggressive management of underlying causes is necessary, such as more aggressive control of the source of infection, dynamic adjustment of the dosage and frequency of renally excreted drugs based on renal function, and strict avoidance of hyperglycemia and hypoglycemia.

[0010] Given the high dimensionality and complex interaction of clinical data, machine learning has become an important tool for accurately predicting AKI and has promising clinical applications. However, most existing machine learning models mainly focus on predicting the overall risk and prognosis of hospitalized patients with AKI[5]. Little attention is paid to severe AKI, which has greater clinical significance. Most existing models for predicting severe AKI are based on retrospective databases such as MIMIC-IV[6]. They have inherent defects such as missing data, inability to control confounding factors, and inconsistent diagnostic criteria for AKI when constructing predictive models. On the other hand, due to the limitations of the times, the generalization ability of these models is limited. Models trained on old data may not be applicable to current clinical practice and cannot be extended to the patient population in my country. There are a few machine learning models that rely on self-built databases in my country, but these databases have small sample sizes or are only for specific populations, which limits the generalization ability and clinical applicability of predictive models[7]. In addition, some studies excluded patients with a history of chronic kidney disease during patient screening.

[0011] Therefore, based on recent prospective, large-sample clinical research data in my country, using the latest definition of ADQI as the diagnostic criteria for SA-AKI, screening out key characteristic variables that are easily accessible in clinical practice, and then constructing an accurate model for early prediction of the risk of severe SA-AKI, has become a key technical problem that urgently needs to be solved in the field of prevention and treatment of severe SA-AKI.

[0012] References

[0013] [1] Peerapornratana S, Manrique-Caballero CL, Gómez H, Kellum JA. Acute kidney injury from sepsis: current concepts, epidemiology,pathophysiology, prevention and treatment. Kidney Int. 2019;96(5):1083–1099.

[0014] [2] White KC, Serpa-Neto A, Hurford R, et al. Sepsis-associated acute kidney injury in the intensive care unit: incidence, patient characteristics, timing, trajectory, treatment, and associated outcomes. A multicenter, observational study. Intensive Care Med. 2023, 49(9): 1079-1089.

[0015] [3] Zarbock A, Nadim MK, Pickkers P, et al. Sepsis-associated acutekidney injury: consensus report of the 28th Acute Disease Quality Initiative workgroup. Nature Reviews Nephrology. 2023, 19: 401-417.

[0016] [4] Poston JT, Koyner JL. Sepsis associated acute kidney injury.British Medical Journal. 2019, 364: k4891.

[0017] [5] Cama-Olivares A, Braun C, Takeuchi T, et al. Systematic Reviewand Meta-Analysis of Machine Learning Models for Acute Kidney Injury RiskClassification. J Am Soc Nephrol. 2025, 36(10): 1969-1983.

[0018] [6] Shi T, Lin Y, Zhao H, et al. Artificial intelligence models forpredicting acute kidney injury in the intensive care unit: a systematicreview of modeling methods, data utilization, and clinical applicability.JAMIA Open. 2025, 8 (4), ooaf065.

[0019] [7] Zhou Y, Feng J, Mei S, et al. Machine learning models forpredicting acute kidney injury in patients with sepsis-associated acuterespiratory distress syndrome. Shock. 2023;59(3): 352-359. Summary of the Invention

[0020] Currently, there is a lack of models based on large-scale prospective clinical studies in China for the early prediction of the risk of severe SA-AKI in ICU sepsis patients. To address this issue, this study, relying on a recent prospective, multicenter, large-scale cohort study conducted in China, used the latest ADQI definition as the diagnostic criteria for SA-AKI and included 1715 ICU sepsis patients. Clinical variables were collected within 24 hours after the diagnosis of ICU sepsis, and the six feature variables with the strongest association with severe SA-AKI were screened. Various machine learning algorithms were used to construct an early prediction model. The constructed models all showed good discrimination and robust performance. Combined with the SHAP method interpretation framework, it can effectively achieve global feature importance assessment and local interpretation of individual patient risk. The specific technical solution is as follows:

[0021] Firstly, constructing an early risk prediction model for severe SA-AKI.

[0022] Step S1: Patient Enrollment and Cohort Construction

[0023] This prospective, consecutive enrollment study included 2418 ICU patients with sepsis meeting the diagnostic criteria for Sepsis-3. Based on inclusion and exclusion criteria, 703 patients were excluded, resulting in a final cohort of 1715 ICU patients with sepsis. Patients were divided into a non-severe SA-AKI group and a severe SA-AKI group based on whether severe SA-AKI developed within 24 hours to 7 days after sepsis diagnosis. According to the AQDI consensus definition, 670 patients (39.07%) developed severe SA-AKI (i.e., KDIGO stage 2-3 AKI).

[0024] Step S2: Data Collection

[0025] The system collected clinical variables within 24 hours of the diagnosis of sepsis, including: patient demographics, baseline creatinine levels, past medical history, site and etiology of infection, SOFA score, vital signs, routine laboratory results, medications, and invasive treatment measures.

[0026] Step S3: Feature Variable Selection

[0027] Remove clinical variables with missing values ​​greater than 20%. For two clinical variables that are correlated to some extent, determine whether they need to be removed based on their clinical significance.

[0028] 1715 patients were randomly assigned to a training set (80%) and an independent test set (20%). All feature selection was performed on the training set, using the Boruta algorithm and LASSO independently. The six clinical variables most strongly associated with severe SA-AKI were ultimately selected: SOFA score, history of chronic kidney disease, central venous oxygen saturation (ScvO2), peripheral blood mononuclear cell count, history of chronic heart failure, and serum total bilirubin level.

[0029] Step S4: Model building, validation, and performance evaluation

[0030] Based on the six clinical variables mentioned above, six machine learning prediction models were constructed using six algorithms: Logistic Regression, Random Forest, Support Vector Classification, Lightweight Gradient Boosting Machine, Extreme Gradient Boosting, and Neural Network. The models were retrained using the complete training set data. The final models were rigorously evaluated on a fully reserved internal validation set. Model performance was assessed using ROC-AUC, accuracy, sensitivity, specificity, and F1-score. The results showed that all six algorithms constructed models exhibited good predictive performance. Among the six algorithms, the Support Vector Machine (SVM) model showed the best overall performance, followed by the Logistic Regression algorithm.

[0031] Step S5: Model Interpretation Analysis

[0032] The SHAP method was used to perform interpretive analysis on the trained prediction model, generating a feature importance ranking chart to visually demonstrate the contribution of the six clinical variables to the model output. A SHAP summary chart was drawn to further reveal the global influence pattern and direction of each feature on the prediction results.

[0033] Secondly, creating nomograms based on the Logistic regression model.

[0034] Based on the regression coefficients of each variable, a linear transformation is used to convert them into corresponding nomogram scores. Then, based on the correspondence between the total score and the probability of the outcome event, an intuitive and visual nomogram is constructed. This nomogram allows for the rapid reading of the predicted probability of a patient developing severe SA-AKI based on the actual values ​​of each variable.

[0035] Thirdly, an examination of predictive models based on 5 or 7 clinical variables.

[0036] A logistic regression model was constructed to predict severe SA-AKI based on five clinical variables (SOFA, history of chronic kidney disease, ScvO2, history of chronic heart failure, and total bilirubin level, excluding peripheral blood mononuclear cell count). The ROC curve area under the curve for predicting severe SA-AKI on the training and validation sets was 0.876 and 0.824, respectively. The area under the ROC curve for the validation set was significantly lower than that of the logistic regression model constructed with six clinical variables (0.861).

[0037] A logistic regression model was constructed to predict severe SA-AKI based on five clinical variables (SOFA, history of chronic kidney disease, ScvO2, peripheral blood mononuclear cell count, and total bilirubin level, excluding history of chronic heart failure). The results showed that the area under the ROC curve for predicting severe SA-AKI on the training and validation sets was 0.874 and 0.841, respectively. The area under the ROC curve on the validation set was lower than that of the logistic regression model constructed with six clinical variables (0.861).

[0038] A logistic regression model was constructed based on five clinical variables: SOFA, history of chronic kidney disease, ScvO2, peripheral blood mononuclear cell count, and history of chronic heart failure (without serum total bilirubin level). The model predicted severe SA-AKI on the training and validation sets with ROC curves of 0.874 and 0.848, respectively. The ROC curve area under the validation set was lower than that of the logistic regression model constructed with six clinical variables (0.861).

[0039] Based on seven clinical variables (SOFA, history of chronic kidney disease, ScvO2, peripheral blood mononuclear cell count, history of chronic heart failure, total bilirubin level, and peripheral blood lymphocyte count, with peripheral blood lymphocyte count added), an early predictive logistic regression model for severe SA-AKI was constructed. The results showed that the area under the ROC curve for predicting severe SA-AKI on the training and validation sets was 0.874 and 0.857, respectively. The area under the ROC curve on the validation set was approximately 0.861, which was similar to that of the logistic regression model constructed with six clinical variables, indicating that the added seventh variable, peripheral blood lymphocyte count, did not provide independent predictive information.

[0040] Therefore, compared with other models, the prediction model constructed based on 6 clinical variables in this invention achieved the highest validation set AUC (0.861), demonstrating optimal and robust predictive efficacy.

[0041] Compared with the prior art, the beneficial effects of the present invention are as follows:

[0042] 1. It addresses the issue of relatively few models available both domestically and internationally for early prediction of severe SA-AKI risk;

[0043] 2. Existing models for the diagnosis of SA-AKI are mostly based on the KDIGO criteria (without specifying the onset time of AKI) and are mostly based on changes in serum creatinine, ignoring urine output criteria; this invention uses the latest ADQI consensus as a unified diagnostic standard for SA-AKI.

[0044] 3. Patients with a history of chronic kidney disease were included;

[0045] 4. The sample size is large and highly representative, overcoming the limitations of existing models that are mostly based on retrospective data, small samples, and single populations, and significantly improving the robustness and generalization ability of the model.

[0046] 5. Excellent predictive performance;

[0047] 6. Convenient for clinical translation and highly practical;

[0048] The six clinical variables used in the modeling were routinely collected clinical indicators, which are easy to obtain and facilitate clinicians in quickly assessing individual risks.

[0049] 7. A nodal chart was created, providing good model visibility;

[0050] 8. It has good interpretability and can assist in clinical decision-making.

[0051] With the help of interpretability tools such as SHAP, the specific contribution of each indicator to risk prediction can be given for individual patients, helping doctors understand "why this patient is high-risk" and thus develop targeted interventions. Attached Figure Description

[0052] Figure 1 Research flowchart;

[0053] Figure 2 , 6 ROC curves of a validation set for a machine learning model;

[0054] Figure 3 Ranking of the importance of key feature variables;

[0055] Figure 4 SHAP feature summary diagram;

[0056] Figure 5 , line chart;

[0057] Figure 6 ROC curve of a predictive model based on 5 clinical variables (without peripheral blood mononuclear cell count);

[0058] Figure 7 ROC curve of a predictive model based on 5 clinical variables (lacking a history of chronic heart failure);

[0059] Figure 8 ROC curve of a predictive model based on 5 clinical variables (excluding total bilirubin level);

[0060] Figure 9 ROC curve of the prediction model based on 7 clinical variables (with peripheral blood lymphocyte count added). Detailed Implementation

[0061] The present application will now be described in further detail with reference to specific embodiments. The embodiments given are merely illustrative of the present application and are not intended to limit its scope. The embodiments provided below can serve as a guide for further improvements by those skilled in the art and do not constitute a limitation on the present application in any way.

[0062] Unless otherwise specified, the experimental methods in the following embodiments are conventional methods, performed in accordance with the techniques or conditions described in the literature in this field or in accordance with the product instructions.

[0063] Unless otherwise specified, all materials and reagents used in the following examples are commercially available.

[0064] Example 1, Sample Collection

[0065] like Figure 1 As shown, a prospective, consecutive study was conducted on 2418 patients with sepsis in the ICUs of five general tertiary hospitals in Beijing (Beijing Hospital, Peking Union Medical College Hospital, Beijing Shijitan Hospital, Beijing Jishuitan Hospital, and Beijing Chaoyang Hospital) from August 2023 to January 2026. Sepsis-3 was used as the diagnostic criterion for sepsis. The ethics approval numbers are as follows: Beijing Hospital (2023BJYYEC-150-01); Peking Union Medical College Hospital (K3148 and I-22PJ1104); Beijing Shijitan Hospital (IIT2023-007-002); Beijing Jishuitan Hospital (K2023-195-00); Beijing Chaoyang Hospital (2025-KE-869).

[0066] 1. Inclusion criteria

[0067] It meets the diagnostic criteria for sepsis-3.

[0068] 2. Exclusion criteria

[0069] (1) Multiple ICU admissions;

[0070] (2) Or under the age of 18;

[0071] (3) Or, upon admission to the ICU, the patient was diagnosed with end-stage renal disease or was receiving maintenance dialysis treatment;

[0072] (4) Or has received a kidney transplant or nephrectomy during hospitalization;

[0073] (5) Or a baseline serum creatinine level ≥4.0 mg / dL;

[0074] (6) Or missing key data for AKI diagnosis;

[0075] (7) Or the ICU stay is less than 48 hours;

[0076] (8) Acute kidney injury has occurred before admission to the ICU or within 24 hours after diagnosis of sepsis.

[0077] 3. Patient screening

[0078] A total of 2418 ICU patients meeting the Sepsis-3 criteria were continuously collected. Based on the exclusion criteria, 703 patients were excluded, including: 118 patients with multiple ICU admissions; 34 patients under 18 years of age; 76 patients with end-stage renal disease or maintenance hemodialysis; 30 patients who underwent kidney transplantation or nephrectomy during hospitalization; 152 patients with baseline serum creatinine ≥4.0 mg / dL; 34 patients with missing key information for AKI diagnosis; 36 patients with ICU stay less than 48 hours; and 223 patients who developed AKI before ICU admission or within 24 hours of sepsis diagnosis. A total of 1715 ICU patients with sepsis were ultimately included. According to the AQDI consensus definition, among the 1715 ICU patients with sepsis, 1045 (60.93%) had non-severe SA-AKI, and the remaining 20% ​​(KDIGO) had severe SA-AKI. The number of patients with stage 2-3 AKI was 670 (39.07%); the number of patients with non-severe SA-AKI included 744 patients who did not develop SA-AKI and 301 patients with mild SA-AKI (KDIGO stage 1 AKI).

[0079] 4. Data Collection and Preprocessing

[0080] The system collected clinical variables from patients within 24 hours of sepsis diagnosis, including: patient demographics (sex, age), baseline creatinine level, past medical history (chronic heart failure, chronic obstructive pulmonary disease, chronic liver failure, chronic kidney disease, diabetes, autoimmune diseases, hematological diseases, solid tumors), Sequential Organ Failure Assessment (SOFA) score, site and etiology of infection, vital signs (maximum body temperature, maximum heart rate, maximum respiratory rate, mean hourly urine output), and laboratory tests (peripheral blood leukocyte count, neutrophil count, monocyte count, lymphocyte count, platelet count, total bilirubin level, international normalized ratio, albumin, central venous oxygen saturation (ScvO2), and arterial-venous carbon dioxide partial pressure difference [CO2]. [gap], arterial blood carbon dioxide partial pressure [PaCO2], oxygenation index, lactate), drugs (vasoactive and positive inotropic drugs, hormones, nephrotoxic drugs), invasive treatment measures (renal replacement therapy, mechanical ventilation, intra-aortic balloon counterpulsation [IABP], extracorporeal membrane oxygenation [ECMO]).

[0081] 5. Patient grouping

[0082] Using the occurrence of severe SA-AKI within 24 hours to 7 days after the diagnosis of sepsis as the outcome variable, among the 1715 ICU sepsis patients included, 1045 had non-severe SA-AKI (KDIGO stage 1 AKI) and 670 had severe SA-AKI (KDIGO stage 2-3 AKI) (39.07%).

[0083] Example 2: Feature Variable Selection, Model Construction, Validation, and Performance Evaluation

[0084] 1. Feature variable selection

[0085] Clinical variables with missing values ​​greater than 20% were removed. If the absolute value of the Spearman correlation coefficient between two variables was >0.4, they were considered to have a certain degree of correlation, and the decision to remove them was made based on clinical significance. 1715 patients were randomly assigned to a training set (80%) and an independent test set (20%). All feature selection was performed on the training set. To improve the robustness of feature selection, a 10-fold cross-validation framework was incorporated: the training set was approximately divided into 10 non-overlapping subsets (8 subsets with 137 cases each, and 2 subsets with 138 cases each). In each iteration, 9 subsets were used as the training subset, and the remaining subset was used as the validation set to evaluate feature performance. This process was repeated 10 times to ensure that each subset was used as the validation set once. Based on this, the Boruta algorithm and the Least Absolute Contraction and Selection Operator Regression (LASSO) were applied to screen for feature variables that were strongly associated with severe SA-AKI. The intersection of the screening results of the two algorithms was taken to select the 6 clinical variables that were most strongly associated with severe SA-AKI as the final feature variables included in the model.

[0086] Ultimately, the six clinical variables most strongly associated with severe SA-AKI were selected: SOFA score, history of chronic kidney disease, central venous oxygen saturation (ScvO2), peripheral blood mononuclear cell count, history of chronic heart failure, and total bilirubin level. Peripheral blood lymphocyte count was the seventh clinical variable most associated with severe SA-AKI.

[0087] 2. Model building, validation and performance evaluation

[0088] Based on the above 6 feature variables, 6 algorithms were used to construct 6 machine learning prediction models. The 6 algorithms are: Logistic Regression (LR), Random Forest (RF), Support Vector Machine (SVM), Lightweight Gradient Boosting Machine (LightGBM), Extreme Gradient Boosting (XGBoost), and Artificial Neural Network (ANN).

[0089] All model development was conducted solely on the training set, employing a hierarchical 5-fold cross-validation framework and fine-tuning each algorithm using a pre-defined grid search. The model with the best average cross-validation performance was selected and retrained using the complete training set. The final model was rigorously evaluated on a fully reserved internal validation set, which had been separately partitioned before any feature selection or model training began.

[0090] The model performance was evaluated using AUC, accuracy, sensitivity, specificity, and F1-score. The results are shown in Table 1, and the models constructed using the six algorithms all exhibited good predictive performance.

[0091] Given that severe SA-AKI can lead to serious damage to kidney structure, increased conversion rate to chronic kidney disease, and mortality, early identification and intervention of high-risk patients are crucial. Therefore, while ensuring a certain level of specificity, model selection focuses more on sensitivity (recall) to maximize the screening efficacy for patients with severe SA-AKI. Comparing six algorithms, the model constructed using Support Vector Machine (SVM) showed the best overall performance, with a validation set ROC-AUC of 0.861 (95% CI 0.819-0.900), the highest F1 score (0.754), and good sensitivity (0.836) and specificity (0.756). The second best algorithm was Logistic Regression, with a validation set ROC-AUC of 0.861, sensitivity of 0.813, specificity of 0.775, and F1 score of 0.752, which best meets the clinical needs of this invention for early warning.

[0092] The ROC curves of the validation set for the 6 algorithms are as follows: Figure 2 As shown. ROC-AUC of the training set and validation set, accuracy, sensitivity, specificity, and F1 score of the validation set.

[0093] Table 1. ROC-AUC, accuracy, sensitivity, and specificity of the training and validation sets for the six algorithms.

[0094]

[0095] Example 3: Model Interpretive Analysis

[0096] Taking one preferred implementation method—the Logistic regression model—as an example, the SHAP method is used to perform interpretive analysis on the trained Logistic regression model.

[0097] like Figure 3 As shown in the feature importance ranking plot, the differences in the contribution of each key variable to the model are illustrated. The SOFA score contributes the most to the model prediction, with a mean absolute SHAP value of 1.1932; a history of chronic kidney disease ranks second, with a SHAP value of 0.9437; ScvO2 ranks third, with a SHAP value of 0.6936; peripheral blood mononuclear cell count ranks fourth, with a SHAP value of 0.1545; a history of chronic heart failure ranks fifth, with a SHAP value of 0.1166; and total bilirubin level ranks sixth, with a SHAP value of 0.0430.

[0098] like Figure 4As shown, the SHAP summary plot illustrates the contribution and direction of influence of each key feature on the model output. Each point in the plot represents a sample, with color indicating the level of the feature value (red for high values, blue for low values). The positive or negative SHAP value on the horizontal axis indicates a positive or negative impact on the predicted outcome. The results show that high values ​​or positive histories of SOFA score, chronic kidney disease or chronic heart failure, peripheral blood mononuclear cell count, and serum bilirubin level (red) highly overlap with the positive SHAP value area, suggesting that elevated levels or positive histories can significantly increase the predictive risk of severe SA-AKI. Conversely, high values ​​of ScvO2 are mainly distributed in the negative SHAP region, indicating that high values ​​of ScvO2 are a protective factor against severe SA-AKI.

[0099] Example 4: Creating a nomogram based on a Logistic regression model

[0100] The six selected clinical variables were used as predictors and incorporated into the Logistic regression model to determine the regression coefficients and standard errors of each clinical variable.

[0101] The logistic regression equation is:

[0102] Logit(P) = 2.5965 + 0.2654 × SOFA + 2.5113 × history of chronic kidney disease - 0.0885 × ScvO2 + 0.2537 × monocyte count + 0.3576 × history of chronic heart failure + 0.0024 × total bilirubin

[0103] The constructed Logistic regression model has an area under the ROC curve of 0.878 in the training set and 0.861 in the internal validation set, demonstrating good discriminative ability and robust performance. This indicates that the model can accurately identify high-risk patients who will develop severe SA-AKI, thus providing a reliable basis for clinical decision-making.

[0104] Furthermore, based on the regression coefficients of each variable, a linear transformation is used to convert them into corresponding nomogram scores. Then, based on the correspondence between the total score and the probability of the outcome event, an intuitive and visual nomogram is constructed. This transforms the complex regression equation into an intuitive visualization tool, facilitating clinicians to quickly assess individual risk and achieving a balance between predictive accuracy and clinical applicability. The nomogram is shown below. Figure 5 As shown.

[0105] The first row is the score scale, with scores ranging from 0 to 100;

[0106] The second line is the history of chronic kidney disease. No history corresponds to a score of 0 on the scale, and a history corresponds to a score of 42.5 on the scale.

[0107] The third line is the history of chronic cardiac insufficiency. If there is no history, the corresponding score is 0 points; if there is a history, the corresponding score is 6.25 points.

[0108] The fourth line is the SOFA score. A SOFA score of 2 corresponds to a score of 0 on the scale, and a SOFA score of 24 corresponds to a score of 100 on the scale. The intervals between these scores are averaged.

[0109] The fifth line is the central venous oxygen saturation (ScvO2), in percentage. A ScvO2 of 95 corresponds to a score of 0 on the scale, and a ScvO2 of 35 corresponds to a score of 90 on the scale. The intervals between these values ​​are averaged.

[0110] The sixth line is the monocyte count, unit: ×10 9 / L, Monocyte is 0, corresponding to 0 points on the scale; Monocyte is 4, corresponding to 17.5 points on the scale; the intervals between these values ​​are averaged.

[0111] The seventh line is the total bilirubin, in μmol / L. A total bilirubin of 0 corresponds to a score of 0 on the scale, and a total bilirubin of 500 corresponds to a score of 20 on the scale. The intervals between these values ​​are averaged.

[0112] The eighth row is the Total Points scale;

[0113] The ninth line represents the probability of severe SA-AKI risk, i.e., scvcrc SA-AKI risk, with a probability range of 0.1 to 0.9.

[0114] This nomogram includes the six clinical variables most strongly associated with severe SA-AKI. Each variable has a corresponding score on the Points scale at the top. By summing the six scores and projecting them vertically onto the Total Points scale at the bottom, the individualized predicted probability of the patient developing severe SA-AKI can be visually obtained on the risk axis.

[0115] Example 5: Clinical Application of Nonograph

[0116] Example of clinical application of nomogram 1: Suppose an ICU patient with sepsis has a history of chronic kidney disease and chronic heart failure (corresponding to scores of 42.5 and 6.25 on the scale, respectively). Within 24 hours of sepsis diagnosis, the SOFA score is 13 (corresponding to a score of 50 on the scale), ScvO2 is 65% (corresponding to a score of 45 on the scale), and the peripheral blood mononuclear cell count is 1.0 × 10⁻⁶. 9 / L (corresponding to a score of 5 on the scale), total bilirubin 50 μmol / L (corresponding to a score of 2.5 on the scale). The total score is 151.25, predicting that this patient has a greater than 90% probability of developing severe SA-AKI.

[0117] Example of clinical application of nomogram 2: Suppose an ICU patient with sepsis has no history of chronic kidney disease (corresponding score 0 on the scale) but has a history of chronic heart failure (corresponding score 6.25 on the scale). Within 24 hours of sepsis diagnosis, the SOFA score is 6 (corresponding score 18 on the scale), ScvO2 is 75% (corresponding score 30 on the scale), and the peripheral blood mononuclear cell count is 0.5 × 10⁻⁶. 9 / L (corresponding score 2.5 points), total bilirubin 15 μmol / L (corresponding score 0.8 points). Total score 57.55, predicting a 15% probability of this patient developing severe SA-AKI.

[0118] Example 6: Examination of a predictive model based on 5 clinical features (Comparative Examples 1-3)

[0119] Based on the following five clinical variables, an early predictive logistic model for severe SA-AKI was constructed, and the ROC-AUC was examined.

[0120] 1. Comparative Example 1

[0121] A logistic model was constructed to predict severe SA-AKI based on five clinical variables: SOFA, history of chronic kidney disease, ScvO2, history of chronic heart failure, and total bilirubin level (without peripheral blood mononuclear cell count).

[0122] The logistic regression equation is:

[0123] Logit(P) = 2.3468 + 0.2760 × SOFA + 2.6997 × history of chronic kidney disease - 0.0852 × ScvO2 + 0.3635 × history of chronic heart failure + 0.0030 × total bilirubin

[0124] The results are as follows Figure 6 As shown, the area under the ROC curve for predicting severe SA-AKI on the training and validation sets was 0.876 and 0.824, respectively. The area under the ROC curve for the validation set was significantly lower than that of the Logistic regression model constructed with 6 clinical variables, which was 0.861.

[0125] The training set AUC reflects the model's ability to fit historical data, while the validation set simulates the model's performance on new patients in the real world and is the gold standard for judging its clinical practical value.

[0126] A high AUC on the training set but a low AUC on the validation set indicates that the model is overfitting: it has memorized the training set but has not learned the true patterns, and will fail in actual clinical predictions, indicating that the model's ability is insufficient or the feature selection is inappropriate.

[0127] In this invention, the AUC of Comparative Example 1 and the model of this invention are similar on the training set, indicating comparable learning ability. However, on the validation set, the AUC of this invention is significantly higher, indicating superior generalization ability and stability. Therefore, the model of this invention is more likely to achieve accurate risk warnings in real clinical scenarios and has better predictive performance.

[0128] 2. Comparative Example 2

[0129] A logistic model was constructed to predict severe SA-AKI based on five clinical variables: SOFA, history of chronic kidney disease, ScvO2, peripheral blood mononuclear cell count, and total bilirubin level (without a history of chronic heart failure).

[0130] The logistic regression equation is:

[0131] Logit(P) = 2.5752 + 0.2666 × SOFA + 2.7403 × history of chronic kidney disease - 0.0892 × ScvO2 + 0.3615 × monocyte count + 0.0047 × total bilirubin

[0132] The results are as follows Figure 7 As shown, the area under the ROC curve for predicting severe SA-AKI on the training and validation sets was 0.874 and 0.841, respectively. The area under the ROC curve on the validation set was lower than that of the Logistic regression model constructed with 6 clinical variables, which was 0.861.

[0133] 3. Comparative Example 3

[0134] A logistic model was constructed to predict severe SA-AKI based on five clinical variables: SOFA, history of chronic kidney disease, ScvO2, peripheral blood mononuclear cell count, and history of chronic heart failure (without serum total bilirubin level).

[0135] The logistic regression equation is:

[0136] Logit(P) = 2.4108 + 0.2827 × SOFA + 2.6320 × history of chronic kidney disease - 0.0882 × ScvO2 + 0.3705 × monocyte count + 0.3006 × history of chronic heart failure

[0137] The results are as follows Figure 8As shown, the area under the ROC curve for predicting severe SA-AKI on the training and validation sets was 0.874 and 0.848, respectively. The area under the ROC curve on the validation set was lower than that of the Logistic regression model constructed with 6 clinical variables, which was 0.861.

[0138] Example 7: Examination of a predictive model based on 7 clinical features (Comparative Example 4)

[0139] Based on seven clinical variables (SOFA, history of chronic kidney disease, ScvO2, peripheral blood mononuclear cell count, history of chronic heart failure, total bilirubin level, and peripheral blood lymphocyte count) (with peripheral blood lymphocyte count added), an early predictive logistic regression model for severe SA-AKI was constructed.

[0140] The logistic regression equation is:

[0141] Logit(P) = 2.3866 + 0.2660 × SOFA + 2.6205 × history of chronic kidney disease - 0.0858 × ScvO2 + 0.3702 × monocyte count + 0.3365 × history of chronic heart failure + 0.0041 × total bilirubin - 0.0001 × lymphocyte count

[0142] The results are as follows Figure 9 As shown, the area under the ROC curve for predicting severe SA-AKI on the training and validation sets was 0.874 and 0.857, respectively. The area under the ROC curve on the validation set was approximately 0.861, which is similar to that of the Logistic regression model constructed with six clinical variables, suggesting that the newly added seventh variable, peripheral blood lymphocyte count, did not provide independent predictive information.

[0143] Therefore, compared with other models, the prediction model constructed based on six clinical variables in this invention achieved the highest validation set AUC (0.861), demonstrating optimal and robust predictive power. This result verifies that the strategy of selecting variables based on association strength ranking is reasonable and efficient. Furthermore, based on the principles of model simplicity and clinical applicability, the aforementioned six-variable model was selected as the final prediction tool.

[0144] Although the present invention has been described in detail above with general descriptions and specific embodiments, modifications or improvements can be made to it, which will be obvious to those skilled in the art. Therefore, all such modifications or improvements made without departing from the spirit of the present invention fall within the scope of protection claimed by the present invention.

Claims

1. A method for constructing an early predictive model for the risk of severe SA-AKI in ICU sepsis patients, characterized in that, The steps are as follows: (1) Data collection and grouping: ICU sepsis patients were screened according to the inclusion and exclusion criteria, and clinical variables of the ICU sepsis patients within 24 hours after diagnosis of sepsis were collected. The ICU sepsis patients were randomly assigned to a training set and a validation set. (2) Feature variable screening: The Boruta algorithm and the LASSO algorithm of minimum absolute contraction and selection operator regression were respectively used to screen feature variables with strong association with severe SA-AKI on the training set. The intersection of the screening results of the two algorithms was taken to screen out the 6 clinical variables with the strongest association with severe SA-AKI as the final feature variables to be included in the model. The six clinical variables most strongly associated with severe SA-AKI are: SOFA score, history of chronic kidney disease, central venous oxygen saturation (ScvO2), peripheral blood mononuclear cell count, history of chronic heart failure, and serum total bilirubin level. (3) Model training, validation and evaluation: The support vector machine (SVM) algorithm or logistic regression algorithm is used for model training and validation. The predictive performance is evaluated by ROC-AUC, accuracy, sensitivity, specificity and F1 score. The Logistic regression equation of the aforementioned Logistic regression algorithm is: Logit(P) = 2.5965 + 0.2654 × SOFA + 2.5113 × history of chronic kidney disease - 0.0885 × ScvO2 + 0.2537 × monocyte count + 0.3576 × history of chronic heart failure + 0.0024 × total bilirubin; The term "severe SA-AKI" refers to severe sepsis-related acute kidney injury.

2. The construction method as described in claim 1, characterized in that, Clinical variables for ICU patients diagnosed with sepsis within 24 hours include: patient demographics, baseline creatinine levels, past medical history, Sequential Organ Failure Assessment (SOFA) score, site and etiology of infection, vital signs, laboratory tests, medications, and invasive treatments. The laboratory tests include: peripheral blood leukocyte count, neutrophil count, monocyte count, lymphocyte count, platelet count, total bilirubin level, international normalized ratio, albumin, central venous oxygen saturation (ScvO2), arterial-venous carbon dioxide partial pressure difference (CO2 gap), arterial carbon dioxide partial pressure (PaCO2), oxygenation index, and lactate. The aforementioned medical history includes: chronic heart failure, chronic obstructive pulmonary disease, chronic liver failure, chronic kidney disease, diabetes, autoimmune diseases, hematological diseases, and solid tumors.

3. The construction method as described in claim 1, characterized in that, The inclusion criteria mentioned above are those that meet the diagnostic criteria for sepsis-3. Exclusion criteria include: (1) Multiple ICU admissions; (2) Or under the age of 18; (3) Or, upon admission to the ICU, the patient was diagnosed with end-stage renal disease or was receiving maintenance dialysis treatment; (4) Or has received a kidney transplant or nephrectomy during hospitalization; (5) Or a baseline serum creatinine level ≥4.0 mg / dL; (6) Or missing key data for AKI diagnosis; (7) Or the ICU stay is less than 48 hours; (8) Acute kidney injury has occurred before admission to the ICU or within 24 hours after diagnosis of sepsis.

4. A method for constructing an early predictive model for the risk of severe SA-AKI in ICU sepsis patients, characterized in that, The steps are as follows: (1) Based on the regression coefficients in the regression equation as described in claim 1, convert them into corresponding nomogram scores through linear transformation, and construct an intuitive and visual nomogram according to the correspondence between the total score and the probability of the outcome event. (2) Use the nomogram to predict the risk of severe SA-AKI in ICU sepsis patients at an early stage.

5. The construction method as described in claim 4, characterized in that, The column chart includes: The first row is the Points scale, with a score range of 0 to 100. The second line is the history of chronic kidney disease. If there is no history, the corresponding score on the scale is 0 points; if there is a history, the corresponding score on the scale is 42.5 points. The third line is the history of chronic cardiac insufficiency. If there is no history, the corresponding score is 0 points on the scale; if there is a history, the corresponding score is 6.25 points on the scale. The fourth line is the SOFA score. A SOFA score of 2 corresponds to a score of 0 on the scale, and a SOFA score of 24 corresponds to a score of 100 on the scale. The intervals between these scores are averaged. The fifth line is the central venous oxygen saturation ScvO2, in percentage. ScvO2 of 95 corresponds to 0 points on the scale, and ScvO2 of 35 corresponds to 90 points on the scale. The intervals between these values ​​are averaged. The sixth line is the monocyte count, in units of ×102. 9 / L, Monocyte is 0, corresponding to 0 points on the scale; Monocyte is 4, corresponding to 17.5 points on the scale; the intervals between these values ​​are averaged. The seventh line is the total bilirubin, in μmol / L. A total bilirubin of 0 corresponds to a score of 0 on the scale, and a total bilirubin of 500 corresponds to a score of 20 on the scale. The intervals between these values ​​are averaged. The eighth row is the Total Points scale; The ninth line represents the probability of severe SA-AKI risk, which ranges from 0.1 to 0.

9. In the nomogram, the different clinical variables in rows 2 to 7 correspond to different scores on the score scale in row 1. The sum of the scores of the clinical variables in rows 2 to 7 is projected onto the total score scale in row 8 and the corresponding position of the probability of severe SA-AKI risk in row 9, which is the predicted probability of severe SA-AKI risk.

6. An early prediction model for the risk of severe SA-AKI in ICU sepsis patients obtained by the construction method according to any one of claims 1-5.

7. The early prediction model as described in claim 6, characterized in that, The six clinical variables, ranked from largest to smallest contribution to the early prediction model, are: SOFA score, history of chronic kidney disease, central venous oxygen saturation (ScvO2), peripheral blood mononuclear cell count, history of chronic heart failure, and total bilirubin level.