Intelligent early warning method and system for hospitalization acquired acute kidney injury

The intelligent early warning system, built by integrating multi-source data and constructing multi-algorithm models, solves the problems of delayed HA-AKI early warning and single evaluation dimensions, realizes early and accurate HA-AKI risk warning, supports dynamic personalized assessment and clinical interpretability, and is applicable to patients with different disease severity and departments.

CN122050864APending Publication Date: 2026-05-15RUIJIN HOSPITAL AFFILIATED TO SHANGHAI JIAO TONG UNIV SCHOOL OF MEDICINE
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202610040624.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-01-13
Publication Date
2026-05-15

AI Technical Summary

Technical Problem

Existing HA-AKI early warning technology suffers from problems such as delayed warnings, single evaluation dimensions, data application limited to static data, lack of mining of time-series data, and limitations in application scenarios, which prevent the achievement of early, accurate, and automated risk warnings.

Method used

We constructed an intelligent early warning system by integrating multi-source heterogeneous data, dividing cross-departmental datasets, using two-stage feature engineering, constructing multi-algorithm models, and enhancing interpretability. The system integrates text data, laboratory indicators, medical orders, and imaging reports, dynamically updates features in real time, and uses the LightGBM model to calculate the HA-AKI risk probability.

Benefits of technology

It enables prospective early warning of HA-AKI, significantly advances the intervention window, has excellent predictive accuracy, supports dynamic personalized assessment, improves clinical interpretability, adapts to patient groups with different disease severity and departments, and provides early intervention support.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122050864A_ABST
    Figure CN122050864A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of intelligent early warning of kidney injury, and provides an intelligent early warning method for hospitalization acquired acute kidney injury, which comprises the following steps: S1, carrying out integrated acquisition and standardized preprocessing on multi-source heterogeneous data; s2, performing cross-department data set division and accurate screening of research objects; s3, screening core variables of two-stage feature engineering; s4, time sequence dynamic feature construction and secondary feature optimization are carried out, multi-dimensional time sequence features are constructed for the key variables, and the key variables and the time sequence features are integrated to form an optimized feature data set adaptive to early warning of a future preset time period; s5, multi-algorithm model construction, cross validation and interpretability enhancement are carried out, an early warning model is constructed based on multiple machine learning algorithms, and the generalization ability of the model is evaluated through cross-department verification; and S6, carrying out model clinical integration and automatic grading early warning, and selecting an optimal model to be integrated to a hospital information system. And multi-source data types are fused to provide decision support for clinic before acute kidney injury occurs.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the technical field of intelligent early warning for kidney injury, and in particular to an intelligent early warning method and system for hospitalized acute kidney injury. Background Technology

[0002] Hospital-acquired acute kidney injury (HA-AKI) is a serious complication that occurs frequently among hospitalized patients. Its onset is insidious and its progression is rapid. If it is not warned and intervened in a timely manner, it will significantly increase the difficulty of treatment, prolong the hospital stay, and increase the mortality rate, while also causing excessive consumption of medical resources. Therefore, achieving early and accurate warning of HA-AKI is a key requirement in clinical diagnosis and treatment, and it has important practical significance for improving patient prognosis and optimizing the allocation of medical resources.

[0003] Currently, clinical early warning of HA-AKI mainly relies on two core methods: regular monitoring of serum creatinine levels and recording of urine output. However, this early warning model suffers from an insurmountable diagnostic lag. Furthermore, the risk of HA-AKI is not determined by a single factor, but rather by the combined effects of multiple dimensions, including patient disease progression, dynamic changes in laboratory indicators, medication use, surgical procedures, and other clinical interventions. However, traditional early warning methods lack an effective tool that can automatically and in real-time integrate these multi-dimensional risk factors and quantify the risk of HA-AKI. This results in the timeliness, accuracy, and practicality of early warnings failing to meet clinical needs. There is an urgent need to overcome existing technological bottlenecks to achieve earlier, more accurate, and automated HA-AKI risk early warning. Specifically, existing technologies have the following five significant shortcomings: (1) The problem of delayed early warning is prominent, making it difficult to achieve prospective early warning. Current technologies generally regard elevated absolute serum creatinine or decreased urine output as the core diagnostic criteria for HA-AKI. This criterion is essentially an outcome judgment after kidney damage occurs, rather than a prospective early warning of the risk of disease. On the one hand, an increase in serum creatinine often means that the patient's kidney damage has already occurred, at which point kidney function has been substantially damaged, and the best time for prevention and early intervention has been missed. On the other hand, in the daily clinical management of general wards, patients' hourly urine output is not routinely and systematically recorded. It is difficult to capture early abnormal signals related to the occurrence of HA-AKI in a timely manner through intermittent and non-standardized urine output monitoring, which further exacerbates the delay in early warning. As a result, current technologies cannot achieve a true "early warning" function and can only play the role of "post-diagnosis".

[0004] (2) The evaluation dimensions are too narrow, and the value of multi-source clinical data has not been fully explored. Electronic medical record systems store massive and rich multi-source clinical data, covering patients' basic health information, past medical history, present medical history, various laboratory test results, imaging reports, and medical operation records. Among them, patients' medical history information (such as comorbidities, history of kidney injury, etc.) and imaging reports (such as pleural effusion, descriptions of organ dysfunction, etc.) contain key clinical clues closely related to the risk of HA-AKI. However, existing HA-AKI early warning technologies have failed to fully value and explore the potential value of these multi-source data, and are limited to a few indicators such as serum creatinine and urine output for early warning assessment. This results in a severe limitation of the information input dimensions of the early warning model, which cannot comprehensively and objectively reflect the complex pathogenesis and multi-factor interaction characteristics of HA-AKI, thus affecting the comprehensiveness and reliability of the early warning results.

[0005] (3) Data application is limited to static data and cannot reflect the dynamic changes in the patient's condition. Most existing HA-AKI early warning models rely on clinical information within 24 hours of patient admission. This type of information is static data and can only reflect the patient's physical condition at the beginning of admission, but cannot capture the dynamic evolution of the patient's condition during hospitalization. The risk of HA-AKI is not fixed, but changes dynamically with various factors such as the patient's treatment progress, disease fluctuations, changes in laboratory indicators, and medication adjustments. Early warning models built solely based on static data at the beginning of admission are difficult to accurately adapt to the patient's risk status at different stages of hospitalization, resulting in insufficient accuracy in assessing the HA-AKI risk of the patient throughout hospitalization and failing to provide continuous risk monitoring support for clinical practice.

[0006] (4) Insufficient utilization of features, failing to uncover key risk information in time-series data. Data such as laboratory indicators, medication records, and vital signs generated during clinical diagnosis and treatment all have significant time-series characteristics. These time-series data contain complex patterns and potential risk events related to the occurrence of HA-AKI, such as gradual abnormal changes in laboratory indicators, cumulative effects of drug use, and fluctuation trends in vital signs. These dynamic features are often early warning signals of HA-AKI. However, existing technologies lack the ability to deeply mine and effectively integrate such time-series data, failing to identify valuable risk patterns and key features from the time-series data. This results in the model being unable to capture early signals before the occurrence of HA-AKI, further limiting the foresight and accuracy of the warning, and failing to meet the core clinical need for early risk identification.

[0007] (5) The application scenarios are significantly limited, and the generalization ability is insufficient. Existing HA-AKI early warning models are all constructed and validated based on clinical data of patients in general wards, and no specific adaptation studies and validation work have been carried out for high-risk patient groups such as intensive care units (ICUs). ICU patients are often more critically ill, have more complex underlying diseases, more comorbidities, and more intensive treatment interventions. Their risk of HA-AKI is significantly higher than that of patients in general wards, and the pathogenesis and risk factors are also more complex. Due to the lack of validation in the high-risk scenario of the ICU, the early warning performance of existing models cannot be guaranteed in this population, making it difficult to meet the clinical early warning needs of ICU patients. This seriously restricts the clinical promotion scope and practical application value of existing technologies, and cannot provide comprehensive HA-AKI risk early warning support for hospitalized patients in different departments and with different disease severity.

[0008] In summary, existing HA-AKI early warning technologies have significant shortcomings in terms of early warning timeliness, data utilization completeness, feature mining depth, dynamic adaptation capability, and application scenario coverage. They cannot achieve earlier, more accurate, and automated risk warnings for HA-AKI in hospitalized patients. There is an urgent need for a new technical solution to address these technical issues and fill the gap in existing clinical needs. Summary of the Invention

[0009] To address the aforementioned problems, the present invention aims to provide an intelligent early warning method and system for hospitalized acquired acute kidney injury (HA-AKI). This system integrates multi-source data such as text data, laboratory indicators, medical orders, and imaging reports, dynamically updating time-series characteristics such as laboratory indicators, medication use, and surgical procedures in real time. Furthermore, it employs high-performance machine learning algorithms to calculate the risk probability of HA-AKI, thereby providing clinical decision support before the onset of acute kidney injury and improving patient prognosis.

[0010] The above-mentioned objective of this invention is achieved through the following technical solutions: A smart early warning method for hospitalized acute kidney injury includes the following steps: S1: Integrate and collect multi-source heterogeneous data and perform standardized preprocessing. Collect structured data, unstructured text data and other derived computational data of hospitalized patients. Through data cleaning, normalization and text structuring transformation, construct a standardized patient time-series data system. S2: Perform cross-departmental dataset division and precise screening of research subjects. Divide the training set and test set based on the differences in departmental risk levels, and screen suitable research subject samples according to the HA-AKI diagnostic criteria and strict inclusion and exclusion criteria. S3: Select the core variables for two-stage feature engineering. First, construct an initial feature set by hierarchical processing of missing values, and then use a regularized regression algorithm to select key variables to achieve feature dimensionality reduction and overfitting suppression. S4: Construct dynamic features and optimize secondary features for time series. Build multi-dimensional time series features for key variables and integrate key variables and time series features to form an optimized feature dataset for early warning that is adapted to a future preset time period. S5: Construct, cross-validate, and enhance interpretability of multi-algorithm models; build early warning models based on multiple machine learning algorithms; evaluate the generalization ability of models through cross-departmental validation; and quantify feature contribution by combining interpretability algorithms. S6: Perform clinical integration of models and automated hierarchical early warning, select the optimal model to integrate into the hospital information system, and build an automated early warning closed loop of timed data update - feature processing - risk prediction - clinical decision support.

[0011] Furthermore, in step S1, multi-source heterogeneous data is integrated, collected, and standardized preprocessed. Structured data, unstructured text data, and other derived computational data from hospitalized patients are collected. Through data cleaning, normalization, and text structuring transformation, a standardized patient time-series data system is constructed, specifically as follows: S11: Define the data source and time range: The data comes from the hospital information platform of the medical institution and is included in all inpatient data of the institution, including cardiology, cardiac surgery, respiratory medicine, emergency medicine and critical care medicine, within the preset time period. This ensures that the data covers patient groups with different disease severity and department types, and guarantees the representativeness and diversity of the data, so as to provide comprehensive sample support for subsequent model training and validation. S12: Conduct multi-dimensional data collection: The types of data collected include structured data, unstructured text data, and other derived computational data; Structured data covers: Demographic information: Record ID, gender, age, height, weight, admission time, discharge time, length of hospital stay, admission department, admission diagnosis, discharge department, discharge diagnosis, and discharge status; Laboratory indicators: complete blood count, inflammatory markers, renal function indicators, liver function indicators, cardiac function indicators, coagulation function indicators, and electrolyte indicators; Vital signs data: respiratory rate, systolic blood pressure, diastolic blood pressure, heart rate, body temperature, and blood oxygen saturation; Medical records of medication and treatment: including records of the use of anticoagulants, nonsteroidal anti-inflammatory drugs, diuretics, vasoactive drugs, nephrotoxic antibiotics, and contrast agents, as well as clinical treatment records of mechanical ventilation and renal replacement therapy; Unstructured text data includes present medical history texts and imaging reports containing descriptions of pleural effusion; Other derived calculation data include the ratio of blood urea nitrogen to blood creatinine, the ratio of neutrophils to lymphocytes, and the sequential organ failure SOFA score (excluding the nervous system module). The sequential organ failure SOFA score must be greater than or equal to the preset score. S13: Implement data standardization preprocessing: Structured data preprocessing: Data cleaning algorithms are used to identify and process missing values ​​and detect and remove outliers in structured data. Then, normalization is used to unify the data scale and eliminate the dimensional differences between different indicators. Finally, time window alignment technology is used to align the structured data collected at different time points according to the preset time granularity to form standardized patient time series data. Unstructured text data structuring transformation: Using natural language processing technology, an entity recognition and relation extraction framework is built based on a pre-trained biomedical language model. Semantic parsing is performed on current medical history texts and imaging reports to accurately extract key clinical information, including comorbidities and pleural effusion. The extracted unstructured information is then mapped into structured feature vectors that can be used for model training, achieving homogeneous integration of multi-source data.

[0012] Furthermore, in step S2, cross-departmental dataset partitioning and precise screening of research subjects are performed. Training and test sets are divided based on differences in departmental risk levels. Suitable research subject samples are selected according to the HA-AKI diagnostic criteria and strict inclusion and exclusion rules. Specifically: S21: Cross-departmental differentiated dataset partitioning: Based on the differences in the disease risk level of hospitalized patients in different departments, the data of hospitalized patients in general wards (excluding ICU) are set as the model training set for model parameter fitting and algorithm optimization; the data of hospitalized patients in ICU intensive care units are set as an independent external test set, specifically for verifying the model's generalization ability. By separating and verifying the data from high-risk departments and general departments, the adaptability and predictive reliability of the model in patient groups with different risk levels are ensured. S22: Study Subject Inclusion Screening: Establish uniform inclusion criteria and include only patients who meet all of the following conditions in the study cohort: age ≥18 years; hospital stay >2 days; creatinine measurement ≥2 times during hospitalization, to ensure that the study sample has sufficient clinical data support and follow-up period; S23: Exclusion and Screening of Study Subjects: Strict exclusion criteria were established, and patients meeting any of the following conditions were excluded: serum creatinine ≥353.6 μmol / L within 48 hours of admission; missing key clinical diagnostic and treatment data affecting risk factor assessment; confirmed diagnosis of acute kidney injury (AKI) within 48 hours of admission; receiving renal replacement therapy within 48 hours of admission; or having a history of long-term maintenance dialysis treatment to avoid interfering with the model's early warning judgment of hospitalized AKI. S24: Standardized diagnosis of HA-AKI and calculation of baseline creatinine: The internationally recognized KDIGO standard is used as the basis for the diagnosis of HA-AKI, that is, the absolute value of serum creatinine increases by ≥26.5 μmol / L within 48 hours, or increases to more than 1.5 times the baseline creatinine value. Baseline creatinine was calculated using the MDRD normalization formula, assuming an estimated glomerular filtration rate (eGFR) of 75 mL / min / 1.73 m 2 The MDRD formula is as follows: Baseline creatinine = (75 / [186×(age)) -0.203 )×(0.742 if female)×(1.21 if black)]) -0.887 The variables in the formula are defined as follows: age is the patient's actual age; female is the gender identifier variable, which is 0.742 when the patient is female and 1.0 when the patient is male; black is the race identifier variable, which is 1.21 when the patient is Black and 1.0 when the patient is not Black. By clearly defining the variables in the formula, a unified quantitative calculation of baseline creatinine is achieved, ensuring the consistency and accuracy of the HA-AKI diagnostic criteria.

[0013] Furthermore, in step S3, the core variable selection for two-stage feature engineering is performed. First, an initial feature set is constructed through hierarchical processing of missing values. Then, a regularized regression algorithm is used to select key variables, achieving feature dimensionality reduction and overfitting suppression. Specifically: S31: Stratified handling of missing variable values: Establish a quantitative assessment mechanism for the missing variable rate, and perform statistical analysis on the missing value of the feature variables corresponding to the research object samples after screening in step S2; for low information value variables with a missing value greater than the preset missing value, they are directly removed to avoid interference of invalid data on model performance; for valid feature variables with a missing value less than or equal to the preset missing value, a hybrid imputation scheme combining forward imputation, backward imputation and global median imputation is adopted. By supplementing the data through the correlation between time series data and the reference of the overall distribution characteristics, the data integrity is restored to the maximum extent and the feature quality is guaranteed. S32: Initial feature dataset construction: Based on the standardized time-series data of the training set, a time-series feature alignment mechanism is adopted to construct an initial feature dataset corresponding to each time node, which is used to predict the occurrence status of acute kidney injury (AKI) within a preset time period in the future. This achieves accurate binding between the warning time window and the feature data, and provides targeted feature support for subsequent model training. S33: Selection of key variables for regularized regression: LASSO regression analysis is used as the regularized regression algorithm. Under the lambda.1se criterion, variables in the initial feature dataset are selected for regularization. The interference of redundant variables is suppressed by the regularization penalty mechanism, and only key clinical variables with non-zero regression coefficients are retained. This not only achieves efficient compression of feature dimensions, but also effectively suppresses the risk of overfitting during model training, and improves the generalization ability and prediction stability of the model.

[0014] Furthermore, in step S4, time-series dynamic feature construction and secondary feature optimization are performed. Multi-dimensional time-series features are constructed for key variables, and the key variables and time-series features are integrated to form an optimized feature dataset for early warnings adapted to a preset future time period. Specifically: S41: Construction of a Multi-Dimensional Temporal Feature System: Based on the key clinical variables screened by LASSO in step S3, and combined with the dynamic evolution of HA-AKI incidence risk, targeted temporal features are constructed according to variable type, specifically including: Laboratory indicator time series characteristics: Calculate the dynamic change rate, 3-day rolling average, historical maximum and historical minimum of key variables for each laboratory to accurately capture the fluctuation trend and cumulative effect of indicators over time; Drug use timing characteristics: Statistical analysis of the cumulative number of days of use for various key drugs to quantify the duration of drug exposure on the risk of HA-AKI; Other clinical event time-series features: Extract dynamic event features related to surgery, including postoperative days, to reflect changes in risk over time after clinical interventions; S42: Construction of secondary feature dataset fusion: Using feature fusion technology, the key clinical variables selected in step S3 are deeply integrated with the multi-dimensional time-series features constructed in step S41 to eliminate feature redundancy and enhance risk association information. For the training set and the test set respectively, a secondary feature dataset adapted to the HA-AKI early warning scenario in the future preset time period is reconstructed to provide more targeted and predictive feature support for subsequent model training and validation.

[0015] Furthermore, in step S5, multi-algorithm model construction, cross-validation, and interpretability enhancement are performed. An early warning model is built based on multiple machine learning algorithms. The model's generalization ability is evaluated through cross-departmental validation. The feature contribution is quantified using interpretability algorithms. Specifically: S51: Construction of a multi-algorithm early warning model library: Construct a multi-dimensional machine learning algorithm model library including Logistic Regression (LR), Random Forest (RF), Linear Support Vector Machine (LinearSVC), Neural Network (NNET), K Nearest Neighbors (KNN), Lightweight Gradient Boosting Machine (LightGBM), Extreme Gradient Boosting Machine (XGBoost), and Gradient Boosting Decision Tree (GBDT). Through parallel construction of multiple algorithms, a sufficient algorithmic foundation is provided for subsequent model optimization. S52: Model-oriented training and cross-validation optimization: Using the training set secondary feature dataset constructed in step S4 as input, model parameter training is carried out on patient data in general wards outside the ICU. At the same time, a 5-fold internal cross-validation technique is adopted. By randomly dividing the training set into 5 mutually exclusive subsets, 4 subsets are used as training samples and 1 subset is used as validation samples for iterative training and parameter adjustment, which effectively avoids model overfitting and improves model training stability and parameter optimization accuracy. S53: Cross-departmental model generalization ability assessment: Using the test set secondary feature dataset constructed in step S4 as input, external model validation is carried out on the independent test set of the ICU. A multi-dimensional performance evaluation system is used to comprehensively evaluate the model performance. Evaluation indicators include the area under the receiver operating characteristic curve (AUC-ROC), the area under the precision-recall curve (PR-AUC), the F1 score, the probability prediction error Brier score, and the DCA curve of decision curve analysis. Through cross-departmental data validation, the generalization ability of the model in high-risk patient groups is accurately judged. S54: Model Performance Optimization and Enhanced Clinical Interpretability: Based on the multi-dimensional performance evaluation results of step S53, the early warning model with the best overall performance is selected. For this optimal model, the SHAP value interpretability analysis algorithm is used to quantify the contribution direction and intensity of each clinical feature to the HA-AKI risk prediction results. The analysis scope comprehensively covers the training set, the test set, and multiple key clinical subgroups, including non-ICU / ICU patients, non-elderly / elderly patients, non-cardiology / cardiology patients, and non-sepsis / sepsis patients. This achieves a transparent presentation of the model's prediction logic and improves the credibility of clinical applications.

[0016] Furthermore, in step S6, clinical integration and automated hierarchical early warning of the model are performed. The optimal model is selected and integrated into the hospital information system to construct an automated early warning closed loop of timed data updates, feature processing, risk prediction, and clinical decision support. Specifically: S61: Optimal Model Clinical System Integration: The early warning model with the best overall performance obtained from step S5 is integrated into the hospital information system of medical institutions through a standardized API service interface. An independent HA-AKI risk prediction function module is built to achieve seamless connection between the model and the existing clinical data system, ensuring stable access to the risk prediction function and secure data interaction. S62: Deployment and Implementation of Automated Early Warning Process: Configure the early warning system to a customizable timed data acquisition mode; the system automatically captures the latest clinical data of patients before the corresponding time node through the hospital information system interface at regular intervals, performs standardized preprocessing and feature engineering processes on the acquired data that are completely consistent with the model training phase, and ensures the consistency of data format and feature dimensions; input the processed standardized feature data into the integrated optimal early warning model, and the model automatically calculates and outputs the accurate risk probability of the patient developing HA-AKI in the next early warning cycle; S63: Clinical Decision Support and Tiered Early Warning Implementation: Through a visual clinical decision support interface, the system simultaneously displays to medical staff the patient's HA-AKI risk probability, as well as key risk factors, including the direction and intensity of risk contribution, obtained based on SHAP value analysis. The system has built-in risk grading threshold rules, supporting the preset judgment thresholds for different risk levels according to clinical needs. When the patient's risk probability reaches the corresponding threshold, a tiered early warning is automatically triggered, providing medical staff with intuitive and quantitative decision-making basis, assisting them in timely assessing the patient's condition and adjusting the treatment plan, and achieving early intervention for HA-AKI.

[0017] A hospital-acquired acute kidney injury intelligent early warning system for implementing the intelligent early warning method for hospital-acquired acute kidney injury as described above, comprising: The multi-source data integration and preprocessing module is used to integrate and collect heterogeneous data from multiple sources and perform standardized preprocessing. It collects structured data, unstructured text data and other derived computational data from hospitalized patients, and constructs a standardized patient time-series data system through data cleaning, normalization and text structuring transformation. The cross-departmental dataset partitioning and screening module is used to partition cross-departmental datasets and accurately screen research subjects. It divides the training set and test set based on the differences in departmental risk levels and screens suitable research subject samples according to the HA-AKI diagnostic criteria and strict inclusion and exclusion criteria. The two-stage feature variable screening module is used to screen the core variables for two-stage feature engineering. First, it constructs an initial feature set by hierarchical processing of missing values, and then uses a regularized regression algorithm to screen key variables, thereby achieving feature dimensionality reduction and overfitting suppression. The temporal feature construction and optimization module is used to construct and optimize dynamic temporal features. It constructs multi-dimensional temporal features for key variables and integrates key variables and temporal features to form an optimized feature dataset for early warning that is adapted to a preset time period in the future. Multi-algorithm modeling, verification, and interpretation are used for multi-algorithm model construction, cross-validation, and interpretability enhancement. Early warning models are built based on multiple machine learning algorithms. The generalization ability of the model is evaluated through cross-departmental verification, and the contribution of features is quantified by combining interpretability algorithms. The model clinical integration early warning is used for model clinical integration and automated hierarchical early warning. It selects the optimal model to integrate into the hospital information system and builds an automated early warning closed loop of timed data update, feature processing, risk prediction and clinical decision support.

[0018] A computer device, characterized in that it includes a memory and one or more processors, wherein the memory stores computer code, and when the computer code is executed by the one or more processors, causes the one or more processors to perform the method as described above.

[0019] A computer-readable storage medium, characterized in that the computer-readable storage medium stores computer code, which, when executed, is performed as described above.

[0020] Compared with the prior art, the present invention has at least one of the following beneficial effects: (1) Prominent early warning and significantly earlier intervention window. This invention breaks through the limitations of traditional diagnostic criteria that rely on elevated serum creatinine or decreased urine output by deeply capturing the data evolution pattern before the occurrence of HA-AKI. It can achieve a forward warning in advance within a preset time period (such as 24 hours), which is significantly earlier than the warning time of traditional creatinine diagnostic criteria. This provides clinicians with sufficient time for prevention and early intervention, and effectively avoids the problem of missed treatment opportunities caused by delayed warning.

[0021] (2) Excellent prediction accuracy and superior model performance. This invention innovatively integrates multi-source heterogeneous data such as text data, laboratory indicators, medication prescriptions, and imaging reports, and constructs time-series dynamic features for key variables to fully explore potential risk correlation information in the data. The LightGBM model used has been verified through multi-algorithm comparison and shows that it outperforms existing similar early warning models in terms of accuracy and specificity. It can more accurately quantify the risk of HA-AKI occurrence and solves the problem of low prediction accuracy caused by the single evaluation dimension and insufficient feature utilization of traditional technologies.

[0022] (3) Supports dynamic and personalized assessment to achieve continuous risk monitoring. The early warning model of this invention supports real-time data input and can generate a unique dynamic risk trajectory for each patient based on dynamic information such as changes in the patient's condition, fluctuations in laboratory indicators, and adjustments in medication during hospitalization. This breaks the limitation of traditional models that rely solely on static data within a certain period of hospitalization and enables personalized and continuous risk monitoring of patients throughout their hospitalization, adapting to the risk assessment needs of different patients at different stages of diagnosis and treatment.

[0023] (3) Excellent clinical interpretability, enhancing the credibility and practicality of clinical applications. This invention not only outputs the probability of HA-AKI occurrence but also quantifies the specific contribution direction and intensity of each clinical feature (such as specific drug use, abnormal trends in laboratory indicators, etc.) to the prediction results for individual patients through interpretable artificial intelligence methods such as SHAP value analysis. This design makes the model's decision-making basis completely transparent to clinicians, helping them to transform the prediction results into concrete clinical insights. This not only enhances doctors' trust in the early warning information but also directly guides the formulation of precise clinical intervention measures, effectively improving the clinical application value of the technology. Attached Figure Description

[0024] Figure 1 This is an overall flowchart of the intelligent early warning method for hospitalized acute kidney injury of the present invention; Figure 2 This is a flowchart illustrating the specific implementation of the HA-AKI early warning model of the present invention. Figure 3 The flowchart shows the exclusion process for research populations that meet the inclusion and exclusion criteria of this invention. Figure 4 This is a schematic diagram of the variables selected for LASSO regression analysis in this invention; Figure 5 shows the ROC curve, PR curve, calibration curve and DCA curve of various machine learning algorithms of the present invention on the test set, wherein Figure 5(a) is a schematic diagram of the ROC curve, Figure 5(b) is a schematic diagram of the PR curve, Figure 5(c) is a schematic diagram of the calibration curve and Figure 5(d) is a schematic diagram of the DCA curve. Figure 6(a) is a schematic diagram of the SHAP value analysis results in the whole population. Figure 6(b) is a heatmap of the feature importance of the SHAP value analysis in the whole population and each subgroup. Figure 6(c) is a schematic diagram of the distribution of the ranking of each feature in the whole population and different subgroups. Figure 7 This is a structural diagram of the intelligent early warning system for hospitalized acute kidney injury of the present invention. Detailed Implementation

[0025] To make the objectives, technical solutions, and advantages of the embodiments of this application clearer, the technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.

[0026] Those skilled in the art will understand that, unless specifically stated otherwise, the singular forms “a,” “an,” “the,” and “the” used herein may also include the plural forms. It should be further understood that the term “comprising” as used in this specification means the presence of the stated features, integers, steps, operations, elements, and / or components, but does not exclude the presence or addition of one or more other features, integers, steps, operations, elements, components, and / or groups thereof.

[0027] First Embodiment like Figure 1 As shown in the figure, this embodiment provides an intelligent early warning method for hospitalized acute kidney injury, including the following steps: S1: Integrate and collect multi-source heterogeneous data and perform standardized preprocessing. Collect structured data, unstructured text data and other derived computational data of hospitalized patients. Through data cleaning, normalization and text structuring transformation, construct a standardized patient time-series data system.

[0028] In this embodiment, step S1 specifically includes: S11: Define the data source and time range: The data comes from the hospital information platform of the medical institution and is included in all inpatient data of the institution, including cardiology, cardiac surgery, respiratory medicine, emergency medicine and critical care medicine, within the preset time period. This ensures that the data covers patient groups with different disease severity and department types, and guarantees the representativeness and diversity of the data, so as to provide comprehensive sample support for subsequent model training and validation. S12: Conduct multi-dimensional data collection: The types of data collected include structured data, unstructured text data, and other derived computational data; Structured data covers: Demographic information: Record ID, gender, age, height, weight, admission time, discharge time, length of hospital stay, admission department, admission diagnosis, discharge department, discharge diagnosis, and discharge status; Laboratory indicators: complete blood count, inflammatory markers, renal function indicators, liver function indicators, cardiac function indicators, coagulation function indicators, and electrolyte indicators; Vital signs data: respiratory rate, systolic blood pressure, diastolic blood pressure, heart rate, body temperature, and blood oxygen saturation; Medical records of medication and treatment: including records of the use of anticoagulants, nonsteroidal anti-inflammatory drugs, diuretics, vasoactive drugs, nephrotoxic antibiotics, and contrast agents, as well as clinical treatment records of mechanical ventilation and renal replacement therapy; Unstructured text data includes present medical history texts and imaging reports containing descriptions of pleural effusion; Other derived calculation data include the ratio of blood urea nitrogen to blood creatinine, the ratio of neutrophils to lymphocytes, and the sequential organ failure SOFA score (excluding the nervous system module). The sequential organ failure SOFA score must be greater than or equal to a preset score (e.g., 2 points). S13: Implement data standardization preprocessing: Structured data preprocessing: Data cleaning algorithms are used to identify and process missing values ​​and detect and remove outliers in structured data. Then, normalization is used to unify the data scale and eliminate the dimensional differences between different indicators. Finally, time window alignment technology is used to align the structured data collected at different time points according to the preset time granularity to form standardized patient time series data. Unstructured text data structuring transformation: Using natural language processing technology, an entity recognition and relation extraction framework is built based on a pre-trained biomedical language model. Semantic parsing is performed on current medical history texts and imaging reports to accurately extract key clinical information, including comorbidities and pleural effusion. The extracted unstructured information is then mapped into structured feature vectors that can be used for model training, achieving homogeneous integration of multi-source data.

[0029] Step S1 serves as the data foundation for the entire HA-AKI intelligent early warning method. Its core design goal is to overcome the bottlenecks of traditional early warning technologies, such as single data sources, heterogeneous formats, and inconsistent data quality. Through a systematic "collection-integration-standardization" process, it provides high-quality, highly available, multi-dimensional time-series data support for subsequent model training and early warning analysis. The specific design logic and technical advantages are explained below: From the perspective of defining the data source and scope (S11), this solution specifically selects patient data covering multiple core clinical departments, including cardiology, cardiac surgery, respiratory medicine, emergency medicine, and critical care medicine. The core consideration is that these departments exhibit significant differences in the complexity of patients' conditions and the intensity of treatment interventions (e.g., critical care patients are more critically ill and have a higher risk of HA-AKI, while general internal medicine patients have relatively stable conditions). This allows for a gradient sample distribution of "low-risk-medium-risk-high-risk," effectively avoiding model bias caused by data from a single department. The preset time period ensures both the accumulation of a sufficient sample size (meeting the data volume requirements of machine learning models) and the capture of clinical characteristic differences across different seasons and treatment cycles, further enhancing the representativeness of the data. This design ensures that the sample set upon which the subsequent model training relies comprehensively covers high-incidence scenarios of HA-AKI and populations at different risk levels, laying the foundation for the model's generalization ability.

[0030] In the design of multi-dimensional data acquisition (S12), this solution fully addresses the shortcomings of existing technologies that suffer from "single evaluation dimensions": structured data, as the basic record of clinical diagnosis and treatment, covers demographic information, laboratory indicators, vital signs, and medication records, directly reflecting the patient's baseline status, physiological and pathological changes, and treatment interventions. Demographic information is used to differentiate individual baseline differences (such as the impact of age and gender on renal function), laboratory indicators (such as renal function and inflammatory markers) are directly related to the physiological and pathological mechanisms of kidney injury, and medication records focus on exposure to nephrotoxic and therapeutic drugs, providing a basis for analyzing treatment-related risks; unstructured text data (currently...) Medical history texts and imaging reports contain a wealth of key clinical information that is not recorded in a structured manner (such as latent comorbidities, pleural effusions, and other local lesions). This information often cannot be captured by conventional structured indicators, but may indirectly affect the risk of HA-AKI. Targeted collection can fill the information gaps of traditional techniques. On the other hand, derived computational data (such as the ratio of blood urea nitrogen to serum creatinine and SOFA score ≥2) is a deep mining of the original data. Through quantitative indicator combinations or scoring systems, it can more accurately reflect the organ function status (such as the SOFA score reflecting the risk of multiple organ failure) and the intensity of inflammatory response (neutrophil to lymphocyte ratio), further enriching the dimensions of risk assessment.

[0031] Data standardization preprocessing (S13) is a key step in achieving "homogeneous integration" of multi-source heterogeneous data. For structured data, missing values ​​and outliers are handled using algorithmic identification and targeted repair strategies, which avoids interference from invalid data to the model while maximizing the retention of valid samples (e.g., a hybrid imputation scheme that considers both the correlation between time series data and the overall distribution characteristics). The core purpose of normalization is to eliminate the dimensional differences between different indicators (e.g., height in cm, blood pressure in mmHg), avoiding weight bias caused by differences in numerical scales during model training and ensuring the fairness of the influence of each indicator on the model. Time window alignment technology is used to transform discretely collected structured data (e.g., daily laboratory tests, real-time monitoring of vital signs) into time series data with a unified time granularity, adapting to the subsequent "dynamic early warning" requirements and enabling the model to capture the changing trends of indicators over time. For unstructured text data, a pre-trained biomedical language model is used for entity recognition and relation extraction, designed based on the professional characteristics of clinical texts. The biomedical language model, trained on a large amount of medical text, can accurately identify professional terms such as "comorbidities" and "pleural effusion" and their clinical relationships, avoiding semantic misunderstandings in medical scenarios using general natural language processing technologies. By converting text information into structured feature vectors, the model achieves format unification between text data and structured data, enabling multi-source data to be directly processed by machine learning models. This truly achieves the design goal of "multi-source heterogeneous data fusion," laying the foundation for mining risk associations across data types in subsequent feature engineering.

[0032] In summary, step S1, through its comprehensive design of "precise sampling - comprehensive collection - standardized integration," not only solves the problems of limited data sources, heterogeneous formats, and insufficient quality in traditional technologies, but also constructs a standardized data system that accurately reflects the patient's individual baseline, dynamic physiological and pathological changes, the impact of treatment interventions, and potential risk factors through in-depth integration and temporal processing of multi-dimensional data. This provides solid data support for subsequent feature screening, temporal feature construction, and model training, and is the core prerequisite for ensuring the foresight and accuracy of the entire early warning model.

[0033] S2: Perform cross-departmental dataset partitioning and precise screening of research subjects. Based on the differences in departmental risk levels, divide the training set and test set, and screen suitable research subject samples according to the HA-AKI diagnostic criteria and strict inclusion and exclusion criteria.

[0034] In this embodiment, step S2 specifically includes: S21: Cross-departmental differentiated dataset partitioning: Based on the differences in the disease risk level of hospitalized patients in different departments, the data of hospitalized patients in general wards (excluding ICU) are set as the model training set for model parameter fitting and algorithm optimization; the data of hospitalized patients in ICU intensive care units are set as an independent external test set, specifically for verifying the model's generalization ability. By separating and verifying the data from high-risk departments and general departments, the adaptability and predictive reliability of the model in patient groups with different risk levels are ensured. S22: Study Subject Inclusion Screening: Establish uniform inclusion criteria and include only patients who meet all of the following conditions in the study cohort: age ≥18 years; hospital stay >2 days; creatinine measurement ≥2 times during hospitalization, to ensure that the study sample has sufficient clinical data support and follow-up period; S23: Exclusion and Screening of Study Subjects: Strict exclusion criteria were established, and patients meeting any of the following conditions were excluded: serum creatinine ≥353.6 μmol / L within 48 hours of admission; missing key clinical diagnostic and treatment data affecting risk factor assessment; confirmed diagnosis of acute kidney injury (AKI) within 48 hours of admission; receiving renal replacement therapy within 48 hours of admission; or having a history of long-term maintenance dialysis treatment to avoid interfering with the model's early warning judgment of hospitalized AKI. S24: Standardized diagnosis of HA-AKI and calculation of baseline creatinine: The internationally recognized KDIGO standard is used as the basis for the diagnosis of HA-AKI, that is, the absolute value of serum creatinine increases by ≥26.5 μmol / L within 48 hours, or increases to more than 1.5 times the baseline creatinine value. Baseline creatinine was calculated using the MDRD normalization formula, assuming an estimated glomerular filtration rate (eGFR) of 75 mL / min / 1.73 m 2 The MDRD formula is as follows: Baseline creatinine = (75 / [186×(age)) -0.203 )×(0.742 if female)×(1.21 if black)]) -0.887 The variables in the formula are defined as follows: age is the patient's actual age; female is the gender identifier variable, which is 0.742 when the patient is female and 1.0 when the patient is male; black is the race identifier variable, which is 1.21 when the patient is Black and 1.0 when the patient is not Black. By clearly defining the variables in the formula, a unified quantitative calculation of baseline creatinine is achieved, ensuring the consistency and accuracy of the HA-AKI diagnostic criteria.

[0035] Step S2, as the core link in ensuring data quality control and sample validity for the early warning model, aims to overcome the limitations of traditional early warning technologies, such as single dataset partitioning, ambiguous sample selection, and inconsistent diagnostic standards. Through a complete process design of "scientific partitioning - precise selection - standardized diagnosis," it provides unbiased, effective, and homogeneous research samples for subsequent model training and validation, ensuring that the model can learn the real risk patterns of HA-AKI occurrence. The specific design logic and technical advantages are explained below: From the perspective of cross-departmental differentiated dataset partitioning (S21), the core innovation of this solution lies in constructing a dual dataset for "training-validation" based on "disease level differences in departments." Patients in general wards (non-ICU) have relatively stable conditions, and the risk of HA-AKI occurs in a gradient distribution. The sample size is sufficient, and the clinical scenario is more universal. Using this as the training set allows the model to fully learn the basic risk patterns of HA-AKI occurrence, achieving robust fitting of model parameters and algorithm optimization. In contrast, patients in the intensive care unit (ICU) have critical conditions, complex underlying diseases, and intensive treatment interventions, resulting in a significantly higher risk of HA-AKI than those in general wards. As a high-risk group, using this as an independent external test set can specifically validate the model's generalization ability in high-risk scenarios. This separate design of "general department training + high-risk department validation" specifically addresses the shortcomings of existing technologies that "only validate in general wards and lack promotion to critical care scenarios." By covering patient groups with different risk levels, it ensures that the model can adapt to diverse clinical scenarios in practical applications, enhancing the practical value of the technology.

[0036] In the design of the study subject inclusion screening (S22), each inclusion criterion revolves around "sample validity" and "data sufficiency": the age requirement of ≥18 years is because adult renal function is mature and physiological state is relatively stable, which can exclude confounding factors other than the incomplete development of renal function in minors or the special physiological decline of the elderly (there is no upper age limit, which can be covered by subgroup analysis later), ensuring the homogeneity of the sample; the requirement of hospitalization time >2 days is to ensure that patients have a sufficient hospitalization period to accumulate enough time-series data (such as dynamic changes in laboratory indicators, medication records, etc.) to meet the needs of subsequent time-series feature construction and 24-hour early warning; the standard of ≥2 creatinine measurements is directly related to the diagnostic logic of HA-AKI (which needs to be determined by changes in creatinine), and also provides core data support for the model to capture dynamic changes in renal function, avoiding invalid samples due to the lack of key indicators. The combination of these three criteria ensures that the included study samples have both complete clinical data and meet the technical requirements of HA-AKI early warning.

[0037] The core design of the study subject exclusion screening (S23) is to "purify the study cohort" and remove confounding factors that may interfere with the model's learning of the true risk of HA-AKI. Patients with "serum creatinine ≥353.6 μmol / L within 48 hours of admission", "diagnosed with AKI within 48 hours of admission", "receiving renal replacement therapy within 48 hours of admission", or "receiving long-term maintenance dialysis" are excluded because the kidney damage in these patients may have existed before admission (i.e., community-acquired kidney damage) or their kidney function may be in a special pathological state, rather than "hospital-acquired" acute kidney injury. Excluding these patients ensures that the study cohort focuses on the core objective of "new-onset HA-AKI during hospitalization" and avoids the model learning risk characteristics of non-target scenarios. Patients with "missing key clinical data" are excluded to ensure the integrity of the features of each sample and to avoid bias in model training due to missing key risk factors (such as medication records and laboratory indicators). This ensures that the model learns the association between complete risk factors and the occurrence of HA-AKI.

[0038] The core design of the standardized diagnosis and baseline creatinine calculation for HA-AKI (Solution S24) is to address the issue of "diagnostic consistency" by providing clear label definitions for model training. Adopting the internationally recognized KDIGO standard as the diagnostic criterion for HA-AKI ensures that the determination of HA-AKI conforms to clinical consensus, avoids diagnostic bias caused by custom-defined standards, and improves the credibility and clinical acceptance of research results. The calculation of baseline creatinine uses the MDRD standardized formula and explicitly assumes eGFR = 75 mL / min / 1.73 m 2 The purpose is to standardize the baseline creatinine levels across different patients. Because individual baseline creatinine levels naturally vary (e.g., influenced by age, sex, and race), directly using creatinine at admission as the baseline can lead to inaccurate HA-AKI assessments. Standardizing the calculation using the MDRD formula incorporating variables such as age, sex, and race eliminates the interference of individual baseline differences, achieving a unified quantification of baseline creatinine. The clear definition of each variable in the formula (age being actual age, female / black being sex / race identifiers and corresponding coefficients) further ensures the accuracy and consistency of baseline creatinine calculations. This allows the diagnostic criterion of "serum creatinine increasing by ≥26.5 μmol / L or reaching 1.5 times the baseline within 48 hours" to be strictly adhered to, providing the model with clear and accurate positive and negative sample labels, which is fundamental to ensuring the model's predictive accuracy.

[0039] In summary, step S2, through a comprehensive design that "diversified dataset partitioning ensures model generalization ability, precise screening ensures sample validity, and standardized diagnosis ensures label consistency," systematically solves the problems of limited dataset scenarios, mixed samples, and inconsistent diagnostic standards in existing technologies. It provides high-quality, unbiased, and homogeneous research samples for subsequent feature engineering and model training, which is a key prerequisite for ensuring that the early warning model can learn the real risk patterns of HA-AKI and achieve accurate early warning.

[0040] S3: Select core variables for two-stage feature engineering. First, construct an initial feature set by hierarchical processing of missing values, and then use a regularized regression algorithm to select key variables, thereby achieving feature dimensionality reduction and overfitting suppression.

[0041] In this embodiment, step S3 specifically includes: S31: Stratified handling of missing variable values: Establish a quantitative assessment mechanism for the missing variable rate, and perform statistical analysis on the missing value of the feature variables corresponding to the research object samples after screening in step S2; for low information value variables with a missing rate > preset missing rate value (e.g., >50%), they are directly removed to avoid interference of invalid data on model performance; for valid feature variables with a missing rate ≤ preset missing rate value (e.g., ≤50%), a hybrid imputation scheme combining forward imputation, backward imputation, and global median imputation is adopted. By supplementing through the correlation between time series data and the reference of overall distribution characteristics, the integrity of the data is restored to the maximum extent and the quality of features is guaranteed. S32: Initial feature dataset construction: Based on the standardized time-series data of the training set, a time-series feature alignment mechanism is adopted to construct an initial feature dataset corresponding to each time node, which is used to predict the occurrence status of acute kidney injury (AKI) within a future preset time period (such as the next 24 hours). This achieves accurate binding between the warning time window and the feature data, and provides targeted feature support for subsequent model training. S33: Selection of key variables for regularized regression: LASSO regression analysis is used as the regularized regression algorithm. Under the lambda.1se criterion, variables in the initial feature dataset are selected for regularization. The interference of redundant variables is suppressed by the regularization penalty mechanism, and only key clinical variables with non-zero regression coefficients are retained. This not only achieves efficient compression of feature dimensions, but also effectively suppresses the risk of overfitting during model training, and improves the generalization ability and prediction stability of the model.

[0042] Step S3, as the core link between data preprocessing and secondary feature engineering, aims to overcome the limitations of traditional techniques, such as inconsistent feature quality, redundant dimensions, and weak generalization ability. Through a two-stage feature engineering design—"data quality purification - targeted feature construction - key variable screening"—it provides high-quality, low-redundancy, and strongly correlated core features for subsequent time-series feature construction and model training. This ensures the model can accurately capture key risk signals associated with HA-AKI. The specific design logic and technical advantages are explained below: From the perspective of stratified handling of missing variable values ​​(S31), the core design principle of this scheme is "differentiated processing, balancing data completeness and effectiveness." Missing values ​​are common in clinical data, but variables with different missing rates have significantly different impacts on the model: for variables with a missing rate > a preset threshold (e.g., 50%), they contain very little effective information, and forcibly filling them in would introduce a lot of noise, interfering with the model's learning of the true risk patterns. Therefore, they are directly removed to ensure feature quality from the source. For effective variables with a missing rate ≤ the preset threshold, a hybrid scheme of "forward filling + backward filling + global median filling" is adopted. This scheme is specifically designed based on the characteristics of time series data. Forward filling and backward filling can utilize the correlation between time series data of the same patient (e.g., if a kidney function indicator is missing at a certain time point, the test results at adjacent time points can be referenced) to restore the dynamic trend of the data to the greatest extent. Global median filling serves as a supplement. When there is no effective reference before and after the time series data, it is filled based on the indicator distribution of the entire patient population to avoid the bias caused by a single filling method. This hierarchical processing strategy not only solves the problem of "simple and crude handling of missing values" in traditional technologies, but also preserves the true distribution characteristics and temporal correlation of the data while ensuring data integrity through a hybrid filling scheme, laying a reliable foundation for subsequent feature construction.

[0043] In the design of the initial feature dataset construction (S32), the core logic is "temporal alignment + targeted binding," directly addressing the shortcomings of existing technologies that are "mainly static data with features disconnected from warning targets." This solution, based on standardized time-series data from the training set, constructs corresponding feature sets for each time node and explicitly binds them to the goal of "predicting the occurrence status of AKI within a preset time period (e.g., 24 hours)." Essentially, it constructs a three-dimensional association system of "feature-time-warning target": On the one hand, temporal alignment ensures that each feature corresponds to a specific time node, adapting to the dynamic changes in HA-AKI risk over time and avoiding the limitations of traditional static features in reflecting disease progression; on the other hand, targeted binding of the warning window (e.g., the next 24 hours) allows features to accurately correspond to the prediction target, ensuring that the model learns the correlation between "features at the current time point and the risk of AKI occurrence in the short term," rather than redundant information unrelated to the time dimension. This provides a clear time benchmark for the subsequent construction of time-series features (e.g., 3-day rolling average, rate of change), making the prediction of warning targets more targeted.

[0044] The core design of the regularized regression key variable screening (S33) is "dimensionality reduction and redundancy removal, balancing model performance and generalization ability." After data preprocessing and initial feature construction, the feature dimension is often high (e.g., 89 initial features in this embodiment). High-dimensional data easily leads to model overfitting (i.e., the model overlearns training set noise and performs poorly on new data), and redundant variables can mask the role of key risk factors. This solution chooses LASSO regression as the regularization algorithm. Its core advantage lies in its ability to compress the regression coefficients of unimportant variables to 0 through the regularization penalty mechanism, thereby achieving the effect of "automatically screening key variables," perfectly adapting to the feature screening needs of high-dimensional clinical data. The use of the lambda.1se criterion to determine the penalty strength is the optimal choice after rigorous verification. This criterion achieves a balance between "lowest model complexity" and "optimal predictive performance," avoiding feature redundancy and overfitting caused by too light a penalty, and preventing the loss of key information and underfitting caused by too heavy a penalty. Through this design, 89 initial features in this embodiment were filtered into 24 key variables, which not only achieved efficient compression of feature dimensions, but also retained core clinical variables that are strongly correlated with HA-AKI (such as renal function indicators, key drug use records, etc.), effectively improving the model's generalization ability and predictive stability. At the same time, it focused on core variables for the construction of subsequent time-series features, avoiding redundant calculations of invalid features.

[0045] In summary, step S3, through its two-stage design of "hierarchical processing of missing values ​​to purify data quality, time-series targeted construction of clear feature associations, and regularization screening to simplify core variables," systematically solves the problems of "low feature quality, redundant dimensions, and weak association with early warning targets" in existing technologies. It provides a high-quality foundation of core variables for subsequent time-series dynamic feature construction, ensuring that the subsequent model can focus on key risk factors and accurately learn the inherent laws of HA-AKI occurrence. This is a key prerequisite for ensuring the model's foresight and accuracy.

[0046] S4: Construct dynamic time-series features and optimize secondary features. Build multi-dimensional time-series features for key variables and integrate key variables and time-series features to form an optimized feature dataset that adapts to the future preset time period for early warning.

[0047] In this embodiment, step S4 specifically includes: S41: Construction of a Multi-Dimensional Temporal Feature System: Based on the key clinical variables screened by LASSO in step S3, and combined with the dynamic evolution of HA-AKI incidence risk, targeted temporal features are constructed according to variable type, specifically including: Laboratory indicator time series characteristics: Calculate the dynamic change rate, 3-day rolling average, historical maximum and historical minimum of key variables for each laboratory to accurately capture the fluctuation trend and cumulative effect of indicators over time; Drug use timing characteristics: Statistical analysis of the cumulative number of days of use for various key drugs to quantify the duration of drug exposure on the risk of HA-AKI; Other clinical event time-series features: Extract dynamic event features related to surgery, including postoperative days, to reflect changes in risk over time after clinical interventions; S42: Construction of secondary feature dataset fusion: Using feature fusion technology, the key clinical variables selected in step S3 are deeply integrated with the multi-dimensional time-series features constructed in step S41 to eliminate feature redundancy and enhance risk association information. For the training set and the test set respectively, a secondary feature dataset adapted to the HA-AKI early warning scenario in the future preset time period is reconstructed to provide more targeted and predictive feature support for subsequent model training and validation.

[0048] Step S4, as a deepening and optimization stage of feature engineering, aims to overcome the limitations of traditional techniques that rely primarily on static data and lack sufficient temporal feature mining. Through a design that combines targeted temporal feature construction with deep feature fusion, the static key variables selected in S3 are transformed into high-value features that reflect the dynamic evolution of HA-AKI risk. This provides core support for the subsequent model to accurately capture early risk signals. The specific design logic and technical advantages are explained below: From the perspective of constructing a multi-dimensional temporal feature system (S41), the core innovation of this scheme lies in "customizing the construction of temporal features according to variable type" to accurately match the dynamic risk pattern of HA-AKI onset. The occurrence of HA-AKI is not determined by the state of indicators at a single point in time, but is closely related to the trend of indicator changes over time and the cumulative effect. Taking laboratory indicators as an example, a normal renal function indicator at a single time point does not mean there is no risk. However, if the dynamic change rate continues to rise and the 3-day rolling average exceeds the critical range, it is often an early signal of kidney damage. Therefore, designing time-series features such as "dynamic change rate, 3-day rolling average, and historical maximum / minimum value" can transform the "static value" of the indicator into a "dynamic trend" and capture potential risks that traditional static features cannot cover. For drug use variables, the toxic effects of drugs on the kidneys have a cumulative effect, and the risk difference between short-term use and long-term exposure is significant. Through the time-series feature of "cumulative use days", the association between drug exposure duration and HA-AKI risk can be accurately quantified, avoiding the shortcomings of traditional techniques that only focus on "whether medication is used" while ignoring "duration of medication". For clinical events such as surgery, the risk of kidney damage varies significantly at different time points after surgery (e.g., 3-7 days after surgery is the high-incidence period of AKI). Extracting dynamic event features such as "postoperative days" can reflect the time-dimensional risk changes after clinical intervention measures, allowing the features to accurately correspond with the clinical diagnosis and treatment process. This classification-customized approach to constructing temporal features ensures that the temporal information of each key variable can be fully mined and is highly consistent with the pathogenesis of HA-AKI.

[0049] The core logic behind the design of the secondary feature dataset fusion construction (S42) is "integrating complementary information and strengthening risk correlation." The key clinical variables selected in S3 (such as eGFR and blood urea nitrogen) are the "static foundation" reflecting the patient's basic physiological state and core risk, while the time-series features constructed in S41 are the "dynamic supplement" reflecting the dynamic changes in risk. The two are highly complementary; without static key variables, time-series features would lose their core anchor; without time-series features, static variables cannot reflect the evolution trend of risk. This solution uses feature fusion technology to deeply integrate the two. On the one hand, it eliminates redundant information through redundancy removal (e.g., avoiding information overlap between static values ​​and time-series trend values ​​of the same indicator); on the other hand, it strengthens feature correlation, allowing the basic information of static variables and the dynamic trend of time-series features to form a synergistic effect, so that each feature can accurately point to the core objective of "the risk of HA-AKI occurring in a future preset time period (e.g., 24 hours)." Meanwhile, secondary feature datasets were constructed for the training and test sets respectively, ensuring the consistency of data distribution and avoiding deviations in model training and validation due to differences in feature processing. This provided high-quality and highly targeted feature inputs for subsequent multi-algorithm model training and performance evaluation.

[0050] In summary, step S4, through the design of "classifying and customizing temporal features to mine dynamic risks and integrating complementary features to enhance correlation effectiveness," systematically solves the problems of "insufficient feature utilization and inability to capture temporal patterns" in existing technologies. It transforms the originally isolated static key variables into a multi-dimensional feature system that can comprehensively reflect "basic state - dynamic changes - cumulative effects," allowing the features to be deeply adapted to the pathogenesis and clinical diagnosis and treatment process of HA-AKI, laying the core feature foundation for the subsequent model to achieve "early and accurate" early warning.

[0051] S5: Construct multi-algorithm models, perform cross-validation and interpretability enhancement, build early warning models based on multiple machine learning algorithms, evaluate the generalization ability of the models through cross-departmental validation, and quantify the feature contribution by combining interpretability algorithms.

[0052] In this embodiment, step S5 specifically includes: S51: Construction of a multi-algorithm early warning model library: Construct a multi-dimensional machine learning algorithm model library including Logistic Regression (LR), Random Forest (RF), Linear Support Vector Machine (LinearSVC), Neural Network (NNET), K Nearest Neighbors (KNN), Lightweight Gradient Boosting Machine (LightGBM), Extreme Gradient Boosting Machine (XGBoost), and Gradient Boosting Decision Tree (GBDT). Through parallel construction of multiple algorithms, a sufficient algorithmic foundation is provided for subsequent model optimization. S52: Model-oriented training and cross-validation optimization: Using the training set secondary feature dataset constructed in step S4 as input, model parameter training is carried out on patient data in general wards outside the ICU. At the same time, a 5-fold internal cross-validation technique is adopted. By randomly dividing the training set into 5 mutually exclusive subsets, 4 subsets are used as training samples and 1 subset is used as validation samples for iterative training and parameter adjustment, which effectively avoids model overfitting and improves model training stability and parameter optimization accuracy. S53: Cross-departmental model generalization ability assessment: Using the test set secondary feature dataset constructed in step S4 as input, external model validation is carried out on the independent test set of the ICU. A multi-dimensional performance evaluation system is used to comprehensively evaluate the model performance. Evaluation indicators include the area under the receiver operating characteristic curve (AUC-ROC), the area under the precision-recall curve (PR-AUC), the F1 score, the probability prediction error Brier score, and the DCA curve of decision curve analysis. Through cross-departmental data validation, the generalization ability of the model in high-risk patient groups is accurately judged. S54: Model Performance Optimization and Enhanced Clinical Interpretability: Based on the multi-dimensional performance evaluation results of step S53, the early warning model with the best overall performance is selected. For this optimal model, the SHAP value interpretability analysis algorithm is used to quantify the contribution direction and intensity of each clinical feature to the HA-AKI risk prediction results. The analysis scope comprehensively covers the training set, the test set, and multiple key clinical subgroups, including non-ICU / ICU patients, non-elderly / elderly patients, non-cardiology / cardiology patients, and non-sepsis / sepsis patients. This achieves a transparent presentation of the model's prediction logic and improves the credibility of clinical applications.

[0053] Step S5, as the core modeling and optimization stage of the entire early warning method, aims to overcome the shortcomings of traditional technologies such as "single algorithm, limited validation scenarios, and 'black box' model". Through a full-process design of "multiple algorithm selection - precise training and validation - cross-scenario generalization evaluation - interpretability enhancement", a HA-AKI early warning model with high predictive performance, strong generalization ability, and clinical credibility is constructed. The specific design logic and technical advantages are explained below: From the perspective of constructing a multi-algorithm early warning model library (S51), the core consideration of this approach is "adapting to the complexity of clinical data and ensuring optimal performance coverage." The risk of HA-AKI is influenced by the interaction of multiple dimensions of factors, and different machine learning algorithms have significantly different characteristics: Logistic Regression (LR) has a simple structure and strong interpretability but is difficult to capture nonlinear relationships; Random Forest (RF) and gradient boosting algorithms (LightGBM, XGBoost, GBDT) are good at handling high-dimensional data, capturing feature interactions and nonlinear associations, and are suitable for complex clinical multi-factor association scenarios; Neural Networks (NNET) have strong feature fitting capabilities but are prone to overfitting; K Nearest Neighbors (KNN) depends on sample distribution and is sensitive to high-dimensional data. By constructing a model library covering eight mainstream algorithms, rather than using a single algorithm for modeling, the advantages of different algorithms can be fully utilized. Through subsequent performance comparisons, the model best suited for HA-AKI risk prediction scenarios can be selected. In particular, gradient boosting algorithms such as LightGBM are highly compatible with the characteristics of clinical data in terms of their ability to process time-series features and high-dimensional data, laying the foundation for the superior performance of subsequent models and avoiding the limitations in prediction accuracy caused by the "insufficient adaptability" of traditional single algorithms.

[0054] The core of the model-oriented training and cross-validation optimization (S52) design is "precise training based on data characteristics to avoid overfitting risks." Data from general wards (non-ICU) patients was chosen as the training set because this group has a sufficient sample size (14,308 cases in this example) and a uniform distribution of disease progression, allowing the model to fully learn the underlying risk patterns of HA-AKI and achieve robust parameter fitting. The 5-fold internal cross-validation technique is designed to address the characteristics of clinical data, which are characterized by "high noise and complex feature associations." By randomly dividing the training set into five mutually exclusive subsets and iteratively executing "4 subset training + 1 subset validation," the model can effectively avoid overlearning random noise in the training set. Simultaneously, through multiple rounds of parameter adjustment, the model's adaptability to different data distributions is optimized, ensuring the stability and parameter accuracy of the model training. This combination of "oriented data training + cross-validation optimization" solves the overfitting problem caused by "single training data and insufficient parameter optimization" in traditional models, providing a fundamental guarantee for the model's subsequent generalization ability.

[0055] The core design of the cross-departmental model generalization ability assessment (S53) is to "break through scenario limitations and verify the suitability for high-risk populations." Existing technologies generally only validate models in general ward data, resulting in insufficient generalization ability in high-risk scenarios such as the ICU. ICU patients have critical conditions, complex underlying diseases, and intensive treatment interventions, leading to a significantly higher risk of HA-AKI compared to general ward patients, making them a key population for clinical early warning. This approach selects ICU patient data as an independent external test set to specifically verify the model's predictive performance in high-risk populations, accurately determining whether the model has broken through the "general ward scenario dependence" and truly adapted to diverse clinical risk scenarios. Meanwhile, employing a "multi-dimensional performance evaluation system" rather than a single indicator is to comprehensively measure the model's clinical practical value: AUC-ROC reflects the model's ability to distinguish between "high-risk / low-risk patients"; PR-AUC adapts to the clinical reality of a low proportion of HA-AKI positive samples, avoiding evaluation bias caused by a single AUC-ROC; F1 score balances the model's precision (avoiding false positives for high-risk cases) and recall (avoiding false negatives for high-risk cases); Brier score assesses the degree of calibration of risk probability (ensuring that predicted probabilities match actual risks); DCA curve directly quantifies the model's net clinical benefit (determining whether the model can truly help doctors improve decision-making). This multi-indicator synergistic evaluation ensures that the model not only "predicts accurately" but also meets the core needs of practical clinical applications, avoiding the limitations of traditional "single-indicator evaluation" that results in "performance meeting standards but clinical ineffectiveness."

[0056] The core of the model performance optimization and clinical interpretability enhancement (S54) design is to "balance performance and clinical credibility, breaking down the 'black box' barrier." Based on multi-dimensional evaluation results, the optimal model is selected (in this embodiment, the LightGBM model performs best, with an AUC-ROC of 0.957 on the test set), ensuring that the model is at its best in terms of discrimination ability, accuracy, stability, and clinical benefit. The introduction of SHAP value interpretability analysis addresses a core pain point in the clinical application of AI models—"opaque decision-making logic, making it difficult for doctors to trust them." SHAP value analysis quantifies the "contribution direction" (promoting or inhibiting risk) and "contribution intensity" of each clinical feature (such as the 3-day rolling average of blood urea nitrogen, eGFR, and medication use records) to the predicted outcome for individual patients. This transforms the model's warning basis from "fuzzy probability" to "clear clinical factors." Simultaneously, the analysis covers the training set, test set, and key subgroups such as non-ICU / ICU, elderly / non-elderly, cardiology / non-cardiology, and sepsis / non-sepsis. This verifies the stability of the model's decision-making basis across different populations (e.g., the core influencing features for the entire population and each subgroup are the 3-day rolling average of blood urea nitrogen, eGFR, and the 3-day rolling average of proBNP), while also providing targeted risk interpretation for different clinical scenarios. This combination of "high performance + interpretability" overcomes the clinical implementation barrier of traditional AI models that "only provide results without explanations," enabling doctors to understand and trust the warning information and translate it into precise interventions, significantly enhancing the clinical practical value of the technology.

[0057] In summary, step S5, through a comprehensive design that "ensures performance ceiling through multiple algorithm selection, ensures parameter robustness through targeted training and verification, ensures generalization ability through cross-departmental evaluation, and ensures clinical trust through enhanced interpretability," systematically solves the problems of existing technologies such as single algorithm, limited verification, and lack of interpretability. It constructs an early warning model that combines technological superiority with clinical adaptability, and is the core support for realizing HA-AKI's "precise early warning + clinical implementation."

[0058] S6: Perform clinical integration of models and automated hierarchical early warning, select the optimal model to integrate into the hospital information system, and build an automated early warning closed loop of timed data update - feature processing - risk prediction - clinical decision support.

[0059] In this embodiment, step S6 specifically includes: S61: Optimal Model Clinical System Integration: The early warning model with the best overall performance obtained from step S5 is integrated into the hospital information system of medical institutions through a standardized API service interface. An independent HA-AKI risk prediction function module is built to achieve seamless connection between the model and the existing clinical data system, ensuring stable access to the risk prediction function and secure data interaction. S62: Deployment and Implementation of Automated Early Warning Process: Configure the early warning system to a customizable timed data acquisition mode (e.g., 8:00 AM daily); the system automatically captures the latest clinical data of patients before the corresponding time node through the hospital information system interface at regular intervals, and performs standardized preprocessing and feature engineering processes on the acquired data that are completely consistent with the model training phase to ensure the consistency of data format and feature dimensions; input the processed standardized feature data into the integrated optimal early warning model (e.g., the optimal model in this embodiment is the LightGBM model), and the model automatically calculates and outputs the accurate risk probability of the patient developing HA-AKI in the next early warning cycle (e.g., the next 24 hours); S63: Clinical Decision Support and Tiered Early Warning Implementation: Through a visual clinical decision support interface, the system simultaneously displays to medical staff the patient's HA-AKI risk probability, as well as key risk factors, including the direction and intensity of risk contribution, obtained based on SHAP value analysis. The system has built-in risk grading threshold rules, supporting the preset judgment thresholds for different risk levels according to clinical needs. When the patient's risk probability reaches the corresponding threshold, a tiered early warning is automatically triggered, providing medical staff with intuitive and quantitative decision-making basis, assisting them in timely assessing the patient's condition and adjusting the treatment plan, and achieving early intervention for HA-AKI.

[0060] Step S6 is the core component for the clinical implementation of the entire HA-AKI intelligent early warning method. Its core design goal is to overcome the limitations of traditional technologies, such as "disconnect between the model and clinical system, manual processes, and lack of targeted early warnings." Through a full-chain design of "seamless system integration - automated process deployment - clinical decision support," it transforms validated high-performance early warning models into practical tools that can be directly applied clinically, constructing an automated closed loop of "data-analysis-early warning-intervention." The specific design logic and technical advantages are explained below: From the perspective of optimal model-clinical system integration (S61), the core design of this solution is "standardized adaptation + independent module construction," ensuring the compatibility and security of the model with existing clinical systems. The Hospital Information System (HIS) is the core carrier for clinical data storage and transfer. Integration using standardized API service interfaces avoids adaptation challenges caused by differences in the architecture of different hospital information systems, achieving seamless integration between the model and the HIS system. This ensures both the secure access to patient clinical data (complying with medical data privacy protection requirements) and the stable response of the risk prediction function, avoiding impacts on clinical use due to system compatibility issues. Constructing an independent HA-AKI risk prediction function module, rather than directly embedding it into existing system processes, allows for the independent operation and maintenance of the early warning function without interfering with the hospital's existing diagnostic and treatment data flow and business processes, reducing the implementation cost and risk of the technology. Simultaneously, the independent module design facilitates subsequent functional iteration and parameter optimization based on clinical feedback, improving the sustainability of the technology.

[0061] The core design of the automated early warning process deployment (S62) is "end-to-end unmanned intervention + data consistency guarantee," addressing the pain points of traditional early warning technologies such as "cumbersome manual operation and data processing deviations." The system is configured with a customizable timed data acquisition mode (e.g., 8:00 AM daily), designed based on the work rhythm of clinical diagnosis and treatment. Data is updated and early warnings are generated at fixed daily times, ensuring the timeliness of early warning information (covering risks for the next 24 hours) and allowing medical staff to develop a fixed viewing habit, increasing the utilization rate of early warning information. The "custom" design adapts to the differences in diagnosis and treatment processes in different hospitals and departments (e.g., some ICU departments may require early warnings every 12 hours), enhancing the technology's scenario adaptability. In the data processing stage, the emphasis on "standardized preprocessing and feature engineering processes completely consistent with the model training phase" is crucial to ensuring the accuracy of prediction results. The high performance of the model is based on the specific data processing logic formed during the training phase. Maintaining process consistency during deployment avoids prediction deviations caused by data format conversions and changes in feature dimensions, ensuring that the risk probability calculation for each patient is based on the same standards as the training samples, allowing the model's high performance to be realistically reproduced in clinical scenarios. The entire process of "automatic data capture - automatic data processing - automatic calculation and prediction" completely eliminates the reliance on manual input and processing, which not only improves the efficiency of early warning, but also eliminates the errors that may be caused by manual operation, laying the foundation for large-scale clinical application.

[0062] The core design of the Clinical Decision Support and Tiered Early Warning Implementation (S63) is "clinical presentation + precise intervention," ensuring that early warning information truly serves clinical decision-making. The design of the visual clinical decision support interface abandons the technical output of raw data, instead presenting medical staff with intuitive information on "risk probability + key risk factors." The risk probability uses quantified values ​​(e.g., 72%), allowing doctors to quickly assess a patient's risk level. Key risk factors based on SHAP value analysis (including contribution direction and intensity) transform the model's "black box" decision-making logic into clinically understandable language (e.g., "the 3-day rolling average of blood urea nitrogen is trending upward" and "eGFR is continuously decreasing"), overcoming the clinical trust barrier of traditional AI models that "only provide results without explanations." The system's built-in risk grading threshold rules are designed based on the clinical intervention logic of HA-AKI. Different intervention strategies are corresponding to patients with different risk levels (e.g., high-risk patients need immediate assessment of renal function and adjustment of nephrotoxic drugs, while medium-risk patients need increased monitoring frequency). By automatically triggering graded warnings through preset thresholds, doctors can quickly distinguish priorities, prioritize the treatment of high-risk patients, and optimize the efficiency of medical resource allocation. At the same time, the thresholds can be customized according to clinical needs to adapt to the differences in diagnosis and treatment standards of different hospitals and departments (e.g., some departments have a lower risk tolerance, so the high-risk threshold can be lowered), further improving the clinical adaptability of the technology.

[0063] In summary, step S6, through a full-chain design of "standardized integration to ensure feasibility of implementation, automated processes to improve efficiency, and clinical presentation to enhance decision-making value," systematically solves the problems of "difficult implementation, low efficiency, and poor clinical adaptability" of traditional technologies. It successfully transforms the high-performance, interpretable model built in the previous steps into an automated early warning tool that can be directly applied in clinical practice, realizing the final mile leap from "technology modeling" to "clinical empowerment." This allows early warning of HA-AKI to be truly integrated into the clinical diagnosis and treatment process, providing medical staff with accurate, efficient, and reliable decision support, and ultimately helping to achieve early intervention and prognosis improvement of HA-AKI.

[0064] Second Embodiment like Figure 2 As shown in the figure, this embodiment, combined with specific clinical data and practical procedures, details the implementation process of the intelligent early warning method for hospitalized acquired acute kidney injury based on multi-source data fusion of the present invention: (1) Data preparation The data comes from the information platform of Ruijin Hospital affiliated to Shanghai Jiao Tong University School of Medicine. It specifically extracts data from all inpatients in the hospital's cardiology, cardiac surgery, respiratory, emergency and critical care departments from August 1, 2021 to July 31, 2023, collecting a total of 47,207 patient records.

[0065] The above data were rigorously screened according to the research subject screening criteria described in step S2 of this invention: Inclusion criteria: age ≥ 18 years, length of hospital stay > 2 days, number of creatinine measurements ≥ 2 during hospitalization; Exclusion criteria: serum creatinine ≥353.6 μmol / L within 48 hours of admission, missing key clinical diagnostic and treatment data, confirmed diagnosis of AKI within 48 hours of admission, receiving renal replacement therapy within 48 hours of admission, or having a history of long-term maintenance dialysis treatment.

[0066] After screening, a total of 17,513 patients were ultimately included in the study cohort. Of these, data from 14,308 patients outside the ICU (general wards) were used as the training set for model parameter fitting and algorithm optimization; data from 3,205 patients in the ICU were used as the independent test set for external validation of the model's generalization ability (see appendix for detailed screening procedures). Figure 3 ).

[0067] (2) Feature processing The feature engineering process described in step S3 of this invention is performed on the training set data. First, an initial feature dataset is constructed based on multi-source data, containing a total of 89 feature variables. Then, the LASSO regression analysis algorithm is used to filter the initial features under the lambda-1se criterion, ultimately retaining 24 key clinical variables with non-zero regression coefficients (see Appendix for specific screening results). Figure 4 ).

[0068] Based on this, and following the time-series feature construction requirements of step S4 of this invention, multi-dimensional time-series features are constructed for the aforementioned 24 key variables: laboratory indicator time-series features include dynamic change rate, 3-day rolling average, historical maximum value, and historical minimum value; drug use time-series features include the cumulative number of days of use for various key drugs; other clinical event time-series features include postoperative days related to surgery, etc. The selected key variables are deeply integrated with the constructed time-series features. After eliminating feature redundancy, an optimized feature dataset containing 63 features is finally formed, adapted to the modeling needs of both the training and test sets.

[0069] (3) Model training, validation and selection Using the 63-dimensional optimized feature dataset of the training set as input, and based on the multi-algorithm model library described in step S5 of this invention, eight machine learning early warning models are constructed respectively: Logistic Regression (LR), Random Forest (RF), Linear Support Vector Machine (LinearSVC), Neural Network (NNET), K Nearest Neighbors (KNN), Lightweight Gradient Boosting Machine (LightGBM), Extreme Gradient Boosting (XGBoost), and Gradient Boosting Decision Tree (GBDT). Five-fold internal cross-validation is performed on the training set to optimize the model parameters.

[0070] Cross-validation results showed that the LightGBM model performed best, with an average AUC-ROC of 0.968, significantly outperforming other algorithms. To further verify the model's generalization ability, the LightGBM model was applied to an independent test set (ICU patient data) for performance evaluation. A multi-dimensional evaluation system consisting of AUC-ROC, PR-AUC, F1 score, Brier score, and DCA curve was used. The results showed that the model achieved an AUC-ROC of 0.957, a PR-AUC of 0.556, an F1 score of 0.444, and a Brier score of 0.042 on the test set, demonstrating excellent performance across all metrics and good suitability for high-risk populations (see Figure 5 for detailed evaluation results). Based on the combined results of cross-validation on the training set and external validation on the test set, the LightGBM algorithm was ultimately chosen to construct the HA-AKI intelligent early warning model in this embodiment.

[0071] (4) SHAP-based model interpretability analysis In accordance with the clinical interpretability analysis requirements of step S5 of this invention, SHAP value analysis is performed on the trained LightGBM model. The analysis scope comprehensively covers the training set, the test set, and multiple key clinical subgroups, including non-ICU / ICU patients, non-elderly / elderly patients, non-cardiology / cardiology patients, and non-sepsis / sepsis patients.

[0072] The analysis showed that the top three features contributing most significantly to HA-AKI risk prediction were highly consistent across the entire population and all clinical subgroups: the 3-day rolling mean of blood urea nitrogen (BUN), estimated glomerular filtration rate (eGFR), and the 3-day rolling mean of pro-brain natriuretic peptide (proBNP) (see Figure 6 for detailed analysis results). This result not only confirms that the model has stable and reliable decision-making basis across different patient groups but also clarifies the core role of dynamic changes in renal and cardiac function-related indicators in AKI risk early warning. It effectively breaks down the "black box" barrier of traditional machine learning models, significantly improving the model's clinical interpretability and credibility.

[0073] (5) Examples of early warning applications After the model completes training and validation, the LightGBM early warning model is integrated into the hospital information system according to the deployment scheme described in step S6 of this invention to form an automated hierarchical early warning tool. Its typical operating mode is as follows: Taking an ICU inpatient as an example, the early warning system is configured to automatically start running at 8:00 AM daily. The system automatically extracts the patient's latest clinical data up to 8:00 AM that day (including demographic information, laboratory indicators, medication records, surgical records, imaging reports, and other multi-source data) through the hospital information system interface. Subsequently, it performs a standardized preprocessing process (structured data cleaning, normalization, time window alignment, entity recognition, relation extraction, and structured transformation of unstructured text data) and feature engineering process on the extracted raw data, generating a 63-dimensional standardized feature vector. This feature vector is then input into the integrated LightGBM model, which automatically calculates and outputs a 72% risk probability for the patient to develop HA-AKI within the next 24 hours.

[0074] Based on preset risk grading threshold rules, the system identifies the patient as a high-risk individual, automatically triggering a high-risk warning. This warning information is then displayed to healthcare professionals through a clinical decision support interface. Specifically, it states: "High-risk warning (72% probability of occurrence) – Key risk contributing factors: 1. Increasing trend in the 3-day rolling average of blood urea nitrogen; 2. Continuously decreasing estimated glomerular filtration rate (eGFR); 3. Day 3 post-surgery." Based on this warning information and identified risk factors, healthcare professionals can promptly assess the patient's renal function, adjust the nephrotoxic drug regimen, or implement other targeted interventions to reduce the risk of HA-AKI and improve patient prognosis.

[0075] Third Embodiment like Figure 7 As shown, this embodiment provides a hospital-acquired acute kidney injury intelligent early warning system for executing the hospital-acquired acute kidney injury intelligent early warning method as described in the first embodiment, comprising: The multi-source data integration and preprocessing module is used to integrate and collect heterogeneous data from multiple sources and perform standardized preprocessing. It collects structured data, unstructured text data and other derived computational data from hospitalized patients, and constructs a standardized patient time-series data system through data cleaning, normalization and text structuring transformation. The cross-departmental dataset partitioning and screening module is used to partition cross-departmental datasets and accurately screen research subjects. It divides the training set and test set based on the differences in departmental risk levels and screens suitable research subject samples according to the HA-AKI diagnostic criteria and strict inclusion and exclusion criteria. The two-stage feature variable screening module is used to screen the core variables for two-stage feature engineering. First, it constructs an initial feature set by hierarchical processing of missing values, and then uses a regularized regression algorithm to screen key variables, thereby achieving feature dimensionality reduction and overfitting suppression. The temporal feature construction and optimization module is used to construct and optimize dynamic temporal features. It constructs multi-dimensional temporal features for key variables and integrates key variables and temporal features to form an optimized feature dataset for early warning that is adapted to a preset time period in the future. Multi-algorithm modeling, verification, and interpretation are used for multi-algorithm model construction, cross-validation, and interpretability enhancement. Early warning models are built based on multiple machine learning algorithms. The generalization ability of the model is evaluated through cross-departmental verification, and the contribution of features is quantified by combining interpretability algorithms. The model clinical integration early warning is used for model clinical integration and automated hierarchical early warning. It selects the optimal model to integrate into the hospital information system and builds an automated early warning closed loop of timed data update, feature processing, risk prediction and clinical decision support.

[0076] A computer-readable storage medium stores computer code that, when executed, performs the methods described above. Those skilled in the art will understand that all or part of the steps in the various methods of the above embodiments can be implemented by a program instructing related hardware. This program can be stored in a computer-readable storage medium, which may include: read-only memory (ROM), random access memory (RAM), a magnetic disk, or an optical disk, etc.

[0077] The above description is merely a preferred embodiment of the present invention. The scope of protection of the present invention is not limited to the above embodiments. All technical solutions falling within the scope of the present invention's concept are within the scope of protection of the present invention. It should be noted that for those skilled in the art, any improvements and modifications made without departing from the principles of the present invention should also be considered within the scope of protection of the present invention.

[0078] The technical features of the above embodiments can be combined in any way. For the sake of brevity, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.

[0079] It should be noted that the above embodiments can be freely combined as needed. The above description is only a preferred embodiment of the present invention. It should be pointed out that for those skilled in the art, several improvements and modifications can be made without departing from the principle of the present invention, and these improvements and modifications should also be considered within the scope of protection of the present invention.

Claims

1. A smart early warning method for hospitalized acute kidney injury, characterized in that, Includes the following steps: S1: Integrate and collect multi-source heterogeneous data and perform standardized preprocessing. Collect structured data, unstructured text data and other derived computational data of hospitalized patients. Through data cleaning, normalization and text structuring transformation, construct a standardized patient time-series data system. S2: Perform cross-departmental dataset division and precise screening of research subjects. Divide the training set and test set based on the differences in departmental risk levels, and screen suitable research subject samples according to the HA-AKI diagnostic criteria and strict inclusion and exclusion criteria. S3: Select the core variables for two-stage feature engineering. First, construct an initial feature set by hierarchical processing of missing values, and then use a regularized regression algorithm to select key variables to achieve feature dimensionality reduction and overfitting suppression. S4: Construct dynamic features and optimize secondary features for time series. Build multi-dimensional time series features for key variables and integrate key variables and time series features to form an optimized feature dataset for early warning that is adapted to a future preset time period. S5: Construct, cross-validate, and enhance interpretability of multi-algorithm models; build early warning models based on multiple machine learning algorithms; evaluate the generalization ability of models through cross-departmental validation; and quantify feature contribution by combining interpretability algorithms. S6: Perform clinical integration of models and automated hierarchical early warning, select the optimal model to integrate into the hospital information system, and build an automated early warning closed loop of timed data update - feature processing - risk prediction - clinical decision support.

2. The intelligent early warning method for hospitalized acute kidney injury according to claim 1, characterized in that, In step S1, multi-source heterogeneous data is integrated, collected, and standardized preprocessed. Structured data, unstructured text data, and other derived computational data from hospitalized patients are collected. Through data cleaning, normalization, and text structuring transformation, a standardized patient time-series data system is constructed, specifically as follows: S11: Define the data source and time range: The data comes from the hospital information platform of the medical institution and is included in all inpatient data of the institution, including cardiology, cardiac surgery, respiratory medicine, emergency medicine and critical care medicine, within the preset time period. This ensures that the data covers patient groups with different disease severity and department types, and guarantees the representativeness and diversity of the data, so as to provide comprehensive sample support for subsequent model training and validation. S12: Conduct multi-dimensional data collection: The types of data collected include structured data, unstructured text data, and other derived computational data; Structured data covers: Demographic information: Record ID, gender, age, height, weight, admission time, discharge time, length of hospital stay, admission department, admission diagnosis, discharge department, discharge diagnosis, and discharge status; Laboratory indicators: complete blood count, inflammatory markers, renal function indicators, liver function indicators, cardiac function indicators, coagulation function indicators, and electrolyte indicators; Vital signs data: respiratory rate, systolic blood pressure, diastolic blood pressure, heart rate, body temperature, and blood oxygen saturation; Medical records of medication and treatment: including records of the use of anticoagulants, nonsteroidal anti-inflammatory drugs, diuretics, vasoactive drugs, nephrotoxic antibiotics, and contrast agents, as well as clinical treatment records of mechanical ventilation and renal replacement therapy; Unstructured text data includes present medical history texts and imaging reports containing descriptions of pleural effusion; Other derived calculation data include the ratio of blood urea nitrogen to blood creatinine, the ratio of neutrophils to lymphocytes, and the sequential organ failure SOFA score (excluding the nervous system module). The sequential organ failure SOFA score must be greater than or equal to the preset score. S13: Implement data standardization preprocessing: Structured data preprocessing: Data cleaning algorithms are used to identify and process missing values ​​and detect and remove outliers in structured data. Then, normalization is used to unify the data scale and eliminate the dimensional differences between different indicators. Finally, time window alignment technology is used to align the structured data collected at different time points according to the preset time granularity to form standardized patient time series data. Unstructured text data structuring transformation: Using natural language processing technology, an entity recognition and relation extraction framework is built based on a pre-trained biomedical language model. Semantic parsing is performed on current medical history texts and imaging reports to accurately extract key clinical information, including comorbidities and pleural effusion. The extracted unstructured information is then mapped into structured feature vectors that can be used for model training, achieving homogeneous integration of multi-source data.

3. The intelligent early warning method for hospitalized acute kidney injury according to claim 1, characterized in that, In step S2, cross-departmental dataset partitioning and precise screening of research subjects are performed. Training and test sets are divided based on differences in departmental risk levels. Suitable research subject samples are selected according to the HA-AKI diagnostic criteria and strict inclusion and exclusion rules. Specifically: S21: Cross-departmental differentiated dataset partitioning: Based on the differences in the disease risk level of hospitalized patients in different departments, the data of hospitalized patients in general wards (excluding ICU) are set as the model training set for model parameter fitting and algorithm optimization; the data of hospitalized patients in ICU intensive care units are set as an independent external test set, specifically for verifying the model's generalization ability. By separating and verifying the data from high-risk departments and general departments, the adaptability and predictive reliability of the model in patient groups with different risk levels are ensured. S22: Study Subject Inclusion Screening: Establish uniform inclusion criteria and include only patients who meet all of the following conditions in the study cohort: age ≥18 years; hospital stay >2 days; creatinine measurement ≥2 times during hospitalization, to ensure that the study sample has sufficient clinical data support and follow-up period; S23: Exclusion and Screening of Study Subjects: Strict exclusion criteria were established, and patients meeting any of the following conditions were excluded: serum creatinine ≥353.6 μmol / L within 48 hours of admission; missing key clinical diagnostic and treatment data affecting risk factor assessment; confirmed diagnosis of acute kidney injury (AKI) within 48 hours of admission; receiving renal replacement therapy within 48 hours of admission; or having a history of long-term maintenance dialysis treatment to avoid interfering with the model's early warning judgment of hospitalized AKI. S24: Standardized diagnosis of HA-AKI and calculation of baseline creatinine: The internationally recognized KDIGO standard is used as the basis for the diagnosis of HA-AKI, that is, the absolute value of serum creatinine increases by ≥26.5 μmol / L within 48 hours, or increases to more than 1.5 times the baseline creatinine value. Baseline creatinine was calculated using the MDRD normalization formula, assuming an estimated glomerular filtration rate (eGFR) of 75 mL / min / 1.73 m 2 The MDRD formula is as follows: Baseline creatinine = (75 / [186×(age)) -0.203 )×(0.742 if female)×(1.21 if black)]) -0.887 The variables in the formula are defined as follows: age is the patient's actual age; female is the gender identifier variable, which is 0.742 when the patient is female and 1.0 when the patient is male; black is the race identifier variable, which is 1.21 when the patient is Black and 1.0 when the patient is not Black. By clearly defining the variables in the formula, a unified quantitative calculation of baseline creatinine is achieved, ensuring the consistency and accuracy of the HA-AKI diagnostic criteria.

4. The intelligent early warning method for hospitalized acute kidney injury according to claim 1, characterized in that, In step S3, the core variable selection for two-stage feature engineering is performed. First, an initial feature set is constructed through hierarchical processing of missing values. Then, a regularized regression algorithm is used to select key variables, achieving feature dimensionality reduction and overfitting suppression. Specifically: S31: Stratified handling of missing variable values: Establish a quantitative assessment mechanism for the missing variable rate, and perform statistical analysis on the missing value of the feature variables corresponding to the research object samples after screening in step S2; for low information value variables with a missing value greater than the preset missing value, they are directly removed to avoid interference of invalid data on model performance; for valid feature variables with a missing value less than or equal to the preset missing value, a hybrid imputation scheme combining forward imputation, backward imputation and global median imputation is adopted. By supplementing the data through the correlation between time series data and the reference of the overall distribution characteristics, the data integrity is restored to the maximum extent and the feature quality is guaranteed. S32: Initial feature dataset construction: Based on the standardized time-series data of the training set, a time-series feature alignment mechanism is adopted to construct an initial feature dataset corresponding to each time node, which is used to predict the occurrence status of acute kidney injury (AKI) within a preset time period in the future. This achieves accurate binding between the warning time window and the feature data, and provides targeted feature support for subsequent model training. S33: Selection of key variables for regularized regression: LASSO regression analysis is used as the regularized regression algorithm. Under the lambda.1se criterion, variables in the initial feature dataset are selected for regularization. The interference of redundant variables is suppressed by the regularization penalty mechanism, and only key clinical variables with non-zero regression coefficients are retained. This not only achieves efficient compression of feature dimensions, but also effectively suppresses the risk of overfitting during model training, and improves the generalization ability and prediction stability of the model.

5. The intelligent early warning method for hospitalized acute kidney injury according to claim 1, characterized in that, In step S4, dynamic temporal features are constructed and secondary features are optimized. Multi-dimensional temporal features are constructed for key variables, and the key variables and temporal features are integrated to form an optimized feature dataset for early warnings adapted to a preset future time period. Specifically: S41: Construction of a Multi-Dimensional Temporal Feature System: Based on the key clinical variables screened by LASSO in step S3, and combined with the dynamic evolution of HA-AKI incidence risk, targeted temporal features are constructed according to variable type, specifically including: Laboratory indicator time series characteristics: Calculate the dynamic change rate, 3-day rolling average, historical maximum and historical minimum of key variables for each laboratory to accurately capture the fluctuation trend and cumulative effect of indicators over time; Drug use timing characteristics: Statistical analysis of the cumulative number of days of use for various key drugs to quantify the duration of drug exposure on the risk of HA-AKI; Other clinical event time-series features: Extract dynamic event features related to surgery, including postoperative days, to reflect changes in risk over time after clinical interventions; S42: Construction of secondary feature dataset fusion: Using feature fusion technology, the key clinical variables selected in step S3 are deeply integrated with the multi-dimensional time-series features constructed in step S41 to eliminate feature redundancy and enhance risk association information. For the training set and the test set respectively, a secondary feature dataset adapted to the HA-AKI early warning scenario in the future preset time period is reconstructed to provide more targeted and predictive feature support for subsequent model training and validation.

6. The intelligent early warning method for hospitalized acute kidney injury according to claim 1, characterized in that, In step S5, multi-algorithm model construction, cross-validation, and interpretability enhancement are performed. An early warning model is built based on multiple machine learning algorithms. The model's generalization ability is evaluated through cross-departmental validation. The feature contribution is quantified using interpretability algorithms. Specifically: S51: Construction of a multi-algorithm early warning model library: Construct a multi-dimensional machine learning algorithm model library including Logistic Regression (LR), Random Forest (RF), Linear Support Vector Machine (LinearSVC), Neural Network (NNET), K Nearest Neighbors (KNN), Lightweight Gradient Boosting Machine (LightGBM), Extreme Gradient Boosting Machine (XGBoost), and Gradient Boosting Decision Tree (GBDT). Through parallel construction of multiple algorithms, a sufficient algorithmic foundation is provided for subsequent model optimization. S52: Model-oriented training and cross-validation optimization: Using the training set secondary feature dataset constructed in step S4 as input, model parameter training is carried out on patient data in general wards outside the ICU. At the same time, a 5-fold internal cross-validation technique is adopted. By randomly dividing the training set into 5 mutually exclusive subsets, 4 subsets are used as training samples and 1 subset is used as validation samples for iterative training and parameter adjustment, which effectively avoids model overfitting and improves model training stability and parameter optimization accuracy. S53: Cross-departmental model generalization ability assessment: Using the test set secondary feature dataset constructed in step S4 as input, external model validation is carried out on the independent test set of the ICU. A multi-dimensional performance evaluation system is used to comprehensively evaluate the model performance. Evaluation indicators include the area under the receiver operating characteristic curve (AUC-ROC), the area under the precision-recall curve (PR-AUC), the F1 score, the probability prediction error Brier score, and the DCA curve of decision curve analysis. Through cross-departmental data validation, the generalization ability of the model in high-risk patient groups is accurately judged. S54: Model Performance Optimization and Enhanced Clinical Interpretability: Based on the multi-dimensional performance evaluation results of step S53, the early warning model with the best overall performance is selected. For this optimal model, the SHAP value interpretability analysis algorithm is used to quantify the contribution direction and intensity of each clinical feature to the HA-AKI risk prediction results. The analysis scope comprehensively covers the training set, the test set, and multiple key clinical subgroups, including non-ICU / ICU patients, non-elderly / elderly patients, non-cardiology / cardiology patients, and non-sepsis / sepsis patients. This achieves a transparent presentation of the model's prediction logic and improves the credibility of clinical applications.

7. The intelligent early warning method for hospitalized acute kidney injury according to claim 1, characterized in that, In step S6, clinical integration and automated hierarchical early warning of the model are performed. The optimal model is selected and integrated into the hospital information system to construct an automated early warning closed loop of timed data updates, feature processing, risk prediction, and clinical decision support. Specifically: S61: Optimal Model Clinical System Integration: The early warning model with the best overall performance obtained from step S5 is integrated into the hospital information system of medical institutions through a standardized API service interface. An independent HA-AKI risk prediction function module is built to achieve seamless connection between the model and the existing clinical data system, ensuring stable access to the risk prediction function and secure data interaction. S62: Deployment and Implementation of Automated Early Warning Process: Configure the early warning system to a customizable timed data acquisition mode; the system automatically captures the latest clinical data of patients before the corresponding time node through the hospital information system interface at regular intervals, performs standardized preprocessing and feature engineering processes on the acquired data that are completely consistent with the model training phase, and ensures the consistency of data format and feature dimensions; input the processed standardized feature data into the integrated optimal early warning model, and the model automatically calculates and outputs the accurate risk probability of the patient developing HA-AKI in the next early warning cycle; S63: Clinical decision support and graded early warning implementation: Through a visual clinical decision support interface, the probability of HA-AKI in patients is displayed to medical staff in a synchronous manner, as well as key risk factors, including the direction and intensity of risk contribution, obtained based on SHAP value analysis. The system has built-in risk grading threshold rules, which support the preset judgment thresholds for different risk levels according to clinical needs. When the patient's risk probability reaches the corresponding threshold, a graded warning is automatically triggered, providing medical staff with intuitive and quantitative decision-making basis, assisting them in timely assessing the patient's condition and adjusting the treatment plan, and realizing early intervention of HA-AKI.

8. A hospital-acquired acute kidney injury intelligent early warning system for implementing the intelligent early warning method for hospital-acquired acute kidney injury as described in any one of claims 1-7, characterized in that, include: The multi-source data integration and preprocessing module is used to integrate and collect heterogeneous data from multiple sources and perform standardized preprocessing. It collects structured data, unstructured text data and other derived computational data from hospitalized patients, and constructs a standardized patient time-series data system through data cleaning, normalization and text structuring transformation. The cross-departmental dataset partitioning and screening module is used to partition cross-departmental datasets and accurately screen research subjects. It divides the training set and test set based on the differences in departmental risk levels and screens suitable research subject samples according to the HA-AKI diagnostic criteria and strict inclusion and exclusion criteria. The two-stage feature variable screening module is used to screen the core variables for two-stage feature engineering. First, it constructs an initial feature set by hierarchical processing of missing values, and then uses a regularized regression algorithm to screen key variables, thereby achieving feature dimensionality reduction and overfitting suppression. The temporal feature construction and optimization module is used to construct and optimize dynamic temporal features. It constructs multi-dimensional temporal features for key variables and integrates key variables and temporal features to form an optimized feature dataset for early warning that is adapted to a preset time period in the future. Multi-algorithm modeling, verification, and interpretation are used for multi-algorithm model construction, cross-validation, and interpretability enhancement. Early warning models are built based on multiple machine learning algorithms. The generalization ability of the model is evaluated through cross-departmental verification, and the contribution of features is quantified by combining interpretability algorithms. The model clinical integration early warning is used for model clinical integration and automated hierarchical early warning. It selects the optimal model to integrate into the hospital information system and builds an automated early warning closed loop of timed data update, feature processing, risk prediction and clinical decision support.

9. A computer device, characterized in that, The device includes a memory and one or more processors, wherein the memory stores computer code that, when executed by the one or more processors, causes the one or more processors to perform the method as described in any one of claims 1 to 7.

10. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores computer code, and when the computer code is executed, the method as described in any one of claims 1 to 7 is performed.