Detecting kidney injury based on machine learning

A machine learning model for AKI detection addresses the challenge of delayed and biased detection by integrating renal function data and handling missing outcomes, enabling early and accurate risk assessment for all patients, including discharged individuals, thereby improving emergency care.

WO2025240829A9PCT designated stage Publication Date: 2026-02-05JOHNS HOPKINS UNIVERSITY
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
PCT/US2025/029707
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-05-16
Filing Date
2025-05-16
Publication Date
2026-02-05

AI Technical Summary

Technical Problem

Current techniques for detecting acute kidney injury (AKI) are suboptimal due to the lack of timely detection, especially in emergency department settings, as they often rely on laboratory-based methods that are delayed or costly, and existing machine learning models are biased towards hospitalized patients, neglecting those discharged without recorded outcomes.

Method used

A machine learning model is developed to predict AKI risk by integrating baseline and current renal function data, addressing missing outcomes through techniques like multiple imputation and inverse probability weighting, enabling early detection and treatment for both hospitalized and discharged patients.

Benefits of technology

The model significantly reduces AKI detection delay and improves treatment timeliness by providing actionable risk estimates for all patients, including those discharged from the emergency department, enhancing decision support in emergency care.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure US2025029707_05022026_PF_FP_ABST
    Figure US2025029707_05022026_PF_FP_ABST
Patent Text Reader

Abstract

This document describes a method including reading digital health records each associated with a unique key value representing a patient, each digital health record being structured with a plurality of fields, parsing the plurality of fields to identify the unique key values, and identifying, based on the parsing, a given key value and a digital health record for that given key value. For the identified digital health record, the method includes generating baseline renal function data based on the particular digital health record, comparing the baseline renal function data with current renal function data associated with the given key value, determining that a difference between the baseline renal function data and the current renal function data is greater than a threshold, inputting the current renal function data to a machine learning model, and predicting, by the machine learning model, a propensity of AKI for the given key value.
Need to check novelty before this filing date? Find Prior Art

Description

Attorney Docket No : 44807-0484WO1DETECTING KIDNEY INJURY BASED ON MACHINE LEARNING

[0001] This application claims the benefit of U. S. Provisional Patent Application No. 63 / 648,433 filed on May 16, 2024, the contents of which are fully incorporated herein by reference.TECHNICAL FIELD

[0002] This document relates to systems and methods for using a machine learning model to detect a kidney injury.BACKGROUND

[0003] More than a million people develop acute kidney injury7(AKI) every7year, which results in a large number of death and economic loss. Different from chronic kidney diseases, AKI symptoms often develop rapidly within a few days or hours (e.g., 72 hours). Although AKI outcome improvement is achievable through early recognition and prevention of disease progression, current techniques are often suboptimal due to lack of capability to provide timely detection of new or progressing AKIs.SUMMARY

[0004] This document describes devices, systems, and methods for generating a machine learning model for estimating a probability that a patient will experience acute kidney injury (AKI) or a severity of AKI that a patient is likely to experience, based on patent data collected in a clinical setting, such as the emergency department (ED) of a clinic or hospital. In some cases, the machine learning model can be trained to determine a likelihood that a patient will experience AKI within a period of time (e.g., within 72 hours). One issue with training a machine learning model is that patient records are more complete for patients who are admitted to the hospital after arrival at the ED as compared with records for patients who arrive at the ED and are later discharged. Models trained only with patient records where an AKI outcome is known can be biased towards records from patients that were admitted to the hospital. The techniques described herein include techniques for supplementing training data to account for unknown AKI outcomes for patients who were admitted to the ED, tested, and later discharged without a recorded outcome.Attorney Docket No : 44807-0484WO1

[0005] In one aspect, a system for processing digital records includes a digital network in communication with one or more external systems, at least one processor, and a memory subsystem communicatively coupled to the at least one processor. The memory subsystem stores instructions which, when executed by the at least one processor, cause the at least one processor to perfonn operations comprising reading, by the digital network from a hardware storage device of the one or more external systems, digital health records each associated with a unique key value representing a patient, with each digital health record being structured with a plurality of fields, wherein the one or more external systems authenticate that the system is authorized to access the digital health records. The operations also include parsing, by a parser, the plurality of fields of the digital health records to identify the unique key values, and identifying, based on the parsing, a given key value and a digital health record for that given key value. For the identified digital health record corresponding to an identified patient, the operations include generating, by the at least one processor, baseline renal function data based on the particular digital health record, comparing the baseline renal function data with current renal function data associated with the given key value, determining that a difference between the baseline renal function data and the current renal function data is greater than a threshold, inputting the current renal function data to a machine learning model, predicting, by the machine learning model, a propensity' of acute kidney inj ury (AKI) for the given key value, and determining, by the at least one processor, a treatment of the identified patient based on the propensity of AKI.

[0006] In another aspect, a method for processing digital records includes reading, from a hardware storage device of one or more external systems, digital health records each associated with a unique key value representing a patient, yvith each digital health record being structured with a plurality of fields, wherein the one or more external systems authenticate that access to the digital health records is authorized, parsing the plurality of fields of the digital health records to identify the unique key values, and identifying, based on the parsing, a given key value and a digital health record for that given key value. For the identified digital health record corresponding to an identified patient, the method includes generating, by at least one processor, baseline renal function data based on the particular digital health record, comparing the baseline renal function data with current renal function data associated with the given key value, determining that a difference between the baseline renal function data and the current renal function data is greater thanAttorney Docket No : 44807-0484WO1 a threshold, inputting the current renal function data to a machine learning model, predicting, by the machine learning model, a propensity of acute kidney injury (AKI) for the given key value; and determining, by the at least one processor, a treatment of the patient based on the propensity of AKI.

[0007] In another aspect, a computer-readable medium includes instructions that, when executed by a processor, cause the processor to: read, from a hardware storage device of one or more external systems, digital health records each associated with a unique key value representing a patient, with each digital health record being structured with a plurality of fields, wherein the one or more external systems authenticate that access to the digital health records is authorized; parse the plurality of fields of the digital health records to identify the unique key values; and identify, based on the parsing, a given key value and a digital health record for that given key value. For the identified digital health record corresponding to an identified patient, the instructors cause the processor to generate, by at least one processor, baseline renal function data based on the particular digital health record, compare the baseline renal function data with current renal function data associated with the given key value, determine that a difference between the baseline renal function data and the current renal function data is greater than a threshold, input the current renal function data to a machine learning model, predict, by the machine learning model, a propensity of acute kidney injury (AKI) for the given key value, and determine, by the at least one processor, a treatment of the patient based on the propensity of AKI.

[0008] Particular embodiments of the subject matter described in this document can be implemented to realize one or more of the following advantages. For example, the machine learning model can be trained using some data that includes AKI diagnosis results and other training data that does not include AKI diagnosis results. Prior to using the training data that does not include AKI diagnosis results, a computing system can perfomi one or more actions for addressing the lack of AKI diagnosis results in samples missing these results. For example, techniques can be used to estimate the results (e.g., by imputing or using inverse propensity, or applying blind assumptions). Such techniques result in a machine learning model that more accurately predicts a probability of AKI in a current patient, because training data missing AKI diagnosis results can be skew ed towards datasets corresponding to patients who were discharged from the ED, which are important patients to consider in training the model.Attorney Docket No : 44807-0484WO1

[0009] The details of one or more embodiments of this disclosure are set forth in the accompanying drawings and the description herein. Other features, objects, and advantages of this disclosure will be apparent from the description and drawings, and from the claims.BRIEF DESCRIPTION OF THE DRAWINGS

[0010] FIG. 1 illustrates a block diagram of an example system for detecting kidney injury, according to some implementations.

[0011] FIG. 2 illustrates an example process of predicting AKI using a machine learning model, according to some implementations.

[0012] FIG. 3 illustrates another example process of predicting AKI using a machine learning model, according to some implementations.

[0013] FIG. 4 illustrates a flowchart of an AKI detection process, according to some implementations.

[0014] FIGs. 5A and 5B each illustrate a graph with example curves showing the performance of prediction, according to some implementations.

[0015] FIG. 6 illustrates a flowchart of an example method, according to some implementations.

[0001] FIG. 7 is a block diagram of an example computer system in accordance with one or more implementations.

[0016] Figures are not drawn to scale. Like reference numbers refer to like components.DETAILED DESCRIPTION

[0017] Existing AKI diagnosis techniques are usually laboratory -based. For tests that are routinely available (e.g., creatinine and urine output), it usually takes considerable delay for the AKI signals to manifest and become detectable. On the other hand, for tests that detect AKI signals in a more timely manner, these tests are often not routinely available and are costly.

[0018] Acute kidney injury' (AKI) is a prevalent clinical syndrome directly linked to increased morbidity and mortality. Patients with AKI can face heightened risks of prolonged and costly hospital stays, progression to chronic kidney disease, dialysis, major adverse cardiovascular events, and death. In some cases, diagnosis of AKI involves detecting elevated serum creatinine (sCr) concentration or reduced urine output. Both of these markers, however, represent indirect indicators of kidney function and can lag daysAttorney Docket No : 44807-0484WO1 behind the onset of kidney injury, contributing to under-recognition and delayed diagnosis, especially in emergency department (ED) settings. Early detection and risk stratification of AKI, paired with kidney-focused clinical decision support, can provide reduced AKI incidence and severity in some settings.

[0019] In some embodiments, artificial intelligence (Al) can be applied to routinely collected clinical data holds promise for AKI prediction and prevention. Some machine learning (ML) models, for example, can detect patterns in electronic health record (EHR) data to generate reliable AKI risk estimates in hospitalized patients. These ML models can drive decision support that mitigates or prevents kidney injury. Becasue care trajectories for millions of patients are set in EDs, it can be beneficial to provide decision support in the emergency care setting. To be used in the ED, Al-informed decision support can be broadly applicable with high reliability across a wide range of clinical presentations.

[0020] In the episodic care environment of the ED, a development of ML models to drive decision support can be challenged by missing outcomes data. Although EHR facilitates use automated capture of predictor variables and clinical outcomes in ED patients admitted to the hospital, most ED patients are discharged to the community where outcomes go unmeasured. Common approaches to address this challenge involve excluding patients with missing outcomes data from model development cohorts (complete case analyses), or assuming negative outcomes in those for whom follow-up data are unavailable. The first approach can result in models that are trained to make reliable predictions in ED patients who are hospitalized, the models being naive of the larger proportion of patients who are discharged to the ED. This approach may result in over-estimation of model performance during development and under-estimation of risk across the population when deployed in real-time.

[0021] In some embodiments, an AKI prediction model describe herein integrates seamlessly into ED clinical workflows to provide early, actionable risk estimates to support both disposition decisions and timely AKI mitigation interventions. Unlike previous efforts which focus on hospitalized patients, the AKI prediction model can account for the challenges of missing outcomes data to expand applicability. This means that the AKI prediction model can be used to determine outcomes for discharged patients, an attribute required for utility in the emergency care setting where decision support is needed long before a disposition is determined. Here, In some examples, the AKI prediction model described herein can estimate risk for new or progressive AKI within 72Attorney Docket No : 44807-0484WO1 hours of ED departure, thus improving patient outcomes and advancing real-time decision support in emergency care as compared with solutions that do not account for such missing outcomes.

[0022] In this disclosure, an ML-based mechanism can assist in early AKI diagnosis and treatment in healthcare departments. Implementations of the disclosure can involve obtaining baseline renal function data and current renal function data of a patient and use a trained ML model to predict a propensity of new or progressing AKI. These techniques may, in some cases, significantly reduces a delay of AKI detection and improves a timeliness of diagnosis and treatment as compared with systems that do not assist in early AKI diagnosis based on data from EDs, including data from patients who are discharged from the ED.

[0023] FIG. 1 illustrates a block diagram of an example system 100 for detecting kidney injury, according to some implementations. As illustrated, system 100 has data processing system 101 communicatively coupled to database 105 via a digital network, which can be configured to transmit and receive encrypted patient data. Database 105 can be implemented as a hardware storage device of one or more external systems configured to authenticate access requests. System 100 also has machine learning model 130, which is implemented outside of data processing system 101 and in communication with data processing system 101. In some alternative implementations, machine learning model 130 is integrated within data processing system 101.

[0024] Data processing system 101 has processor 102 and memory subsystem 103 (which can include one or more memory circuits) communicatively coupled to each other.Processor 102 can execute program instructions stored in memory subsystem 103 to perform, e.g., healthcare management applications. For example, when patient 150 is admitted to a hospital, processor 102 can execute applications to digitally record the demographic information (e.g., sex, age, or race) and physiological information (e.g., height, w eight, and vital measurements). Processor 102 can parse, using parser 104, fields stored in digital health records of the patients. From the parsed information, processor 102 can extract various information associated to patient 150. In implementations where machine learning model 130 is implemented within data processing system 101, memory subsystem 103 can also store software code and / or parameters of machine learning model 130.Attorney Docket No : 44807-0484WO1

[0025] Data processing system 101 can exchange data with database 105, which can be configured to store health records of patients 151-153. The exchange can be via one or more application programming interfaces (APIs), such as Fast Healthcare Interoperability Resources (FHIR), Cloud Healthcare API, or GOOGLE CLOUD. When patient 150 is admitted, data processing system 101 can encrypt and transmit the health records of patient 150 to database 105. For each patient, the health records stored in database 105 can be updated real-time, on a regular basis, or upon the occurrence of an event. For example, patient 150 can wear sensors configured to capture physiological information. The sensors can regularly transmit the captured physiological information to data processing system 101, which can in turn update the health records of patient 150 in database 105. This, via data processing system 101. one can inquire into database 105 about historical measurements of physiological information and diagnosis and treatment history of a patient.

[0026] In addition to information of individual patients, database 105 can store statistical data of a larger population as part of the health records. For example, database 105 can store data that indicates the distribution of physiological information across different age groups, sex groups, race groups, or medical history groups.

[0027] Based on health records 110 stored in database 105, data processing system 101 can generate baseline renal function data, which can indicate the expected renal function of patient 150 absent AKI. For example, when patient 150 is admitted, data processing system 101 can inquire about the medical history of patient 150, such as prior renal function data of patient 150 obtained from real-time surveillance. Such data can be organized into a plurality7of fields of the health records and can be parsed by parser 104. The fields can include, e.g., a unique key that identifies the patient, prior medical data, and / or the patient’s demographic information. If a health record 110 includes prior renal function data of the patient as identified by the unique key, then data processing system 101 can extract, using parser 104, patient 150’s prior renal function data from health records 110 to obtain baseline renal function data. If the health records 110 do not include prior renal function data of patient 150. then data processing system 101 can estimate baseline renal function data according to the statistics of a larger population with the same demographic information as patient 150. The generation of baseline renal function data can be based on Kidney Disease Improving Global Outcomes (KDIGO), which provide criteria for staging renal function data for different stages of AKI.Attorney Docket No : 44807-0484WO1

[0028] Data processing system 101 obtains current renal function data of patient 150, which can be obtained when patient 150 is admitted. Data processing system 101 can compare current renal function data with the baseline renal function data to determine any deviation between the two sources of data. If the current renal function data is consistent with the baseline renal function data, then data processing system 101 can determine that the actual kidney function of patient 150 is normal compared to the baseline, so AKI is unlikely to develop within the next few days (e.g., 72 hours). If the current renal function data is inconsistent with the baseline renal function data (e.g., the difference between the current renal function data and the baseline renal function data is greater than a threshold), then data processing system 101 can determine that the actual kidney function of patient 150 is outside the range expected for a normal person. Still, data processing system 101 needs to determine whether patient 150 is at the risk of having AKI soon or having chronic kidney in the long term. To this end, data processing system 101 can input current renal function data 140 of patient 150 to ML model 130, such as an XGBOOST model, which can be trained to predict the propensity of AKI.

[0029] In some cases, it is beneficial for ML model 130 to estimate a probability any new or progressive AKI within 72 hours of ED departure. In some examples, stage 1 AKI is defined as an absolute increase of > 0.3 mg / dL sCr concentration or >1.5 times baseline sCr concentration, stage 2 AKI is defined as an increase of of 2.0-2.9 times baseline sCr concentration, stage 3 AKI is defined as an increase of> 4.0 mg / dL or > 3.0 times baseline or initiation of RRT. In some cases, urine output-based criteria are not included in a definition of AKI because this variable was not reliably recorded in the HER, but this is not required. In some cases, urine output is included. Baseline sCr concentration can be defined as the median of all sCr concentration measurements for a patient in the 180 days prior to an ED encounter, or as the sole sCr concentration measurement in 180 days prior to the ED encounter if only one measurement was recorded. If baseline sCr concentration was unavailable, expected baseline can be estimated using a CKD-EPI formula. For each patient, an initial AKI stage (0 [no AKI], 1, 2, or 3 can be determined by comparing a first sCr concentration measured during the index ED encounter against baseline sCr concentration. Peak sCr concentration within 72 hours of ED departure can be compared to a baseline to determine a highest AKI stage met during the outcome window. Outcomes can be flagged as positive if the peak AKI stage exceeds an initial stage, indicating AKI progression or the finding of any stage of new AKI. In some embodiments, separateAttorney Docket No : 44807-0484WO1 models can be developed to predict new or progressive AKI that meets criteria for moderate (e.g., stage 2) or severe (e.g., stage 3) AKI. If no follow-up sCr concentration was obtained within 72 hours of a patient’s ED departure, this outcome can be flagged as “missing.”

[0030] In some cases, data that is used for predicting AKI can be confined to data routinely collected and stored in the EHR during ED care delivery. Example data elements include patient demographics, comorbidities, chief complaints, initial triage vital signs, routine laboratory results, historical or imputed baseline sCr concentrations, and an initial AKI stage. Demographic data can be restricted to age (e.g., years) and sex (e.g., male or female). Comorbidities can be identified using ICD-10 codes. Chief complaints can be grouped into clinically meaningful categories using a previously defined schema. Triage vitals can include systolic and diastolic blood pressure, temperature, heart rate, respiratory rate, and oxygen saturation. In some cases, laboratory results can be used such as serum or plasma albumin, anion gap, blood urea nitrogen, sCr, glucose, potassium, sodium, lactate, troponin, white blood cell count, hemoglobin, platelets, and calculated anion gap and urine specific gravity. Vitals and laboratory results can be further processed to identify most recent, minimum, and maximum values obtained before, or concurrently with, the first metabolic panel.

[0031] In some examples, the machine learning model 130 includes one or more predictive models, such as XCGBOOST models. In some cases, separate XGBOOST algorithms can be trained to predict each outcome (e.g., whether any new or progressive AKI, whether AKI is likely to be moderate or severe). XGBOOST can, in some examples, represent an ensemble decision tree learning model that leverages sequential gradient descent optimization across decision tree iterations to minimize loss function and improve prediction accuracy, while also employing regularization techniques to avoid overfitting. XGBOOST can, in some examples, handle missing predictor data natively using a sparsity-aware split finding algorithm, making XGBOOST particularly suitable for prediction in contexts where variability exists in the collection of predictor data. In some examples, XGBOOST hyperparameter optimization can be conducted using a grid search to determine the optimal settings, including a subsample ratio of training instances, regularization terms, number of estimators, maximum depth, and learning rate parameters for column subsampling.Attorney Docket No : 44807-0484WO1

[0032] In some cases, data processing system 101 can address missing predictor data in one or more training datasets stored by database 105. For example, missing categorical data can include ED arrival mode, sex, chief complaint, and comorbidities. These can be assigned to a “null” category. Missing continuous values such as age, laboratory values, and vital signs can be retained as null and handled natively by XGBOOST. In cases where baseline sCr is missing, an estimated baseline can be calculated as described above in Outcomes and this imputed value can be used as a baseline.

[0033] Data processing system 101 can, in some examples, address missing outcome data in one or more training datasets stored by database 105. There are several example techniques for handling such missing outcome data. For example, in some embodiments only complete cases (e.g.. cases where repeat sCr was measured within 72 hours of ED departure) are included in the model training cohort for training the machine learning model 130. In some embodiments, every case is included in the model training cohort, and cases where sCr was not measured within 72 hours of ED departure can be assumed to be outcome negative (e.g., no new or progressive AKI).

[0034] In some embodiments, multiple imputation can be used to estimate unknown outcomes. For example, every case can be included in training, but observations from complete cases can be used to impute outcomes for incomplete cases. Under this approach, a preliminary' XGBoost model can be trained to estimate AKI outcome probability using data from complete cases including outcomes, then employed to generate an outcome probability' for each incomplete case missing an outcome. Probabilities can be used to impute 20 outcomes per incomplete case, and these outcomes can be randomly sampled to generate a full complete case training cohort with equal representation for all encounters.

[0035] In some embodiments, inverse probability weighting (IPW) is used to estimate unknown outcomes. For example, complete cases can be included in training, with each complete case being weighted based on a likelihood of outcome missingness. This approach can enable pseudorepresentation of incomplete cases. Complete cases where missingness of outcome data can be expected, but not observed, can be given more weight than cases where presence of outcome data was expected. Using an entire dataset (e.g., complete and incomplete cases, a separate XGBOOST model can be developed to estimate the likelihood of outcome missingness for each case. This model can be used to assign propensity scores to all complete cases, and these propensity scores can then be incorporated into AKI prediction models.Attorney Docket No : 44807-0484WO1

[0036] The data processing system 101 can generate a series of machine learning models to predict new or progressive AKI using clinical data collected from diverse EDs within a large health system, including machine learning model 130. In some cases, machine learning model 130 can be optimized to inform EHR-embedded decision support in the emergency care setting, where a range of clinical severity is wide, patients present at variable points along their trajectory of illness, and clinical decisions are made using incomplete information and under significant time pressure. To meaningfully inform decision-making at an earliest point in the care continuum, predictive tools can be capable of operating under these unique conditions.

[0037] A modeling approach used by data processing system 101 can be driven by usercentered design considerations. In some cases, machine learning model 130 can be limited to predictor data accumulated within hours of presentation to the healthcare system, and predictions were made at the time of an initial kidney function assessment, the timepoint most pertinent for emergency clinicians. Additionally, or alternatively, machine learning model 130 can integrate surveillance and prediction. Prior to generating a prediction, algorithms can surveil an EHR to calculate each patient’s baseline kidney function and determine whether the patient currently meets consensus criteria for AKI. In cases where historical data is not available, it is possible to determine whether current kidney function is below expected based on age and sex. By employing a composite outcome for prediction (new or progressive AKI), it can be possible to estimate ongoing risk for all patients, irrespective of AKI status at arrival. This combined approach can enable emergency clinicians to act promptly on existing AKI that might otherwise go unnoticed, while also anticipating potential deterioration of kidney function. This can prevent easily diagnosable AKI is from being overlooked and undertreated in both the inpatient and emergency care settings.

[0038] Data processing system 101 can generate machine learning model 130 based on consideration and inclusion of discharged patients. Some models can make reliable predictions in patients who are hospitalized after their ED encounter, do not consider discharged patients. Many kidney-relevant decisions (e.g., diagnostic evaluation and medication prescribing) can be made well in advance of disposition. This means that one or more AKI prediction models described herein (e.g., machine learning model 130) can be applicable to all ED patients to be useful for any. One modeling approach can enableAttorney Docket No : 44807-0484WO1 reliable predictions across a full spectrum of ED encounters and expand applicability to all of the approximately 140 million ED encounters that occur in the United States annually.

[0039] At least four different techniques for supplementing patient data to generate missing outcome data can be used. These include incomplete case exclusion, negative outcome assumption, multiple imputation, and inverse probability weighting. In some cases, explicitly accounting for missing outcomes can yield predicted probabilities more consistent with observed and estimated outcome rates. Notably, inverse probability weighting can emerge as a practical and effective solution, enabling robust model training while preserving performance across development and validation cohorts. This methodological rigor can underscore an importance of addressing missing outcomes to enhance the reliability and fairness of machine learning models in clinical practice.

[0040] In some embodiments, machine learning model 130 can use predictors routinely available during ED care, including demographic characteristics, chief complaints, vital signs, and laboratory' values to generate actionable risk estimates early in the clinical workflow. Provision of kidney-focused decision support in the ED can be beneficial for both high and low-acuity patients. The system including machine learning model 130 can help clinicians to establish treatment plans for a majority of hospitalized patients and provide definitive care to a larger number discharged into the community. In some cases, machine learning model 130 addresses key challenges such as missing outcome data and incorporating a diverse population of clinical encounters resulting in either hospitalization or discharge from the ED. By providing early and reliable risk estimates, machine learning model 130 can provide clinical decision-making and reduce the burden of AKI across varied patient populations.

[0041] FIG. 2 illustrates an example process 200 of predicting AKI using a machine learning model, according to some implementations. Process 200 utilizes the machine learning algorithm of multiple imputation to process original data 205, which is a combination of i) data 210 of patients with know n AKI diagnosis results (whether positive or negative), and ii) data 215 of patients without known AKI diagnosis results. Data 210 are considered complete cases while data 215 are considered incomplete cases. The demographics, comorbidities, vital signs, and / or laboratory results of the patients in each group are known to the machine learning model as predictors and are obtained from, e.g., health records stored in a database, such as database 105. As illustrated, each row of data 210 and 215 represents a patient. The X column of data 210 and 215 represents theAttorney Docket No : 44807-0484WO1 predictors and the Y column represents the AKI diagnosis result, with 0 indicating negative, 1 indicating positive, and “?” indicating unknown.

[0042] Data 205, 210, and 215 can be structured to encapsulate one or more values of the current renal function data of each patient. As discussed above, an example of the current renal function data is creatinine value. Accordingly, data 205, 210, and 215 can have one or more predictor fields to represent the numeric value of creatinine measurements of a patient. Data 205, 210, and 215 can have one or more other predictor fields to indicate, e.g., the age, sex, or priory kidney injury history, of the patient.

[0043] In process 200, complete case data 210 is used to derive and train a preliminary XGBOOST model 220. For each patient, model 220 is trained to predict the probability Y that the patient has new or progressing AKI based on information in the predictor fields corresponding to that patient. Applying trained model 220 to incomplete case data 215 as testing data, the probability Y for each patient in the incomplete case group ii) to have new or progressing AKI can be obtained as table 225. Model 220 can then compare the probability' Y for each patient with a threshold probability, which can be predetermined (e.g., 0.72 in the illustrated example). If the probability' Y for a patient is greater than the threshold probability', then the patient can be determined as AKI positive; otherwise the patient can be determined as AKI positive.

[0044] Model 220 can be used to perform multiple (e.g., 20) imputations on datasets having both complete case data and incomplete case data. Each imputation can result in a table 230 with probability values for the patients in the dataset and correspondingly a table 235 with binary AKI diagnosis results. After performing an imputation, a set of imputed data is generated. For example, after performing the 20 imputations, 20 sets of imputed data 240-1 to 240-20 are generated. The 20 sets of imputed data 240-1 to 240-20 can be combined as final imputed data 255, which can be used to develop and train a full XGBOOST model and evaluate the performance of the full XGBOOST model.

[0045] FIG. 3 illustrates another example process 300 of predicting AKI using a machine learning model, according to some implementations. Similar to process 200, process 300 uses original data 305, which is a combination of i) data 310 of patients with known AKI diagnosis results (whether positive or negative), and ii) data 315 of patients without known AKI diagnosis results. Data 310 are considered complete cases while data 315 are considered incomplete cases. Different from the multiple imputation algorithm in processAttorney Docket No : 44807-0484WO1200, process 300 utilizes the machine learning algorithm of inverse propensity weighting (IPW) to process original data 305.

[0046] As illustrated, IPW model 320, which can be an XGBOOST model, is trained with a sample set of training data obtained from original data 305. The trained IPW model 320 is capable to tell the propensity (e.g., probability7) of each patient having new or progressing AKI. The propensity can be represented by a score 325 ranging from 0 to 1. with 0 indicating full certainty of the patient being AKI negative and 1 indicating full certainty7of the patient being AKI positive.

[0047] Inverse propensity score 330 is then calculated for each patient. Based on inverse propensity score 330, a weight 345 is calculated. Weight 345 and complete case data 310 are used to develop a full XGBOOST model 335. The prediction performance of model 335 is evaluated at 340.

[0048] FIG. 4 illustrates a flowchart of an AKI detection process 400, according to some implementations. Process 400 can be implemented by a data processing system, e.g., system 100 of FIG. 1, with access to health records of patients of an emergency department (ED) to predict AKI of a patient.

[0049] At 401, the data processing system obtains the blood urea nitrogen (BUN) value of a patient. The BUN value can be measured when the patient arrives at the ED and the result can be ready within a few hours after the measurement. Using the measured BUN value, the data processing system can obtain the patient’s current serum creatinine (sCr) value as cunent renal function data.

[0050] At 402, the data processing system determines whether baseline renal function data, such as a baseline sCr value measured on the patient prior to the patient’s current admission to ED, is available in the health records of the patient. If Yes, then the patient’s prior sCr value is used as baseline renal function data. Otherwise, the data processing system calculates at 403 the baseline renal function data using the demographic information of the patient based on, e.g., a Chronic Kidney Disease Epidemiology (CKD- EPI) Creatinine Equation. Either way, the baseline renal function data is obtained at 404.

[0051] At 405, the baseline renal function data obtained at 404 is used for detecting AKI based on, e.g., AKI diagnosis criteria published by Kidney Disease: Improving Global Outcomes (KDIGO).Attorney Docket No : 44807-0484WO1

[0052] At 406, the data processing system obtains the sCr value of the patient from the blood urea nitrogen (BUN) measurement. The sCr value can be considered an initial sCr value at 407 after the patient’s admission to the ED.

[0053] Additionally, at 408 and 409, the data processing system attempts to obtain the sCr values measured within 72 hours after the patient is discharged from ED. If the sCr value is present at 409, then the peak sCr value during the period of measurement is obtained at 410. The peak sCr value and the initial sCr value can be used to detect any development of kidney function disorders at 405. If the sCr value is not present at 409, then the data processing system labels the patient’ s data as outcome missing at 411. The patient’s data can thus be used as an incomplete case in the training and testing of a machine learning model, such as the operations described with reference to FIGs. 2 and 3. In some implementations, the data processing system can be configured to assume that a patient with missing outcome is AKI negative. This assumption can be referred to as naive baseline. In some alternative implementations, the data processing system can be configured to exclude patients with missing outcomes from machine learning training datasets, which results in all data in the training dataset being complete cases.

[0054] With inputs from 405, 407, and 410, the data processing system at 405 can utilize a machine learning model, such as the machine learning model in FIG. 2 or 3, to determine a) at 412, the initial (e.g., at the time of ED arrival) stage of AKI of the patient, and / or b) at 413, the peak stage of AKI during the 72 hour period after the patient leaves the ED. The determination at 405 can involve, e.g., comparing the current renal function data with the baseline renal function data to determine a kidney disorder, and predicting the probability of the patient having AKI. For example, if it is determined at 414 that the peak stage is more severe than the initial stage, then the data processing system can consider that the patient has new or progressing AKI at 415. Otherwise, the data processing system can consider that the patient does not have new or progressing AKI at 416.

[0055] Implementations according to process 400 make an improvement to existing AKI diagnosis techniques by reducing the delay from the patient’s arrival at the ED to the availability of diagnosis results. The reduction of diagnosis delay is due to, among others, the obtaining of baseline renal function data and current renal function data within very short time. As discussed above, the baseline renal function data can be obtained from a database that communicates with the data processing system in real time. Also, the current renal function data of a patient can be based on BUN values and sCr values, theAttorney Docket No : 44807-0484WO1 measurements of which are routinely available in most EDs and the results of which are usually available within a few hours. These sources of baseline renal function data and current renal function data typically do not require complex data structures, and can be conveniently provided to train a machine learning model for predicting AKI without demanding for excessive computing resources. Using the trained machine learning model, healthcare professionals can timely evaluate the risk for a patient to develop AKI and take measures accordingly, even before AKI symptoms manifest. This could reduce the possibility of an AKI patient being discharged prematurely due to the failure of timely detecting upcoming AKI.

[0056] FIGs. 5A and 5B each illustrate a graph with example curves showing the performance of prediction, according to some implementations. Each of FIGs. 5 A and 5B has four curves showing the relationship between a false positive AKI prediction rate versus a true positive AKI prediction rate corresponding to a) a prediction method using the naive baseline, b) a prediction method using only complete cases, c) the multiple imputation algorithm, and d) the IPW algorithm. Each of predictions in a)-d) can be made using an external validation dataset. For example, the external validation dataset includes data from patients who visited one or more EDs over a past period of time, including who were admitted and who were discharged. The data include compete cases for which AKI outcomes 72 hours after the visits were known, and incomplete cases for which AKI outcomes 72 hours after the visits were unknown. After making predictions using random samples from the dataset, a data processing system can plot these curves and evaluate the AKI prediction performance based on an area-under-curve analysis. The curves in FIG. 5A represent prediction performances based on random samples from both complete and incomplete cases, whereas the curves in FIG. 5B represent prediction performances based on random samples from complete cases only.

[0057] From the curves in FIGs. 5A and 5B, it can be observed that the naive baseline method may deviate from other prediction methods in terms of performance. On the other hand, both the multiple imputation and the IPW algorithms show consistent prediction performances and thus can be used to develop a machine learning model for predicting AKI based on missing outcomes. Using the machine learning model and using outcome data from prior health encounters of patients, AKI diagnosis can be conducted more efficient and accurate for both hospitalized and discharged population, which improves upon the current emergency healthcare techniques.Attorney Docket No : 44807-0484WO1

[0058] FIG. 6 illustrates a flowchart of an example method, according to some implementations. For clarity’ of presentation, the description that follows generally describes method 600 in the context of the other figures in this description. It will be understood that method 600 can be performed, for example, by any suitable system, environment, software, hardware, or a combination of systems, environments, software, and hardware, as appropriate, such as system 100 of FIG. 1. One or more steps of method 600 can be substantially the same as or similar to the operations described with reference to FIGs. 1 -3. In some implementations, various steps of method 600 can be run in parallel, in combination, in loops, or in any order.

[0059] At 602, method 600 involves reading, from a hardware storage device of one or more external systems, digital health records each associated with a unique key value representing a patient, with each digital health record being structured with a plurality of fields, wherein the one or more external systems authenticate that access to the digital health records is authorized.

[0060] At 604, method 600 involves parsing the plurality of fields of the digital health records to identify the unique key values.

[0061] At 606, method 600 involves identifying, based on the parsing, a given key value and a digital health record for that given key value.

[0062] At 608, method 600 involves, for the identified digital health record corresponding to an identified patient, generating, by at least one processor, baseline renal function data based on the particular digital health record.

[0063] At 610, method 600 involves comparing the baseline renal function data with current renal function data associated with the given key value.

[0064] At 612, method 600 involves determining that a difference between the baseline renal function data and the current renal function data is greater than a threshold.

[0065] At 614, method 600 involves inputting the current renal function data to a machine learning model.

[0066] At 616, method 600 involves predicting, by the machine learning model, a propensity of acute kidney injury (AKI) for the given key value.

[0067] At 618, method 600 involves determining, by the at least one processor, a treatment of the patient based on the propensity7of AKI.

[0068] To further illustrate, AKI detection according to one or more implementations can be performed in operations 1-11 of the following example scenario.Attorney Docket No : 44807-0484WO1

[0069] 1. A patient arrives to the emergency department (ED).

[0070] 2. ED triage is perforated. During this process, data are collected and added to the medical record, such as: demographics, reason for visit, arrival mode, vital signs.

[0071] 3. These data are merged to a standing electronic medical record that already contains information about the patient: previously measured labs, medical history7.

[0072] 4. Additional actions may be taken on the patient by the clinical team including ordering and perforaiance of diagnostic tests, ordering and administration of treatments (e g., medications), repeat measurement of vital signs.

[0073] 5. When the first marker of kidney function (current creatinine level) is populated in the electronic health record, the algorithms are activated.

[0074] 6. The detection algorithm is activated first:

[0075] a. This algorithm surveils the electronic health record of the individual patient, and searches for every measure of kidney function (creatinine level) made in the preceding 6 months.

[0076] i. If a single measurement has been made, this is used as ‘baseline.'

[0077] ii. If multiple measurements have been made, the median of these is calculated and this is used as ‘baseline. ’

[0078] iii. If no measurements have been made, ‘expected baseline’ is imputed using consensus equations endorsed by the international group Kidney Disease Improving Global Outcomes (KDIGO) for calculation of normal kidney function.

[0079] b. Current kidney function (creatinine level) is compared to baseline above. Using consensus KDIGO criteria for AKI, patients are classified as having AKI or not having AKI. For patients who have AKI, they are staged by severity of AKI (stage 1, 2 or 3) using consensus KDIGO criteria.

[0080] c. If AKI status and stage determinations were made using an imputed‘expected baseline’ as described above, and indicator of uncertainty will be included, as we are unable to determine whether the abnormality7is truly acute or chronic.

[0081] 7. Next, the prediction algorithm is activated. This machine learning algorithm estimates the probability that each patient will develop AKI (or progress to a higher stage of AKI if already meeting criteria for stage 1 or 2 AKI) based on a wide array of information collected from the electronic health record. This information includes all data collected at ED triage (demographics, reason for visit, arrival mode, vital signs),Attorney Docket No : 44807-0484WO1 additional vital signs measured during the encounter, diagnostic test results and preexisting medical history.

[0082] 8. The algorithms send detection and prediction results to the electronic health record, which are used to drive kidney -protective clinical decision support. Patient status determinations can be as follows:

[0083] 9. Algorithm output is populated in the electronic health record.

[0084] 10. For patients identified as having new AKI, abnormal kidney function without definitive baseline, or high risk for new or progressive AKI within 72 hours, decision support is provided to promote nephroprotective treatment (eg, identify and reverse cause of AKI, optimize hemodynamic status and fluid balance, avoid further nephrotoxic exposures).

[0085] 11. The unified algorithm is triggered each time renal function (creatinine level) is measured and new calculations will be generated each time.

[0086] In some implementations, a tree-based machine learning (ML) model (e.g., XGBOOST) is trained using timestamped and matched predictor data (those described in #7 above) and outcome data (kidney function tests [creatinine] performed in the 72-hour outcome window) from several hundred thousand prior emergency department encounters. The ML model learns from subtle patterns in these data to make reliable predictions in future encounters. This is done through high-dimensional analysis of prior encounter data - wherein ML ‘learns’ about the importance of individual variables as predictors of new or progressive AKI and about the importance of complex relationships between variables. The specific variables that are used as predictors were chosen because they are routinelyAttorney Docket No : 44807-0484WO1 collected in the emergency department and have known relationships with kidney function based on inventors' understanding of human pathophysiology and previously published reports.

[0087] In some implementations, this AKI prediction model is adjusted to optimize prediction performance in the emergency department where many patients are expected to be discharged and could be under-represented in training data. These adjustments are achieved during training by calculating and applying 'weights’ (inverse probability weighting) to each patient in the training dataset. Weights (0.0-100.0) are determined based on the likelihood of predicted outcome missingness. Each patient in the training set is assigned a weight. This facilitates more reliable prediction across the entire cohort. Weights are generated by a separate tree-based ML prediction model that uses similar prediction variables to estimate the probability that the outcome will be missing (unmeasured). This is usually due to having been discharged.

[0088] FIG. 7 is a block diagram of an example computer system 700 in accordance with implementations of the present disclosure. The system 700 includes a processor 710, a memory 720, a storage device 730, and one or more input / output interface devices 740. Each of the components 710, 720, 730, and 740 can be interconnected, for example, using a system bus 750.

[0089] The processor 710 is capable of processing instructions for execution within the system 700. The term “execution” as used here refers to a technique in which program code causes a processor to cany out one or more processor instructions. In some implementations, the processor 710 is a single-threaded processor. In some implementations, the processor 710 is a multi-threaded processor. The processor 710 is capable of processing instructions stored in the memory 720 or on the storage device 730. The processor 710 may execute operations such as those described with reference to FIGs. 1-6 described herein.

[0090] The memory 720 stores information within the system 700. In some implementations, the memory 720 is a computer-readable medium. In some implementations, the memory 720 is a volatile memory’ unit. In some implementations, the memory 720 is a non-volatile memory unit.

[0091] The storage device 730 is capable of providing mass storage for the system 700. In some implementations, the storage device 730 is a non-transitory computer-readable medium. In various different implementations, the storage device 730 can include, forAttorney Docket No : 44807-0484WO1 example, a hard disk device, an optical disk device, a solid-state drive, a flash drive, magnetic tape, or some other large capacity storage device. In some implementations, the storage device 730 may be a cloud storage device, e.g., a logical storage device including one or more physical storage devices distributed on a network and accessed using a network. In some examples, the storage device may store long-tenn data. The input / output interface devices 740 provide input / output operations for the system 700. In some implementations, the input / output interface devices 740 can include one or more of a network interface devices, e.g., an Ethernet interface, a serial communication device, e g., an RS-232 interface, and / or a wireless interface device, e.g., an 702. 11 interface, a 3G wireless modem, a 4G wireless modem, a 5G wireless modem, etc. A network interface device allows the system 700 to communicate, for example, transmit and receive data. In some implementations, the input / output device can include driver devices configured to receive input data and send output data to other input / output devices, e.g., keyboard, printer and display devices 770. In some implementations, mobile computing devices, mobile communication devices, and other devices can be used.

[0092] A server can be distributively implemented over a network, such as a server farm, or a set of widely distributed servers or can be implemented in a single virtual device that includes multiple distributed devices that operate in coordination with one another. For example, one of the devices can control the other devices, or the devices may operate under a set of coordinated rules or protocols, or the devices may be coordinated in another fashion. The coordinated operation of the multiple distributed devices presents the appearance of operating as a single device.

[0093] In some examples, the system 700 is contained within a single integrated circuit package. A system 700 of this kind, in which both a processor 710 and one or more other components are contained within a single integrated circuit package and / or fabricated as a single integrated circuit, is sometimes called a microcontroller. In some implementations, the integrated circuit package includes pins that correspond to input / output ports, e.g., that can be used to communicate signals to and from one or more of the input / output interface devices 740.

[0094] Although an example processing system has been described in FIG. 7, implementations of the subject matter and the functional operations described in this specification can be implemented in digital electronic circuitry, in tangibly embodied computer software or firmware, in computer hardware, including the structures disclosedAttorney Docket No : 44807-0484WO1 in this specification and their structural equivalents, or in combinations of one or more of them. Software implementations of the described subject matter can be implemented as one or more computer programs. Each computer program can include one or more modules of computer program instructions encoded on a tangible, non-transitory, computer-readable computer-storage medium for execution by, or to control the operation of, data processing apparatus. Alternatively, or additionally, the program instructions can be encoded in / on an artificially generated propagated signal. In an example, the signal can be a machine-generated electrical, optical, or electromagnetic signal that is generated to encode information for transmission to a suitable receiver apparatus for execution by a data processing apparatus. The computer-storage medium can be a machine-readable storage device, a machine-readable storage substrate, a random or serial access memory device, or a combination of computer-storage mediums.

[0095] The terms “data processing apparatus,” “computer,” and “computing device” (or equivalent as understood by one of ordinary skill in the art) refer to data processing hardware. For example, a data processing apparatus can encompass all kinds of apparatus, devices, and machines for processing data, including by way of example, a programmable processor, a computer, or multiple processors or computers. The apparatus can also include special purpose logic circuitry’ including, for example, a central processing unit (CPU), a field programmable gate array (FPGA), or an application specific integrated circuit (ASIC). In some implementations, the data processing apparatus or special purpose logic circuitry (or a combination of the data processing apparatus or special purpose logic circuitry ) can be hardware- or software-based (or a combination of both hardware- and software-based). The apparatus can optionally include code that creates an execution environment for computer programs, for example, code that constitutes processor fimiware, a protocol stack, a database management system, an operating system, or a combination of execution environments. The present disclosure contemplates the use of data processing apparatuses with or without conventional operating systems, for example LINUX, UNIX, WINDOWS, MAC OS, ANDROID, or IOS.

[0096] A computer program, which can also be referred to or described as a program, software, a software application, a module, a software module, a script, or code, can be written in any form of programming language. Programming languages can include, for example, compiled languages, interpreted languages, declarative languages, or procedural languages. Programs can be deployed in any form, including as standalone programs,Attorney Docket No : 44807-0484WO1 modules, components, subroutines, or units for use in a computing environment. A computer program can, but need not, correspond to a file in a file system. A program can be stored in a portion of a file that holds other programs or data, for example, one or more scripts stored in a markup language document, in a single file dedicated to the program in question, or in multiple coordinated files storing one or more modules, sub programs, or portions of code. A computer program can be deployed for execution on one computer or on multiple computers that are located, for example, at one site or distributed across multiple sites that are interconnected by a communication network. While portions of the programs illustrated in the various figures may be shown as individual modules that implement the various features and functionality through various objects, methods, or processes, the programs can instead include a number of sub-modules, third-party services, components, and libraries. Conversely, the features and functionality of various components can be combined into single components as appropriate. Thresholds used to make computational determinations can be statically, dynamically, or both statically and dynamically detennined.

[0097] The methods, processes, or logic flows described in this specification can be performed by one or more programmable computers executing one or more computer programs to perform functions by operating on input data and generating output. The methods, processes, or logic flows can also be performed by, and apparatus can also be implemented as. special purpose logic circuitry, for example, a CPU, an FPGA, or an ASIC.

[0098] Computers suitable for the execution of a computer program can be based on one or more of general and special purpose microprocessors and other kinds of CPUs. The elements of a computer are a CPU for performing or executing instructions and one or more memory devices for storing instructions and data. Generally, a CPU can receive instructions and data from (and write data to) a memory. A computer can also include, or be operatively coupled to, one or more mass storage devices for storing data. In some implementations, a computer can receive data from, and transfer data to, the mass storage devices including, for example, magnetic, magneto optical disks, or optical disks. Moreover, a computer can be embedded in another device, for example, a mobile telephone, a personal digital assistant (PDA), a mobile audio or video player, a game console, a GNSS sensor or receiver, or a portable storage device such as a universal serial bus (USB) flash drive.Attorney Docket No : 44807-0484WO1

[0099] The term “computer-readable medium” includes, but is not limited to, portable or non-portable storage devices, optical storage devices, and various other mediums capable of storing, containing, or carrying instruction(s) and / or data. A computer-readable medium may include a non-transitory medium in which data can be stored and that does not include carrier waves and / or transitory7electronic signals propagating wirelessly or over wired connections. Examples of a non-transitory medium may include, but are not limited to, a magnetic disk or tape, optical storage media such as compact disk (CD) or digital versatile disk (DVD), flash memory, memory or memory devices. A computer-readable medium may have stored thereon code and / or machine-executable instructions that may represent a procedure, a function, a subprogram, a program, a routine, a subroutine, a module, a software package, a class, or any combination of instructions, data structures, or program statements. A code segment may be coupled to another code segment or a hardware circuit by passing and / or receiving information, data, arguments, parameters, or memory contents. Information, arguments, parameters, data, etc. may be passed, forw arded, or transmitted via any suitable means including memory sharing, message passing, token passing, network transmission, or the like.

[0100] Computer readable media (transitory or non-transitory, as appropriate) suitable for storing computer program instructions and data can include all forms of permanent / non- permanent and volatile / non-volatile memory, media, and memory devices. Computer readable media can include, for example, semiconductor memory devices such as random access memory (RAM), read only memory (ROM), phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), and flash memory7devices. Computer readable media can also include, for example, magnetic devices such as tapes, cartridges, cassettes, and intemal / removable disks. Computer readable media can also include magneto optical disks and optical memory7devices and technologies including, for example, digital video disc (DVD), CD ROM, DVD+ / -R. DVD-RAM, DVD-ROM, HD-DVD, and BLURAY. The memory can store various objects or data, including caches, classes, frameworks, applications, modules, backup data, jobs, web pages, web page templates, data structures, database tables, repositories, and dynamic information. Types of objects and data stored in memory can include parameters, variables, algorithms, instructions, rules, constraints, and references. Additionally, the memory can include logs, policies, security or access data,Attorney Docket No : 44807-0484WO1 and reporting files. The processor and the memory can be supplemented by, or incorporated in, special purpose logic circuitry.

[0101] While this specification includes many specific implementation details, these should not be construed as limitations on the scope of what may be claimed, but rather as descriptions of features that may be specific to particular implementations. Certain features that are described in this specification in the context of separate implementations can also be implemented, in combination, in a single implementation. Conversely, various features that are described in the context of a single implementation can also be implemented in multiple implementations, separately, or in any suitable sub-combination. Moreover, although previously described features may be described as acting in certain combinations and even initially claimed as such, one or more features from a claimed combination can, in some cases, be excised from the combination, and the claimed combination may be directed to a sub-combination or variation of a sub-combination.

[0102] Particular implementations of the subject matter have been described. Other implementations, alterations, and permutations of the described implementations are within the scope of the following claims as will be apparent to those skilled in the art. While operations are depicted in the drawings or claims in a particular order, this should not be understood as requiring that such operations be performed in the particular order shown or in sequential order, or that all illustrated operations be performed (some operations may be considered optional), to achieve desirable results. In certain circumstances, multitasking or parallel processing (or a combination of multitasking and parallel processing) may be advantageous and performed as deemed appropriate.

[0103] Moreover, the separation or integration of various system modules and components in the previously described implementations should not be understood as requiring such separation or integration in all implementations, and it should be understood that the described program components and systems can generally be integrated together in a single software product or packaged into multiple software products.

[0104] Accordingly, the previously described example implementations do not define or constrain the present disclosure. Other changes, substitutions, and alterations are also possible without departing from the spirit and scope of the present disclosure.

Claims

Attorney Docket No : 44807-0484WO1CLAIMSWhat is claimed is:

1. A system for processing digital records, the system comprising: a digital network in communication with one or more external systems; at least one processor; and a memory subsystem communicatively coupled to the at least one processor, the memory subsystem storing instructions which, when executed by the at least one processor, cause the at least one processor to perform operations comprising: reading, by the digital network from a hardware storage device of the one or more external systems, digital health records each associated with a unique key value representing a patient, with each digital health record being structured with a plurality of fields, wherein the one or more external systems authenticate that the system is authorized to access the digital health records; parsing, by a parser, the plurality of fields of the digital health records to identify the unique key values; identifying, based on the parsing, a given key value and a digital health record for that given key value; for the identified digital health record corresponding to an identified patient: generating, by the at least one processor, baseline renal function data based on the particular digital health record; comparing the baseline renal function data with current renal function data associated with the given key value; determining that a difference between the baseline renal function data and the current renal function data is greater than a threshold; inputting the current renal function data to a machine learning model; predicting, by the machine learning model, a propensity of acute kidney injury (AKI) for the given key value; and determining, by the at least one processor, a treatment of the identified patient based on the propensity of AKI.Attorney Docket No : 44807-0484WO12. The system of claim 1, wherein the identified digital health record comprises prior renal function data associated with the given key value, and wherein generating the baseline renal function data comprises extracting the prior renal function data from the plurality7of fields of the identified digital health record.

3. The system of claim 1, wherein the identified digital health record does not comprise prior renal function data associated with the given key value, and wherein generating the baseline renal function data comprises: extracting demographic information associated with the given key value from the plurality of fields of the identified digital health record; and obtaining the baseline renal function data from the one or more external systems based on the demographic information.

4. The system of claim 1, wherein the machine learning model comprises an Extreme Gradient Boosting (XGBOOST) model.

5. The system of claim 4, further comprising obtaining, from the one or more external systems, a set of digital health records of a plurality of patients, wherein the set of digital health data comprises a first set with AKI diagnosis results and a second set without AKI diagnosis results.

6. The system of claim 5, the operations further comprising: deriving a preliminary XGBOOST model based on the first set; generating a set of imputed data based on the second set using the preliminaryXGBOOST model; and training the XGBOOST model using the set of imputed data and the second set.

7. The system of claim 5, further comprising: determining a plurality' of inverse propensity scores based on the second set; calculating a plurality of weights based on the plurality' of inverse propensity7scores; and training the XGBOOST model using the first set and the plurality7of weights.Attorney Docket No : 44807-0484WO18. The system of claim 1, wherein the current renal function data comprises at least one of: an albuminuria value of the patient, a blood urea nitrogen (BUN) value of the patient, or a creatinine value of the patient,9. The system of claim 1 , wherein the identified digital health record comprises at least one of: demographic information of the identified patient, an active medical problem of the identified patient, a chief complaint of the identified patient, or AKI stage information of the identified patient.

10. The system of claim 1, the operations further comprising: plotting a curve showing a performance of prediction of the machine learning model; and evaluating the performance based on an area-under-the-curve analysis.

11. The system of claim 1, wherein the digital health records are encrypted.

12. The system of claim 1, further comprising an application programming interfaces (APIs) that communicatively couples the at least one processor to the digital network.

13. A method for processing digital records, the method compnsing: reading, from a hardware storage device of one or more external systems, digital health records each associated with a unique key value representing a patient, with each digital health record being structured with a plurality of fields, wherein the one or more external systems authenticate that access to the digital health records is authorized; parsing the plurality of fields of the digital health records to identify the unique key values; identifying, based on the parsing, a given key value and a digital health record for that given key value;Attorney Docket No : 44807-0484WO1 for the identified digital health record corresponding to an identified patient: generating, by at least one processor, baseline renal function data based on the particular digital health record; comparing the baseline renal function data with current renal function data associated with the given key value; determining that a difference between the baseline renal function data and the current renal function data is greater than a threshold; inputting the current renal function data to a machine learning model; predicting, by the machine learning model, a propensity of acute kidney injury (AKI) for the given key value; and determining, by the at least one processor, a treatment of the patient based on the propensity of AKI.

14. The method of claim 13, wherein the identified digital health record comprises prior renal function data associated with the given key value, and wherein generating the baseline renal function data comprises extracting the prior renal function data from the plurality of fields of the identified digital health record.

15. The method of claim 13, wherein the identified digital health record does not comprise prior renal function data associated with the given key value, and wherein generating the baseline renal function data comprises: extracting demographic information associated with the given key value from the plurality of fields of the identified digital health record; and obtaining the baseline renal function data from the one or more external systems based on the demographic information.

16. The method of claim 13, wherein the machine learning model comprises an Extreme Gradient Boosting (XGBOOST) model.

17. The method of claim 16, further comprising obtaining, from the one or more external systems, a set of digital health records of a plurality of patients, wherein the set of digital health data comprises a first set with AKI diagnosis results and a second set without AKI diagnosis results.Attorney Docket No : 44807-0484WO118. The method of claim 17, further comprising: deriving a preliminary XGBOOST model based on the first set; generating a set of imputed data based on the second set using the preliminary XGBOOST model; and training the XGBOOST model using the set of imputed data and the second set.

19. The method of claim 17, the further comprising: determining a plurality' of inverse propensity scores based on the second set; calculating a plurality' of weights based on the plurality7of inverse propensity7scores; and training the XGBOOST model using the first set and the plurality7of weights.

20. The method of claim 13, further comprising: plotting a curve showing a performance of prediction of the machine learning model; and evaluating the performance based on an area-under-the-curve analysis.

21. A computer-readable medium comprising instructions that, when executed by a processor, cause the processor to: read, from a hardware storage device of one or more external systems, digital health records each associated with a unique key value representing a patient, with each digital health record being structured with a plurality' of fields, wherein the one or more external systems authenticate that access to the digital health records is authorized; parse the plurality of fields of the digital health records to identify the unique keyvalues; identify, based on the parsing, a given key value and a digital health record for that given key value; for the identified digital health record corresponding to an identified patient: generate, by at least one processor, baseline renal function data based on the particular digital health record; compare the baseline renal function data with current renal function data associated with the given key value;Attorney Docket No : 44807-0484WO1 determine that a difference between the baseline renal function data and the current renal function data is greater than a threshold; input the current renal function data to a machine learning model; predict, by the machine learning model, a propensity of acute kidney injury(AKI) for the given key value; and determine, by the at least one processor, a treatment of the patient based on the propensity of AKI.