Systems and methods for predicting renal function decline
Patent Information
- Application Number
- JP2024510262
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2021-08-18
- Filing Date
- 2022-08-17
- Publication Date
- 2025-07-23
AI Technical Summary
Current methods for predicting chronic kidney disease (CKD) progression are limited, particularly in early stages, and lack accuracy and applicability across all stages of CKD, missing opportunities for early intervention and resource allocation.
Development of machine learning models, such as random survival forest models, trained on comprehensive medical laboratory data to predict CKD progression, including a 40% decline in eGFR or renal failure, applicable from early stages (G1-G5) using electronic health records, enabling personalized risk assessment and treatment guidance.
The models provide accurate, personalized predictions of CKD progression, facilitating timely interventions and resource allocation, improving patient outcomes and healthcare efficiency by identifying high-risk individuals for early-stage CKD.
Smart Images

Figure 00000000_0000_ABST
Abstract
Description
[Background technology]
[0001] Cross-reference to related applications
[0001] This application claims priority to U.S. Provisional Patent Application No. 63 / 234,535, filed August 18, 2021, entitled "SYSTEMS AND METHODS FOR PREDICTING KIDNEY FUNCTION DECLINE," which is incorporated herein by reference in its entirety.
[0002]
[0002] Chronic kidney disease (CKD) currently affects more than 850 million adults worldwide and is associated with high morbidity and mortality, as well as high medical costs. To illustrate, in 2009, the treatment of CKD, e.g., end-stage renal disease (ESRD), required $40 billion in costs in the United States alone. Although only a small proportion of CKD patients progress to renal failure, much of the excess morbidity and costs associated with CKD are driven by individuals who progress to more advanced stages of CKD before reaching organ failure requiring dialysis.
[0003]
[0003] Providing resource-efficient and appropriate treatment for CKD patients benefits those suffering from the disease and improves resource allocation in an increasingly burdened health care system. Accurate prediction of an individual's risk of CKD progression can improve patient experience and outcomes through shared knowledge and shared decision-making, enhance care by better matching the risks and hazards of treatment to the risk of disease progression, and / or improve the efficiency of the health system by driving better alignment between resource allocation and individual risk. Summary of the Invention
[0004]
[0004] Therefore, there is a need for improved techniques for predicting the risk of progression to CKD in an individualized manner. [Brief description of the drawings]
[0005]
[0005] In order to explain the manner in which the above-cited advantages and features, as well as other advantages and features, can be obtained, the subject matter briefly described above will now be more particularly described with reference to specific embodiments which are illustrated in the accompanying drawings, the embodiments being more specifically and in detail described and explained through the use of the accompanying drawings, with the understanding that these drawings merely illustrate exemplary embodiments and therefore should not be considered as limiting the scope thereof. [Figure 1]
[0006] An example computing environment, including an example computing system that can incorporate and / or be used to implement the disclosed embodiments, is illustrated. [Diagram 2]
[0007] 1 illustrates a conceptual representation of an example of a machine learning model trained on a training data set that includes medical laboratory data and configured to generate a prediction of progression of chronic kidney disease. [Figure 3A]
[0008] 1 shows an example flow diagram depicting acts associated with generating a chronic kidney disease progression prediction. [Figure 3B] 1 shows an example flow diagram depicting acts associated with generating a chronic kidney disease progression prediction. [Figure 3C] 1 shows an example flow diagram depicting acts associated with generating a chronic kidney disease progression prediction. [Figure 3D] 1 shows an example flow diagram depicting acts associated with generating a chronic kidney disease progression prediction. [Figure 4]
[0009] Here are some reported examples related to predicting the progression of chronic kidney disease. [Diagram 5]
[0010] Schematic of an example cohort of patients from which a machine learning model is generated to train a dataset. [Figure 6A]
[0011] A table is shown containing a description of the baseline cohort sample, including the various test results included in the medical laboratory data for each patient. [Figure 6B]
[0012] A table containing an overview of variable missingness in the baseline cohort as described in Figure 6A is shown. [Figure 7]
[0013] An excerpt of tariff codes used to regulate dialysis and kidney transplantation. [Figure 8]
[0014] 1 is a table outlining the variable importance for each variable included in a machine learning model training a dataset. [Figure 9]
[0015] 1 shows a conceptual representation of an example training data set comprising a 10-variable medical laboratory data set. [Figure 10]
[0016] 13 is a graph showing an example calibration plot for a machine learning model configured as a random forest model (e.g., using a training data set as shown in FIG. 9 for a two-year time period). [Figure 11]
[0017] FIG. 13 is a graph showing an example calibration plot for a machine learning model configured as a random forest model (e.g., using a training data set as shown in FIG. 9 for a 5-year time period). [Figure 12]
[0018] FIG. 13 shows an example calibration plot for a machine learning model configured as a Cox model (e.g., for a 2-year time period, using a training data set as shown in FIG. 9). [Figure 13]
[0019] FIG. 13 shows an example calibration plot for a machine learning model configured as a Cox model (e.g., for a 5-year time period, using a training data set as shown in FIG. 9). [Figure 14]
[0020] 1 illustrates an example of machine learning trained on a training data set that includes nine-variable medical laboratory data and is configured to generate a prediction of chronic kidney disease progression. [Figure 15]
[0021] FIG. 15 is a graph illustrating an example calibration plot for a machine learning model configured as a Cox model using the training data set shown in FIG. 14, for a time period of 2 years. [Figure 16]
[0022] FIG. 15 is a graph illustrating an example calibration plot for a machine learning model configured as a Cox model using the training data set shown in FIG. 14, for a time period of 5 years. [Figure 17]
[0023] An example of a training data set is shown, which includes a 16-22 variable medical laboratory data set. [Figure 18]
[0024] For example, a graph showing an example calibration plot for a machine learning model using a training data set as shown in FIG. 17 for a two-year time period is shown. [Figure 19] For example, a graph showing an example calibration plot for a machine learning model using a training data set as shown in FIG. 17 for a two-year time period is shown. [Figure 20] For example, a graph showing an example calibration plot for a machine learning model using a training data set as shown in FIG. 17 for a two-year time period is shown. [Figure 21]
[0025] An example of a training data set is shown, which includes at least a 15-variable medical laboratory data set. [Figure 22]
[0026] FIG. 22 is a graph showing an example calibration plot for a machine learning model using a training data set such as that shown in FIG. 21, for a two-year time period. [Diagram 23]
[0027] FIG. 24 is a graph showing an example calibration plot for a machine learning model using a training data set such as that shown in FIG. 23, for a five-year time period. [Figure 24]
[0028] FIG. 1 shows a table illustrating one summary example of performance evaluation statistics for various examples of machine learning models disclosed herein and configured as Cox models. [Diagram 25]
[0029] FIG. 1 shows calibration plots for various examples of machine learning models disclosed herein and configured as Cox models. [Figure 26A]
[0030] A table showing various summary examples of performance evaluation statistics for various examples of machine learning models configured as Random Forest models is shown. [Figure 26B] FIG. 26B shows a table depicting various summary examples of performance evaluation statistics for various examples of machine learning models configured as random forest models. [Figure 27A]
[0031] FIG. 13 shows an example calibration plot for a Random Forest model in a subgroup analysis of patients with diabetes. [Figure 27B]
[0032] FIG. 13 shows an example calibration plot for a Random Forest model in a subgroup analysis of non-diabetic patients. [Figure 27C]
[0033] FIG. 1 shows an example of a calibration plot for a random forest model in a subgroup analysis of CKD patients at different stages of disease. [Figure 27D] Figure 27D shows an example calibration plot for the Random Forest model in a subgroup analysis of CKD patients at different stages of disease. [Figure 28]
[0034] FIG. 1 illustrates aspects of a validation cohort used to externally validate an example of a Random Survival Forest model that generates CKD progression predictions. [Figure 29]
[0035] Figure 1 shows an overview of the degree of missingness for the laboratory panel used to develop an example random survival forest model that generates CKD progression predictions. [Diagram 30]
[0036] Tariff codes identifying dialysis and transplantation are outlined to generate a training data set for constructing an example of a random survival forest model that generates CKD progression predictions. [Diagram 31]
[0037] Variable importance is shown for an example of a 22-variable survival forest generating CKD progression predictions. [Diagram 32]
[0038] We provide an overview of baseline descriptive statistics for the training cohort, internal testing cohort, and external validation cohort for constructing an example random survival forest model generating CKD progression predictions. [Diagram 33]
[0039] Figure 1 shows the AUC and Brier scores at 1 to 5 years for an example of a random survival forest model with 22 variables to generate a CKD progression prediction. [Diagram 34]
[0040] AUC and Brier scores for the internal testing and external validation cohorts of an example random survival forest model with 22 variables to generate CKD progression predictions. [Figure 35A]
[0041] Figure 2 shows various calibration charts for an example of a Random Survival Forest model with 22 variables to generate CKD progression predictions in year 2. [Figure 35B] FIG. 35B shows various calibration charts for an example of a Random Survival Forest model with 22 variables to generate CKD progression predictions in the second year. [Diagram 36]
[0042] Figure 1 shows a performance summary of an example of a Random Survival Forest model with 22 variables for generating CKD progression predictions. [Figure 37A]
[0043] Figure 1 shows various calibration charts for an example of a Random Survival Forest model with 22 variables to generate CKD progression predictions at 5 years. [Figure 37B] Figure 1 shows various calibration charts for an example of a Random Survival Forest model with 22 variables to generate CKD progression predictions at 5 years. [Figure 38]
[0044] We present the results of a heapmap model that generates CKD progression predictions. [Figure 39]
[0045] We present the results of a medical model that generates CKD progression predictions. DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS
[0006]
[0046] The disclosed embodiments are directed to improvements in systems, methods, and / or frameworks for training and / or utilizing machine learning models to predict CKD progression and / or guide practitioners in medical decisions for patients at risk for CKD progression.
[0007]
[0047] The Kidney Failure Risk Equation (KFRE) is an internationally validated risk prediction method to predict the risk of progression of kidney failure for individual CKD patients. However, KFRE has important limitations, such as being applicable only to late CKD stages (G3-G5) and considering only the clinical outcome of kidney failure requiring dialysis. In early stages of CKD, kidney failure is a rare event, even though progression to more advanced stages is not rare. In these early stages, a 40% decline in GFR is medically meaningful to both patients and physicians, and allows sponsors to design feasible randomized controlled trials in all stages of CKD.
[0008]
[0048] In addition, new disease-modifying therapies for CKD that slow progression are available, but they have been primarily studied in patients with preserved renal function. The use of these therapies may be beneficial, especially in individuals with high-risk early CKD, where the benefit of dialysis prevention can be substantial and cost-effective. To apply disease-modifying therapies for CKD to individuals with high-risk early CKD, models can be implemented that predict a 40% decline in eGFR, or a composite outcome of renal failure or a 40% decline in eGFR, and can be applied to patients in all stages of CKD (G1-G5). When such models are based on laboratory data, they can be used through electronic health records or laboratory information systems and are not subject to the coding variability and complexity often seen in CKD. At least some of the disclosed embodiments involve the derivation and external validation of a novel laboratory-based machine learning predictive model that accurately predicts a 40% decline in eGFR or renal failure in patients (e.g., CKD G1-G5 patients). Technical advantages
[0049] The disclosed embodiments facilitate various technical advantages over existing systems and methods related to CKD progression prediction, particularly in being able to predict chronic kidney disease (CKD) progression for patients at any stage of CKD (or for patients without CKD or with unknown CKD status). Additionally, predictions generated by the present disclosure may be based on a composite outcome (e.g., not just renal failure) of either a 40% decline in eGFR and / or renal failure. Predictions generated according to at least some embodiments of the present disclosure may provide a risk score for patients experiencing either outcome.
[0009]
[0050] In patients with CKD, the disclosed methods can be used to inform a variety of important medical decisions, such as, by way of non-limiting example, informing renal referral triage, assessing the need for more intensive clinical care, determining timing of modality education, dialysis access planning, and / or other. The disclosed embodiments generate CKD progression predictions and can be implemented in a variety of ways, such as to generate CKD progression predictions for individual patients (e.g., when implemented in an electronic health record or linked software solution and / or in response to an individual physician request) and / or to facilitate batch processing of patients in a patient database (e.g., a hospital or clinic database).
[0010]
[0051] At least some of the disclosed embodiments include models predicting individual outcomes (risk of 40% decline in eGFR or risk of renal failure) or composite outcomes (risk of developing renal failure or 40% decline in eGFR), which can be applied to patients screened for or at all stages of CKD (G1-G5). Systems and / or methods providing such features are urgently needed. At least some of the models disclosed herein can be utilized to risk stratify patients with early stage disease (G1-G3) who are at high risk for CKD progression, to inform enrollment of patients (at any CKD stage) in clinical trials, and / or to guide the administration of treatments that can modify disease progression, such as sodium-glucose cotransporter-2 (SGLT2) inhibitors or mineralocorticoid receptor antagonists (MRAs). Systems and techniques for predicting CKD progression
[0052] Attention is now directed to Figure 1, which illustrates example components of a computing system 110 that may include and / or be used to implement aspects of the disclosed invention. Figure 1 illustrates various machine learning (ML) modules and data types associated with the inputs and outputs of the machine learning models.
[0011]
[0053] As used herein, a machine learning model or module refers to any combination of software and / or hardware components operable to facilitate processing using a machine learning model or other artificial intelligence-based structure / architecture. For example, one or more processors may be configured to run a variety of neural networks, including, but not limited to, a random forest model, a random survival forest model, a Cox proportional hazards model, a single layer neural network, a feed forward neural network, a radial basis function network, a deep feed-forward network, a recurrent neural network, a long-short term memory (LSTM), a sparse neural network, a sparse ... The present invention may include and / or utilize hardware components and / or computer executable instructions operable to execute functional blocks and / or processing layers configured in the form of: a deep memory network, a gated recurrent unit, an autoencoder neural network, a variational autoencoder, a denoising autoencoder, a sparse autoencoder, a Markov chain, a Hopfield neural network, a Boltzmann machine network, a constrained Boltzmann machine network, a deep belief network, a deep convolutional network (or convolutional neural network), a deconvolutional neural network, a deep convolutional inverse neural network, a generative adversarial network, a liquid state machine, an extreme learning machine, an echo state network, a deep residual network, a Kohonen network, a support vector machine, a neural Turing machine, and / or others.
[0012]
[0054] The example shown in FIG. 1 illustrates a computing system 110 as part of a computing environment 100, which may include a third party system(s) 120 in communication with the computing system 110 (through a network 130). In one embodiment, the computing system 110 is configured to train and / or configure a machine learning model (e.g., a CDK predictive model) to generate a CKD progression prediction for one or more patients. The machine learning model may additionally or alternatively be trained / configured to generate treatment, monitoring, or other caring recommendations for one or more patients. The computing system 110 of FIG. 1 may additionally or alternatively be configured to run a machine learning model, such as a CKD predictive model trained / configured as described herein.
[0013]
[0055] The computing system 110 of FIG. 1 includes one or more processor(s) (such as one or more hardware processor(s)) 112 and storage (i.e., hardware storage device(s) 140) that stores computer readable instructions 118. The hardware storage device(s) 140 may contain any number of data types and any number of computer readable instructions 118 such that the computing system 110 is configured to implement one or more of the disclosed embodiments when the computer readable instructions 118 are executed by the one or more processor(s) 112. The hardware storage device(s) 140 may also include physical, tangible storage means. The computing system 110 is also shown to include user interface(s) 114 and input / output (I / O) device(s) 116.
[0014]
[0056] As shown in FIG. 1, the hardware storage device(s) 140 are shown as one storage unit. However, it will be appreciated that the hardware storage device(s) 140 may be implemented as distributed storage and distributed across various separate and sometimes distant systems and / or third party system(s) 120. The computing system 110 may also comprise a distributed system, where one or more of the components of the computing system 110 may be separate from one another and are maintained / run by different discrete systems, each performing different tasks. In one instance, multiple distributed systems perform similar and / or shared tasks to implement the disclosed functionality, such as in a distributed cloud environment.
[0015]
[0057] In the example of Figure 1, hardware storage device(s) 140 can store different data types including training data set 141, medical laboratory data 142, patient information 143, and CKD progression prediction data 144. As shown in Figure 1, the storage (e.g., hardware storage device(s) 140) can include computer readable instructions 118. The computer readable instructions 118 may be usable to facilitate training / configuring and / or executing (e.g., for generating CKD progression predictions) one or more of the models and / or modules (e.g., machine learning model 145) shown in Figure 1.
[0016]
[0058] The machine learning model 145 can be trained using a training data set 141. The training data set 141 can include medical laboratory data for a cohort of patients (e.g., included in the medical laboratory data 142) and / or other patient information (e.g., included in the patient information 143). The training data set 141 can be applied to the machine learning model (e.g., the machine learning model 145) to train the machine learning to generate a CKD progression prediction. In an embodiment, the training data set 141 includes (i) a first set of medical laboratory data related to a plurality of patients, (ii) an age of each patient included in the plurality of patients, and (iii) a gender of each patient included in the plurality of patients. The first set of medical laboratory data can include a variety of labs / measurements associated with a particular patient, including, but not limited to, estimated glomerular filtration rate (eGFR), urinary albumin / creatinine ratio (ACR), urea, serum sodium, serum chloride, serum hemoglobin, serum potassium, glucose, serum albumin, alkaline phosphatase, serum phosphate, serum bicarbonate, serum magnesium, serum calcium, aspartate aminotransferase (AST), alanine aminotransaminase (ALT), bilirubin, gamma-glutamyl transferase (GGT), hematocrit, platelet count, and / or others.
[0017]
[0059] Various laboratory data / measurements associated with various patients included in the training cohort can be collected (or may have already been collected) at one or more time points or over one or more time periods (e.g., from samples or measurements obtained from each individual patient during one or more patient-physician interactions over time, such as over multiple consecutive medical appointments to obtain a series of samples or measurements over a period of time (e.g., a week, a month, etc.)). For example, various clinical tests are ordered for the patient on the first day of a visit with the physician. As another example, the patient may provide one or more blood test results on the first day and then submit a urine sample for testing on another day. Alternatively, a particular test may require samples from multiple days over a time period of a week or a month, or even a year.
[0018]
[0060] In one embodiment, one time point is used for each set of lab values included in the training and / or testing data. For example, in one instance, the time point is defined by the eGFR lab measurement, and all other lab values are selected from labs within 365 days of eGFR lab measurements.
[0019]
[0061] Medical laboratory data 142 may be collected from a patient based on one or more samples obtained from the patient at one or more single periods of time (e.g., derived from samples or measurements obtained from each particular patient at one time during a patient-physician interaction, such as obtaining one sample or measurement (e.g., blood or urine sample) during a single medical appointment). The one or more samples may include various results from different blood, urine, and other laboratory tests.
[0020]
[0062] In one embodiment, the laboratory tests utilized to obtain the measurements represented in training data set 141 are routine laboratory tests that patients would typically have performed during regular outpatient visits. For example, at least some of the measurements represented in training data set 141 may include one or more measurements obtained in connection with a urine chemistry test (e.g., urine creatinine, urine albumin, urine ACR), a comprehensive metabolic panel (e.g., eGFR, glucose, calcium, sodium, albumin, potassium, bicarbonate, chloride, urea, phosphate / phosphorus, magnesium, liver enzymes), a complete blood count (e.g., hemoglobin, hematocrit, platelet count), a liver panel (e.g., ALT, AST, ALKP, GGT, bilirubin), and / or a uric acid test.
[0021]
[0063] In some instances, one or more of the measurements represented in the training data set 141 are not measured directly, but are derived or inferred from other measurements. By way of example, a urine ACR measurement for a particular patient may be converted from a urine protein-creatinine test or a urine dipstick test.
[0022]
[0064] It will be appreciated that, for the purposes of this disclosure, one or more measurements for one or more patients represented in training data set 141 may be missing or removed from training data set 141. As a non-limiting example, if training data set 141 includes medical laboratory data 142 for patient A and patient B, patient A may have laboratory data / measurements that are not available for patient B, such as when a urine chemistry test and complete blood count were performed on both patient A and patient B, but a liver panel was performed only on patient A. Nevertheless, medical laboratory data 142 represented in training data set 141 may be deemed to include one or more measurements related to a urine chemistry test, a complete blood count, and a liver panel, even if a liver panel was not obtained for patient B. In this regard, a set of laboratory data / measurements can be represented in the training data set 141 by one patient combination in the training cohort (e.g., patient A and patient B) even if one or more lab data / measurements in the set of laboratory data / measurements are missing for one or more patients in the set of laboratory data / measurements, and even if there is no single patient in the training cohort for which all lab data / measurements in the set of laboratory data / measurements are present (as long as each of the lab data / measurements in the set of laboratory data / measurements is included for at least one patient included in the training cohort).
[0023]
[0065] In one embodiment, the medical lab data 142 for the training data set 141 has missing values for at least some patients represented in the medical lab data 142. In one instance, the training data set 141 supplements the missing values / measurements by utilizing imputed data. The imputed data can be imputed using any suitable technique (e.g., adaptive tree imputation, proximity techniques, regression imputation, mean substitution, and / or others). For example, the training data set 141 may include eGFR, urinary ACR, urea, potassium, hemoglobin, platelet count, albumin, calcium, glucose, bilirubin, sodium, bicarbonate, and / or GGT for its associated cohort of patients with a degree of value imputation of 30% or less (e.g., any of the above measurements may include imputed values for 30% or less of the patients in the cohort).
[0024]
[0066] The training data set 141 may also include additional information related to a plurality of patients (i.e., a cohort of patients), such as patient outcome information (e.g., included in the patient information 143). Such patient outcome information may also include information about whether and / or when a patient experienced a decline in eGFR (e.g., 40% or other decline), renal failure (e.g., requiring dialysis or a kidney transplant), and / or other medical outcomes related to CKD. The patient information 143 may also additionally or alternatively include a CKD stage of one or more patients. The CKD stage may include stage G1, stage G2, stage G3, stage G4, or stage G5. The stage may also be selected from a plurality of sub-stages corresponding to each of the aforementioned stages, in some instances (e.g., sub-stages of stage G1, etc.). The patient information 143 may also include the patient's sex and / or gender, the patient's age at the time each sample was collected from each patient, history of other diseases / conditions, family history of conditions, previous medical treatments / surgeries, and / or other relevant information such as blood pressure, temperature, oxygen level, reflex testing, and / or other vitals. Such variables, however, may not be required in certain embodiments and may be omitted.
[0025]
[0067] The training data set 141 can be used to train the machine learning model 145 in a variety of ways (e.g., using supervised learning techniques, unsupervised learning techniques, combinations thereof, and / or others). As an illustrative example, to build a random forest model, the system can build a decorrelated tree by randomly sampling (e.g., bootstrap sampling) the original training data set (e.g., training data set 141), fitting the model to a randomly sampled (e.g., smaller) data set, and aggregating the predictions. As another example, to build a random survival forest model, the system can randomly select a subset of features and / or a threshold for evaluation at each node for aggregation.
[0026]
[0068] After training the machine learning model 145, the machine learning model 145 can be utilized (run or execute) to generate a CKD progression prediction (e.g., CKD progression prediction data 144) for a particular patient (e.g., for a new patient). For example, in addition to the medical laboratory data 142 for the new patient, patient information (e.g., age and gender) for the new patient can also be obtained. The medical laboratory data for the new patient may include one or more of the laboratory data / measurements discussed above in association with the medical laboratory data 142 for the training data set 141. By way of example, medical laboratory data for a new patient may include one or more of estimated glomerular filtration rate (eGFR), urinary albumin / creatinine ratio (ACR), urea, serum sodium, serum chloride, serum hemoglobin, serum potassium, glucose, serum albumin, alkaline phosphatase, serum phosphate, serum bicarbonate, serum magnesium, serum calcium, aspartate aminotransferase (AST), alanine aminotransaminase (ALT), bilirubin, gamma-glutamyl transferase (GGT), hematocrit, platelet count, and / or others. Laboratory data / measurements for new patients may include one or more components of urine chemistry (e.g., urine creatinine, urine albumin, urine ACR), comprehensive metabolic panel (e.g., eGFR, glucose, calcium, sodium, albumin, potassium, bicarbonate, chloride, urea, phosphate / phosphorus, magnesium, liver enzymes), complete blood count (e.g., hemoglobin, hematocrit, platelet count), liver panel (e.g., ALT, AST, ALKP, GGT, bilirubin), and / or uric acid.
[0027]
[0069] The age, sex, and medical laboratory data for the new patient can be utilized as input to a (trained) machine learning model 145 to generate CKD progression prediction data 144 for the new patient. The CKD progression prediction data 144 can indicate the risk of the new patient to experience CKD progression, such as in the form of at least a 40% decline in eGFR. In an embodiment, the CKD progression prediction additionally or alternatively indicates the risk of CKD progression in the form of renal failure. Illustratively, the CKD progression prediction data 144 can indicate the risk of a composite CKD progression outcome occurring, the composite outcome including a 40% decline in eGFR or renal failure (e.g., the patient experiencing a 10 ml / min / 1.73 m 2 (An eGFR below 0.05 mg / kg / day will result in the need for long-term dialysis, or a kidney transplant.) As noted above, the machine learning model 145 can be utilized to generate such CKD progression prediction data 144 even for patients who are in the early stages of CKD, such as stage G1 or stage G2, or substages thereof (e.g., for patients who are not in a more end-stage CKD stage than G3).
[0028]
[0070] The CKD progression prediction (e.g., CKD progression prediction data 144) can indicate the risk of CKD progression occurring within a particular amount of time (e.g., from a time point associated with an input data set for a new patient, such as a time point associated with an eGFR measurement for the new patient). By way of non-limiting example, the amount of time associated with the CKD progression prediction may be 2 years, 5 years, or other amounts of time (e.g., 6 months, 1 year, 18 months, 3 years, 4 years, etc.).
[0029]
[0071] In an embodiment, separate machine learning models 145 (e.g., separate random forest models) are trained to generate CKD progression predictions associated with different time scales (e.g., one model for 2-year CKD progression predictions, a separate model for 5-year CKD progression predictions, etc.). In an embodiment, one machine learning model 145 (e.g., one random survival forest model) is trained to generate CKD progression predictions associated with different time scales. Illustratively, a time scale or a particular amount of time (e.g., 2 years, 5 years, or any amount of time or days) can be combined with gender, age, and medical laboratory data for a new patient and provided as input to the machine learning model 145, which can generate a CKD progression prediction for the input time scale or particular amount of time.
[0030]
[0072] 1 also illustrates example additional modules that may be stored on hardware storage device(s) 140 and / or otherwise associated with computing system 110. The additional modules may include one or more of a data retrieval module 151, a data transformation module 152, a training module 153, a validation module 155, and / or an implementation module 156.
[0031]
[0073] As used herein, the term "module" may refer to any combination of hardware components or software objects, routines, or methods that may configure the computing system 110 to perform a particular act. Illustratively, the different components, modules, engines, devices, and / or services described herein may be implemented using one or more objects or processors executing (e.g., as separate threads) on the computing system 110. While FIG. 1 illustrates various independent modules, it will be understood that the characterization of the modules is at least somewhat arbitrary. In at least one embodiment, the various modules described herein may be combined, divided, or eliminated in configurations other than those explicitly described or illustrated. For example, any of the functionality described herein with reference to any particular module may be performed using any number and / or combination of processing units, software objects, modules, instructions, computing centers (e.g., computing centers separate from the computing system 110), and the like. This specification provides individual modules for clarity and explanation, but is not intended to be limiting.
[0032]
[0074] The data retrieval module 151 can be configured to locate and access data sources, databases, and / or storage devices containing one or more data types, from which the data retrieval module 151 can extract a set or subset of data to use as training data. The data retrieval module 151 can receive data from databases and / or hardware storage devices, and the data retrieval module 151 is configured to reformat or otherwise modify the received data for use as training data. Additionally or alternatively, the data retrieval module 151 can communicate with one or more remote systems (e.g., third party system(s) 120) that include third party data collections and / or data sources. In one instance, these data sources include patient lab test results and other patient information portals.
[0033]
[0075] The data retrieval module 151 can access electronically stored information including medical laboratory data 142, patient information 143, and / or CKD progression prediction data 144. The data retrieval module 151 can be configured as a smart module and can learn the optimal data set extraction process to obtain a sufficient amount of data in a timely manner and to retrieve data that is most applicable to the desired application for which a machine learning model / module is to be trained. For example, the data retrieval module 151 can learn which databases and / or data sets will generate training data that will train a model (e.g., for a particular query or a particular task) and increase the accuracy, efficiency, and / or effectiveness of the model in a desired chronic kidney disease prediction technique.
[0034]
[0076] The data retrieval module 151 may locate, select, and / or store raw recorded source data when communicating with one or more ML module(s) and / or models included in the computing system 110. In such instances, other modules in communication with the data retrieval module 151 may receive retrieved (i.e., extracted, pulled, etc.) data from one or more data sources to further enhance the received data and / or apply it to downstream processes. For example, the data retrieval module 151 may communicate with the training module 153 and / or the implementation module 156. The data retrieval module 151 may also be configured to retrieve a training data set (e.g., training data set 141) that includes medical laboratory data 142 and patient information 143.
[0035]
[0077] In one example, data conversion module 152 is configured to convert any raw data retrieved by data retrieval module 151 into workable data for inclusion in training data set 141 .
[0036]
[0078] In one example, the training module 153 communicates with one or more of the data retrieval module 151, the data conversion module 152, the validation module 154, and / or the implementation module 156. In such an embodiment, the training module 153 is configured to receive one or more training data sets (e.g., the training data set 141) via the data retrieval module 151. After receiving the training data related to a particular application or task, the training module 153 can train one or more models on the training data. The training module 153 can be configured to train the models by unsupervised training and / or supervised training. The training module 153 is configured to train the machine learning model 145 to generate a chronic kidney disease progression prediction by applying the training data set 141 including the medical laboratory data 142 and the patient information 143 to generate the CKD progression prediction data 144 as an output.
[0037]
[0079] In one embodiment, the training data set 141 is split into a training data set and a validation data set. The validation module 155 is configured to utilize the validation data set to test the machine learning model 145 for accuracy and precision in predicting CKD progression. For example, a random forest model can be fitted using any desired demographic and laboratory variables and using the Random Forest for Survival, Regression, and Classification (RF-SRC) package in R. Illustratively, the available data can be split into a training (e.g., 70%) data set and a testing / validation (e.g., 30%) data set. Parameters can include a node size (or other size) of 15 and a number of trees (or other number of trees) equal to 60. Additional or alternative random forest or random survival forest (or other) models can also be used within the scope of the present disclosure.
[0038]
[0080] The computing system 110 includes an implementation module 156 that communicates with any one of the modules and / or ML models 145 (or all models / modules) included in the computing system 110 such that the implementation module 156 is configured to implement, initiate, or execute the functionality of one or more of these modules. In one example, the implementation module 156 is configured to operate the data retrieval module 151 such that the data retrieval module 151 can retrieve data at the appropriate time and generate training data for the training module 153. The implementation module 156 can facilitate process communication and timing of communication between one or more of the modules and can be configured to implement and / or operate the machine learning model 145 configured as a CKD progression prediction model.
[0039]
[0081] The computing system can be in communication with third party system(s) 120. The third party system 120 includes one or more processor(s) 122, one or more of the computer readable instructions 118, and one or more hardware storage device(s) 124. The third party system(s) 120 can further include a database that contains data that can be used as training data, e.g., medical laboratory data that is not included in the local storage. Additionally or alternatively, the third party system(s) 120 can also include a machine learning system that is external to the computing system 110.
[0040]
[0082] 2 illustrates an example of a machine learning model 230 (e.g., machine learning model 145 of FIG. 1) trained on a training data set 210 (e.g., training data set 141). The training data set 210 includes medical laboratory data 220A / 220B (e.g., medical laboratory data 142) and patient information (e.g., patient information 143), including CKD stage 214A / 214B, gender 216A / 216B, and age 218A / 218B for a number of patients (e.g., patient A 212A and patient B 212B). The machine learning model 230 is configured to generate a chronic kidney disease progression prediction 280 (e.g., CKD progression prediction data 144) for a new patient 242. The medical laboratory data 220A includes at least eGFR 222A for patient A, and may also include additional laboratory data / measurements for patient A (as indicated by oval 224A). Similarly, medical lab data 220B includes at least eGFR 222B for patient B, and may also include additional lab data / measurements for patient B (as shown by oval 224B). Training data set 210 includes data for any number of patients (as shown in FIG. 2 by the ovals associated with training data set 210).
[0041]
[0083] The training data set 210 is then input to the machine learning model 230 to train the machine learning model 230 to generate a CKD progression prediction, thereby resulting in a CKD progression prediction model 270. A new input data set 240 related to a new patient 242 (e.g., a patient not included in the training data set 210, or a patient for whom a CKD progression prediction is desired) is input as an input to the CKD progression prediction module 270 to generate a CKD progression prediction 280 for the new patient 242. The input data set 242 includes a CKD stage 244, a gender 246, an age 248, and medical laboratory data 250 for the new patient. The medical laboratory data 250 (for the new patient 242) includes at least one eGFR 262 based on one or more samples obtained from the new patient (e.g., eGFR 262 at a point in time or a period of time obtained from samples and / or information obtained from / about the new patient during a patient-physician appointment, during a day, within an hour, etc.). Additionally, the medical laboratory data 250 for the new patient 242 may also include one or more other laboratory data / measurements (as indicated by oval 264). The CKD progression prediction 280 includes a risk score for the new patient to develop a 40% drop in eGFR 282 and / or renal failure 284 within a specified time frame (e.g., within 2 years or within 5 years).
[0042]
[0084] As noted above, a time frame or a particular amount of time 290 associated with a CKD progression prediction 280 can be provided as an input to the CKD progression prediction model 270, such as when the CKD progression prediction model 270 is implemented as a random survival forest model. In some instances, the input time frame or particular amount of time 290 is not provided as an input, and instead the CKD progression prediction model 270 is selected from multiple CKD progression prediction models, each associated with a different time frame or particular amount of time.
[0043]
[0085] The following discussion now refers to a number of methods (e.g., computer-implementable or system-implementable methods) and / or method acts that may be performed in accordance with the present disclosure. Although the method acts are discussed in a particular order and shown in the flow charts as occurring in a particular order, no particular ordering is required unless specifically stated or required, as one act is dependent on other acts being completed before the act is performed. It will be appreciated that certain embodiments of the present disclosure may omit one or more of the acts described herein. The various acts described herein may be performed utilizing one or more computing system components (e.g., hardware processor(s) 112, hardware storage device(s) 140, instructions, and / or modules, etc.) described above.
[0044]
[0086] FIG. 3A illustrates an example of a flow diagram 300 illustrating acts associated with generating a machine learning model for predicting CKD progression.
[0045]
[0087] Act 302 of flow diagram 300 includes accessing a training data set, the training data set including (i) a first set of medical laboratory data related to a plurality of patients, (ii) an age for each patient in the plurality of patients, and (iii) a gender for each patient in the plurality of patients. The first set of medical laboratory data is indicative of at least estimated glomerular filtration rate (eGFR), urinary albumin / creatinine ratio (ACR), urea, serum sodium, serum chloride, serum hemoglobin, serum potassium, glucose, serum albumin, alkaline phosphatase (ALKP), serum phosphate, serum bicarbonate, serum magnesium, serum calcium, aspartate aminotransferase (AST), alanine aminotransaminase (ALT), bilirubin, gamma-glutamyl transferase (GGT), hematocrit, and platelet count for a combination of patients in the plurality of patients.
[0046]
[0088] Act 304 of the flow diagram 300 includes generating a machine learning model by applying a training data set to an untrained model. The machine learning model is configured to generate a chronic kidney disease (CKD) progression prediction for the new patient by applying an input data set related to the new patient to the machine learning model. The input data set includes an age of the new patient, a gender of the new patient, and a second set of medical laboratory data, the second set being indicative of one or more of eGFR, urinary ACR, urea, serum sodium, serum chloride, serum hemoglobin, serum potassium, glucose, serum albumin, ALKP, serum phosphate, serum bicarbonate, serum magnesium, serum calcium, AST, ALT, bilirubin, GGT, hematocrit, and platelet count for the new patient.
[0047]
[0089] It should be appreciated, in light of the present disclosure, that the medical laboratory data utilized as input to the machine learning models can take a variety of forms, and that the machine learning models can treat the input data in a variety of ways. By way of example, any of the measurements can include continuous measurements, categorical measurements, transformed / modified measurements (e.g., log-transformed measurements), mathematically modified measurements (e.g., squared, cubed, etc.), and the like.
[0048]
[0090] In one example, the machine learning model includes a random survival forest model configured to receive a time period input (e.g., days, months, years, etc.) in addition to an input data set to generate a CKD progression prediction for an input time period (e.g., likelihood of CKD progression occurring, such as a 40% drop in eGFR and / or renal failure within an input time period). In one example, the machine learning model includes a random forest model configured to generate a CKD progression prediction for a particular time period. Multiple models can also be generated to generate CKD progression predictions for different time scales.
[0049]
[0091] 3B-3D show example flow charts 310, 320, and 330, respectively, illustrating acts associated with generating a CKD progression prediction for a new patient.
[0050]
[0092] Act 312 of the flow diagram 310 of FIG. 3B includes accessing a machine learning model configured to generate a chronic kidney disease (CKD) progression prediction. The machine learning model is trained on a training data set, the training data set including (i) a first set of medical laboratory data related to a plurality of patients, (ii) an age of each patient in the plurality of patients, and (iii) a gender of each patient in the plurality of patients. The first set of medical laboratory data is indicative of at least estimated glomerular filtration rate (eGFR), urinary albumin / creatinine ratio (ACR), urea, serum sodium, serum chloride, serum hemoglobin, serum potassium, glucose, serum albumin, alkaline phosphatase (ALKP), serum phosphate, serum bicarbonate, serum magnesium, serum calcium, aspartate aminotransferase (AST), alanine aminotransaminase (ALT), bilirubin, gamma-glutamyl transferase (GGT), hematocrit, and platelet count for a combination of patients in the plurality of patients.
[0051]
[0093] In one embodiment, the machine learning model includes a random survival forest model. The first set of medical laboratory data may include one or more imputed values in place of missing data values. In one example, the first set of medical laboratory data is indicative of eGFR, urinary ACR, urea, potassium, hemoglobin, platelet count, albumin, calcium, glucose, bilirubin, sodium, bicarbonate, and GGT with a value imputation degree of 30% or less.
[0052]
[0094] Act 314 of the flow diagram 310 includes generating a CKD progression prediction for the new patient by inputting an input data set related to the new patient into a machine learning model. The CKD progression prediction for the new patient is based on an output of the machine learning model resulting from inputting the input data set related to the new patient into the machine learning model. The input data set includes an age of the new patient, a gender of the new patient, and a second set of medical laboratory data, the second set being indicative of one or more of eGFR, urinary ACR, urea, serum sodium, serum chloride, serum hemoglobin, serum potassium, glucose, serum albumin, ALKP, serum phosphate, serum bicarbonate, serum magnesium, serum calcium, AST, ALT, bilirubin, GGT, hematocrit, and platelet count for the new patient. As used herein, "urine ACR" can include components of the urine ACR, such as direct urine ACR measurements, analytical or estimated urine ACR, and / or urine albumin, urine creatinine, urine protein, and / or qualitative urine albumin (e.g., from a dipstick).
[0053]
[0095] In some instances, the new patient is not associated with a CKD stage beyond G3. In some embodiments, the CKD progression prediction includes predicting the risk of the new patient to develop renal failure or to experience a decline in eGFR of about 40% or more for the new patient. In some instances, the risk of renal failure is predicted for the new patient to (i) require long-term dialysis, (ii) require a kidney transplant, or (iii) experience a decline of 10 ml / min / 1.73 m 2 The patient is at risk of developing a glomerular filtration rate lower than 100 mg / kg / day.
[0054]
[0096] The CKD progression prediction may also indicate the risk of CKD progression occurring within a particular amount of time from a time period associated with the inputted data set for the new patient (e.g., the amount of time from an eFGR measurement associated with the new patient). In one embodiment, the particular amount of time is provided as an input to the machine learning model to generate the CKD progression prediction, such as when the machine learning model is implemented as a random survival forest model. The particular amount of time may include 2 years, 5 years, or any amount of time.
[0055]
[0097] The urinary ACR for one or more of the patients or for new patients may be converted from a urinary protein-creatinine test or a urinary general substance qualitative / semiquantitative test.
[0056]
[0098] Act 316 of the flow chart 310 includes determining that the CKD progression prediction indicates a prediction of the risk of the new patient developing CKD within a particular time period that meets one or more prediction risk thresholds. The one or more prediction risk thresholds can also be based on a particular time period associated with the CKD progression prediction (e.g., different timelines may have different sets of thresholds). In one example, a CKD progression prediction of 2% or more (e.g., indicating that the new patient has a 2% likelihood of CKD progression in the form of a 40% drop in eGFR or renal failure) in a two-year time period can be associated with a "medium" risk classification for the new patient, and a CKD progression prediction of 10% or more can be associated with a "high" risk classification for the new patient. As another example, a CKD progression prediction of 5% or more can be associated with a "medium" risk classification for the new patient, and a CKD progression prediction of 25% or more can be associated with a "high" risk classification for the new patient in a five-year time period. Additional or alternative threshold structures for the same or different time scales are within the scope of this disclosure.
[0057]
[0099] One or more of Acts 318A to 318D may be performed based on the execution of Act 316. Act 318A includes generating a notification that the new patient may require renal interventional therapy. Act 318B includes generating a recommendation for renal interventional therapy for the new patient based on the CKD progression prediction. Act 318C includes generating a recommendation for a monitoring frequency of CKD progression for the new patient based on the CKD progression prediction. Act 318D includes performing renal interventional therapy on the new patient. In accordance with Act 316, the selection of Acts 318A, 318B, 318C, and / or 318D to be performed in response to the CKD progression prediction meeting one or more thresholds may also be based on one or more other factors, such as a particular time period associated with the CKD progression prediction (e.g., 2 years or 5 years), the particular threshold(s) met (e.g., whether the patient is classified as "moderate" or "high" risk), and / or at least a portion of the set of laboratories for the new patient (e.g., used as part of the input data set to generate the CKD progression prediction for the new patient).
[0058]
[0100] Various illustrative examples relating to Acts 318A to 318D are now discussed. In one example, the execution of Act 318A may also include generating a notification to the new patient of complications that may develop with CKD. This may be based on personalized patient laboratory data / measurements and / or other patient data for the new patient.
[0059]
[0101] For example, in response to determining that a new patient is male and has a hemoglobin below about 130 g / L, or female and has a hemoglobin below about 120 g / L, Act 318A may involve generating a notification for the new patient indicating that anemia is a potential complication.
[0060]
[0102] As another example, in response to determining that a new patient has greater than about 5 mEq / L potassium, Act 318A may involve generating a notification for the new patient indicating that hyperkalemia is a potential complication.
[0061]
[0103] As another example, in response to determining that the new patient has serum bicarbonate below about 22 mEq / L, Act 318A may involve generating a notification for the new patient indicating that metabolic acidosis is a potential complication.
[0062]
[0104] As another example, in response to determining that the new patient has phosphorus greater than about 1.6 mg / dL and / or calcium less than about 2.1 mmol / L or greater than about 2.7 mmol / L, Act 318A may involve generating a notification indicating that CKD mineral bone disease (CKD-MBD) is a potential complication for the new patient.
[0063]
[0105] In one example, the recommendations generated pursuant to Act 318B may be based on individualized patient laboratory data / measurements and / or other patient data for the new patient and / or may be based on comorbidities as noted above with respect to Act 318A.
[0064]
[0106] For example, if a new patient has an age greater than about 50 years and a 2 In response to determining that the new patient has an eGFR lower than about 100 mg / mmol or a urinary ACR greater than about 3 mg / mmol, Act 318B may also involve generating a suggestion to prescribe a statin (and / or other cholesterol treatment) to the new patient.
[0065]
[0107] As another example, a new patient may develop a rate of approximately 30 mL / min / 1.73 m 2In response to determining that the patient has an eGFR less than 0.05 mg / kg and is classified as being at "high" risk of CKD progression in accordance with Act 316, Act 318B may also involve generating a recommendation to refer the new patient to nephrology.
[0066]
[0108] As another example, in response to determining that a new patient is classified as being at “moderate” or “high” risk of CKD progression pursuant to Act 316, Act 318B provides that the new patient be treated with renin-angiotensin-aldosterone system (RAAS) inhibition (e.g., the new patient is receiving potassium greater than about 5 mEq / L or greater than about 15 mL / min / 1.73 mEq / L). 2 New patients should have an eGFR of approximately 15 mL / min / 1.73 m 2 RAAS inhibition may be strongly recommended if the patient has an eGFR greater than about 5 mEq / L and a urinary ACR greater than about 3 mg / mmol), nonsteroidal mineralocorticoid receptor antagonist (MRA) therapy (e.g., if a new patient has potassium greater than about 5 mEq / L or a urinary ACR greater than about 25 mL / min / 1.73 m 2 New patients should be treated with 100 mL / min / 1.73 m unless they have an eGFR lower than 2 Approximately 60mL / min / 1.73m 2 If new patients have an eGFR in the range of 10 to 200 mL / min / 1.73 m2, 10 mg daily may be recommended. If new patients have an eGFR higher than approximately 60 mL / min / 1.73 m2, 20 mg daily may be recommended), and / or a sodium-glucose cotransporter-2 (SGLT2) inhibitor (e.g., if new patients have an eGFR higher than approximately 20 mL / min / 1.73 m2, 20 mg daily may be recommended). 2 The method may also involve generating a recommendation to receive treatment with a lower eGFR than the recommended level (unless the patient has an eGFR lower than 100 mg / kg).
[0067]
[0109] In another example, in response to determining that anemia is a potential complication for a new patient (as discussed above with reference to Act 318A), Act 318B may involve generating a recommendation to obtain iron studies, such as ferratin, serum iron, and / or total iron binding capacity (TIBC), for the new patient (e.g., at regular monitoring intervals, as discussed below with reference to Act 318C).
[0068]
[0110] As another example, in response to determining that hyperkalemia is a potential complication for the new patient (as discussed above with reference to Act 318A), Act 318B may involve generating a recommendation that the patient undergo a low potassium diet (if the new patient has potassium in the range of about 5 mEq / L to 5.5 mEq / L) and / or receive hyperkalemia monitoring and / or treatment in accordance with clinical practice guidelines (e.g., if the new patient has potassium greater than about 5.5 mEq / L).
[0069]
[0111] As another example, in response to determining that metabolic acidosis is a potential complication for a new patient (as discussed above with reference to Act 318A), Act 318B may involve generating a recommendation that the patient undergo metabolic acidosis monitoring and / or treatment in accordance with clinical practice guidelines.
[0070]
[0112] As another example, in response to determining that CKD-MBD is a potential complication for a new patient (as discussed above with reference to Act 318A), Act 318B may involve generating a recommendation that the patient undergo a low phosphorus diet.
[0071]
[0113] In one example, Act 318B recommends a target blood pressure of approximately 130 / 80 mmHg (or approximately 60 mL / min / 1.73 m for new patients).2 The method may also include the step of recommending one or more blood pressure goals to the new patient, such as a target systolic blood pressure of about 120 mmHg if the patient has an eGFR lower than about 100 mmHg or a urinary ACR greater than about 3 mg / mmol.
[0072]
[0114] In one instance, the recommendations generated pursuant to Act 318C may be based on individualized patient laboratory data / measurements and / or other patient data for the new patient and / or based on comorbidities as noted above with reference to Act 318A.
[0073]
[0115] For example, according to Act 316, new patients are classified as being at “high” risk for CKD progression and have a median risk of approximately 60 mL / min / 1.73 m 2 In response to determining that the new patient has an eGFR lower than 100%, Act 318C may also involve generating a recommendation that the new patient undergo CKD monitoring at least four times (or more) each year.
[0074]
[0116] As another example, according to Act 316, new patients are classified as being at "high" risk for CKD progression and have a CKD rate of approximately 60 mL / min / 1.73 m 2 In response to determining that the new patient has an eGFR greater than 3x, Act 318C may also involve generating a recommendation that the new patient undergo CKD monitoring three (or more) times per year.
[0075]
[0117] As another example, according to Act 316, new patients are classified as having a "moderate" risk of CKD progression and have a mean blood flow of approximately 45 mL / min / 1.73 m 2 In response to determining that the new patient has an eGFR lower than 3x, Act 318C may also involve generating a recommendation that the new patient undergo CKD monitoring three (or more) times per year.
[0076]
[0118] As another example, according to Act 316, new patients are classified as having a "moderate" risk of CKD progression and have a mean blood flow of approximately 45 mL / min / 1.73 m 2In response to determining that the new patient has an eGFR greater than 100%, Act 318C may also involve generating a recommendation that the new patient undergo CKD monitoring twice (or more) annually.
[0077]
[0119] As another example, in response to determining in accordance with Act 316 that the new patient is classified as having a "low" risk of CKD progression (e.g., the new patient is not classified as either "moderate" or "high" risk), Act 318C may involve generating a recommendation that the new patient undergo CKD monitoring annually (or more).
[0078]
[0120] Act 318D may include implementing one or more of the recommendations discussed above with reference to Acts 318B and / or 318C (e.g., RAAS inhibition, blood pressure control, SGLT2 inhibitors, MRA treatment), and / or others (e.g., nephrology consultation, home dialysis, and / or kidney transplant).
[0079]
[0121] FIG. 4 illustrates an example report including various components discussed above with reference to Acts 314, 316, 318A, 318B, and / or 318C, such as a CKD progression prediction 402 (indicating a 22% risk of CKD progression over a 5-year time horizon, characterized as "moderate" based on meeting thresholds above 5% and below 25%), potential complications of CKD 404, recommended treatments 406 and additional recommendations 408, a referral recommendation to a nephrologist 410, blood pressure target recommendations 412, and monitoring frequency recommendations 414.
[0080]
[0122] Reports similar (at least in some respects) to the report shown in Figure 4 may also be generated in response to a request made by a physician or following initial care performed (e.g., as a routine procedure for patients who meet certain criteria). It will be appreciated, in light of this disclosure, that reports according to this disclosure may include additional or alternative components and may take on a variety of forms / formats.
[0081]
[0123] Turning attention to FIG. 3C, FIG. 3C illustrates that act 322 of flow diagram 320 includes accessing a machine learning model configured to generate a chronic kidney disease (CKD) progression prediction. The machine learning model is trained on a training data set, the training data set including (i) a first set of medical laboratory data related to a plurality of patients, (ii) an age of each patient included in the plurality of patients, and (iii) a gender of each patient included in the plurality of patients. The first set of medical laboratory data is indicative of a urinary albumin / creatinine ratio (ACR), an estimated glomerular filtration rate (eGFR), urea, hemoglobin, albumin, hematocrit, glucose, phosphate, bicarbonate, gamma-glutamyl transferase (GGT), platelet count, magnesium, and chloride for at least one combination of patients included in the plurality of patients.
[0082]
[0124] Act 324 of the flow diagram 320 includes generating a CKD progression prediction for the new patient by inputting an input data set related to the new patient into a machine learning model. The CKD progression prediction for the new patient is based on an output of the machine learning model resulting from inputting the input data set related to the new patient into the machine learning model. The input data set includes an age of the new patient, a gender of the new patient, and a second set of medical laboratory data, the second set including one or more of a urine chemistry test, a comprehensive metabolic panel, a complete blood count, a liver panel, or a uric acid test for the new patient.
[0083]
[0125] In one embodiment, the second set of medical laboratory data includes one or more of a urine chemistry test, a comprehensive metabolic panel, and a complete blood count for the new patient. Although not shown in FIG. 3C, the flow chart 320 may further include acts similar to acts 316, 318A, 318B, 318C, and / or 318D for execution based on the CKD progression prediction generated according to act 324.
[0084]
[0126] Act 332 of the flow diagram 330 of Figure 3D includes accessing a machine learning model configured to generate a chronic kidney disease (CKD) progression prediction. The machine learning model is trained on a training data set, the training data set including (i) a first set of medical laboratory data related to a plurality of patients, (ii) an age of each patient in the plurality of patients, and (iii) a gender of each patient in the plurality of patients. The first set of medical laboratory data is indicative of at least a urinary albumin / creatinine ratio (ACR), estimated glomerular filtration rate (eGFR), urea, and hemoglobin for a combination of patients in the plurality of patients.
[0085]
[0127] Act 334 of flow diagram 330 includes generating a CKD progression prediction for the new patient by inputting an input data set related to the new patient into a machine learning model. The CKD progression prediction for the new patient is based on an output of the machine learning model resulting from inputting the input data set related to the new patient into the machine learning model. The input data set includes an age of the new patient, a gender of the new patient, and a second set of medical laboratory data, the second set including one or more of a urine chemistry test, a comprehensive metabolic panel, a complete blood count, a liver panel, or a uric acid test for the new patient.
[0086]
[0128] In an embodiment, the second set of medical laboratory data includes one or more items of a urine chemistry test for the new patient. In one example, the second set of medical laboratory data includes one or more items of a urine chemistry test and a comprehensive metabolic panel for the new patient. Although not shown in FIG. 3D, the flow chart 330 may further include acts similar to acts 316, 318A, 318B, 318C, and / or 318D for execution based on the CKD progression prediction generated according to act 334.
[0087]
[0129] As previously noted, various types of machine learning models can be implemented to facilitate the generation of CKD progression predictions for patients in accordance with the present disclosure. The following discussion refers to example implementations of various Random Forest and Random Survival Forest models for generating CKD progression predictions. Random Forest model example(s)
[0130] Figure 5 shows a schematic of an example of the selection of a cohort of patients from which the training data set for the machine learning model was generated. The study development cohort was derived from administrative data in Manitoba, Canada (population 1.4 million at the time) using data from the Manitoba Centre for Health Policy (MCHP). MCHP is a research unit within the Department of Community Health Sciences at the University of Manitoba that maintains a population-based repository of data on health activities and other social determinants of health across all individuals in the province. The training data set included all adult (age 18 and over) individuals in the province with outpatient eGFR testing available between April 1, 2006 and December 31, 2016, and a valid Manitoba Health Registry for at least 1 year pre-index. For example, the CKD-EPI formula was used to calculate eGFR from available serum creatinine testing. In addition, patients were asked to include demographic information regarding age and sex, as well as urinary albumin-to-creatinine ratio (ACR) or protein-to-creatinine ratio (PCR) test results. Patients with a history of renal failure (dialysis or transplant) were excluded. Data were de-identified using scrambled personal health information numbers.
[0088]
[0131] In this example study, the system identified 6,717,522 serum creatinine tests between April 1, 2006 and December 31, 2016, of which 3,574,628 were performed in outpatient settings. From this, the system was able to identify 634,133 unique individuals with at least one computable eGFR measurement and a valid health registration. After narrowing down the requirement that a urine ACR test (or converted PCR test) be valid, the system arrived at a total cohort size of 77,196 individuals for both the training and testing datasets (Figure 5). The training dataset included complete follow-up in 61,353 individuals (42,947 in training and 18,406 in examination) to assess outcomes at 2 years, and an additional 35,736 individuals (54,037 in training and 23,159 in examination) to assess outcomes at 5 years.
[0089]
[0132] In one example embodiment, the mean age of the baseline cohort was 59.3 years (± 17.0) and patients had a mean blood glucose level of 82.2 (± 27.2) ml / min / 1.73 m 2 The median ACR after including converted PCR was 1.1 mg / mmol (interquartile range 0.5 to 4.7 mg / mmol). 47.7% of patients were male, 45.2% had diabetes, and 69.9% had hypertension. 5.2%, 3.6%, and 2.6% had a history of congestive heart failure, stroke, or myocardial infarction, respectively. When divided into training and examination groups, characteristics were similar.
[0090]
[0133] 6A shows a table containing a description of the cohort discussed above with reference to FIG. 5, including various test results contained in the medical laboratory data for each patient. The various test results were classified as independent and dependent variables included in a training data set (e.g., training data set 141).
[0091]
[0134] The training data set included age, sex, eGFR, and urinary ACR as previously described. Starting from the first recorded eGFR during the study period, moving to the last available test in a 6-month window, and calculating the average of the tests during this period, baseline eGFR was calculated as the average of all available eGFR results. The index date of the patient was considered to be the date of the last eGFR in this 6-month period. Age was determined at the date of index eGFR, and sex was determined using linkage to the Manitoba Health Insurance Registry, which contains date of birth and other demographic data. When a urinary ACR test was not available, available urinary protein / creatinine (PCR) tests were converted to the corresponding urinary ACR using published and validated formulas. The closest result within 1 year of the index date was selected (pre- or post-index). Due to variables skewed distribution, urinary ACR was log-transformed.
[0092]
[0135] In addition to the variables already described, other relevant laboratory variables with low missingness (<15% or <30%) were included in the model construction. These included serum sodium, serum chloride, serum hemoglobin, urea, serum potassium, glucose, AST, ALT, bilirubin, GGT, hematocrit, and / or platelet count. The closest values within 1 year of the index date were selected (pre- or post-). The model constructed using these variables is called the "10-variable model" (age, sex, and the aforementioned laboratory data (labs)).
[0093]
[0136] When applied in the Cox proportional hazards model, multiple imputation (n=5) was applied using SA PROC MI. The Random Forest model allows for variables to be missing due to the observation that having a "missing value" is treated as a split value for the variable when determining branching using SAS PROC HPFOREST. An additional Random Forest model is evaluated that includes six additional variables that allow for any degree of missingness: serum albumin, alkaline phosphatase, serum phosphate, serum bicarbonate, serum magnesium, and serum calcium. This model is referred to as the 16-variable model. Laboratory data included in the training dataset can be extracted from the Shared Health Diagnostic Services of Manitoba (DSM) Laboratory Information System.
[0094]
[0137] The outcome for at least some of the disclosed embodiments is a prediction and / or risk score for a 40% decline in eGFR or renal failure for a patient. Within the training data set, a 40% decline in eGFR was determined as the first eGFR test that was a 40% or greater decline from baseline eGFR, followed by a second confirmatory test at least one month later if the patient did not die or develop renal failure in that month period. The date on which the 40% decline occurred is considered the first of these qualifying tests. Renal failure is defined as any of three conditions: initiation of long-term dialysis, acceptance of a transplant, or a decrease in eGFR<10ml / min / 1.73m 2Dialysis was defined as any 2 claim in the Manitoba Health Services Database for long-term dialysis, transplantation was defined as any 1 claim in the Manitoba Health Services Database for transplantation, and hospitalization in the Discharge Abstract Database (DAD) was defined by the procedure code corresponding to kidney transplantation (1PC85 or 1OK85 using Canadian Classification of Health Interventions (CCI) codes). An overview of tariff codes identifying dialysis and transplantation is shown in Figure 7.
[0095]
[0138] Figure 6B is a table outlining the missingness of different variables in the baseline cohort. When applied in the Cox proportional hazards model, the system applied multiple imputation for variables with missingness less than 30% using SAS PROC MI. When applied in the random forest model, the system applied imputation for missing data using the missing data algorithm. All included laboratory data were extracted from the Shared Health Diagnostic Services of Manitoba (DSM) Laboratory Information System and did not include any values recorded during hospitalization events adjudicated by linkage to the Discharge Information Database (DAD).
[0096]
[0139] The outcome date for a 40% decline in eGFR or renal failure was determined based on the first of these events. Figure 8 shows a table outlining the variable importance for each variable included in the training data set of the machine learning model. Specifically, the table shows that for the example random forest model, the variables that had the highest impact in generating accurate CKD progression predictions included urinary ACR, eGFR, urea, and hemoglobin. Age and gender were also significant variables.
[0097]
[0140] 9 conceptually illustrates an example of a training data set 910 including patient information (e.g., gender 916A, 916B, age 918A, 918B) and medical laboratory data for each patient included in the training data set 910. As shown, medical laboratory data 920A associated with patient A 912A includes measurements of eGFR 922A, urinary ACR 924A, serum sodium 926A, serum chloride 928A, serum hemoglobin 932A, urea 934A, serum potassium 936A, and glucose 938A. Similarly, as shown, medical laboratory data 920B associated with patient B 912B includes measurements of eGFR 922B, urinary ACR 924B, serum sodium 926B, serum chloride 928B, serum hemoglobin 932B, urea 934B, serum potassium 936B, and glucose 938B. The ovals indicate that any number of patients may be included in the training data set 910. As noted above, some measurements may be missing for one or more patients represented in the training data set 910.
[0098]
[0141] A random forest model can be fitted using the R package Fast Unified Random Forest (RF-SRC) for survival, regression, and classification, which uses survival forest with right-censored survival. To accomplish this, the data is split into training (70%) and testing (30%) data sets. The model was evaluated for accuracy using the time-dependent area under the receiver operating characteristic (ROC) curve, Brier score, and calibration plots of observed versus predicted risk. Additionally, in this particular example, the system evaluated the sensitivity, specificity, negative predictive value (NPC), and positive predictive value (PPV) for the top 10%, 15%, and 20% of patients by estimated risk (high risk) and in the lower 50%, 45%, and 30% of estimated risk (low risk).
[0099]
[0142] To assess generalizability, the system evaluated the model in subpopulations of the test cohort: (1) patients with diabetes, (2) patients without diabetes, and (3) patients with eGFR < 60 ml / min / 1.73 m 2 or patients with CKD as defined by a urine ACR > 3 mg / mmol (including converted urine PCR test), and (4) an eGFR of 30 to 60 ml / min / 1.73 m 2 or eGFR > 60 ml / min / 1.73 m 2 and urinary ACR>3 mg / mmol (including converted urine PCR test). See Figures 27A-B. The final grown 22-variable forest was used to assess variable importance of the included parameters.
[0100]
[0143] We also developed Cox proportional hazards models in the training dataset: (1) a model including variables with at most 30% missing (11-variable model), and (2) a model including the variables age, sex, eGFR, and urinary ACR for comparison with the Kidney Failure Risk Prediction Equation (KFRE). Model discrimination was assessed using Harrell's C statistic, accuracy was assessed using Brier scores, and calibration was assessed using plots of observed versus predicted risk probability in the laboratory dataset. Analyses were performed using SAS version 9.3 (Cary, NC) and R version 4.1.0. Statistical significance was identified a priori using alpha = 0.05.
[0101]
[0144] Random forest models were also fitted using SAS PROC HPFOREST and internally validated using SAS PROC HP4SCORE using a variety of demographic and laboratory variables. One statistical analysis examined out-of-bag (OOB) misclassification rates versus the number of leaves selected in the model. Measures of accuracy for predicting outcomes at 2 and 5 years were assessed for the random forest models. Measures included the area under the receiver operating characteristic (ROC) curve, Brier score, and calibration plots of observed and predicted risks by risk decile of predicted probability.
[0102]
[0145] In addition, other parameters including sensitivity, specificity, negative predictive value (NPV), and positive predictive value (PPV) were evaluated at cutoffs of 1% and 10% in the 2-year model and 5% and 25% in the 5-year model. These cutoffs were selected because they are clinically significant and correspond approximately to the bottom 60% and top 10% of individuals classified by predicted risk score. To assess squared error loss, measures of variable importance were calculated using the random branch assignments (RBA) method in SAS PROC HP4SCORE.
[0103]
[0146] For example, FIG. 10 is a graph showing an example calibration plot for a machine learning model configured as a random forest model, for example, using the training data set shown in FIG. 9 to predict decline within a two-year time period. FIG. 11 is a graph showing an example calibration plot for a machine learning model configured as a random forest model, for example, using the training data set shown in FIG. 9 for a five-year time period. As is evident from the graphs shown in FIGS. 10-11 for the example implementation, the five-year prediction (FIG. 11) correlated more closely with the observed outcome than the two-year prediction (FIG. 10), but both predictive models provided useful predictive metrics that can guide patient care and / or treatment / prevention decisions.
[0104]
[0147] This study also analyzed various expanded Cox proportional hazards models in a training data set with the aforementioned variables to predict the risk of developing a 40% decline or renal failure outcome, and subsequently validated these internally in a test set. At 2 and 5 years, model discrimination was assessed using Harrell's C statistic, precision was assessed using Brier scores, and calibration was assessed by deciles of predicted risk using plots of observed versus predicted risk probability. All analyses were performed using SAS version 9.4 (Cary, NC). Statistical significance was identified a priori using alpha = 0.05.
[0105]
[0148] For example, FIG. 12 is a graph showing an example of a calibration plot for a machine learning model configured as a Cox model using, for example, the training data set shown in FIG. 9 for a 2-year time period. FIG. 13 is a graph showing an example of a calibration plot for a machine learning model configured as a Cox model using, for example, the training data set shown in FIG. 9 for a 5-year time period. As is evident from the graphs shown in FIGS. 12-13 for the example implementation, the 2-year predictions (FIG. 12) correlated more closely with the observed outcomes than the 5-year predictions (FIG. 13), but both prediction models provided useful predictive metrics that could guide patient care and / or treatment / prevention decisions. Furthermore, the 10-variable Cox model was more highly correlated with the observed outcomes at 2 years (FIG. 12) compared to the 10-variable Random Forest model (FIG. 10).
[0106]
[0149] 14 conceptually illustrates an example of a training data set 1410 including patient information (e.g., gender 1416A, 1416B, age 1418A, 1418B) and medical laboratory data for each patient included in the training data set 1410 that can be used to form a nine-variable model predicting CKD progression. The training data set 1410 is similar to the training data set 910 of FIG. 9, but with the urinary ACR measurements removed. As shown, the medical laboratory data 1420A associated with patient A 1412A includes measurements of eGFR 1422A, serum sodium 1426A, serum chloride 1428A, serum hemoglobin 1432A, urea 1434A, serum potassium 1436A, and glucose 1438A. Similarly, as shown, medical laboratory data 1420B associated with patient B 1412B includes measurements of eGFR 1422B, serum sodium 1426B, serum chloride 1428B, serum hemoglobin 1432B, urea 1434B, serum potassium 1436B, and glucose 1438B. Any number of patients may be included in the training data set 1410. As noted above, some measurements may be missing for one or more patients represented in the training data set 1410.
[0107]
[0150] FIG. 15 is a graph showing an example of a calibration plot for a machine learning model configured as a Cox model, for example, using a training data set as shown in FIG. 14, for a 2-year time period. FIG. 16 is a graph showing an example of a calibration plot for a machine learning model configured as a Cox model, for example, using a training data set as shown in FIG. 14, for a 5-year time period. As is evident from the graphs shown in FIG. 15-FIG. 16 for this implementation example, the 2-year prediction (FIG. 15) correlated more closely with the observed outcomes than the 5-year prediction (FIG. 16), but both prediction models provided useful predictive metrics that could guide patient management and / or treatment / prevention decisions. It should be noted that the 2-year and 5-year predictions using the 9-variable model (FIG. 15 and FIG. 16) produced similar correlation results as the 2-year and 5-year predictions using the 10-variable model (FIG. 12 and FIG. 13), and the omission of the ACR still provided closely correlated predictive power for either time frame.
[0108]
[0151] 17 shows an example of a training data set 1710 including a medical laboratory data set of 16 to 22 variables that can be used to train a machine learning model configured to generate a chronic kidney disease progression prediction. The training data set 1710 is an example of the training data set 910 in FIG. 9 (including gender 1716A and 1716B and age 1718A and 1718B for patient A 1712A and patient B 1712B, respectively), where additional measurements are included in the medical laboratory data for at least some of the patients included in the training data set 1710.
[0109]
[0152] As shown, medical laboratory data 1720A associated with patient A 1712A includes measurements of eGFR 1722A, urinary ACR 1724A, serum sodium 1726A, serum chloride 1728A, serum hemoglobin 1732A, urea 1734A, serum potassium 1736A, glucose 1738A, serum albumin 1721A, alkaline phosphatase 1723A, serum phosphate 1725A, serum bicarbonate 1727A, serum magnesium 1729A, and serum calcium 1731A.
[0110]
[0153] Similarly, as shown, medical laboratory data 1720B associated with patient B 1712B includes measurements of eGFR 1722B, urinary ACR 1724B, serum sodium 1726B, serum chloride 1728B, serum hemoglobin 1732B, urea 1734B, serum potassium 1736B, glucose 1738B, serum albumin 1721B, alkaline phosphatase 1723B, serum phosphate 1725B, serum bicarbonate 1727B, serum magnesium 1729B, and serum calcium 1731B. In one embodiment, patient A medical laboratory data 1720A and patient B medical laboratory data 1720B further include AST, ALT, bilirubin, GGT, hematocrit, and / or platelet count 1740A and 1740B, respectively. Any number of patients may be included in the training data set 1710. As noted above, some measurements may be missing for one or more patients represented in the training data set 1710.
[0111]
[0154] In one embodiment, the machine learning model trained using the training dataset 1710 is configured as a 22-variable model, i.e., the input dataset for a new patient may include as many as 22 different laboratory data points / measurements (or potentially even more).
[0112]
[0155] FIG. 18 is a graph showing an example of a calibration plot for a machine learning model using, for example, 16 variables of a training data set as shown in FIG. 17 for a 2-year time period. FIG. 19 is a graph showing an example of a calibration plot for a machine learning model using, for example, 16 variables of a training data set as shown in FIG. 17 for a 5-year time period. As is evident from the graphs shown in FIG. 18 and FIG. 19 for this implementation example, the 5-year prediction (FIG. 19) correlated more closely with the observed outcomes than the 2-year prediction (FIG. 18), but both prediction models provided useful predictive metrics that can guide patient management and / or treatment / prevention decisions. It should further be noted that for the 2-year prediction, the 16-variable model (FIG. 18) showed improved correlation when compared to the 10-variable model (FIG. 10). However, for the 5-year prediction, both the 16-variable model (FIG. 19) and the 10-variable model (11) showed substantially similar performance for the 40% prediction threshold. The 16-variable model (Figure 19) provided more stable correlations with lower percentage thresholds than the 10-variable model (Figure 11).
[0113]
[0156] FIG. 20 is a graph showing the calibration plot for the 22-variable random forest model for prediction of 40% decline in eGFR or renal failure at 5 years.
[0114]
[0157] 21 shows an example of a training data set 2110 including 15 to 21 variables of medical laboratory data that can be used to train a machine learning model configured to generate a chronic kidney disease progression prediction. The training data set 2110 is an example of the training data set 1710 of FIG. 17 (including gender 2116A and 2116B, and age 2118A and 2118B for patient A 2112A and patient B 2112B, respectively), with the exception that it excludes urea ACR measurements for each patient included in the training data set 2110.
[0115]
[0158] As shown, medical laboratory data 2120A associated with patient A 2112A includes measurements of eGFR 2122A, serum sodium 2126A, serum chloride 2128A, serum hemoglobin 2132A, urea 2134A, serum potassium 2136A, glucose 2138A, serum albumin 2121A, alkaline phosphatase 2123A, serum phosphate 2125A, serum bicarbonate 2127A, serum magnesium 2129A, and serum calcium 2131A.
[0116]
[0159] Similarly, as shown, medical laboratory data 2120B associated with patient B 2112B includes measurements of eGFR 2122B, serum sodium 2126B, serum chloride 2128B, serum hemoglobin 2132B, urea 2134B, serum potassium 2136B, glucose 2138B, serum albumin 2121B, alkaline phosphatase 2123B, serum phosphate 2125B, serum bicarbonate 2127B, serum magnesium 2129B, and serum calcium 2131B. In one embodiment, patient A medical laboratory data 2120A and patient B medical laboratory data 2120B further include AST, ALT, bilirubin, GGT, hematocrit, and / or platelet count 2140. The training data set 2110 may include any number of patients. As noted above, some measurements may be missing for one or more patients represented in the training data set 2110.
[0117]
[0160] FIG. 22 is a graph showing an example of a calibration plot for a machine learning model using, for example, a training data set (15 variables) as shown in FIG. 21 for a 2-year time period. FIG. 23 is a graph showing an example of a calibration plot for a machine learning model using, for example, a training data set (15 variables) as shown in FIG. 21 for a 5-year time period. As shown in the graphs illustrated in FIG. 22 and FIG. 23 for this implementation example, the 5-year prediction (FIG. 23) correlated more closely with the observed outcomes than the 2-year prediction (FIG. 22), but both prediction models provided useful predictive metrics that could guide patient management and / or treatment / prevention decisions. Furthermore, for the 5-year prediction, the 15-variable model (FIG. 23) performed similarly to the 16-variable model (FIG. 19), suggesting that removing the ACR did not significantly affect the predictions provided by the models.
[0118]
[0161] Figure 24 is a table showing an example of summary performance evaluation statistics for various examples of machine learning models having 4 to 11 variables and configured as Cox models disclosed herein. As shown in Figure 24, various models were evaluated for predicted performance at 5 years. The variables considered included age, eGFR, log-transformed ACR, hematocrit, potassium, chloride, glucose, sodium, urea, male sex, and platelet count.
[0119]
[0162] In other studies (not shown), the system evaluated a Cox proportional hazards model in a cohort that had complete available follow-up at years 2 and 5 for comparison to the output of a random forest model: In this study cohort, for predicting outcomes at year 2, the Cox proportional hazards model had a C-statistic of 0.8492 (SE 0.007) in the baseline model, decreasing to 0.8151 (0.006) at year 5.
[0120]
[0163] In models that omitted urinary ACR (e.g., 9 and 15 variable models), the system observed C-statistics of 0.8266 (0.008) at 2 years and 0.7942 (0.006) at 5 years. In models that applied 2 years of follow-up to the cohort, the Brier score was 0.0298 (0.001) for predicting outcomes of eGFR decline or renal failure, and in models that applied 5 years of follow-up to the cohort, the Brier score was 0.0832 (0.002) in the tested cohort. In models that omitted urinary ACR, the Brier score was 0.0305 (0.001) for predicting outcomes at 2 years and 0.0855 (0.002) for predicting outcomes at 5 years.
[0121]
[0164] Figure 25 shows the calibration plots for Cox proportional hazards models, including a 4-variable model and an 11-variable model. Both models performed well and predicted risk with high accuracy. Different Cox proportional hazards models were evaluated with a maximum follow-up time of 5 years for outcomes of 40% decline in eGFR or renal failure, and censoring for death and loss to follow-up. These included (1) an 11-variable model including all variables with 30% or less missingness, namely, age, eGFR, male, urinary ACR, platelet count, potassium, hematocrit, serum chloride, glucose, serum sodium, and urea, and (2) a 4-variable model including age, eGFR, male, and urinary ACR. For the 11-variable Cox model, the Harrell's C statistic was 0.849 (95% confidence interval of 0.837 to 0.861) and the Bryan score was 4.4 (2.4 to 6.3), indicating that it was correctly calibrated at all risk levels. Similarly, the four-variable Cox model yielded a similar calibration with a Harrell's C statistic of 0.829 (0.816 to 0.842) and a Bryan score of 4.5 (2.5 to 6.5), as shown in Figure 25.
[0122]
[0165] FIG. 26A is a table showing an example of a summary of performance evaluation statistics for various examples of machine learning models configured as random forest models. For the random forest model with 10 variables, the system observed excellent discrimination with the area under the ROC being 0.8406 (SE 0.0080) at 2 years and 0.7966 (0.0069) at 5 years. In terms of accuracy, the system observed a Bryan score of 0.029 (SE 0.001) at 2 years and 0.077 (0.002) at 5 years. For the baseline model at 2 and 5 years, the system observed excellent calibration. For the 16-variable random forest, the C statistic was 0.8697 (0.007) for predicting outcome at 2 years and 0.8190 (0.006) for predicting outcome at 5 years. When ACR was excluded from the model, the C statistic was 0.8597 (0.007) at year 2 and 0.8014 (0.007) at year 5. Additional model metrics and calibration plots for the 16-variable and 15-variable (excluding ACR) models are shown in the corresponding figures.
[0123]
[0166] Figure 26B is another table showing the model performance summary for the Random Forest model (22-variable version of the machine learning model described above). A low risk was determined when the risk was between 1.2% and 2.6%. A high risk was determined when the risk was between 9% and 17%. Performance was evaluated in a test cohort of 23,159 patients. Even with the Random Forest model with 22 variables, the system showed excellent discrimination, with a time-dependent area under the receiver operating characteristic (AUROC) curve of 86.9 (95% CI 85.8 to 88.1) for up to 5 years of follow-up, and a Bryan score of 4.2 (2.5 to 6.0). The observed results included good calibration. Similar performance was observed in all subgroups: diabetes (AUROC: 86.3, Brier: 5.2), no diabetes (AUROC: 87.1, Brier: 3.1), CKD (AUROC: 83.5, Brier: 7.7), and CKD stage G1-G3 (AUROC: 79.8, Brier: 6.7).
[0124]
[0167] Statistics for sensitivity, specificity, and positive predictive value were evaluated in patients at high risk (top 10, 15, and 20% risk scores, respectively). In this evaluation, in the top 10% of risk scores, the sensitivity was 47% (17% 5-year risk threshold), the specificity was 93%, and the positive predictive value was 36%. In the top 15% (12% 5-year risk threshold), the sensitivity was 59%, the specificity was 89%, and the positive predictive value was 30%. In the top 20% (9% 5-year risk threshold), the model had a sensitivity of 67%, a specificity of 84%, and a positive predictive value of 26%.
[0125]
[0168] Similarly, the system evaluated sensitivity, specificity, and negative predictive value in low-risk patients (the bottom 50, 45, and 30% of patients, respectively). In the lowest 50% of patients (2.6% 5-year risk threshold), the model had a sensitivity of 91%, a specificity of 53%, and a negative predictive value of 99%. In the lowest 45% of patients (2.1% 5-year risk threshold), the model had a sensitivity of 93%, a specificity of 48%, and a negative predictive value of 99%. Finally, in the lowest 30% of patients (1.2% 5-year risk threshold), the model had a sensitivity of 96%, a specificity of 32%, and a negative predictive value of 99%.
[0126]
[0169] Figures 27A-D show various calibration plots for the 22-variable model constructed as a random forest model in various subgroups. For example, Figure 27A shows the calibration plot for the diabetic subgroup. Figure 27B shows the calibration plot for the non-diabetic subgroup. Figure 27C shows the calibration plot for the non-diabetic subgroup. 2 Figure 27D shows the calibration plot for the subgroup of patients with CKD stages G1-G3 (e.g., eGFR between 30-60 ml / min / 1.73 m^2, or eGFR>60 ml / min / 1.73 m^2, and urine ACR>3 mg / mmol with converted urine PCR). Random Survival Forest model example(s)
[0170] To develop an example of a random survival forest model that generates CKD progression predictions, a construction cohort was derived from administrative data in Manitoba, Canada (population 1.4 million) using data from the Manitoba Centre for Health Policy. All adult (age 18 years or older) individuals in the province who had an available outpatient eGFR test between April 1, 2006 and December 31, 2016 and a valid Manitoba Health Registry for at least 1 year pre-index were identified. eGFR was calculated from available serum creatinine tests using the CKD-Disease Epidemiology Collaboration equation. Included patients were further required to have complete demographic information on age and sex, including at least one urinary ACR or protein-to-creatinine ratio (PCR) test result. Patients with a history of renal failure (dialysis or transplant) were excluded. The cohort discussed above with reference to Figure 5 was used to construct the random survival forest model.
[0127]
[0171] The validation cohort was derived from the Alberta Heath database. This database contains information on demographic data, laboratory data, hospitalizations, and physician claims for all patients in the province of Alberta, Canada (population 4.4 million). Formal laboratory coverage of creatinine measurements and ACR / PCR values was completed in 2005. However, additional laboratory values were only fully imputed after 2009. Therefore, we identified a cohort of individuals with at least one calculable eGFR, valid health registration, and ACR (or imputed PCR) value starting from April 1, 2009 through December 31, 2016. One-third of the external cohort was randomly sampled to perform the final analysis and reduce imputation time. Patients with a history of renal failure (dialysis or transplant) were excluded. Figure 28 shows an aspect of the validation cohort used to externally validate the random survival forest model.
[0128]
[0172] To construct the random survival forest model, all candidate models included age, sex, eGFR, and urinary ACR (e.g., as previously described). Baseline eGFR was calculated as the average of all available outpatient eGFR results, starting from the first recorded eGFR during the study period and working forward to the last available test in a 6-month window, and the average of tests over this period was calculated. The index date of the patient was considered to be the date of the last eGFR in this 6-month period. Age was determined as the date of index eGFR, and sex was determined using linkage to the Manitoba Health Insurance Registry, which records birthdate and other demographic data. When a urinary ACR test was not available, the available urinary PCR test was converted to the corresponding urinary ACR using a published and validated formula. The closest result was selected within 1 year before or after the index date. To handle the skewed distribution, urinary ACR was log-transformed.
[0129]
[0173] In addition to the variables already described (age, sex, eGFR, and urinary ACR), the utility of additional laboratory results from chemistry panel, liver enzymes, and complete blood count panel was evaluated for inclusion in the random forest model for survival. The closest value within 1 year of the index date was selected for inclusion. When necessary, distributional transformations were applied. The final random survival forest model included eGFR, urinary ACR, and 18 additional laboratory results (i.e., urea, serum sodium, serum chloride, serum hemoglobin, serum potassium, glucose, serum albumin, alkaline phosphatase, serum phosphate, serum bicarbonate, serum magnesium, serum calcium, AST, ALT, bilirubin, GGT, hematocrit, and platelet count). Figure 29 shows an overview of the missingness for the laboratory panel. The random forest model applied imputation for missing data using an adaptive tree imputation method.
[0130]
[0174] All laboratory data included were extracted from the Manitoba Laboratory Information System's Shared Health Diagnostics Service; none were included if they were recorded during a hospitalization event, as determined by linkage to the hospital discharge database (admission tests). For the validation cohort, Alberta Health Laboratory data were extracted from the Alberta Kidney Disease Network. Of the 18 laboratory tests used in the Manitoba model, 16 laboratory tests were also routinely collected from the Alberta Kidney Disease Network. Tests that were not available (aspartate aminotransferase and gamma-glutamyl transferase) were treated as missing data.
[0131]
[0175] The primary outcomes in this case study were a 40% decline in eGFR or renal failure. A 40% decline in eGFR was determined at the first eGFR test where there was a 40% or greater decline from baseline eGFR in the laboratory data, and a second confirmatory test result was required between 90 days and 2 years after the first test, unless the patient died or developed renal failure within 90 days after the first test result that revealed a 40% or greater decline. Thus, if a patient's eGFR represents a 40% decline at one time, and the patient dies within 90 days, it is treated as an event. It is also treated as an event if they develop renal failure during this time period. Renal failure is determined by the initiation of long-term dialysis, receipt of a transplant, or a decline in eGFR < 10 ml / min / 1.73 m 2 Dialysis was defined as any 2 claim in the Manitoba Health Services Database for long-term dialysis, transplantation was defined as any 1 claim in the Manitoba Health Services Database for kidney transplantation, and hospitalization in the Discharge Information Database (DAD) was defined by the corresponding procedure code for kidney transplantation (1PC85 or 1OK85 using the Canadian Classification of Health Interventions code or International Classification of Diseases, Ninth Revision, procedure code 55.6). A summary of tariff codes identifying dialysis and transplantation is shown in Figure 30.
[0132]
[0176] The outcome date of a 40% decline in eGFR or renal failure was determined based on the first of these events. Patients were followed until reaching the composite end point described above, death (determined by linkage to the Manitoba Health Registry), for a maximum of 5 years, or until loss to follow-up.
[0133]
[0177] A 40% decline in eGFR was identified using laboratory creatinine measurements as described for the Manitoba cohort previously. Renal failure was defined similarly, with minor modifications required for administrative data sets with different structures (see Figure 30). Long-term dialysis and kidney transplants were identified using the North and South Alberta Renal Program Databases, the provincial registry of renal replacement, and any one code for hemodialysis, peritoneal dialysis, or transplant. (Note: Because the registry began in 2001, physicians requested that their data be used when excluding individuals who had previously undergone transplant or dialysis.) These data were source linked to the provincial laboratory repository by a unique coded patient identifier.
[0134]
[0178] Baseline characteristics for the construction (internal training and testing) and external validation cohorts were summarized by descriptive statistics. Random forest models were constructed using the R package Fast Unified Random Forest for Survival, Regression, and Classification using survival forest with right-censored data. Data were split into training (70%) and testing (30%) data sets in a single split, then validated in the external cohort. Models were assessed for accuracy using the area under the receiver operating characteristic curve, Brier score, and calibration plots of observed versus predicted risk. Area under the receiver operating characteristic curve and Brier score were assessed at 1-year intervals for predicting outcomes at years 1 through 5, and calibration plots were assessed at years 2 and 5. Model hyperparameters were optimized using the tune.rfsrc function, and comparisons of maximum terminal node size and number of variables were used to possibly split at each node to out-of-bag error rates from the Random Forest package for Survival, Regression, and Classification. In addition, sensitivity, specificity, negative predictive value (NPV), and positive predictive value (PPV) were evaluated for the top 10%, 15%, and 20% of patients predicted to be at highest risk (high risk), including the bottom 50%, 45%, and 30% of patients predicted to be at lowest risk (low risk). These metrics were evaluated at 2 and 5 years. The risk of progression versus predicted probability was visualized and plotted over 2 and 5 years. The final grown 22-variable survival forest was used to evaluate the variable importance of the included parameters, as shown in Figure 31.
[0135]
[0179] To assess robustness, the models were evaluated for 5-year prediction of the primary outcome defined by CKD stage and presence or absence of diabetes in subpopulations of the testing and validation cohorts. For sensitivity analysis, two comparison models were considered: (i) a Cox proportional hazards model was evaluated using a guideline-based risk definition with a 3-level definition of albuminuria as a categorical predictor and 5 stages of eGFR as comparators (heatmap model); (ii) a Cox proportional hazards model including the variables eGFR, urinary ACR, diabetes, hypertension, stroke, myocardial infarction, age, and sex (clinical model). In addition, the models were evaluated in an external validation cohort in which laboratory values were included only for the year before the index date.
[0136]
[0180] Analyses were performed using R version 4.1.0. Statistical significance was specified a priori using a factor of 1 / 4 0.05. For the construction cohort (training and testing), a total sample size of 77,196 was used, with 54,037 assigned to the training data set (70%) and 23,159 assigned to the testing data set. In the validation cohort, a total of 321,396 individuals were identified, and a random subset of 107,097 were selected for evaluation. A detailed overview of the cohort selection process for both the construction and validation cohorts is shown in Figure 5 and Figure 28.
[0137]
[0181] The mean age of the cohort was 59.3 years, and the mean eGFR was 82.2 ml / min / 1.73 m 2 and median urinary ACR was 1.1 mg / mmol. Of the patients, 48% were male, 45% had diabetes, 70% had hypertension, 5% had a history of congestive heart failure, 4% had a previous stroke, and 3% had a previous myocardial infarction (similar between the test and training cohorts).
[0138]
[0182] The validation cohort was somewhat younger, with a mean age of 55.5 years and a mean eGFR of 86.0 ml / min / 1.73 m 2and median ACR was 0.8 mg / mmol. The validation cohort had a higher proportion of male patients (53%), 41% of patients had diabetes, 51% had hypertension, 5% had congestive heart failure, 5% had prior stroke, and 5% had prior myocardial infarction. A summary of baseline descriptive statistics is shown in Figure 32.
[0139]
[0183] A random survival forest model with 22 variables was found to have an AUC of 0.90 (0.89-0.92) for 1-year prediction of the primary outcome and 0.84 (0.83-0.85) for 5-year prediction when evaluated in the test cohort. The Brier score was 0.02 (0.01-0.02) for 1-year prediction of the primary outcome and 0.07 (0.06-0.09) for 5-year prediction. The AUC and Brier scores for 1-5 years are shown in Figure 33. The AUC and Brier scores were similar in multiple predefined subgroups (Figure 34). The model exhibited excellent calibration at both 2 and 5 years in both the internal and external test cohorts (see Figures 35A and 35B). In addition, it was observed that the association between the occurrence of the primary outcome event increased as the predicted probability generated by the random forest algorithm increased.
[0140]
[0184] Statistics for sensitivity, specificity, and PPV were evaluated in high-risk patients (top 10%, 15%, and 20% of risk scores, respectively). For predicting the primary outcome at 2 years, patients falling in the top decile (14% 2-year risk threshold) were found to have a sensitivity of 58%, specificity of 92%, and PPV of 25%. Similarly, for the top 15% of patients (10% 2-year risk threshold), sensitivity was found to be 69%, specificity was 87%, and PPV was 20%. For the top 20% of patients (7% 2-year risk threshold), sensitivity was 76%, specificity was 83%, and PPV was 16%. Using a 30% threshold to identify high- and intermediate-risk patients, 87% of individuals had an event within 2 years and 77% within 5 years were identified.
[0141]
[0185] For low-risk patients, the bottom 50% of patients (1.95% 2-year risk threshold) were found to have a sensitivity of 94%, specificity of 52%, and NPV of >99%. For the lowest 45% of risk scores (1.61% 2-year risk threshold), the sensitivity was 95%, the specificity was 47%, and the NPV was >99%. Finally, for the lowest 30% of risk scores (0.85% 2-year risk threshold), the sensitivity was 97%, the specificity was 31%, and the NPV was >99%. These statistics were also examined for predicting outcomes at 5 years and were found to have similar accuracy (see Figure 36).
[0142]
[0186] Urinary ACR (including converted PCR) was the most influential variable in the Random Forest model, followed by eGFR, urea, hemoglobin, age, serum albumin, hematocrit, and glucose. As noted above, a detailed summary of the model inputs, ranked according to importance, is shown in Figure 31.
[0143]
[0187] Performance was found to be similar when evaluated in the external validation cohort, with AUC decreasing from 0.87 (0.86-0.89) for 1-year predictions to 0.84 (0.84-0.85) for 5-year predictions, and Brier scores of 0.01 (0.01-0.01) at 1 year and 0.04 (0.04-0.04) at 5 years (Figure 33). In the external validation cohort, the overall risk decreased at both 2 and 5 years, but the model exhibited excellent calibration (Figures 37A and 37B), and the association between the rank of the risk score and the probability of the composite outcome was similarly high.
[0144]
[0188] In addition, patients with and without diabetes, CKD stages G1 to G3, and eGFR < 60 ml / min / 1.73 m 2Analysis of the subgroups yielded outcomes similar to those of the internal testing cohort (Figure 34). Diagnostic accuracy was observed in the external validation cohort similar to that of the construction cohort, as assessed by sensitivity, specificity, NPV, and PPV (Figure 36).
[0145]
[0189] In comparator analyses, the Heatmap model performed less well than the 22-variable Random Survival Forest model in the construction cohort (C-statistic 0.78 vs. 0.84 at year 5, Figure 38) as well as the clinical model (C-statistic 0.81 at year 5, P<0.001, Figure 39). When considering only laboratory values in the 12 months prior to the index date, model evaluation results were unchanged against the Random Forest model (1-year AUC 0.87, 0.86-0.88; 5-year AUC 0.84, 0.83-0.85). conclusion
[0190] At least some of the disclosed embodiments provide an externally validated laboratory-based predictive model for the outcome of renal failure or 40% decline in eGFR. The disclosed model can be based entirely on a single time point measure of routinely collected laboratory data and can predict the subject's outcome (CKD progression) with higher accuracy than current standard of care models or commercially available models that intend to test for novel biomarkers and / or use machine learning methods. The models disclosed herein, when taken together, can be implemented in clinical and research settings.
[0146]
[0191] At least some of the disclosed machine learning models using random forest or random survival forest appear to perform better than commercially available machine learning models, such as RenalytixAI. Compared to the RenalytixAI tool, at least some of the disclosed models have the advantage of having external validation in an independent population, thus reducing the risk of overfitting. This step is particularly important for machine learning models, which tend to overfit the development population when derived on a small data set with many predictors, and often do not generalize well. Furthermore, at least some of the disclosed models only require laboratory data that can be easily mapped, making them easier to implement at scale than models that require multiple electronic health record fields and data types, such as the RenalytixAI tool.
[0147]
[0192] Finally, at least some of the disclosed models, in contrast to RenalytixAI, do not require (and can explicitly omit) any measurement or use as input of novel or proprietary biomarkers, and therefore, at least some of the disclosed models can be implemented in a routine laboratory setting or using already collected laboratory data.
[0148]
[0193] The disclosed models have important clinical and research implications. From a clinical perspective, physicians could use at least some of the disclosed models in the clinic to assess whether patients are early in the CKD process (eGFR > 60 ml / min / 1.73 m 2) can identify patients at high risk of progression over the next 5 years. Given the effect of interventions such as SGLT2 inhibitors on the slope of eGFR in this population, these patients may be able to forestall or even completely prevent the lifetime onset of renal failure, as opposed to delaying the time to dialysis if interventions are implemented later during the disease progression. In addition, as newer therapies emerge, such as finerenone, they may have the added benefit of slowing CKD progression. However, such new and / or in-development therapies have been extensively studied in patients with preserved renal function, and in order to maximize efficacy while reducing cost burden and polypharmacy, these therapies may be initially reserved for intermediate and high risk subgroups. Implementing the disclosed model can facilitate the targeted and efficient guidance of the use of such new therapies for at-risk patients.
[0149]
[0194] From a research perspective, various large clinical trials use a 40% decline in eGFR or renal failure as primary outcomes, and validation of at least a portion of the disclosed models in these trial data sets may also help highlight risk treatment interactions. Future trials currently in the planning or enrollment stages may use at least a portion of the disclosed models to help improve the quality of the test population and generate a reasonable number of outcomes in a reasonable time frame.
[0150]
[0195] Strengths of at least some of the embodiments discussed above include external validation. This is particularly important for machine learning models because they can overfit small data sets with many predictor variables. In addition to this issue, at least some of the disclosed models have been externally validated and found to be rigorous in cohorts with complete missingness on two variables. Further strengths include novel research methods, including random forest methodology on two detailed data sets, the results of which have been proven to be generalizable to multiple renal outcomes and interventions. A notable strength is that it relies only on routinely collected laboratory data, allowing for rapid integration into electronic health records and laboratory information systems.
[0151]
[0196] In conclusion, we disclose a machine learning model that uses routinely collected laboratory data to predict CKD progression (40% decline in eGRF or renal failure) with high accuracy in all CKD patients (even those in early stages of CKD, e.g., G1 or G2). Further terms and definitions
[0197] The present invention may be embodied in other specific forms without departing from its spirit or essential characteristics. The described embodiments are to be considered in all respects only as illustrative and not limiting. The scope of the present invention is therefore indicated by the appended claims rather than by the foregoing description. All changes that come within the meaning and range of equivalency of the claims are intended to be embraced therein. Furthermore, elements described in connection with any embodiment illustrated and / or described herein may in any way be combined with elements described in connection with any other embodiment illustrated and / or described herein.
[0152]
[0198] The terms "approximately," "about," and "substantially," as used herein, refer to an amount or condition that is close to a stated amount or condition and still performs a desired function or achieves a desired result. For example, the terms "approximately," "about," and "substantially" may refer to an amount or condition that deviates from a stated amount or condition by less than 10%, or by less than 5%, or by less than 1%, or by less than 0.1%, or by less than 0.01%.
[0153]
[0199] In an embodiment, the time period (or point in time or time frame) refers to a minute, an hour, a day, a week, or a year. Alternatively, in an embodiment, the time period refers to a time period such as over multiple hours, over multiple days, over multiple weeks, or over multiple years, the time period having a first start time and a second end time that is after the first start time. Typically, the input data set for a new patient as described herein includes medical laboratory data (typically lab data (labs) ordered from a single doctor's visit or a series of related and / or collective doctor's visits designed to diagnose and / or treat a specific set of symptoms or a specific disease, e.g., CKD) based on one or more samples obtained from the patient during a single examination session. Further computer system details
[0200] Embodiments of the present invention may comprise or utilize a special purpose or general purpose computer (e.g., computing system 110), including computer hardware, as discussed in more detail below. Embodiments within the scope of the present invention also include physical computer-readable media and other computer-readable media for carrying or storing computer-executable instructions and / or data structures. Such computer-readable media may be any available media that can be accessed by a general purpose or special purpose computer system. A computer-readable medium (e.g., hardware storage device 140 of FIG. 1) that stores computer-executable instructions (e.g., computer-readable instructions 118 of FIG. 1) is a physical hardware storage medium / device, and excludes transmission media. A computer-readable medium that carries computer-executable or computer-readable instructions (e.g., computer-readable instructions 118) in one or more carrier waves or signals is a transmission medium. Thus, by way of example and not limitation, embodiments of the present invention may include at least two distinctly different kinds of computer-readable media: physical computer-readable storage media / devices and transmission computer-readable media.
[0154]
[0201] A physical computer readable storage medium / device is hardware and includes RAM, ROM, EEPROM, CD-ROM, or other optical disk storage (such as CDs, DVDs, etc.), magnetic disk storage, or other magnetic storage devices, or any other hardware that can be used to store desired program code means in the form of computer-executable instructions or data structures and that can be accessed by a general purpose or special purpose computer.
[0155]
[0202] A "network" (e.g., network 130 in FIG. 1) is defined as one or more data links that enable the transport of electronic data between computer systems and / or modules and / or other electronic devices. When information is transferred or provided to a computer over a network or other communications connection (either hardwired, wireless, or a combination of hardwired or wireless), the computer properly views the connection as a transmission medium. Transmission media can include a network and / or data link that can be used to carry or otherwise transmit the desired program code means in the form of computer-executable instructions or data structures that can be accessed by a general purpose or special purpose computer. Combinations of the above are also included within the scope of computer-readable media.
[0156]
[0203] Furthermore, program code means in the form of computer executable instructions or data structures can be automatically transferred from a transmission computer readable medium to a physical computer readable storage medium (or vice versa) when reaching various computer system components. For example, computer executable instructions or data structures received over a network or data link can be buffered in RAM within a network interface module (e.g., a "NIC") and eventually transferred to the computer system's RAM and / or to a less volatile computer readable physical storage medium within the computer system. Thus, computer readable physical storage media can be included in computer system components that also utilize transmission media (or even in computer system components that primarily utilize transmission media).
[0157]
[0204] Computer-executable instructions include, for example, instructions and data that cause a general purpose computer, special purpose computer, or special purpose processing device to perform a certain function or group of functions. Computer-executable instructions may be, for example, binaries, intermediate format instructions such as assembly language, or source code. Although the subject matter has been described above in language specific to structural features and / or methodological acts, it should be understood that the subject matter defined in the appended claims is not necessarily limited to the features and acts set forth above. On the contrary, the described features and acts are disclosed as example forms of implementing the claims.
[0158]
[0205] However, those skilled in the art will appreciate that the present invention may be practiced in a network computing environment with many types of computer system configurations, including personal computers, desktop computers, laptop computers, message processors, handheld devices, multiprocessor systems, microprocessor-based or programmable consumer electronics, network PCs, minicomputers, mainframe computers, mobile phones, PDAs, pagers, routers, switches, and the like. The present invention may also be practiced in a distributed system environment, in which local and remote computer systems are linked through a network (either by hardwired data links, wireless data links, or a combination of hardwired and wireless data links) to both perform tasks. In a distributed system environment, program modules may be located in both local and remote memory storage devices.
[0159]
[0206] Alternatively, or in addition, the functionality described herein may be performed, at least in part, by one or more hardware logic components. For example, and without limitation, exemplary types of hardware logic components that may be used include field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), application specific standard products (ASSPs), systems on a chip (SOCs), complex programmable logic devices (CPLDs), and the like.
Claims
**Claim 1** A method comprising: accessing a machine learning model configured to generate a prediction of the progression of chronic kidney disease (CKD), wherein the machine learning model is trained on a training data set comprising (i) a first set of medical laboratory data related to a plurality of patients, (ii) the age of each patient included in the plurality of patients, and (iii) the gender of each patient included in the plurality of patients, wherein the first set of medical laboratory data indicates estimated glomerular filtration rate (eGFR), urine albumin / creatinine ratio (ACR), urea, serum sodium, serum chloride, serum hemoglobin, serum potassium, glucose, serum albumin, alkaline phosphatase, serum phosphate, serum bicarbonate, serum magnesium, serum calcium, aspartate aminotransferase (AST), alanine aminotransferase (ALT), bilirubin, gamma-glutamyl transferase (GGT), hematocrit, and platelet count for at least one combination of patients included in the plurality of patients; generating a CKD progression prediction for the new patient by inputting an input data set related to the new patient into the machine learning model, wherein the CKD progression prediction for the new patient is based on the output of the machine learning model obtained by inputting the input data set related to the new patient into the machine learning model, wherein the input data set includes the age of the new patient, the gender of the new patient, and a second set of medical laboratory data, wherein the second set indicates one or more of eGFR, urine ACR, urea, serum sodium, serum chloride, serum hemoglobin, serum potassium, glucose, serum albumin, alkaline phosphatase (ALKP), serum phosphate, serum bicarbonate, serum magnesium, serum calcium, AST, ALT, bilirubin, GGT, hematocrit, and platelet count for the new patient; A method comprising the above steps. **Claim 2** The method according to claim 1, wherein the new patient is not associated with a CKD stage of G3 or later. **Claim 3** The method according to claim 1, wherein the machine learning model includes a random survival forest model. **Claim 4** The method according to claim 1, wherein the CKD progression prediction indicates a risk of CKD progression within a specific amount of time from a time period associated with the input data set for the new patient.
5. The method according to claim 4, wherein the specific amount of time is supplied as an input to the machine learning model that generates the CDK progression prediction.
6. The method according to claim 4, wherein the specific amount of time includes 2 years or 5 years.
7. The method according to claim 1, wherein the urinary ACR for one or more of the plurality of patients or the new patient is converted from a urine protein / creatinine test or a semi-quantitative qualitative test for general substances in urine.
8. The method according to claim 1, wherein the CKD progression prediction includes a prediction of the risk that the new patient will develop renal insufficiency or a risk that a reduction of 40% or more in eGFR will occur in the new patient.
9. The method according to claim 8, wherein the risk of renal insufficiency for the new patient is (i) the risk of requiring long-term dialysis, (ii) the risk of requiring kidney transplantation, or (iii) an indication that there is a risk of glomerular filtration rate less than 10 ml / min / 1.73 m 2 A method comprising an indication that the glomerular filtration rate is less than 10 ml / min / 1.73 m
10. The method according to claim 1, further comprising: determining that the CKD progression prediction indicates a prediction of the risk that the new patient will develop CKD within a specific time period that meets one or more prediction risk thresholds; (i) generating a notification that the new patient may need renal intervention; (ii) generating a recommendation for renal intervention for the new patient based on the CKD progression prediction; (iii) generating a recommendation for the monitoring frequency of CKD progression for the new patient based on the CKD progression prediction, or (iv) performing renal intervention on the new patient.
11. The method according to claim 10, wherein the one or more prediction risk thresholds are based on the specific time period associated with the CKD progression prediction.
12. The method according to claim 10, wherein the recommendation for renal intervention or the recommendation for the monitoring frequency of CKD progression is further based on at least a part of the second set of medical laboratory data related to the new patient.
13. The method according to claim 10, wherein the renal intervention therapy includes one or more of renin-angiotensin-aldosterone system (RAAS) inhibition, blood pressure management, sodium-glucose cotransporter-2 (SGLT2) inhibitor, mineralocorticoid receptor antagonist (MRA) therapy, or preparation for nephrology consultation, home dialysis, dialysis access, or kidney transplantation.
14. The method according to claim 1, wherein the first set of the medical laboratory data includes one or more substitution values instead of missing values.
15. The method according to claim 14, wherein the first set of the medical laboratory data shows estimated glomerular filtration rate (eGFR), urinary albumin / creatinine ratio (ACR), urea, potassium, hemoglobin, platelet count, albumin, calcium, glucose, bilirubin, sodium, bicarbonate, and GGT with a value substitution degree of 30% or less.
16. A system, one or more processors, one or more hardware storage devices storing instructions executable by the one or more processors, comprising, the instructions configure the system to access a training data set, the training data set includes (i) a first set of medical laboratory data related to a plurality of patients, (ii) the age of each patient included in the plurality of patients, and (iii) the gender of each patient included in the plurality of patients, the first set of the medical laboratory data shows estimated glomerular filtration rate (eGFR), urinary albumin / creatinine ratio (ACR), urea, serum hemoglobin, glucose, and hematocrit for at least one combination of patients included in the plurality of patients, the instructions further configure the system to generate a machine learning model by applying the training data set to an untrained model, the machine learning model is configured to generate a chronic kidney disease (CKD) progression prediction for the new patient by inputting an input data set related to the new patient into the machine learning model, the input data set includes the age of the new patient, the gender of the new patient, and a second set of medical laboratory data, and the second set shows one or more of eGFR, urinary ACR, urea, serum hemoglobin, glucose, and hematocrit for the new patient.
17. The system according to claim 16, wherein the machine learning model includes a random survival forest model.
18. One or more hardware storage devices storing instructions executable by one or more processors of the system, the instructions causing the system to access a machine learning model configured to generate a chronic kidney disease (CKD) progression prediction, the machine learning model being trained on a training data set including (i) a first set of medical laboratory data related to a plurality of patients, (ii) the age of each patient included in the plurality of patients, and (iii) the gender of each patient included in the plurality of patients, the first set of medical laboratory data indicating, for at least one combination of patients included in the plurality of patients, urinary albumin / creatinine ratio (ACR), estimated glomerular filtration rate (eGFR), urea, and hemoglobin, generate a CKD progression prediction for the new patient by inputting an input data set related to the new patient into the machine learning model, the CKD progression prediction for the new patient being based on the output of the machine learning model obtained by inputting the input data set related to the new patient into the machine learning model, the input data set including the age of the new patient, the gender of the new patient, and a second set of medical laboratory data, the second set including one or more items of a urinalysis, comprehensive metabolic panel, complete blood count, liver panel, or uric acid test for the new patient.
19. The one or more hardware storage devices according to claim 18, wherein the one or more items of a urinalysis, comprehensive metabolic panel, complete blood count, liver panel, or uric acid test for the new patient include at least measurements of ACR, eGFR, urea, hemoglobin, glucose, and hematocrit.
20. The one or more hardware storage devices according to claim 19, wherein the first set of medical laboratory data further indicates glucose and hematocrit for at least one combination of patients included in the plurality of patients.