Methods and systems for monitoring and treating post cardiac surgery patients

By integrating CQL with distributional RL and IQN, the models address the limitations of current protocols by providing personalized and adaptive treatment plans for cardiac surgery patients, effectively managing glucose and hemodynamics to improve patient safety and outcomes.

WO2026072677A1PCT designated stage Publication Date: 2026-04-02MT SINAI SCHOOL OF MEDICINE
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Filing Date
2025-09-24
Publication Date
2026-04-02

AI Technical Summary

Technical Problem

Current post-operative management protocols for cardiac surgery patients, particularly in the ICU, fail to account for individual patient variability, leading to high rates of hyperglycemic and hypoglycemic episodes, and are ineffective in managing acute kidney injury (AKI), with existing machine learning models lacking real-time adaptability and personalization.

Method used

Integration of conservative Q-learning (CQL) with distributional reinforcement learning (RL) and implicit quantile networks (IQN) to develop personalized treatment plans for insulin titration and hemodynamic management, adapting to rapidly changing clinical conditions.

Benefits of technology

The integrated models provide accurate and safe insulin dosing and hemodynamic management, maintaining optimal glucose and hemodynamic parameters, reducing the risk of adverse events like hypoglycemia and AKI, and improving patient outcomes.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure US2025047727_02042026_PF_FP_ABST
    Figure US2025047727_02042026_PF_FP_ABST
Patent Text Reader

Abstract

The present disclosure provides machine learning based systems, methods, and non-transitory media for diagnosing and treating patients that have undergone cardiac surgery. Included herein are pre-trained system, methods, and non-transitory media for diagnosing and treating post-surgical cardiac patients based on post-surgical clinical information to address and / or prevent postoperative hyperglycemia or acute kidney injury (AKI) that may arise in the days immediately after surgery. Also included are system, methods, and non-transitory media for diagnosing and treating post-surgical cardiac patients for other ailments that may arise during the period following such surgery.
Need to check novelty before this filing date? Find Prior Art

Description

Attorney Docket Number: MS-0045-01-WOMETHODS AND SYSTEMS FOR MONITORING AND TREATING POST CARDIACSURGERY PATIENTSSTATEMENT REGARDING FEDERALLY SPONSORED RESEARCHThis work was supported by National Institutes of Health (NIH) grant K08DK131286.RELATED APPLICATION

[0001] This application claims priority to U.S. Provisional Patent Application No. 63 / 698.447, filed September 24, 2024, entitled “GLUCOSE: A Distributional Reinforcement Learning Model for Optimal Glucose Control After Cardiac Surgery" the content of which is incorporated herein in its entirety.FIELD

[0002] The present disclosure relates generally to systems, methods, and non-transitory media for use in diagnosing and treating patients after cardiac surgery.BACKGROUND

[0003] Cardiac surgery often causes substantial metabolic and physiological stress, which can manifest as postoperative hyperglycemia, acute kidney injury' (AKI), or other ailments.

[0004] Post-operative hyperglycemia after cardiac surgery is common, occurring in 60- 80% patients with diabetes, and over 50% of non-diabetic patients. It is associated with higher rates of post-operative infections, acute kidney injury', cardiac arrhythmias, longer length of stay, and higher mortality. Due to its significance, the Society of Thoracic Surgeons (STS) recommends maintaining blood glucose levels below 180 mg / dL after cardiac surgery.

[0005] Acute kidney injury, defined by a rise in serum creatinine, occurs in over one-third of patients following cardiac surgery and is associated w ith 4-fold increase in mortality and doubling of health-care costs. Outcomes are much worse when AKI persists for over 48 hours, also known as persistent AKI. Management of each of these conditions in the post-operative setting is challenging.

[0006] For glucose control, one study found that only 15% of patients had appropriate control, defined as glucose level between 70 mg / dL to 180 mg / dL, within the first day afterAttorney Docket Number: MS-0045-01-WO cardiac surgery' (Williams, J. B. et al., J. Crit. Care 42, 328-333 (2017)). The early postoperative period, when patients are critically ill and require care in intensive care unit (ICU) settings, is highly dynamic with rapidly changing clinical characteristics of patients. Currently, post-operative glucose management involves titration of regular insulin based on hospital specific protocols and the experience of treating clinicians. However, due to the highly dynamic nature of this early post-operative period, some treatment regimens maybe more suitable for certain patients or only effective for a limited time as their condition evolves. This can lead to high rates of hyperglycemic and hypoglycemic episodes, both associated with worse outcomes, as these protocolized regimens often fail to account for individual patient variability in real-world settings.

[0007] Hyperglycemia early after cardiac surgery is associated with higher rates of postoperative infections, acute kidney injury', cardiac arrhythmias, longer length of stay, and higher mortality. This underscores the significance of glucose control in the early post-operative period. Moreover, research indicates that the harmful effects of hyperglycemia are dosedependent, with longer exposure and higher glucose levels leading to worse outcomes. Therefore, both the severity and duration of hyperglycemia should be managed. To achieve the target glucose levels, most cardiac surgery' centers employ institutional protocols for managing hyperglycemia. However, a significant challenge in early postoperative glucose management is that insulin, the primary treatment for hyperglycemia, has a narrow therapeutic window. Since these protocolized regimens often fail to account for individual patient variability in real- world settings, hypoglycemia becomes a significant risk, particularly with intensive insulin dosing schemes.

[0008] Hypoglycemia, defined as a blood glucose level <70 mg / dL, can trigger increased sympathetic activity leading to increased heart rate or arrhythmias, impairment of autonomic cardiac reflexes, poor neurological outcomes, and death. Hypoglycemia is seen in 5-21% patients after cardiac surgery, prompting a more conservative insulin dosing which, in turn, can result in persistent hyperglycemia. Thus, not surprisingly, these protocols frequently fall short, with only 15% of patients reaching the recommended glucose levels without hypoglycemia on the first day after surgery, which is a significant and dynamic period after cardiac surgery.

[0009] Clinicians review over 1300 data points per patient each day, making it difficult to effectively use all this information for clinical decision-making. An algorithm that can systematically process and interpret these data points can significantly enhance clinicianAttorney Docket Number: MS-0045-01-WO workflow while improving patient outcomes. Prior algorithmic approaches to inpatient insulin management primarily involve institution specific sliding scale doses, focused on glucose prediction, or use static daily insulin dose estimation. Sliding scale insulin regimens are standard across most institutions, but they are reactive and non-personalized, providing the same dose for a given glucose regardless of patient-specific factors, a practice that can be both ineffective and dangerous (Moghissi, E. S. et al., Diab Care 32, 1119-1131 (2009)). Nguyen et al. developed a supervised machine learning model to predict the total daily dose of insulin to improve upon weight-based dosing guidelines (Nguyen, M. et al., J. Am. Med Inf. Assoc. 28, 2212-2219 (2021)). However, this approach excluded ICU patients and does not provide realtime dosing recommendations.

[0010] Alternatively, while there exist many supervised machine learning models for inpatient glucose prediction to address such challenges in glycemic management, including in ICU settings, these models forecast glucose trends rather than recommending sequential, personalized insulin dosing strategies (Fitzgerald. O. et al.. J. Am. Med Inf. Assoc. 28, 1642- 1650 (2021); Zale, A. & Mathioudakis, N„ Curr. Diab Rep. 22, 353-364 (2022)). Several modeling approaches have been explored for glucose prediction and control, including stochastic modeling frameworks that leverage variable-length time-stamped data to capture seasonal glucose patterns. For example, a seasonal stochastic local modeling approach (“Glucose Prediction under Variable-Length Time-Stamped Daily Events7’) has been proposed to address inter-day variability in glucose regulation (Montaser, E. et al.. Sensors (Basel) 21, 3188 (2021)). While these models offer valuable predictive capabilities, they often lack adaptive decision-making mechanisms for real-time insulin dosing.

[0011] Acute kidney injury (AKI), defined by a rise in serum creatinine, occurs in over one-third of patients following cardiac surgery' and is associated with a 4-fold increase in mortality and doubling of health-care costs. Outcomes are much worse w hen AKI persists for over 48 hours, also known as persistent AKI. Therefore, the management of AKI to prevent its further development is of importance. The current protocols are primarily supportive, with hemodynamic management as the cornerstone of AKI management. The treatment for AKI management focuses on a three-pronged approach: intravenous (IV) fluids, vasopressors, and inotropes. No specific pharmacologic therapy has been shown to consistently prevent the progression of AKI. This is significant in the early post-operative period after cardiac surgery when patients have a very' dynamic hemodynamic profile requiring simultaneous use of IVAttorney Docket Number: MS-0045-01-WO fluids, vasopressors, and inotropes. These interventions when used as a bundle have shown benefit in single-center trials but have failed to reproduce success in multi-center settings. Furthermore, implementation of these bundles in real-world settings remains alarmingly low, with compliance rates under 12%. The limitations of this "one-size-fits-all" approach are increasingly evident. Recent work has shown that AKI is a heterogeneous syndrome, with distinct phenotypes that exhibit variable responses to therapy. Notably, even when phenotypes are defined by widely available clinical data rather than specialized biomarkers, differences in disease trajectory and treatment response persist.

[0012] Therefore, there is a need for dynamic, data-driven, personalized methods and systems that clinicians and caregivers can utilize to manage and treat post-operative cardiac patients, particularly in the days following immediately after surgery.SUMMARY

[0013] The present disclosure fulfills the foregoing and / or other needs by providing machine learning models that combine conservative Q-leaming models (CQL) with implicit quantile networks (IQN) that are modified and trained with curated information to avoid adverse outcomes while optimizing adherence to the observed behavior of experts that treat such patients.

[0014] Reinforcement learning (RL) is a t pe of machine learning wfiere an agent leams to make decisions by performing actions in an environment to maximize cumulative rewards. RL algorithms receive feedback in the form of rew ards or penalties based on the actions taken, allowing the agent to improve its policy over time. This iterative learning process enables the agent to develop policies that maximize cumulative reward over time. In healthcare, where clinical trajectories evolve dynamically, RL is particularly well-suited for guiding timesensitive interventions such as titration of intravenous fluids, vasopressors, and inotropes in response to continuously changing physiological states. This adaptability makes RL particularly well-suited for tasks that involve complex decision-making and real-time adjustments, such as insulin titration or the titration of intravenous fluids, vasopressors, and inotropes in response to continuously changing physiological states during the dynamic postoperative environment of cardiac patients.

[0015] The present disclosure provides RL-based techniques for insulin titration and hemodynamic management that addresses the limitations of current post-surgery protocols forAttorney Docket Number: MS-0045-01-WO cardiac patients. By continually learning from individual patient data, the RL-based technique(s) of the present disclosure provides personalized treatment plans that account for specific patient variability and maintain glucose and hemodynamic parameters in optimal or otherwise suitable ranges. Additionally, the present models’ capabilities to adapt to rapidly changing clinical characteristics enables insulin dosing or hemodynamic management to remain optimal or otherwise suitable as patient conditions evolve. In comparison, offline RL, where the agent learns from a fixed dataset without further interaction with the environment, often fails to account for the full spectrum of possible patient outcomes. As a result, offline RL techniques do not adequately address the diverse risk profiles associated with different patient actions, potentially compromising the safety and effectiveness of clinical interventions.

[0016] In some implementations, the present disclosure addresses these limitations by integrating a CQL model with distributional RL, which characterizes the entire distribution of potential outcomes rather than just the expected reward. This methodology provides a more comprehensive understanding of the risks and benefits associated with various actions, allowing for more nuanced decision-making under uncertainty. By considering the full range of potential patient responses, distributional RL can enhance the personalization and safety of insulin titration and hemodynamic management protocols, enabling optimal or otherwise suitable dosing as patient conditions change.

[0017] The present disclosure also utilizes a curated data set and select computational constraints for the models to optimize or otherwise improve patient safety while also encompassing the full scope of possible treatment regimens. In addition, the models demonstrate accuracy and efficacy in the models’ recommended treatment protocols.

[0018] One aspect of the present disclosure provides a method for generating an output for the administration of an insulin therapy for a post operative cardiac patient, the method comprising: receiving patient information including blood glucose levels of said patient; providing said patient information as input to a conservative Q-leaming (CQL) model that has been trained on data from cardiac patients using a reward function that includes a safety component that penalizes blood glucose levels below about 70 mg / dL and above about 180 mg / dL and a progress component that rewards time where the blood glucose levels are between about 70 mg / dL and about 180 mg / dL for determining whether to administer insulin to said patient; wherein said CQL model is integrated with a distributional reinforcementAttorney Docket Number: MS-0045-01-WO learning (RL) model; and generating, with said CQL model based on said patient information, an output for the administration of an insulin dose.

[0019] In certain embodiments of said method, said insulin dose is for administration during the first 24 hours after cardiac surgery. In certain embodiments, said distributional RL model incorporates an implicit quantile network (IQN). In certain embodiments, said insulin dose is for administration during a first 24 hours after cardiac surgery. In certain embodiments, said patient information is selected from a group comprising: demographics, comorbidities, laboratory values, vital signs, medications, mechanical ventilation status, body mass index (BMI), race, sequential organ failure assessment (SOFA) scores, insulin dosing, vasopressor and inotrope doses, and patient history. In certain embodiments, said method further comprises providing subsequent patient information to said integrated model; and generating, with said integrated model based on said subsequent patient information, an output indicative of subsequent insulin doses at least on an hourly basis. In certain embodiments, said CQL model includes a multilayer perceptron (MLP) matrix having three 512-dimension hidden layers. In certain embodiments, said CQL model utilizes a Markov decision process (MDP) and wherein states of said MDP are selected from: demographics, comorbidities, laboratory values, vital signs, medications, mechanical ventilation, and sequential organ failure assessment (SOFA) scores. In certain embodiments, said states are provided in hourly increments. In certain embodiments, said method further comprises calculating a ratio of glucose change to insulin dose for each hour. In certain embodiments, there is a maximum reward of +0.2 within a 140- 180mg / dL range and penalties become increasingly negative, down to -I, outside of that range. In certain embodiments, the maximum number of insulin doses per hour is 10 doses. In certain embodiments, said CQL model is trained on prior patient information that excludes information from patients who: died within first 24 hours of intensive care unit (ICU) admission; had ambiguous medication administration information; or lack information concerning available glucose levels within first three hours of documented ICU admission time after surgery. In certain embodiments, said CQL model is trained on prior patient information that excludes information from: patients who received short acting insulins. In certain embodiments, said patient information is extracted as multidimensional discrete time series in 1 hour time intervals and wherein the information features are summed or averaged as clinically appropriate. In certain embodiments, during training of the CQL model, the CQL model learning rate is set at le'4; the critic learning rate is set at 3e'4; the discount factor y is set at 0.67; and a dropout rateAttorney Docket Number: MS-0045-01-WO of p=0.01 is applied. In certain embodiments the output from said method is used to treat a patient in need thereof.

[0020] One aspect of the present disclosure provides a system comprising: one or more processors; and one or more non-transitory computer-readable media, coupled to the one or more processors and storing instructions which, when executed by the one or more processors, cause the one or more processors to perform or control performance of operations comprising: receiving patient information including blood glucose levels of said patient; providing said patient information as input to a conservative Q-leaming (CQL) model that has been trained on data from cardiac patients using a reward function that includes a safety component that penalizes blood glucose levels below about 70 mg / dL and above about 180 mg / dL and a progress component that rewards time where the blood glucose levels are between about 70 mg / dL and about 180 mg / dL for determining whether to administer insulin to said patient; wherein said CQL model is integrated with a distributional reinforcement learning (RL) model; and generating, with said CQL model based on said patient information, an output for the administration of an insulin dose.

[0021] In certain embodiments of said system, said insulin dose is for administration during the first 24 hours after cardiac surgery. In certain embodiments, said distributional RL model incorporates an implicit quantile network (IQN). In certain embodiments, said insulin dose is for administration during a first 24 hours after cardiac surgery. In certain embodiments, said patient information is selected from a group comprising: demographics, comorbidities, laboratory values, vital signs, medications, mechanical ventilation status, body mass index (BMI), race, sequential organ failure assessment (SOFA) scores, insulin dosing, vasopressor and inotrope doses, and patient history. In certain embodiments, said method further comprises providing subsequent patient information to said integrated model; and generating, with said integrated model based on said subsequent patient information, an output indicative of subsequent insulin doses at least on an hourly basis. In certain embodiments, said CQL model includes a multilayer perceptron (MLP) matrix having three 512-dimension hidden layers. In certain embodiments, said CQL model utilizes a Markov decision process (MDP) and wherein states of said MDP are selected from: demographics, comorbidities, laboratory7values, vital signs, medications, mechanical ventilation, and sequential organ failure assessment (SOFA) scores. In certain embodiments, said states are provided in hourly increments. In certain embodiments, said method further comprises calculating a ratio of glucose change to insulinAttorney Docket Number: MS-0045-01-WO dose for each hour. In certain embodiments, there is a maximum reward of +0.2 within a 140- 180mg / dL range and penalties become increasingly negative, down to -1, outside of that range. In certain embodiments, the maximum number of insulin doses per hour is 10 doses. In certain embodiments, said CQL model is trained on prior patient information that excludes information from patients who: died within first 24 hours of intensive care unit (ICU) admission; had ambiguous medication administration information; or lack information concerning available glucose levels within first three hours of documented ICU admission time after surgery. In certain embodiments, said CQL model is trained on prior patient information that excludes information from: patients who received short acting insulins. In certain embodiments, said patient information is extracted as multidimensional discrete time series in 1 hour time intervals and wherein the information features are summed or averaged as clinically appropriate. In certain embodiments, during training of the CQL model, the CQL model learning rate is set at le'4; the critic learning rate is set at 3e'4; the discount factor y is set at 0.67; and a dropout rate of p=0.01 is applied. In certain embodiments said output is used to treat a patient in need thereof.

[0022] One aspect of the present disclosure provides a system for determining an insulin therapy for a post operative cardiac patient comprising: one or more processors; and one or more non-transitory computer-readable media, coupled to the one or more processors and storing instructions which, when executed by the one or more processors, cause the one or more processors to perform or control performance of operations comprising: receiving patient information regarding a post operative cardiac patient, including blood glucose levels of said patient: providing said patient information as input to a conservative Q-leaming (CQL) model that has been trained on data from cardiac patients using a reward function that includes a safety component that penalizes blood glucose levels below about 70 mg / dL and above about 180 mg / dL and a progress component that rewards time where the blood glucose levels are betw een about 70 mg / dL and about 180 mg / dL for determining whether to administer insulin to said patient; w herein said CQL model is integrated with a distributional reinforcement learning (RL) model; wherein said distributional RL model incorporates an implicit quantile netw ork (1QN); wherein said CQL model includes a multilayer perceptron (MLP) matrix having three 512- dimension hidden layers; wherein said CQL model utilizes a Markov decision process (MDP); wherein the states of said MDP are selected from: demographics, comorbidities, laboratory values, vital signs, medications, mechanical ventilation, and sequential organ failure assessment (SOFA) scores; generating from said CQL model based on subsequent patientAttorney Docket Number: MS-0045-01-WO information an output; and wherein said output provides information for diagnosis or treatment. In certain embodiments, said output said information was for diagnosis or treatment of said patient.

[0023] One aspect of the present disclosure provides a method for hemodynamic management in a post operative cardiac patient, the method comprising: receiving patient information for adult patients that were admitted after cardiac surgery; providing said patient information as input to a conservative Q-leaming (CQL) model that has been trained on training features from cardiac patient training data, wherein said cardiac patient training data is from patients admitted to an intensive care unit (ICU) after cardiac surgery and wherein said model has been integrated with an implicit quantile network (IQN); generating, with said CQL model, based on said patient information, an output regarding whether said patient is at risk of acute kidney injury (AKI), wherein said output includes a list of treatment actions, optionally including administering one or more medications.

[0024] In certain embodiments, said training features include one or more of: demographics, anthropometries, vital signs, laboratory’ values, fluid balance, ions, and sequential organ failure assessment (SOFA) Score. In certain embodiments, said medications include one or more of: IV fluids, vasopressors, inotropes, and nephrotoxins. In certain embodiments, the reward function for the CQL is defined as: a terminal penalty of -5 at a final time step for cardiac patient training data from a patient that developed persistent AKI (pAKI) in a first 5 days of ICU stay; a terminal reward of +5 at the final time step for cardiac patient training data from a patient that did not develop pAKI in the first 5 days of ICU stay; a penalty of -2.5 at any time-step t for any patient training data showing a drop in mean arterial pressure bel ow 65mm Hg at a next step t+ 1 ; and a reward of + 1 at any time-step t for any pati ent training data showing that the patient does not meet the criteria for AKI. In certain embodiments, wherein during training of the CQL model, the discount factor rate is set at 0.99 and a conservativeness coefficient is set to 0.25. In certain embodiments, the training features that changed over time are restricted to information obtained within the earlier of 72 hours after ICU admission or up until ICU discharge, the features are segmented on an hourly basis, features for a patient are excluded where more than 30% of the patient information is missing, and forward-fill imputation and k-nearest neighbors are applied for any remaining missing information. In certain embodiments, information from patients with end stage kidney disease, patients who died within 72 hours of admission, and patients with no recorded values ofAttorney Docket Number: MS-0045-01-WO creatinine or IV fluids is excluded from the training features used to train the model. In certain embodiments, the IQN uses 62 sample quantiles and 32 of these are designated as greedy quantiles. In certain embodiments, said CQL network contains four hidden layers of 512 neurons each and a rectified linear unit (ReLU) activation function. In certain embodiments, a patient is treated based on said output.

[0025] One aspect of the present disclosure provides a system comprising: one or more processors; and one or more non-transitory processor-readable media, coupled to the one or more processors and storing instructions which, when executed by the one or more processors, cause the one or more processors to perform or control performance of operations comprising: receiving patient information for adult patients that were admitted after cardiac surgery; providing said patient information as input to a conservative Q-leaming (CQL) model that has been trained on training features from cardiac patient training data, wherein said cardiac patient training data is from patients admitted to an intensive care unit (ICU) after cardiac surgery' and wherein said model has been integrated with an implicit quantile network (IQN); generating, with said CQL model based on said patient information, an output regarding whether said patient is at risk of acute kidney injury' (AKI), wherein said output includes a list of treatment actions, optionally including administering one or more medications.

[0026] In certain embodiments, said training features include one or more of: demographics, anthropometries, vital signs, laboratory' values, fluid balance, ions, and sequential organ failure assessment (SOFA) Score. In certain embodiments, said medications include one or more of: IV fluids, vasopressors, inotropes, and nephrotoxins. In certain embodiments, the reward function for the CQL is defined as: a terminal penalty of -5 at a final time step for cardiac patient training data from a patient that developed persistent AKI (pAKI) in a first 5 days of ICU stay; a terminal reward of +5 at the final time step for cardiac patient training data from a patient that did not develop pAKI in the first 5 days of ICU stay; a penalty of -2.5 at any time-step t for any patient training data showing a drop in mean arterial pressure below 65mm Hg at a next step t+1; and a reward of +1 at any time-step t for any patient training data showing that the patient does not meet the criteria for AKI. In certain embodiments, wherein during training of the CQL model, the discount factor rate is set at 0.99 and a conservativeness coefficient is set to 0.25. In certain embodiments, the training features that changed over time use information obtained within the earlier of 72 hours after ICU admission or up until ICU discharge, the features are segmented on an hourly basis, features for a patientAttorney Docket Number: MS-0045-01-WO are excluded where more than 30% of the patient information is missing, and forward-fill imputation and k-nearest neighbors are applied for any remaining missing information. In certain embodiments, information from patients with end stage kidney disease, patients who died within 72 hours of admission, and patients with no recorded values of creatinine or IV fluids is excluded from the training features used to train the model. In certain embodiments, the IQN uses 62 sample quantiles and 32 of these are designated as greedy quantiles. In certain embodiments, said CQL network contains four hidden layers of 512 neurons each and a rectified linear unit (ReLU) activation function. In certain embodiments, a patient in need thereof is treated based on the output of said system.BRIEF DESCRIPTION OF THE DRAWINGS

[0027] Figure 1 provides the study overview as follows: a) Schema for development, testing, and selection of embodiments of the GLUCOSE model; b) Overview of clinician validation study.

[0028] Figure 2 provides performance of an embodiment of the GLUCOSE model as follows: a) Off policy evaluation (OPE) counterfactual estimated performance of the model computed by fitted Q estimation (FQE) (solid lines) compared to the returns by the treating clinicians (dotted lines) with 95% CI; b) Average time in range and average glucose levels: Across all datasets, patients who received insulin doses similar to those recommended by the GLUCOSE model had the highest average time in range (TIR), defined as the number of hours with blood glucose between 70-180 mg / dL. This trend was observed in internal testing (Medical Information Mart for Intensive Care (MIMIC-IV)), external validation (eICU), and a more recent external validation dataset at Mount Sinai (Sinai), covering Jan 1. 2023, through Apr 30, 2025. TIR decreased as the difference between model-recommended and clinician- administered doses increased, suggesting that the model suggests different amounts of insulin to improve glycemic control, while patients who had the most similar recommendations to the model achieved the most optimal blood glucose control. All analyses were conducted on a HIPAA-compliant server hosted on Mount Sinai Scientific Computing and Data’s Minerva High Performance Computing Platform, c) Average Insulin Doses: The action distribution of the histogram highlights expected dosing patterns. In the left panel (MIMIC), and the center panel (eICU), the GLUCOSE model consistently recommends lower average insulin doses at low glucose values, reflecting a guideline-aligned strategy that prioritizes minimizing the risk of hypoglycemia. Additionally, more insulin is given at higher glucose values. Therefore, thisAttorney Docket Number: MS-0045-01-WO overall demonstrates a proactive approach to correcting significant hyperglycemia while minimizing the risk of hypoglycemia. The right panel (Sinai) shows results from a more recent validation cohort at a large tertiary academic medical center that adopted the most recent Society of Thoracic Surgeons guidelines between Jan 1, 2023, and Apr 30, 2025. This distribution more closely matches the GLUCOSE model’s recommendations, indicating improved compliance to this guideline by treating physicians. All analyses were conducted on a HIPAA-compliant server hosted on Mount Sinai Scientific Computing and Data’s Minerva High Performance Computing Platform.

[0029] Figure 3 provides the human validation study results for an embodiment of the GLUCOSE model as follows: a) MAE of clinician groupings and GLUCOSE relative to an endocrinologist baseline with standard error of the mean; and b) Blinded ratings across safety, effectiveness, and acceptability of all clinicians by a blinded senior intensivist panel with standard error of the mean.

[0030] Figure 4 provides case examples for an embodiment of the GLUCOSE model as follows: representative case examples to compare the GLUCOSE model’s insulin dosing recommendations compared with actual clinician administered insulin in internal testing (a-c) and external validation (d-f) patients. Lines indicate glucose levels and insulin doses. Bands indicate glycemic ranges.

[0031] Figure 5 provides a time in range analy sis stratified by subgroup for an embodiment of the GLUCOSE model as follows: average time in range ± the standard error of the mean shown via solid lines, and dashed lines indicate the average glucose level corresponding to that subgroup.

[0032] Figure 6 provides feature importances for an embodiment of the GLUCOSE model as follows: a) ranked feature importances by mean absolute SHapley Additive exPlanations (SHAP) value and b) relative contributions of varying categorizations of features.

[0033] Figure 7 provides an analysis of predictions of an embodiment of the GLUCOSE model in excluded subsets as follows: average insulin values in varying clinically relevant ranges are shown in each of the subsets, and 95% confidence intervals are shown for each bar.

[0034] Figure 8 provides a CONSORT diagram for an embodiment of the GLUCOSE model as follows: inclusion / exclusion criteria applied during data preprocessing.Attorney Docket Number: MS-0045-01-WO

[0035] Figure 9 provides the overall structure of the RL-based personalization mechanism for an embodiment of the RENAL model.

[0036] Figure 10 provides a comparison of the results from an embodiment of the RENAL model and from clinician Off policy evaluation (OPE) results.

[0037] Figures I la and 1 lb provide a comparison of the results from an embodiment of the RENAL model vs clinician actions over time.

[0038] Figures 12a and 12b provide a comparison of the results from an embodiment of the RENAL model vs clinician actions across MAPs.

[0039] Figure 13 provides the U-curve analysis of the results of an embodiment of the RENAL model.

[0040] Figure 14 provides an individual assessment of the per patient behavior timeline of clinician versus an embodiment of the RENAL model.

[0041] Figure 15 provides a SHapley Additive exPlanations (SHAP) analysis of a trained RENAL model’s features.

[0042] Figure 16 provides the consort diagram for an embodiment of the RENAL model.

[0043] Figures 17a-17e provide the Off policy evaluation (OPE) analysis by acute kidney injury AKI status in the first 3 days, by sex, race, BMI, and age.

[0044] Figure 18 provides a schematic of a patient diagnosis and treatment system, according to an implementation of the present disclosure.

[0045] Figure 19 provides an illustration of a computing device used in the implementation of Figure 19.

[0046] Figure 20 provides an illustration of a method used in the implementation of Figure 19.

[0047] Figure 21 provides an illustration of certain embodiments of processing models used in Figure 19.Attorney Docket Number: MS-0045-01-WO

[0048] For the purpose of illustrating the disclosure, there are depicted in drawings certain embodiments and implementations. However, the precise arrangements and instrumentalities of the embodiments and implementations depicted in the drawings are not intended to limit the disclosure herein.DETAILED DESCRIPTION

[0049] Example embodiments and implementations consistent with the teachings included in the present disclosure are directed to diagnostic systems 100 and methods 1 100 configured to generate outputs and signals for treating post-surgical cardiac patients.

[0050] Included in example embodiments of the present disclosure are models and systems for diagnosing and treating post-surgical cardiac patients, including by managing blood glucose levels (GLUCOSE models) and / or hemodynamics (RENAL models). Certain embodiments of both the GLUCOSE and RENAL models integrate conservative Q learning (CQL), an offline RL algorithm, with distributional reinforced learning (RL) (specifically implicit quantile networks (IQN)), to process data from post-surgical cardiac patients and generate outputs for use in diagnosing and / or treating said patients.

[0051] The GLUCOSE model provides embodiments of the present disclosure that are useful for the management of glucose levels in post-operative cardiac surgery patients and that are useful for the prevention of hyperglycemia in cardiac patients. Embodiments of the GLUCOSE model include the systems, methods and non-transitory media are provided below.

[0052] The RENAL model provides embodiments of the present disclosure that are useful for hemodynamic management in post-operative cardiac surgery patients, including for the management and prevention of acute kidney injury (AKI) and persistent AKI. Embodiments of the RENAL model include the systems, methods and non-transitory media are provided below.

[0053] An illustration of certain embodiments of the models and systems is provided in Figure 18, which provides in an implementation consistent with the disclosure. In the figure, the patient diagnostic and treatment system 100 includes a hardware-based processor 102, a memory 104 configured to store instructions and configured to provide the instructions to the hardware-based processor 102, a communication interface 106, an input / output device 108,Attorney Docket Number: MS-0045-01-WO and a set of models 110-116 configured to implement the instructions provided to the hardwarebased processor 102. The set of models 110-116 optionally includes: a conservative Q learning model (CQL) 110, a distributional reinforcement learning (RL) model 112; one or more implicit quantile learning networks (IQN) 114; one or more multilayer perceptron matrices 116. In an implementation, the instructions are code written in any known programming language.

[0054] An implementation of certain methods consistent with the disclosure as provided in Figure 18 provides: receiving subsequent patient data 124 as input into a conservative Q- leaming (CQL) model 110 that has been trained on cardiac patient training data 122; wherein said CQL model 110 is integrated with a distributional reinforcement learning (RL) model 112, which includes one or more implicit quantile learning networks 114, wherein said CQL model 110 includes a multilayer perceptron matrix 116, and generating, with said CQL model that w as previously trained on said patient data 122 and based on said subsequent patient data 124. an output for the administration of an insulin dose to said patient.

[0055] An implementation of certain methods consistent with the disclosure as provided in Figure 19 provides: receiving subsequent patient data 124 as input into a conservative Q- leaming (CQL) model 110 that has been trained on cardiac patient training data 122; wherein said CQL model 110 is integrated with a distnbutional reinforcement learning (RL) model 112. which includes one or more implicit quantile learning networks 114, and generating, with said CQL model that was previously trained on said patient data 122 and based on said subsequent patient data 124, an output for the hemodynamic management, which optionally includes administering medicine.

[0056] To further the evaluation of uncertainty' and risk a CQL 110 is integrated with distributional reinforcement learning (RL) 112, an approach that characterizes the entire distribution of potential outcomes rather than focusing only on the expected reward. By capturing the full range of possible patient responses, this methodology enables more nuanced decision-making, particularly for rare but significant events such as hypoglycemia, which traditional RL methods may underestimate.

[0057] In one implementation, the patient diagnostic and treatment system 100 is operatively connected to a data source 120 through a network. For example, the netw ork is the Internet. In another example, the network is an internal netw ork or intranet of an organization.Attorney Docket Number: MS-0045-01-WOIn a further example, the network is a heterogeneous or hybrid network including the Internet and the intranet. The data source 120 transmits, conveys, or otherwise provides training data 122 and patient data 124 to the diagnostic system 100.

[0058] In one implementation, the data source 120 is in proximity to the diagnostic system 100. In another implementation, the data source 120 is remote from the diagnostic system 100. In a further implementation, the training data 122 is obtained from a database of patient information. For example, the training data 122 optionally includes data obtained from the Medical Information Mart for Intensive Care (MIMIC-IV); Salzburg Intensive Care Database (SICdb); and the eICU Collaborative Research Database (elCU-CRD). In one implementation, the training data 122 are formatted as structured extensible markup language (XML) files including both raw data and metadata associated with patient identifiers, time, place, indication, and characteristics such as diagnoses of the patients associated with each of the training data 122. In a further implementation, the training data 122 optionally includes data from a pilot study. In another implementation, the training data 122 are formatted in any know n data format.

[0059] In one implementation, the patient data 124 is received in real time. Such real time data 124 is temporarily or permanently stored in the data source 120, and is transmitted, conveyed, or otherwise provided to the communication interface 106 of the diagnostic system 100. In another implementation, the patient data / information 124 data are formatted as structured extensible markup language (XML) files including both raw data and metadata associated with patient identifiers, time, place, indication, and additional patient information received from the data source 120. In further implementations, the patient data 124 are formatted in any known data format. In one implementation, the output 126 is an alert, a notification, or a message output from the input / output device 108 containing a treatment recommendation. For example, the input / output device 108 includes a display or monitor configured to visually display the treatment recommendation as an output 126 to a doctor, a technician, or a patient. The output 126 is a text message or an image representing the diagnostic state of the patient, such as the patient corresponding to the patient data 124, as requiring insulin, hemodynamic management, or intervention to treat an emergent cardiac condition.

[0060] In another example, the input / output device 108 includes an audio speaker configured to output or signal 126 using an audible sound, corresponding to the patientAttorney Docket Number: MS-0045-01-WO treatment recommendation , to a doctor, a technician, or a patient. In a further example, the input / output device 108 include both a display and an audio speaker, and the signal or output 126 providing the patient diagnosis or treatment recommendation includes a video or animation with audio conveying that the patient, corresponding to the patient 124, requires insulin, hemodynamic management, or other intervention to treat an emergent cardiac condition.

[0061] Figure 19 illustrates an embodiment of the present disclosure and provides a schematic of a computing device 200 including a processor 202 having code therein, a memory 204, and a communication interface 206. Optionally, the computing device 200 can include a user interface 208, such as an input device, an output device, or an input / output device. The processor 202, the memory 204, the communication interface 206, and the user interface 208 are operatively connected to each other via any known connections. Any component, combination of components, and modules of the system 100 in Figure 19 can be implemented by a respective computing device 200. For example, each of the hardwarebased processor 102, the memory 104, the communication interface 106, the input / output device 108, and the set of modules 110-116 shown in Figure 18 can be implemented by a respective computing device 200 shown in Figure 19 and described below'.

[0062] In embodiments of the present disclosure, the computing device 200 can include different components. In certain implementations, the computing device 200 can include additional components. In another alternative implementation, the functions of a given component can instead be carried out by one or more multiple different components. The computing device 200 can be implemented by a virtual computing device, using a cloud computing environment, or by by a plurality of any known computing devices.

[0063] The processor 202 can be a hardware-based processor implementing a system, a sub-system, or a module. The processor 202 can include one or more general-purpose processors. Alternatively, the processor 202 can include one or more special-purpose processors. The processor 202 can be integrated in whole or in part with the memory 204, the communication interface 206, and the user interface 208. In another alternative implementation, the processor 202 can be implemented by any known hardware-based processing device such as a controller, an integrated circuit, a microchip, a central processing unit (CPU), a microprocessor, a system on a chip (SoC), a field-programmable gate array (FPGA), or an application-specific integrated circuit (ASIC). In addition,Attorney Docket Number: MS-0045-01-WO the processor 202 can include a plurality of processing elements configured to perform parallel processing. In a further alternative implementation, the processor 202 can include a plurality of nodes or artificial neurons configured as an artificial neural network. The processor 202 can be configured to implement any known machine learning (ML) based devices, any known artificial intelligence (Al) based devices, and any known artificial neural networks, including a recursive neural network (RNN) or a convolutional neural network (CNN).

[0064] The memory 204 can be implemented as a non-transitory computer-readable storage medium such as a hard drive, a solid-state drive, an erasable programmable read-only memory (EPROM), a universal serial bus (USB) storage device, a floppy disk, a compact disc read-only memory (CD-ROM) disk, a digital versatile disc (DVD), cloud-based storage, or any known non-volatile storage.

[0065] The code of the processor 202 can be stored in a memory internal to the processor 202. The code can be instructions implemented in hardware. Alternatively, the code can be instructions implemented in software. The instructions can be machine-language instructions executable by the processor 202 to cause the computing device 200 to perform the functions of the computing device 200 described herein. Alternatively, the instructions can include script instructions executable by a script interpreter configured to cause the processor 202 and computing device 200 to execute the instructions specified in the script instructions. In another alternative implementation, the instructions are executable by the processor 202 to cause the computing device 200 to execute an artificial neural network. The processor 202 can be implemented using hardware or software, such as the code. The processor 202 can implement a system, a sub-system, or a module, as described herein.

[0066] The memory 204 can store data in any known format, such as databases, data structures, data lakes, or network parameters of a neural netw ork. The data can be stored in a table, a flat file, data in a filesystem, a heap file, a B+ tree, a hash table, or a hash bucket. The memory 204 can be implemented by any known memory, including random access memory (RAM), cache memory. register memory, or any other known memory device configured to store instructions or data for rapid access by the processor 202, including storage of instructions during execution.

[0067] The communication interface 206 can be any known device configured to perform the communication interface functions of the computing device 200 described herein.Attorney Docket Number: MS-0045-01-WOThe communication interface 206 can implement wired communication between the computing device 200 and another entity. Alternatively, the communication interface 206 can implement wireless communication between the computing device 200 and another entity. The communication interface 206 can be implemented by an Ethernet, Wi-Fi, Bluetooth, or USB interface. The communication interface 206 can transmit and receive data over a network and to other devices using any known communication link or communication protocol.

[0068] The user interface 208 can be any known device configured to perform user input and output functions. The user interface 208 can be configured to receive an input from a user. Alternatively, the user interface 208 can be configured to output information to the user. The user interface 208 can be a computer monitor, a television, a loudspeaker, a computer speaker, or any other known device operatively connected to the computing device 200 and configured to output information to the user. A user input can be received through the user interface 208 implementing a keyboard, a mouse, or any other known device operatively connected to the computing device 200 to input information from the user. Alternatively, the user interface 208 can be implemented by any known touchscreen. The computing device 200 can include a server, a personal computer, a laptop, a smartphone, or a tablet.

[0069] Referring to Figure 18, the patient diagnostic and treatment system 100 is configured to receive the training data 122 at the communication interface 106 from the data source 120, and to store the received training data 122 in the memoiy 104. The training data 122 is processed by the processor 102 to train the conservative Q learning model 110, which is integrated with a distributional reinforcement learning model, that optionally includes implicit quantile learning (IQN) networks 112. Once trained, the patient diagnosis and treatment system 100 receives, from the data source 120, the data 124 corresponding to a patient at the communication interface 106. The patient diagnosis and treatment system 100 stores the received patient data 124 in the memory 104. As described below, the trained vision transformer system 100 processes the patient data 124 to generate and output, from the input / output device 108, a treatment recommendation 126 of the patient corresponding to the patient input data 124.

[0070] Referring to Figure 20, it provides an implementation consistent with the present disclosure, showing a computer-based method 1100 that implements the patient diagnosis and treatment system 100 configured to generate a patient diagnosis 126 from the patient data 124.Attorney Docket Number: MS-0045-01-WOThe computer-based method 1110 includes receiving training data 122 in step 1102, training at least a conservative Q learning model 110, which is integrated with a distributional reinforcement learning (RL) model 112, using the patient training data 122 and the patient data 124 and generating an output including information for patient diagnosis and / or treatment 1126. Alternatively, the computer-based method 1110 includes receiving training data 122 in step 1104, training at least a conservative Q learning model 110, which is integrated with a distributional reinforcement learning (RL) model 112, which incorporates implicit quantile networks 114, using the patient training data 122 and the patient data 124 and generating an output including information for patient diagnosis and / or treatment 1126.

[0071] In an implementation consistent with the disclosure, a non-transitory computer- readable storage medium stores instructions executable by a processor to implement the vision transformer system 100 and method 1100 configured to generate a patient diagnosis 126 from the patient data 124. The instructions include receiving the training data 122, training at least the conservative Q learning (CQL) model 110, which is integrated with at least one distributional reinforcement learning (RL) model 112, using the training data 122, receiving data 124 of a patient to be diagnosed, and generating an output 126 for diagnosis and / or treatment of said patient. In an implementation, the CQL model optionally includes at least one multilayer perceptron matrix 116. In a further implementation, the RL model 112 further includes implicit quantile (IQN) networks 114.

[0072] As shown in Figure 21, in certain embodiments, the CQL model 110, which is integrated with a distributional reinforcement learning (RL) model 112, also uses a multi-layer perceptron (MLP) network 116, which optionally includes three 512-dimension hidden layers. Optionally, at least one implicit quantile network (IQN) 114 is incorporated into the distributional RL model 112 in certain embodiments, thereby leveraging the strengths of distributional RL to better model the variability' in patient responses, thereby improving the robustness of those embodiments.

[0073] Unlike other distributional methods that require explicitly defining the number of quantiles, IQN adds a layer which flexibly leams the full return distribution by sampling from continuous quantile values during training, allowing it to approximate the entire outcome distribution without fixed bins. This enables a more comprehensive representation of potential outcomes while improving upon its non-distributional counterparts. As a result, such modelsAtorney Docket Number: MS-0045-01-WO can beter capture clinical uncertainty and make decisions around nuanced risk profiles, particularly in setings with high variability.

[0074] In certain embodiments, a batch training sampling strategy was employed for offline RL, thus avoiding overregularization by low-return actions, allowing the learned policy to reflect more high-retum trajectories.

[0075] Embodiments of the RENAL model utilize a CQL 110 for similar reasons. The CQL is designed to mitigate overestimation of action values in offline setings by penalizing actions not well supported by the training data, thereby reducing the likelihood of unsafe or clinically implausible recommendations for hemodynamic management. In certain embodiments, the CQL is integrated with the IQN 114, which enables the model to encompass the entire distribution of possible future outcomes. This richer representation is useful for clinical decision-making, as it enables the model to account for rare but significant risks, such as acute kidney injury or hypotension. IQN achieves this by learning a quantile-based approximation of the return distribution. Instead of relying on a fixed number of bins, IQN samples quantiles from a continuous distribution, allowing it to flexibly capture a wide range of potential outcomes, including rare but significant risks, such as acute kidney injury or hypotension, which would be missed otherwise. This structure not only allows the model to understand the full spectrum of possible future states but also enables fine-tuning of risk preferences via greedy quantile selection.

[0076] In certain embodiments, an implementation of the RENAL model uses 64 sampled quantiles, with 32 designated as greedy quantiles to guide action selection. This combination of CQL and IQN enables embodiments of the RENAL model to make more informed and risksensitive decisions, especially for interventions with narrow therapeutic indices. In certain embodiments, the CQL implementation of this model used an encoder network with four hidden layers of 512 neurons each and rectified linear unit (ReLU) activation function.

[0077] A GLUCOSE model was trained and validated on two separate ICU datasets: the MIMIC -IV database and the elCU-CRD database. MIMIC-IV was used as the development cohort and split into training and internal testing sets. elCU-CRD was used as the external validation dataset. Based on the inclusion and exclusion criteria, information from 6,148 patients was included in the training dataset and from 649 patients was included in an external validation dataset. The mean age of patients in the training dataset was 67.8 ± 11.6 years withAttorney Docket Number: MS-0045-01-WO71.1% males, and in the external validation dataset was 67.0 ± 11.3 years with 67.2% males. At least one hypoglycemic event ( < 70 mg / dL) occurred in 7.6% of patients in the training dataset and among 7.2% of patients in the external validation dataset. Similarly, at least one hyperglycemic event ( > 180 mg / dL) occurred in 47.8% of patients in the training dataset and 47.3% of patients in the external validation dataset. The baseline characteristics of the patients are shown in Tables 2 and 3. The overall structure of the design, development and testing of an embodiment of the GLUCOSE model is illustrated in Figure 1.

[0078] The dataset included all adult patients (age >18 years) who were admitted to ICU after cardiac surgery. ICD-9-PCS and ICD-10-PCS codes were used to identify patients who underwent cardiac surgery in MIMIC-IV database (Table 4). The elCU-CRD database does not include 1CD-9-PCS or ICD-10-PCS procedure codes. The inputs included patients admitted to the ICU after cardiac surgery7using the “admissiondx” table that provides the primary diagnosis for ICU admissions. The inputs excluded patients who died within first 24 hours of ICU admission, had ambiguous medication administration information such that it did not allow calculation of the exact dose of medication administered, or did not have available glucose levels within first three hours of documented ICU admission time after surgery. To develop a system that would provide personalized administration of regular insulin, patients who received other short acting insulins (aspart, lispro, NPH, insulin 70 / 30) were also excluded.

[0079] An embodiment of the RENAL model used data from Medical Information Mart for Intensive Care (MIMIC-IV) database as the for the training and Salzburg Intensive Care Database (SICdb) as the external validation cohort. MIMIC-IV is a publicly available clinical data developed by MIT laboratory for computational physiology7. It contains data from ICUs at Beth Israel Deaconess Medical Center from years 2008-2019. SICdb on the other hand is a database of 27,000 ICU admissions to the University Hospital Salzburg over 2013-2021. The overall structure of the design, development, and testing of embodiments of the RENAL model is shown in Figure 9.

[0080] The dataset included adult patients (>18 years) admitted to the ICU after cardiac surgery. In MIMIC-IV, cardiac surgery patients were identified using appropriate ICD-10-PCS codes. In SICdb, cardiac surgery patients were identified based on the use of "heartsurgerybeginoffset" variable provided in the database. Patients with end stage kidney disease and those with no recorded values of creatinine or IV fluids w ere excluded. The model was trained on rich, high-resolution electronic health record data encompassing vital signs,Attorney Docket Number: MS-0045-01-WO laboratory' values, medication administration, and clinical interventions over time. The training data focused specifically on decisions hemodynamic management of patients in the time since their admission to ICU post-surgery and the lesser of seventy-two hours and ICU discharge. Patients who died within the first 72-hour period were also excluded to avoid including cases with non-representative clinical trajectories for training of the RL model (see Table 3; Figure 16).

[0081] Unless defined otherwise, all technical and scientific terms used herein have the same meaning as is commonly understood by one of ordinary' skill in the art to which this disclosure belongs. All patents and publications referred to herein are incorporated by reference in their entireties.Definitions

[0082] When ranges are used herein to describe, for example, physical or chemical properties such as molecular weight or chemical formulae, all combinations and subcombinations of ranges and specific embodiments therein are intended to be included. Use of the term “about” when referring to a number or a numerical range means that the number or numerical range referred to is an approximation within experimental variability (or within statistical experimental error), and thus the number or numerical range may vary. The variation is typically from 0% to 15%, or from 0% to 10%, or from 0% to 5% of the stated number or numerical range.

[0083] As used interchangeably herein, the term “classifier” or “model” refers to a machine learning model or algorithm.

[0084] In some embodiments, a model is a reinforcement learning (RL) model. Nonlimiting examples of RL models include Q-leaming. conservative Q-Leaming, SARSA, deep Q-Networks (DQN), and policy gradient methods like actor-critic and proximal policy optimization (PPO). In some embodiments, a model is a distributional RL model. Nonlimiting examples of distributional RL models include implicit quantile networks, categorical deep Q- networks (DQN), quantile regression DQN, and fully parameterized quantile function.

[0085] Neural networks. In some embodiments, the model is a neural network (e.g., a convolutional neural network and / or a residual neural netw ork). Neural network algorithms, also known as artificial neural networks (ANNs), include convolutional and / or residual neuralAttorney Docket Number: MS-0045-01-WO network algorithms (deep learning algorithms). Neural networks can be machine learning algorithms that may be trained to map an input data set to an output data set, where the neural network comprises an interconnected group of nodes organized into multiple layers of nodes. For example, the neural network architecture may comprise at least an input layer, one or more hidden layers, and an output layer. The neural network may comprise any total number of layers, and any number of hidden layers, where the hidden layers function as trainable feature extractors that allow mapping of a set of input data to an output value or set of output values. As used herein, a deep learning algorithm (DNN) can be a neural network comprising a plurality of hidden layers (e.g., two or more hidden layers). Each layer of the neural network can comprise a number of nodes (or “neurons’"). A node can receive input that comes either directly from the input data or the output of nodes in previous layers, and perform a specific operation (e.g., a summation operation). In some embodiments, a connection from an input to a node is associated with a parameter (e.g., a weight and / or weighting factor). In some embodiments, the node may sum up the products of all pairs of inputs, xi, and their associated parameters. In some embodiments, the weighted sum is offset with a bias, b. In some embodiments, the output of a node or neuron may be gated using a threshold or activation function, f, which may be a linear or non-linear function. The activation function may be, for example, a rectified linear unit (ReLU) activation function, a leaky ReLU activation function, or other function such as a saturating hyperbolic tangent, identity, binary' step, logistic, arcTan, softsign, parametric rectified linear unit, exponential linear unit, softPlus. bent identity, softExponential, sinusoid, sine, gaussian, or sigmoid function, or any combination thereof.

[0086] The weighting factors, bias values, and threshold values, or other computational parameters of the neural network, may be “taught” or “learned” in a training phase using one or more sets of training data. For example, the parameters may be trained using the input data from a training data set and a gradient descent or backward propagation method so that the output value(s) that the ANN computes are consistent with the examples included in the training data set. The parameters may be obtained from a back propagation neural network training process.

[0087] F or the avoidance of doubt, it is intended herein that particular features (for example integers, characteristics, values, uses, diseases, formulae, compounds or groups) described in conjunction with a particular aspect, embodiment or example of the disclosure are to be understood as applicable to any other aspect, embodiment or example described herein unlessAttorney Docket Number: MS-0045-01-WO incompatible therewith. Thus, such features may be used where appropriate in conjunction with any of the definition, claims or embodiments defined herein. All the features disclosed in this specification (including any accompanying claims, abstract and drawings), and / or all of the steps of any method or process so disclosed, may be combined in any combination, except combinations where at least some of the features and / or steps are mutually exclusive. The disclosure is not restricted to any details of any disclosed embodiments. The disclosure extends to any novel one. or novel combination, of the features disclosed in this specification (including any accompanying claims, abstract and drawings), or to any novel one, or any novel combination, of the steps of any method or process so disclosed.

[0088] Moreover, as used herein, the term “about’7means that dimensions, sizes, formulations, parameters, shapes and other quantities and characteristics are not and need not be exact, but may be approximate and / or larger or smaller, as desired, reflecting tolerances, conversion factors, rounding off, measurement error and the like, and other factors known to those of ordinary skill in the art. In general, a dimension, size, formulation, parameter, shape or other quantity or characteristic is “about” or “approximate” whether or not expressly stated to be such. It is noted that embodiments of very different sizes, shapes and dimensions may employ the described arrangements.

[0089] In embodiments of the present disclosure, the methods, systems, and various implementations, systems, methods, and non-transitoiy computer-readable media generate “outputs” for patient treatment. Said “outputs” can include information for administering and / or withholding treatment over various periods of time and at various intervals.

[0090] Datasets: the Medical Information Mart for Intensive Care (MIMIC-IV); Salzburg Intensive Care Database (SICdb); and the eICU Collaborative Research Database (elCU-CRD) data sets used to train the GLUCOSE and RENAL models are publicly available. A Pilot Evaluation of the GLUCOSE model was performed at Mount Sinai and oversight was provided by Mount Sinai Assurance Lab. Names, dates, and identifying details were modified to protect patient privacy, in accordance with HIPAA regulations. Any necessary patient consents were obtained for the pilot study, and all data for the pilot study was stored on secure servers and anonymized per HIPAA requirements.EXAMPLESAttorney Docket Number: MS-0045-01-WOEXAMPLE 1 :Building the models: State Space, Action Space, and Rewards

[0091] RL frames decision-making problems as Markov decision processes (MDPs). An MDP is typically defined by a sequence of tuples (st,a,r,st-i ) at each time step t, where stis the observed feature vector representing the state at time t, a is the action taken, and r is the reward received for taking action aa in state st. The resulting states reflects the system's condition at the next time step following the action.

[0092] GLUCOSE model state space. For certain embodiments, features were derived from demographics, comorbidities, laboratory values, vital signs, medications, mechanical ventilation, and SOFA scores binned into hourly time-steps to develop the state space. The prior four hours of glucose values, when available, were also incorporated into the RL model. To provide additional context, information was included on glucose level changes during this period and calculated the ratio of glucose change to insulin dose for each hour, with the minimum insulin dose set at 0. 1 for this calculation.

[0093] RENAL model state space. For certain embodiments, the states included features derived from laboratory values, vital signs, medications (vasopressors, inotropes. and nephrotoxins), demographics, anthropometries (height, weight, BMI), SOFA score and IV fluids. Values of mean arterial pressure (MAP) and amount of IV fluids, vasopressors, and inotropes received in the 3-hours prior as part of the state space at time-step t were also included. In total, there were sixty-eight states, segmented into one-hour time-windows.

[0094] GLUCOSE model action space. In certain embodiments of the GLUCOSE model, actions are defined as the amount of regular insulin administered each hour, utilizing a continuous action space. For ease of interpretability, the recommended insulin doses were rounded to the nearest integer. This practice aligns the model’s recommendations with practical clinical standards and facilitates the clinical implementation of suggested doses. Additionally, the insulin doses recommended and observed by the GLUCOSE models were capped at a maximum of 10 units per hour. This decision was based on both data-driven considerations and clinical safety parameters. The mean hourly insulin dose in MIMIC-IV was 2.2 ± 3.0 units / hour, with doses exceeding 10 units / hour accounting for only 2.29% of all administered doses. From a clinical perspective, given insulin’s narrow therapeutic index, this cap serves as a safeguard to minimize the risk of hypoglycemia. It also reinforces the importance of maintaining aAttorney Docket Number: MS-0045-01-WO physician-in-the-loop framework for cases that may necessitate higher doses, ensuring safety and clinician oversight.

[0095] RENAL model action space. In certain embodiments of the RENAL model, the actions included recommending doses of IV fluids, inotropes, and vasopressors discretized into clinically meaningful categories. The bins were developed based on content expertise of nephrologists and critical care physicians (GK, GNN, AS). Seven discrete actions for IV fluids and vasopressors each, and six actions for inotropes were developed. This resulted in a total of 294 discrete actions for those embodiments of RENAL model.

[0096] GLUCOSE model reward. In the certain embodiments of the GLUCOSE model, the reward is designed to maximize or otherwise increase glycemic control while strongly discouraging behavior that would result in both hypoglycemia (glucose <70 mg / dL) and hyperglycemia (glucose >180 mg / dL). A maximum reward of +0.2 within the 140-180 mg / dL range was provided and penalties become increasingly negative, down to -1, outside that range. A reward with relatively low7magnitude was chosen to improve training stability, and a relatively negative reward was chosen to disincentive any out of range values. The reward is outlined in Eq (1):8 -1, x < 703x / 175 - 2.2, 70 < x < 140 r = 0.2, 140 < x < 180 (1)-0.03x + 5.6, 180 < x < 220-1, x > 220

[0097] To provide the safety of insulin dosing recommendations, considering insulin’s narrow therapeutic window, an exponentially increasing penalty was implemented in certain embodiments to discourage large overcorrections and promotes more cautious dosing adjustments. rt = rt - (0.001a2t) (2)Attorney Docket Number: MS-0045-01-WO

[0098] RENAL reward. In certain embodiments of the RENAL model, a clinically informed reward function was implemented to guide the agent toward preventing pAKI while maintaining hemodynamic stability. The reward structure is as follows:

[0099] Terminal Reward: If a patient develops pAKI in the first 5 days of ICU stay, a terminal penalty of -5 is assigned at the final time-step. Conversely, if the patient does not develop pAKI during this time-period, a terminal reward of +5 is given at the time time-step.

[0100] Intermediate Rewards: At any time-step the model receives a penalty of -2.5 if the patient's mean arterial pressure falls below 65mmHg at the next step. At any time-step if the patient does not meet criteria for AKI. the model receives a reward of + 1.EXAMPLE 2:Training the models

[0101] Training an embodiment of the GLUCOSE model. To train the policies, the development dataset was split into 85% training and 15% internal testing sets. Since RL is particularly subject to high stochastic training variation, sampling and stochastic biases were mitigated by training multiple models on subsampled 80% splits of the training set. Each training run and sampling split utilized a unique seed to ensure different training sets and distinct sampling order while maintaining reproducibility. Training of the model continued until substantial and significant improvements over the clinician policies were observed in the 15% internal testing set. One embodiment of the GLUCOSE model had the highest lower bound of the 95% CI for estimated performance returns in the internal testing set, which was then evaluated on the external validation set (Fig. 1).

[0102] Training was conducted in batches of 256, with actor and critic learning rates of le4and 3e4. respectively. The discount factor y was set at 0.67, corresponding to a 3 h effective horizon (calculated as 1 / (1 -y)). Discount factors are problem specific, and the choice of a lower discount factor is critical in the context of this problem. Higher discount factors, such as those exceeding 0.95, extend the effective horizon beyond the episode length, resulting in future rewards being weighted nearly as heavily as immediate rewards. Glucose levels can change rapidly, which could lead to suboptimal policy development as the model may either overly prioritize distant future rewards unaffected by the current state or become insensitive to immediate low reward states. A lower discount factor was chosen to align the temporal focusAttorney Docket Number: MS-0045-01-WO of the model to ensure it remains responsive to rapidly changing glucose levels. A dropout rate of p = 0. 1 was also adopted during training to improve policy generalization. All models were implemented in Python 3.8.2 using d3rlpy.Feature Extraction and pre-processing

[0103] In an embodiment, information was extracted about patient demographics (age, sex, race), comorbidities (history of diabetes, hypertension, end stage renal disease, chronic obstructive pulmonary disease, asthma, prior myocardial infarction, congestive heart failure, Elixhauser comorbidity score), laboratory values (complete blood count, comprehensive metabolic panel, coagulation studies, and blood gases), vital signs (systolic blood pressure, diastolic blood pressure, mean arterial pressure, heart rate, respiratory rate, temperature, and oxygen saturation), vasopressor and inotrope doses, mechanical ventilation status, and SOFA scores. The data was extracted as multidimensional discrete time series in 1-h time intervals, with features summed or averaged as clinically appropriate. Features with over 30% missingness were excluded. In line with standard approach to handling missingness in these data, forward fill imputation was used for all features with k-nearest neighbor (k-NN) imputation (k = 5) to impute any remaining missing data. The first 24 hours of data for each patient was utilized. All features were checked for outliers using a frequency histogram and descriptive statistics. Errors were corrected as appropriate, such as conversion of temperature to Fahrenheit to Celsius. All features across all datasets were normalized into range [0, 1] based on the training set to improve training stability.

[0104] The target outcome was appropriate glucose control, defined as an hourly glucose level between 70-180 mg / dL, in the first day after cardiac surgery. Consequently, the recording of timesteps began from the availability of the first glucose level measurement after admission to ICU.

[0105] Training an embodiment of the RENAL model. Features routinely available to clinicians at bedside were extracted for development of an embodiment of the RENAL model. These included demographics (age, sex), anthropometries (height, weight, BMI), vital signs, laboratory values, fluid balance, medications (vasopressors, inotropes, IV fluids and nephrotoxins), and SOFA Score. IV fluids included both crystalloids and colloids. The vasopressor and inotrope doses were converted into their norepinephrine and dobutamine equivalent doses respectively.Attorney Docket Number: MS-0045-01-WD

[0106] Data were extracted from ICU admission up to the earlier of either 72 hours or ICU discharge. Biologically implausible values were removed. The resulting time-series data were segmented into hourly bins, with features either summed or averaged as clinically appropriate. Features with more than 30% missingness were excluded. Following prior literature, forwardfill imputation and k-nearest neighbors (kNN) imputation were applied for any remaining missing values.

[0107] An embodiment of the RENAL model was developed using data from MIMIC-IV (development dataset) and validated it in SICdb (external validation dataset). The specific breakdown is data from 6623 post cardiac surgery patients in MIMIC-IV dataset, which was externally validated on 2230 patients from SICDB dataset. The average age in the development dataset for patients was 68.07 years with 71.87% being male. On the other hand, the average for the external validation set was 67.15 years, where 67.15% being males. AKI occurred in 2749 (41.51%) patients in the development cohort and 443 (19.87%) in the external validation cohort. Persistent AKI occurred among 842 (30.6%) patients with AKI in development cohort and 190 (42.9%) patients with AKI in external validation cohort. Among those with AKI in the development cohort, the initial stage was stage 1 in 88.8% of patients, stage 2 in 7.0%, and stage 3 in 4.2%. In the external validation cohort, the initial stage was stage 1 in 75.7% of patients, stage 2 in 5.4%, and stage 3 in 18.9%. The patient characteristic information for use in developing this embodiment of the RENAL model is provided in Table 6.

[0108] The development dataset was split into three subsets: 70% training, 15% validation, and 15% internal testing. This split ensured that the model could be trained effectively while also being evaluated on unseen data to assess generalizability. Dunng training, a discount factor rate of 0.99 was chosen to emphasize long-term rewards and reflect the clinical importance of avoiding adverse outcomes over time. The conservativeness coefficient (alpha) was set to 0.25 which controls the degree to which the learned policy diverges from the observed clinician behavior. This was chosen to encourage safe exploration without straying too far from established practices. This embodiment of the model was trained using a batch size of 512, an initial learning of 0.001 with a learning rate scheduler that progressively reduced the rate to le-8.Attorney Docket Number: MS-0045-01-WOEXAMPLE 3:Evaluating the models

[0109] Evaluating an embodiment of the GLUCOSE model. To mitigate the sampling and stochastic biases inherent in offline RL multiple models were trained, in line with prior literature, until no significant improvements in the RL policies were observed. Two hundred independent models were trained, and the model with the maximal lower bound of the 95% CI of mean estimated performance returns within the internal testing was set as the GLUCOSE model for additional testing. Bootstrapping was applied across all episodes in the datasets to generate 95% confidence intervals by sampling with replacement 10.000 times. The performance was estimated using fitted Q estimation (FQE) on both internal testing set and external validation dataset.

[0110] The estimated performance returns of this embodiment of the GLUCOSE model were compared at the lower bound of its 95% CI with the upper bound of clinicians' 95% CI (Figure 2a) using FQE for off policy evaluation (OPE), illustrating the differences in average estimated performance after evaluating 200 policies. The lines reflect the 95% confidence intervals of the mean performance for the observed clinicians’ behavior in internal testing and external validation, respectively, while their non-dotted counterparts reflect the estimated performance of this GLUCOSE model through OPE (Fig. 2a). The best GLUCOSE model resulted in a mean estimated performance return of 0.0 [-0.07,0.06] in the internal testing set and -0.63 [-0.74, -0.52] in the external validation dataset, showing significant improvements over the clinician returns of -1.29 [-1.37, -1.20] in the internal testing set and -1.02 [-1.16, - 0.89] in the external validation dataset.

[0111] Policy performance was further examined by analyzing the time in range (TIR) (70- 180 mg / dL) in relation to the difference of this GLUCOSE model's dosing recommendations and clinician-administered doses (Fig. 2b). A calculation was made of the cumulative differences as the model’s predicted insulin doses minus the observed insulin doses over the first 24 hours of ICU stay:Attorney Docket Number: MS-0045-01-WO

[0112] Cumulative differences are positive when the RL model recommends more insulin than what was administered, while negative differences indicate that the model predicted a lower insulin dose compared to what was observed.

[0113] In the internal testing set, 27.3% of the time patients received insulin doses from clinicians identical to those recommended by the GLUCOSE model, while in the external validation dataset, this occurred 20.3% of the time. As shown in Fig. 2b, patients who received insulin doses like those suggested by the GLUCOSE model had the highest average TIR in both the internal testing set and external validation dataset. The TIR decreased as the difference of model recommended doses minus clinician administered doses increased, indicating that the model identifies areas for improvement in insulin management. For example, at more negative cumulative differences, where the average glucose is also lower, the GLUCOSE model suggests less insulin to mitigate the risk of hypoglycemia (Figure 2b, Figure 2c). Conversely, at more positive cumulative differences, where average blood glucose is higher and average TIR is worse, the model suggests higher insulin doses to avoid hyperglycemia. To explore subgroup-specific performance, a TIR analysis was conducted and was stratified by sex, race, and diabetic status (Figure 5). Across all subgroups, the GLUCOSE model achieved the highest average TIR when its recommended insulin dose matched the clinician-administered dose. This consistent pattern across all groups suggests that this GLUCOSE model performs well across these subgroups of patient populations.

[0114] The action distribution of Figure 2c further illustrates these dosing patterns. Across glucose ranges below 180 mg / dL, this GLUCOSE model consistently recommends lower average insulin doses than clinicians, reflecting a guideline-aligned strategy that prioritizes maintaining glucose between 140 and 180 mg / dL while reducing the risk for hypoglycemia. The recommended insulin dose starts to increase after this threshold and surpasses the clinician doses when glucose levels were above 200 mg / dL, demonstrating a proactive approach to correcting significant hyperglycemia aligning with the recommendations to avoid hyperglycemia while minimizing the risk of hypoglycemia.

[0115] To characterize clinician-model disagreement, an analysis was performed of the clinical features associated with the top and bottom deciles of absolute differences between clinician-administered insulin and the GLUCOSE model-predicted insulin, corresponding to the highest and lowest disagreement, respectively. Across both internal and external cohorts, the largest disagreements occurred near glucose values of approximately 140 mg / dL. ThisAttorney Docket Number: MS-0045-01-WO range lies near the lower boundary of the 140-180 mg / dL target recommended for glucose management among critically ill patients. In the internal testing set. clinicians administered an average of 5.9 units of insulin in high-disagreement cases compared to the model’s 1.8 units, with a mean glucose of 142 mg / dL. Similarly, in the external validation set, clinicians gave 6.7 units versus the model’s 1.6 units at a mean glucose of 139.5 mg / dL.

[0116] To gain insight into a model’s representations and ensure its clinical interpretability, feature importances were derived for the GLUCOSE model using SHapley Additive exPlanations (SHAP) (Figure 6). The interpretability of machine learning models is a factor in clinical care, where the rationale behind a model’s predictions should be clear so as to provide patient safety and informed decision-making. SHAP, a method grounded in cooperative game theory, assesses the contribution of each feature to a model’s prediction by analyzing all possible combinations of feature values. In this study, permutation SHAP was employed to estimate these contributions, as it provides a model-agnostic framework for elucidating model outputs.

[0117] This analysis revealed that the most heavily weighted features align well with clinical intuition. Notably, recent and historical glucose measurements emerged as key predictors, underscoring the value of capturing real-time trends. Additionally, indicators of patient acuity, which may influence stress-induced hyperglycemia, such as the use of and duration of mechanical ventilation, sequential organ failure assessment (SOFA) score, elixhauser comorbidity index, and the ty pe of surgery, were weighed heavily. These findings demonstrate that this GLUCOSE model uses clinically relevant information in its decision making.

[0118] The clinical validity of this GLUCOSE model was further assessed in three separate phases of human evaluations. In the first phase, the hourly insulin dosing recommended by the GLUCOSE model was compared to that administered to patients in both internal testing and external validation datasets. Two senior endocrinologists, each with over ten years of clinical experience, provided their own hourly insulin dosing recommendations for the first day after cardiac surgery' for 10 randomly selected patients in each cohort. Using the average hourly insulin doses recommended by endocrinologists as reference, the insulin doses recommended by this GLUCOSE model were compared to those administered to these patients.Attorney Docket Number: MS-0045-01-WO

[0119] To allow the endocrinologists to provide the most accurate dosing schemes to use as a reference, the endocrinologists were provided with the entire time series of patient data, including the insulin doses actually administered by the treating clinicians, and resultant glucose levels recommended by this GLUCOSE model, which unlike the endocrinologists had access only to the current state, to those actually administered by clinicians using the average endocrinologist doses as the reference. Across both datasets, this GLUCOSE model achieved significantly lower mean absolute error (MAE) in hourly insulin dosing, indicating that its dosing scheme more closely aligned with the recommendations of the endocrinologists. In the internal testing set, this GLUCOSE model had an MAE of 0.9 units compared to the treating clinician's 1.97 units MAE (p < 0.001). In the external validation dataset, this GLUCOSE model had an MAE of 1.90 units compared to the treating clinician’s MAE of 2.24 units (p = 0.003).

[0120] In the second phase, two senior cardiac intensivists (> 5 years’ experience), two junior cardiac intensivists (< 5 years’ experience), and two cardiac intensive care unit nurse practitioners provided their recommendations for hourly insulin doses in the first day after cardiac surgery for the same patients. These clinicians were also provided with the entire time series of data, actual insulin administration record, and glucose levels to allow them to generate their most retrospectively optimal possible human policies. This information was then compared with the GLUCOSE model’s recommended doses, which again had access to a single state of information at the current timestep, to those recommended by these 6 clinicians, with the endocrinologist recommendations as the reference (Figure 3a).

[0121] In internal testing, this GLUCOSE model achieved an MAE of 0.90 units compared to that of senior intensivists’ 0.82-unit MAE (p = 0.57), junior intensivists’ 1.15-unit MAE (p = 0.25), and nurse practitioners’ 1.23-unit MAE (p = 0.21). In external validation, this GLUCOSE model achieved an MAE of 1.90 units compared to that of senior intensivists’ 1.58 units (p = 0.32), junior intensivists 2. 15 units (p = 0.53), and nurse practitioners 2.28 units (p = 0.38). Although the differences in MAE did not reach statistical significance, this GLUCOSE model demonstrates a trend toward lower MAEs than that of junior intensivists and nurse practitioners when compared against endocrinologists as the reference.

[0122] In the third phase, a blinded evaluation of this GLUCOSE model and all 8 clinician dosing recommendations using an expert panel of 2 separate senior intensivists to assess practical safety7, effectiveness, and acceptability7of the model’s recommended insulin doses.Attorney Docket Number: MS-0045-01-WO(Senior intensivists were chosen for this phase because, in practice, these frontline clinicians are frequently responsible for making rapid decisions regarding glucose control in critically ill patients.) The two additional senior intensivists used a 5-point Likert scale to assess the safety (to reduce hypoglycemia), effectiveness (if the regimen would bring glucose into an acceptable range), and acceptability' (if the regimen would be acceptable in a clinical scenario) of each recommended insulin regimen for the same group of 20 patients. They used the following 5- point Likert scale questions:

[0123] QI (Safety ) — How much risk for hypoglycemia does this regimen put the patient at? 1. Very high risk 2. High risk 3. Medium risk 4. Low risk 5. Minimal risk

[0124] Q2 (Effectiveness) — How effective is this regimen in bringing the glucose level within an acceptable range? 1. Not effective at all 2. Slightly effective 3. Moderately effective 4. Very effective 5. Extremely effective

[0125] Q3 (Acceptability) — How acceptable would this regimen be to you in clinical settings? 1. Strongly unacceptable 2. Unacceptable 3. Neutral 4. Acceptable 5. Strongly acceptable

[0126] In both internal testing and external validation datasets, the GLUCOSE model’s rated safety effectiveness, and acceptability demonstrated either comparable performance or statistically significant improvements over all human policies (Figure 3b). Notably, the GLUCOSE model performed at or above the level of senior cardiac surgery intensivists across all domains in both the internal testing and external validation datasets. This demonstrates the GLUCOSE model’s reliability7and consistent high-level performance across diverse clinical scenarios.

[0127] To illustrate the GLUCOSE model’s real time decision making, a comparison was made between representative case examples comparing the model’s insulin dosing recommendations with actual clinician-administered doses (Figure 4). Overall, the GLUCOSE model consistently demonstrated dynamic and personalized insulin dosing strategies, adapting to changes in glucose trajectories. Across randomly7selected internal testing and external validation patients, the GLUCOSE model provided timely insulin adjustments, often moderating dosing to avoid overshooting glucose targets. These cases highlight how the GLUCOSE model responds to evolving patient conditions and targets an optimal glucose range more in line with STS guidelines.Attorney Docket Number: MS-0045-01-WO

[0128] The GLUCOSE model’s recommendations were also evaluated using the part of external validation dataset that was excluded from the primary analysis due to presence of ambiguous insulin administration data, which prevented direct comparisons or the calculation of OPE. These patients had documented insulin, vasopressor, or inotrope administration but with insufficient information to determine the exact timing or dose - an issue commonly reported in the multicenter elCU-CRD database. Based on the affected medication, this evaluation was performed separately in patients with only ambiguous insulin data (3,001 patients) and in those with ambiguous data for both insulin and vasopressors / inotropes (1,804 patients). As there were only 33 patients with non-ambiguous insulin but ambiguous vasopressor / inotrope data, these were excluded from this analysis.

[0129] The external validation cohort, the ambiguous insulin subset, and the ambiguous medication subset were similar in terms of age, gender, and weight. However, these additional subsets included a higher proportion of white patients (66.6% vs 81.1% vs 89.7%, p < 0.001) and fewer patients with type 2 diabetes (9.1% vs. 3.1% vs. 1.2%, p < 0.001). While there were statistically significant differences in the average glucose levels (134.5 mg / dL vs. 132.2 mg / dL vs. 130.6 mg / dL, p < 0.001) these small differences are not clinically meaningful. Full demographic analysis can be found in Tables 2 & 3.

[0130] Due to the lack of accurately recorded insulin in these subgroups. OPE or direct comparisons were could not be performed because they depend on accurate insulin administration records. However, the overall distribution of model recommended actions was comparable across datasets (Figure 7). Although there were statistically significant differences, all differences in average insulin across all glucose ranges were less than half a unit and therefore not clinically significant (Figure 7).

[0131] Comparisons of categorical features were performed using Chi-squared test and continuous features using t test and Kruskal-Wallis test, as appropriate. All significance levels were set at a = 0.05. To compare insulin doses recommended for administration by clinicians and the GLUCOSE model before hypo- and hyperglycemic episodes, Mann-Whitney U tests were used given the skewed distributions. To evaluate the accuracy of the insulin dosing schemes, a calculation was made of the mean absolute error (MAE) between the insulin doses recommended by various dosing schemes to those provided by endocrinologists. MAEs were determined by subtracting the endocrinologists’ recommendations from the doses suggested by clinicians and the GLUCOSE model’s method for each hourly dose administered orAttorney Docket Number: MS-0045-01-WO recommended for the 20 retrospectively reviewed patients. A t test was then performed to identify any significant differences in MAE between the observed clinicians’ dosing and the endocrinologist’s suggested dosing, and MAE between the GLUCOSE model’s suggested dosing and the endocrinologist’s suggested dosing. To assess differences in the average insulin doses across subsets of excluded patients, ANOVA tests were used for group-wise assessment and two-sided t-tests for pairwise assessment. All statistical tests were conducted using Python 3.8.2 using SciPy.

[0132] A Pilot evaluation of an embodiment of the GLUCOSE model was performed at Mount Sinai over a 6-week period. The inclusion criteria are the same as was used previously. Model integration and data flow: The study automatically extracted EMR data hourly: demographics, vitals, labs (including glucose), SOFA score, ventilator and vasopressor status, and prior insulin doses. The silent inference was set at GLUCOSE model runs each hour on the current state, generation of a recommended insulin dose (rounded to nearest unit, capped at 10 U / hr). Model recommendations are generated only until the earliest of: 24 hours post-ICU admission, patient transfer out of the ICU, or administration of any non-regular insulin formulation. The system stores timestamped pairs of (model -recommended dose, clinician- administered dose, glucose level). The model’s outputs are not visible to treatment teams; usual care proceeds per institutional protocol. Endpoints and metrics: The feasibility metrics were set at: % of ICU hours with complete data and successful model execution and time latency from data availability to model output (< 5 min target). The performance metrics were: 5 points Linkert scale stores for safety, efficacy, and applicability' by senior intensivists. Time-in-range difference was stratified by clinician dosing discrepancy with that of GLUCOSE recommended dose. Statistical analysis plan: For descriptive analyses: Summarize feasibility metrics as proportions (e.g., % of ICU hours with complete data) and medians (IQR) for latency; and report intensivist ratings (Likert scores) for safety, efficacy, and applicability as means ± SD and medians (IQR). For comparative analyses, evaluate time in range (70-180mg / dL) for glucose values based on deviations of clinician dosing from GLUCOSE recommendations. For the sample size, there was no formal power calculation, and the study anticipated -100 patients (-2,400 patient-hours) to yield stable estimates of feasibility and rating distributions. The results of the pilot study are incorporated into Figure 2.

[0133] Evaluating an embodiment of the RENAL model. Upon completion of training, an embodiment of the RENAL model’s performance was evaluated using both offline policyAttorney Docket Number: MS-0045-01-WO evaluation (OPE) and clinical evaluation. As a method of offline evaluation, fitted Q evaluation (FQE), an established OPE method, was employed to compare total rewards of RENAL versus that of clinicians in preventing persistent AKI. These estimates were then compared to those obtained for the historical clinician policy to assess their relative effectiveness in preventing persistent AKI following cardiac surgery. The RENAL model achieved a mean overall reward of 54.38 ± 1.33 vs. 18.69 ± 1.10 for clinicians in the internal test set, and 76.48 ± 0.78 vs. 21.33 ± 1.05 in the external validation set (Figure 10). To assess generalizability, an evaluation was conducted of the performance of the RENAL model in SICdb, an independent, international, external validation cohort using FQE. A comparison was also made for continuous features using t-test or Kruskal-Wallis test as appropriate, and discrete features using Chi-squared test, at the alpha was set at 0.05.

[0134] To illustrate this model’s individualized dosing behavior, hour-by-hour heatmaps were generated for thirty randomly selected patients from our external cohort, stratified by the presence or absence of AKI and pAKI. As shown in Figure Ila, this RENAL model more frequently recommended lower volumes of IV fluids compared to clinicians. Overall, the most commonly recommended fluid volume category by this RENAL model was 0-50 mL, accounting for 83.29% of cases, whereas clinicians recommended this volume in only 30.74% of cases (p < 0.001).Similarly, when vasopressors were recommended, this RENAL model most often selected a moderate dose of 0.02-0.05 mcg / kg / min, comprising 27.50% of all recommendations, compared to 8.28% in clinician decisions (p < 0.001). For inotropes, this RENAL model most frequently recommended a dose of 5-7.5 mcg / kg / min and 2.5-5 mcg / kg / min (both at 1.73% of all cases). In contrast, clinicians most commonly recommended 0-2.5 mcg / kg / min (4.95%), followed by 2.5-5 mcg / kg / min (1.69%: p < 0.001).

[0135] Similar trends were observed among patients who ultimately developed pAKI (Figure lib). In this subgroup, this RENAL model again favored lower fluid volumes, recommending 0-50 mL in 73.96% of cases compared to 27.9% by clinicians (p < 0.001). For vasopressors, the most frequently recommended dose by this RENAL model was 0.02- 0.05 mcg / kg / min in 32.47% of cases, versus 10.06% by clinicians (p < 0.001). When inotropes were used, this RENAL model recommended a higher dose of 5-7.5 mcg / kg / min in 5.41% of cases, compared to 1.26% by clinicians (p < 0.001).

[0136] To assess the clinical impact of this RENAL model’s treatment recommendations, an evaluation was performed of the association between deviations from model-suggestedAttorney Docket Number: MS-0045-01-WO doses and the occurrence of pAKI over time. These recommendations were examined across various MAP thresholds. Specifically, an examination was made of the magnitude of deviations between clinician actions and this RENAL model’s recommendations and an evaluation was made of their association with pAKI risk by plotting U-shaped dose-risk curves. This allowed the examination of whether deviations from the recommended doses of IV fluids, vasopressors, and inotropes were associated with an increased risk of persistent AKI.

[0137] As shown in Figure 12a, this model’s recommendations across different MAP values demonstrated a consistent pattern. Specifically, among patients with MAP <65 mm Hg, this RENAL model most frequently recommended an IV fluid volume of 0-50 mL, accounting for 77.98% of cases, compared to clinicians, whose most recommended volume was 50- 100 mL (28.58% of cases; p < 0.001 ). In this subgroup, fluid volumes greater than 500 mL were recommended far less often by the RENAL model (3.63%) than by clinicians (10.3%; p < 0.001). For vasopressors, the RENAL model most frequently recommended a moderate dose (0.02-0.05 mcg / kg / min) in 40.14% of cases, compared to 13.58% by clinicians (p < 0.001). Additionally, the model showed a greater preference for higher inotrope doses, recommending 5-7.5 mcg / kg / min in 3.54% of cases versus only 0.74% by clinicians (p < 0.001).

[0138] A similar pattern was observed among patients with MAP <65 mm Hg who subsequently developed pAKI (Figure 12b). In this group, this embodiment of the RENAL model again favored lower fluid volumes, recommending 0-50 mL in 67.68% of cases, while clinicians most frequently chose 50-100 mL (30.94%; p < 0.001). This model also continued to favor moderate vasopressor dosing (0.02-0.05 mcg / kg / min in 36.09% of cases), whereas clinicians most often used higher doses (0.1 -0.2 mcg / kg / min in 19.8% of cases; p < 0.001. For inotropes, the model recommended a 5-7.5 mcg / kg / min dose in 7.86% of cases versus 1.25% by clinicians (p < 0.001). The analysis of model vs, clinician concordance of actions relay similar results.

[0139] A calculation was made of the difference between the administered dose and the model-recommended dose at each decision point and then averaged these differences across all time points to derive a patient-level mean dose deviation for each treatment modality. The relationship betw een the magnitude of this average deviation and the incidence of pAKI was then analyzed. U-shaped curves were constructed by plotting the risk of pAKI against the average dose deviation. The 95-percentile confidence intervals are constructed based on criticalAttorney Docket Number: MS-0045-01-WO t-value and standard error of the mean. This analysis allowed an evaluation of whether close adherence to this model’s recommended dosing was associated with reduced risk of pAKI.

[0140] Figure 13 presents the effect of concordance, defined as agreement between the RL model’s recommended action and the action taken by the clinician, on the risk of persistent AKI across subgroups in the internal test cohort (MIMIC-IV) and the external validation cohort (SICdb). In the MIMIC-IV cohort, concordance was associated with a lower risk of persistent AKI, with an overall odds ratio (OR) of 0.75 (95% CI, 0.71-0.79) using the standard variance estimate and 0.75 (95% CI, 0.62-0.90) with the sandwich estimator. Subgroup-specific ORs were consistent, ranging from 0.72 to 0.76 across age (<65 vs >65 years), sex (male vs female), and race (White vs Other). In the SICdb validation cohort, the direction of effect was similar, with an overall OR of 0.83 (95% CI, 0.79-0.88) by standard estimation and 0.83 (95% CI, 0.72-0.97) with the sandwich estimator. Subgroup analyses showed ORs between 0.83 and 0.86, demonstrating consistent results with those observed in MIMIC-IV across age and sex. Concordance was defined as agreement between the RL model’s recommended action and the action taken by the clinician; for our internal test dataset, the concordance rate was 41 .85% and for external validation it was 22.49%. Odds ratios (95% Cis) were estimated from weighted pooled logistic regression models with inverse probability7of treatment weighting, using both standard and sandwich variance estimation. Subgroup analyses by age, sex. and race in SICdb showed results consistent with those observed in MIMIC-IV

[0141] To reduce short-term variability and capture cumulative clinical impact, an evaluation was performed on dose deviations aggregated over a4-hour window. This approach revealed consistent U-shaped dose-risk curves, with the lowest pAKI risk observed when clinician actions closely aligned with the model’s recommendations, and increasing risk associated with larger deviations in either direction.

[0142] To evaluate the recommendations of this embodiment of the RENAL model at individual patient level there was a random sampling of thirty patients - ten who developed pAKI, ten who only developed AKI, and ten who never developed AKI. A heatmap was created from trajectories of MAP and clinical actions performed by clinicians (Fig 11-a) and the model’s hour-by-hour recommendations (Figure 11-b). Visual inspection of the heatmaps of the model’s actions revealed several notable patterns. The model generally delivers more graded, anticipatory support, ramping up low-dose vasopressors or inotropes just as MAP begins to drift downward and tapering fluids sooner once perfusion stabilizes. By comparison,Attorney Docket Number: MS-0045-01-WO clinician actions frequently show larger IV fluid surges or delayed inotrope support, especially among patients who ultimately developed pAKI

[0143] When visually comparing those panels to this RENAL model’s timelines, a pattern emerges where the model generally delivers more graded, anticipatory7support, ramping up low-dose vasopressors or inotropes just as MAP begins to drift downward and tapering fluids sooner once perfusion stabilizes. By comparison, bedside practice frequently shows larger IV fluid surges or delayed vasopressor support, especially among patients who ultimately developed pAKI In AKI-only and no-AKI cases the same patterns appear where clinicians give broader, less finely tuned interventions, whereas the model's suggestions follow each patient’s hemodynamic curve more closely. Together, these side-by-side heat-maps underscore the model’s capacity for smoother, physiology-driven dosing and highlight exactly where and when its guidance would have diverged from real-world care.

[0144] Explainability is useful in RL for clinical decision-making to provide transparency and clinician trust, particularly when recommending high-risk interventions. SHAP was applied to assess global interpretability of this embodiment of the RENAL model. SHAP was selected for its strong theoretical foundation and ability to quantify the average contribution of each input feature to the model’s overall predictions. Shapely analysis stems from cooperative games in game theory, and it is used to fairly distribute the payoffs in cooperative game to each contributing player. In RL, it is used to derive the contributions of each feature to the overall output in the model. SHAP values w ere computed across the entire test set. This allowed the identification of the most influential clinical variables driving the model’s recommendations and assess alignment with clinical expectations. A heatmap was created by using matplotlib library and indexing hourly prescriptions from clinicians and the equivalent median values from the model recommendations.

[0145] As shown in Figure 15, the top contributors to the model’s decision-making were largely hemodynamic and fluid-balance variables with cumulative fluid balance (mean |SHAP| ~ 6.05) with by far the strongest signal, followed by the net fluid balance and vasopressor dose in the previous timestep. SOFA score, serum creatinine, inotrope dose in the previous timestep, baseline creatinine and MAP were also ranked highly. Other vital signs and demographic features occupied the middle of the importance spectrum. This pattern confirms that the model is basing its IV fluid and vasopressor suggestions primarily on each patient’s ongoing balance and perfusion status, with secondary7tuning from overall illness severity7and baselineAttorney Docket Number: MS-0045-01-WO physiology. The local SHAP values show how the influence of features for recommendations of clinical actions evolved over time.

[0146] Table 1. The Population cohort ICD Codes

[0147] Table 2. Patient Characteristics. Percentages may exceed 100% for types of surgery as patients may undergo both coronary’ artery bypass grafting and valvular surgery, including multiple valves, simultaneously.Attorney Docket Number: MS-0045-01-WOAttorney Docket Number: MS-0045-01-WO

[0148] TABLE 3 Population characteristics of excluded external validation patients. Characteristics of all patients excluded from external validation due to ambiguous medication information. Information across demographic information, vitals, laboratory data, scores, and medications shown.Attorney Docket Number: MS-0045-01-WOAttorney Docket Number: MS-0045-01-WO

[0149] TABLE 4 ICD-9-PCS and ICD-10-PCS codes. Table of codes used to identifyCABG, valve repair, and valve replacement procedures.Attorney Docket Number: MS-0045-01-WO

[0150] TABLE 5. Admission diagnosis reasons. Admission diagnosis reasons used in elCU-CRD in the absence of available ICD-9-PCS and ICD-10-PCS to identify post-cardiac surgery patient.Attorney Docket Number: MS-0045-01-WOTable 6. Patient characteristic information for use in developing an embodiment of the RENAL model.

Claims

Attorney Docket Number: MS-0045-01-WOCLAIMSWhat is claimed is:1 . A method for generating an output for the administration of an insulin therapy for a post operative cardiac patient, the method comprising: receiving patient information including blood glucose levels of said patient; providing said patient information as input to a conservative Q-leaming (CQL) model that has been trained on data from cardiac patients using a reward function that includes a safety component that penalizes blood glucose levels below about 70 mg / dL and above about 180 mg / dL and a progress component that rewards time where the blood glucose levels are between about 70 mg / dL and about 180 mg / dL for determining whether to administer insulin to said patient; wherein said CQL model is integrated with a distributional reinforcement learning (RL) model; and generating, with said integrated model based on said patient information, an output for the administration of an insulin dose to said patient.

2. The method of claim 1 , wherein said distributional RL model incorporates an implicit quantile network (IQN).

3. The method of claim 2, wherein said insulin dose is for administration during a first 24 hours after cardiac surgery.

4. The method of claim 2, wherein said patient information is selected from a group comprising: demographics, comorbidities, laboratory values, vital signs, medications, mechanical ventilation status, body mass index (BMI), race, sequential organ failure assessment (SOFA) scores, insulin dosing, vasopressor and inotrope doses, and patient history.

5. The method of claim 4, further comprising: providing subsequent patient information to the CQL model; and generating, with said CQL model based on said subsequent patient information, an output indicative of subsequent insulin doses at least on an hourly- basis.Attorney Docket Number: MS-0045-01-WO6. The method of claim 4. wherein said CQL model includes at least one multilayer perceptron (MLP) matrix having three 512-dimension hidden layers.

7. The method of claim 4, wherein said CQL model utilizes a Markov decision process (MDP) and wherein states of said MDP are selected from: demographics, comorbidities, laboratory' values, vital signs, medications, mechanical ventilation, and sequential organ failure assessment (SOFA) scores.

8. The method of claim 7, wherein there is a maximum reward of +0.2 within a 140-180mg / dL range and penalties become increasingly negative, own to -1 , outside of that range.

9. The method of claim 8, wherein said patient information is extracted as features providing multidimensional discrete time series in 1 hour time intervals, and wherein said features are summed or averaged as clinically appropriate.

10. A method of using the output provided by the method of claim 7 to treat a patient.I L A system comprising: one or more processors; and one or more non-transitory computer-readable media, coupled to the one or more processors and storing instructions which, when executed by the one or more processors, cause the one or more processors to perform or control performance of operations comprising: receiving patient information including blood glucose levels of said patient; providing said patient information as input to a conservative Q-leaming (CQL) model that has been trained on data from cardiac patients using a reward function that includes a safety component that penalizes blood glucose levels below' about 70 mg / dL and above about 180 mg / dL and a progress component that rewards time where the blood glucose levels are between about 70 mg / dL and about 180 mg / dL for determining whether to administer insulin to said patient; wherein said CQL model is integrated with a distributional reinforcement learning (RL) model; and generating, with said integrated model based on said patient information, an output for the administration of an insulin dose.

12. The system of claim 11, wherein said distributional RL model incorporates an implicit quantile network (IQN).

13. The system of claim 12, wherein said insulin dose is for administration during a first 24 hours after cardiac surgeryAttorney Docket Number: MS-0045-01-WO14. The system of claim 12, wherein said patient information is selected from a group comprising: demographics, comorbidities, laboratory values, vital signs, medications, mechanical ventilation, body mass index (BMI), race, sequential organ failure assessment (SOFA) scores, insulin dosing, vasopressor and inotrope doses, and patient history.

15. The system of claim 14, wherein the operations further comprise: providing subsequent patient information to the CQL model; and generating, with said CQL model based on said patient information, output indicative of subsequent insulin doses at least on an hourly basis.

16. The system of claim 14, wherein said CQL model includes at least one multilayer perceptron (MLP) matrix having three 512-dimension hidden layers.

17. The system of claim 14, wherein said CQL model utilizes a Markov decision process (MDP), and wherein the states of said MDP are selected from: demographics, comorbidities, laboratory values, vital signs, medications, mechanical ventilation, and sequential organ failure assessment (SOFA) scores.

18. The system of claim 14, wherein there is a maximum reward of +0.2 within a 140- 180 mg / dL range and penalties become increasingly negative, down to -1, outside of that range.

19. A method of using the system of claim 14 to treat a patient in need thereof.

20. A system for determining an insulin therapy for a post operative cardiac patient comprising: one or more processors; and one or more non-transitory computer-readable media, coupled to the one or more processors and storing instructions which, when executed by the one or more processors, cause the one or more processors to perform or control performance of operations comprising: receiving patient information regarding a post operative cardiac patient, including blood glucose levels of said patient; providing said patient information as input to a conservative Q-leaming (CQL) model that has been trained on data from cardiac patients using a reward function that includes a safety component that penalizes blood glucose levels below about 70 mg / dL and above about 180 mg / dL and a progress component that rewards time where the blood glucose levels are between about 70 mg / dL and about 180 mg / dL for determining whether to administer insulin to said patient;Attorney Docket Number: MS-0045-01-WO wherein said CQL model is integrated with a distributional reinforcement learning (RL) model; wherein said distributional RL model incorporates an implicit quantile network (IQN); wherein said CQL model includes a multilayer perceptron (MLP) matrix; wherein said CQL model utilizes a Markov decision process (MDP); wherein the states of said MDP are selected from: demographics, comorbidities, laboratory values, vital signs, medications, mechanical ventilation, and sequential organ failure assessment (SOFA) scores; generating from said CQL model based on subsequent patient information an output; and wherein said output provides information for diagnosis or treatment of a patient.

21. A method for treating a patient in need thereof using the system of claim 20.

22. A method for hemodynamic management in a post operative cardiac patient, the method comprising: receiving patient information for adult patients that were admitted after cardiac surgery7; providing said patient information as input to a conservative Q-leaming (CQL) model that has been trained on training features from cardiac patient training data, wherein said cardiac patient training data is from patients admitted to an intensive care unit (ICU) after cardiac surgery and wherein said model has been integrated with an implicit quantile network (IQN); generating, with said CQL model based on said patient information, an output regarding whether said patient is at risk of acute kidney injury (AKI), wherein said output includes a list of treatment actions, optionally including administering medications.

23. The method of claim 22, wherein said training features include one or more of: demographics, anthropometries, vital signs, laboratory values, fluid balance, ions, and sequential organ failure assessment (SOFA) score.

24. The method of claim 22, wherein said medications include one or more of: IV fluids, vasopressors, inotropes, and nephrotoxins.

25. A method of treating a patient in need thereof using the method of claim 24.

26. A system comprising:Attorney Docket Number: MS-0045-01-WO one or more processors; and one or more non-transitory processor-readable media, coupled to the one or more processors and storing instructions which, when executed by the one or more processors, cause the one or more processors to perform or control performance of operations comprising: receiving patient information for adult patients that were admitted after cardiac surgery; providing said patient information as input to a conservative Q-leaming (CQL) model that has been trained on training features from cardiac patient training data, wherein said cardiac patient training data is from patients admitted to an intensive care unit (ICU) after cardiac surgery and wherein said model has been integrated with an implicit quantile network (IQN); and generating, with said CQL model based on said patient information, an output regarding w hether said patient is at risk of acute kidney injury (AKI), wherein said output includes a list of treatment actions, optionally including administering medications.

27. The system of claim 26, wherein said training features include: demographics, anthropometries, vital signs, laboratory values, fluid balance, medications, and sequential organ failure assessment (SOFA) score.

28. The system of claim 26, wherein said medications included: IV fluids, vasopressors, inotropes, and nephrotoxins.

29. A method of using the system of claim 28 to treat a patient in need thereof.