System and method for predicting clinical conditions

A transformer-based machine learning system for predicting clinical events in hospitalized patients addresses the limitations of reactive care by integrating diverse EHR data for proactive risk management, improving hypoglycemia prediction and reducing healthcare burdens.

WO2026156182A1PCT designated stage Publication Date: 2026-07-23CEDARS SINAI MEDICAL CENT
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
CEDARS SINAI MEDICAL CENT
Filing Date
2026-01-15
Publication Date
2026-07-23

Smart Images

  • Figure US2026011454_23072026_PF_FP_ABST
    Figure US2026011454_23072026_PF_FP_ABST
Patent Text Reader

Abstract

A system and method for predicting a clinical condition or clinical event in an individual are provided. In some aspects, the method includes receiving electric health record (EHR) data associated with the individual, and providing the EHR data as input to a machine learning (ML) model configured to predict a clinical condition or clinical event. The method also includes receiving an output from the ML model indicative of a risk of the individual developing the clinical condition or clinical event a within a prediction window, and generating a report indicating the risk of the individual.
Need to check novelty before this filing date? Find Prior Art

Description

SYSTEM AND METHOD FOR PREDICTING CLINICAL CONDITIONSCROSS-REFERENCE TO RELATED APPLICATIONS

[0001] This application claims the benefit of, and priority to, U.S. Provisional Patent Application No. 63 / 745,479 filed on January 15, 2025, and U.S. Provisional Patent Application No. 63 / 801,606 filed on May 7, 2025, each of which is hereby incorporated by reference herein in its entirety.TECHNICAL FIELD

[0002] The present disclosure relates generally to systems and methods for monitoring and controlling patient conditions and complications, and more specifically to a system and method for predicting a clinical condition or clinical event in an individual.BACKGROUND

[0003] Hospitalized patients face significant risks from a variety of unpredictable clinical conditions and events, including abnormal blood glucose levels, acute kidney injury, sepsis, and other serious complications. These events impose substantial burdens on both patient health outcomes and healthcare systems. For example, hypoglycemic events affect approximately 20% of hospitalized patients on insulin therapy and are associated with increased morbidity, prolonged hospital stays, higher readmission rates, and increased mortality. Hypoglycemia is particularly critical and often unpredictable for hospitalized patients with diabetes, with severe episodes demonstrating hazard ratios for mortality as high as 1.847 compared to patients without hypoglycemia. Other clinical conditions and events may also present critical issues during patient care.

[0004] The complexity of managing these clinical conditions is compounded by numerous contributing factors. Clinical, social, and economic variables interact in ways that create substantial barriers to effective prediction and prevention. For instance, inpatient hypoglycemia can result from prior hypoglycemic events, intensive insulin therapy, end-stage renal disease, cognitive impairment, female sex, advanced age (>75 years), social determinants such as low income and homelessness, medication interactions, variations in stress and nutritional intake during hospitalization, and individual metabolic variability. Additionally, standard inpatient insulin dosing protocols fail to account for the full spectrum of individual variability across patient populations, and many risk factors remain incompletely understood or difficult to14918-2941-55562065472-001007WOPTquantify in clinical practice. Other issues may also present during the treatment and management of other clinical conditions and events.

[0005] Current systems and approaches to managing patient care are often reactive, and lacks a predictive component that allows for advanced patient treatment. Likewise, conventional methods of managing patient care often do not utilize computer-implemented technologies such as electronic health record (EHR) systems and artificial intelligence.SUMMARY

[0006] According to some implementations of the present disclosure, a computer-implemented method for predicting a clinical condition or a clinical event in an individual is given. The method includes obtaining, via a communications interface of a computing system and at least one processor of the computing system, electronic health record (EHR) data associated with the individual. The method also includes providing, by the at least one processor, the EHR data as input data to a trained machine learning (ML) model configured to predict a clinical condition or clinical event based at least in part on the EHR data. The method also includes generating, by the at least one processor, and based on an application of the trained ML model to at least the input data, output data indicative of a risk of the individual developing the clinical condition or clinical event a within a prediction window. The method also includes generating, by the at least one processor, and based at least on the output data, a report indicating the risk of the individual.

[0007] According to some implementations of the present disclosure, a system for predicting a clinical condition or a clinical event is provided. The system comprises a communications interface configured to obtain electronic health record (EHR) data associated with the individual; a storage configured to store the obtained EHR data associated with the individual; a memory storing instructions; and at least one processor coupled to the memory, the at least one processor being configured to execute the instructions. The processor is configured to provide the EHR data as input data to a trained machine learning (ML) model configured to predict a clinical condition or clinical event based at least in part on the EHR data. The processor is configured to generate, based on an application of the trained ML model to at least the input data, output data indicative of a risk of the individual developing the clinical condition or clinical event a within a prediction window. The processor is configured to generate, based at least on the output data, a report indicating the risk of the individual.

[0008] According to some implementations of the present disclosure, a computer-implemented method of training a machine learning (ML) model is provided. The method comprises partitioning, by one or more processors of a computing system, an input dataset 24918-2941-55562065472-001007WOPTaccording to at least a plurality of identifiers associated with a plurality of individuals. The method also includes performing operations, by the one or more processors, on the input dataset to generate a regularized dataset. The method also includes performing operations, by the one or more processors, that train an ML model based at least in part on the input dataset to generate a trained ML model configured to generate output data indicative of a risk of the individual developing a clinical condition or clinical event a within a prediction window, wherein the training of the ML model includes optimizing hyperparameters of the ML model using a neural architecture search

[0009] The above summary is not intended to represent each implementation or every aspect of the present disclosure. Additional features and benefits of the present disclosure are apparent from the detailed description and figures set forth below.BRIEF DESCRIPTION OF THE DRAWINGS

[0010] The disclosure, and its advantages and drawings, may be understood from the following description of representative embodiments together with reference to the accompanying drawings. The drawings depict only representative embodiments, and should not be considered as limitations on the scope of the various embodiments or claims.

[0011] FIG. l is a block diagram of an example system for predicting a clinical condition or clinical event, according to aspects of the present disclosure;

[0012] FIG. 2 is a flowchart setting forth steps of a method, according to aspects of the present disclosure;

[0013] FIG. 3 is an illustration of an example architecture for a machine learning (ML) model, according to aspects of the present disclosure;

[0014] FIG. 4 is a graph comparing performance of ML models, according to aspects of the present disclosure;

[0015] FIG. 5 is a graph illustrating validation loss of an ML model, according to aspects of the present disclosure;

[0016] FIG. 6 is a graph comparing training categorization loss without and with quantification embedding, according to aspects of the present disclosure;

[0017] FIG. 7 is a graph illustrating validation loss of an ML model, according to aspects of the present disclosure;

[0018] FIG. 8 is a graph illustrating model finetuning, in accordance with aspects of the present disclosure.

[0019] FIG. 9 is a diagram of a modular transformer construction system, according to 34918-2941-55562065472-001007WOPTaspects of the present disclosure.

[0020] While specific implementations and embodiments are shown in the drawings and described in detail herein by way of example, various modifications and alternative forms are possible. Hence, it should be understood that the present disclosure is not intended to be limited to particular forms disclosed. Rather, the present disclosure covers all modifications, equivalents, and alternatives falling within the spirit and scope of the present disclosure as defined by the appended claims.DETAILED DESCRIPTION

[0021] Certain clinical conditions or events can result in serious complications among hospitalized patients. For instance, hypoglycemic events affecting approximately 20% of hospitalized patients on insulin therapy are linked to increased morbidity, prolonged hospital stays, and higher readmission rates, which can impose a substantial burden on both patient health outcomes and healthcare systems. In particular, hypoglycemia (defined as a blood glucose (BG) level below 70 mg / dL), can be a critical and often unpredictable event for hospitalized patients with diabetes.

[0022] The American Diabetes Association (ADA) 2024 guidelines emphasize a need for providing individualized inpatient glycemic management. Yet conventional care is typically reactive, with treatment plans being adjusted only after a clinical event has occurred. However, various clinical, social and economic factors can contribute to clinical conditions and events. For instance, a prior hypoglycemic event (e.g., 84% of inpatients with a BG<40 mg / dL had a prior BG<70 mg / dL in the same hospital stay), intensive insulin therapy, end-stage renal disease, cognitive impairment, female sex, age >75 years, and social determinants, such as low income and homelessness, may contribute to hypoglycemia. Additionally, many risk factors are not well understood, and may introduce further complexity in treating hypoglycemia and complications stemming from hypoglycemia. For example, recent work found that some inpatient medications can have small but significant impacts on BG. Furthermore, during a hospital stay, insulin requirements can vary significantly due to changes in stress and interruptions in nutritional intake, and standard inpatient insulin dosing protocols fail to consider the full spectrum of individual variability. Other barriers, such as provider confidence, knowledge gaps, complex workflows and clinical inertia may also complicate or impede effective glycemic management.

[0023] While some efforts have been made to anticipate hypoglycemia and other clinical conditions and events, current predictive approaches fail to achieve clinically useful metrics,44918-2941-55562065472-001007WOPTand therefore fall short of accurately predicting future events. In particular, shortcomings of current predictive approaches include (1) focusing on isolated factors (e.g., only BG and insulin doses), (2) failing to leverage temporal dependencies, and (3) lacking flexibility in integrating diverse data sources, such as laboratory, medications, diet, demographic, and other variables. Hence, there remains no widely accepted tool for proactive prediction in hospital. Therefore, there is an urgent need for improved approaches to predict future inpatient conditions and events, such as hypoglycemia, that can account for diverse influences on risk and respond in a clinically meaningful way.

[0024] The present disclosure recognizes that machine learning (ML) can transform management of inpatient diabetes, and other clinical conditions, by leveraging electronic health records (EHR), and drawing from past clinician decisions and patient outcomes to inform individualized care. Unlike statistical models, ML can be more adept in automatic feature selection, can be less reliant on underlying distribution assumptions, and can perform well when nonlinear relationships are present. Recent advancements in Deep Learning (DL), particularly Long Short-Term Memory (LSTM) networks, have improved the prediction of BG levels using outpatient continuous glucose monitoring (CGM) data compared to statistical and standard ML models.

[0025] However, outpatient DL models fit with CGM data do not generalize to the inpatient population. Also, previous ML models that use EHR data have several shortcomings, and challenges remain in obtaining accurate prediction. For instance, prior models oversimply data into single-token representations by tokenizing specific diagnosis codes across a visit, or by combining data into a single token that represents a single visit. Additionally, focus on diagnosis codes omits certain patient data, such as socio-demographic information, laboratory results, medication dosages, and dietary intake. Further, inpatient glucose monitoring relies on point- of-care (POC) measurements taken every 4-6 hours rather than CGM data, which is typically used only in outpatient settings and primarily for patients with type 1 diabetes. And so, models trained on CGM data (not inpatient POC blood glucose data) to predict blood glucose provide only short-term predictions (at 15, 30, 45 minutes), whereas longer-term predictions are clinically more meaningful to allow time for interventions.

[0026] The present disclosure improves upon conventional technologies for monitoring and controlling clinical conditions and events, along with associated complications, by providing a system and method for predicting a clinical condition or event in an individual using ML. As appreciated from description herein, the ML approach described may provide a number of advantages, and may be applied to a variety of applications, including predicting 54918-2941-55562065472-001007WOPTglycemic conditions or events in hospitalized patients, such as patients with diabetes and stress hyperglycemia. Further, the present system and method provide a framework for real-time comprehensive management platform for clinical conditions and clinical events that can be integrated into clinical care.

[0027] As appreciated from description herein, in some aspects, an ML model may be generated based on a foundation model configuration, where the ML model may be trained with self-supervision using clinical data (e.g., medications, laboratory values, procedures, diet orders, and so forth) and tuned to predict various clinical events and clinical conditions, as well as associated complications. In some aspects, various model interpretation methods can be used to connect predictor variables and reveal actionable causes of inpatient complications, thereby allow for prevention and / or treatment.

[0028] FIG. 1 shows a block diagram of an example system 100 for predicting a clinical condition or event of an individual. The system 100 may include one or more processing devices 110 configured to carry out steps, in accordance with aspects of the present disclosure. In general, a processing device 110 may include one or more processors 112, one or more memory 114, one or more output devices 116, one or more input devices 118, or a combination thereof, as well as other components. For example, a processing device 110 may include a computer, a laptop, a tablet, a smart device (e.g., smart phone), a mobile device (e.g., mobile telephone, personal digital assistant (PDA), etc.), a server, a mainframe, awearable device, and so forth. In some example embodiments, the processing device 110 may be a computing system associated with a treatment device of a patient. For example, the processing device could be a monitor or computing system integrated into a medical treatment apparatus. In some embodiments, the processing device(s) 110 may be incorporated into a monitoring and / or treatment device, system, or apparatus.

[0029] A processor 112 can include any processing device, such as a computer processing unit (CPU), graphical processing unit (GPU), microprocessor, digital signal processor, microcontroller, application specific integrated circuit (ASIC), programmable logic device (PLD), field programmable logic devices (FPLD), programmable gate array (PGA), field programmable gate array (FPGA), and so forth. The processor(s) 112 may be configured to carry out various tasks to operate a processing device 110. Further, in some embodiments, the processor(s) 112 may be configured to generate, train, apply, and / or utilize an ML model, in accordance with aspects of the present disclosure. In particular, the ML model may be configured to generate and output prediction data, such as prediction data indicative of a clinical event or condition, based on an application of the ML model. For example, the ML 64918-2941-55562065472-001007WOPTmodel may be configured to generate and output data indicative of a glycemic event or condition. This may be done by using the processor 112 to apply the ML model to a set of input data.

[0030] In some implementations, the processor(s) 112 may be configured generate, train, apply, and / or utilize an ML model with a transformer architecture that includes an encoder-only structure, and dynamic, quantity-specific encodings with advanced temporal attention mechanisms. For example, in some embodiments, the transformer architecture may include categorical embedding (e.g., of medication, laboratory, diet, and so forth), values embedding (e.g., medication dosage, laboratory results, and so forth), and time embedding. In some example embodiments, categorical embedding, values embedding, and time embedding is implemented utilizing an Attention with Linear Biases (ALiBi) architecture. ALiBi is configured to capture temporal relationships between events and clinical measurements stored within the patient’s EHR system 120. The transformer architecture may also include a multihead attention mechanism configured to identify complex relationships among the heterogeneous clinical data types received from the EHR system 120, enabling the model to learn sophisticated patterns in how different clinical factors interact to influence clinical risk. For example, the multi-head attention mechanism may be used to learn patterns across a plurality of user records associated with a plurality of users in the EHR system 120. The transformer architecture may also incorporate a dual-loss configuration. For instance, in some embodiments, the transformer architecture may include dual -loss functions with (1) masked quantity prediction (e.g., mean squared error loss for predicting insulin dosages, laboratory values, and measurement results) and (2) masked categorical event prediction (e.g., crossentropy loss for predicting categorical labels such as medication names, laboratory test types, and diet interruptions). The dual-loss configuration may also use complementary loss functions. Such dual-loss configuration may enhance representation learning for certain clinical risk factors, such as hypoglycemia risk factors, supporting downstream fine-tuning for prediction.

[0031] The memory 114 can include any suitable memory device and / or machine-readable medium capable of storing, encoding, and / or carrying a set of instructions for execution by a processing device and that cause the processing device to perform and / or implement any of the features discussed herein, including solid-state memories, optical media, magnetic media, random access memory (RAM), read only memory (ROM), a floppy disk, a hard disk, a CD ROM, a DVD ROM, flash memory, or other computer readable medium that is read from and / or written to by a magnetic, optical, or other reading and / or writing system that is coupled 74918-2941-55562065472-001007WOPTto the processing device, can be used for the memory or memories. In some embodiments, the memory 114 may include machine-readable instructions for carrying out various steps of methods disclosed herein, and / or other methods. The memory 114 can also store data associated with steps of methods described herein.

[0032] The output device 116 may include various output devices, such as any type of display device (e.g., monitor, touchscreen, and so forth). A display device can include any display technology, including but not limited to one or more display devices using Liquid Crystal Display (LCD) or Light Emitting Diode (LED) technology. In some implementations, an output device 116 may be used to display or report any data and information associated with the features disclosed herein, including the results of analysis and / or prediction data or output generated by a machine learning model, recommendations, and so forth, according to description herein.

[0033] In some implementations, an output device 116 may be used to communicate, provide, and / or store various data, information, and / or signals, for instance, via a wired and / or wireless communication network. For example, in some implementations, the processor(s) and / or output device 116 may generate one or more signal, based on prediction data or output generated by a ML model, and transmit the signal(s) to a device, system, apparatus, and so forth, configured to carry out various tasks responsive to such received signal(s), such as initiating, pausing, or terminating a treatment (e.g., for treating a glycemic condition or event). In some example embodiments, a communications interface 119 may be used to communicate with various devices other than the processing device 110. For example, the communications interface 119 may use a wired, wireless, or other type of network to communicatively couple the processing device 110 and the treatment apparatus 124 and / or the EHR system 120.

[0034] The input device 118 may include various input devices, such as any type of mouse, keyboard, microphone, camera, and so forth. The input device 118 may also include any type of touch-sensitive screen, touch-sensitive pad, a motion sensor, or another computing device with a respective input device communicatively coupled with the system 100 and configured to serve as an input device. In some embodiments, the input device 118 may include at least one electronic interface configured to receive various data associated with the individual. The input device 118 may also include a user interface that may allow a user to interact with the system 100 and / or processing device 110 for any suitable purpose, including providing user input, initiating, pausing, and / or terminating analysis, training, and / or execution of an ML model, adjusting one or more parameters of analysis, model, and so forth.84918-2941-55562065472-001007WOPT

[0035] In some embodiments, the system 100 may include an EHR system 120. The EHR system 120 may be operated using various systems, devices, computers, servers, databases, and other hardware. In some embodiments, the EHR system 120 may store a number N of patient records 122 that may include various data and information, such as admission data, demographic data, medication data, laboratory data, measurement data, diet data, prediction data, and so forth. For example, when an individual is admitted, various information about them (e.g., medical history, current measurements of various parameters, etc.) can be obtained and stored in a patient record 122. As information and / or measurements are generated during the patient’s stay in the hospital (e.g., amounts and / or types of medication administered to the individual, blood glucose and / or other measurements, carbohydrate amounts consumed by the individual, etc.), such information and / or measurements can be stored in the patient record 122. As illustrated in FIG. 1, in some embodiments, a patient record 122 may include prediction data, generated according to aspects of the present disclosure. Such prediction data may include data and information indicative of clinical condition or event predicted for the patient, such as a glycemic condition or event.

[0036] As illustrated in FIG. 1, the EHR system 120 may be connected or connectable to a processing device 110. As such, the processing device 110 (and / or other processing devices or processors) may store data in and / or access data from the EHR system 120. For instance, in some implementations, the processing device 110 may utilize data from patient records 122 to train an ML model, according to aspects of the present disclosure. In some implementations, the processing device 110 may store prediction data, for instance, indicative of a predicted clinical condition or event, such that a clinician with access to the EHR system may be informed of such predicted condition or event.

[0037] In some embodiments, the processing device 110 may be connected or connectable, via wired and / or wireless communication, to a treatment apparatus 124, such as an apparatus configured to administer insulin to a patient (e.g., an insulin pump). For instance, based upon a prediction for a clinical condition or event (e.g., a hypoglycemic condition or event) for a patient, the processing device 110 and / or output device 116 may be configured to generate and transmit one or more signal to the treatment apparatus 124. Responsive to the transmitted signal(s), the treatment apparatus 124 may initiate, modify, pause, and / or cease administration of a treatment. For example, responsive to transmitted signal(s) an insulin pump may initiate, modify, pause, and / or cease administration of insulin.

[0038] Referring to FIG. 2, a flowchart setting forth steps of a process 200, according to aspects of the present disclosure is illustrated. The process 200 may be carried out using any 94918-2941-55562065472-001007WOPTsuitable device, apparatus, or system, such as the system 100 described with reference to FIG.1. In some embodiments, steps of the process 200 may be implemented as instructions stored in non-transitory computer-readable media, as a program, firmware or software, and executed by various general-purpose, programmed or programmable computers, processors or other processing devices. In other embodiments, steps of the process 200 may be hardwired in an application-specific computer, server, processor, dedicated system, or module. Although the process 200 is illustrated and described as a sequence of steps, it is contemplated that the steps may be performed in any order or combination, need not include all illustrated steps, and may include additional steps. The process 200 and / or steps therein may be carried out any number of times, such as intermittently, periodically, and / or subject to user input.

[0039] The process may begin at process block 202 with receiving EHR data associated with an individual or patient. As described, EHR data may include various data and information generated and / or collected during any previous diagnoses, monitoring, and / or treatment. For instance, EHR data may include admission data, demographic data, medication data, laboratory data, measurement data, diet data, and so forth. Hence, the EHR data may include various static and / or sequential data, collected at various points in time or over various periods of time. In some implementations, EHR data may be accessed from an EHR system, database, memory, or other data storage location. The EHR data associated with the individual or patient may include data included in one of a plurality of patient records 122 included in the EHR system 120.

[0040] As indicated by process block 204, received or accessed EHR data may then be provided as input to an ML model configured to predict a clinical condition or clinical event, according to aspects of the present disclosure. In some example embodiments, an ML model based on a transformer architecture may be particularly used recognize important patterns across varied timeframes. For example, an ML model based on a transformer architecture may be used to recognize patterns across a number of days of treatment of a patient, or a plurality of time periods across a time period. Such attribute may be important for certain clinical condition or event prediction, and particularly in situations where various clinical, temporal, and demographic factors may converge to result in the clinical condition or clinical event. Hence, in some implementations, the ML model utilized at process block 204 may be based on self-supervised transformer-based foundation model. In some aspects, the ML model may include multi-step pretraining and dynamic encodings, and alternative loss functions with prediction strategies that may enhance flexibility and prevent oversimplification of data into single-token representations. In some applications, the ML model may be trained for predicting 104918-2941-55562065472-001007WOPTa clinical condition or clinical event, such as hypoglycemia in hospitalized patients with diabetes and stress hyperglycemia, for instance.

[0041] Preprocessing of the input data may also be performed to improve a quality, a property, or a composition of the dataset. For example, one or more preprocessing operations may be performed on the EHR data prior to providing the EHR data to the ML model at process block 204. In some aspects, raw EHR data received at process block 202 may include heterogeneous fields such as admission data, demographic data, medication administration records, laboratory results, diet orders, vital signs, and point-of-care blood glucose measurements, each recorded at irregular time intervals and with varying units and coding systems. Accordingly, the processing device 110 may be configured to normalize, encode, and temporally align these data elements. Data preprocessing may also include data cleaning and normalization steps. For example, medication and laboratory records may be mapped to standardized concept identifiers (e.g., using formulary or laboratory code dictionaries), duplicate or obviously erroneous entries may be removed or corrected, and values may be excluded according to predefined clinical thresholds or user configured parameters. Quantitative variables, such as medication doses, laboratory results, and vital signs may be standard scaled with respect to the training dataset, thereby producing normalized quantities for use with the ML model.

[0042] At block 204, the process includes receiving output from the ML model indicative of a clinical condition or event. For example, the process may include receiving output from the ML model that is indicative of a clinical condition or clinical event for the individual. In some implementations, the output may comprise prediction data representing an estimated risk that the individual will develop a particular clinical condition or experience a particular clinical event within a defined prediction window. For example, the prediction data may be a risk of hypoglycemia within a 4-hour window or up to 36 hours, or more. The time periods presented are merely exemplary, and any time period or combination of time periods could be used as a defined prediction window. The prediction data may include one or more probabilities associated with different outcomes, such as a probability that a measurement of a physical characteristic of the individual may fall below, exceed, or remain within a threshold range. For example, the prediction data may include a probability that a blood glucose level will fall below, exceed, or remain within a predetermined threshold range, and may optionally include confidence scores, classification labels, or other metrics that characterize prediction quality. In some embodiments, the prediction data may also include auxiliary information, such as feature importance indicators, risk factor scores, or temporal risk trajectories derived from the internal 114918-2941-55562065472-001007WOPTrepresentations of the transformer architecture, thereby allowing a clinician to understand which medications, laboratory values, diet orders, or demographic factors contributed most strongly to the predicted risk. The prediction data generated at process block 206 may be stored in association with the individual’s patient record 122 in the EHR system 120, and may be updated intermittently, periodically, or in response to new EHR data becoming available.

[0043] At block 208, the process includes generating one or more reports and / or signals to facilitate monitoring and treatment of the clinical condition or clinical event. For example, the process may include generating the one or more reports and / or signals based at least in part on the output data or the prediction data generated at process block 206. In some implementations, the report may be generated for display on the output device 116, and may include a visual or textual indication of the predicted risk (e.g., numeric probability, risk category, color-coded alert), the relevant prediction window, and optional recommendations for clinical action such as adjusting insulin dosing, ordering confirmatory laboratory tests, modifying diet orders, or increasing the frequency of monitoring. For example, the processor 112 may perform operations that cause the presentation of a graphical user interface by the output device 116. The report may be generated and presented via the graphical user interface in accordance with operations performed by the processor 112.

[0044] In some embodiments, the processor(s) 112 may further generate one or more electronic control signals configured to be transmitted, via a communications interface, to a treatment apparatus 124, such as an insulin pump, an insulin monitor, an intravenous or intramuscular infusion pump, or another infusion device, to initiate, modify, pause, or terminate a treatment in accordance with predetermined clinical rules or user-defined protocols. A “treatment” may include any method, process, or use of a device or composition performed on or for a human or animal subject by, with, or through a medical device that is designed to cure, alleviate, mitigate, prevent, or reduce the likelihood, severity, or progression of a disease, disorder, injury, physiological condition, or symptom, including prophylaxis and palliative care. Examples of treatments include administering, dispensing, delivering, infusing, or otherwise providing a therapeutic agent, fluid, biologic, or composition; applying, generating, or modulating physical energies or stimuli (e.g., electrical, electromagnetic, optical, acoustic, thermal, mechanical, or pneumatic); and monitoring and automatically or semi-automatically adjusting one or more therapy parameters or device operations that directly influence the therapy delivered to the subject. The devices, systems, and apparatuses used to administer such treatment may be communicatively coupled with the system 100. For example, responsive to a predicted increase in hypoglycemia risk above a threshold level, the system 100 may 124918-2941-55562065472-001007WOPTautomatically issue a signal to reduce or temporarily suspend insulin delivery, and / or issue a notification to a clinician for review before implementation of a treatment adjustment. In some implementations, the content and timing of reports and signals generated at process block 208 may be configurable to align with institutional workflows, thereby enabling integration into existing clinical decision support systems while reducing alert fatigue and supporting real-time, risk-adjusted interventions.

[0045] Referring to FIG. 3, an example of an ML model with a transformer architecture, in accordance with aspects of the present disclosure, is shown. The transformer architecture includes an encoder layer 302, which may include a multi-head attention mechanism 302A, a first addition and normalization (Add & Norm) layer 302D, a feed forward network 302E, and a second add & norm layer 302F. The multi-head attention mechanism 302A may also include a time embedding 302B and an attention matrix 302C. In some aspects, the transformer architecture may include token or categorical embedding 304 for EHR data, such as medication, laboratory, diet, and so forth, and quantitative or values embedding 306 for EHR data, such as medication dosage, laboratory results, and so forth. In some implementations, quantities encoded as sinusoidal embeddings may be added to token embeddings, and provided as input into the encoder layer 302.

[0046] As shown in FIG. 3, in some implementations, the encoder layer may include the multi-head attention mechanism 302A that encodes time in a time embedding 302B. The time embedding 302B may use Attention with Linear Biases (ALiBi) embeddings. The ALiBi embeddings may include time associated with various events. The multi-head attention mechanism 302A may be followed by a first Add & Norm layer 302D, a Feed Forward Network 302E, and a second Add & Norm layer 302F, as shown in FIG. 3.

[0047] The multi-head attention mechanism 302A also includes the attention matrix 302C. Generating the attention matrix 302C may include query -key dot products that quantify similarity between given pairs of clinical event tokens in the input sequence. The attention matrix 302C includes similarity scores for a plurality of token pairs included in the EHR data associated with an individual, where higher scores indicate greater contextual relevance between events. The similarity scores are modified by ALiBi temporal biases 302B that apply linear penalties proportional to time differences between events. For example, clinically proximal events (e.g., insulin given at close to a blood glucose measurement) receive positive bias while temporally distant events (e.g., a meal eaten by the individual more than two days before the blood glucose measurement) receive negative bias.134918-2941-55562065472-001007WOPT

[0048] In some example embodiments, the multi-head attention mechanism 302A may include multiple parallel attention matrices. Each attention matrix may help the ML model train to recognize or predict district clinical relationships. For example some heads may be configured to predict medication-laboratory interactions, some heads may be configured to predict diet-blood glucose interactions, and other heads may further be configured to recognize static demographic influences.

[0049] The transformer architecture may also incorporate a dual-loss configuration with quantity loss 308, and categorical loss 310. For instance, in some embodiments, the transformer architecture may include dual-loss functions with (1) quantity loss 308, such as masked quantity prediction (e.g., insulin dosages or laboratory values) and (2) categorical loss 310, such as masked categorical event prediction (e.g., laboratory results and diet interruptions). Such dual-loss configuration may enhance representation learning for certain clinical risk factors, for example, hypoglycemia risk factors, and support downstream finetuning for prediction.

[0050] An example training regimen is given. In some example embodiments, for pretraining the ML model, a comprehensive EHR dataset covering, for example, 8 years of hospital admissions and associated EHR data associated with each individual included in the hospital admissions dataset, may be utilized. Data elements may include time series data, such as inpatient medications, laboratory results, and diet orders. Other data elements may also be included in the example EHR dataset. Data elements may also include static data (e.g., unchanging with time), such as demographics (e.g. age, sex, race), past medical history (e.g. chronic kidney disease), and social history including Social Determinants of Health (SDoH). In particular, laboratory and medications may include (1) laboratory or medication identity, (2) quantity of the medication or laboratory result, and (3) time. Each may be encoded separately and then added to the encoded vector fed into the encoder layer 302. Using the categorical embedding 304, medication or laboratory identity may be encoded as learned embeddings. Using the values embedding 306, quantities for medication doses or laboratory results may be standard scaled to zero mean and unit variance, and then encoded with sinusoidal or learned numeric representations that are traditionally used to encode sequence position. Using the time embedding 302B, ALiBi position embedding techniques may be used to encode temporal distance of clinical events for medications, laboratory results, and diet orders.

[0051] In an example implementation, an individual may have a blood glucose (BG) result of 120 mg / dL at 8pm on a particular day after being admitted at 1pm that same day. BG values would be considered a token with a learned embedding vector representation. The tokens 144918-2941-55562065472-001007WOPTassociated with the BG values would be encoded with the categorical embedding 304. The value of 120 mg / dL would be standard scaled to zero mean and unit variance in reference to all BG labs in the training dataset, resulting in a number close to zero indicating positive or negative standard deviations from the mean. The scaled result may then be multiplied by 1000 and then encoded as a vector using the sinusoidal positional embedding strategy of the values embedding 306. Sinusoidal embedding may then be added to learned embedding of BG laboratory token from the categorical embedding 304 to represent a given laboratory result. Time may be represented as a number of hours from patient admittance to an event (i.e., 7 hours in this example), encoded in the time embedding 302B. During the attention scoring portion of multi-head attention, the ALiBi position encoding strategy of the time embedding 302B may penalize all other time series events by their distance from the 7-hour mark for this lab example. In that way, a medication given upon admittance (in this example, at hour value 0) may be biased to have a lesser impact on the attention score matrix than medications given within an hour of the laboratory (hour values between 6 and 8). Static variables like sex and race may be encoded for the same attention head but not given a position encoding, only an encoding for variable type and variable value that are added together (for example a “sex” encoding and “male” encoding). In some example embodiments, static variables may be exempt from ALiBi attention bias of the time embedding 302B, allowing them to attend to and be attended by all variables equally.

[0052] In some implementations, a transformer architecture with an encoder-only configuration may be utilized, where the architecture includes a series of encoder layers similar to the encoder layer 302, each with at least one respective multi-head attention 302 A sublayer and a feed forward network 302E sublayer. Pre-training on language may use a dual loss configuration, namely a first loss such as the quantity loss 308 based on random tokens that are masked and a second loss such as the categorical loss 310 for assessing accuracy of sentence pairs. In some implementations, a triple loss configuration specific to a dataset for pre-training may be used in the ML model. The first loss may focus on predicting masked quantities, such as medication doses and laboratory values, where the exact medication or laboratory category are unmasked. This loss may consist of a regression problem and use mean squared error (MSE) loss. The second loss may follow a standard masking procedure of encoder models, where token types are masked and then predicted in a classification problem using negative log likelihood (NLL) loss. Notably, the quantity paired with the event type may still be encoded, thus potentially informing the expected token event. A third categorical loss may be based on masking the values of static variables, such as whether that patient was male or female, also 154918-2941-55562065472-001007WOPTusing NLL loss. The three losses may be added together before backpropagation and model weight updates.

[0053] For pre-training, a training dataset may be received or accessed, and the training dataset divided into training, validation, and test subsets (e.g., using GroupShuffleSplit) based on a unique patient identifier (rather than admission) to avoid information leak between train and test data. Regularization methods, including dropout and weight regularization, may be applied to mitigate overfitting. Model structure along with hyperparameters may be optimized using a Neural Architecture Search (NAS). Parameters to be optimized may include batch size, embedding dimension, number of layers, number of heads, feedforward dimension, dropout, and ALiBi slope degree or “m” variable. In some variations to the model architecture, a Big Bird attention may be utilized for longer token sequences and variations on numeric embeddings, including for time, quantities, and position.

[0054] To demonstrate clinical relevance at predicting hypoglycemic events, a subset of patient data may be used for fine-tuning the ML model. In some aspects, another layer may be added to the ML model and trained after pre-training. In some implementations, the ML model may be fine-tuned to predict BG<70 mg / dL within a 4-hour prediction window, for instance, or a probability thereof, using a sequence of events in the prior 5 days, for example. Such taskspecific layer may serve as a binary classifier, enabling real-time predictions of hypoglycemia risk to guide clinical interventions. Model performance may be assessed with bootstrapping from predictions using Fl, precision, recall, and AUPR. Fl scores between the tuned transformer and the developed optimized LSTM) may be compared statistically using a t-test if metrics are normally distributed, otherwise a Wilcoxon Rank-Sum Test.

[0055] Referring to FIG. 4, a graph 400 shows performance of different ML models in an example implementation. The graph 400 is a graph comparing performance of various machine learning models for predicting hypoglycemic events in hospitalized patients, according to aspects of the present disclosure. The vertical axis represents Fl score (ranging from approximately 0.55 to 0.90), a composite statiscial metric that balances precision and recall and is a measure of predictive performance. The Fl score is calculated by finding a harmonic mean of the precision and the recall of the predictions made by each ML model. The precision is calculated by dividing the number of true positive results for a clinical event predicted by each system by the number of all samples predicted to be positive. The recall is calculated by dividing the number of true positive results by the number of all samples that should have been identified as positive. The horizontal axis displays six ML model architectures, including Logistic Regression, XGBoost, Simple RNN, Long short-term memory RNN (LSTM), LSTM 164918-2941-55562065472-001007WOPTwith synthetic minority over-sampling technique (SMOTE), and LSTM with SMOTE and a tuner. Each ML model architecture was used to predict clinical events and evaluated using identical EHR data, prediction task (hypoglycemia within a 4-hour prediction window using 5-day lookback), and evaluation methodology (bootstrap resampling with 95% confidence intervals).

[0056] By way of example, a dataset of EHR data from inpatients aged >18 years treated at hospitals from was analyzed using the present approach. The dataset included 103,824 admissions meeting specific criteria, such as at least one anti-hyperglycemic medication and a length of stay >24 hours, and encompassed a wide range of variables: -14,000 medications, 4,000 labs, 84 diet orders, 25,000 diagnosis codes, 30,500 admission diagnoses, and 11 sociodemographic factors, alongside over 3 million POC BG measurements (approximately 30 POC BG per admission). Notably, 20% of admissions involved at least one hypoglycemic event, with an average of 2.7 events per admission. Demographically, this cohort differs from the Los Angeles County population, with a higher median age and greater representation of White and Black or African American individuals.

[0057] Using a subset of these features, a bi-directional LSTM model was trained with Synthetic Minority Oversampling Technique (SMOTE) and Bayesian hyperparameter tuning (i.e., LSTM+SMOTE+tuner) to predict hypoglycemia in a 4-hour prediction window using a 5-day lookback window. The model achieved higher Fl scores when compared to the other models shown in graph 400, calculated on the training dataset using bootstrap resampling, as compared to logistic regression (LR), XGBoost, simple recurrent neural network (RNN), and LSTM + / - SMOTE without tuning. For example, the LSTM+SMOTE+tuner architecture had an Fl score of 0.79 when evaluated on a held-out 0% test set, outperforming previous hypoglycemia prediction ML model (precision 0.72 and recall 0.59). However, the LSTM's sequential nature limits its ability to capture the diverse temporal patterns that characterize hypoglycemia risk.

[0058] FIG. 5 is a graph 500 illustrating validation loss convergence of a dual-loss transformer foundation model during self-supervised pre-training, according to aspects of the present disclosure. The vertical axis represents combined Mean Squared Error (MSE) and Negative Log Likelihood (NLL) loss (ranging from approximately 0.54 to 0.66), where MSE measures prediction error for masked quantitative values (medication doses, laboratory results) and NLL measures classification error for masked categorical events (medication / laboratory identities). The horizontal axis displays training steps (2.5K to 20. OK174918-2941-55562065472-001007WOPTincrements), representing iterative optimization over the comprehensive EHR dataset comprising 103,824 encounters with self-supervised masking of clinical time series data.

[0059] In one example shown in FIG. 500, the system described herein was used to build a prototype foundation model trained on 103,824 encounters with self-supervision to predict masked laboratory and medications (loss #1, NLL) and masked quantities of laboratory and medications (loss #2, MSE loss). The initial dataset for training included medication and laboratory data for patients with diabetes and stress hyperglycemia who had POC BG measurements during their stay, with an initial focus on 23 medications and laboratory relevant to hypoglycemic events. This architecture included an encoder-only transformer, with 6 encoder layers, each layer consisting of multi-head attention and feed forward network sublayers. The model used a batch size of 32, an embedding dimension of 512, 8 attention heads for each multi -head attention sublayer, a feedforward network dimension of 1024, and a dropout rate of 0.1. Input token vectors were shortened to a maximum of 256 tokens for prototype testing brevity.

[0060] In the example, 15% of tokens were randomly masked for both the classification and regression losses. In some example embodiments, no token was masked for both losses simultaneously. Medication doses and laboratory result quantities were encoded with sinusoidal embedding strategies, and time between events was encoded using the ALiBi embedding strategy. In some example implementations, running the model architecture with a dual-loss strategy may demonstrate clear, steady improvement over several epochs (as illustrated in FIG. 5, showing total loss with two loss functions added together, NLL loss for classification of medications or labs and MSE loss for predicting the dosage or results, respectively, of those events.)

[0061] As illustrated in graph 600 of FIG. 6, quantifying values can have a meaningful impact on classification of specific medications and laboratory results. For instance, direct tokenization of medication dosage and laboratory result quantities as a sinusoidal embedding, and adding quantification embeddings to the token embeddings of specific medications clearly improved model performance, even from the perspective of the categorical NLL.

[0062] In another example, graph 700 of FIG. 7 shows a combination of three losses (y axis of graph 700) and number of training steps (x axis of graph 700) during the self-supervised training phase where values are randomly masked from clinical time courses and predicted as the model's output. Loss indicates distance of predictions from the truth, where a lower loss is closer to the truth. The first point in FIG. 7 is omitted for clarity. The values of loss in FIG. 7 illustrate the validation set loss instead of the training loss, where validation data was used to 184918-2941-55562065472-001007WOPTassess generalization performance. As appreciated from FIG. 7, the model can generalize to unseen data and not yet overfitting. As such, the transformer foundation model, according to embodiments described herein, can learn to predict masked values.

[0063] In yet another example, graph 800 of FIG. 8 illustrates that after this self-supervised training phase, a transformer model, according to embodiments described herein, can be finetuned to predict specific clinical events, such as whether a next blood sugar value is going to be associated with a hypoglycemia event or not. This may be measured using a classification metric called Fl score (y axis of graph 800), which is the harmonic mean of precision and recall, as explained above with reference to FIG. 4, and that is plotted versus the training steps (x axis of graph 800). As illustrated in FIG. 8, the model improves as it sees more iterations of the training data. Further the validation Fl shows that the model can generalize to unseen data and is not yet overfitting.

[0064] As appreciated from description herein, the present approach provides a number of advantages and improvements over conventional approaches. For instance, systems and methods described herein may serve as a cornerstone for a comprehensive, real-time glycemic management platform. By anticipating hypoglycemia risk before it manifests, an ML model approach, as described herein, may empower clinicians to proactively manage BG levels, may facilitate individualized inpatient BG management, and / or provide seamless integration with clinical workflows. Additionally, the present approach may find application in clinical decision support for other inpatient complications, such as acute kidney injury and sepsis, where realtime, risk-adjusted decision-making is crucial. Furthermore, the present approach may have wider impact across other prediction tasks for improving care of inpatients with diabetes, infection, poor wound healing, and delays in surgery or discharge.

[0065] Further, conventional approaches to clinical condition or clinical event prediction rely on models that do not leverage a sequence of events that occurs during the course of patient treatment. To overcome such shortcomings, the present disclosure introduces a new ML approach. Among other advantages and benefits, the present approach provides encoding of quantities as sinusoidal embeddings. Quantitative values in some transformer models are often categorized into ranges for masking loss. By contrast, in the present approach, quantity embeddings (e.g., encoded as sinusoidal embeddings) may be utilized and added to token embeddings. This approach may avoid choosing meaningful ranges for each medication and laboratory by hand, for instance. Also, the present approach provides encoding of time as ALiBi embeddings. ALiBi embeddings typically require tokens in the attention mechanism to prioritize attending to neighboring tokens in a sequence. Rather than using ALiBi to encode 194918-2941-55562065472-001007WOPTsequential position, in some aspects, the present approach may encode distance between events, for instance, in hours, biasing tokens to attend more closely to events that can happen within similar time frames. Linear bias toward recent events may simulate and capture real-world conditions in a clinical setting. For example, meals from 40 hours ago may be less likely to impact BG than meals 1 hour ago, and an ML model, according to the present disclosure, may prioritize recent events better to capture such distinction. Also, the present approach provides multi-dimensional data integration, allowing for incorporating diverse EHR data (e.g., medications, laboratory, diet, and so forth), learning of complex, multi-faceted interactions among variables, and providing a more comprehensive risk profile compared to approaches operating on isolated factors.

[0066] In some example embodiments, a modular transformer construction system is given. Conventional transformer models lack flexibility, and are often tailored to common use cases and strategies, requiring any experimental deviation to be built from scratch.

[0067] Referring to FIG. 9, a modular transformer construction system 900 is shown. To address this shortcoming, a modular transformer construction system 900 for building transformers is provided. The modular system comprises a configuration interface 902 enabling selection from a library of interchangeable components including embedding modules 904, attention modules 906, and loss modules 908. The embeddings modules 904 may include various types of embeddings available to be used in the constructed transformer. Categorical embeddings 904A, values embeddings 904B, and time embeddings 904E may be used, and their respective parameters configured via the interface 902. In addition, sinusoidal embeddings 904C and learned embeddings 904D used in other models may be available. The present modular transformer construction system 900 enables easy use of certain embeddings such as the time embedding 904E ALiBi and rotary positional embeddings, enabling a wider spread of positional embedding strategies than possible with previous approaches. In some aspects, positional embeddings can also be adapted to non-positional numeric embeddings in the present modular software. Further, the present modular software may enable, via the attention module 906, use of Big Bird attentions 0, a sparse attention strategy designed for longer sequences. Each of these enabled strategies may be conveniently swapped in and out, enabling a flexible foundation for transformer model development that can be added to over time without compromising the underlying flexibility. This provides capabilities outside of standard transformer models.204918-2941-55562065472-001007WOPTALTERNATIVE IMPLEMENTATIONS

[0068] Alternative implementation 1. A computer-implemented method for predicting a clinical condition or a clinical event in an individual, the method comprising: obtaining, via a communications interface of a computing system and at least one processor of the computing system, electronic health record (EHR) data associated with the individual; providing, by the at least one processor, the EHR data as input data to a trained machine learning (ML) model configured to predict a clinical condition or clinical event based at least in part on the EHR data; generating, by the at least one processor, and based on an application of the trained ML model to at least the input data, output data indicative of a risk of the individual developing the clinical condition or clinical event a within a prediction window; and generating, by the at least one processor, and based at least on the output data, a report indicating the risk of the individual.

[0069] Alternative implementation 2. The computer-implemented method of implantation 1, wherein the method further comprises accessing the EHR data from a storage or memory of an EHR system associated with a healthcare provider of the individual.

[0070] Alternative implementation 3. The computer-implemented method of implementation 2, wherein the method further comprises performing operations, by the at least one processor, that establish a second trained ML model configured to predict a (i) glycemic condition of the individual, (ii), a glycemic event of the individual, or (iii) any combination of (i)-(ii).

[0071] Alternative implementation 4. The computer-implemented method of implementation 1, wherein trained ML model uses (i) a transformer architecture with an encoder-only structure (ii) a transformer architecture with a decoder-only structure, (iii) a transformer architecture with an encoder-decoder structure, or (iv) any combination of (i)-(iii).

[0072] Alternative implementation 5. The computer-implemented method of implementation 4, wherein the trained ML model includes categorical embeddings and values embeddings associated with the EHR data.

[0073] Alternative implementation 6. The computer-implemented method of implementation 5, wherein the values embeddings comprise sinusoidal embeddings.

[0074] Alternative implementation 7. The computer-implemented method of implementation 5, wherein the at least one processor provides the categorical embeddings and the values embeddings as input to an encoder layer of the transformer architecture.

[0075] Alternative implementation 8. The computer-implemented method of implementation 5, wherein the transformer architecture of the trained ML model uses a multihead attention mechanism that includes a time embedding.214918-2941-55562065472-001007WOPT

[0076] Alternative implementation 9. The computer-implemented method of implementation 5, wherein the method further comprises generating the transformer architecture to include a quantity loss function and a categorical loss function.

[0077] Alternative implementation 10. The computer-implemented method of implementation 9, wherein the time embedding includes an Attention with Linear Biases (ALiBi) embedding configured to encode a temporal distance between clinical events such that attention weights are biased towards at least a portion of the clinical events.

[0078] Alternative implementation 11. The computer-implemented method of implementation 10, wherein the method further includes generating the transformer architecture to include a static-variable classification loss function.

[0079] Alternative implementation 12. The computer-implemented method of implementation 1, generating, by the at least one processor, and based at least in part on the output data, one or more signals configured to cause a treatment apparatus to carry out a task; and transmitting, by the at least one processor and via the communications interface, the one or more signals to the treatment apparatus.

[0080] Alternative implementation 13. The computer-implemented method of implementation 12, wherein the one or more signals is configured to cause the treatment apparatus to, responsive to the one or more signals, (i) initiate a treatment, (ii) pause a treatment, (iii) terminate a treatment, or (iv) any combination of (i)-(iii).

[0081] Alternative implementation 14. The computer-implemented method of implementation 13, wherein the treatment apparatus comprises an insulin pump, and the one or more signals cause the insulin pump to, responsive to the one or more signals, (i) initiate a treatment, (ii) pause a treatment, (iii) terminate a treatment, or (iv) any combination of (i)-(iii).

[0082] Alternative implementation 15. The computer-implemented method of implementation 1, wherein the output data includes a recommendation of (i) a treatment dose, (ii) a treatment dose adjustment, or (iii) any combination of (i) or (ii) based on the risk of the individual.

[0083] Alternative implementation 16. A system for predicting a clinical condition or a clinical event in an individual, the system comprising: a communications interface configured to obtain electronic health record (EHR) data associated with the individual; a storage configured to store the obtained EHR data associated with the individual; a memory storing instructions; and at least one processor coupled to the memory, the at least one processor being configured to execute the instructions to: provide, the EHR data as input data to a trained machine learning (ML) model configured to predict a clinical condition or clinical event based 224918-2941-55562065472-001007WOPTat least in part on the EHR data; generate, based on an application of the trained ML model to at least the input data, output data indicative of a risk of the individual developing the clinical condition or clinical event a within a prediction window; and generate, based at least on the output data, a report indicating the risk of the individual.

[0084] Alternative implementation 17. The system of implementation 16, wherein the EHR data is obtained from a storage or memory of an EHR system associated with a healthcare provider of the individual.

[0085] Alternative implementation 18. The system of implementation 16, wherein the trained ML model uses (i) a transformer architecture with an encoder-only structure (ii) a transformer architecture with a decoder-only structure, (iii) a transformer architecture with an encoder-decoder structure, or (iv) any combination of (i)-(iii).

[0086] Alternative implementation 19. The system of implementation 18, wherein the trained ML model includes categorical embeddings and values embeddings associated with the EHR data.

[0087] Alternative implementation 20. The system of implementation 19, wherein the values embeddings comprises sinusoidal embeddings.

[0088] Alternative implementation 21. The system of implementation 19, wherein the at least one processor is further configured to provide the categorical embeddings and the values embeddings as input to an encoder layer of the transformer architecture.

[0089] Alternative implementation 22. The system of implementation 19, wherein the transformer architecture of the trained ML model uses a multi-head attention mechanism that includes a time embedding.

[0090] Alternative implementation 23. The system of implementation 19, wherein the transformer architecture of the trained ML model includes a quantity loss function and a categorical loss function.

[0091] Alternative implementation 24. The system of implementation 16, wherein the at least one processor is further configured to generate, based at least on the output data, one or more signals configured to cause a treatment apparatus to carry out a task; and transmit the one or more signals to a treatment apparatus via the communications interface.

[0092] Alternative implementation 25. The system of implementation 24, wherein the one or more signals are configured to cause the treatment apparatus to, responsive to the one or more signals, (i) initiate a treatment, (ii) pause a treatment, (iii) terminate a treatment using the treatment apparatus, or (iv) any combination of (i)-(iii).234918-2941-55562065472-001007WOPT

[0093] Alternative implementation 26. A computer-implemented method of training a machine learning (ML) model, the method comprising: partitioning, by one or more processors of a computing system, an input dataset according to at least a plurality of identifiers associated with a plurality of individuals; performing operations, by the one or more processors, on the input dataset to generate a regularized dataset;

[0094] performing operations, by the one or more processors, that train an ML model based at least in part on the input dataset to generate a trained ML model configured to generate output data indicative of a risk of the individual developing a clinical condition or clinical event a within a prediction window, wherein the training of the ML model includes optimizing hyperparameters of the ML model using a neural architecture search.

[0095] One or more elements or aspects or steps, or any portion(s) thereof, from one or more of any of the claims below can be combined with one or more elements or aspects or steps, or any portion(s) thereof, from one or more of any of the other claims or combinations thereof, to form one or more additional implementations and / or claims of the present disclosure.

[0096] While the present disclosure has been described with reference to one or more particular embodiments or implementations, those skilled in the art will recognize that many changes may be made thereto without departing from the spirit and scope of the present disclosure. Each of these implementations and variations thereof is contemplated as falling within the spirit and scope of the present disclosure. It is also contemplated that additional implementations according to aspects of the present disclosure may combine any number of features from any of the implementations described herein.244918-2941-55562065472-001007WOPT

Claims

CLAIMSWhat is claimed is:

1. A computer-implemented method for predicting a clinical condition or a clinical event in an individual, the method comprising:obtaining, via a communications interface of a computing system and at least one processor of the computing system, electronic health record (EHR) data associated with the individual;providing, by the at least one processor, the EHR data as input data to a trained machine learning (ML) model configured to predict a clinical condition or clinical event based at least in part on the EHR data;generating, by the at least one processor, and based on an application of the trained ML model to at least the input data, output data indicative of a risk of the individual developing the clinical condition or clinical event within a prediction window; andgenerating, by the at least one processor, and based at least on the output data, a report indicating the risk of the individual.

2. The computer-implemented method of claim 1, wherein the method further comprises accessing the EHR data from a storage or memory of an EHR system associated with a healthcare provider of the individual.

3. The computer-implemented method of claim 1, wherein the method further comprises performing operations, by the at least one processor, that establish a second trained ML model configured to predict a (i) glycemic condition of the individual, (ii), a glycemic event of the individual, or (iii) any combination of (i)-(ii).

4. The computer-implemented method of claim 1, wherein trained ML model uses (i) a transformer architecture with an encoder-only structure (ii) a transformer architecture with a decoder-only structure, (iii) a transformer architecture with an encoder-decoder structure, or (iv) any combination of (i)-(iii).

5. The computer-implemented method of claim 4, wherein the trained ML model includes categorical embeddings and values embeddings associated with the EHR data.254918-2941-55562065472-001007WOPT6. The computer-implemented method of claim 5, wherein the values embeddings comprise sinusoidal embeddings.

7. The computer-implemented method of claim 5, wherein the at least one processor provides the categorical embeddings and the values embeddings as input to an encoder layer of the transformer architecture.

8. The computer-implemented method of claim 5, wherein the transformer architecture of the trained ML model uses a multi-head attention mechanism that includes a time embedding.

9. The computer-implemented method of claim 5, wherein the method further comprises generating the transformer architecture to include a quantity loss function and a categorical loss function.

10. The computer-implemented method of claim 8, wherein the time embedding includes an Attention with Linear Biases (ALiBi) embedding configured to encode a temporal distance between clinical events such that attention weights are biased towards at least a portion of the clinical events.

11. The computer-implemented method of claim 10, wherein the method further includes generating the transformer architecture to include a static-variable classification loss function.

12. The computer-implemented method of claim 1, wherein the method further comprises:generating, by the at least one processor, and based at least in part on the output data, one or more signals configured to cause a treatment apparatus to carry out a task; andtransmitting, by the at least one processor and via the communications interface, the one or more signals to the treatment apparatus.

13. The computer-implemented method of claim 12, wherein the one or more signals is configured to cause the treatment apparatus to, responsive to the one or more signals, (i) initiate a treatment, (ii) pause a treatment, (iii) terminate a treatment, or (iv) any combination of (i)-(iii).264918-2941-55562065472-001007WOPT14. The computer-implemented method of claim 13, wherein the treatment apparatus comprises an insulin pump, and the one or more signals cause the insulin pump to, responsive to the one or more signals, (i) initiate a treatment, (ii) pause a treatment, (iii) terminate a treatment, or (iv) any combination of (i)-(iii).

15. The method of claim 1, wherein the output data includes a recommendation of (i) a treatment dose, (ii) a treatment dose adjustment, or (iii) any combination of (i) or (ii) based on the risk of the individual.

16. A system for predicting a clinical condition or a clinical event in an individual, the system comprising:a communications interface configured to obtain electronic health record (EHR) data associated with the individual;a storage configured to store the obtained EHR data associated with the individual; a memory storing instructions; andat least one processor coupled to the memory, the at least one processor being configured to execute the instructions to:provide the EHR data as input data to a trained machine learning (ML) model configured to predict a clinical condition or clinical event based at least in part on the EHR data;generate, based on an application of the trained ML model to at least the input data, output data indicative of a risk of the individual developing the clinical condition or clinical event within a prediction window; and generate, based at least on the output data, a report indicating the risk of the individual.

17. The system of claim 16, wherein the EHR data is obtained from a storage or memory of an EHR system associated with a healthcare provider of the individual.

18. The system of claim 16, wherein the trained ML model uses (i) a transformer architecture with an encoder-only structure (ii) a transformer architecture with a decoder-only structure, (iii) a transformer architecture with an encoder-decoder structure, or (iv) any combination of (i)-(iii).274918-2941-55562065472-001007WOPT19. The system of claim 18, wherein the trained ML model includes categorical embeddings and values embeddings associated with the EHR data.

20. The system of claim 19, wherein the values embeddings comprise sinusoidal embeddings.

21. The system of claim 19, wherein the at least one processor is further configured to provide the categorical embeddings and the values embeddings as input to an encoder layer of the transformer architecture.

22. The system of claim 19, wherein the transformer architecture of the trained ML model uses a multi-head attention mechanism that includes a time embedding.

23. The system of claim 19, wherein the transformer architecture of the trained ML model includes a quantity loss function and a categorical loss function.

24. The system of claim 16, wherein the at least one processor is further configured to generate, based at least on the output data, one or more signals configured to cause a treatment apparatus to carry out a task; andtransmit the one or more signals to a treatment apparatus via the communications interface.

25. The system of claim 24, wherein the one or more signals are configured to cause the treatment apparatus to, responsive to the one or more signals, (i) initiate a treatment, (ii) pause a treatment, (iii) terminate a treatment using the treatment apparatus, or (iv) any combination of (i)-(iii).

26. A computer-implemented method of training a machine learning (ML) model, the method comprising:partitioning, by one or more processors of a computing system, an input dataset according to at least a plurality of identifiers associated with a plurality of individuals;284918-2941-55562065472-001007WOPTperforming operations, by the one or more processors, on the input dataset to generate a regularized dataset; andperforming operations, by the one or more processors, that train an ML model based at least in part on the input dataset to generate a trained ML model configured to generate output data indicative of a risk of the individual developing a clinical condition or clinical event a within a prediction window, wherein the training of the ML model includes optimizing hyperparameters of the ML model using a neural architecture search.294918-2941-55562065472-001007WOPT