Predicting the rate of hypoglycemia using machine learning systems.
A machine learning system using electronic medical records predicts hypoglycemic events to optimize basal insulin selection, addressing underestimation issues and reducing healthcare costs by personalizing insulin treatment.
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- SANOFI SA(FR)
- Filing Date
- 2025-07-09
- Publication Date
- 2026-06-01
AI Technical Summary
Current methods struggle to accurately predict hypoglycemic events in diabetes patients, leading to suboptimal glycemic management and increased healthcare costs due to the varying effects of different basal insulins, with existing studies underestimating real-world event rates and costs being heterogeneous and difficult to pool.
A machine learning system trained on electronic medical records to predict hypoglycemic events using structured and unstructured data, employing techniques like natural language processing to identify relevant covariates and filter medical records, and utilizing algorithms such as generalized linear regression and artificial neural networks to determine the appropriate basal insulin for individual patients.
The system provides personalized insulin recommendations, reducing hypoglycemic events and associated costs by accurately predicting event rates and costs, thereby improving patient health outcomes and guiding prescription decisions.
Smart Images

Figure 0007868236000001 
Figure 0007868236000002 
Figure 0007868236000003
Abstract
Description
Technical Field
[0001] Claims of Priority This application claims the benefit of U.S. Provisional Patent Application No. 62 / 689,005, filed Jun. 22, 2018, the entire content of which is incorporated herein by reference.
Background Art
[0002] Machine learning is a subset of artificial intelligence in the field of computer science and often uses statistical techniques to give a computer the ability to "learn" (i.e., gradually improve performance on a specific task) from data without being explicitly programmed.
Summary of the Invention
Means for Solving the Problems
[0003] Generally, one aspect of the invention described herein is embodied as a method that includes an operation of receiving data indicative of medical records of a patient diagnosed with type 1 diabetes. The method includes an operation of using a machine learning system to determine a prediction rate of hypoglycemic events, the machine being trained using data indicative of medical records of a plurality of patients and corresponding rates of hypoglycemic events for each patient. The method also includes an operation of generating a prediction rate for a patient.
[0004] The above and other embodiments may, depending on the circumstances, include one or more of the following functions, either alone or in combination: Each of a group of patients may use the same type of basal insulin. The method may include the operation of determining a second predictive rate of hypoglycemic events using a second machine learning system, the second machine being trained with data showing medical records of a second group of patients and the corresponding hypoglycemic event rates for each of the second patients, each of the second group of patients using a second type of basal insulin, which is different from a first type of basal insulin, and the method may further include the operation of comparing the first predictive rate with the second predictive rate. The method may include the operation of recommending basal insulin for a patient based on the comparison. The method may include the operation of determining multiple predictive rates of hypoglycemic events for a second group of patients by providing data corresponding to the respective medical records of each of the second group of patients to a machine learning system; the operation of identifying one or more covariates in the data correlated with the predictive rates of hypoglycemic events, based on that data and the multiple predictive rates of hypoglycemic events; and the operation of generating a report identifying one or more covariates and the corresponding predictive rates of hypoglycemic events. The method may include the operation of determining multiple predictive rates of hypoglycemic events for a second group of patients by providing data corresponding to the respective medical records of each of the second group of patients to a machine learning system, each of the second group of patients having the same covariates, and the method may further include the operation of generating a report identifying the covariates and the corresponding predictive rates of hypoglycemic events.
[0005] The disclosure also provides a computer-readable storage medium connected to one or more processors, which, when executed by one or more processors, causes one or more processors to perform operations by implementing the methods provided herein.
[0006] This disclosure further provides a system for carrying out the methods provided herein. The system includes one or more processors and a computer-readable storage medium connected to one or more processors and storing instructions, which, when executed by one or more processors, are processed by one or more processors. The operations are performed by carrying out the methods provided herein.
[0007] It should be understood that the embodiments provided herein may include any combination of the embodiments and functions described herein. That is, the embodiments provided herein are not limited to combinations of embodiments and functions specifically described herein, but also include any other suitable combination of embodiments and functions provided herein.
[0008] Details of one or more embodiments of the subject matter described herein are briefly outlined in the accompanying drawings and the following embodiments for carrying out the invention. Other features, aspects, and advantages of the subject matter will become apparent from the embodiments for carrying out the invention, the drawings, and the claims. [Brief explanation of the drawing]
[0009] [Figure 1] This figure shows the environment in which a machine learning model is trained to predict the expected rate of hypoglycemic events. [Figure 2] This flowchart shows an example of a process for classifying a single event as an ED / outpatient, inpatient (secondary), or inpatient (primary). [Figure 3] This figure shows an exemplary process for determining the hypoglycemia rate for various covariates. [Figure 4] This figure shows an example of determining the hypoglycemia rate for various covariates. [Figure 5] This is a flowchart illustrating an example of the process of generating a pre-trained machine learning model using patient data. [Modes for carrying out the invention]
[0010] The same reference number and name in various drawings refer to the same element.
[0011] Diabetes mellitus (DDI) is the seventh leading cause of death and the leading cause of prevalence in the United States. An estimated 29.1 million people in the U.S. are affected by DDI, with 1.4 million new cases diagnosed each year. The number of affected individuals is projected to exceed 54.9 million by 2030. In 2012, the total cost of diagnosed diabetes in the U.S. was $245 billion ($176 billion in direct healthcare costs and $69 billion in lost productivity).
[0012] Given the significant and increasing burden of the disease, complications associated with diabetes are becoming increasingly important for effective prevention and management. Hypoglycemia is a frequent, and sometimes fatal, adverse effect of insulin and oral antidiabetic drugs (OADs) in patients with diabetes. In addition to the current risks posed by hypoglycemia, recurrent episodes can lead to anxiety about future episodes, and have been shown to constitute both patient-led and physician-led barriers to optimal glycemic management. The resulting elevated hemoglobin A1c (HbA1c) levels have been linked to an increased risk of microvascular (and sometimes macrovascular) complications.
[0013] It is estimated that patients with type 1 diabetes (T1DM) experience an average of two mild hypoglycemic events per week and one severe event per year. However, event rates in randomized clinical trials for type 1 patients range from 0.15 severe events to 88.3 non-severe events per patient per year. Event rates for patients with type 2 diabetes (T2DM) vary considerably between studies, ranging from 0.05 to 26.6 severe and non-severe events per patient per year. However, studies suggest that these studies significantly underestimate the true real-world event rates of hypoglycemia, particularly with respect to severe events.
[0014] The average cost for hypoglycemic events is also heterogeneous across studies and difficult to pool due to differences in the definition of hypoglycemia and cost estimation methods. Current estimates for the average cost for hospitalized hypoglycemic events range from $2,205 to $17,564. The average outpatient cost per hypoglycemic event ranges from $148 to $501.
[0015] Different basal insulins exhibit different rates of hypoglycemia. For example, numerous studies have shown the superiority of insulin glargine 100 units / mL (Lantus) over protamine insulin in relation to severe hypoglycemic events. Some basal insulins, such as insulin glargine 300 units / mL (Toujeo), exhibit flatter, longer-term pharmacokinetic and pharmacodynamic profiles compared to others with sustained glucose control exceeding 24 hours.
[0016] Therefore, the burden of hypoglycemia in the United States is significant, and substantial benefits can be gained by identifying the appropriate type of basal insulin for each patient. The resulting cost savings for payers due to reduced healthcare costs for these patients can help guide payers in prescription decisions and drug pricing negotiations.
[0017] Figure 1 shows an environment 100 in which a machine learning model is trained to predict the expected rate of hypoglycemic events. The system described herein uses electronic medical record data (EMR) 102 to train the machine learning system (EMR is obtained from a number of different sources, including, but not limited to, hospital records and physician records). In some embodiments, the EMR may include information about demographic and socioeconomic categories, coded diagnoses and procedures, prescribed and administered medications, laboratory results, and clinical management data. In some embodiments, the EMR is processed by a processor 104. For example, the EMR may include both structured and unstructured data. Structured data may include information such as the date of the visit and the patient's name. Unstructured data may include free-form text added by the physician (e.g., physician's notes, visit summaries). The EMR processor can extract facts from the unstructured data of the EMR using techniques such as natural language processing and transform the unstructured data into structured EMR data 106.
[0018] In some embodiments, the EMR processor 104 can filter portions of the medical records. For example, the EMR processor 104 can select only the medical records of patients that share values for specific variables (called covariates). Examples of covariates include, for example, sex, geographical region, rate, age group, insurance company, years since diagnosis, HbA1c range, body mass index, blood pressure range, diabetic complications, alcohol and / or drug use, and any other physiological or demographic characteristics. The EMR processor 104 can be, for example, one or more computer systems as described below.
[0019] In some embodiments, the same manually created covariates are used for descriptive analysis, hypoglycemia rate prediction modeling, and cost estimate analysis. The predetermined covariates may be based on an expert clinical review of the literature around hypoglycemic events and cost predictors. The predetermined covariates are defined using ICD-9 and ICD-10 diagnostic codes, laboratory values, and drug names and / or National Drug Code (NDC) codes. An EHR dataset (not insurance claims) is used to identify covariates unless the covariate is cost-related, in which case insurance claim data is used.
[0020] A default lookback period of one year prior to treatment initiation is used, although this period can also vary depending on how long the covariates are assumed to persist. For example, the lookback period for cancer can be five years. The reason is that if a patient was diagnosed with cancer more than five years ago and there has been no subsequent diagnosis, it is unlikely that the patient has active cancer at the index date. For covariates such as gender and race that do not affect the likelihood of a patient being captured in the dataset based on the length of their medical history, the lookback period can be eight years or limited only by the amount of data available. Even for irreversible and chronic conditions, the maximum lookback period can be eight years or limited only by the amount of data available.
[0021] An example of the categories of covariates used in both the hypoglycemia rate model and the cost estimate analysis and their lookback periods can be as follows: 1. Demographics a. The lookback period for these covariates is eight years.
[0022] 2. Socioeconomics b. The lookback period for these covariates is eight years.
[0023] 3. Comorbidities c. The lookback period for these covariates ranges from one year (for reversible / acute conditions) to eight years (for irreversible / chronic conditions).
[0024] d. Charlson Comorbidity Index (CCI) score as a distinct covariate within the comorbidity category. The CCI is a measure of a patient's comorbidity status, including diabetic complications, associated with the period prior to the indicator (including the diabetic complications category).
[0025] 4. Diabetic complications e. The review period for these covariates is 8 years.
[0026] 5. Diabetic disease status a. The review periods for these covariates are 1 year for pre-hypoglycemic events and 8 years for diabetes over a known period in the dataset.
[0027] 6. Use of medications a. The review period for these covariates is one year.
[0028] The additional covariates are included as part of the cost estimation analysis only (and not the hypoglycemia rate prediction), because these covariates are predicted to be drivers of hypoglycemia-related costs (rather than hypoglycemia rates): 1. Physician's specialization: f. The review period is 2 years for all available data on the “physician specialization” covariate, and 2 years for “most commonly prescribed insulin by physicians.”
[0029] 2. Previous average cost of hypoglycemia events g. The review period for these covariates is one year.
[0030] Another set of covariates is determined. This set of covariates is not prespecified but includes all comorbidities, procedures, and prescriptions that existed one year prior to the patient's index date (collectively referred to below as "markers") which are included in the predictive model.
[0031] In some embodiments, the following methods are used for unsupervised covariate generation.
[0032] 1. The distance between markers is defined: a. Each marker is associated with a vector containing the set of patients for which the marker is true or false.
[0033] b. In this case, the distance between markers becomes the Jackard distance between vectors.
[0034] 2. Next, clusters were generated based on these distances: a. Hierarchical clustering based on the distance matrix was performed to obtain a cluster hierarchy. The average distance between clusters was used to link the clusters hierarchically.
[0035] 3. The hierarchy levels for extracting clusters are defined: a. The "mismatch" method was used to determine at which level of the hierarchy we wanted to extract clusters. "Mismatch" refers to the mismatch in mean distance between linked clusters: a large value suggests that the clusters should not be linked.
[0036] 4. A cluster is selected that "demonstrates" a large number of patient treatments: a. Of the 1500 clusters formed, 100 were selected based on the treatment of most patients.
[0037] b. Patient treatment was said to “indicate” a cluster if, in the year prior to the indicator date, the patient had any of the diagnoses, procedures, or prescriptions (e.g., markers) that constituted a cluster.
[0038] 5. Medical rationalization of clusters: a. Clinical experts then examine the resulting clusters against medical logic to enable the parameters used for unsupervised cluster generation.
[0039] In some embodiments, the EMR records are filtered, for example, by a filter 116 after generating the structured medical records 106, but before generating the training data 108.
[0040] Structured EMR data 106 is used to generate training records 108 (collectively, training sets). These training records may represent available data in the structured EMR data. For example, in some embodiments, a portion of the structured EMR records is used to train a machine learning system, and the remaining records are used to enable the trained machine learning system. In some embodiments, training records are created at the "patient treatment" level, where the patient is undergoing basal insulin therapy. Thus, a large number of training records are created for each individual patient. The unit of analysis is "patient treatment," defined as the period during which the patient is observed for basal insulin therapy in the dataset (the period between the treatment indicator and the end of treatment observation). Hypoglycemic events are the target endpoint only within this patient treatment period.
[0041] The treatment index date is defined as either the very first start of any basal insulin prescription; or the change in prescription from one basal insulin to another. Baseline basal insulins include: Gla-300, Gla-100, IDet, IDeg, and NPH. The index basal insulins included in the study were: Gla-300, Gla-100, IDet, and IDeg.
[0042] For the purpose of calculating the rate, "period" is interpreted as the period obtained by subtracting the total length of stay of all hospitalized patients during this period from the patient treatment period mentioned above.
[0043] In some embodiments, the indicator date is the date of the first prescription of the BI, or one of the base dates. This is the date of the prescription change from one insulin to another. Treatment completion is defined as the end of the follow-up period in the dataset, the prescription change from benchmark basal insulin to another BI, or one year after the treatment benchmark date (whichever comes first).
[0044] In one embodiment, patient treatment is excluded from training data if it meets any of the following criteria: 1. Treatment using multiple types of basal insulin: i.e., patient treatment initiated within one week (before or after) the initiation of another treatment in the same patient.
[0045] 2. Patient treatment with any inactivity period longer than 270 days within the 365 days prior to the indicator date (inactivity is defined as the absence of timestamped data in the relevant tables in the dataset).
[0046] 3. Treatment of patients whose treatment period is less than one day.
[0047] Furthermore, hospital stays are often excluded from the patient treatment period because patients are typically switched to standard basal insulin according to the hospital provider's prescription guidelines upon admission. Therefore, hypoglycemic events during hospital stays are not considered to be attributable to the standard basal insulin.
[0048] The training record may contain information about hypoglycemic events. In some embodiments, hypoglycemic events are counted within the patient treatment period. The period for determining the hypoglycemia rate was interpreted as the patient treatment period minus the total duration of hospital stays during this period. Hypoglycemic events can be the expected output of the training set. For example, the training set can be used to train a machine learning system to determine the expected number of hypoglycemic events within a fixed period (e.g., 1 month, 6 months, 1 year, 5 years, etc.). Alternatively, the training set can be used to train a machine learning system to determine the expected number of hypoglycemic events within a period (e.g., 1 month, 6 months, 1 year, 5 years, etc.).
[0049] The hypoglycemic event rate can include both severe and non-severe events. Figure 2 is a flowchart of an example of the process for classifying hypoglycemic events as severe or non-severe. The definition of "severe" hypoglycemia may include, for example, ICD-9 / 10 codes that are inherently severe, such as the administration of intramuscular glucagon. Furthermore, EMR's natural language processing is used to identify hypoglycemia. In terms of severity, any hypoglycemic event that was not severe is defined as "non-severe."
[0050] In some embodiments, a single event is defined as hypoglycemia if any of the following criteria are met: 1. ICD-9 and 10 Hypoglycemia Diagnostic Codes 2. Laboratory plasma glucose level ≤70 mg / dL 3. Administration of intramuscular glucagon 4. NLP Output In some embodiments, an NLP-recognized hypoglycemic event is defined as any mention of hypoglycemia, except for those marked with negative emotion or a historical event. For example, a “mention” of hypoglycemia may be any event relating to the common expression “*hypoglycemia*,” unless the term is precisely one of “hypoglycemia awareness,” “hypoglycemia unawareness,” or “neonatal hypoglycemia.” A “negative emotion” may be any marking in which the mention of hypoglycemia was negative, for example, indicating that hypoglycemia did not occur. A historical event is a record indicating that the mention is about a past event (e.g., “the patient has a history of hypoglycemia”). In some embodiments, a list of negative emotion and historical keywords is used to filter out any accompanying mentions of hypoglycemia.
[0051] In some embodiments, a maximum number of single hypoglycemic events are counted per calendar day. For example, if a single hypoglycemic event is recorded at multiple locations in treatment or under multiple defining criteria, only one event is counted.
[0052] In some embodiments, a hypoglycemic event is defined as severe if any of the following conditions are met: 1. Hypoglycemia is severe by default according to ICD-9 or ICD-10 diagnostic codes (ICD-9 249.30; 250.30; 250.31; 251.0; ICD-10 E08.641; E09.641; E10.641; E11.641; E13.641; E15).
[0053] 2. The ICD code for hypoglycemia is flagged as being diagnosed at admission, the primary reason for treatment at discharge, or present at admission.
[0054] 3. The onset of hypoglycemia occurred on the same day as the emergency department (ED) visit or the patient's admission to the hospital.
[0055] 4. Plasma glucose level is <54 mg / dL.
[0056] 5. Intramuscular glucagon was administered.
[0057] 6. NLP references to hypoglycemia are accompanied by severity descriptors, including severity terms (e.g., "severe") and attributes (e.g., "emergency").
[0058] 7. The NLP hypoglycemic event occurred on the same day as the (ED) home visit or hospital admission.
[0059] The machine learning environment 110 may include a machine learning trainer 112. The machine learning trainer 112 can train a machine learning model 114 to predict the expected rate of hypoglycemic events for different patients. The trained machine learning system can be used in a variety of different ways, including identifying the most appropriate basal insulin associated with the patient's hypoglycemic outcomes, thereby improving the patient's health.
[0060] In general, machine learning can encompass a wide variety of different techniques used to train machines to perform specific tasks without being specifically programmed to do so. Machines are trained using different machine learning techniques, including, for example, supervised learning, unsupervised learning, and reinforcement learning. In supervised learning, the machine is provided with a target input and a corresponding output. The machine tunes its function to provide the desired output when given an input. Supervised learning is commonly used, for example, to teach a computer to solve a problem where the outcome is deterministic, and the training set 108 is used to train a trained machine learning model 114 to predict the likelihood of a hypoglycemic event occurring in a given patient or group of patients. In contrast, in unsupervised learning, an input is provided, but no corresponding desired output is provided. Unsupervised learning is commonly used in classification problems, such as customer segmentation (for example, segmenting patients into different groups based on characteristics associated with hypoglycemic events). Reinforcement learning describes an algorithm in which a machine makes decisions using a trial-and-error method. Feedback informs the machine when a good or bad choice has been made. Next, the machine adjusts its algorithm accordingly.
[0061] During the training process, different algorithms are used, including, in particular, generalized linear regression (GLM). Poisson GLM is an algorithm used to model discrete counts based on individual inputs.
[0062] To develop a trained machine learning system that can accurately predict the rate of hypoglycemic events, the machine learning model 114 may be trained with unambiguous information. However, patient treatment for diabetes (and other medical conditions) can be fluid. For example, a patient may switch from one type of basal insulin to another. Therefore, in some embodiments, the EMR 102 (and thus the corresponding training data 108) of some patients may be excluded from the training data.
[0063] Figure 2 is a flowchart illustrating an example of a process for classifying a single event as an ED / outpatient (Result 210), an inpatient (secondary) (Result 212), or an inpatient (primary) (Result 214).
[0064] A patient is defined as a primary inpatient (Result 214) if all of the following criteria are met: 1. When an event is linked to a visit schedule and the visit type is not ED (Process 202). 2. Hypoglycemic events are not only identified by natural language processing (step 204). 3. The issue can be identified using the diagnostic sheet (step 206). 4. The diagnosis will be marked as “Discharge Diagnosis,” “Admission Diagnosis,” or “Present at Admission.”
[0065] A patient is defined as a secondary hospitalization patient if a single event meets all of the following criteria: 1. Hypoglycemic events are not only identified by NLP (Step 204). 2. If the event is found using the diagnostic chart, the diagnosis is not marked as “Discharge Diagnosis,” “Admission Diagnosis,” or “Present at Admission” (Step 206). 3. The event is linked to the visit schedule using the PTID, the visit type is hospitalized, and the date of the hypoglycemic event is between one day before the visit start date and one day after the visit end date (this provides a buffer for linking hypoglycemic events to visits, considering that the dataset often lacks exact date matching) (Step 208).
[0066] In some embodiments, as described above, secondary hospitalization events are excluded because the patient is often switched to a different basal insulin and the patient's dosage is changed during the patient's hospital stay. Therefore, any hypoglycemic events during this period are not considered to be attributable to the patient's usual insulin.
[0067] For example, an event is defined as an outpatient / ED (Result 210) if it satisfies any of the following conditions: 1. The event is linked to the visit schedule using the PTID, the visit type is ED, and the hypoglycemic event day is between 1 day minus the visit start date and 1 day plus the visit end date (Step 202).
[0068] 2. The events are identified using only natural language processing (step 204).
[0069] 3. If an event is identified using a diagnostic form, it is not identified using NLP alone, the diagnosis is not marked as discharge diagnosis, admission diagnosis, or present at admission, and is not linked to the visitation form (step 206).
[0070] Training records are used to train machine learning systems once they are created. Different types of machine learning models are trained.
[0071] For example, a trained learning model is embodied as a generalized linear model. Different types of generalized linear models may be appropriate for various scenarios. A zero-plus negative binomial distribution GLM (zNBGLM) was used. The reason is that zNBGLM discretely counts events occurring over a given period, and the probability of hypoglycemia counts observed in the data is the highest. This is because it models by estimating the hypoglycemic event rate per patient. One disadvantage of zNBGLM is that the model allows for too many degrees of freedom and tends to overfit data to small segments, thereby reducing generalization performance. Another type of generalized linear model is Poisson GLM. Poisson GLM is well-suited to modeling discrete counts but does not tolerate "excessive variance" (i.e., it suppresses variance to be equal to the mean). In Poisson GLM, the number of hypoglycemic events was used as the target variable (outcome), and the length of observation was used as the offset variable.
[0072] In another example, a trained learning model can be embodied as an artificial neural network. An artificial neural network (ANN), or connectionist system, is a computing system inspired by the biological neural networks that make up the brains of animals. An ANN is based on a collection of connected units or nodes called artificials. Each connection can transmit signals from one artificial neuron to another, much like the synapses in a biological brain. The receiving artificial neuron can process the signal and then send it to further artificial neurons connected to it.
[0073] In a typical ANN embodiment, the signals at the junctions between artificial neurons are real numbers, and the output of each artificial neuron is calculated by some nonlinear function of the sum of its inputs. The junctions between artificial neurons are called "edges." Artificial neurons and edges may have weights that are adjusted as learning progresses (for example, each input to an artificial neuron is weighted separately). The weights increase or decrease the strength of the signal at the junction. Artificial neurons may have a threshold such that a signal is emitted only when the aggregate signal crosses that threshold. The transfer function along an edge usually has an S-shape, but it can also take the form of other nonlinear functions, piecewise linear functions, or step functions. Generally, artificial neurons are gathered into multiple layers. Different layers can perform different kinds of transformations on their inputs. The signal travels from the first layer (input layer) to the last layer (output layer), sometimes crossing each layer multiple times.
[0074] In some embodiments, a machine learning system is used to identify hypoglycemic event rates and hypoglycemic costs. Lasso regression (LASSO) regularization is used to select variables. To validate the model, the model is developed ("trained") for 80% of each treatment-specific cohort (referred to as the "training set"). Ten-fold cross-validation is used to signal model selection and model parameter optimization. The model is then validated for the remaining 20% of each treatment-specific cohort (internal validation). Bootstrapping is used to assess the variability of the model estimates (i.e., to generate confidence intervals).
[0075] Once a machine learning system is trained, it can be used to identify patients who are likely to experience fewer hypoglycemic events when treated with one type of basal insulin compared to another. For example, a model can be trained for each type of basal insulin and for the number of severe and non-severe hypoglycemic events. Each model can then be applied to the entire population treated with basal insulin to obtain insulin-specific hypoglycemia rate predictions (i.e., estimates of the hypoglycemia rate in the overall population if all patients were using a particular basal insulin).
[0076] Next, the system can compare the rate of hypoglycemia among patients based on another variable.
[0077] Figure 3 shows an example of using a trained machine learning model. Patient EMR310 is processed. This generates input 312. Input 312 is provided to each of the trained machine learning models, in this example, the Gla-300 trained machine learning model, the Gla-100 trained machine learning model 304, the IDet trained machine learning model 306, and the IDeg trained machine learning model 308. Each model can produce one output. For example, the Gla-300 trained machine learning model 302 produces the Gla-300 output 314, the Gla-100 trained machine learning model 304 produces the Gla-100 output 316, the IDet trained machine learning model 306 produces the IDet output 318, and the IDeg trained machine learning model 308 produces the IDeg output 320.
[0078] In some embodiments, as described above, two pre-trained machine learning models are generated for each type of basal insulin. The first pre-trained machine learning model is trained to determine the expected rate of severe hypoglycemic events. The second pre-trained machine learning model is trained to determine the expected rate of non-severe hypoglycemic events.
[0079] The results of processing these covariates using a machine learning system are analyzed to identify correlations between different covariates and different hypoglycemia rates for different types of basal insulin. For example, a linear regression model is used to identify correlations between different covariates and the outcomes predicted by this model.
[0080] In some embodiments, the covariates of interest are also determined by the machine learning system. For example, the machine learning system is trained to cluster individuals based on, for instance, the rate and severity of hypoglycemic events across multiple variables. In this way, the machine learning system can identify covariates that might otherwise go unnoticed.
[0081] In some embodiments, EMR and training datasets are used to construct cost models for predicting the cost of hypoglycemic events in a T2DM population. In some embodiments, the datasets used for treatment cost modeling included all hypoglycemic events in EHRs for T2DM patients who were at least 18 years old at the time of the event and had linked billing data. Severe hypoglycemic events with a cost of $0 were excluded. The study period and covariates were the same as those described for the hypoglycemia prediction models, except where data limitations prevented the creation of covariates.
[0082] Gradient-boosted trees (which improve the performance of subsequent trees using prediction errors from previous decision trees), which had previously been successful in cost forecasting, were used for cost estimation; these gradient-boosted trees made it possible to capture the complex nonlinear relationships underlying hypoglycemia costs. The cost estimator was applied to subgroups identified as drivers of the hypoglycemia differential rate, and the cost per hypoglycemic event was estimated for each subgroup. Where a key defining variable was missing for a subgroup due to data limitations, the total model cost estimate for one hypoglycemic event was used for that subgroup. Cost savings at the subgroup level were calculated by applying the subgroup-specific cost estimates of hypoglycemic events to the delta hypoglycemic event rate between the comparison criterion and the reference BI.
[0083] Figure 4 shows an example of determining the hypoglycemia rate for various covariates. Input data 402 provided to the machine learning system and corresponding output data 404 from the machine learning system described above are provided to the statistical analysis system 406. The statistical analysis system can identify correlations and relationships between different variables in the output data using various statistical techniques. The correlations and relationships are presented as a report 408.
[0084] Another application of trained machine learning models involves identifying the appropriate type of basal insulin for a particular patient. For example, a trained machine learning system could analyze a patient's medical records. It can be retrieved and entered.
[0085] Medical records were provided to each of the trained machine learning models. As described above, each model was used to predict the number of severe or non-severe hypoglycemic incidents (or the probability of a hypoglycemic incident occurring).
[0086] The system can suggest a basal insulin based on the model's results. In some embodiments, the system can suggest a basal insulin with a lower risk of severe hypoglycemic events (or non-severe hypoglycemic events). In some embodiments, the system can determine the cost savings by using a particular basal insulin regimen compared to another associated with hypoglycemic events and suggest a solution that reduces hypoglycemia-related costs. In some embodiments, the system can suggest a basal insulin that reduces the hypoglycemic event rate, but if two different basal insulins produce results within a threshold (e.g., within 1%, 5%, or 10% effect), the system can suggest the basal insulin with the lower cost.
[0087] In another embodiment, a trained machine learning model is used to identify patients who are more likely to experience hypoglycemic events. For example, the system can access a patient's medical records, identify the type of basal insulin the patient is using, and process the medical records using a corresponding trained machine learning model. The trained machine learning model generates indicators of the likelihood or expected frequency and / or severity of hypoglycemic events. If the indicators exceed a threshold (e.g., more than one event per week, or a greater than 20% probability of a severe hypoglycemic event), the patient and / or the patient's physician are notified.
[0088] Figure 5 is a flowchart of an example of process 500, which generates a trained machine learning model using patient data. Process 500 is performed by one or more computer systems as described below.
[0089] Process 500 receives data from 502 that shows the medical records of patients diagnosed with diabetes mellitus.
[0090] Process 500 is 504, which uses a machine learning system to determine the expected rate of hypoglycemic events, and the machine is trained with data showing the medical records of multiple patients and the corresponding hypoglycemic event rates for each patient. In some embodiments, multiple patients are taught using the same type of basal insulin.
[0091] Process 500 is 506, which generates a predicted rate for the patient.
[0092] In some embodiments, process 500 may include determining a second predicted rate of hypoglycemic events using a second machine learning system, the second machine being trained with data showing medical records of a second group of patients and the corresponding rate of hypoglycemic events for each of the second patients, each of the second group of patients using a second type of basal insulin, which is different from a first type of basal insulin; process 500 may further include comparing the first predicted rate with the second predicted rate.
[0093] In some embodiments, procedure 500 may include recommending basal insulin for the patient based on comparison.
[0094] In some embodiments, the process 500 is performed on the respective medical records of a second group of patients. This may include: determining multiple predicted rates of hypoglycemic events for a second set of patients by providing corresponding data to a machine learning system; identifying one or more covariates in the data correlated with the predicted rates of hypoglycemic events, based on that data and the multiple predicted rates of hypoglycemic events; and generating a report that identifies one or more covariates and the corresponding predicted rates of hypoglycemic events.
[0095] In some embodiments, the process 500 includes determining multiple predicted rates of hypoglycemic events for a second group of patients by providing a machine learning system with data corresponding to the respective medical records of the second group of patients, each of the second group of patients having the same covariates, and the process 500 may further include generating a report that identifies the covariates and the corresponding predicted rates of hypoglycemic events.
[0096] The embodiments and functional operations of the subject matter described herein are implemented as digital electronic circuits, tangibly embodied computer software or firmware, computer hardware (including structures disclosed herein and their structural equivalents), or one or more combinations thereof. The embodiments of the subject matter described herein are implemented as one or more computer programs (i.e., one or more modules consisting of computer program instructions encoded in tangible, non-temporary program carriers for execution by or for controlling the operation of a data processing device). Computer storage media may be machine-readable storage devices, machine-readable storage boards, random or serial access memory devices, or one or more combinations thereof.
[0097] The term “data processing device” refers to data processing hardware and encompasses all types of devices, machines, and equipment for processing data, including programmable processors, computers, or multiple processors or computers. A device may also be a dedicated logic circuit (e.g., an FPGA (Field-Programmable Gate Array) or an ASIC (Application-Specific Integrated Circuit)), or otherwise. In addition to hardware, a device may, in some cases, include code that creates the execution environment for computer programs (e.g., code that constitutes processor firmware, protocol stacks, database management systems, operating systems, or one or more of these).
[0098] Computer programs, also called or written as programs, software, software applications, modules, software modules, scripts, or code, are written in any form of programming language, including compiled or interpreted languages, or declarative or procedural languages. Computer programs are deployed in any form, including as standalone programs or as modules, components, subroutines, or other units suitable for use in a computer environment. Computer programs may, but may not, correspond to files in a file system. A program may be stored in a part of a file that holds other programs or data (e.g., a markup language document, a single file dedicated to the program in question, or one or more scripts stored in multiple collaborative files, e.g., files that store one or more modules, subprograms, or parts of code). A computer program is deployed so that it can be run on one computer, or on multiple computers located in one place, or distributed across multiple locations and interconnected by a data communication network.
[0099] The processing and logical flows described herein operate on input data and generate output. It is executed by one or more programmable computers, which in turn run one or more computer programs to perform various functions. The processing and logic flow is also executed by dedicated logic circuits (e.g., FPGAs (Field-Programmable Gate Arrays) or ASICs (Application-Specific Integrated Circuits)), and the device itself is also implemented as one of these dedicated logic circuits.
[0100] A computer suitable for running computer programs can be based on a general-purpose or dedicated microprocessor, or both, or any other type of central processing unit. Generally, the central processing unit receives instructions and data from read-only memory or random-access memory, or both. Essential elements of a computer are the central processing unit for executing or running instructions and one or more memory devices for storing instructions and data. Generally, a computer also includes one or more mass storage devices (e.g., magnetic disks, magneto-optical disks, or optical disks) for storing data, or is operablely connected to receive, transfer, or both data to and from mass storage devices, although a computer does not have to have such devices. Furthermore, a computer may be embedded in another device (for example, to name just a few, a mobile phone, a digital information terminal (PDA), a mobile voice or video player, a game console, a Global Positioning System (GPS) receiver, or a portable storage device (e.g., a Universal Serial Bus (USB) flash drive).
[0101] Computer-readable media suitable for storing computer program instructions and data include, for example, all forms of non-volatile memory on media and memory devices, including semiconductor memory devices (e.g., EPROM, EEPROM, and flash memory devices), magnetic disks (e.g., internal hard disks or removable disks), magneto-optical disks, and CD-ROM and DVD-ROM disks. Processors and memory are complemented by or incorporated into dedicated logic circuits.
[0102] To provide user interaction, the embodiments of the subject matter described herein are implemented on a computer having a display device for displaying information to the user (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) and a keyboard and pointing device (e.g., a mouse or trackball) that allows the user to provide input to the computer. Other types of devices may be used similarly to provide user interaction; for example, the feedback provided to the user may be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback), and input from the user may be received in any form, including acoustic, voice, or tactile input. In addition, the computer may interact with the user by sending documents to and receiving documents from a device used by the user, for example, by sending a web page to a web browser on the user's device in response to a request received from that web browser.
[0103] Embodiments of the subject matter described herein are implemented in a single computer system that includes a backend component (e.g., as a data server), a middleware component (e.g., an application server), or a frontend component (e.g., a client computer having a graphical user interface or a web browser that allows a user to interact with one embodiment of the subject matter described herein), or any combination of one or more such backend, middleware, or frontend components. The components of this system are digital They are interconnected by any form or medium of data communication (e.g., communication networks). Examples of communication networks include local area networks (LANs) and wide area networks (WANs) (e.g., the Internet).
[0104] A computing system may include a client and a server. The client and server are generally geographically separated and typically interact via a communication network. The client-server relationship arises from computer programs running on each computer and having a client-server relationship with each other. In some embodiments, the server transmits data (e.g., an HTML page) to a user device (e.g., to display data to a user interacting with the user device and to receive user input from that user), and this user device functions as a client. Data generated on the user device (e.g., the results of user interactions) is received from the user device by the server.
[0105] This specification includes details of many specific embodiments, which should be interpreted not as limitations on the scope of any invention or the scope of claims, but rather as descriptions of functions that may be specific to a particular embodiment of a particular invention. Some functions described herein in the context of separate embodiments may also be performed when combined as a single embodiment. Conversely, some functions described in the context of a single embodiment may also be performed separately as multiple embodiments or as any appropriate subcombination. Furthermore, while some functions are described above as working in combination and are initially claimed as such, one or more functions in a claimed combination may, in some cases, be removed from the combination, and the claimed combination may be a subcombination or a variation of a subcombination.
[0106] Similarly, while the operations are depicted in a specific order in the drawings, this should not be understood as requiring that such operations be performed in a specific order or sequence as shown, or that all of the illustrated operations be performed, in order to obtain the desired results. In certain situations, multitasking and parallel processing may be advantageous. Furthermore, the separation of the various system modules and components in the embodiments described above should not be understood as requiring such separation in all embodiments, and the described program components and systems should generally be understood as being integrated as a single software product or packaged into multiple software products.
[0107] Specific embodiments of the subject matter have been described. Other embodiments are within the scope of the appended claims. For example, the operations enumerated in the claims may be performed in a different order to obtain the desired results. As one example, the processes depicted in the appended figures do not necessarily have to be in the specific order or sequence shown to obtain the desired results. In some cases, multitasking and parallel processing may be advantageous.
Claims
1. A method performed by a computer system: To obtain data showing the medical records of patients diagnosed with diabetes mellitus; Identifying the type of insulin the patient uses from medical records; The process involves selecting a first machine learning model from among multiple models within a machine learning system that is trained on a specific type of insulin, wherein the first machine learning model is trained using training data that includes medical records of multiple training patients and data showing the rate of corresponding hypoglycemic events for each training patient. Each of the multiple training patients used the same type of insulin identified in the patient's medical records. Each model in the machine learning system is trained on its own type of insulin from among several types; To determine the predictive rate of hypoglycemic events in patients by processing patients' medical records using a first machine learning model; Comparing the prediction rate to a predetermined threshold; In response to a determination that the prediction rate exceeds a predetermined threshold, a notification will be sent to the patient or physician. The method comprising the above.
2. The method according to claim 1, wherein determining the predictive rate of hypoglycemic events includes determining the frequency or severity of hypoglycemic events, or both, based on the patient's medical records.
3. Receiving multiple medical records from multiple patients; At least one covariate is shared across all multiple medical records. To determine the predictive rate of hypoglycemic events corresponding to at least one covariate by using a machine learning system on the medical records of multiple patients; To generate a report that identifies at least one covariate and the determined predictive rate corresponding to at least one covariate, The method according to claim 1, further comprising:
4. The method according to claim 3, wherein each patient in a group of patients uses the same type of insulin that the patient uses, and the predictive rate is determined by using a first machine learning model for each of a group of medical records.
5. The method according to claim 3, wherein a first patient among a group of patients uses a first type of insulin, a second patient among a group of patients uses a second type of insulin, and the predictive rate is determined at least in part by running a group of machine learning models, each trained on one of the first and second types of insulin, respectively.
6. The method according to claim 3, wherein the covariates include one or more of the following for each patient: demographics, socioeconomics, comorbidities, diabetic complications, diabetic disease status, drug use, sex, age group, insurance company, body mass index, blood pressure range, alcohol or drug use.
7. The method according to claim 1, wherein the notification indicates that the prediction rate exceeds a predetermined threshold.
8. A non-temporary computer that, when executed by one or more computers, stores one or more instructions that power one or more computers. It is a readable medium, To obtain data showing the medical records of patients diagnosed with diabetes mellitus; Identifying the type of insulin the patient uses from medical records; The process involves selecting a first machine learning model from among multiple models within a machine learning system that is trained on a specific type of insulin, wherein the first machine learning model is trained using training data that includes medical records of multiple training patients and data showing the rate of corresponding hypoglycemic events for each training patient. Each of the multiple training patients used the same type of insulin identified in the patient's medical records. Each model in the machine learning system is trained on its own type of insulin from among several types; To determine the predictive rate of hypoglycemic events in patients by processing patients' medical records using a first machine learning model; Comparing the prediction rate to a predetermined threshold; In response to a determination that the prediction rate exceeds a predetermined threshold, a notification will be sent to the patient or physician. The non-temporary computer-readable media, including the above.
9. A non-temporary computer-readable medium according to claim 8, wherein determining the predictive rate of hypoglycemic events includes determining the frequency or severity of hypoglycemic events, or both, based on the patient's medical records.
10. The order is, Receiving multiple medical records from multiple patients; At least one covariate is shared across all multiple medical records. To determine the predictive rate of hypoglycemic events corresponding to at least one covariate by using a machine learning system on the medical records of multiple patients; To generate a report that identifies at least one covariate and the determined predictive rate corresponding to at least one covariate, A non-temporary computer-readable medium according to claim 8, further comprising:
11. Each patient in a group of patients uses the same type of insulin they use, and the prediction rate is determined by using a first machine learning model for each of the multiple medical records. A non-temporary computer-readable medium as defined in claim 10.
12. A non-temporary computer-readable medium according to claim 10, wherein a first patient of a group of patients uses a first type of insulin, a second patient of a group of patients uses a second type of insulin, and the predictive rate is determined at least in part by running a group of machine learning models, each trained on one of the first and second types of insulin, respectively.
13. The non-temporary computer-readable medium according to claim 10, wherein at least one covariate includes one or more of the demographics, socioeconomics, comorbidities, diabetic complications, diabetic disease status, and drug use of each patient.
14. The notification is a non-temporary computer-readable medium according to claim 8, indicating that the prediction rate exceeds a predetermined threshold.
15. A system including one or more computers and one or more recording devices storing one or more instructions that, when executed by one or more computers, cause one or more computers to operate, To obtain data showing the medical records of patients diagnosed with diabetes mellitus; Identifying the type of insulin the patient uses from medical records; The process involves selecting a first machine learning model from among multiple models within a machine learning system that is trained on a specific type of insulin, wherein the first machine learning model is trained using training data that includes medical records of multiple training patients and data showing the rate of corresponding hypoglycemic events for each training patient. Each of the multiple training patients used the same type of insulin identified in the patient's medical records. Each model in the machine learning system is trained on its own type of insulin from among several types; To determine the predictive rate of hypoglycemic events in patients by processing patients' medical records using a first machine learning model; Comparing the prediction rate to a predetermined threshold; In response to a determination that the prediction rate exceeds a predetermined threshold, a notification will be sent to the patient or physician. The system including the above.
16. The system according to claim 15, wherein determining the predictive rate of hypoglycemic events includes determining the frequency or severity of hypoglycemic events, or both, based on the patient's medical records.
17. The order is, Receiving multiple medical records from multiple patients; At least one covariate is shared across all multiple medical records. To determine the predictive rate of hypoglycemic events corresponding to at least one covariate by using a machine learning system on the medical records of multiple patients; To generate a report that identifies at least one covariate and the determined predictive rate corresponding to at least one covariate, The system according to claim 15, further comprising:
18. The system according to claim 17, wherein each patient in a group of patients uses the same type of insulin that the patient uses, and the prediction rate is determined by using a first machine learning model for each of a group of medical records.
19. The system according to claim 17, wherein a first patient among a group of patients uses a first type of insulin, a second patient among a group of patients uses a second type of insulin, and the predictive rate is determined at least in part by running a group of machine learning models, each trained on one of the first and second types of insulin, respectively.
20. The system according to claim 15, wherein the notification indicates that the prediction rate exceeds a predetermined threshold.
21. A method performed by a computer system, This involves receiving data showing the medical records of multiple first patients, Each of the multiple first patients was diagnosed with diabetes mellitus and was using first-type insulin; Determining the predictive rate of each first hypoglycemic event for each first patient by processing medical records using a first machine learning model trained with first training data, which includes data showing the first medical records of multiple first training patients and the corresponding hypoglycemic event rates for each first training patient, wherein each of the first training patients uses a first type of insulin; Identifying one or more first covariates that correlate with the predictive rate of a first hypoglycemic event based on a first predictive rate in the medical records of multiple first patients; To generate a report showing the identified first covariate and the correlation between the identified first covariate and the predictive rate of the first hypoglycemic event, The method comprising the above.
22. In the medical records of multiple first patients, one or more second covariates corresponding to the predictive rate of a second hypoglycemic event, Further including identifying one or more second covariates distinct from one or more first covariates, The report further demonstrates the identified second covariate and the respective correlations between the identified second covariate and the predictive rate of the second hypoglycemic event. The method according to claim 21.
23. This involves obtaining data showing the medical records of multiple second patients diagnosed with diabetes mellitus, The use of a second type of insulin, which is different from the first type of insulin; Determining the predictive rate of each second hypoglycemic event for each second patient by processing the medical records of multiple second patients using a second machine learning model trained with second training data that includes data showing the second medical records of multiple second training patients and the corresponding hypoglycemic event rates for each second training patient, Each of the multiple second training patients uses a second type of insulin; Identifying one or more second covariates that correlate with the predictive rate of a second hypoglycemic event in the medical records of multiple second patients, based on the predictive rate of the second event, The report further demonstrates the identified second covariate and the correlation between the identified second covariate and the predictive rate of the second hypoglycemic event, The method according to claim 21, further comprising:
24. The method according to claim 21, wherein one or more first covariates are identified using a linear regression model.
25. The covariates are the demographics, socioeconomics, comorbidities, and diabetic complications of each first patient. The method according to claim 21, comprising one or more of the following: diabetic disease status, drug use, sex, age range, insurance company, body mass index, blood pressure range, alcohol or drug use.
26. The method according to claim 21, wherein the first prediction rate includes the respective severity of the hypoglycemic event predicted for each patient.
27. The method according to claim 26, wherein the first machine learning model is further trained to cluster the first patients based on the predictive value and severity of hypoglycemic events across different covariates.
28. A non-temporary computer-readable medium containing one or more instructions that, when executed by one or more computers, operate one or more computers, This involves receiving data showing the medical records of multiple first patients, Each of the multiple first patients was diagnosed with diabetes mellitus and was using first-type insulin; Determining the predictive rate of each first hypoglycemic event for each first patient by processing medical records using a first machine learning model trained with first training data, which includes data showing the first medical records of multiple first training patients and the corresponding hypoglycemic event rates for each first training patient, wherein each of the first training patients uses a first type of insulin; Identifying one or more first covariates that correlate with the predictive rate of a first hypoglycemic event based on a first predictive rate in the medical records of multiple first patients; To generate a report showing the identified first covariate and the correlation between the identified first covariate and the predictive rate of the first hypoglycemic event, The non-temporary computer-readable media, including the above.
29. The operation is, In the medical records of multiple first patients, one or more second covariates corresponding to the predictive rate of a second hypoglycemic event, Further including identifying one or more second covariates distinct from one or more first covariates, The report further demonstrates the identified second covariate and the respective correlations between the identified second covariate and the predictive rate of the second hypoglycemic event. A non-temporary computer-readable medium according to claim 28.
30. The operation is, This involves obtaining data showing the medical records of multiple second patients diagnosed with diabetes mellitus, The use of a second type of insulin, which is different from the first type of insulin; Determining the predictive rate of each second hypoglycemic event for each second patient by processing the medical records of multiple second patients using a second machine learning model trained with second training data that includes data showing the second medical records of multiple second training patients and the corresponding hypoglycemic event rates for each second training patient, Each of the multiple second training patients uses a second type of insulin; Identifying one or more second covariates that correlate with the predictive rate of a second hypoglycemic event in the medical records of multiple second patients, based on the predictive rate of the second event, The report further demonstrates the identified second covariate and the correlation between the identified second covariate and the predictive rate of the second hypoglycemic event, A non-temporary computer-readable medium according to claim 28, further comprising:
31. A non-temporary computer-readable medium according to claim 28, wherein one or more first covariates are identified using a linear regression model.
32. A non-temporary computer-readable medium according to claim 28, wherein the covariates include one or more of the demographics, socioeconomics, comorbidities, diabetic complications, diabetic disease status, drug use, sex, age range, insurance company, body mass index, blood pressure range, alcohol or drug use of each first patient.
33. The non-temporary computer-readable medium according to claim 28, wherein the first prediction rate includes the respective severity of hypoglycemic events predicted for each patient.
34. The non-transient computer-readable medium according to claim 33, wherein the first machine learning model is further trained to cluster the first patients based on the respective predictive rates and severity of hypoglycemic events across different covariates.
35. A system including one or more computers and one or more recording devices storing one or more instructions that, when executed by one or more computers, power one or more computers, This involves receiving data showing the medical records of multiple first patients, Each of the multiple first patients was diagnosed with diabetes mellitus and was using first-type insulin; Determining the predictive rate of each first hypoglycemic event for each first patient by processing medical records using a first machine learning model trained with first training data, which includes data showing the first medical records of multiple first training patients and the corresponding hypoglycemic event rates for each first training patient, wherein each of the first training patients uses a first type of insulin; Identifying one or more first covariates that correlate with the predictive rate of a first hypoglycemic event based on a first predictive rate in the medical records of multiple first patients; To generate a report showing the identified first covariate and the correlation between the identified first covariate and the predictive rate of the first hypoglycemic event, The system including the above.
36. The operation is, In the medical records of multiple first patients, one or more second covariates corresponding to the predictive rate of a second hypoglycemic event, Further including identifying one or more second covariates distinct from one or more first covariates, The report further demonstrates the identified second covariate and the respective correlations between the identified second covariate and the predictive rate of the second hypoglycemic event. The system according to claim 35.
37. This involves obtaining data showing the medical records of multiple second patients diagnosed with diabetes mellitus, The use of a second type of insulin, which is different from the first type of insulin; Determining the predictive rate of each second hypoglycemic event for each second patient by processing the medical records of multiple second patients using a second machine learning model trained with second training data that includes data showing the second medical records of multiple second training patients and the corresponding hypoglycemic event rates for each second training patient, Each of the multiple second training patients uses a second type of insulin; Identifying one or more second covariates that correlate with the predictive rate of a second hypoglycemic event in the medical records of multiple second patients, based on the predictive rate of the second event, The report further demonstrates the identified second covariate and the correlation between the identified second covariate and the predictive rate of the second hypoglycemic event, The system according to claim 35, further comprising:
38. The system according to claim 35, wherein one or more first covariates are identified using a linear regression model.
39. The system according to claim 35, wherein the covariates include one or more of the demographics, socioeconomics, comorbidities, diabetic complications, diabetic disease status, drug use, sex, age range, insurance company, body mass index, blood pressure range, alcohol or drug use of each first patient.
40. The first prediction rate includes the predicted severity of hypoglycemic events for each patient. The system according to claim 35, wherein the first machine learning model is further trained to cluster the first patients based on the respective predictive rates and severity of hypoglycemic events across different covariates.