Systems and Methods for Machine Learning Approaches to the Management of Healthcare Populations
The system addresses the limitations of existing models by using machine learning to analyze comprehensive health data and care gaps, providing actionable recommendations for heart failure management, and achieving improved patient outcomes and resource allocation.
Patent Information
- Application Number
- JP2022528558
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Priority Date
- 2019-11-15
- Filing Date
- 2020-11-16
- Publication Date
- 2025-06-30
- Estimated Expiration
- 2040-11-16
AI Technical Summary
Current machine learning models for managing heart failure populations are limited by their small, systematically collected datasets and lack of generalizability to wide and heterogeneous populations, as well as their failure to provide clinically relevant actionable outcomes.
A system and method using machine learning techniques that incorporate comprehensive health information, including clinical variables, diagnostic test measurements, and evidence-based care gaps, to provide ranked lists of patients for targeted interventions, thereby improving resource allocation and patient outcomes.
The proposed system achieves high accuracy in predicting 1-year all-cause mortality in heart failure patients and efficiently prioritizes patients for interventions, leading to improved clinical outcomes and resource allocation within value-based care models.
Smart Images

Figure 0007700114000009 
Figure 0007700114000010 
Figure 0007700114000011
Abstract
Description
Technical Field
[0001] Cross - reference to Related Applications This application claims the benefit of priority of U.S. Provisional Application No. 62 / 936,374, filed on November 15, 2019.
Background Art
[0002] The present invention relates to systems and methods for analyzing and managing a heart failure population. Heart failure (HF) has a lifetime prevalence of 1 in 3 in the United States, results in approximately 1 million hospitalizations per year in the United States, and is the cause of death for 1 in 8 people. The estimated annual cost of HF in the United States is $30.7 billion, an amount that is expected to more than double to $69.7 billion by 2030, with an average annual burden of $244 per U.S. citizen.
[0003] In response to these rising costs, new models of healthcare and reimbursement are being developed. In these "value - based care" models, the management of many chronic conditions, such as heart failure, is expanding beyond the encounter between an individual patient and a physician and instead treating the disease at the population level. The overall goal of such models is to reduce / suppress costs by improving patient outcomes, maintaining efficient management of patients, and providing care that reduces the frequency of high - cost / high - acuity encounters. To optimize this type of management at the population level, it is necessary to identify and stratify patients in need of intervention and, ideally, to identify effective means for identifying and deploying appropriate interventions. Currently, there is a severe lack of data - driven models that have been recognized as effective for supporting these population health goals.
[0004] Data science techniques, including machine learning, are well-suited to assist with these tasks. For example, one of the first papers on this topic in 1995 showed that a neural network could use echocardiogram data to predict the one-year mortality rate in 95 heart failure patients with greater accuracy than linear models or clinical judgment. Since then, numerous additional studies involving thousands of patients have placed high expectations on machine learning for predicting hospitalizations, readmissions, or deaths in patients with heart failure.
[0005] Previously published models that use machine learning for risk prediction in heart failure patients have two major limitations regarding their usefulness in optimizing clinical population health management. First, most models use small, systematically collected and annotated datasets (such as from clinical trials) or focus on important but narrow clinical settings (such as in-hospital mortality during hospitalization for heart failure due to acute decompensation). Such approaches are effective and appropriate within their respective constraints but are not necessarily generalizable to the wide and heterogeneous heart failure populations characterized in "real world" clinical data. The second limitation is that the published findings using machine learning models do not lead to clinically relevant actionable outcomes.
[0006] Therefore, there is a need for a system that provides clinically relevant actionable treatment recommendations for patients who should receive evidence-based care but are not, and that can be generalized to wide and heterogeneous heart failure populations. SUMMARY OF THE INVENTION PROBLEMS TO BE SOLVED BY THE INVENTION
[0007] The present disclosure includes systems and methods for machine learning techniques for the management of heart failure populations. More specifically, the present disclosure provides systems and methods for providing clinically relevant actionable treatment recommendations for patients who should receive evidence-based care but have not received it, which can be generalized to a wide and heterogeneous heart failure population. The present disclosure provides systems and methods for generating a ranked list of patients ordered by the highest estimated benefit of providing other resources such as additional treatments and / or medications to more efficiently provide resources to patients. **Means for Solving the Problems**
[0008] Some embodiments of the present disclosure provide a method for providing treatment recommendations for patients to a physician. The method includes receiving health information related to a patient, determining a first risk score of the patient based on the health information using a trained predictor model, determining a second risk score of the patient based on the health information and at least one artificially reduced care gap included in the health information using the predictor model, determining a predicted risk reduction score based on the first risk score and the second risk score, determining a patient classification based on the predicted risk reduction score, and outputting a report based on at least one of the first risk score, the second risk score, or the predicted risk reduction score.
[0009] Next, to achieve the foregoing and related objects, the present invention includes the features fully described hereinafter. The following description and drawings show specific exemplary aspects of the present invention in detail. However, these aspects merely show some of the various ways in which the principles of the present invention can be employed. Other aspects, advantages, and novel features of the present invention will become apparent from the following detailed description of the present invention when considered in conjunction with the drawings. **Brief Description of the Drawings**
[0010]
Figure 1
Figure 2
Figure 3A
Figure 3B
Figure 4
Figure 5A
Figure 5B
Figure 6A
Figure 6B
Figure 7
Figure 8
Figure 9
[0011] While the invention is susceptible to various modifications and alternative forms, specific embodiments thereof are shown by way of example in the drawings and will herein be described in detail. It should be understood, however, that the description herein of specific embodiments is not intended to limit the invention to the particular forms disclosed, but on the contrary, the intention is to cover all modifications, equivalents, and alternatives falling within the spirit and scope of the invention as defined by the appended claims.
[0012] Next, various aspects of the invention will be described with reference to the accompanying drawings. It should be understood, however, that the drawings and the following detailed description thereof are not intended to limit the claimed subject matter to the particular forms disclosed. On the contrary, the intention is to cover all modifications, equivalents, and alternatives falling within the spirit and scope of the claimed subject matter.
[0013] As used herein, terms such as "component" and "system" are intended to refer to a computer-related entity, either hardware, a combination of hardware and software, software, or software in execution. For example, a component can be, but is not limited to, a process running on a processor, a processor, an object, an executable, a thread of execution, a program, and / or a computer. By way of illustration, both an application running on a computer and the computer can be a component. One or more components can reside within an execution process and / or thread of execution, and a component can be located on one computer and / or distributed between two or more computers or processors.
[0014] The term "exemplary" as used herein means to serve as an example, an instance, or an illustration. Any aspect or design described herein as "exemplary" should not necessarily be construed as being more preferred or advantageous than other aspects or designs.
[0015] Hereinafter in this specification, unless otherwise indicated, the following terms and expressions will be used in this disclosure as described. The term "provider" is used to denote the entity that operates the entire system disclosed herein and will most often include a company or other entity that runs a server and maintains a database, and these companies or entities employ people with many different skill sets necessary to build, maintain, and adapt the disclosed system to accommodate new data types, new medical and treatment findings, and other needs. Exemplary provider personnel can include researchers, clinical trial designers, oncologists, neurologists, psychiatrists, data scientists, and many other personnel with specialized skill sets.
[0016] The term "physician" is used to generally denote any healthcare provider including, among others, primary care physicians, medical specialists, oncologists, neurologists, nurses, and medical assistants, but is not limited thereto.
[0017] The term "researcher" is used to generally denote any person who conducts research including radiologists, data scientists, or other healthcare providers, but is not limited thereto. A person may be both a physician and a researcher, and another person may work only with one of those qualifications.
[0018] Furthermore, the disclosed subject matter can be implemented as a system, method, apparatus, or article of manufacture that uses programming and / or engineering techniques to create software, firmware, hardware, or any combination thereof for controlling a computer or processor-based device to implement the aspects detailed herein. As used herein, the term "article of manufacture" (or alternatively, "computer program product") is intended to encompass a computer program accessible from any computer-readable device, carrier, or medium. For example, the computer-readable medium can include, but is not limited to, magnetic storage devices (e.g., hard disks, floppy disks, magnetic strips, etc.), optical disks (e.g., compact disc (CD), digital versatile disc (DVD)), smart cards, and flash memory devices (e.g., cards, sticks, etc.). Additionally, it should be recognized that a carrier wave can be employed to carry computer-readable electronic data such as that used when sending and receiving an e-mail or accessing a network such as the Internet or a local area network (LAN). Transitory computer-readable media (such as carrier waves and signal-based) should be considered separate from non-transitory computer-readable media such as those described above. Of course, those skilled in the art will recognize that many modifications can be made to this configuration without departing from the scope or spirit of the claimed subject matter.
[0019] In the present disclosure, ARB indicates an angiotensin II receptor blocker, ACEI indicates an angiotensin-converting enzyme inhibitor, ARA indicates an aldosterone receptor antagonist, ARNI indicates an angiotensin receptor neprilysin inhibitor, AUC indicates the area under the receiver operating characteristic curve, EBBB indicates an evidence-based beta blocker, ECG indicates an electrocardiogram, and EHR indicates an electronic health record.
[0020] The inventors utilized a large-scale 20-year retrospective dataset derived from a health system (Geisinger) adopted in the early days of electronic health record (EHR) technology to develop a predictive model for all patients with heart failure using machine learning. This model included a comprehensive set of input variables, including six "care gap" metrics. A "care gap" is defined as the difference between the recommended best care and the care actually provided.
[0021] This novel incorporation of evidence-based care gaps into a predictive model represents a methodology for driving clinical behavior from machine learning models (predicting not only risk but also a reduction in risk or "benefit" as a result of the behavior). Moreover, through population health management efforts, it is demonstrated how such findings can be utilized to simultaneously stratify risk and treatment benefit at the individual patient level to efficiently deploy healthcare resources.
[0022] Methods EHR Data Collection Patients with heart failure were identified from Geisinger's EHR over a 19-year period (January 2001 - February 2019). Heart failure was defined using the validated eMERGE phenotype. All clinical encounters from 6 months prior to the heart failure diagnosis date, including outpatient clinic visits, hospitalizations, emergency department visits, laboratory tests, and cardiac diagnostic tests (e.g., echocardiogram or electrocardiogram), were identified as independent samples.
[0023] Model Input Figure 1 is a flow for training a model to predict 1-year all-cause mortality using EHR data, and to predict mortality risk when simulating care gap reduction / treatment by artificially narrowing the care gap and when not doing so. A machine learning model was used to integrate clinical variables, measurements from diagnostic tests (e.g., echocardiography and electrocardiography), and care gap variables based on evidence from electronic health records to investigate 1-year all-cause mortality in a large cohort of heart failure patients. The average area under the ROC curve (AUC) from the "split-by-year" training scheme was reported to evaluate model performance. Then, using the model with the best performance, the risk reduction (potential benefit) by artificially narrowing the care gap in a prospective prediction set was estimated, and the efficiency of benefit-driven patient prioritization was evaluated. Figure 1 identifies various exemplary machine learning models that can be used as part of this method, including logistic regression ("LR"), random forest ("RF"), and XGBoost. In Figure 1, the abbreviation BP indicates blood pressure.
[0024] A total of 80 variables were collected from the EHR, namely, 8 clinical variables (age, gender, height, weight, smoking status, heart rate, systolic blood pressure, and diastolic blood pressure), loop diuretic use, 12 biomarkers (hemoglobin, eGFR, CKMB, lymphocytes, HDL, LDL, uric acid, sodium, potassium, NT-proBNP, troponin T, A1c), 44 non-redundant echocardiogram variables, 9 ECG measurements (such as QRS duration), and 6 care gap variables (described later) (see Figure 1). Experimental values (Lab values), vital signs, and ECG measurements closest to the in-person visit within a 6-month window were extracted. All echocardiogram measurements recorded in the Xcelera database within 12 months of the in-person visit were extracted. If no measurement was available within the specified time window, the variable was set to missing. EHR data preprocessing / cleaning is described in more detail in the "EHR Data Preprocessing" section below. It is understood that these variables are just one of many possible sets of variables that can be used to train a similar model. Moreover, additional data types such as medical image data, medical signal data (e.g., electrocardiograms), and genomic data can be used as inputs to the model.
[0025] EHR Data Preprocessing Regarding the physiological limits of echocardiogram variables, they were defined with the assistance of cardiologists with expertise in echocardiography. Data cleaning included the removal of 1) redundant variables directly derived from other variables and 2) values outside the physiologically possible range defined by cardiologists that could be due to human error and include physiologically impossible values (e.g., LVEF < 0% or > 100%, height and weight < 0). The removed values were then set to missing.
[0026] Since the predictive model requires a complete dataset, missing data for continuous variables were imputed using two steps. First, if a complete value was found in an adjacent encounter, missing values between encounters for individual patients were linearly interpolated. Next, to ensure that sufficient samples were available for imputation of each measurement, measurements missing in over 90% of the samples were discarded, and the remaining missing values were imputed using robust, multiple imputation by chained equations (MICE).
[0027] Missing values for the diastolic function (represented as a categorical variable) were imputed by training a one-vs-all logistic regression classifier from all samples where the diastolic function was available. The diastolic function was reported as an ordinal variable based on the level of abnormality, with -1 for normal, 0 for abnormal (without grade reporting), and 1, 2, and 3 for grade I, II, and III diastolic dysfunction, respectively.
[0028] Care gap variable Six actionable interventions (care gap variables) based on evidence, namely, 1) influenza vaccination, 2) target hemoglobin A1c (<8%), 3) target BP (blood pressure <140 / 90 mmHg), 4) evidence-based beta-blockers (EBBB), 5) active angiotensin-converting enzyme inhibitors (ACEI), angiotensin II receptor blockers (ARB), or angiotensin receptor neprilysin inhibitors (ARNI), and 6) active aldosterone receptor antagonists (ARA) were introduced into a machine learning model to investigate their association with patient outcomes. These care gap variables were defined with the assistance of cardiologists and pharmacists with expertise in heart failure. The detailed inclusion / exclusion criteria are listed in Table 1 below. The blinded chart review validation of each care gap variable is described in detail in the section "Care Gap Validation" below. It is understood that there are other treatments or interventions that can be input into the model, such as medication, clinic visits, provider visits to the patient's home, etc., in addition to the listed care gap variables. Note that for new therapies or medications for which outcomes have not yet been obtained in a large retrospective clinical dataset, data showing the effect of the therapy on the specific outcome of interest can be used until sufficient data are incorporated to generate a new model to facilitate the most accurate machine learning model training.
Table 1
[0029] Care Gap Validation To verify the accuracy of the defined care gap variables, two reviewers independently reviewed 50 - 100 charts for each care gap variable manually in a blinded manner. Specifically, the questionnaire was created in REDCap for each care gap variable such that the questions covered patient demographics (e.g., whether the patient has heart failure), gap expansion / narrowing situations (e.g., whether the patient's most recent A1c < 8%), and exclusion criteria (e.g., whether the patient is allergic to the influenza vaccine). While balancing positive and negative cases for each criterion, 50 - 100 cases were randomly selected from the inventors' database for each care gap. For example, in the case of the influenza vaccine, 25 cases with an expanded gap (not vaccinated against the influenza vaccine) were expected, and 25 cases with a narrowed gap (vaccinated against the influenza vaccine) were expected. The number of cases was determined based on how many criteria / questions were included for each care gap. It should be noted that since medication - related gaps that are rare in the EHR contain many exclusion criteria, the balance of cases based on exclusion criteria was not taken, and only that representative cases were included. For the selected cases, the reviewers were provided with the patient's medical record number (MRN, unique identifier) and the date of encounter. The reviewers then filled out the questionnaire by reviewing the patient's chart in EPIC using the provided date of encounter as the reference date. This was used as the ground truth and compared with the care gap data calculated by the inventors. An overview of the review results is presented in Table 2 below. In Table 2, N / A means there are no inclusion / exclusion restrictions for the gap.
Table 2
[0030] Primary outcome Using a machine learning model, the all-cause mortality one year after the relevant face-to-face date was predicted. The survival period was calculated from the date of death (cross-referenced with the national death index database in monthly units) or from the last in-person encounter that could be confirmed from the EHR. This is an example of a single clinically relevant endpoint, but it is understood that additional endpoints include, but are not limited to, hospitalizations, visits to the emergency department or clinic, total cost of care, adverse outcomes such as stroke or heart attack, etc.
[0031] Training and evaluation of the machine learning model First, a linear logistic regression classifier was used for its simplicity (specifically, to examine the directionality of the relationship between model inputs and primary outcomes), and then compared with the performance of non-linear models including random forest and XGBoost (a scalable gradient tree boosting system). Assuming these non-linear models, the predictive accuracy was improved by capturing more complex non-linear relationships between input variables. The model with the best performance was selected for subsequent analysis of the care gap reduction effect. The model was evaluated using the "annual" form of cross-validation described in the "Machine learning model evaluation" section below.
[0032] Machine learning model evaluation As the forward prediction dataset (clinically "actionable" dataset), the most recent face-to-face encounters in all surviving patients with heart failure (as of February 9, 2019) were excluded. All remaining samples (face-to-face encounters) for which the outcome status was known were used for model evaluation.
[0033] The random split method misrepresents the "real world" scenario, so to evaluate the proposed model, the inventors deviated from the conventional cross - validation method. Instead, after the "year - by - year" procedure, the samples were split into a training set (past) and a test set (future). To deploy the model, the model was trained on all available data prior to the current date and applied to the most recent encounters of the patients, so that the model can be retrospectively evaluated as if it was deployed at a given date. For each year (e.g., 2010), the cut - off date was set to January 1st of that year (2010 / 1 / 1), so that all encounters before the cut - off date were used for training, and the first encounter of a given patient after the cut - off (but within the calendar year of 2010 / 1 / 1 - 2010 / 12 / 31) was used for testing. This method was repeated between 2010 and 2018.
[0034] The area under the receiver operating characteristic (ROC) curve (AUC) was obtained from the test set, and the overall model performance was reported as the mean AUC and standard deviation across all training years. The mean importance and ranking of each individual variable across all training years were obtained to identify the most important variables. Open - source Python packages "scikit - learn" (version 0.20.0) and "XGBoost" (version 0.80) were used to implement the machine - learning pipeline and evaluate the model.
[0035] After the training stage, the optimal set of hyperparameters was obtained and used to retrain the entire dataset to obtain the final model. Then, the final model was used on the submitted actionable prediction dataset (the most recent encounters from all surviving patients as of 2019 / 2 / 9) to obtain the likelihood score for each individual patient. This likelihood score, called the risk score, was scaled to the range of 0 - 1, where higher values correspond to a higher mortality risk.
[0036] During training, risk scores were obtained for each individual sample in the test set. These risk scores were binned into 20 groups increasing by 0.05 from 0 to 1, and the true mortality rate was calculated using ground truth from the samples within that group. The average event rate over all training years for a particular bin was used to estimate the event rate as a function of the computer-calculated risk scores in the prediction set. This enabled the mapping of risk scores to the mortality event rate.
[0037] Benefit prediction for surviving patients by simulation of care gap reduction To investigate the effect of care gap reduction on improving patient outcomes, the care gap was artificially reduced (i.e., the value was changed from 1 = widened / untreated to 0 = narrowed / treated) while keeping all other variables unchanged. The care gap did not narrow in patients who met the exclusion criteria for that care gap (e.g., patients with bradycardia who could not be treated by EBBB). First, logistic regression was used to estimate the associated directionality of each care gap variable with the predicted mortality risk (e.g., receiving an influenza vaccine is associated with a decrease in mortality risk). Care gaps with a negative or undetermined relationship with the outcome (i.e., the target BP described below) were not narrowed at all. For care gaps with a positive relationship with the outcome, care gap reduction was simulated using a non-linear model that maximizes performance by artificially narrowing the gap and recalculating the risk scores using the same model.
[0038] After the simulation, the change in the risk score, i.e., the difference between the baseline risk score with a widening care gap and the risk score with a narrowing care gap, was calculated for each patient and further converted into an estimated benefit, i.e., a reduction in the estimated mortality rate. Then, the cumulative sum of the benefits from all patients was used to generate the estimated number of lives that could be saved by narrowing the care gap. In some embodiments, the risk score with a widening care gap and / or the risk score with a narrowing care gap can be provided and used by a physician and / or provider to estimate the death risk of a particular patient. In this way, the physician and / or provider can estimate whether a patient is likely to die within that year (or within another time period), thereby enabling appropriate resources, such as palliative care physicians, to be provided to the patient at the appropriate time.
[0039] Results Study population A total of 24,740 patients (median age 76 years, 45% female) with heart failure who had 945,404 face-to-face encounters in the EHR that met the inclusion criteria were identified. The face-to-face encounters are used as predictive inputs into the model in this scenario, but it should be noted that the predictive inputs can be differently configured, for example, by using "episodes" where multiple face-to-face encounters are grouped or combined in another form at one point in time for which the prediction is made. Tables 3 and 4 below show the summary statistics. On average, each patient had 38 face-to-face encounters (interquartile range (IQR): 10 - 49). The median follow-up period was 3.4 years (IQR: 1.4 - 6.3 years) using the reverse Kaplan-Meier method, and 12,594 patients (51%) had a recorded death. Data are reported as median [interquartile range] or as a percentage. [Table 3] [Table 4A]
Table 4B
Table 4C
[0040] Of the 12,146 patients alive as of February 2, 2019, at their most recent in-person visit, 9,474 (78%) had at least one care gap that had widened, and 501 (4%) had four or more care gaps that had widened. Figure 2 is a graph of the number of patients with each gap that is widened / untreated or narrowed / treated. The sum of those groups represents the number of patients eligible for the gap (i.e., who met the eligibility criteria). Depending on the gap, 20 - 74% of eligible patients had a widened gap. Additional details are available in Table 5 below. In Table 5, the percentage of non-eligible and widened percentages are calculated based on the number of in-person visits included (i.e., in-person visits while the patient was eligible and thus met the eligibility criteria for taking the medicine). In Figure 2, EBBB (also described in Figure 1) is the abbreviation for evidence-based beta blocker, ACEI is the abbreviation for angiotensin-converting enzyme inhibitor, ARB is the abbreviation for angiotensin II receptor blocker, ARNI is the abbreviation for angiotensin receptor neprilysin inhibitor, and ARA is the abbreviation for aldosterone receptor antagonist.
Table 5
[0041] Prediction accuracy of all-cause mortality using machine learning All three machine learning models predicted all-cause mortality at one year with an AUC exceeding 0.70, and the non-linear models achieved higher average AUCs (Random Forest: 0.76±0.02, XGBoost: 0.77±0.03) compared to linear logistic regression (0.73±0.02, Figure 3). Figure 3A is a graph of the average AUCs of the linear and non-linear models. Both non-linear models demonstrated superior performance in predicting all-cause mortality at one year compared to linear logistic regression (LR), and XGBoost (XGB) had the highest average AUC. Figure 3B is a graph of the area under the curve from 2010 to 2018 for the linear and non-linear models.
[0042] Figure 4 is a graph of the top 20 variable rankings using XGBoost. In addition to commonly used clinical variables (age, weight) and biomarkers (HDL, LDL), echocardiogram variables were found to be very important in predicting all-cause mortality at one year in patients with heart failure. See Table 4 above for variable explanations. The variable importance ranking using XGBoost demonstrated that 15 out of the top 20 variables were echocardiogram measurements. The results of logistic regression demonstrated that 5 out of 6 care gap variables (all except the target BP) had a positive association such that an expanded gap was expected to be associated with a higher risk of all-cause mortality at one year. Using only these 5 variables, the effect of reducing the care gap in subsequent models was predicted.
[0043] Prediction of the benefit of reducing the care gap Since the XGBoost model has the highest AUC in the retrospective data, XGBoost was selected as the final model to predict the benefit of reducing the care gap for surviving patients. The distribution of the risk scores is shown in Figures 5A - 5B. Figure 5A is a graph of the average mortality rate corresponding to each risk score bin derived from the training data over all training years. Figure 5B is a graph of the distribution of the predicted risk scores in the prediction set (surviving patients), which is then transformed into the predicted mortality rate using the relationship shown in Figure 5A. The number of face - to - face encounters included in each training / validation fold for each year is included in Table 6 below. Based on the estimated mortality rate, it was predicted that 2,662 out of 12,146 surviving patients (21.9%) would die within one year. The decrease in the validation set in 2018 is due to the insufficient follow - up period (< 1 year) for the surviving patients at the time of data collection (February 9, 2019).
Table 6
[0044] As a result of artificially reducing five care gaps that are positively correlated with the mortality rate, it was predicted that 2,495 (20.5%) patients would die within one year. This resulted in a predicted absolute risk reduction of 1.4% in the mortality rate (range: 0 - 31%, absolute). Assuming that all five care gaps can be reduced, it is estimated that an additional 167 patients (6.3% of 2,662) would be able to survive beyond one year.
[0045] The relationship between risk and benefit (risk reduction) was further investigated by comparing the predicted benefits among several subgroups. Figure 6A is a scatter plot of risk scores and corresponding benefits for individual patients in the prediction set (N = 12,146). The negative reduction in mortality reflects a harmful effect on mortality risk that narrows the care gap, as predicted by the non-linear XGBoost model for a small percentage of patients. Figure 6B is a graph of the average mortality before and after the care-shortening simulation in the selected group. Note that the risk is not equivalent to the benefit, as the predicted benefits of narrowing the care gap are not the same for patients at the same similarly high mortality risk level.
[0046] Figure 6B shows that the overall average benefit was predicted to be relatively small, with a low mortality risk at baseline (risk score < 0.2) and a low benefit after narrowing the care gap (reduction in mortality < 5%) ( "Low Risk, Low Benefit"), mainly driven by a large group of patients. However, there was a subgroup of patients predicted to have a high risk of mortality (risk score > 0.5) and also predicted to have a high benefit after narrowing the gap (reduction in mortality > 10%) ( "High Risk, High Benefit"). However, not all high-risk patients were predicted to have a high benefit, as evidenced by another subgroup of patients with a similarly high risk at baseline but only a small reduction in risk after narrowing the care gap ( "High Risk, Low Benefit").
[0047] Prioritization of patients for efficient narrowing of the care gap through population health management To illustrate the potential value of machine learning in optimizing the allocation of care team resources in this setting, we plotted the number of lives saved against the number of patients receiving an intervention (assuming that all eligible gaps were subsequently closed) for several different prioritization strategies. If we assume that we could convene and deploy a population health management team to narrow the care gap, the effectiveness of this effort would depend on effective guidance on which patients to target first in a ranked order. Strategy 1: Random prioritization Strategy 2: Randomly prioritize any patient with at least one wide care gap Strategy 3: Rank patients by the number of wide care gaps Strategy 4: Stratify patients using the Seattle Heart Failure risk score Strategy 5: Stratify patients according to the predicted "benefit" (i.e., reduction in mortality risk) of the XGBoost model
[0048] Figure 7 is a graph of the number of lives saved estimated by various stratification techniques during a simulation of care gap reduction using XGBoost. Prioritizing patients according to predicted benefit is the most efficient method of resource allocation based on the highest predicted patient survival rate (y-axis) relative to the number of patients requiring treatment (x-axis). Note that the slope of the plotted lines is inversely proportional to the number requiring treatment, and thus the steeper the line, the more efficient the patient prioritization. The slight decrease in the number of lives saved at the right end of the line corresponding to the "Benefit Driven" model reflects patients for whom closing the care gap had a predicted negative impact on mortality risk, as shown in Figure 6A.
[0049] Figure 7 demonstrates that the proposed machine learning benefit stratification model (Strategy 5) was the most efficient. That is, benefit stratification had the steepest slope among all the prioritization strategies and thus maximized the total number of predicted saved lives for a given number of patient interventions in a resource-constrained environment.
[0050] Discussion Optimized population health management requires new data-driven approaches for allocating healthcare resources, especially within new value-based care models. This investigation has made considerable progress towards the development of such an approach for heart failure by carefully combining curated clinical data with machine learning. The model incorporates important clinical variables, quantitative measurements from common diagnostic tests such as echocardiography and electrocardiogram, and interventions based on evidence in the form of "care gaps". As a result, it has been shown that machine learning models using these inputs can achieve excellent accuracy in predicting 1-year all-cause mortality in patients with heart failure. Furthermore, the explicit representation of clinical care gaps in the model represents a new paradigm for guiding clinical behavior using machine learning. Specifically, the present disclosure shows how these care gap inputs can be used to predict the risk reduction associated with specific interventions at the individual patient level.
[0051] These model predictions can provide guidance to integrated health systems that function to efficiently allocate scarce clinical resources (e.g., care teams) to the patients who need them most. Importantly, most published models and clinical scoring systems rely heavily on risk prediction, which can be used to prioritize the allocation of healthcare resources. However, risk is not equivalent to benefit, and thus patients with the same risk of 1-year mortality may have very different predicted benefits from an intervention. Therefore, simply deploying resources based on risk is likely to be inefficient, as demonstrated by the superiority of the predictive performance of predictive models over the Seattle Heart Failure score for prioritizing interventions for patients.
[0052] Comparison with Other Predictive Machine Learning Models in Heart Failure In recent years, several studies have been published using machine learning to predict outcomes (mostly survival) in patients with heart failure. These studies have used a variety of methods, from traditional classifications (e.g., logistic regression, random forest) to custom-developed algorithms (contrast pattern-assisted logistic regression with a probabilistic loss function), to predict mortality in heart failure. The reported accuracy (AUC) varies from 0.61 to 0.94, but most cluster around 0.75 to 0.8.
[0053] On the surface, the model performance is comparable to these previous investigations. However, there are some critical differences that should be noted when they reflect more difficult prediction tasks achieved by the predictive models presented by this disclosure. First, the model is designed for forward implementation in a "real-world" clinical setting when reflected in both the training / validation scheme and the prospective randomized clinical trial initiated using this model. Thus, the approach relied on clinical EHR data (as opposed to data collected during a controlled clinical trial) and tolerated its associated issues (e.g., incomplete and / or erroneous data). Second, most previous investigations have focused on stratification by maintained or decreased ejection fraction, or on specific subgroups of heart failure such as patients with acute decompensation, or on prediction in specific settings such as in-hospital or post-hospitalization mortality. The inventors' analysis has a broad focus on all patients with heart failure, taking into account both inpatient and outpatient encounters, and in this case also reflects the need for a population health management approach that is continuously updated.
[0054] Assuming this more difficult prediction task, it is notable that the model performance was consistent with previous investigations. This achievement was mainly driven by two attributes of the dataset. Above all, the sample size of the investigation was an order of magnitude larger (nearly one million encounters from 24,000 patients) compared to previous investigations (most of which were in the hundreds to thousands), thereby enabling a more generalizable model with a reduced likelihood of overfitting. In addition, this model included a comprehensive set of patient characteristics (input variables) that included data from diagnostic tests such as echocardiograms, which are very important for predicting all-cause mortality in the setting of heart failure and a more general cardiology population (note that some patients from previous investigations involving 171,510 patients were included in the current investigation). In contrast, most previous investigations were limited to basic clinical information (demographics, vital signs), results from laboratory tests, and comorbidities. There was only one investigation that reported an AUC of 0.72 for all-cause mortality, including additional diagnostic measurements from echocardiography and electrocardiograms, despite a small patient sample (n = 397), further underscoring the importance of these quantitative diagnostic data.
[0055] Another major weakness of most previous studies is the lack of actionable model results that can be clinically used. Therefore, although a number of accurate models have been developed over the past decade to predict the outcomes of patients with heart failure, few have truly affected clinical practice. In a recent study, an attempt was made to address this problem by evaluating the association between treatment (various medications) and outcomes among four subgroups of heart failure identified using unsupervised clustering in a retrospective dataset. The authors of the study showed significant differences in outcomes and different responses to medications among the four subgroups, which could potentially help define effective treatment strategies specific to each subgroup. Consistent with that study, this concept was taken a step further by introducing six evidence-based interventions (care gaps) into a machine learning model and using these variables as actionable "levers" within the model to predict the outcomes of individual patients after clinical actions. By artificially narrowing these care gaps, it is predicted that an additional 167 patients could potentially survive beyond one year.
[0056] Despite these interventions (care gaps) being recommended in national guidelines based on proven benefits (such as even influenza vaccination being associated with a reduction in all-cause mortality in heart failure), the prevalence of widespread care gaps remains a significant problem in the medical field. For example, in patients with heart failure, the utilization rate of life-prolonging therapies is surprisingly low, with only 57% receiving ACE inhibitors, 34% receiving evidence-based beta blockers, and only 32% receiving mineralocorticoid antagonists. This problem is very complex and is unlikely to be solved by relying on individual providers to change their practice. However, new, value-based care models are likely to be more effective in addressing this problem by creating organized care teams. These teams will require accurate and reliable data science, such as that presented in this disclosure, to effectively allocate resources.
[0057] The "BP in goal" care gap, contrary to guidelines based on evidence from observational studies showing that low blood pressure is associated with a reduced risk of adverse events in heart failure, surprisingly had a negative relationship with outcomes. However, the "blood pressure paradox" has also been noted in numerous studies where low blood pressure or significant changes (increases or decreases) in blood pressure have been associated with poor outcomes. In the current study, a linear logistic regression model demonstrated an inconsistent relationship between blood pressure and survival, i.e., a negative relationship in one training year and a positive relationship in other training years, although on average a slight negative relationship was demonstrated (data not shown). In the present disclosure, a machine learning model configured to predict 1-year all-cause mortality with excellent accuracy in a large cohort of patients with heart failure is presented. As a result of leveraging nearly one million encounters from over 24,000 patients, it is shown that these models can be used not only to stratify patients by risk but also to efficiently prioritize patients based on the predicted benefits of clinically relevant evidence-based interventions. This approach may prove useful in supporting heart failure population health management teams within a new, value-based payment model. It is also contemplated that models can be generated that are configured to predict all-cause mortality for time periods other than 1 year, including 6 months, 2 years, 3 years, 4 years, 5 years, or other appropriate time periods. Additionally, as described above, it is possible to train predictive machine learning models using additional clinically relevant endpoints.
[0058] Next, referring to FIG. 8, an exemplary method 100 is shown for predicting the all-cause mortality in patients suffering from heart failure over a predetermined time period (i.e., one year), as well as for providing treatment recommendations for patients to a physician. Method 100 predicts a patient's risk score based on a machine learning model trained with respect to clinical variables (e.g., demographics and labs), electrocardiogram measurements, echocardiogram measurements, and care gap variables based on the evidence described above. Method 100 can be employed in a population health analysis module relied upon by a care team including a physician to prioritize patients who should be receiving evidence-based care but are not.
[0059] At 102, method 100 can receive health information related to a patient. The health information can include at least a portion of the EHR related to the patient. The EHR can be stored in a provider's database. In some embodiments, the health information can include eight clinical variables (age, gender, height, weight, smoking status, heart rate, systolic blood pressure, and diastolic blood pressure), use of loop diuretics, twelve biomarkers (hemoglobin, eGFR, CKMB, lymphocytes, HDL, LDL, uric acid, sodium, potassium, NT-proBNP, troponin T, A1c), forty-four non-redundant echocardiogram variables, nine ECG measurements (such as QRS duration), and six care gap variables, including the aforementioned eighty variables. In some embodiments, the health information may not include a target BP. Method 100 can then proceed to 104.
[0060] In 104, method 100 can determine a first risk score for a patient based on health information using a trained predictor model. The trained predictor model can be a linear model such as linear logistic regression as described above, or a non-linear model such as random forest or XGBoost. The predictor model can be trained to predict a risk score for all-cause mortality for a predetermined time period, such as one year, although it is understood that the model can be trained to predict all-cause mortality for other time periods, namely, six months, two years, three years, four years, five years, or other appropriate time periods, or for other appropriate clinical endpoints. Method 100 can provide at least a portion of the health information to the model and receive the first risk score from the model. The first risk score can represent a baseline score for the patient's actual predicted mortality risk. Then, method 100 can proceed to 106.
[0061] In 106, method 100 can determine a second risk score for a patient using a predictor model based on health information and at least one artificially shrunk care gap included in the health information. Method 100 can artificially shrink an appropriate care gap by changing the value of each expanded care gap from 1 = expanded / untreated to 0 = shrunk / treated while leaving all other variables included in the health information unchanged. Method 100 may not shrink a particular care gap in a patient who meets the exclusion criteria for that care gap. For example, a patient suffering from bradycardia who may not be treatable by EBBB will not have the EBBB care gap shrunk. Then, method 100 can provide the model with the health information modified to shrink any appropriate care gap and receive the second risk score from the model. The second risk score can represent a simulated score corresponding to what the predicted mortality risk of the patient would be if all appropriate expanded care gaps were shrunk. For some patients, in 106, since the care gap is either already shrunk or cannot be shrunk for patients who meet the exclusion criteria for the specific care gap described above, the method may not be able to shrink any care gap, and in this case, the second risk score will be the same as the first risk score. Then, method 100 can proceed to 108.
[0062] In 108, method 100 can determine a predicted risk reduction score based on the first risk score and the second risk score. Method 100 can calculate the predicted risk reduction score by determining the difference between the first risk score and the second risk score. Then, method 100 can proceed to 110.
[0063] In 110, method 100 can determine a patient classification based on a predicted risk reduction score. Method 100 can determine a patient classification by comparing a patient's predicted risk reduction score to the predicted risk reduction scores of other groups of patients. The other groups of patients can include other patients treated by a provider. Method 100 can determine the rank of a patient's predicted risk reduction score compared to a group of patients (i.e., using strategy 5 described above). For example, method 100 can determine that a predicted risk reduction score of 0.3 is the 500th highest predicted risk reduction score out of 10,000 patients. Then, method 100 can proceed to 112.
[0064] In 112, method 100 can generate and output a report based on at least one of a first risk score, a second risk score, or a predicted risk reduction score. For example, the report can include the raw first risk score, the raw second risk score, and the raw predicted risk reduction score. The report can include the raw rank of the patient's predicted risk reduction score compared to a group of patients (e.g., the predicted risk reduction score is the 500th highest predicted risk reduction score out of 10,000 patients), or the percentile rank of the predicted risk reduction score (e.g., the predicted risk reduction score is at the 95th percentile of all the provider's patients). The report can include any suitable graph and / or chart generated based on the first risk score, the second risk score, and / or the predicted risk reduction score. The report can be displayed to a physician using a display such as a computer monitor or screen essential for a tablet computer, smartphone, laptop computer, etc. In some embodiments, the report can be output to a storage device including a memory. In some embodiments, the report can include the raw first risk score and the second raw risk score. The first risk score and the second risk score can be used by a physician and / or provider to estimate the risk of death of a patient. In this way, a physician and / or provider can estimate whether a patient is likely to die within that year (or other time period), thereby enabling appropriate resources such as a palliative care physician to be provided to the patient at the appropriate time.
[0065] Next, referring to FIG. 9, an exemplary system 210 for implementing the foregoing disclosure is shown. System 210 can include one or more computing devices 212a, 212b that communicate with each other, as well as with server 214 and one or more databases or other data repositories 216, via, for example, the Internet, an intranet, Ethernet, a LAN, a WAN, etc. The computing devices can also communicate with additional computing devices 212c, 212d through a separate network 218. Particular attention is paid to computing device 212a, although each computing device can include a processor 220, one or more computer-readable media drives 222, a network interface 224, and one or more I / O interfaces 226. Device 212a can also include a memory 228 containing instructions that configure the processor to execute an operating system 230 and a population health analysis module 232 for predicting one-year all-cause mortality in patients suffering from heart failure and providing treatment recommendations for the patients to physicians, as described herein. The population health analysis module 232 can be used to perform at least a portion of the method 100 described above with reference to FIG. 8.
[0066] The foregoing methodology for driving clinical action based on a predicted reduction in risk (i.e., benefit) can be applied to the management of any particular population (other than the heart failure population) in healthcare, including, but not limited to, diabetes, lung disease, kidney disease, rheumatic disease, musculoskeletal disease, endocrine disorders, etc. Further, this methodology can be extended to predict a reduction in risk for any particular clinical outcome of interest, including, but not limited to, additional adverse clinical events such as mortality, stroke, or heart attack, hospitalizations, total cost of care, or other healthcare utilization metrics.
[0067] Accordingly, as described herein, the present disclosure provides a system and method for providing clinically relevant actionable treatment recommendations for patients who should be receiving evidence-based care but are not, which can be generalized to a wide range of heterogeneous heart failure populations.
[0068] The present disclosure may admit of various modifications and alternative forms, but specific embodiments are shown by way of example in the drawings and are described in detail herein. It should be understood, however, that the present disclosure is not intended to be limited to the particular forms disclosed. Rather, the present disclosure is to cover all modifications, equivalents, and alternatives falling within the spirit and scope of the present disclosure as defined by the appended claims.
[0069] This specification discloses the present disclosure, including the best mode, using examples, and enables any person skilled in the art to practice the present disclosure, including making and using any device or system and performing any incorporated methods. The patentable scope of the present disclosure is defined by the claims and may include other examples that occur to those skilled in the art. Such other examples are intended to be within the scope of the claims if they have structural elements that do not differ from the literal language of the claims or if they include equivalent structural elements that do not differ substantially from the literal language of the claims.
[0070] Finally, it is expressly contemplated that any of the methods or steps described herein can be combined, excluded, or reordered. Accordingly, this specification should be construed as illustrative only and not in any way limiting the scope of the present disclosure.
Description of the Reference Numerals
[0071] 100 Method 210 System 212a, 212b, 212c, 212d Computing Device 214 Server 216 Database or other data repository 218 Network 220 Processor 222 Computer-readable media drive 224 Network interface 226 I / O interface 228 Memory 230 Operating system 232 Population health analysis module
Claims
**Claim 1** A method for providing treatment recommendations for a patient, comprising: receiving health information related to the patient; using a trained predictor model to determine a first risk score for the patient based on the health information; using the predictor model to determine a second risk score for the patient based on the health information and at least one artificially reduced care gap included in the health information; determining a predicted risk reduction score based on the first risk score and the second risk score; determining a patient classification based on the predicted risk reduction score, the patient classification including both a risk factor related to the first risk score and a benefit factor related to the second risk score; outputting a report based on at least one of the first risk score, the second risk score, or the predicted risk reduction score A method comprising the above steps. **Claim 2** The method according to claim 1, wherein the health information includes diagnostic tests. **Claim 3** The method according to claim 2, wherein the diagnostic test includes at least one of an echocardiogram or an electrocardiogram. **Claim 4** The method according to claim 1, further comprising removing redundant health information and removing physiologically impossible health information before determining the first risk score. A method according to claim 1, further including the above step. **Claim 5** The method according to claim 1, further comprising imputing missing health information using one or more of linear imputation from related health information or robust multiple imputation by chained equations before determining the first risk score. A method according to claim 1, further including the above step. **Claim 6** The method according to claim 1, further comprising discarding health information with at least one sample threshold number of missing values before determining the first risk score. A method according to claim 1, further including the above step. **Claim 7** The method according to claim 1, wherein the predictor model is a linear model. **Claim 8** The method according to claim 7, wherein the linear model is a linear logistic regression model. **Claim 9** The method according to claim 1, wherein the predictor model is a non-linear model. **Claim 10** The method according to claim 9, wherein the non-linear model is one of a random forest or XGBoost. **Claim 11** The method of claim 1, wherein the at least one artificially reduced care gap comprises a plurality of artificially reduced care gaps. **Claim 12** The method of claim 1, wherein the trained predictor model is selected from among a plurality of trained predictor models. **Claim 13** The method of claim 12, wherein the annual procedure applied to each of the plurality of trained predictor models is used to determine which model is the best model. **Claim 14** The method of claim 13, wherein the best model is retrained using an optimal set of hyperparameters. **Claim 15** The method of claim 1, wherein the step of determining a patient classification comprises comparing the predicted risk reduction score to predicted risk reduction scores of a plurality of other patients. **Claim 16** The method of claim 15, wherein the step of determining a patient classification further comprises ranking the patient relative to the plurality of other patients. **Claim 17** The method of claim 1, wherein the patient is part of a population of patients with heart failure. **Claim 18** The method of claim 1, wherein the patient is part of at least one population of patients with diabetes, lung disease, kidney disease, rheumatic disease, musculoskeletal disease, or endocrine disorder. **Claim 19** The method of claim 1, wherein the first risk score, the second risk score, and the predicted risk reduction score are related to the occurrence of an event within a predetermined time period. **Claim 20** The method of claim 19, wherein the event is the mortality of the patient. **Claim 21** The method of claim 19, wherein the predetermined time period is one year. **Claim 22** The method of claim 19, wherein the at least one artificially reduced care gap has a positive relationship to the occurrence of the event. **Claim 23** The method of claim 1, wherein the patient classification comprises evaluating the predicted risk reduction score relative to the number of patients in need of treatment. **Claim 24** The method of claim 1, wherein the report comprises treatment recommendations for the patient. **Claim 25** The method of claim 24, wherein the treatment recommendations comprise palliative care. **Claim 26** The step of allocating resources to the patient based on the patient classification further comprising the method of claim 1. A computer program for causing a computer to execute the method according to any one of claims 1 to 26.
Citation Information
Patent Citations
Self-evolving predictive models
JP2016519807A
System and method for predicting the vulnerability of coronary artery plaques from image data of a patient's inherent anatomical structure
JP2017503561A
Prediction of Cardiovascular Risk Events and Its Use
JP2017530356A
Data collection device and data collection method
JP2018175850A
Using data imputation to determine and rank of risks of health outcomes
US20110105852A1