Method and system for prediction of atherosclerotic cardiovascular disease (ASCVD)
A computer-implemented method using age and laboratory test results simplifies the prediction of cardiovascular risk, addressing the underutilization of current risk models by automating the process and providing accurate ASCVD risk assessments.
Patent Information
- Application Number
- PCT/CA2024/051497
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2023-11-17
- Filing Date
- 2024-11-13
- Publication Date
- 2025-05-22
AI Technical Summary
Current risk models for predicting atherosclerotic cardiovascular disease (ASCVD) are underutilized in clinical practice due to their complexity and the time-consuming process of collecting and inputting risk factor information.
A computer-implemented method and system that uses only age and predefined laboratory test results to predict cardiovascular risk, eliminating the need for clinical variables and physician input, and can be automated to generate risk estimates embedded within laboratory reports.
The method accurately predicts cardiovascular risk without clinical risk factors, simplifying risk estimation and reporting, and has been validated in an external primary care cohort with performance comparable to traditional models.
Smart Images

Figure IMGF000015_0001 
Figure IMGF000029_0001 
Figure IMGF000030_0001
Abstract
Description
METHOD AND SYSTEM FOR PREDICTION OF ATHEROSCLEROTIC CARDIOVASCULAR DISEASE (ASCVD)FIELD
[0001] The present disclosure relates methods and systems for predicting cardiovascular risk. BACKGROUND
[0002] Guidelines recommend using risk models that estimate the risk of atherosclerotic cardiovascular disease (ASCVD) to guide treatment decisions for preventative therapy (1). However, risk models are underutilized in routine clinical practice, primarily because physicians perceive them to be cumbersome and time consuming (2). Typically, risk models require physicians to collect risk factor information from medical history, physical measurements, as well as laboratory tests and then input these factors into risk calculators (3). Even when risk calculators are embedded within electronic health records, they are rarely automated, and thus do not seamlessly integrate into clinical workflow (4).
[0003] A model that utilizes only laboratory tests and does not rely on clinical variables or physician input may represent a significant advantage over traditional models. This is because a laboratory-based model could be automated to generate risk estimates that could be embedded within the laboratory report. This could potentially eliminate the need for any additional work by physicians and provide actionable decisions at the point of care. For instance, most laboratories now calculate and provide the estimated glomerular filtration rate (eGFR) using age, sex, and serum creatinine results directly on the report returned to the ordering clinician. The eGFR is an established marker of renal risk and routine reporting has led to improved detection of chronic kidney disease at the population level (5).SUMMARY
[0004] In one of its aspects, a computer-implemented method for predicting risk of a cardiovascular event, the method comprising: generating an input data set from cardiovascular event data, wherein the cardiovascular event data is acquired from a plurality of patients; generating a trained predictive model, based on patient age and at least one predefined predictor variable; receiving new subject patient data;with the trained predictive model, estimating the risk of the cardiovascular event associated with the subject patient data based on the age of the new subject patient and at least one predefined predictor variable associated with the new subject patient; and generating a report comprising a cardiovascular risk assessment.
[0005] In another aspect, a system for determining the cardiovascular health of a patient, the system comprising: a computer system comprising a hardware processor and a memory device on which instructions are encoded to cause the hardware processor to perform the operations of: generating an input data set from cardiovascular event data, wherein the cardiovascular event data is acquired from a plurality of patients; generating a trained predictive model, based on age and at least one predefined predictor variable; receiving new subject patient data; applying the trained predictive model on the new subject patient data to predict cardiovascular risk; and generating a report comprising a cardiovascular risk assessment.
[0006] In another aspect, a system for generating a trained predictive model for predicting a cardiovascular event, the system comprising: a computer system comprising a hardware processor and a memory device on which instructions are encoded to cause the hardware processor to perform the operations of: receiving patient data comprising at least one of demographic data, cardiovascular risk data, and laboratory data acquired from a plurality of patients; preprocessing the patient data; extracting features from the datasets; selecting features for subsequent use during a training phase, wherein the selected features comprise age and at least one predefined predictor variable for predicting the cardiovascular event;selecting a predictive model and iteratively training the predictive model with the training data sets and validating the predictive model with the testing data sets to generate a trained predictive model; and storing the model on the memory device and / or outputting the model for future use on new subject patient data.
[0007] In another aspect, a computer-readable medium comprising instructions stored thereon executable by a hardware processor to perform the operations of: receiving subject patient data comprising at least one of demographic data, cardiovascular risk data, and laboratory data; preprocessing the patient data; selecting features from the subject patient data, wherein the selected features comprise age and at least one predefined predictor variable for predicting the cardiovascular event; applying a trained predictive model trained using the selected features on the new subject patient data to predict a cardiovascular risk.
[0008] In another aspect, a method for using age and serum laboratory parameters to predict cardiovascular risk in a patient, whereby the cardiovascular risk assessment is automatically embedded within a clinical workflow, and the outcome is reported to a medical professional, a clinician or a third party.
[0009] Advantageously, models are generated and used to predict cardiovascular risk, such as atherosclerotic cardiovascular disease (ASCVD) to guide treatment decisions for preventative therapy. These models scan accurately predict cardiovascular risk without clinical risk factors, and utilize only demographics and laboratory tests simplify risk estimation and reporting. In one example, the method described herein leverages population-level data to derive sex-specific prediction models for ASCVD utilizing age as well as routine serum laboratory tests. These models are validated in an external primary care cohort and their performance is compared to that of the Pooled Cohort Equations (PCEs). Patients with high predicted risks may be candidates for statin therapy, while patients with low predicted risks can be treated with diet and lifestyle optimization alone.
[0010] Furthermore, there is provided a novel laboratory-based prediction model for heart disease and strokes. The models can accurately predict the onset of heart attack, stroke, ordeath from either cause using an individual’s age as well as parameters from routine serum blood work that is usually ordered as part of an annual assessment by the primary care provider (e.g. complete blood count (CBC), cholesterol panel, glucose). The models were developed in nearly 4 million Ontario residents and validated in a primary care cohort of 60,000 patients cared for primary care physicians in Ontario, Canada. The models’ accuracy was excellent and equivalent to models that are more complex with historical, medication, and physical measurement variables. The advantage of these models is that they lend themselves to automation. All the necessary parameters would be available to laboratory at the time of measurement, and hence risk can be estimated by the laboratory and embedded directly in reports sent to physicians who can in turn, focus their efforts on decision-making. These models may improve utilization of statins and can substantially transform care delivery for millions of individuals who are screened annually for preventative therapy. Routine reporting of ASCVD risk may improve the detection and management of treatment-eligible patients in primary prevention. As such, the models can predict the risk of a cardiovascular event, such as, an acute coronary syndrome, stable or unstable angina, arterial revascularization, stroke, transient ischemic attack, peripheral arterial disease, atherosclerotic cardiovascular disease, heart failure, atrial fibrillation, myocardial infarction, and sudden coronary death.BRIEF DESCRIPTION OF THE DRAWINGS
[0011] Figure 1 shows a top-level diagram of an overall system architecture for predicting cardiovascular risk in a patient;
[0012] Figure 2a shows a framework for derivation cohort creation;
[0013] Figure 2b shows a framework for validation cohort creation;
[0014] Figure 3 shows age specific percentiles of predicted risk in women, in which percentiles of predicted risk in the interval validation cohort are stratified by age; green cells correspond to low-risk (< 2.5%), yellow correspond to borderline (2.5 to < 3.75%) and red correspond to intermediate to high-risk (> 3.75%) of ASCVD at 5 years;
[0015] Figure 4 shows age specific percentiles of predicted risk in men, where percentiles of predicted risk in the interval validation cohort are stratified by age; green cells correspond to low-risk (<2.5%), yellow cells correspond to borderline-risk (2.5 to < 3.75%) and red cells correspond to intermediate to high-risk (>3.75%) of ASCVD at 5 years;
[0016] Figure 5 shows graphs associated with the proportion of risk factors by treatment categories, in which the proportion of individuals with risk factors stratified by 5-year predicted ASCVD risk using the CANHEART Lab Model is shown; sex-specific models include leukocytes as a predictor; risk categories correspond to low (< 2.5%), borderline (2.5% to < 3.75%) and intermediate to high-risk (>3.75%) adapted from the AHA / ACC guidelines; in which dyslipidemia is defined by an LDL greater than 3.5 mmol / L or non-HDL greater than 4.3 mmol / L; hypertension in the validation set is defined by a baseline systolic blood pressure > 140mmHg or treatment with anti -hypertensive therapy; and CKD is defined by an eGFR < 60mL / min / ;
[0017] Figure 6 shows calibration plots associated with the calibration of the CANHEART Models using differential white blood cell counts; the calibration plots depict the mean predicted and observed 5-year risks across deciles of predicted risk in the internal validation (blue squares) and external validation cohort (red squares) for women (solid squares) and men (open squares). Sex-specific models include neutrophil lymphocyte ratio and monocytes as a predictor; the dotted line represents perfect agreement between observed and predicted risks;
[0018] Figure 7 shows calibration plots associated with the calibration of the CANHEART Models without Complete Blood Count; the calibration plots depict the mean predicted and observed 5-year risks across deciles of predicted risk in the internal validation (blue diamonds) and external validation cohort (red diamonds) for women (solid diamonds) and men (open diamonds); sex-specific models do not include complete blood count predictors; the dotted line represents perfect agreement between observed and predicted risks;
[0019] Figure 8 shows decision curves of laboratory -based models, in which the plots depict the net benefit for laboratory-based models as well as strategies of treating all or no individuals with statins in the internal validation cohort; the 5 -year treatment probability threshold ranges from 1 to 5%, with the higher net benefit being favourable;
[0020] Figure 9 shows plots of the laboratory model in the entire development cohort (i.e. test and training sets, blue line) and the validation cohort (red line) are depicted; sex-specific models include leukocytes as a predictor; overlapping area of the kernel density function is estimated.
[0021] Figure 10 shows a flowchart outlining the steps for predicting cardiovascular risk in a patient, in one example; and
[0022] Figure 11 shows a block diagram of an example of a machine upon which any one or more of the techniques (e.g., methodologies) discussed herein can be performed.DETAILED DESCRIPTION
[0023] The following detailed description refers to the accompanying drawings. Wherever possible, the same reference numbers are used in the drawings and the following description to refer to the same or similar elements. While embodiments of the disclosure may be described, modifications, adaptations, and other implementations are possible. For example, substitutions, additions, or modifications may be made to the elements illustrated in the drawings, and the methods described herein may be modified by substituting, reordering, or adding stages to the disclosed methods. Accordingly, the following detailed description does not limit the disclosure. Instead, the proper scope of the disclosure is defined by the appended claims.
[0024] Moreover, it should be appreciated that the particular implementations shown and described herein are illustrative of the invention and are not intended to otherwise limit the scope of the invention in any way. Indeed, for the sake of brevity, certain sub -components of the individual operating components, and other functional aspects of the systems may not be described in detail herein. Furthermore, the connecting lines shown in the various figures contained herein are intended to represent exemplary functional relationships and / or physical couplings between the various elements. It should be noted that many alternative or additional functional relationships or physical connections may be present in a practical system.
[0025] Referring to Figure 1, there is shown a top-level diagram of an overall system architecture 10 for predicting cardiovascular risk in a patient. In one example, the system 10 comprises a machine 12 which receives cardiovascular data associated with plurality of patients from a plurality of data sources 14. The cardiovascular data pertains to events, such as strokes, heart disease, heart attacks and fatal cardiovascular events. The data sources 14 may include laboratories, testing facilities, and hospitals. The machine 12 may comprise a computing device comprising a processor, such as a central processing unit (CPU). In one example, the machine 12 implements a model, such as a predictive model, to estimate the cardiovascular risk on given subject patient data, and generates a report comprising the assessed cardiovascular risk. In one example, the report is output via a graphical user interface 20, or made available to other computing devices or systems, as will be described in more detail below.
[0026] The term computing device refers to data processing hardware and encompass all kinds of apparatus, devices, and machines for processing data, including by way of example, a programmable processor, a computer, or multiple processors or computers.
[0027] The machine 12 comprises an interface, as part of an I / O module for communicating with other systems via a communications network 26, in a distributed environment. Generally, the interface comprises logic encoded in software and / or hardware in a suitable combination and operable to communicate with the network 26. More specifically, the interface may comprise software supporting one or more communication protocols associated with communications.
[0028] Client terminals 30 (e.g., remotely located workstations) may request services related to cardiovascular risk assessment, and access the results over the communications network 26. Accordingly, the machine 12 may provide software as a service (SaaS) to the client terminals 30, or provide an application for local download to the client terminals 30, and / or provide functions using a remote access session to the client terminals 30, such as through a web browser. Accordingly, the assessments may be accessible for clinicians on an on-demand basis from the client terminals 30. System 10 may also comprise data storage 32, which is configured to maintain one or more datasets, including data structures storing linkages and other data, such as medical images, libraries, models, and rules. Data storage 32 may be a relational database, a flat data storage, a non-relational database, among others.
[0029] Data Sources
[0030] In one study conducted in Ontario, Canada, the Ontario Laboratory Information System (OLIS) was utilized for model derivation. This is a population-wide repository of laboratory test results from more than 90% of Ontario hospitals, outpatient and public health testing facilities and contains more than 3 billion records. The Electronic Medical Records Primary Care (EMRPC) cohort was used for external validation. EMRPC is composed of more than 400,000 patients managed by more than 350 primary care physicians. The EMRPC database contains comorbidities, medications, physical measurements and laboratory results extracted from primary care encounters (6, 7). Using unique encoded identifiers at ICES (formerly the Institute for Clinical Evaluative Sciences), the Ontario Laboratory Information System (OLIS) and EMRPC were linked to the Canadian Institute for Health Information Discharge Abstract Database to ascertain pre-existing comorbidities as well as follow-up of non-fatal hospitalization events across the province. Fatal cardiovascular events were ascertained through linkage to the Ontario Registrar General Database(8). The use of data in this project was authorized under section 45 of Ontario’s Personal Health Information Protection Act, which does not require review by a Research Ethics Board. Study design and analysis was conducted in accordance with TRIPOD guidelines for prognostic models.Example TRIPOD guidelines are shown below:*Items relevant only to the development of a prediction model are denoted by D, items relating solely to a validation of a prediction model are denoted by V, and items relating to both are denoted D;V. We recommend using the TRIPOD Checklist in conjunction with the TRIPOD Explanation and Elaboration document.
[0031] Derivation Cohort
[0032] The derivation cohort included all residents aged 40 to 75 years with serum lipid, complete blood count (CBC), creatinine and glucose measurements performed on the same day in the outpatient setting from April 1st, 2009 to December 31st, 2015. Laboratory measurements during hospitalization or within 30 days of discharge were excluded since they were unlikely to represent primary prevention assessments. The earliest measurement of same day laboratory tests was set as the index date. Patients with hospital admissions for cardiovascular disease (i.e. myocardial infarction, stroke, coronary revascularization, heart failure, and peripheral arterial disease) in the 20 years prior to the index date were excluded. Individuals with metastatic cancer, moderate-to-severe liver disease, dementia, solid organ transplant, dialysis-dependent renal disease, as well as long-term care residents were also excluded since long-term lipid lowering therapy may not be appropriate because of competing comorbidities. Individuals with extreme laboratory values (<0.1 or >99.9 percentile) were excluded because they might have represented spurious results or be related to severe comorbidities. Individuals with missing components for the CBC (ex. mean corpuscular volume) or lipid panel (ex. triglycerides) were excluded. Finally, individuals within the external validation cohort (see below) were excluded so that the validation cohort would remain independent of the derivation cohort.
[0033] External Validation Cohort
[0034] The external validation cohort was constructed from a cohort of individuals who visited their primary care physician between January 1, 2010 and December 31, 2014. Our objective was to perform external validation of the laboratory-based models alongside traditional models that rely on clinical variables. Hence, we included individuals if they had same-day serum lipid, CBC, creatinine and glucose measurements occurring within the 90 days after a blood pressure assessment. The earliest date of laboratory testing was set as the index date (9). Individuals with less than 1 year of medical record data prior to the index date were excluded. Then, the same exclusions for cardiovascular disease, comorbidities and extreme or missing laboratory values from the derivation cohort were applied to generate the final external validation cohort.
[0035] Serum Laboratory Predictors
[0036] Laboratory predictors included guideline-recommended screening panels for individuals undergoing risk estimation (1, 10). We included individuals who underwent fasting or non-fasting measurements since they have similar prognostic value (11). From the lipid panel, the predictors included serum total cholesterol, high-density lipoprotein cholesterol (HDL-C), and triglycerides. Components from the biochemistry panel included serum glucose and the eGFR which was calculated using serum creatinine without adjustments for Black individuals (12, 13). Components of the CBC that have been independently validated as risk markers for ASCVD included hemoglobin, mean corpuscular volume (MCV), platelets, and leukocytes (14-17). Since reporting of differential leukocyte counts is routine in some jurisdictions such as Ontario, we considered additional models that replaced leukocytes with monocytes and the neutrophil lymphocyte ratio based on previously reported association with ASCVD (14). The predefined predictor variables comprised at least one of serum total cholesterol, high-density lipoprotein cholesterol, triglycerides, hemoglobin, mean corpuscular volume, platelets, leukocytes, monocytes, neutrophils, lymphocytes, estimated glomerular filtration rate, and glucose.
[0037] Outcomes
[0038] The outcome of interest was ASCVD ascertained using validated administrative algorithms (18-20). ASCVD was defined as a composite of hospitalization for myocardial infarction (ICD-10: 121-122), stroke (160, 161, 163 [excluding 163.6], 164, H34.1), death from ischemic heart disease (120-125) or death from cerebrovascular disease (160-169). Individuals werefollowed to the first of: an ASCVD event, death from a non-ASCVD cause, or December 31st, 2018 (8, 9).
[0039] Statistical Analysis
[0040] Model Development
[0041] Sex-specific multivariable Fine-Gray subdistribution hazards models, herein referred to as the CANHEART Lab Models, were developed in the derivation cohort (21). The incidence of ASCVD was regressed on age and the pre-specified laboratory predictor variables. All predictors were centered on the mean values and then modelled using restricted cubic splines (22). Predictor importance was assessed as each variables’ Wald / 2relative to the overall / 2in the multivariable model. Separate models derived i) using age alone, ii) without CBC, and iii) with differential leukocyte counts were generated for comparison to the CANHEART Lab Model on internal validation. Percentile charts were generated for predicted risk at 5 years for every age and sex combination in the derivation cohort (23).
[0042] Internal Validation
[0043] Discrimination was assessed quantitively by estimating the C-statistic (24). Calibration was quantitatively assessed using the calibration slope (25). Optimism in the C-statistic and calibration slope were estimated using 200 bootstrap resamples (22). Calibration was also assessed as the relative difference in mean predicted and observed risk (i.e. discordance) as well as absolute difference in risk. Calibration was assessed graphically through plots of predicted and observed risk. Decision curves were constructed to compare net benefit for all laboratory-based models across 5 -year treatment thresholds of 1 to 5% (26).
[0044] External Validation
[0045] Overlap in the distributions of the linear predictors in the internal and external validation samples was estimated (27, 28). Discrimination and calibration were assessed using the same measures as the internal validation set. Overall performance was also assessed using a Brier Score adapted for survival data with right censoring (29, 30). We generated 95% confidence intervals (CI) for calibration measures using 200 bootstrap resamples. Further details on model derivation and validation are presented in the Supplementary Methods.
[0046] Comparison to Traditional Models
[0047] In all individuals without missing smoking status, the CANHEART Lab Models were compared to the PCEs. Since the validation cohort has not yet completed 10 years of follow-up,predicted ASCVD risk at 5-years was estimated using the PCEs and SO(t) at 5 years for the pooled cohorts (31). Overall performance (i.e. Brier score) and discrimination (i.e. C-statistic) between the laboratory models and the PCEs were reported as they are less influenced by model miscalibration (32). We elected not to compare calibration or net benefit, since the PCEs have been reported to overestimate risk in contemporary cohorts (7, 33, 34). All analyses were conducted at ICES using SAS version 9.7 and R version 3.6.1 using the survival, cmprsk, rms, dca, boot, riskRegression and overlapping packages.
[0048] RESULTS
[0049] Derivation Cohort
[0050] There were 4,610,826 Ontario residents between 40 and 75 years of age with outpatient serum laboratory values. After applying exclusions, there were 3,993,644 individuals remaining that comprised the derivation cohort (Supplementary Figure 2A). Baseline characteristics are presented in Table 1. The mean age was 54 years. Hypertension was the most common cardiovascular risk factor. There were 57.0% of women and 58.1% of men with fasting bloodwork.Table 1. Baseline CharacteristicsWomen MenExternal ExternalDerivation DerivationBaseline Characteristics Validation Validation Cohort Cohort Cohort CohortN=2,160,497 N=18,342 N=l,833,147 N=13,355DemographicsAge, years , mean (SD) 54.92 (9.45) 54.96 (9.33) 54.69 (9.26) 54.61 (9.15)1040685 866315>55, n (%) (48.2) 8887 (48.5) (47.3) 6236 (46.7)Baseline Risk Factors718201 643241Hypertension, n (%)* (33.2) 3917 (21.4) (35.1) 3070 (23.0) 216715Diabetes, n (%) 196899 (9.1) 1248 (6.8) (11.8) 1185 (8.9)Chronic Kidney Disease (eGFR < 60 mL / min), n (%) 84433 (3.9) 656 (3.6) 63568 (3.5) 370 (2.8)125.51 130.09Systolic Blood Pressure, mmHg, mean (SD) (17.88) - (17.04)Active Smoker, n (%)+ 2459 (13.4) 2243 (16.8)Women MenExternal ExternalDerivation . Derivation .Baseline Characteristics _ . Validation _ . ValidationCohort , CohortCohort CohortN=2,160,497 N=18,342 N=l,833,147 N=13,355Baseline Laboratory Variables 1231932 10881 1064413Fasting, n (%) (57.0) (59.3) (58.1) 8159 (61.1)Hemoglobin, g / L, mean (SD) 13.43 (1.08) 13.53 (1.02) 15.03 (1.07) 15.10 (1.02)Mean Corpuscular Volume, fL, mean (SD) 90.35 (5.71) 91.29 (5.22) 90.84 (5.30) 91.55 (4.90)254.28 251.47 225.27 223.68Platelets, xlO9 / L, mean (SD) (59.43) (57.80) (52.99) (51.71)Leukocytes, xlO9 / L, mean (SD) 6.32 (1.78) 6.23 (1.76) 6.59 (1.77) 6.52 (1.76)NeutrophikLymphocyte Ratio, mean (SD) 2.02 (0.91) 2.06 (0.91) 2.13 (0.98) 2.21 (1.00)Monocytes, xlO9 / L, mean (SD) 0.46 (0.15) 0.47 (0.15) 0.53 (0.17) 0.54 (0.17)90.43 89.43 89.54 89.68 eGFR, mL / min, mean (SD) (15.54) (15.00) (14.71) (13.93)Total Cholesterol, mean (SD) 202.25 204.48 194.55 197.16 mg / dL (39.72) (38.77) (40.70) (39.25) mmol / L 5.23 (1.03) 5.29 (1.00) 5.03 (1.05) 5.10 (1.01)High-Density Lipoprotein Cholesterol, mg / dL, mean (SD) 59.35 61.59 47.98 49.10 mg / dL (15.68) (16.24) (12.75) (13.00) mmol / L 1.53 (0.41) 1.59 (0.42) 1.24 (0.33) 1.27 (0.34)Triglycerides, mean (SD) 120.75 117.02 146.07 142.22 mg / dL (71.58) (71.14) (94.94) (92.84) mmol / L 1.36 (0.81) 1.32 (0.80) 1.65 (1.07) 1.61 (1.05)Glucose, mean (SD) 98.63 94.49 105.27 100.44 mg / dL (26.65) (22.16) (32.15) (27.63) mmol / L 5.48 (1.48) 5.25 (1.23) 5.85 (1.79) 5.58 (1.54)* indicates individuals with hypertension on anti-hypertensive therapy in the external validation cohort+ missing in 16.9% of patients in the external validation cohort eGFR - estimated glomerular filtration rate, SD - standard deviation
[0051] Table 2. Validation of the CANHEART Laboratory Model
[0052] Model Development
[0053] Over a median follow-up of 7.8 years, there were a total of 40,759 ASCVD events in women and 71,664 in men. The cumulative incidence of ASCVD at 5 years was 1.09% and 2.47% in women and men, respectively. The regression coefficients, centering values, knot locations and baseline survival for the models are presented in Supplementary Tables 1 and 2. In both sexes, age provided the greatest contribution to model fit, as can be seen in Figures 2A and 2B. Age and sexspecific percentiles of predicted risk are presented in Figures 3 and 4.
[0054] Internal Validation
[0055] The C-statistics for the CANHEART Lab Models was 0.77 in women and 0.71 in men while the calibration slope was 1.00 for both sexes (Table 2). Optimism in the C-statistic and calibration slope was less than 0.01 in both sexes. Predicted risk of ASCVD was closely matched to observed risk with discordance being less than 1%. Calibration was also good across the spectrum of predicted risk (Figure 3) and among clinically relevant subgroups (Figure 4). The proportion of individuals with modifiable clinical risk factors increased with higher categories of predicted ASCVD risk (Figure 5).
[0056] Results for additional laboratory-based models are presented in Supplementary Tables 3 - 8 and Figures 6-7. In comparison to the CANHEART Lab Models, the C-statistics were lower and the difference was statistically significant for all models except the model incorporating differential leukocyte counts in women (Supplementary Table 9). Similarly, net benefit was comparable for the CANHEART Lab Models and the model that incorporated differential leukocyte counts but lower at all treatment thresholds for other laboratory-based models (Figure 8).
[0057] External Validation Cohort
[0058] There were 36,978 eligible primary care patients in the external validation cohort. After similar exclusions were applied, 31,697 remained for validation assessment (Supplementary Figure 2B). The mean age was 54 years in both sexes (Table 1). Fifty-nine percent of women and 61% of men had fasting bloodwork. There were 7.6% of women and 12.0% of men on statin therapy at baseline.
[0059] External Validation
[0060] Over a median follow-up of 7.3 years, there were 279 ASCVD events in women and 451 in men. The cumulative incidence of ASCVD at 5 years was 0.91% in women and 2.06% in men. This was 17% lower than the incidence of ASCVD in the derivation cohort for both sexes (Table 2). The overlap agreement between the CANHEART Lab Model’s linear predictor in derivation and external validation cohorts was greater than 97% in both sexes (Figure 9). The C- statistic in the external validation cohort was 0.72 for women and men. The mean predicted risk was 11.4% higher relative to observed risk in women (1.01% vs. 0.91%; absolute difference 0.103%) and 13.5% higher in men (2.34% vs. 2.06%; absolute difference 0.278%). Calibration plots revealed that overestimation was greater for high-risk patients, but overall, well calibrated for low-risk patients (Figure 3 and Figure 4). Enrichment of modifiable clinical risk factors was noted in both sexes as ASCVD risk increased (Figure 5).
[0061] Comparison to the Pooled Cohort Equations
[0062] Among the 26,342 individuals in the external validation cohort with complete information on smoking history, the difference in the C-statistics between the CANHEART Lab Models and the PCEs was -0.01 (95% CL -0.03 to 0.01) in women and -0.01 (95% CL -0.04 to 0.02), which was not statistically significant. Similarly, the difference in the Brier scores for the CANHEART Lab Models compared to the PCEs was not statistically significant in women (- 0.00006, 95% CI: -0.00015 to 0.00002) or men (-0.00006, 95% CI: -0.0008 to 0.0002).
[0063] Additional Models
[0064] In another example, machine learning models are employed to reliably predict cardiovascular risk in patients and / or predict future outcomes of individuals with cardiovascular disease. These models take into account the data observation mechanisms and training procedures of a number of different algorithms, and their efficacy may be verified using the test set with other classification models.
[0065] Figure 10 shows flow diagram 100 depicting a computer-implemented method for predicting cardiovascular risk in a patient in accordance with various implementations described herein. It should be understood that while the operational flow diagram indicates a particular order of execution of the operations, in other implementations, the operations might be executed in a different order. Further, in some implementations, additional operations or blocks may be added to the method. Likewise, some operations or blocks may be omitted.
[0066] At step 102, the datasets used for developing the detection models was taken from the the Ontario Laboratory Information System (OLIS) data repository with patient demographic data and clinical data, as described above. The dataset contains a plurality of attributes, such as age, sex, blood pressure, cholesterol, diabetes etc. and other risk factors, and associated integer values
[0067] At step 104, the datasets undergo a pre-processing stage which generally involves smoothing, standardization, and aggregation. This stage may comprise one or more actions, such as eliminating null values, filtering for denoizing, and removing any outliers present in the datasets, thereby minimizing distorted measurements, which ultimately makes predictions more dependable.
[0068] At step 106, subsets of datasets are randomly selected to either train the predictive models or to assess the predictive models e.g., a 80: 20 split.
[0069] At step 108, features that contain the information that is used to make decisions for predicting cardiovascular risk is extracted from the datasets. Such features may include demographic data (e.g., age, among others); baseline risk factors (e.g., diabetes, blood pressure, smoker, among others) and baseline laboratory variables (e.g., glucose, triglycerides, eGFr, among others).
[0070] At step 110, a subset of features or attributes for making the best predictions are selected. In one example, each input feature is multiplied by some value, or weight, such that during the training step, the weights are updated until the best model is found.
[0071] At step 112, the training data set and the feature vectors are used to fully train one or more predictive models. In one example, different machine learning classifiers or algorithms are used for building the predictive models, such as, supervised learning algorithms, unsupervised learning algorithms and reinforcement learning algorithms. Examples of supervised learning algorithm systems include support vector machine, decision tree, linear regression, logistic regression, naive Bayes, ^-nearest neighbor, random forest, AdaBoost, XGBoost, and neuralnetwork methods. Examples of unsupervised learning algorithm systems include K-means, mean shift, affinity propagation, hierarchical clustering, DBSCAN (density -based spatial clustering of applications with noise), Gaussian mixture modeling, Markov random fields, ISODATA (iterative self-organizing data), and fuzzy C-means systems. Examples of reinforcement learning algorithm systems include Maja and Teaching-Box systems. Generally, training the predictive models involves optimizing the parameters of a predicitive system to minimize the loss function. In addition to the training step, the predictive models also undergoes validation using test datasets.
[0072] At step 114, following training and validation of the predictive models in step 112, a trained predictive models are outputted and are stored in a trained model database. The trained models are used to predict cardiovascular risk when provided with subject patient data. As such, the trained predictive models can predict the risk of a cardiovascular event, such as, an acute coronary syndrome, stable or unstable angina, arterial revascularization, stroke, transient ischemic attack, peripheral arterial disease, atherosclerotic cardiovascular disease, heart failure, atrial fibrillation, myocardial infarction, and sudden coronary death.
[0073] Figure 11 illustrates a block diagram of an example of a machine 12 upon which any one or more of the techniques (e.g., methodologies) discussed herein can be performed. In alternative embodiments, the machine 12 can operate as a standalone device or are connected (e.g., networked) to other machines. In a networked deployment, the machine 12 can operate in the capacity of a server machine, a client machine, or both in server-client network environments. In an example, the machine 12 can act as a peer machine in peer-to-peer (P2P) (or other distributed) network environment. The machine 12 is a personal computer (PC), a tablet PC, a mobile device, a web appliance, or any machine capable of executing instructions (sequential or otherwise) that specify actions to be taken by that machine. In various embodiments, machine 12 can perform one or more of the processes described above. Further, while only a single machine is illustrated, the term “machine” shall also be taken to include any collection of machines that individually or jointly execute a set (or multiple sets) of instructions to perform any one or more of the methodologies discussed herein, such as cloud computing, software as a service (SaaS), other computer cluster configurations.
[0074] Examples, as described herein, can include, or can operate on, logic or a number of components, modules, or mechanisms (all referred to hereinafter as “modules”). Modules are tangible entities (e.g., hardware) capable of performing specified operations and is configured orarranged in a certain manner. In an example, circuits are arranged (e.g., internally or with respect to external entities such as other circuits) in a specified manner as a module. In an example, the whole or part of one or more computer systems (e.g., a standalone, client or server computer system) or one or more hardware processors are configured by firmware or software (e.g., instructions, an application portion, or an application) as a module that operates to perform specified operations. In an example, the software can reside on a non-transitory computer readable storage medium or other machine-readable medium. In an example, the software, when executed by the underlying hardware of the module, causes the hardware to perform the specified operations.
[0075] Accordingly, the term “module” is understood to encompass a tangible entity, be that an entity that is physically constructed, specifically configured (e.g., hardwired), or temporarily (e.g., transitorily) configured (e.g., programmed) to operate in a specified manner or to perform part or all of any operation described herein. Considering examples in which modules are temporarily configured, each of the modules need not be instantiated at any one moment in time. For example, where the modules comprise a general-purpose hardware processor configured using software, the general-purpose hardware processor is configured as respective different modules at different times. Software can accordingly configure a hardware processor, for example, to constitute a particular module at one instance of time and to constitute a different module at a different instance of time.
[0076] Machine 12 can include a hardware processor 202 (e.g., a central processing unit (CPU), a graphics processing unit (GPU), a hardware processor core, or any combination thereof), a main memory 204, and a static memory 206, some or all of which can communicate with each other via an interlink 208 (e.g., bus). The machine 12 can further include a display unit 20, an alphanumeric input device 212 (e.g., a keyboard), and a user interface (UI) navigation device 214 (e.g., a mouse). In an example, the display unit 20, input device 212 and UI navigation device 214 are a touch screen display. The machine 12 can additionally include a storage device (e.g., drive unit) 216, a signal generation device 218 (e.g., a speaker), a network interface device 220, and one or more sensors 221, such as an accelerometer, or other sensor. The machine 12 can include an output controller 228, such as a serial (e.g., universal serial bus (USB), parallel, or other wired or wireless (e.g., infrared (IR), near field communication (NFC), etc.) connection to communicate or control one or more peripheral devices (e.g., a printer, card reader, etc.).
[0077] The storage device 216 can include a machine readable medium 222 on which is stored one or more sets of data structures or instructions 224 (e.g., software) embodying or utilized by any one or more of the techniques or functions described herein. The instructions 224 can also reside, completely or at least partially, within the main memory 204, within static memory 206, or within the hardware processor 202 during execution thereof by the machine 12. In an example, one or any combination of the hardware processor 202, the main memory 204, the static memory 206, or the storage device 216 can constitute machine readable media. While the machine readable medium 222 is illustrated as a single medium, the term "machine readable medium" can include a single medium or multiple media (e.g., a centralized or distributed database, and / or associated caches and servers) configured to store the one or more instructions 224.
[0078] The term “machine readable medium” can include any medium that is capable of storing, encoding, or carrying instructions for execution by the machine 12 and that cause the machine 12 to perform any one or more of the techniques of the present disclosure, or that is capable of storing, encoding, or carrying data structures used by or associated with such instructions. Nonlimiting machine-readable medium examples can include solid-state memories, and optical and magnetic media. Specific examples of machine-readable media can include: nonvolatile memory, such as semiconductor memory devices (e.g., Electrically Programmable Read- Only Memory (EPROM), Electrically Erasable Programmable Read-Only Memory (EEPROM)) and flash memory devices; magnetic disks, such as internal hard disks and removable disks; magneto-optical disks; Random Access Memory (RAM); Solid State Drives (SSD); and CD- ROM and DVD-ROM disks. In some examples, machine readable media can include non- transitory machine-readable media. In some examples, machine readable media can include machine readable media that is not a transitory propagating signal.
[0079] The instructions 224 can further be transmitted or received over a communications network 226 using a transmission medium via the network interface device 220. The machine 12 can communicate with one or more other machines utilizing any one of a number of transfer protocols (e.g., frame relay, internet protocol (IP), transmission control protocol (TCP), user datagram protocol (UDP), hypertext transfer protocol (HTTP), etc.). Example communication networks can include a local area network (LAN), a wide area network (WAN), a packet data network (e.g., the Internet), mobile telephone networks (e.g., cellular networks), Plain Old Telephone (POTS) networks, and wireless data networks (e.g., Institute of Electrical andElectronics Engineers (IEEE) 602.11 family of standards known as Wi-Fi®, IEEE 602.16 family of standards known as WiMax®), IEEE 602.15.4 family of standards, a Long Term Evolution (LTE) family of standards, a Universal Mobile Telecommunications System (UMTS) family of standards, peer-to-peer (P2P) networks, among others. In an example, the network interface device 220 can include one or more physical jacks (e.g., Ethernet, coaxial, or phonejacks) or one or more antennas to connect to the communications network 226. In an example, the network interface device 220 can include a plurality of antennas to wirelessly communicate using at least one of single-input multiple-output (SIMO), multiple-input multiple-output (MIMO), or multipleinput single-output (MISO) techniques. In some examples, the network interface device 220 can wirelessly communicate using Multiple User MIMO techniques.
[0080] Examples, as described herein, can include, or can operate on, logic or a number of components, modules, or mechanisms. Modules are tangible entities (e.g., hardware) capable of performing specified operations and are configured or arranged in a certain manner. In an example, circuits are arranged (e.g., internally or with respect to external entities such as other circuits) in a specified manner as a module. In an example, the whole or part of one or more computer systems (e.g., a standalone, client, or server computer system) or one or more hardware processors are configured by firmware or software (e.g., instructions, an application portion, or an application) as a module that operates to perform specified operations. In an example, the software can reside on a machine-readable medium. In an example, the software, when executed by the underlying hardware of the module, causes the hardware to perform the specified operations.
[0081] Accordingly, the term “module” is understood to encompass a tangible entity, be that an entity that is physically constructed, specifically configured (e.g., hardwired), or temporarily (e.g., transitorily) configured (e.g., programmed) to operate in a specified manner or to perform part or all of any operation described herein. Considering examples in which modules are temporarily configured, each of the modules need not be instantiated at any one moment in time. For example, where the modules comprise a general-purpose hardware processor configured using software, the general-purpose hardware processor is configured as respective different modules at different times. Software can accordingly configure a hardware processor, for example, to constitute a particular module at one instance of time and to constitute a different module at a different instance of time.
[0082] Method examples described herein can be machine or computer-implemented at least in part. Some examples can include a computer-readable medium or machine-readable medium encoded with instructions operable to configure an electronic device to perform methods as described in the above examples. An implementation of such methods can include code, such as microcode, assembly language code, a higher-level language code, or the like. Such code can include computer readable instructions for performing various methods. The code can form portions of computer program products. Further, in an example, the code can be tangibly stored on one or more volatile, non-transitory, or non-volatile tangible computer-readable media, such as during execution or at other times. Examples of these tangible computer-readable media can include, but are not limited to, hard disks, removable magnetic disks, removable optical disks (e.g., compact discs and digital video discs), magnetic cassettes, memory cards or sticks, random access memories (RAMs), read only memories (ROMs), and the like.
[0083] Each of the non-limiting aspects or examples described herein can stand on its own, or can be combined in various permutations or combinations with one or more of the other examples.
[0084] The above detailed description includes references to the accompanying drawings, which form a part of the detailed description. The drawings show, by way of illustration, specific embodiments in which the invention can be practiced. These embodiments are also referred to herein as “examples.” Such examples can include elements in addition to those shown or described. However, the present inventors also contemplate examples in which only those elements shown or described are provided. Moreover, the present inventors also contemplate examples using any combination or permutation of those elements shown or described (or one or more aspects thereof), either with respect to a particular example (or one or more aspects thereof), or with respect to other examples (or one or more aspects thereof) shown or described herein.
[0085] In this document, the terms “a” or “an” are used, as is common in patent documents, to include one or more than one, independent of any other instances or usages of “at least one” or “one or more.” In this document, the term “or” is used to refer to a nonexclusive or, such that “A or B” includes “A but not B,” “B but not A,” and “A and B,” unless otherwise indicated. In this document, the terms “including” and “in which” are used as the plain-English equivalents of the respective terms “comprising” and “wherein.” Also, in the following claims, the terms “including” and “comprising” are open-ended, that is, a system, device, article, composition,formulation, or process that includes elements in addition to those listed after such a term in a claim are still deemed to fall within the scope of that claim. Moreover, in the following claims, the terms “first,” “second,” and “third,” etc. are used merely as labels, and are not intended to impose numerical requirements on their objects.
[0086] The above description is intended to be illustrative, and not restrictive. For example, the above-described examples (or one or more aspects thereof) can be used in combination with each other. Other embodiments can be used, such as by one of ordinary skill in the art upon reviewing the above description. The Abstract is provided to allow the reader to quickly ascertain the nature of the technical disclosure. It is submitted with the understanding that it will not be used to interpret or limit the scope or meaning of the claims. Also, in the above Detailed Description, various features can be grouped together to streamline the disclosure. This should not be interpreted as intending that an unclaimed disclosed feature is essential to any claim. Rather, inventive subject matter can he in less than all features of a particular disclosed embodiment. Thus, the following claims are hereby incorporated into the Detailed Description as examples or embodiments, with each claim standing on its own as a separate embodiment, and it is contemplated that such embodiments can be combined with each other in various combinations or permutations. The scope of the invention should be determined with reference to the appended claims, along with the full scope of equivalents to which such claims are entitled.
[0087] The flowchart and block diagrams in the Figures illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of the present invention. In this regard, each block in the flowchart or block diagrams may represent a module, segment, or portion of instructions, which comprises one or more executable instructions for implementing the specified logical fiinction(s). In some alternative implementations, the functions noted in the block may occur out of the order noted in the figures. For example, two blocks shown in succession may, in fact, be executed substantially concurrently, or the blocks may sometimes be executed in the reverse order, depending upon the functionality involved. It will also be noted that each block of the block diagrams and / or flowchart illustration, and combinations of blocks in the block diagrams and / or flowchart illustration, can be implemented by special purpose hardware-based systems that perform the specified functions or acts or carry out combinations of special purpose hardware and computer instructions.
[0088] Benefits, other advantages, and solutions to problems have been described above with regard to specific embodiments. However, the benefits, advantages, solutions to problems, and any element(s) that may cause any benefit, advantage, or solution to occur or become more pronounced are not to be construed as critical, required, or essential features or elements of any or all the claims. As used herein, the terms "comprises," "comprising," or any other variations thereof, are intended to cover a non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements does not include only those elements but may include other elements not expressly listed or inherent to such process, method, article, or apparatus. Further, no element described herein is required for the practice of the invention unless expressly described as "essential" or "critical."
[0089] The preceding detailed description of exemplary embodiments of the invention makes reference to the accompanying drawings, which show the exemplary embodiment by way of illustration. While these exemplary embodiments are described in sufficient detail to enable those skilled in the art to practice the invention, it should be understood that other embodiments may be realized and that logical and mechanical changes may be made without departing from the spirit and scope of the invention. For example, the steps recited in any of the method or process claims may be executed in any order and are not limited to the order presented. Thus, the preceding detailed description is presented for purposes of illustration only and not of limitation, and the scope of the invention is defined by the preceding description, and with respect to the attached claims.
[0090] DISCUSSION
[0091] In this population-based study of nearly 4 million individuals, we developed the CANHEART Lab models, which are sex-specific prediction models for ASCVD. The model predictors included age as well as routine laboratory test results from the CBC, lipid panel and serum biochemistry. They were internally and externally validated with good calibration in an independent primary care cohort. Furthermore, additional measures of accuracy including the Brier score and C-statistics were similar between the CANHEART Lab Models and the PCEs in the same primary care cohort. Our findings, therefore, suggest that ASCVD can be accurately predicted without clinical risk factors, which may greatly simplify risk prediction in routine clinical practice.
[0092] Risk models are under-utilized in routine clinical practice, leading to many missed opportunities to institute therapies to prevent ASCVD (35). A common theme across studies examining practice patterns in preventative medicine is that risk estimation is felt to be time consuming and lacks automation (2, 4, 36). A survey of physicians in the United States suggested that only 33% of physicians routinely used risk models. The most important physician-level barrier reported by primary care providers was a lack of time (2). In another nationwide survey exclusively of primary care physicians in the United States, 42% reported never using risk calculators while an additional 20% reported rarely using them, once again citing time as the most important barrier (36). More recently, an academic internal medicine clinic at a tertiary safety net hospital reported on a quality improvement initiative to increase ASCVD risk estimation using the PCEs (4). Over a 12-month period and four-step quality-improvement strategy that included integration in electronic health record platforms, ASCVD risk reporting increased by 19%, but the overall completion rate remained low at 33%. The lack of automation and integration of risk estimators into electronic platforms remained a crucial barrier to adoption.
[0093] The CANHEART Lab models potentially circumvents these limitations since they use only age and routine laboratory test results and could be implemented at the time of laboratory testing. For instance, laboratory requisitions could include options for ASCVD risk estimation, laboratory vendors could estimate risk with the models on behalf of physicians, and risk could be reported directly alongside laboratory results with interpretive comments (37). This may lead to greater uptake because it could alleviate physicians from the time burden imposed by traditional models and allow them to focus their time on decision-making and patient-risk discussions. This is supported by an observational study across four primary care clinics in Canada that examined laboratory-initiated reporting of Framingham risk (38). When laboratories were provided with the risk factor information to report ASCVD risk alongside lipid panel results, statin prescription rates by primary care providers increased by up to 25%.
[0094] The CANHEART Lab models demonstrated excellent internal and external validity. Despite using only laboratory tests, the models were well calibrated in individuals with and without traditional risk factors (e.g., diabetes, smoking history, etc.). While we did note a small degree of overestimation on external validation, this may be explained by the fact that the validation cohort was a lower risk cohort compared to the derivation cohort. The 17% higher incidence of ASCVD at 5 years in the derivation cohort relative to the external validation cohort was consistent with the11-13% overestimation we observed with the models. This degree of overestimation is generally lower than previous studies for clinical models such PCEs that were derived in historical and higher risk natural history cohorts. This includes a previous validation study using the same EMRPC cohort, in which the PCEs overestimated risk by more than 100% (9, 39). Furthermore, overestimation was greatest in patients in the highest estimated risk treatment categories, which has also been observed for clinical models such as the PCEs. This may be explained by the effects of preventative treatment initiation overtime where observed risks are lower than would have been expected in natural history cohorts that were studied before the availability of statins (33). Arguably, overestimation in high-risk treatment categories may be less relevant for decisionmaking since these patients would be candidates for therapy even if overestimation was not present.
[0095] The ability to accurately predict ASCVD with routine laboratory tests extends prior knowledge for risk modelling in primary prevention. Using the Atherosclerosis Risk in Communities study, investigators have previously developed a prediction model using age, sex, race, N-terminal pro-B-type natriuretic peptide, high-sensitivity cardiac troponin T, and high- sensitivity C-reactive protein (40). While it demonstrated improved discrimination compared with the PCEs, it was only validated among individuals (>70 years) and for outcomes that included heart failure. The Multi-ethnic Study of Atherosclerosis group also developed a model with similar accuracy as the PCEs but included test results not commonly used in clinical practice (e.g. homocysteine) (41). In contrast, simpler models have been developed that use inputs from CBCs and serum electrolytes (42, 43). However, these models were not developed exclusively for primary prevention and they predict all-cause mortality, rather than ASCVD which is the more relevant endpoint in prevention (44).
[0096] Even though we did not have information on statin use in our derivation cohort, we believe the CANHEART Lab models remain valid and applicable to contemporary practice for several reasons. First, the models remained well calibrated on external validation in patients prescribed and not prescribed statins at baseline. This is consistent with the newly developed Systematic Coronary Risk Evaluation 2 model which does not incorporate baseline statin treatment as a predictor yet demonstrates good calibration across a diverse number of contemporary cohorts (45). Second, statin therapy was considered as a covariate for the PCEs, but did not demonstrate a significant association with ASCVD outcomes and was therefore not retained in the final model(46). Perhaps this is because risk estimates in primary prevention assessments are more dependent on the cholesterol level rather than treatment itself. This is supported by studies in clinical trial populations and genetic cohorts, where low-density lipoprotein cholesterol is linearly associated with ASCVD outcomes irrespective of treatment (47). Third, excluding individuals treated with statins in contemporary risk score derivation or validation studies can have detrimental effects on model performance by excluding individuals with risk factor profdes that are higher risk for ASCVD (48). Many individuals on statins discontinue treatment over time (49). Furthermore, the majority individuals on statins do not achieve treatment targets, implying a residual ASCVD risk (50).
[0097] Although guidelines recommend engaging individuals in risk discussions, absolute risk estimates may be challenging for some individuals to conceptualize (1, 51). When short-term estimates are low, it is conceivable that individuals may perceive false reassurance of low future cardiovascular risk despite having a high level of risk factors and high lifetime ASCVD risk relative to their peers. To circumvent this, we constructed risk tables accompanying the CANHEART Lab models that provide a complementary measure of risk relative to age and sex- matched individuals in the Ontario population (23). For the clinical vignette of a 50-year-old male presented in Supplementary Table 2, the 5-year risk of ASCVD is 2.60%, which is borderline for treatment initiation. However, relative to all 50-year-old males, this risk estimate is above the 80th percentile, which is more concerning and may help inform risk discussions.
[0098] Our results should be interpreted in the context of limitations. First, the CANHEART Lab Models predicts ASCVD risk over the short-term (i.e. 5 years). While longer time frames may increase utilization of statins in younger individuals, shorter prediction times remain concordant with benefits observed in clinical trials (10, 52). Second, implementation may be less feasible in countries where laboratory testing is not readily accessible (53). Third, while data on emerging risk-enhancers such as Lipoprotein(a) results were not available at the population-level, incorporating additional laboratory parameters can be accomplished with relative ease and no loss in the potential for automation (54). Fourth, calibration may be less optimal when considering individuals residing in countries with markedly different rates of ASCVD than Ontario. Further research is needed to explore strategies to improve transportability in these scenarios such as crosscountry recalibration (55).
[0099] In conclusion, ASCVD risk can be accurately predicted with age and routine serum laboratory tests and the performance of these models are comparable to clinical models. Future studies are needed to determine whether automating these models in daily practice improves prescribing of preventative treatments according to clinical practice guidelines.
[0100] SUPPLEMENTARY TABLESSupplementary Table 1. Predictor Means and Knot Locations for the CANHEART LaboratoryModel with Leukocytes.MCV - mean corpuscular volume, eGFR - estimated glomerular filtration rate, HDL - high-density lipoprotein, TC- total cholesterol.Supplementary Table 2. Regression Equation Parameters for the CANHEART Laboratory Model with Leukocytes.The pooled regression coefficients and baseline survival are presented for women and men. A case example is provided demonstrate risk estimation by first i) calculating the corresponding spline terms for the non-linear terms using the knot locations specified in Supplementary Table I using the formulas:Where Stand Sfc2is the Is* and 2"*1spline terms with k=4 knots. The knot locations are given by tk. X represents the value of the predictor and (X — t*)+ equals (X — t*)+ when X — tk> 0 and 0 otherwise. Then ii) the product between each term and the regression coefficients is calculated, followed by iii) taking the sum of these products to obtain the linear predictor, and finally iv) estimating the 5-year incidence of ASCVD as 1 — (Baseline Survival)^ ^u”ear f,redlc“n':i.Supplementary Table 3. Predictor Means and Knot Locations for the CANHEART Laboratory Model with White BloodCell Differential.Supplementary Table 4. Regression Equation Parameters for the CANHEART Laboratory Model with White Blood CellDifferential. The pooled regression coefficients and baseline survival are presented for women and men. Spline terms and knot locations are specified in Supplementary Table 3.Supplementary Table 5. Validation of the CANHEART Laboratory Model with White Blood Cell Differential.Supplementary Table 6. Predictor Means and Knot Locations for the CANHEART Laboratory Model without CompleteBlood Count.Supplementary Table 7. Regression Equation Parameters for the CANHEART Laboratory Model without CompleteBlood Count. The pooled regression coefficients and baseline survival are presented for women and men without complete bloodSupplementary Table 8. Validation of the CANHEART Laboratory Model without Complete Blood Cell Count.* Discordance is estimated as 100 x (mean predicted event risk - observed event risk) / observed event risk+ Absolute difference is estimated as the difference between mean predicted and observed event riskSupplementary Table 9. Comparison of Laboratory-Based Models on Internal Validation
[0101] References
[0102] The following references, to the extent that they provide exemplary procedural or other details supplementary to those set forth herein, are specifically incorporated herein by reference:1. Grundy SM, Stone NJ, Bailey AL, Beam C, Birtcher KK, Blumenthal RS, et al. 2018AHA / ACC / AACVPR / AAPA / ABC / ACPM / ADA / AGS / APhA / ASPC / NLA / PCNAGuideline on the Management of Blood Cholesterol: A Report of the American College of Cardiology / American Heart Association Task Force on Clinical Practice Guidelines.Circulation. 2019;139(25):el082-el43.2. Shillinglaw B, Viera AJ, Edwards T, Simpson R, Sheridan SL. Use of global coronary heart disease risk assessment in practice: a cross-sectional survey of a sample of U.S. physicians. BMC Health Serv Res. 2012; 12:20.3. North F, Fox S, Chaudhry R. Clinician time used for decision making: a best case workflow study using cardiovascular risk assessments and Ask Mayo Expert algorithmic care process models. BMC Med Inform Decis Mak. 2016;16:96.4. Bakhai S, Bhardwaj A, Sandhu P, Reynolds JL. Optimisation of lipids for prevention of cardiovascular disease in a primary care. BMJ Open Qual. 2018;7(3):e000071.5. Richards N, Harris K, Whitfield M, O'Donoghue D, Lewis R, Mansell M, et al. The impact of population-based identification of chronic kidney disease using estimated glomerular filtration rate (eGFR) reporting. Nephrol Dial Transplant. 2008;23(2):556-61.6. Tu K, Widdifield J, Young J, Oud W, Ivers NM, Butt DA, et al. Are family physicians comprehensively using electronic medical records such that the data can be used for secondary purposes? A Canadian perspective. BMC Med Inform Decis Mak. 2015; 15 :67.Sud M, Sivaswamy A, Chu A, Austin PC, Anderson TJ, Naimark DMJ, et al. Population- Based Recalibration of the Framingham Risk Score and Pooled Cohort Equations. J Am Coll Cardiol. 2022;80(14): 1330-42. Tu JV, Chu A, Donovan LR, Ko DT, Booth GL, Tu K, et al. The Cardiovascular Health in Ambulatory Care Research Team (CANHEART): using big data to measure and improve cardiovascular health and healthcare services. Circ Cardiovasc Qual Outcomes. 2015;8(2):204-12. Ko DT, Sivaswamy A, Sud M, Kotrri G, Azizi P, Koh M, et al. Calibration and discrimination of the Framingham Risk Score and the Pooled Cohort Equations. CMAJ. 2020;192(17):E442-E9. Pearson GJ, Thanassoulis G, Anderson TJ, Barry AR, Couture P, Dayan N, et al. 2021 Canadian Cardiovascular Society Guidelines for the Management of Dyslipidemia for the Prevention of Cardiovascular Disease in Adults. Can J Cardiol. 2021. Langsted A, Freiberg JJ, Nordestgaard BG. Fasting and nonfasting lipid levels: influence of normal food intake on lipids, lipoproteins, apolipoproteins, and cardiovascular risk prediction. Circulation. 2008;l 18(20):2047-56. Diao JA, Inker LA, Levey AS, Tighiouart H, Powe NR, Mamai AK. In Search of a Better Equation - Performance and Equity in Estimates of Kidney Function. N Engl J Med.202I;384(5):396-9. Levey AS, Stevens LA, Schmid CH, Zhang YL, Castro AF, 3rd, Feldman HI, et al. A new equation to estimate glomerular filtration rate. Ann Intern Med. 2009;150(9):604-12.Home BD, Anderson JL, John JM, Weaver A, Bair TL, Jensen KR, et al. Which white blood cell subtypes predict increased cardiovascular risk? J Am Coll Cardiol. 2005;45(10): 1638-43. Lee G, Choi S, Kim K, Yun JM, Son JS, Jeong SM, et al. Association of Hemoglobin Concentration and Its Change With Cardiovascular and All -Cause Mortality. J Am Heart Assoc. 2018;7(3). Samak MJ, Tighiouart H, Manjunath G, MacLeod B, Griffith J, Salem D, et al. Anemia as a risk factor for cardiovascular disease in the atherosclerosis risk in communities (aric) study. Journal of the American College of Cardiology. 2002;40(l):27-33. Lassale C, Curtis A, Abete I, van der Schouw YT, Verschuren WMM, Lu Y, et al.Elements of the complete blood count associated with cardiovascular disease incidence: Findings from the EPIC-NL cohort study. Sci Rep. 2018;8(l):3290. Austin PC, Daly PA, Tu JV. A multicenter study of the coding accuracy of hospital discharge administrative data for patients admitted to cardiac care units in Ontario. Am Heart J. 2002;144(2):290-6. Lee DS, Stitt A, Wang X, Yu JS, Gurevich Y, Kingsbury KJ, et al. Administrative hospitalization database validation of cardiac procedure codes. Med Care. 2013;51(4):e22-6. Porter J, Mondor L, Kapral MK, Fang J, Hall RE. How Reliable Are Administrative Data for Capturing Stroke Patients and Their Care. Cerebrovasc Dis Extra. 2016;6(3):96-106. Austin PC, Fine JP. Practical recommendations for reporting Fine-Gray model analyses for competing risk data. Stat Med. 2017;36(27):4391-400.Harrell FE. Regression Modeling Strategies: Springer-Verlag; 2006. Navar AM, Pencina MJ, Mulder H, Elias P, Peterson ED. Improving patient risk communication: Translating cardiovascular risk into standardized risk percentiles. Am Heart J. 2018;198: 18-24. Wolbers M, Blanche P, Koller MT, Witteman JC, Gerds TA. Concordance for prognostic models with competing risks. Biostatistics. 2014;15(3):526-39. Van Calster B, McLernon DJ, van Smeden M, Wynants L, Steyerberg EW, Topic Group 'Evaluating diagnostic t, et al. Calibration: the Achilles heel of predictive analytics. BMC Med. 2019;17(l):230. Vickers AJ, Van Calster B, Steyerberg EW. Net benefit approaches to the evaluation of prediction models, molecular markers, and diagnostic tests. BMJ. 2016;352:i6. Pastore M, Calcagni A. Measuring Distribution Similarities Between Samples: A Distribution-Free Overlapping Index. Front Psychol. 2019;10: 1089. Riley RD, Ensor J, Snell KI, Debray TP, Altman DG, Moons KG, et al. External validation of clinical prediction models using big datasets from e -health records or IPD meta-analysis: opportunities and challenges. BMJ. 2016;353:i3140. Gerds TA, Schumacher M. Consistent estimation of the expected Brier score in general survival models with right-censored event times. Biom J. 2006;48(6): 1029-40. Graf E, Schmoor C, Sauerbrei W, Schumacher M. Assessment and comparison of prognostic classification schemes for survival data. Statistics in Medicine. 1999; 18(17- 18):2529-45.Muntner P, Colantonio LD, Cushman M, Goff DC, Jr., Howard G, Howard VJ, et al. Validation of the atherosclerotic cardiovascular disease Pooled Cohort risk equations.JAMA. 2014;311(14): 1406-15. Gerds TA, Cai T, Schumacher M. The performance of risk prediction models. Biom J. 2008;50(4):457-79. Cook NR, Ridker PM. Calibration of the Pooled Cohort Equations for Atherosclerotic Cardiovascular Disease: An Update. Ann Intern Med. 2016; 165(11):786-94. Van Calster B, Vickers AJ. Calibration of risk prediction models: impact on decision- analytic performance. Med Decis Making. 2015;35(2): 162-9. Hobbs FD, Jukema JW, Da Silva PM, McCormack T, Catapano AL. Barriers to cardiovascular disease risk scoring and primary prevention in Europe. QJM. 2010;103(10):727-39. Eaton CB, Galliher JM, McBride PE, Bonham AJ, Kappus JA, Hickner J. Family physician's knowledge, beliefs, and self-reported practice patterns regarding hyperlipidemia: a National Research Network (NRN) survey. J Am Board Fam Med. 2006; 19(l):46-53. White-Al Habeeb NMA, Higgins V, Venner AA, Bailey D, Beriault DR, Collier C, et al. Canadian Society of Clinical Chemists Harmonized Clinical Laboratory Lipid Reporting Recommendations on the Basis of the 2021 Canadian Cardiovascular Society Lipid Guidelines. Can J Cardiol. 2022;38(8): 1180-8.Naugler C, Cook C, Morrin L, Wesenberg J, Venner AA, Campbell N, et al. Statin Prescriptions for High-Risk Patients Are Increased by Laboratory -Initiated Framingham Risk Scores: A Quality-Improvement Initiative. Can J Cardiol. 2017;33(5):682-4. Rana JS, Tabada GH, Solomon MD, Lo JC, Jaffe MG, Sung SH, et al. Accuracy of the Atherosclerotic Cardiovascular Risk Equation in a Large Contemporary, Multiethnic Population. J Am Coll Cardiol. 2016;67(18):2118-30. Saeed A, Nambi V, Sun W, Virani SS, Taffet GE, Deswal A, et al. Short-Term Global Cardiovascular Disease Risk Prediction in Older Adults. J Am Coll Cardiol. 2018;71(22):2527-36. Akintoye E, Briasoulis A, Afonso L. Biochemical risk markers and 10-year incidence of atherosclerotic cardiovascular disease: independent predictors, improvement in pooled cohort equation, and risk reclassification. Am Heart J. 2017;193:95-103. Home BD, May HT, Muhlestein JB, Ronnow BS, Lappe DL, Renlund DG, et al. Exceptional mortality prediction by risk scores from common laboratory tests. Am J Med. 2009; 122(6):550-8. Truslow JG, Goto S, Homilius M, Mow C, Higgins JM, MacRae CA, et al. Cardiovascular Risk Assessment Using Artificial Intelligence -Enabled Event Adjudication and Hematologic Predictors. Circ Cardiovasc Qual Outcomes. 2022;15(6):e008007. Sud M, Chu A, Austin PC, Naimark DJ, Thanassoulis G, Wijeysundera HC, et al. Impact of Outcome Definitions on Cardiovascular Risk Prediction in a Contemporary Primary Prevention Population. Eur Heart J Qual Care Clin Outcomes. 2022.group Sw, collaboration ESCCr. SCORE2 risk prediction algorithms: new models to estimate 10-year risk of cardiovascular disease in Europe. Eur Heart J. 2021;42(25):2439- 54. Goff DC, Jr., Lloyd-Jones DM, Bennett G, Coady S, D'Agostino RB, Gibbons R, et al. 2013 ACC / AHA guideline on the assessment of cardiovascular risk: a report of the American College of Cardiology / American Heart Association Task Force on Practice Guidelines. Circulation. 2014;129(25 Suppl 2):S49-73. Packard C, Chapman MJ, Sibartie M, Laufs U, Masana L. Intensive low-density lipoprotein cholesterol lowering in cardiovascular disease prevention: opportunities and challenges. Heart. 2021; 107(17): 1369-75. Lloyd- Jones DM, Goff DC, Jr. Need for Better Methodology in Assessing Pooled Cohort Equations. J Am Coll Cardiol. 2017;69(3):365-6. Vinogradova Y, Coupland C, Brindle P, Hippisley-Cox J. Discontinuation and restarting in patients on statin treatment: prospective open cohort study using a primary care database. BMJ. 2016;353 :i3305. Wong ND, Young D, Zhao Y, Nguyen H, Caballes J, Khan I, et al. Prevalence of the American College of Cardiology / American Heart Association statin eligibility groups, statin use, and low-density lipoprotein cholesterol control in US adults using the National Health and Nutrition Examination Survey 2011-2012. J Clin Lipidol. 2016; 10(5): 1109- 18. Grundy SM, Stone NJ, Bailey AL, Beam C, Birtcher KK, Blumenthal RS, et al. 2018 AHA / ACC / AACVPR / AAPA / ABC / ACPM / ADA / AGS / APhA / ASPC / NLA / PCNAGuideline on the Management of Blood Cholesterol: A Report of the American College of Cardiology / American Heart Association Task Force on Clinical Practice Guidelines. J Am Coll Cardiol. 2018. Efficacy and safety of more intensive lowering of LDL cholesterol: a meta-analysis of data from 170 000 participants in 26 randomised trials. The Lancet.2010;376(9753): 1670-81. Ueda P, Woodward M, Lu Y, Hajifathalian K, Al-Wotayan R, Aguilar-Salinas CA, et al. Laboratory-based and office-based risk scores and charts to predict 10-year risk of cardiovascular disease in 182 countries: a pooled analysis of prospective cohorts and health surveys. The Lancet Diabetes & Endocrinology. 2017;5(3): 196-213. Willeit P, Kiechl S, Kronenberg F, Witztum JL, Santer P, Mayr M, et al. Discrimination and net reclassification of cardiovascular risk with lipoprotein(a): prospective 15 -year outcomes in the Bruneck Study. J Am Coll Cardiol. 2014;64(9):851-60. group SOw, collaboration ESCCr. SCORE2-OP risk prediction algorithms: estimating incident cardiovascular event risk in older persons in four geographical risk regions. Eur Heart J. 2021;42(25):2455-67.
Claims
CLAIMS:
1. A computer-implemented method for predicting risk of a cardiovascular event, the method comprising: generating an input data set from cardiovascular event data, wherein the cardiovascular event data is acquired from a plurality of patients; generating a trained predictive model, based on patient age and at least one predefined predictor variable; receiving new subject patient data; with the trained predictive model, estimating the risk of the cardiovascular event associated with the subject patient data based on the age of the new subject patient and at least one predefined predictor variable associated with the new subject patient; and generating a report comprising a cardiovascular risk assessment.
2. The method of claim 1, wherein the at least one predefined predictor variable value associated with the patient is obtained from a biological sample.
3. The method of claim 1 or 2, wherein the at least one predefined predictor variable comprises at least one of serum total cholesterol, high-density lipoprotein cholesterol, triglycerides, hemoglobin, mean corpuscular volume, platelets, leukocytes, monocytes, neutrophils, lymphocytes, estimated glomerular filtration rate, and glucose.
4. The method of claim 1, wherein the at least one predefined predictor variable is modelled using at least one of restricted cubic splines, linear splines, b-splines, fractional polynomials.
6. The method of claim 5, wherein predicting the risk of the cardiovascular event is integrated into a clinical workflow.
7. The method of any one of claims 1 to 6, wherein the cardiovascular event is at least one of an acute coronary syndrome, stable or unstable angina, arterial revascularization, stroke, transient ischemic attack, peripheral arterial disease, atherosclerotic cardiovascular disease, heart failure, atrial fibrillation, myocardial infarction, and sudden coronary death.
8. The method of any one of claims 1 to 7, wherein the predictive model comprises at least one of machine learning algorithms selected from a group comprising support vector machine, decision tree, linear regression, logistic regression, naive Bayes, k- nearest neighbor, random forest, AdaBoost, XGBoost and neural network methods.
9. A system for determining the cardiovascular health of a patient, the system comprising: a computer system comprising a hardware processor and a memory device on which instructions are encoded to cause the hardware processor to perform the operations of: generating an input data set from cardiovascular event data, wherein the cardiovascular event data is acquired from a plurality of patients; generating a trained predictive model, based on age and at least one predefined predictor variable; receiving new subject patient data; applying the trained predictive model on the new subject patient data to predict cardiovascular risk; and generating a report comprising a cardiovascular risk assessment.
10. The system of claim 9, wherein the predictive model is trained and optimized using data acquired from a plurality of patients and evaluated on a test set of patients.
11. The system of claim 10, wherein the at least one predefined predictor variable comprises at least one of serum total cholesterol, high-density lipoprotein cholesterol, triglycerides, hemoglobin, mean corpuscular volume, platelets, leukocytes monocytes, neutrophils, lymphocytes, estimated glomerular filtration rate, and glucose.
12. The system of claim 11, wherein predicting risk is integrated into a clinical workflow.
13. The system of any one of claims 9 to 12, wherein the cardiovascular event is at least one of an acute coronary syndrome, stable or unstable angina, arterial revascularization, stroke, transient ischemic attack, peripheral arterial disease, atherosclerotic cardiovascular disease, heart failure, atrial fibrillation, myocardial infarction, and sudden coronary death.
14. The system of any one of claims 9 to 13, wherein the predictive model comprises at least one of machine learning algorithms selected from a group comprising support vector machine, decision tree, linear regression, logistic regression, naive Bayes, ^-nearest neighbor, random forest, AdaBoost, XGBoost, and neural network methods.
15. A system for generating a trained predictive model for predicting a cardiovascular event, the system comprising: a computer system comprising a hardware processor and a memory device on which instructions are encoded to cause the hardware processor to perform the operations of: receiving patient data comprising at least one of demographic data, cardiovascular risk data, and laboratory data acquired from a plurality of patients; preprocessing the patient data;extracting features from the datasets,; selecting features for subsequent use during a training phase, wherein the selected features comprise age and at least one predefined predictor variable for predicting the cardiovascular event; selecting a predictive model and iteratively training the predictive model with the training data sets and validating the predictive model with the testing data sets to generate a trained predictive model; and storing the model on the memory device and / or outputting the model for future use on new subject patient data.
16. The system of claim 15, wherein the trained predictive model predicts the cardiovascular event from the new subject patient data based on the age of the patient and at least one predefined predictor variable value associated with the patient.
17. The system of claim 16, wherein the at least one predefined predictor variable value associated with the patient is obtained from a biological sample.
18. The system of claim 17, wherein the biological sample is at least one of a blood product, urine.
19. The system of any one of claims 15 to 18, wherein the at least one predefined predictor variable comprises at least one of serum total cholesterol, high-density lipoprotein cholesterol, triglycerides, hemoglobin, mean corpuscular volume, platelets, leukocytes, monocytes, neutrophils, lymphocytes, estimated glomerular filtration rate, and glucose.
20. The system of any one of claims 15 to 19, wherein the cardiovascular event is at least one of an acute coronary syndrome, stable or unstable angina, arterial revascularization, stroke, transient ischemic attack, peripheral arterial disease,atherosclerotic cardiovascular disease, heart failure, atrial fibrillation, myocardial infarction, and sudden coronary death.
21. The system of any one of claims 15 to 20, wherein the trained predictive model comprises at least one of machine learning algorithms selected from a group comprising support vector machine, decision tree, linear regression, logistic regression, naive Bayes, ^-nearest neighbor, random forest, AdaBoost, XGBoost, and neural network methods.
22. The system of any one of claims 15 to 21, further comprising a step of partitioning the dataset into training data sets and testing data sets using predefined splits.
23. A computer-readable medium comprising instructions stored thereon executable by a hardware processor to perform the operations of: receiving subject patient data comprising at least one of demographic data, cardiovascular risk data, and laboratory data; preprocessing the patient data; selecting features from the subject patient data, wherein the selected features comprise age and at least one predefined predictor variable for predicting the cardiovascular event; and applying a trained predictive model trained using the selected features on the new subject patient data to predict a cardiovascular risk.
Citation Information
Patent Citations
Healthcare Information Technology System for Predicting Development of Cardiovascular Conditions
US20110202486A1