Systems and methods for predicting heart disease events
Patent Information
- Application Number
- PCT/CA2026/050247
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2025-02-18
- Filing Date
- 2026-02-17
- Publication Date
- 2026-08-27
Smart Images

Figure CA2026050247_27082026_PF_FP_ABST
Abstract
Description
SYSTEMS AND METHODS FOR PREDICTING HEART DISEASE EVENTS CROSS-REFERENCE TO RELATED APPLICATIONS
[0001] This application claims all benefit including priority to U.S. Provisional Patent Application 63 / 759,824, filed February 18, 2025, and entitled “SYSTEMS AND METHODS FOR PREDICTING HEART DISEASE EVENTS”; the entire contents of which are hereby incorporated by reference herein.FIELD
[0002] This disclosure relates to heart disease, and more specifically relates to automated prediction of heart disease events.BACKGROUND
[0003] Heart failure (HF) is a global health crisis affecting an estimated 64.3 million individuals worldwide. The annual global cost of HF is estimated at $346 billion, with a substantial portion attributed to hospitalizations, missed work, and healthcare service utilization. HF is an episodic disease that significantly reduces life expectancy, with a median survival of 3.2 years for women and 1.7 years for men. Periods of stability are frequently disrupted by acute exacerbations requiring intervention. Despite recent medical advances, HF patients continue to experience a high risk of adverse outcomes, indicating a critical need for better prognostic markers to enhance risk stratification and enable more effective and timely interventions.
[0004] Risk evaluation in HF relies on intermittent and static clinical evaluations, making accurate and timely risk stratification a challenge. Many cardiac centers use cardiopulmonary exercise testing (CPET) to estimate 1-year prognosis, enabling stratification of patients based on established exercise parameters. However, while serial CPET trends are informative, its high cost, limited geographic accessibility, and significant patient burden make daily or even weekly monitoring impractical. The six-minute walk test (6MWT) offers an alternative measurement of exercise intolerance in HF patients. However, its prognostic value compared to CPET remains uncertain, and the high variability of unsupervised 6MWT limits its widespread use in free-living remotemonitoring. The New York Heart Association (NYHA) functional classification is still commonly used in practice but is subjective and fails to capture a patient's fluctuating clinical status. Assessment of biomarkers such as N-terminal pro-B-type Natriuretic Peptide (NT-proBNP) provides prognostic value, but remains too static. Thus, while these traditional tools for HF risk stratification provide essential insights into patient health, they cannot also account for the episodic progression of HF.
[0005] Remote patient monitoring (RPM) modalities move away from static and intermittent measurements of patient status to daily measurements and have shown up to a 15% reduction in total hospitalization. The effectiveness of RPM has been demonstrated in studies through improved patient outcomes and reduced healthcare utilization when these systems are proactive and correctly implemented. Nonetheless, the demands of manual data entry for patients and the constant oversight by clinicians hinder adherence and broader implementation of these programs.
[0006] Thus, there is a need for improved or alternative solutions that address one or more of the limitations of existing solutions.SUMMARY
[0007] In accordance with an aspect, there is provided a computer-implemented system for predicting heart disease events, the system comprising: a processing subsystem that includes one or more processors and one or more memories coupled with the one or more processors, the processing subsystem configured to cause the system to: collect biometric data of a particular user over a collection time period spanning a plurality of days, the biometric data including photoplethysmography sensor data; generate biometric features based on the collected biometric data; supplement the biometric features with clinical features for the particular user; provide, to a machine learning model trained to predict heart disease events, the biometric features as supplemented by the clinical features; and receive, from the machine learning model, an output predictive of a heart disease event.
[0008] In some embodiments, the biometric data is collected via a wearable device.
[0009] In some embodiments, the output includes a metric of VO2 regression.
[0010] In some embodiments, the collection time period includes at least 10 days.
[0011] In some embodiments, the collection time period includes at least 30 days.
[0012] In some embodiments, the biometric features provided to the machine learning model reflect biometric data spanning the collection time period.
[0013] In some embodiments, the clinical features include demographic features and / or drug-use features of the particular user.
[0014] In some embodiments, the demographic features include at least one of a sex, an age, or an ethnicity of the particular user.
[0015] In some embodiments, the machine learning model includes a sequence of transformer blocks, each of the transformer blocks along the sequence configured to generate representations of the biometric data in decreasing temporal resolutions and configured for causal-self attention to provide temporal causality across the temporal resolutions.
[0016] In some embodiments, the temporal resolutions decrease by approximately half for each of the transformer blocks along the sequence.
[0017] In some embodiments, the clinical features are provided to at least two of the transformer blocks in the sequence of transformer blocks.
[0018] In some embodiments, the prediction is for a prediction time period subsequent to the collection time period.
[0019] In some embodiments, the biometric data includes measurements of at least one of step count, exercise time, distance traveled, stand time, active energy burned, basal energy burned, heart rate, heart rate variability, and 02 saturation.
[0020] In some embodiments, the biometric features include at least one of sub-daily, daily, or multi-day aggregations.
[0021] In accordance with another aspect, there is provided a computer-implemented method for predicting heart disease events, the method comprising: collecting biometric data of a particular user over a collection time period spanning a plurality of days, the biometric data including photoplethysmography sensor data; generating biometric features based on the collected biometric data; supplementing the biometric features with clinical features for the particular user; providing, to a machine learning model trained to predict heart disease events, the biometric features as supplemented by the clinical features; and receiving, from the machine learning model, an output predictive of a heart disease event.
[0022] In some embodiments, the biometric data is collected via a wearable device.
[0023] In some embodiments, the output includes a metric of VO2 regression.
[0024] In some embodiments, the collection time period includes at least 10 days.
[0025] In some embodiments, the collection time period includes at least 30 days.
[0026] In some embodiments, the biometric features provided to the machine learning model reflect biometric data spanning the collection time period.
[0027] In some embodiments, the clinical features include demographic features and / or drug-use features of the particular user.
[0028] In some embodiments, the demographic features include at least one of a sex, an age, or an ethnicity of the particular user.
[0029] In some embodiments, the machine learning model includes a sequence of transformer blocks, each of the transformer blocks along the sequence configured to generate representations of the biometric data in decreasing temporal resolutions and configured for causal-self attention to provide temporal causality across the temporal resolutions.
[0030] In some embodiments, the temporal resolutions decrease by approximately half for each of the transformer blocks along the sequence.
[0031] In some embodiments, the clinical features are provided to at least two of the transformer blocks in the sequence of transformer blocks.
[0032] In some embodiments, the prediction is for a prediction time period subsequent to the collection time period.
[0033] In some embodiments, the biometric data includes measurements of at least one of step count, exercise time, distance traveled, stand time, active energy burned, basal energy burned, heart rate, heart rate variability, and 02 saturation.
[0034] In some embodiments, the biometric features include at least one of sub-daily, daily, or multi-day aggregations.
[0035] In accordance with another aspect, there is provided a non-transitory computer-readable medium or media having stored thereon machine interpretable instructions which, when executed by a processing system, cause the processing system to perform a method for predicting heart disease events, the method comprising: collecting biometric data of a particular user over a collection time period spanning a plurality of days, the biometric data including photoplethysmography sensor data; generating biometric features based on the collected biometric data; supplementing the biometric features with clinical features for the particular user; providing, to a machine learning model trained to predict heart disease events, the biometric features as supplemented by the clinical features; and receiving, from the machine learning model, an output predictive of a heart disease event.
[0036] Many further features and combinations thereof concerning embodiments described herein will appear to those skilled in the art following a reading of the instant disclosure.BRIEF DESCRIPTION OF THE DRAWINGS
[0037] In the figures,
[0038] FIG. 1 is a schematic diagram of a system for predicting heart disease events, in accordance with an embodiment;
[0039] FIG. 2 depicts a machine learning model architecture, in accordance with an embodiment;
[0040] FIG. 3A and FIG. 3B depict machine learning model training and inference, respectively, in accordance with an embodiment;
[0041] FIG. 4 is a flowchart showing a computer-implemented method for predicting heart disease events, in accordance with an embodiment;
[0042] FIG. 5 is a schematic diagram of a computing device, in accordance with an embodiment;
[0043] FIG. 6A, FIG. 6B, FIG. 6C, FIG. 6D and FIG. 6E show aspects of the study design validating the TRUE-HF model, in accordance with an embodiment;
[0044] FIG. 7A, FIG. 7B, and FIG. 7C show the results of using biometric data to estimate cardiorespiratory fitness, in accordance with an embodiment;
[0045] FIG. 8A, FIG. 8B, FIG. 8C, FIG. 8D and FIG. 8E show TRUE-HF model and a reduced sensor variant (TRUE-HF-RS) predicting decline in daily pVO2prior to unplanned health care use, in accordance with an embodiment;
[0046] FIG. 9 shows a time-dependent performance analysis of the TRUE-HF-RS Model, in accordance with an embodiment;
[0047] FIG. 10A, FIG. 10B, and FIG. 10C show the prediction of 6MWTD, in accordance with an embodiment;
[0048] FIG. 11A and FIG. 11B show the prediction of unplanned healthcare utilization, in accordance with an embodiment;
[0049] FIG. 12A and FIG. 12B show model robustness to the removal of unsupervised exercise data, in accordance with an embodiment; and
[0050] FIG. 13A and FIG. 13B show saliency values from features of the model in accordance with an embodiment.
[0051] These drawings depict exemplary embodiments for illustrative purposes, and variations, alternative configurations, alternative components and modifications may be made to these exemplary embodiments.DETAILED DESCRIPTION
[0052] Disclosed herein are systems and methods for predicting heart disease events. Embodiments of such systems and methods may utilize a combination of biometric data and clinical data for a particular user. Biometric data may be collected for the particular user over a collection time period spanning a plurality of days (e.g., 10 days, 30 days, etc.). Biometric data may be collected using one or more wearable devices. Conveniently, the use of wearable devices enables free-living remote monitoring of the particular user. The collected biometric data may be supplemented with clinical features for the particular user. The biometric features (as supplemented by the clinical features) may be provided to a machine learning model trained to predict heart disease events. An output predictive of a heart disease event may be received from the machine learning model. The term “heart disease event” may be referred to as “heart failure event” or “unplanned healthcare utilization” throughout this disclosure.
[0053] Some embodiments of the disclosed systems and methods may be used for predicting a user’s cardiopulmonary fitness and changes in their fitness over time.
[0054] Some embodiments of the disclosed systems and methods may be suitable for remote patient management (RPM). RPM has significant potential applicability for on-demand healthcare, aiming to mitigate hospitalizations and enhance survival among patients with heart failure (HF). Historically, RPM has focused on summarizing vital signs and daily weights; however, these measures have not significantly decreased hospitalization risk or improved survival, as adherence to RPM low, and weight gain often appears too late to serve as an effective early warning sign.
[0055] To address some of these limitations, some embodiments of the disclosed systems and methods implement a novel machine learning model configured to predict heart disease events. In some embodiments, this machine learning model is an autoregressive attention machine learning model. Embodiments of such machinelearning model may be referred to herein as a “TRUE-HF” model. In some embodiments, this machine learning model is configured to predict daily changes in pVO2and detect early signs of HF deterioration.
[0056] VO2regression, representing a decline in a patient’s estimated or measured oxygen consumption capacity, may serve as a indicator of worsening cardiac function and emerging hemodynamic instability. In heart failure and related cardiovascular conditions, reductions in peak VO2reflect diminished cardiac output or impaired peripheral oxygen extraction, both of which are physiologic changes that can precede symptomatic deterioration. Because peak VO2captures the integrative performance of cardiac, pulmonary, and peripheral muscular systems, downward trends in daily or near-daily VO2estimates may provide early evidence of decompensation before overt clinical signs appear.
[0057] FIG. 1 is a schematic diagram of a prediction system 100 for predicting heart disease events, in accordance with an embodiment.
[0058] Prediction system 100 is configured for electronic communication with a wearable device 50 to receive biometric data therefrom. Biometric data may include data reflecting of various sensor measurements (and derivatives) obtained at wearable device. For example, such biometric data may include data obtained from a photoplethysmography sensor at wearable device 50. In some embodiments, the biometric data may include data obtained from one or more other sensors at wearable device 50 such as, e.g., on-device electrocardiogram, accelerometer, gyroscope, barometer, GPS, ambient light sensor, altimeter, compass, temperature sensor, or the like. In some embodiments, the biometric data may include data derived from one or more sensor measurements. Examples of biometric data may include, for example, energy burned, heart rate, distance traversed, step count, exercise time, idle time, oxygen saturation, and heart rate variability. In some embodiments, biometric data may be obtained from a device-specific API such as, e.g., Apple HealthKit or Android Health Connect.
[0059] In the depicted embodiment, wearable device 50 is a smart watch. The smart watch may, for example, be an Apple Watch. The smart watch may be another type of smart watch such as from another vendor (e.g., Samsung, Google, or the like, and be configured for interoperability with mobile devices). In some embodiments, wearable device 50 may be another type of device in the form of a ring, a band, a patch, or the like. In some embodiments , wearable device 50 may have a relatively reduced sensor suite, such as a FitBit, providing only heart rate and step count data to the prediction system 100.
[0060] Prediction system 100 and wearable device 50 may exchange electronic communication by way of a communication link 70. Communication link 70 may be a wireless link, e.g., Bluetooth, Bluetooth Low Energy, Near-field Communication (NFC), or the like. In some embodiments, communication link 70 may be a wired link suitable for transmitting data. In some embodiments, such wired link may also serve to supply power to wearable device 50.
[0061] In the depicted embodiment, prediction system 100 includes a biometric data interface 102, a clinical data interface 104, a prediction engine 106, and an electronic datastore 108.
[0062] Biometric data interface 102 is configured to receive biometric data from wearable device 50, e.g., by way of communication link 70. The biometric data may include measurements of at least one of step count, exercise time, distance traveled, stand time, active energy burned, basal energy burned, heart rate, heart rate variability, 02 saturation, or underlying measurements indicative of any of the above.
[0063] In some embodiments, biometric data interface 102 may apply pre-processing, e.g., to decode the data, reformat the data, normalize the data, filter the data, remove outlying data, or the like. Biometric data interface 102 stores the biometric data (as optionally pre-processed) in electronic datastore 108.
[0064] Clinical data interface 104 is configured to receive clinical data regarding the user who wears wearable device 50. Such clinical data may correspond, for example, todemographic features or drug-use features of that user. The demographic features may include, for example, at least one of a sex, an age, an ethnicity, a weight, a height of the particular user. The drug-user features may include the types of drugs, dosage, etc. taken by the particular user. Clinical data interface 104 stores the clinical data in electronic datastore 108.
[0065] In some embodiments, clinical data interface 104 may receive at least some clinical data by way of user input, e.g., from a user interface presented at prediction system 100. In some embodiments, clinical data interface 104 may receive at least some clinical data from an electronic medical record system, e.g., by way of a network connection.
[0066] Prediction engine 106 is configured to make predictions of heart failure events for a user. As detailed here, such predictions may be based on a combination of biometric data and clinical data for that user. Prediction engine 106 includes a machine learning model that is trained to make such predictions.
[0067] In some embodiments, the prediction system 100 constructs a temporal feature representation that encapsulates wearable-derived biometric features across a multi-day window and aligned clinical features for integrated analysis by the prediction engine 106 and underlying machine learning models. In some embodiments, the temporal feature representation is an ordered, temporally indexed data structure comprising a sequence of time-bucket feature vectors and a clinical feature vector. Each time-bucket feature vector summarises biometric measurements acquired by the wearable device 50 within a fixed sub-daily interval (e.g., every 90 minutes) using summary data suitable for the measurement type. For continuous physiological signals such as heart rate, heart rate variability and oxygen saturation, statistical descriptors such as mean, median, minimum, maximum, and standard deviation within the interval may be used. For activity-derived measurements such as step count, exercise time, distance travelled, stand time, active energy burned, or basal energy burned, interval sums may be used to capture accumulated activity within the interval. The temporal feature representation aggregates the time-bucket feature vectors into dailyrepresentations and further into a sliding multi-day window. In some embodiments, the window comprises at least ten days and, in further embodiments, at least thirty days.
[0068] In some embodiments, the temporal feature representation includes, in addition to sub-daily feature vectors, pre-computed daily aggregates and multi-day trend descriptors, such as rolling means, rolling standard deviations, or linear trend coefficients over predetermined horizons, which enable the prediction engine 106 to utilize both fine-grained and coarse-grained temporal information.
[0069] In some embodiments, the temporal feature representation includes metadata that specifies one or more of the temporal boundaries of the window, the sub-daily bucket duration, and the ordering of time-bucket feature vectors, thereby allowing the prediction engine 106 and any underlying model to preserve temporal dependencies. In some embodiments, the temporal feature representation also includes flags or quality indicators for time-bucket feature vectors to reflect missing or censored data, for example where forward-filling imputation has been applied to maintain autoregressive integrity of the series.
[0070] The temporal feature representation may be augmented by the clinical data interface 104 with a clinical feature vector representing user-specific context that modulates the interpretation of biometric features. Clinical features may include demographic characteristics such as sex, age and ethnicity, anthropometric attributes such as weight and height, and medication-related information including drug class and dosage. In some embodiments, the clinical feature vector is stored separately from, but associated with, the time-bucket feature sequence by a common user identifier and timestamp metadata. In some embodiments, to preserve temporal causality, only baseline clinical features or features known not to vary over the prediction horizon are included in the clinical feature vector provided to the model for a given longer-trend window.
[0071] In the depicted embodiment, the machine learning model is a contextualized deep learning model configured to retain and analyze temporal trends across 30-days of biometric data of a user, as received from wearable device 50. In this embodiment, themachine learning model is configured to (i) construct increasingly comprehensive temporal representations of the data using a bottom-up approach, extending from 90-minute intervals to full-day aggregation, (ii) integrate user-specific clinical data directly into the biometric features learned by the deep learning model through affine transformation, allowing for adaptive feature calculations, such as, e.g., modulating features for users on beta-blockers or other drugs, and; (iii) explicitly consider relevant temporal constraints, e.g., recognizing that daily activities are influenced by preceding days not vice versa, and uses this to make ongoing predictions for each day.
[0072] In some embodiments, the machine learning model may be operated to provide near-continuous daily monitoring through dynamic next-day predictions.
[0073] The machine learning model is configured to process 30 days of 90-minute summaries of collected biometric data, starting with sequential 90-minute summaries and assembling them into progressively larger time windows, allowing the model to learn temporal relationships at different temporal resolutions while improving processing efficiency. This approach is embodied in the model architecture depicted in FIG. 2 for the TRUE-HF model, which optimizes the feature map and reduces the temporal resolution through pooling, in accordance with an embodiment. In some embodiments, another model architecture for self-attention-based models may be used. In some embodiments, the model architecture may be a convolutional neural network architecture, a feed forward neural network architecture, a recurrent neural network architecture, or the like.
[0074] As shown in FIG. 2, in an embodiment, the TRUE-HF model includes a plurality of transformer blocks 120 configured to learn temporal relationships in the extended data sequence and reduce the representations and temporal resolutions (more oversized windows). As depicted, the length of time is reduced deeper in the model through pooling.
[0075] The TRUE-HF model integrates demographic, drug and other clinical features data with time-series data through affine transformation, where these clinical features can affect the interpretation of biometric features. As shown, clinical features areprovided to at least two of transformer blocks 120. In some embodiments, clinical features are provided to each transformer block 120. The TRUE-HF model includes a linear layer configured to generate a prediction of outcomes (e.g., peak VO2regression or other desired outcomes).
[0076] During operation, biometric features are generated based on the collected biometric data. For example, biometric data are first tokenized through 1 -dimensional convolutions to collapse features along the temporal domain in the TRUE-HF model. The input is then processed through a transformer block 120 that includes a transformer layer followed by a pooling layer. The transformer layer learns relationships within the temporal resolution in each transformer block 120. Subsequently, the pooling layer aggregates pairs of consecutive time points, merging data across the time domain (e.g., 90 to 180 minutes).
[0077] In some embodiments, two or more transformer blocks 120 operate in series. In one specific embodiment, a sequence of four transformer blocks 120 is used to analyze time resolutions of 90 minutes, 180 minutes, 360 minutes, 720 minutes, and 1,440 minutes (daily resolution). The final prediction layer then aggregates information across the previous 30 days, inclusive, to predict a current day’s measurement. The prediction may be, for example, VO2regression.
[0078] To enhance the model's understanding of the collected biometric data, the TRUE-HF model incorporates user-specific clinical data (described above), allowing the model to learn different operations based on various input attributes rather than treating all inputs equally. This is achieved by including within transformer block 120 a featurewise linear modulation (FiLM) block that modulates activations in the neural networks based on clinical features. Only demographic information from the baseline clinical visit was used by the model to maintain temporal causality, which allows the model to learn different features and temporal relationships depending on user-specific medications, thereby calibrating the model to each user.
[0079] FIG. 3A and FIG. 3B depict a machine learning model training framework 300 and an inference framework 350, respectively, in accordance with an embodiment. Asshown in FIG. 3A and FIG. 3B, the temporal learning is constrained to move forward only ( / .e. autocorrelation) in the TRUE-HF model architecture by incorporating a causal self-attention mechanism. This mechanism introduces auto-regressive properties into temporal learning, enhancing the model's predictive capabilities regarding daily clinical outcomes. The architecture provides that predictions for any given day are influenced only by data from that day and preceding days, simulating a realistic clinical scenario where future information is inherently unknown. In some embodiments of the TRUE-HF model, causal self-attention is integrated into each transformer layer, ensuring temporal causality within and across each time resolution.
[0080] Casual predictions act temporally forwardly, where day 1 predicts day 2, day 2 predicts day 3, etc. Only the last prediction (next-day prediction) is considered during inference.
[0081] Self-supervised learning is leveraged in the training framework 300 based on linear imputed outcomes of CPET pVO2based on onboard to three-month offboard values. This approach learns directly from the data and is applicable when dealing with extensive datasets where only a portion possesses ground truth labels. In some data sets, explicit daily labels are absent, and only baseline and three-month offboard clinical visits provided clinical status for particular target outcomes (e.g., CPET pVO2or clinical 6MWT). Consequently, in the training data, three months are provided between the baseline and follow-up visit, and each user has a minimum of 88 unlabeled days for every two labeled days. The additional 88 unlabeled days are used for training purposes. By using the additional 88 unlabeled days only for model training purposes, the amount of training data is enhanced to improve the model's predictive capability without compromising the integrity of model evaluation.
[0082] The TRUE-HF model utilizes linear interpolation of clinical outcomes recorded during the initial baseline assessment and subsequent follow-up visits, thereby establishing daily outcomes. These provide, on average, a 44-fold increase in the training data for the model per user.
[0083] As noted above, the TRUE-HF model processes 30-day input windows and outputs 30 daily tokens, each preserving autoregressive properties. By leveraging this output, the causal mechanism of the model is used to enforce autoregressive properties, where each predicted token informs the next. This self-supervised setup supports the model in predicting the daily interpolated outcomes on the tokens ( / .e., each of the 30 tokens from the TRUE-HF model predicts its corresponding next-day outcome), enabling prospective daily prediction in patients as opposed to a single prediction after an extended period of time (e.g., once every three-months).
[0084] Each of biometric data interface 102, clinical data interface 104, and prediction engine 106 may be implemented using a suitable combination of software and hardware components. In some embodiments, such software components may be implemented in whole or in part using conventional programming languages such as Java, J#, C, C++, C#, Perl, Python, Visual Basic, Ruby, Scala, etc. Such software components of system 100 may be in the form of one or more executable programs, scripts, routines, statically / dynamically linkable libraries, or servlets.
[0085] Electronic datastore 108 may include a combination of non-volatile and volatile memory. In some embodiments, electronic datastore 108 stores collected biometric data and / or clinical data for a particular user. In some embodiments, electronic datastore 108 stores one or more trained machine learning models including, e.g., model weights, hyperparameters, etc.
[0086] FIG. 4 is a flowchart showing a computer-implemented method for predicting heart disease events by the prediction system 100, in accordance with an embodiment.
[0087] At block 402, biometric data of a particular user is collected over a collection time period spanning a plurality of days. The biometric data include photoplethysmography sensor data and may further include measurements of step count, exercise time, distance traveled, stand time, active and basal energy burned, heart rate, heart rate variability, and oxygen saturation, as described herein. In some embodiments, the biometric data may be acquired from sensors on a user-wearable device, such as a smart watch. In some embodiments, the collection time periodincludes at least ten days, and in further embodiments at least thirty days, enabling the system to capture temporal trends in the user’s physiological patterns.
[0088] At block 404, biometric features are generated based on the collected biometric data. This may include pre-processing, normalization, temporal aggregation (e.g., 90-minute summaries), outlier removal, or other feature-engineering operations suitable for preparing continuous wearable signals for downstream analysis. In some embodiments, the biometric data collected may be aggregated into a temporal patient feature representation bound by a collection window of at least ten days, and in further embodiments, at least thirty days.
[0089] At block 406, the biometric features are supplemented with clinical features of the particular user to augment the temporally indexed data structure. Such clinical features may include demographic characteristics (e.g., sex, age, ethnicity), drug-use or medication-related information, and other baseline clinical attributes.
[0090] At block 408, the biometric features and clinical features are provided to a machine learning model trained to predict heart disease events. In some embodiments, the prediction engine 106 continuously updates the temporal feature representation for the current day by sliding the multi-day window forward and appending the latest sub-daily and daily aggregates derived from wearable sensor data. The prediction engine 106 then provides the temporal feature representation to a temporally-aware machine learning model that processes the temporal feature representation to generate an output predictive of a heart disease event, such as a metric indicative of peak oxygen uptake regression or other indicators of cardiopulmonary decline. The model is configured such that predictions for any given day depend only upon biometric data from that day and preceding days as organised in the temporal state construct, thereby enforcing temporal causality in both training and inference.
[0091] In the embodiment as depicted in FIG. 4, the machine learning model includes a sequence of transformer blocks 120, each configured to generate representations of the biometric data in decreasing temporal resolutions. The temporal resolution may decrease by approximately half at each successive transformer block 120, enabling themodel to learn temporal relationships across progressively broader timescales. Each transformer block 120 is further configured for causal self-attention, ensuring that predictions for a given day depend only on that day and preceding days, thereby preserving temporal causality. In some embodiments, clinical features are provided to at least two transformer blocks 120 to modulate intermediate representations through feature-wise transformations.
[0092] At block 410, the system receives, from the machine learning model, an output predictive of a heart disease event. In some embodiments, this output includes or is based on a metric of VO2regression or another indicator of cardiopulmonary fitness decline.
[0093] It should be understood that steps of one or more of the blocks depicted in FIG.4 may be performed in a different sequence or in an interleaved or iterative manner. Further, variations of the steps, omission or substitution of various steps, or additional steps may be considered.
[0094] FIG. 5 is a schematic diagram of computing device 500 which may be used to implement prediction system 100, in accordance with an embodiment. As depicted, computing device 500 includes at least one processor 502, memory 504, at least one I / O interface 506, and at least one network interface 508.
[0095] Each processor 502 may be, for example, any type of general-purpose microprocessor or microcontroller, a digital signal processing (DSP) processor, an integrated circuit, a field programmable gate array (FPGA), a reconfigurable processor, a programmable read-only memory (PROM), or any combination thereof.
[0096] Memory 504 may include a suitable combination of any type of computer memory that is located either internally or externally such as, for example, randomaccess memory (RAM), read-only memory (ROM), compact disc read-only memory (CDROM), electro-optical memory, magneto-optical memory, erasable programmable read-only memory (EPROM), and electrically-erasable programmable read-only memory (EEPROM), Ferroelectric RAM (FRAM) or the like.
[0097] Each I / O interface 506 enables computing device 500 to interconnect with one or more input devices, such as a keyboard, mouse, camera, touch screen and a microphone, or with one or more output devices such as a display screen and a speaker.
[0098] Each network interface 508 enables computing device 500 to communicate with other components, to exchange data with other components, to access and connect to network resources, to serve applications, and perform other computing applications by connecting to a network (or multiple networks) capable of carrying data including the Internet, Ethernet, plain old telephone service (POTS) line, public switch telephone network (PSTN), integrated services digital network (ISDN), digital subscriber line (DSL), coaxial cable, fiber optics, satellite, mobile, wireless (e.g. Wi-Fi, WiMAX), SS7 signaling network, fixed line, local area network, wide area network, and others, including any combination of these.
[0099] For simplicity only, one computing device 500 is shown but prediction system 100 may include multiple computing devices 500. The computing devices 500 may be the same or different types of devices. The computing devices 500 may be connected in various ways including directly coupled, indirectly coupled via a network, and distributed over a wide geographic area and connected via a network (which may be referred to as “cloud computing”).
[0100] For example, a computing device 500 may be a server, network appliance, set-top box, embedded device, computer expansion module, personal computer, laptop, personal data assistant, cellular telephone, smartphone device (e.g., an iPhone or Android phone), LIMPC tablets, video display terminal, gaming console, or any other computing device capable of being configured to carry out the methods described herein.Experimental StudyMethods
[0101] An experimental study was conducted using an embodiment of prediction system 100. All results or the models described herein in connection with this study relate to this embodiment of prediction system 100.
[0102] Between December 2019 and April 2024, 217 participants were enrolled in a prospective observational cohort study. The study enrolled HF outpatients aged 18 years and older who could adequately comprehend English independently or with caregiver assistance. Purposeful sampling was used to ensure diverse demographics in the patient cohort. Research coordinators provided study information at the time of consent and contacted patients for follow-up. Participants without Apple Watch Series 6 were supplied with an Apple Watch and iPhone.
[0103] All patients underwent formal CPET, comprehensive bloodwork, clinical examination, and a supervised 6MWT during onboard and three-month offboard clinical visits. Apple Watch HealthKit data was collected during clinical visits and free-living observations over the three-month observational period. Participants in the cohort were instructed to capture pedestrian activities by using Apple Watch if they were going to partake in the activity; however, they were not instructed to perform additional daily pedestrian activities (patients were asked to complete monthly at home self-supervised 6MWD and Tecumseh cube tests). Patients completed daily self-administered surveys about their health status and unscheduled events including hospital admissions, clinical visits, or intravenous furosemide treatments. Unscheduled events were manually verified using electronic health records (EHR) and physician notes when available. Protocol
[0104] Purposeful sampling was used to ensure diverse demographics in the patient cohort, as it pertains to sex, self-reported race, and NYHA class. The inclusion criteria for participation was adult (>17 years of age), ambulatory HF patients currently followed by the University Health Network. Patients require literacy in English.
[0105] During the initial enrolment visit, patients were provided with guidance on setting up their Apple Watch. During this session, patients were also educated on howto use the Apple Watch and ensure that the iPhone and Apple Watch are synchronized. The appropriate eCRF and ECG applications were downloaded and available on iPhone. Apple Watch ECG has only been validated for patients above the age of 22, therefore only participants older than 22 years of age were asked to download the Apple ECG app.
[0106] During the free-living observation period, patients were instructed to complete daily surveys in which they assessed the following symptoms: increasing shortness of breath, leg swelling, palpitations, chest pain, lightheadedness, and fainting. Additionally, the daily surveys assessed if a patient required the following in the last 24 hours: changes to medication, intravenous Lasix, unscheduled health visit, emergency room visit, and / or admission to hospital. Finally, daily surveys also asked patients if they performed physical activities in the past 24 hours that were not recorded by iPhone or Apple Watch. On a monthly basis, monthly fitness tests were also conducted, where patients were instructed to partake in a monthly unsupervised 6MWT and Tecumseh cube test. Instructional videos could be used asynchronously to support these tests.
[0107] All demographic and clinical measurements recorded during the onboard and three-month offboard clinical visits were captured in Research Electronic Data Capture (REDCap), ensuring data transcription for the study cohort. In collaboration with Apple developers, an electronic case report form iOS mobile application was developed using the Swift programming language to communicate study information, conduct daily surveys, and gather HealthKit wearable data from patients securely and de-identified.
[0108] Patients with non-adherence, defined by wearing their Apple Watch <10 days throughout the 90-day free-living period, were excluded.Wearable data
[0109] The following data were collected from Apple Watch (wearable device 50) at prediction system 100 through HealthKit during the 90-day free-living period ( / .e., excluding onboard and offboard days): Step Count, Exercise Time, Distance Traveled, Stand Time, Active Energy Burned, Basal Energy Burned, Heart Rate, Heart RateVariability, and O2 Saturation. A standardized preprocessing protocol addressed the varied temporal resolution of the different data types.
[0110] In particular, abnormal data record errors were removed using an outlier approach. Records with values greater than three standard deviations from the population mean for each data type were removed. Next, a coherent and meaningful representation of the wearable data was constructed to integrate large-scale data for downstream usage. Specifically, first, the disparate data streams were normalized by synthesizing 90-minute aggregated metrics (mean, median, minimum, maximum, and standard deviation) of Heart Rate, Heart Rate Variability, and 02 Saturation. For measures of Step Count, Exercise Time, Distance Traveled, Stand Time, Active Energy Burned, and Basal Energy Burned, the sum was used instead of the mean to more accurately capture the overall exercise quality of these HealthKit data types during the 90-minute time window. Subsequently, intra-patient forward-filling imputation was employed to maintain a conservative estimation of variable trajectories during periods of sensor non-recording or loss of signal, thereby preserving the autoregressive integrity of the time series for model training.Prediction model
[0111] An embodiment of the TRUE-HF model was designed to assess changes over time using wearable data combined with clinical demographics: sex, race, age, medication dosages, weight, and height. In this embodiment, the model takes temporal trends across 30-days of wearable data as input to predict the patient’s next day, providing a near-continuous assessment of wearable-derived clinical outcomes fluctuations. In this embodiment, the model processes data bottom-up, constructing longer temporal representations from 90-minute intervals up to full-day aggregations and integrating demographics directly into the wearable features. The primary target variables were CPET pVO2and 6MWT distance (6MWTD) at the 3-month offboard. A separate TRUE-HF model was trained for each outcome (TRUE-HF pVO2model and TRUE-HF 6MWTD model, respectively) using a similar training process.
[0112] Iterative model development and tuning were performed using cross-validation within the first 154 patients, while the final 63 patients offboarded were used solely for held-out testing. Cross-validation within the first 150 patients was performed during interim analyses to develop the model without exposure to the independent held-out test. All model performance results reported herein were calculated from that held-out test set.
[0113] To mitigate the lack of daily clinical CPET and 6MWT outcome labels, selfsupervised learning targeting linear approximated values of each test across the study was utilized. This, in turn, increased the training data by approximately 44-fold without greatly limiting the model. Using random drop-out during training, the TRUE-HF model was trained to handle up to 30% missing data per 30-day window, accounting for sporadic wear and charging needs.
[0114] In some embodiments, the TRUE-HF model includes an ensemble models (e.g., 5, 10, or more), each trained using the substantially the same framework but with different random seeds. In such embodiments, the average prediction from these models is used to make predictions.Model validation and outcomes
[0115] The TRUE-HF pVO2and TRUE-HF 6MWTD models were compared against clinically measured CPET pVO2and 6MWTD, as well as to Apple VO2Max and sixMinuteWalkTestDistance, respectively. Note, certain conditions or medications that limit heart rate may cause an overestimation of Apple VO2Max algorithm-as communicated in Apple’s user interface.
[0116] To further assess the model’s prediction accuracy in detecting changes in fitness levels over time, the model’s ability to detect significant declines at three-month offboard was measured against onboard clinical measurements. This approach ensures sensitivity to meaningful declines, beyond overall correlation with offboard test results. For CPET pVO2, a clinically significant three-month drop from onboard to offboard CPET pVO2was defined as a >10% reduction chosen because a >6-10% decrease inpVO2is associated with an increased risk of medium-to-long-term hospitalization or death in HF patients. To classify patients as either a >10% or <10% drop in pVO2, a percentage difference was calculated between the last model prediction ( / .e. TRUE-HF prediction the day before the clinical visit) to onboard CPET pVO2. The same assessment method was used for the TRUE-HF 6MWTD model; correlation measures and drops in 6MWTD and classification of a drop in 6MWTD (10% reduction in distance walked), respectively, were tested.
[0117] A secondary objective was to evaluate the association between TRUE-HF predicted declines in daily pVO2with unplanned healthcare utilization during the 90-day follow-up period. This association was also compared to associations of other well-known traditional static risk factors measured at baseline. Unplanned healthcare utilization was defined as hospitalization, unscheduled clinical visits, or urgent intravenous furosemide treatment during the three-month study period. The first prediction made by the TRUE-HF pVO2model required 30-days of data. Hence, this objective was only evaluated amongst patients free from unplanned healthcare utilization in the first 31 days of the study.Statistical analysis
[0118] The primary and secondary analyses and methods were pre-specified before data evaluation and performed on the held-out cohort (n=63) only at completion of the study. To conservatively estimate variable trajectories during sensor non-recording, intra-patient forward-filling imputation was used to maintain the autoregressive integrity of the time-series for model training.
[0119] The primary analysis involved comparing the TRUE-HF model's observed and predicted values using Spearman’s coefficient, Pearson R, and mean absolute error (MAE). The Area Under the Receiver Operating Characteristic (AUROC) was further used to evaluate the diagnostic accuracy of correctly predicting a significant drop in outcome measurements. Based on Delong's test, a two-sided p-value of <0.05 was considered significant between models. Confidence intervals for AUROC calculation were computed using 1,000 stratified bootstraps and stratified data resampling.
[0120] The secondary analysis consisted of time-varying extended Simon and Makuch Kaplan-Meier cumulative risk curves to assess the association for an observed decrease of >10% in their TRUE-HF predicted daily pVO2and unplanned healthcare utilization. In this analysis, the cohort was stratified into two groups: 1) patients with a >10% reduction in pVO2; 2) patients with <10% reduction in pVO2. Cumulative incidence curves for unplanned healthcare utilization were constructed, accounting for the time-varying nature of the independent variable. An extended Cox proportional hazards model quantified the hazard ratio (HR) and the strength of association between time-varying occurrence of a >10% drop in TRUE-HF daily pVO2and subsequent clinical events.
[0121] Secondary analysis AUROC was drawn using the maximum % drop detected by the model from the first model prediction to the prediction on the day for each patient preceding any event or three-month offboard clinical visit. This secondary AUROC provided a supplementary perspective on the model’s discriminative capacity. DeLong’s one-tailed test was performed to evaluate the a priori hypothesis that continuous monitoring methods would surpass static measures in discriminative performance. Multiple testing correction was applied using the Benjamini-Hochberg method (target FDR = 0.05) and is reported as adjusted p-values. Landmark-based time-dependent AUROC analysis was performed during external validation to assess the model’s ability to predict outcomes based on the largest percentage drop identified up to each landmark time. This analysis was performed in 10-day intervals, with 30-day outcome windows repeated. It offered a complementary view of the differences in pVO2trajectories between the event and no-event groups over time.
[0122] Qualitative results for the trends observed with TRUE-HF predictions were created using LOWESS-smoothed trajectories, generated by forward-filling censored or missing data across the entire study window using the last observed value prior to censoring or event. This ensured all patients contributed data across the complete study timeline and thus prevented artificial fluctuations due to changes in cohort composition over time.Resu / ts
[0123] FIG. 6A, FIG. 6B, FIG. 6C, FIG. 6D and FIG. 6E show aspects of the study design validating the TRUE-HF model, in accordance with an embodiment.
[0124] FIG. 6A shows a consort diagram and schematic representation of the TRUE-HF observational cohort study design. 960 patients were initially assessed for eligibility, but 187 remained in the final analysis. Reasons for patient ineligibility or nonparticipation included: patients were deemed ineligible or not approached for the following reasons: patients were unable to give consent, had documentation requesting not to be approached for research, were on dialysis, had documented non-compliance, had active psychiatric problems, were new patients to the clinic (de novo), or were part of stratified / purposeful sampling. Additionally, the last 50 patients were recruited to match the CPET pVO2distribution in previous cohorts. Within the non-compliance group (insufficient data), one patient received an LVAD and one patient died. The 15 patients who could not complete the three-month offboard had sufficient data and were included in the unplanned healthcare utilization analysis.
[0125] As shown in FIG. 6B, the study framework incorporates onboard, free-living, and three-month offboard clinical visits. As shown in FIG. 6C, the TRUE-HF model employs a 30-day sliding window for outcome prediction. Data collection spans a three-month observation period, integrating wearable metrics from Apple Watch and clinical assessments during onboard and three-month offboard clinical visits to evaluate clinical outcomes. Additionally, changes in near-continuous daily pVO2predictions from TRUE-HF were associated with key clinical events including CPET pVO2declines, hospitalization, emergency visits, and IV furosemide administration.
[0126] FIG. 6D shows a distribution of key HealthKit variables, illustrating the number of HealthKit recordings collected during the study. In some embodiments, all nine wearable-derived features presented in FIG. 6D (ActiveEnergyBurned, BasalEnergyBurned, HeartRate, Distance, StepCount, AppleStandPlusTime, AppleExerciseTime, OxygenSaturation, and HeartRateVariabilitySDNN) were included as model inputs.
[0127] The TRUE-HF model was trained on the wearable data from the first 163 patients offboarded from the clinical study.
[0128] Between December 2019 and April 2024, 217 patients were enrolled in the TRUE-HF study with a median observation period of 94.5 days. Of the 217 patients, 191 completed the study, of whom 187 completed a three-month offboard CPET. During the study period, 25 patients experienced decreased CPET pVO2from onboard to offboard, and 32 had unplanned healthcare utilization. Among the 32 total events in the cohort, 19 were confirmed via clinical record to be directly HF-related through agreement of two board certified cardiologists, 4 were confirmed as non-cardiac related, and 9 were selfreported hospitalizations. The 9 self-reported hospitalizations were patient-reported outcomes that occurred in patients without centralized linking of healthcare records and thus could not be verified against EHR or physician notes.
[0129] Over the free-living window, 60,597,545 Apple Watch biometric measurements were collected, covering 26,040 days of recordings. Thirty patients were excluded from the primary analysis because they died (n=2), received left ventricular assist device (LVAD) (n=2), withdrew (n=7), could not complete the offboarding process (n=4), or had insufficient device usage (n=15) (FIG. 6A). Among those with insufficient device usage, the median Apple Watch wear time was four days. Patient demographics, baseline clinical variables, and cardiac medications are summarized in Table 1. The study cohort predominantly had HFrEF (82.5%) with reasonable adherence to guideline-directed medical therapy.
[0130] Table 1 below shows demographics and characteristics of the TRUE-HF study cohort (N=217). Data are n (%) or median (IQR).Table 1<><> <> <> <>
[0131] In Table 1, ACE=angiotensin-converting enzyme. ARBs=angiotensin II receptor blockers. ARNIs=angiotensin receptor-neprilysin inhibitors.
[0132] During the study period, 25 patients experienced decreased CPET pVO2from onboard to offboard, and 32 had unplanned healthcare utilization. The unscheduled events were manually verified using EHR and physician notes where available. Nine out of 32 events were patient-reported outcomes that occurred in patients without centralized linking of healthcare records.
[0133] The last 50 successful offboards were allocated for an independent held-out test (patient-wise split) to be conducted for evaluation only at study completion, reducing the risk for test-set overfitting bias from repeated model refinement against the test-set. Successful offboards were defined as those completing a three-month offboard CPET. As a result, the TRUE-HF model was trained exclusively within the first 154 offboarded patients and evaluated on the held-out cohort (n=63) at the end of the study. Five patients in the held-out cohort were excluded from evaluation due to insufficient device usage. Of the remaining 58 held-out patients, 50 completed the study and a three-month offboard CPET. Eight of the 58 held-out patients experienced unplanned healthcare utilization: five unexpected hospital admissions, two unscheduled clinicalvisits, and one required intravenous furosemide treatment. Among these, four events were adjudicated as HF-related, three were self-reported, and one was determined to be unrelated to HF.
[0134] Table 2 shows demographics and characteristics of the TRUE-HF study cohort (N=217), divided by training and held-out validation sets. Data are n (%) or median (IQR). In the held-out validation set, 50 patients completed the three-month offboard CPET.Table 2<
[0135] FIG. 6E shows an independent patient cohort used for external validation. The patient cohort is from the NIH Alloflls Research Program, which uses Fitbit wearable devices. Out of the cohort of 32,147 patients, only 1,664 provide wearable data. A further 400 patients did not meet HER data standards, 767 did not meet initial event criteria, and 304 did not meet wearable data standards. Consequently, 193 patients were identified as closely matching the TRUE-HF cohort inclusion criteria, including confirmed diagnosis of HF documented in electronic health records (EHR), sufficient wearable data coverage (defined as wearable use on at least 40% of study days, non-consecutive), and detailed EHR records available for event adjudication.
[0136] However, regarding HF severity, NYHA class was missing in the EHR for 1653 / 1664 (99%) of all possible Alloflls HF patients with wearable data, with BNP missing in 917 / 1664 (55%). Therefore, to more closely reflect the clinical severity observed in the TRUE-HF population — where 77% were >NYHA II — an inclusion criteria was added requiring at least one previously recorded unplanned healthcare event before enrollment. Thus, to establish a consistent starting point across patients in the AllofUs cohort, an onboarding time point was established as the first point at which both the index event had occurred and wearable data had become available for eachpatient. This ensured that the included patients had symptomatic diseases and thus fit a similar underlying distribution, allowing for the external validation of the readmission risk analysis.
[0137] A more stringent definition of unplanned healthcare utilization criterion was adopted in AllofUs, restricting events to only inpatient hospital admissions or need for intravenous furosemide administration, as: 1) outpatient visits were excluded as unplanned events because they could not be verified as scheduled or unscheduled; and 2) data on self-reported outcomes were not available. Taken together, the above steps ensured a more comparable external population in terms of risk and enabled longitudinal tracking of subsequent events, supporting a readmission-based validation design.
[0138] The Alloflls external validation cohort had a median observation period of 120 days, during which 20 patients experienced at least one unplanned healthcare utilization event. In order to provide a calibration and adjustment period for patients initiating wearable use and to account for a refractory interval following the initial clinical encounter, a seven-day buffer was introduced after the baseline (“onboarding”) date, which was not used in TRUE-HF. To ensure timely model predictions aligned with TRUE-HF, which only requires a total of 30 days before pVO2predictions, while accommodating the necessary 7-day buffer period, predictions were allowed to begin after collecting 20 days of wearable data ( / .e., from day 27 onwards, approximating the 30-day calibration period used in TRUE-HF).
[0139] FIG. 7A, FIG. 7B, and FIG. 7C show that biometric data can be used to estimate cardiorespiratory fitness, in accordance with an embodiment. The TRUE-HF framework is used to predict absolute CPET pVO2(L / min) measurements, and regression analysis was performed of the TRUE-HF pVO2model against the three-month offboard absolute CPET pVO2measurement as a clinically validated performance indicator in the held-out test set (n = 50 successful offboards). Absolute pVO2(L / min) was chosen as this measure removes the need for linear indexing to body weight and alleviates the confounding effects of body mass on the model's predictedoutcomes. Wearable-derived CPET absolute pVO2(L / min) from the TRUE-HF model and the Apple pVO2algorithm (VO2Max) using Apple Watch in the validation set (n=63) are compared. The first 154 patients were used to train the TRUE-HF model and the performance of this model on the validation set (n=63), of which 50 patients completed the study with a three-month offboard CPET.
[0140] FIG. 7A depicts scatterplot representations of model predictions compared to the measured ground truth of pVO2during three-month offboard clinical visits. Model performance was evaluated using Spearman’s coefficient, Pearson R, and mean absolute error (MAE) on the held-out test set, comparing observed and predicted values for the offboard CPET test. Strong concordance can be observed in FIG. 7A between the TRUE-HF pVO2model’s predictions the day before the patients’ three-month offboard date and the gold-standard three-month follow-up CPET pVO2measurements.
[0141] FIG. 7B depicts Spearman’s correlation to the CPET pVO2measured at three-month offboard for TRUE-HF (n=50) and Apple’s Model (n=35). Apple’s pVO2was not measured for 15 participants.
[0142] FIG. 7C depicts an AUROC curve detecting a 10% drop in clinical pVO2from onboard to three-month offboard measurement. AUROC was calculated by comparing each model's last (temporally closest to offboard) prediction to the onboard CPET pVO2detected over three-months against actual observed declines in pVO2. The TRUE-HF model showed a higher Pearson correlation, Spearman’s correlation, and a lower mean absolute error (MAE) than Apple’s pVO2. Based on Delong's two-tailed test, the TRUE-HF model achieved significantly higher AUROC than Apple’s pVO2algorithm for detecting drops in clinical pVO2(p<0.01). Note that certain medications that modulate heart rate may result in overestimation by Apple’s pVO2algorithm.
[0143] The primary analysis found strong concordance between the TRUE-HF pVO2model’s predictions the day before the patients’ three-month offboard date and the three-month offboard CPET pVO2measurements (FIG. 7A). TRUE-HF achieved a Pearson’s R of 0.85 and an MAE of 0.25 L / min in the 50 patients with a successful three-month offboard. Spearman’s coefficient measures (FIG. 7B) also suggest theenhanced capability of the TRUE-HF model to rank HF patients accurately based on their health status. The TRUE-HF model pVO2predictions showed significant accuracy for detecting three-month declines in pVO2, achieving an AUROC of 0.82 (95% Cl: 0.69-0.92) compared to Apple’s algorithm with an AUROC of 0.52 (95% Cl: 0.39-0.70; p < 0.01) (FIG. 7C). Finally, TRUE-HF consistently generated daily predictions during free-living across all HF patients. Conversely, Apple’s algorithm produced fewer predictions given design requirements for physical activity, averaging 4.6 predictions for patients with pVO2<1 L / min, 6.3 for pVO2between 1-2 L / min, and 18.0 for pVO2>3 L / min over the observation period.
[0144] A secondary exploratory analysis was also performed to investigate a previously unexamined potential relationship between wearable-derived daily pVO2and near-term unplanned healthcare utilization (defined above) in both the TRUE-HF held-out and AllofUs cohorts. In this study, the TRUE-HF pVO2model used 30 days of data, and hence, the retrospective risk- stratification analysis for each patient in the held-out TRUE-HF test set commenced on day 31.
[0145] FIG. 8A, FIG. 8B, FIG. 8C, FIG. 8D and FIG. 8E show TRUE-HF model and its reduced sensor variant (TRUE-HF-RS) predicting decline in daily pVO2prior to unplanned health care use, in accordance with an embodiment. FIG. 8A and FIG. 8B show that TRUE-HF predicts lower pVO2prior to the event date, as evidenced by retrospective analysis on TRUE-HF detection of on-study hospitalization in the validation set (n=63), of which 58 patients had sufficient data for analysis (>10 days, 58 / 63).
[0146] In the TRUE-HF held-out cohort, all held-out test patients with sufficient data, including those unable to offboard, were included, resulting in 58 patients analyzed. All patients were censored at time of unplanned healthcare utilization or three-month follow-up visit. Receiver operator curves (ROC) were calculated using the maximum percentage of the drop in TRUE-HF model wearable-derived daily pVO2from the first model prediction to the prediction on that day for each patient. The curves were censored and compared to canonical clinical static measures at time of onboard (CPETpVO2, 6MWT distance (6MWTD), NYHA class, and NT-pro BNP), and clinical HF models including Seattle Heart Failure Model (SHFM), Meta-Analysis Global Group in Chronic (MAGGIC), and PRognostic Evaluation During Invasive CaTheterization for Heart Failure (PREDICT-HF).
[0147] FIG. 8A depicts ALIROC of unplanned healthcare utilization prediction of the TRUE-HF model against the above-mentioned canonical clinical static measures. AU ROC was calculated using the maximum change in TRUE-HF pVO2between the first model prediction and drops before the event for TRUE-HF. For NT-proBNP, 6MWT, and NYHA, AUROC was calculated using the corresponding onboard measures. The circled point on the AUROC curve represents the sensitivity and specificity of a 10% TRUE-HF pVO2drop.
[0148] As shown in FIG. 8A, the predictive power of daily pVO2predictions was measured by AUROC of 0.77 (95% Cl: 0.62-0.90) for predicting HF events. TRUE-HF model’s ability to predict unplanned healthcare utilization was statistically higher than baseline clinical metrics: CPET pVO2(p = 0.04, adjusted p = 0.05), NT-proBNP (p<0.01, adjusted p = 0.01), 6MWTD (p = 0.05, adjusted p = 0.05), NYHA (p = 0.01, adjusted p = 0.02) based on a Delong’s one-tailed test and adjusted p-values from Benjamini-Hochberg correction for False Discovery Rate (0.05). The risk stratification analysis focused on the prespecified >10% drop in TRUE-HF daily pVO2is represented as a specific point on the ROC curve (indicating the sensitivity and specificity at this threshold) in FIG. 8A. At this >10% drop: TRUE-HF sensitivity was 88%, specificity was 62%, positive predictive value was 27%, and negative predictive value was 97% across the complete observation period.
[0149] FIG. 8B depicts time-varying extended Simon and Makuch Kaplan-Meier cumulative risk curve for TRUE-HF model of 10% drop detected for event-free survival during the study; shadowed areas indicate 95% confidence intervals. TRUE-HF drop group is a time-dependent covariate, which more accurately reflects the impact of timedependent predictions on survival.
[0150] The secondary analysis shown in FIG. 8A and FIG. 8B demonstrated that patients with decreases in the TRUE-HF model’s daily wearable-derived pVO2were at significantly higher risk of unscheduled healthcare utilization on-study. The study’s risk stratification analysis focused on a >10% drop in daily TRUE-HF pVO2, represented as a specific point on the AU ROC curve (indicating the sensitivity and specificity at this threshold) in FIG. 8A. The predictive power of daily TRUE-HF pVO2predictions achieved an AUROC of 0.77 (95% Cl: 0.63-0.89) for predicting these events. TRUE-HF model’s ability to predict unplanned healthcare utilization was statistically higher than CPET pVO2(p=0.04), NT-proBNP (p<0.01), 6MWTD (p=0.05), and NYHA class (p=0.01). During free-living, a TRUE-HF wearable-derived pVO2drop was found in 26 / 58 patients. Seven (26.9%) of the 26 patients in the drop group experienced unplanned healthcare utilization, compared with one (3.1%) of 32 participants in the nodrop detected group (hazard ratio 8.51 (95% Cl 1.11-21.88, p=0.01, as shown in Table 3 below)). Drops detected by the TRUE-HF model preceded unplanned healthcare utilization by a median time of 7.4 days (IQR 4.5-8.5).
[0151] FIG. 8C shows longitudinal TRUE-HF daily pVO2model predictions in patients with unplanned healthcare utilization events versus those without. Longitudinal qualitative analysis of forward-filled LOWESS-smoothed TRUE-HF predicted trajectories demonstrated that patients experiencing an unplanned healthcare event exhibited a progressive decline in daily TRUE-HF pVO2compared to event-free patients, who maintained stable or slightly increasing pVO2values. Individual patient trajectories revealed heterogeneity in TRUE-HF predictions, although the overall trend showed a consistent downward shift relative to baseline.
[0152] Extended Cox proportional hazards models quantified the unadjusted hazard ratio (HR) for the association between continuous-change TRUE-HF daily pVO2covariate and subsequent unplanned healthcare utilization, demonstrating that each 10% decrease in pVO2was associated with a significantly elevated risk (HR 3.62, 95% Cl: 1.37-9.55, p<0.01; as shown in Table 3 below). Table 3 shows Cox proportional hazard ratios and Wald test p-value for time-varying TRUE-HF 10% drops in daily pVO2and onboard comparator risk scores and clinical measures including CPET pVO2, NT-proBNP, 6MWTD, and NYHA functional class. The Cox proportional hazard ratio and Wald test (p-value) is also reported for time-varying TRUE-HF drops in daily pVO2in the Alloflls cohort. The time-varying Cox model for TRUE-HF achieved a significant hazard ratio for both TRUE-HF (3.619, p-value: 0.009) and AllofUs (1.323, p-value: 0.026).Table 3
[0153] Sensitivity analyses adjusting for independent covariates: age, sex, ethnicity, BMI, smoking, or LVEF confirmed that the association remained significant. Table 4 shows sensitivity analysis of TRUE-HF unplanned healthcare utilization across different demographic and covariates. Each row represents a separate Cox proportional hazards model minimally adjusted for one demographic or clinical variable.Table 4
[0154] Extended Cox proportional hazards models were used to quantify the unadjusted hazard ratio (HR) for the association between continuous-change TRUE-HF daily pVO2covariate and subsequent unplanned healthcare utilization, demonstrating that each 10% decrease in pVO2was associated with a significantly elevated risk (HR 3.62, 95% Cl: 1.37-9.55, p<0.01). Sensitivity analyses adjusting for independent covariates such as age, sex, ethnicity, BMI, smoking, or LVEF confirmed that the association remained significant
[0155] Further analysis of common clinical risk measurements did not achieve significant hazard ratios. Table 5 shows Cox proportional hazard ratios and Wald test (p-value) table for all assessments, for time-varying TRUE-HF drops in pVO2and onboard measures: CPET pVO2, NT-proBNP, 6MWTD, and NYHA functional class. The time-varying Cox model for TRUE-HF achieved a significant hazard ratio of 8.507 (p=0.013).Table 5
[0156] A reduced-sensor variant of the TRUE-HF pVO2model was trained to enable compatibility with data available in Alloflls. Embodiments of the reduced-sensor variant require only heart rate and step count data as inputs. To address the difference in available data, a knowledge-distillation approach was employed, specifically a teacherstudent training strategy. In this approach, predictions generated by a more comprehensive teacher model (TRUE-HF model) serve as training targets (pseudolabels) for a streamlined student model that accommodates the reduced feature set. Given the substantial feature gap between the original TRUE-HF model and the AllofUs- compatible variant, an intermediate “teacher-assistant” model was introduced to facilitate a smoother knowledge transfer and to mitigate performance degradation.
[0157] Initially, the original TRUE-HF model (teacher), trained on the complete wearable and clinical feature set from the TRUE-HF cohort, provided pseudo-label predictions for training a teacher-assistant model. This teacher-assistant model maintained all original wearable features but utilized a reduced clinical feature set matched to the AllofUs cohort. For each training batch, pseudo-labels were generated from a randomly selected member of the original TRUE-HF 10-model ensemble to enhance robustness and generalizability. Subsequently, the ensemble of trained teacher-assistant models provided pseudo-labels to train the final AllofUs-compatible TRUE-HF reduced sensor model (TRUE-HF-RS), which relied exclusively on HeartRate, StepCount, and the reduced clinical feature set. This systematic adaptation ensured effective and gradual transfer of learned predictive relationships from the TRUE-HF cohort to the reduced sensor model for external validation on the AllofUs cohort.
[0158] The model can be trained exclusively on the same training set as the TRUE-HF model (n = 154). Table 6 shows performance of reduced-sensor model, TRUE-HF-RS, on TRUE-HF held-out set (n = 63) and AllofUs Validation (n = 193). Where appropriate, a threshold of >10% TRUE-HF-RS pVO2drop was used.Table 6
[0159] In the internal held-out (n = 63), the reduced-sensor variant (TRUE-HF-RS) was able to predict unplanned healthcare utilization with an AUROC of 0.57 (95% Cl: 0.35-0.79) and achieved a HR of 1.79 (95% Cl: 0.77-4.16, p = 0.18).
[0160] External validation in the AllofUs cohort using the TRUE-HF-RS model reveals significant positive trends in risk characterization of unplanned healthcare utilization.FIG. 8D shows time-varying Simon and Makuch Kaplan-Meier cumulative risk curves for event-free survival in AllofUs patients, comparing individuals with and without detected >10% drops in daily predicted pVO2. Consistent with the trends observed in the internal validation set, each 10% decrease in predicted daily pVO2was associated with a higher risk of unplanned healthcare use among HF patients (HR 1.32, 95% Cl: 1.03-1.69; p =0.03). In contrast, those without an event had progressively increasing pVO2over time. The TRUE-HF-RS model predicted a median unplanned healthcare utilization of 21 days (IQR, 9.25-67.25) before events.
[0161] Repeated sensitivity analyses for Alloflls found that the HR remained statistically significant against each individual covariate. Table 7 shows sensitivity analysis of Alloflls unplanned healthcare utilization across different demographic and covariates with TRUE-HF-RS. Due to NIH policies, groups less than 20 cannot be reported. As such, Ethnicity represents a binary variable denoting whether an individual is or is not White. Each row of Table 7 represents a separate Cox proportional hazards model minimally adjusted for one demographic or clinical variable.Table 7
[0162] Across the complete observational period in Alloflls HF patients, the TRUE-HF-RS model at the pre-specified threshold of >10% decrease in daily pVO2showed a sensitivity of 50%, a specificity of 59%, a PPV of 12%, and a NPV of 91% for unplanned healthcare use (FIG. 7D). Notably, the high NPV (91%) in Alloflls is consistent with the primary findings and underscores the model’s accuracy in identifying low-risk HF patients, providing confidence in its ability to rule out individuals unlikely to experience near-term events. This risk stratification capability enables targeted monitoring of those likely to benefit from early intervention.
[0163] Qualitatively, similar trends were also observed in patients with and without events compared to the TRUE-HF cohort. FIG. 8E shows longitudinal predicted daily pVO2LOWESS trajectories in AllofUs cohort patients, comparing those with and without unplanned healthcare utilization events. Solid lines illustrate smoothed population-level trends with 95% confidence intervals. LOWESS-smoothed trajectories, generated by forward-filling missing or censored data across the study window, highlighted a sustained decline among event patients. As shown in FIG. 8E, there is a decrease in predicted pVO2over time in patients with events compared to those without events. Aligning with the observed trends in AllofUs, the time-dependent landmark AUROC evaluated over month-long windows was 0.51 (95% Cl: 0.33-0.68) at day 10, versus 0.72 (95% Cl: 0.39-0.92) by day 80.
[0164] FIG. 9 shows a time-dependent performance analysis of the TRUE-HF-RS Model, in accordance with an embodiment. For a plurality of predetermined temporal landmarks, FIG. 9 depicts the model’s discrimination capability and classification characteristics when used to predict impending adverse healthcare utilization events in patients monitored via wearable sensor devices.
[0165] The horizontal axis represents a sequence of “landmark times,” each corresponding to a point in the longitudinal monitoring interval at which the system retrospectively evaluates model performance based solely on data accumulated up to that time.
[0166] The vertical axis represents one or more time-dependent performance metrics calculated at each of the landmark times. The AU ROC curve demonstrates a progressive improvement in model discrimination as additional sensor data becomes available, with peak performance occurring at approximately the 50-day landmark. Thereafter, the AUROC maintains a stable plateau, indicating sustained predictive capability over extended monitoring periods.
[0167] The sensitivity curve represents the true-positive detection rate associated with identifying patients at elevated near-term risk of unplanned healthcare utilization based on declines in daily wearable-derived physiological estimates. The sensitivity curve shows substantial temporal variation, including a marked increase near the 50-day landmark time, where sensitivity approaches unity. This indicates that, at this stage of accumulated observational data, the model reliably identifies nearly all individuals who subsequently experience adverse clinical events.
[0168] The specificity curve represents the model’s true-negative rate. In contrast to sensitivity, specificity exhibits a gradual downward trend over the examined interval. This reduction in specificity reflects an increased frequency of precautionary positive classifications as more longitudinal data becomes incorporated into the prediction mechanism.
[0169] The shaded bands surrounding each curve represent confidence intervals demonstrating statistical variability within the monitored population. The widening and narrowing of these regions across the landmark timeline indicate changes in the certainty of the model’s performance estimates as more or less data is available for analysis.
[0170] Collectively, FIG. 9 demonstrate that the TRUE-HF-RS Model exhibits improved discriminative capacity and high sensitivity when provided with sufficient longitudinal wearable sensor data, while simultaneously illustrating the trade-off between elevated sensitivity and reduced specificity. This performance profile enables the system to function as an early-warning mechanism for detecting patient-specificphysiologic deterioration prior to the occurrence of clinically significant unplanned healthcare utilization events.
[0171] The proposed systems and methods are extended to 6MWTD as an additional variable that can be predicted using wearable devices and further examined for its prediction of unplanned healthcare utilization. FIG. 10A, FIG. 10B, and FIG. 10C show Al prediction of 6MWTD from the TRUE-HF model and Apple’s 6MWTD algorithm using Apple Watch in the validation set (n=63), in accordance with an embodiment. The first 154 patients were used to train the TRUE-HF model and the performance of this model on the validation set (n=63), of which 49 patients completed the study with a three-month offboard 6MWT.
[0172] A separate TRUE-HF model was trained for 6MWTD and compared to Apple’s algorithm, the sixMinuteWalkTestDistance. Similar to the CPET pVO2assessment, a drop in 6MWTD (10% reduction in 6MWTD from baseline to three-month follow-up) and healthcare utilization were evaluated. Wearable data from the 30 days before the offboard visit to predict offboard CPET values. This ensured all model inputs reflected free-living conditions, unaffected by structured CPET or tests conducted on the visit day.
[0173] FIG. 10A depicts scatterplot representations of model predictions compared to the measured ground truth of 6MWTD. The sixMinuteWalkTestDistance in HealthKit performed comparably to the 6MWT study model across all correlation metrics and effectively detected drops in 6MWT performance.
[0174] FIG. 10B depicts Spearman’s correlation of the TRUE-HF model and Apple’s model.
[0175] FIG. 10C depicts an AUROC curve of detecting an actual 10% drop in 6MWTD. TRUE-HF and Apple’s 6MWTD algorithm performed similarly in detecting 6MWTD changes.The Apple 6MWT algorithm in HealthKit performed comparably to the 6MWT study model across all correlation metrics, and effectively detected drops in 6MWT performance.
[0176] FIG. 11A and FIG. 11B show the prediction of unplanned healthcare utilization, in accordance with an embodiment. As shown, TRUE-HF and Apple Algorithm 6MWT are predictive in retrospective analysis on TRUE-HF detection of unplanned healthcare utilization. In the validation set (n=63), 58 patients had sufficient data for TRUE-HF analysis (>10 days, 58 / 63). In comparison, 53 patients had Apple SixMinuteWalkDistance predictions (53 / 63).
[0177] FIG. 11A depicts AUROC of unplanned healthcare utilization prediction of the TRUE-HF 6MWT model and Apple Algorithm 6MWT. AUROC was calculated using the maximum change in 6MWT between the first model prediction and drops before the event. The circled point on the AUROC curve represents the sensitivity and specificity of a 10% drop.
[0178] FIG. 11B depicts Cox proportional hazard ratios and Wald test (p-value) results for all assessments for time-varying TRUE-HF 6MWT drops and time-varying Apple Algorithm 6MWT drops.
[0179] As shown in FIG. 11A and FIG. 11B, findings indicate that wearables can reliably track daily physical activity and predict the 6MWTD measured in the clinic. Notably, the HR for the remote measurement of 6MWTD was 1.65 (95% Cl: 0.63-4.31) which exceeds the HR for static 6MWTD measurements.
[0180] Additional ablation and feature importance experiments were conducted on the TRUE-HF model.
[0181] Firstly, wearable data collected during the monthly unsupervised 6MWT and Tecumseh Cube tests was masked to evaluate the contribution of structured exercise sessions. All wearable data collected during these structured sessions were masked by excluding measurements within a 90-minute window surrounding the start and end times of each test.
[0182] FIG. 12A and FIG. 12B show model robustness to the removal of unsupervised exercise data, in accordance with an embodiment. Specifically, FIG. 12A shows a distribution of the difference in TRUE-HF pVO2when wearable data from unsupervised 6MWT and Tecumseh Cube tests were excluded; and FIG. 12B shows all predictions made with TRUE-HF pVO2using all wearable data versus predictions made excluding unsupervised exercise data.
[0183] As shown in FIG. 12A and FIG. 12B, the exclusion resulted in negligible differences in predictive performance (MAE change of 0.00575 in predicted pVO2). Specifically, regression performance and classification accuracy for detecting 10% drops in predicted pVO2, as well as predicting unplanned healthcare utilization, remained unchanged, with AUROCs of 0.82 (95% Cl: 0.69-0.92) and 0.77 (95% Cl: 0.62-0.90), respectively. The strong linear correlation shown in FIG. 12B indicates that model predictions are not dependent on structured unsupervised tests, highlighting robustness to variations in activity context.
[0184] Secondly, variation in TRUE-HF prediction performance with shorter input window durations (20 days and 10 days) during both training and evaluation was examined. Table 8 shows the results of training the TRUE-HF model using fewer input days; and Table 9 shows the results of training the TRUE-HF model on 30-day windows and reduced number of days in the inference input:Table 8Table 9
[0185] Reduced model performance was observed, specifically: unplanned healthcare utilization at 20 day windows had an ALIROC of 0.61 (95% Cl: 0.40-0.80), unplanned healthcare utilization at 10 day windows had an ALIROC of 0.61 (95% Cl: 0.39-0.82), compared to the complete 30-day input window having an ALIROC of 0.77 (95% Cl: 0.62-0.90), indicating that the 30-day input window may provide more reliable information.
[0186] Thirdly, saliency-based analyses were conducted to enhance interpretability and better understand the TRUE-HF model predictions. FIG. 13A and FIG. 13B show saliency values from features of the model. Specifically, FIG. 13A is a bee-swarm plot visualizing per-feature saliency values derived from the TRUE-HF pVO2model. Each point represents the average saliency of an individual sample for a given HealthKit variable across a single day. The x-axis denotes the saliency magnitude (positive indicating contribution towards higher pVO2predictions). Mean saliency contributions sort features along the y-axis. Hues reflect the normalized raw feature value of each sample. FIG. 13B is a bee-swarm plot visualizing the top 10 clinical feature saliency values derived from the TRUE-HF pVO2model for each prediction. The x-axis denotes the saliency magnitude (positive indicating contribution towards higher pVO2predictions). Mean saliency contributions sort features along the y-axis. Hues reflect the normalized raw feature value of each sample.
[0187] As shown in FIG. 13A and FIG. 13B, wearable sensor data, especially AppleExerciseTime, StepCount, and OxygenSaturation, demonstrated contributions to model predictions, with higher activity levels generally associated with increased predicted pVO2. Clinical variables (e.g., age, sex, ethnicity, and furosemide dosage) also influenced predictions but exhibited context-dependent effects, modulated by concurrent wearable data, with no consistent directional trends observed independently.
[0188] As wearable technologies have, at times, shown reduced performance on darker skin, subgroup analysis between white and non-white participants in both cohorts showed comparable model performance across both the TRUE-HF and Alloflls validation sets. Table 10 shows subgroup analysis of white and non-white participants in TRUE-HF and AllofUs. In TRUE-HF, Spearman with clinical CPET pVO2, MAE, and pVO2AU ROC are chosen as they relate to the primary endpoint. In AllofUs, unplanned healthcare utilization performance is measured via Sensitivity and HR since clinical CPET is not available. Notably, AllofUs is performed with the TRUE-HF-RS model. Even though, wearable technologies have at times shown reduced performance on darker skin, subgroup analysis between white and non-white participants in both cohorts showed comparable model performance across both the TRUE-HF and AllofUs validation sets.Table 10
[0189] Finally, Table 11 shows the results of the TRUE-HF model using clinical and / or wearable model inputs. Ablation analyses assessing the independent and combined contributions of wearable and clinical inputs demonstrated that wearable data alone provided substantial predictive capacity (unplanned healthcare utilization had an ALIROC of 0.60 (95% Cl: 0.42-0.78)). In contrast, predictions based solely on clinical variables were notably weaker (unplanned healthcare utilization having an AU ROC of 0.52 (95% Cl: 0.32-0.71)). However, integration of both wearable and clinical data ( / .e., TRUE-HF model) yielded the strongest performance across all outcomes, highlighting their complementary predictive value.Table 11
[0190] In the above-described TRUE-HF observational cohort study, deep learning based wearable-derived biometrics from Apple Watch demonstrated a strong concordance with gold-standard CPET measurements. Further, a significant decline inwearable-derived daily pVO2occurred before unplanned HF healthcare utilization events, suggesting serial monitoring using wearables may be useful to forecast incipient HF acute events. Finally, despite strong concordance between wearable predicted and clinical 6MWTD, neither was associated with worsening HF.
[0191] Alternative monitoring devices and methods are restricted by their inability to integrate broader contextual health data in day-to-day life, which could provide more nuanced insights into a patient’s health trajectory. In contrast, the TRUE-HF model extended the observational time to a rolling 30-day window and leveraged wearable technology data from free-living daily activities, enhancing its predictions' accuracy, reliability, and continuity.
[0192] The findings herein are an essential addition to risk evaluation in HF for several reasons. First, the findings herein demonstrated the promise of wearables for accurate serial pVO2assessment in an outpatient HF population, expanding the ability for pVO2assessment from the clinic to a daily free-living environment. The TRUE-HF model achieved a high level of concordance with gold standard CPET measurements of pVO2, achieving an MAE of 0.25 L / min. This measurement error for the predictions falls close to the test-retest confidence interval (95% Cl 0.19-0.23) previously reported for HF patients. This is further exemplified by the TRUE-HF model's high accuracy in predicting decreases in CPET pVO2(10% drops) between study onboard to offboard. The high accuracy of wearable-derived pVO2using TRUE-HF enables a shift from static, infrequent clinical CPET to frequent, dynamic daily pVO2.
[0193] Further, the findings show that shifting from infrequent to daily pVO2monitoring allows for detecting previously undiscovered rapid declines in daily pVO2(10% drops) that preceded unplanned healthcare utilization. These declines in TRUE-HF wearable-derived pVO2preceded unplanned healthcare utilization by a median of 7.4 day lead time. In the held-out retrospective validation, these drops were strongly associated with unplanned healthcare utilization, with a significant hazard ratio of 8.51 compared to patients without detected drops.
[0194] In this study, the TRUE-HF model significantly enhances early detection of deterioration, acting as an early warning signal to identify patients at risk before clinical symptoms appear. At a 10% drop, the TRUE-HF model demonstrated a sensitivity of 88%, specificity of 62%, and NPV of 97% in predicting unplanned healthcare utilization. The sensitivity and NPV underscore the TRUE-HF model’s effectiveness in accurately identifying low-risk patients, thereby effectively ruling out individuals unlikely to require immediate intervention and supporting early low-risk intervention (a clinical touchpoint) for higher-risk individuals. In this clinical context, prioritizing sensitivity is crucial, as the early identification of at-risk patients enables more proactive measures, such as expedited clinical visits, which can reduce the risk of unplanned healthcare utilization. Thus, using drops in TRUE HF daily pVO2as an EWS could enable on-demand care and early interventions to improve HF outcomes. In contrast, practical, conventional metrics, such as NT-proBNP levels, CPET pVO2, 6MWTD, or NYHA functional class, failed to provide the same proactive early warning.
[0195] Continuous model-derived changes in daily pVO2were strongly associated with unplanned healthcare utilization, with a hazard ratio of 3.62 per 10% decrease (95% Cl 1.37-9.55, p<0.01). Static measures inherently capture a single snapshot of physiological status, whereas the TRUE-HF model and variants capture evolving physiological changes closer to the time of clinical events, which may enhance predictive performance. The near-continuous monitoring of pVO2through the TRUE-HF model represents a more dynamic and sensitive risk stratification method. The Multisensor Non-invasive Remote Monitoring for Prediction of HF Exacerbation (LINK-HF) study demonstrated that a worn chest patch was able to predict HF-related hospitalizations. Together, the results show how wearable-based approaches can impact HF management, offering insights into risk prediction and early intervention.
[0196] The results of the TRUE-HF observational clinical study provide strong evidence to support the potential of consumer wearable data, such as the ones derived from Apple Watch, in estimating cardiorespiratory fitness and building remote monitoring methods for HF. HF patient wearable data in a free-living environment analyzed using a deep learning (DL) transformer model demonstrated strongconcordance with gold-standard CPET measurements. Wearable-derived daily cardiopulmonary fitness assessments enabled new insights that significant drops in daily pVO2forecast imminent unplanned healthcare utilization events in HF outpatients, with external validation of a reduced sensor model maintaining a positive risk relationship. Finally, ancillary analysis found wearable data showed strong concordance with 6MWTD.
[0197] The findings indicate that reliance on specific exercise thresholds to trigger health predictions, e.g., as used by the Apple Watch pVO2algorithm, poses a significant barrier to tracking health in sicker patients, whereas using longer-term trends yields more consistent measurements. HF patients often suffer from exercise intolerance related to deconditioned muscles, chronic anemia, reduced cardiac reserve, peripheral vasoconstriction, and other comorbidities that may limit their ability to trigger intensitybased exercise thresholds. For instance, at the time of the study, Apple’s pVO2algorithm required a patient to reach at least 60% of their maximum heart rate and certain activity parameters before an observation is reported. Thus, there were far fewer predictions in sicker HF patients with Apple’s pVO2, which generates 4.6 predictions over 90 days in patients with a CPET pVO2<1L / min, a lower frequency than the daily predictions generated by the TRUE-HF pVO2model. Similarly, for HF patients with a CPET pVO2ranging from 1L / min to 2L / min, the Apple algorithm produced 6.3 predictions. In healthier HF patients (CPET pVO2> 3L / min), Apple Watch recorded an average of 18.0 predictions over the same duration. In contrast, the TRUE-HF pVO2model focussed on trends across a larger time window and provided consistent daily predictions for all patients in the study’s cohort, improving equitability in sicker patients.
[0198] The systems and methods described herein additionally addresses issues arising from the volume of data associated with remote monitoring with wearables: remote monitoring with wearables can generate upwards of 5 gigabytes of data per patient per week. Analyzing this data to identify biomarker signatures is best suited for modern Al techniques, particularly deep learning. The transformer model framework was utilized to interpret wearable data sequences and measure functional changes in cardiorespiratory fitness among HF patients — an area not previously explored. TheTRUE-HF model employs temporal pooling, allowing it to assess 30-days of wearables data (90-minute segments) and aggregate them into increasingly larger periods while retaining temporal trends and relationships. The TRUE-HF model may look at how each hour fits into the rest of that day, how that day fits into the week, and how that week fits into 30 days. In contrast, other Al methods, like random forest, treat each moment independently; thus, they struggle with understanding how a Tuesday workout might impact activities on Thursday.
[0199] The contributions of various data inputs within the TRUE-HF model were evaluated. The results demonstrated that excluding structured monthly exercise tests (unsupervised 6-minute walk and Tecumseh Cube tests) had a minimal impact on prediction accuracy, underscoring that TRUE-HF predominantly leverages naturalistic, daily-living physiological signals rather than periodic, structured tests. Although ablation studies highlighted that shorter windows retained meaningful predictive capacity, the longer input window-length of 30-days provided notable gains in predicting unplanned healthcare utilizations. Saliency analyses provided insights into model decision-making, highlighting wearable metrics, such as Apple Exercise Time, Step Count, and Oxygen Saturation, as consistently influential factors. Meanwhile, clinical variables, including age, sex, ethnicity, and medication usage, provided contextually relevant information that enhanced predictions. These findings collectively reinforce the capability and clinical relevance of TRUE-HF and its variants, highlighting their potential as proactive tools in remote monitoring and management of HF patients.
[0200] The foregoing discussion provides many example embodiments of the inventive subject matter. Although each embodiment represents a single combination of inventive elements, the inventive subject matter is considered to include all possible combinations of the disclosed elements. Thus, if one embodiment comprises elements A, B, and C, and a second embodiment comprises elements B and D, then the inventive subject matter is also considered to include other remaining combinations of A, B, C, or D, even if not explicitly disclosed.
[0201] The embodiments of the devices, systems and methods described herein may be implemented in a combination of both hardware and software. These embodiments may be implemented on programmable computers, each computer including at least one processor, a data storage system (including volatile memory or non-volatile memory or other data storage elements or a combination thereof), and at least one communication interface.
[0202] Program code is applied to input data to perform the functions described herein and to generate output information. The output information is applied to one or more output devices. In some embodiments, the communication interface may be a network communication interface. In embodiments in which elements may be combined, the communication interface may be a software communication interface, such as those for inter-process communication. In still other embodiments, there may be a combination of communication interfaces implemented as hardware, software, and combination thereof.
[0203] Throughout the foregoing discussion, numerous references will be made regarding servers, services, interfaces, portals, platforms, or other systems formed from computing devices. It should be appreciated that the use of such terms is deemed to represent one or more computing devices having at least one processor configured to execute software instructions stored on a computer readable tangible, non-transitory medium. For example, a server can include one or more computers operating as a web server, database server, or other type of computer server in a manner to fulfill described roles, responsibilities, or functions.
[0204] The technical solution of embodiments may be in the form of a software product. The software product may be stored in a non-volatile or non-transitory storage medium, which may be a compact disk read-only memory (CD-ROM), a USB flash disk, or a removable hard disk. The software product includes a number of instructions that enable a computer device (personal computer, server, or network device) to execute the methods provided by the embodiments.
[0205] The embodiments described herein are implemented by physical computer hardware, including computing devices, servers, receivers, transmitters, processors, memory, displays, and networks. The embodiments described herein provide useful physical machines and particularly configured computer hardware arrangements.
[0206] Of course, the above-described embodiments are intended to be illustrative only and in no way limiting. The described embodiments are susceptible to many modifications of form, arrangement of parts, details and order of operation. The disclosure is intended to encompass all such modification within its scope, as defined by the claims.
Claims
WHAT IS CLAIMED IS:
1. A computer-implemented system for predicting heart disease events, the system comprising:a processing subsystem that includes one or more processors and one or more memories coupled with the one or more processors, the processing subsystem configured to cause the system to:collect biometric data of a particular user over a collection time period spanning a plurality of days, the biometric data including photoplethysmography sensor data;generate biometric features based on the collected biometric data,; supplement the biometric features with clinical features for the particular user;provide, to a machine learning model trained to predict heart disease events, the biometric features as supplemented by the clinical features; andreceive, from the machine learning model, an output predictive of a heart disease event.
2. The computer-implemented system of claim 1, wherein the biometric data is collected via a wearable device.
3. The computer-implemented system of claim 1 or claim 2, wherein the output includes a metric of VO2 regression.
4. The computer-implemented system of any one of claims 1 to 3, wherein the collection time period includes at least 10 days.
5. The computer-implemented system of any one of claims 1 to 3, wherein the collection time period includes at least 30 days.
6. The computer-implemented system of any one of claims 1 to 5, wherein the biometric features provided to the machine learning model reflect biometric data spanning the collection time period.
7. The computer-implemented system of any one of claims 1 to 6, wherein the clinical features include demographic features and / or drug-use features of the particular user.
8. The computer-implemented system of claim 7, wherein the demographic features include at least one of a sex, an age, or an ethnicity of the particular user.
9. The computer-implemented system of any one of claims 1 to 8, wherein the machine learning model includes a sequence of transformer blocks, each of the transformer blocks along the sequence configured to generate representations of the biometric data in decreasing temporal resolutions and configured for causal-self attention to provide temporal causality across the temporal resolutions.
10. The computer-implemented system of claim 9, wherein the temporal resolutions decrease by approximately half for each of the transformer blocks along the sequence.
11. The computer-implemented system of claim 9 or claim 10, wherein the clinical features are provided to at least two of the transformer blocks in the sequence of transformer blocks.
12. The computer-implemented system of any one of claims 1 to 11, wherein the prediction is for a prediction time period subsequent to the collection time period.
13. The computer-implemented system of any one of claims 1 to 12, wherein the biometric data includes measurements of at least one of step count, exercise time, distance traveled, stand time, active energy burned, basal energy burned, heart rate, heart rate variability, and 02 saturation.
14. The computer-implemented system of any one of claims 1 to 13, wherein the biometric features include at least one of sub-daily, daily, or multi-day aggregations.
15. A computer-implemented method for predicting heart disease events, the method comprising:collecting biometric data of a particular user over a collection time period spanning a plurality of days, the biometric data including photoplethysmography sensor data;generating biometric features based on the collected biometric data; supplementing the biometric features with clinical features for the particular user;providing, to a machine learning model trained to predict heart disease events, the biometric features as supplemented by the clinical features; andreceiving, from the machine learning model, an output predictive of a heart disease event.
16. The computer-implemented method of claim 15, wherein the biometric data is collected via a wearable device.
17. The computer-implemented method of claim 15 or claim 16, wherein the output includes a metric of VO2 regression.
18. The computer-implemented method of any one of claims 15 to 17, wherein the collection time period includes at least 10 days.
19. The computer-implemented method of any one of claims 15 to 18, wherein the collection time period includes at least 30 days.
20. The computer-implemented method of any one of claims 15 to 19, wherein the biometric features provided to the machine learning model reflect biometric data spanning the collection time period.
21. The computer-implemented method of any one of claims 15 to 20, wherein the clinical features include demographic features and / or drug-use features of the particular user.
22. The computer-implemented method of claim 21, wherein the demographic features include at least one of a sex, an age, or an ethnicity of the particular user.
23. The computer-implemented method of any one of claims 15 to 22, wherein the machine learning model includes a sequence of transformer blocks, each of the transformer blocks along the sequence configured to generate representations of the biometric data in decreasing temporal resolutions and configured for causal-self attention to provide temporal causality across the temporal resolutions.
24. The computer-implemented method of claim 23, wherein the temporal resolutions decrease by approximately half for each of the transformer blocks along the sequence.
25. The computer-implemented method claim 23 or claim 24, wherein the clinical features are provided to at least two of the transformer blocks in the sequence of transformer blocks.
26. The computer-implemented method of any one of claims 15 to 26, wherein the prediction is for a prediction time period subsequent to the collection time period.
27. The computer-implemented method of any one of claims 15 to 26, wherein the biometric data includes measurements of at least one of step count, exercise time, distance traveled, stand time, active energy burned, basal energy burned, heart rate, heart rate variability, and 02 saturation.
28. The computer-implemented method of any one of claims 15 to 27, wherein the biometric features include at least one of sub-daily, daily, or multi-day aggregations.
29. A non-transitory computer-readable medium or media having stored thereon machine interpretable instructions which, when executed by a processing system, cause the processing system to perform a method for predicting heart disease events, the method comprising:collecting biometric data of a particular user over a collection time period spanning a plurality of days, the biometric data including photoplethysmography sensor data;generating biometric features based on the collected biometric data;supplementing the biometric features with clinical features for the particular user;providing, to a machine learning model trained to predict heart disease events, the biometric features as supplemented by the clinical features; andreceiving, from the machine learning model, an output predictive of a heart disease event.