Obstructive sleep apnea outcome prediction system and method

A machine learning system analyzing ventilatory, hypoxic, and arousal domains in polysomnography data effectively predicts short-term and long-term obstructive sleep apnea outcomes, addressing the limitations of traditional indices by providing accurate risk stratification and survival trajectory analysis.

WO2026060166A1PCT designated stage Publication Date: 2026-03-19MT SINAI SCHOOL OF MEDICINE
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Filing Date
2025-09-11
Publication Date
2026-03-19

AI Technical Summary

Technical Problem

Existing methods for assessing obstructive sleep apnea severity, such as the apnea-hypopnea index, are limited in predicting short-term and long-term adverse outcomes due to subjectivity, variability, and failure to account for the complexity of the disorder's pathophysiology, and alternative metrics in individual domains do not fully represent overall severity.

Method used

A computer-implemented system using machine learning models, specifically XGBoost, analyzes polysomnography data across ventilatory, hypoxic, and arousal domains to predict short-term and long-term outcomes by integrating features into non-linear regression models, providing a comprehensive assessment of obstructive sleep apnea severity.

Benefits of technology

The system accurately predicts short-term outcomes like excessive daytime sleepiness and long-term outcomes like all-cause mortality, outperforming traditional metrics by providing a more nuanced understanding of obstructive sleep apnea severity through improved risk stratification and survival trajectory prediction.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure US2025046015_19032026_PF_FP_ABST
    Figure US2025046015_19032026_PF_FP_ABST
Patent Text Reader

Abstract

A computer-implemented system and method are provided for predicting consequences of obstructive sleep apnea. At least one computing device can access information associated with overnight polysomnography and can extract features from the information. The extracted features are across three physiological domains. Moreover, each of the features can be represented across the three physiological domains as histograms, that are applied to two extreme gradient boosting (XGBoost) models. Moreover, the computing device(s) can predict, as a function of the two XGBoost models, short-term excessive daytime sleepiness (EDS) and long-term all-cause mortality. The histograms can be applied by the at least one computing device to a gradient-boosting survival model and the at least one computing device can evaluate, as a function of the gradient-boosting survival model, survival trajectories, and can evaluate, using artificial intelligence, feature importance of the information associated with overnight polysomnography for predicting a consequence of obstractive sleep apnea.
Need to check novelty before this filing date? Find Prior Art

Description

OBSTRUCTIVE SLEEP APNEA OUTCOME PREDICTION SYSTEM AND METHODCROSS-REFERENCE TO RELATED PATENT APPLICATIONS

[0001] This application is based on and claims priority to U.S. Provisional Patent Application 63 / 694,166, filed September 12, 2024, as well as to U.S. Provisional Patent Application 63 / 818,449, filed June 5, 2025, each of which is hereby incorporated by reference, as if expressly set forth in its respective entirety herein.FIELD OF THE INVENTION

[0002] The present disclosure relates, generally, to data processing operations and management and, more particularly, to predictive capacities by applying machine learning and artificial intelligence.BACKGROUND

[0003] Obstructive sleep apnea is a common sleep disorder, which is thought to affect approximately one billion individuals worldwide. Affected individuals can be associated with short-term and long-term adverse effects, including excessive daytime sleepiness, neurocognitive impairment, and cardio-cerebrovascular morbidity. Known techniques to determine existence and severity quantification of obstructive sleep apnea rely on the apnea- hypopnea index, which factors the frequency of respiratory events defined by apneas (no airflow) and hypopneas (reduced airflow associated with either hypoxia or arousal). Severity can be variably defined by either event frequency or immediate physiological consequences of apneas and hypopneas, i.e., oxygen desaturation and arousal from sleep. The apnea-hypopnea index defines hypopneas by the presence of an associated desaturation or an arousal, rather than by relying on a ventilatory disturbance alone. Further, the value of the apnea-hypopnea index is influenced by several physiological variables (e.g., sleep position and sleep stage), device-dependent technological variables (e.g., absent or poor-quality signals), scoring rules, and by highly subjective determinations of “baselines” for 02 saturation and flow.

[0004] Unfortunately, the apnea-hypopnea index is subjective by its nature and known for having limitations when predicting short-term and long-term adverse outcomes. The apnea-hypopnea index is thus limited to describing the rate of consequence-defined respiratory events. In addition, apnea-hypopnea index fails to segregate ventilatory' disturbances from the consequences of hypopneas (such as arousals and / or desaturations as these are intrinsically tied to each definition of the apnea-hypopnea index). As such, apnea-hypopnea index is affected by multiple sources of potential “noise”. Indeed, remarkable inter-individual variability in type and duration of respiratory events exists across subjects with similar apnea-hypopnea index. While treatment of obstructive sleep apnea improves the common short-term sequelae of daytime sleepiness and impaired vigilance, degree of clinical improvement after treatment has not been predicted by the apnea-hypopnea index. Long-term consequences of obstructive sleep apnea , such as cerebrovascular or ventilatory distribution outcomes and neurocognitive outcomes, are also poorly predicted by baseline apnea-hypopnea index or changes in apnea- hypopnea index.

[0005] Over time, obstructive sleep apnea has become understood as a heterogeneous disorder with a complex underlying pathophysiology. Using apnea-hypopnea index as a sole metric when measuring the frequency of apnea and hypopnea events during sleep can be insufficient to capture this complexity. In other words, a single apnea-hypopnea index value is unlikely to characterize all domains of obstructive sleep apnea, limiting its ability to predict adverse outcomes.

[0006] Some alternative metrics that focus on different physiological domains impacted by obstructive sleep apnea (e.g., ventilation, hypoxia, and arousal) have recently been proposed to improve characterizing obstructive sleep apnea severity. Metrics in the respective domains can be independently associated with long-term adverse outcomes of obstructive sleep apnea. For example, both hypoxic burden (the area under the apnea / hypopneas associated- desaturation curve) and arousal burden (the cumulative duration of all arousal events relative to total sleep time) can be associated with the future all-cause and cardiovascular mortality. While these metrics can be useful alternatives for assessing obstructive sleep apnea severity' in each domain, none can fully represent an overall severity of the disorder.

[0007] It is with respect to these and other considerations that the disclosure made herein is presented.SUMMARY

[0008] In one or more implementation of the present disclosure, a computer- implemented system and method are provided for predicting consequences of obstructive sleep apnea. At least one computing device can access information associated with overnight polysomnography. The at least one computing device can extract features from the information, wherein the extracted features are across three physiological domains. Moreover, the at least one computing device can represent each of the features across the three physiological domains as histograms and apply the histograms to two extreme gradient boosting (XGBoost) models. Moreover, the at least one computing device can predict, as a function of the two XGBoost models, short-term excessive daytime sleepiness (EDS) and long-term all-cause mortality. The histograms can be applied by the at least one computing device to a gradient- boosting survival model and the at least one computing device can evaluate, as a function of the gradient-boosting survival model, survival trajectories, and can evaluate, using artificial intelligence, feature importance of the information associated with overnight polysomnography for predicting a consequence of obstructive sleep apnea.

[0009] These and other aspects, features, and advantages can be appreciated from the accompanying description of certain embodiments of the invention and the accompanying drawing figures and claims.BRIEF DESCRIPTION OF THE DRAWINGS

[0010] Various features, aspects and advantages of the invention can be appreciated from the following detailed description and the accompanying drawing figures, in which:

[0011] FIG. 1 illustrates signal flow, in accordance with an example implementation of the present disclosure;

[0012] FIG. 2 illustrates example breath amplitude and processed output, in accordance with an example implementation of the present disclosure;

[0013] FIG. 3 shows hypoxic distribution of two subjects, in accordance with an example implementation of the present disclosure;

[0014] FIG. 4 illustrates ventilatory burden percentages in connection with respective cohorts, , in accordance with an example implementation of the present disclosure;

[0015] FIG. 5 illustrates example bell shaped curves associated with ventilatory distribution for subjects, in accordance with an implementation of the present disclosure;

[0016] FIG. 6 illustrates distributions of example ventilatory burden percentages, in accordance with an example implementation of the present disclosure;

[0017] FIG. 7 shows relationships between ventilatory burden, apnea-hypopnea index, hypoxic burden, and anthropometric / demographic variables, in connection with a respective cohort, in accordance with an example implementation of the present disclosure;

[0018] FIG. 8 shows resulting survival probability curves over time, in connection with a respective cohort, in accordance with an example implementation of the present disclosure;

[0019] FIG. 9 shows resulting survival probability curves over time in connection with a respective cohort associated with an adjusted model, in accordance with an example implementation of the present disclosure;

[0020] FIG. 10 illustrates cross-validation and feature importance, in accordance with an example implementation of the present disclosure;

[0021] FIG. 11 shows model output associated with respective cohorts, in accordance with an example implementation of the present disclosure;

[0022] FIG. 12 illustrates cross-validation and feature importance in connection with model output, in accordance with an example implementation of the present disclosure;

[0023] FIG. 13 shows model output associated with mean AUROC values in connection with respective models, in accordance with an example implementation of the present disclosure;

[0024] FIG. 14 shows predicted survival trajectories, in accordance with an example implementation of the present disclosure;

[0025] FIG. 15 illustrates data assessment associated with sleep monitoring, in accordance with an example implementation of the present disclosure;

[0026] FIGS. 16, 17A, and 17B illustrate output associated with example tracing, in accordance with an example implementation of the present disclosure;

[0027] FIG. 18 illustrates a process and features, in an accordance with an example implementation of the present disclosure;

[0028] FIG. 19 illustrates an example medical clinical record and corresponding JSON extraction, generated in accordance with an example implementation of the present disclosure;

[0029] FIG. 20 is a diagram showing an example hardware arrangement that is configured for providing the systems and methods disclosed herein; and

[0030] FIG. 21 shows an example information processor and / or user computing device that can be used to implement the techniques shown and described herein.DETAILED DESCRIPTION OF THE EMBODIMENTS

[0031] By way of introduction, the present disclosure includes methods and systems applying a technological framework for configuring one or more computing devices to predict short-term and long-term medical outcomes associated with obstructive sleep apnea evaluation. The framework includes a framework comprising artificial intelligence and machine learning and processing a plurality of physiological features, including across ventilatory, hypoxic, and arousal domains, can be analyzed in an integrated fashion for evaluation and use in generating the predictions.

[0032] In one or more implementations of the present disclosure, one or more computing devices operate by applying a comprehensive and interpretable approach including risk stratification. For example, risk stratification can be based on the severity across physiological domains (e.g., poor ventilation but with low hypoxemia and high arousal “burden”). Risk stratification can include the frequency of respiratory events defined by apneas and hypopneas and an additional automated measure of ventilatory burden. As used herein, ventilatory burden generally represents a proportion of overnight breaths having less than 50% normalized amplitude, which can be also correlated with future prediction of all-cause and cardiovascular mortality. Physiologically, obstructive sleep apnea results in individuals’ multiple pathophysiological changes, such as severe hypoxemia, high arousals, and / or poor ventilation, rather than just one. By evaluating overall severity of obstructive sleep apnea across at least the three physiological domains, which are known to be impacted by obstructive sleep apnea, and evaluating obstructive sleep apnea association with relevant outcomes can be included in a basis for short-term and long-term assessment.

[0033] Accordingly, the present disclosure provides a framework that, when implemented, can harmonize respective domains of obstructive sleep apnea and overcomelimitations of known methods, such as linear interrelationships be tween features of combined measures. Such linearity of different domains is not particularly applicable in real-world operations. The present disclosure includes machine learning (ML) models, which are fashioned to model the non-linear relationships between the different components. In one or more implementations of the present disclosure, machine learning models can be an extension of non-linear regression models that are otherwise rarely used in statistical analyses. The present disclosure applies machine learning including in connection with complex multivariate survival models and can predict an individual's survival trajectory, thereby providing crucial information for intervention.

[0034] The present disclosure includes ML models that are developed to combine known physiological measures across three domains that are impacted by obstructive sleep apnea (ventilatory / hypoxic / arousal). Further, the present disclosure includes systems and methods to evaluate each of a plurality of domains’ respective utility to predict short-term and long-term outcomes of obstructive sleep apnea. Still further, the present disclosure includes a separate ML model that incorporates a measurement representing time-to-event and evaluates its utility to predict survival trajectories of obstructive sleep apnea patients.

[0035] Methodologies associated with the present disclosure are supported by and applicable to respective epidemiological cohorts and clinical cohorts. Data associated therewith can be processed to a) derive the normative range of ventilatory burden, b) asses the relationship between degree of upper airway obstruction and ventilatory burden, and c) assess the relationship between ventilatory burden and all-cause and cardiovascular (“Cventilatory distribution”) mortality with and without hypoxic burden (hypoxic burden). Results of the methodologies have been determined to be within 95th percentile of ventilatory burden in asymptomatic healthy subjects. Results for respective cohorts include 25.2% and 26.7% respectively (ventilatory burden), 5.5[3.5-9.7]%, ventilatory burden =9.8[6.4-15.6]%; median [interquartile range]). Ventilatory burden can be associated with the degree of upper airway obstruction in a dose-response manner (ventilatory burden untreated=31 .6(27.1)%, ventilatory burden treated=7.2(4.7)%, ventilatory burdens sub-optimally-treated=17.6(18.7)%, ventilatory burden off-treatment=41.6(18.1)%) and exhibited low’ night-to-night variability (ICC[2,l]=0.89). Ventilatory burden can be predictive of all-cause and C ventilatory distribution mortality in a respective cohort, before and after adjusting for covariates including hypoxic burden. While apnea-hypopnea index can be predictive of all-cause mortality, it is not required to be associated with C ventilatory distribution mortality in a respective cohort.

[0036] The present disclosure can include one or more computing devices configured by executing instructions, such as stored on processor readable media, to factor excessive daytime sleepiness as a short-term outcome The excessive daytime sleepiness can be measured, for example, by the Epworth Sleepiness Scale in respective datasets (e.g., Sleep Heart Health Study (SHHS) and Multi-ethnic Study of Atherosclerosis (MESA) datasets). A total score of equal to or greater than 10 can be coded as a positive outcome, or “excessive daytime sleepiness” (represented by a value of 1). Conversely, a score less than 10 can be coded as a negative outcome, or “non-excessive daytime sleepiness “ (represented by a value of 0). Data used can be derived from respective cohorts (e.g., 4,000+ subjects, 1200+ subjects, etc.) and (e.g., n = 1246). Flow of data is listed in Table 1, below.

[0037] Polysomnography signals of interest in connection with one or more implementations of the present disclosure, electroencephalography (EEG), airflow, and SpO2 (see, for example, FIG. 1). A signal of interest can include an overnight airflow signal using the Nasal Cannula / Pressure Transducer system, though not necessarily segmented into periods of sleep / wake, but by valid signal. This approach mimics steps for calculating the apnea-hypopnea index when an electroencephalogram (“EEG”) is not available. Alternatively, in cases where the nasal cannula is not available, airflow can be derived using a digital differentiator applied to the thoracic and abdominal respiratory inductance plethysmography (RIP) signal.Physiological measures - Ventilatory Distribution

[0038] The ventilatory distribution can be based on the ventilatory burden, in which the amplitude of each breath can be measured by averaging the flow during the middle one- third of inspiration to avoid errors from snoring or flow-limited breaths overnight. Apneas can be considered as breaths with zero amplitude at the current respiratory rate. A moving average of the breath’s mid-flow amplitude can be calculated using the previous 10 breaths, excluding those with excessive amplitude or signs of inspiratory flow limitation. Flow-limited breaths can be identified automatically based on their shape using an in-house algorithm. A breath’s amplitude can be normalized to this moving average, with “normal” breaths, i.e., no reduction in flow having a normalized amplitude of approximately 100%. Hypopncas shows a gradual decline in amplitude, while resuming flow after an event resulted in amplitudes greater than 100%. A histogram of breath amplitudes (ranging from 0 to 200%, in bins of 5%) then can be created and defined as ventilatory distribution.

[0039] In one or more implementations of the present disclosure, the amplitude of each single breath can measured as the average value of flow during the middle 1 / 3rdof inspiration (see, for example, FIG. 2, at A). Using the middle l / 3rdof inspiration for measuring breath amplitude can be effective to overcome erroneous measurements due to snoring and its associated artifacts, while avoiding short peaks of flow at the initiation of flow-limited breaths. In the case of apneas, breaths can be imputed as zero-amplitude breaths at the then-current respiratory rate, as described previously (see, for example, FIG. 2, at B). A moving “average” of breath mid-flow amplitude can be derived from ten prior breaths, not necessarily consecutive, that meet the following conditions: 1) an amplitude of less than 150% of the previously calculated average; devoid of suspicion of inspiratory' flow limitation determined by flow shape: i.e., a flow limited breath may not enter into calculation of the moving “average” breath amplitude or duration. Identification of breaths with inspiratory flow limitation by shape can be done automatically using an approach, such as shown and described herein. A given single breath’s amplitude can be then taken as the mid-flow amplitude normalized to themoving “average.” As an example, an initial run of equal amplitude non-flow limited (normal) breaths, after processing, would have approximately 100% normalized amplitude, and a gradual decline in amplitude shown during hypopneas (see, for example, FIG. 2, at C). The resumption of flow at the end of events often resulted in >100% amplitude breaths. A normalized histogram consisting of breath amplitudes ranging from 0 to a cap of 200% (bin- width=5) can be calculated and defined as the overnight ventilatory distribution. From this distribution, ventilatory burden can be defined as the percentage of overnight breaths with <50% amplitude (see for example, FIG. 2, at D). The 50% cutoff can be guided by empirical observation that the majority of breaths within “classical” hypopneas are shown to be <50% in a respective definition of amplitude.

[0040] Desaturation events associated with >=3% desaturation events can be identified automatically based on nadir SpO2 values and the peak SpO2 value prior to a desaturation. SpO2 signal can be pre-processed by discarding wake periods identified from sleep / wake manual scoring and removing artifacts (such as values outside the 56-100% oxygen saturation range). A Savitzky-Golay filter can be applied to fill in the gaps and create a continuous SpO2 signal. SpO2 nadirs can be automatically detected by looking for drops greater than 3% over at least 3 seconds, which can be considered potential “events.” The peaks before and after each event can be identified, and the area between them and the nadir computed. Hypoxic distribution (“HD”) can be then defined as the histogram of desaturation areas calculated above. FIG. 3 shows an example of two subjects, Subject A and Subject B, with different HD profiles. As described previously, the hypoxic distribution may be better able to discriminate individuals’ hypoxia profiles.Hypoxic Burden

[0041] Hypoxic burden can be defined on the saturation V time curv e as a cumulative sum of areas, each bounded below by a desaturation event (>3% oxygen desaturation), and above by the left and right peaks associated with that desaturation event, for example, as shown and set forth herein. A calculation of hypoxic burden can result from automatic identification of a >3% desaturation nadir and its associated peaks as opposed to the previously described approach wherein manually marked respiratory events can be used to identify corresponding desaturation events.Arousal Distribution

[0042] Using an arousal-related measure is useful to incorporate changes in EEG observed due to obstructive sleep apnea. For example, the derivation of an automated measure termed ASWAK, which characterizes the slow-wave / K-complex coupling, and can be a significant predictor of objective sleepiness. Briefly, ASWAK quantifies the bursts in delta activity surrounding K-complexes overnight. Given that K-complexes can have a significant role in arousal mechanisms, the delta bursts surrounding K-complex quantify the ability to maintain sleep despite arousal stimuli. ASWAK also shows a dose-response relationship with subjective sleepiness and lapses in vigilance. Since one of the outcomes of interest is sleepiness, which can be known also to impact all-cause mortality (either due to or despite of cardiovascular risk), ASWAK can be included as an “arousal burden” measure. The overnight ASWAK value can be represented as a histogram and defined as the arousal distribution (see, for example, FIG. 3), which shows the example of two subjects, Subject A and Subject B, with different arousal distribution profiles.

[0043] In one or more implementations of the present disclosure, several ML models (including Logistic regression, Support Vector Machines, and Gradient Boosted Trees) can be selected, and the XGBoost model can be used for evaluating a prediction of binary' outcomes (excessive daytime sleepiness and all-cause mortality), based on the highest Area Under Operator Receiver Curve (AUROC) value on predicting short- and long-term outcomes. See, for example, Table 2, below.The predictive performances of different machine learning models treated with ventilatory / hypoxic / arousal distributions on binary' short- and long-term outcomesAbbreviations: AUROC, Area Under Operator Receiver Curve; EDS, excessive daytime sleepiness.

[0044] XGBoost provides a suitable decision-tree-based ensemble ML algorithm that uses a gradient-boosting framework. A set of weak learners can be built on subsets of data to create a strong learner and currently dominates in performance for structured or tabular datasets in classification predictive modeling tasks. For apnea-hypopnea index model, the univariate Logistic Regression model can be trained for both outcomes. Note that the apnea-hypopnea index value includes the apnea-hypopnea index that includes respiratory events associated with either >=3% desaturation and / or EEG arousal. In addition, the gradient-boosting survival (GBS) model can be employed to predict survival trajectories from the same inputs (ventilatory, hypoxic, and arousal distributions). Due to class imbalance in the short- and longterm outcomes (e.g., excessive daytime sleepiness and all-cause mortality), where the positive cases are significantly fewer than negative cases, the Synthetic Minority Oversampling Technique (SMOTE) can be applied to achieve a reasonable class balance (40-60%) for the training sample. The 10-fold cross-validation can be used for validating the models. For excessive daytime sleepiness prediction, MESA dataset can be used as a test set for external validation. All model training and validation can be performed using PYTHON.Features

[0045] In one implementation of the present disclosure, a total of 60 combined metrics can be derived from PSG data representing ventilatory distribution (20 features), hypoxic distribution (20 features), and arousal distribution (20 features) as input features, which can be used for training the ML models for binary outcome predictions (i.e., excessive daytime sleepiness and all-cause mortality). For the survival ML model in predicting all-cause mortality, time-to-event can be added. Considering the significant associations between demographics and the outcomes of obstructive sleep apnea, baseline demographics (age, gender, BMI, and smoking status) can be also included in a separate survival ML model. To compare the performance between combined metrics and single value of apnea-hypopnea index, a separate model also can be evaluated that used apnea-hypopnea index as the sole predictor.Model Evaluations

[0046] The performances of ML models with binary outcomes can be evaluated using the Area Under the Receiver Operating Characteristic Curve (AUROC). The AUROC can be a useful performance metric for binary classification tasks. It compares the false positive rate (1 - specificity) with the true positive rate (sensitivity) across various predicted risk score thresholds. An AUROC value greater than 0.9 indicates that the classifier performs excellently in distinguishing between the two classes. In addition, SHapley Additive exPlanations (SHAP) analysis can be used to assess feature importance in the prediction models. SHAP can be based on the game theoretically optimal Shapley values and commonly used to evaluate the feature significance in non-linear models, such as XGBoost. Still further, SHAP values can be used for model interpretation to determine the most important features for predicting each adverse outcome.

[0047] For evaluating the ML model where the output can be a survival trajectory across time, the C-index, time-dependent AUROC, and Integrated Brier score can be used. The C -index metric can be used for summarizing model performance across different configurations, accounting for both the occurrence and timing of outcomes. A C-index of 0.5 indicates random predictions, 1.0 denotes perfect predictions, and 0.0 represents consistently incorrect predictions. The time-dependent AUROC for survival extends the traditional AUROC analysis to handle time-varying disease status, where sensitivity and specificity become time-dependent measures. The Integrated time-dependent Brier score, an extension of the mean squared error for right-censored data, assesses the accuracy of survival model over time. Further, the Integrated Brier score can be calculated using time points between the 10th and 90th percentiles, and the lower score (close to 0) indicated better predictive accuracy.Results

[0048] A total of 5262 subjects, including 4016 subjects from SHHS (mean ± SD age: 63.4 ± 11.2 yrs) and 1246 from MESA (age: 69.1 ± 8.9 yrs), can be analyzed. See Table 1 for details of the data flow and breakdown. Table 3, below, shows the subject characteristics and the polysomnographic features.

[0049] Out of 4016 subjects in a respective cohort, 4013 have survival time data (927(23.1%) have a fatal event), whereas 3551 have ESS data (1006 (30.3%) had excessive daytime sleepiness). In another respective cohort, 235 out of 1246 (18.9) subjects have excessive daytime sleepiness.

[0050] Subject characteristics 1, DAYFUNNormal, DAYFUN Symptomatic and the SHHS datasets are shown in Table 4 and Table 5, respectively, and a total of 34,446,696 breaths across all subjects (N=5182) analyzed. Derivation of the overnight ventilatory distribution and the resulting value of ventilatory burden took an average 2.2±1.1 minutes, respectively.Ventilatory Burden In Asymptomatic Healthy Subjects

[0051] Ventilatoiy burden for subjects in a respective cohort can be 5.5[3.5— 9.7]%. FIG. 4. for example, shows a median (interquartile range), which indicates a median value of the proportion of overnight breaths with less than 50% normalized amplitude. Further, in another cohort, the 95th percentile of ventilatory burden can be 25.2%, indicating that 95% of subjects have less than 25.2% of overnight breaths with less than 50% normalized amplitude. Ventilator}' burden for subjects have been shown to be 9.8[6.5-15.6]% and the 95th percentile of ventilator}' burden 26.7% (see, for example, FIG. 4).

[0052] The average shape of the ventilatory distribution for subjects can be determined to be a bell-shaped curve centered around 100% normalized amplitude (see, for example, FIG. 5, at A). The average shape of the ventilatory distribution for subjects in DATTUNNormal cohort can be determined to be also a bell-shaped curve centered around 100% normalized amplitude (see, for example, FIG. 5, at B) and similar to that observed in the Sau Paulo Epidemiologic Sleep StudyNormal cohort (similarity=0.49).Dose-Response Relationship Between Ventilatory Burden and Severity of Obstruction Altered by Therapeutic and Sub-Therapeutic CPAP

[0053] In the DAYFUN Symptomatic cohort, at baseline (pre-treatment), newly diagnosed obstructive sleep apnea subjects have a ventilatory burden of 31.6[i9.1-46.2]% (median[iqr]), which decreased to 7.2[4.8-9.5]% after three months of CPAP treatment, increased to 17.6[7.2- 25.9]% on suboptimal CPAP and to 41.6[26.9-44.9]% after two-night off CPAP (see, for example, FIG. 6). The shape of the ventilatory distribution in DAYFUN Symptomatic subjects at baseline indicated a greater proportion of apneic and hypopneic breaths and a diminished peak at 100% normalized amplitude (see, for example, FIG. 5, at C). The ventilatory distribution appeared to normalize, becoming more similar to that in DAYFUNNormal and the Sau Paulo Epidemiologic Sleep StudyNormal cohorts with CPAP treatment (similarity DAYFUN=0.45, similarity Sau Paulo Epidemiologic Sleep Study=0.74; (see, for example, FIG. 5, at D). On suboptimal CPAP, the shape of the ventilatory distribution indicated a greater proportion of apneic and hypopneic breaths than that observed on optimal CPAP (see, for example, FIG 5, at E). The shape of the average overnight ventilatory distribution after two-night CPAP withdrawal can be determined to be similar to that at baseline (similarity=0.64, (see, for example, FIG. 5, at F). Pairwise differences assessed using estimatedmarginal means indicated that, except for the baseline (diagnostic) vs. off-CPAP pair, all other pairs can be shown to be significantly different (see, for example, FIG. 5, at B; Table 6).Night-To-Night Variability in Ventilatory Burden

[0054] Ventilatory burden exhibited low night-to-night variability across in-lab (N=70) and at-home (N=135) studies (ICCin-lab=0.89, ICCat-home=0.86).Relationship Between Ventilatory Burden, Apnea-Hypopnea Index, and Hypoxic Burden

[0055] Across the DAYFUN Symptomatic cohort, ventilatory burden can be determined to be moderately correlated with apnea-hypopnea index (rho SHHS=0.29, p<0.01; rho DAYFUN=0.6, p<0.01). The hypoxic burden metric can be determined to be moderately correlated with ventilatory burden (rho=0.21, p<0.01), but correlated more strongly with apnea-hypopnea index (rho=0.76, p<0.01) in the SHHS cohort. The relationship between ventilatory burden, apnea-hypopnea index, hypoxic burden, and anthropometric / demographic variables (age, gender, BMI) in the SHHS cohort is depicted in FIG. 7. Ventilatory burden can be determined to be correlated with age (rho=0.24, p<0.001), BMI (rho=0.12, p<0.001), but not gender. Due to the selection bias of low apnea-hypopnea index no correlation between apnea-hypopnea index, ventilatory burden, and hypoxic burden can be assessed, e.g., in cohorts.

[0056] In a respective cohort, N=497 subjects having a low overall apnea-hypopnea index (<5) demonstrate REM-predominant obstructive sleep apnea (REM-apnea-hypopnea index > 15; with 3% desaturations and / or arousal). Across these subjects, ventilatory burden is determined to be greater than that observed in respective cohorts (27.9[19.8-37.2]%).Relationship Between Ventilatory Burden and Daytime Sleepiness and Hypertension

[0057] Ventilatory' burden is determined to be significantly higher in sleepy subjects (N=1467) as compared to non-sleepy subjects (30.7 vs. 32. 1; z=-2.48, p=0.01; median), as well as in hypertensive subjects (N=2,042) as compared to normotensive subjects (29.3 vs 33.2; z=- 8.6; p<0.001; median). After adjusting for known covariates, logistic regression models indicated that ventilatory burden quintiles 3-5 (compared to quintile 1) can be associated with greater odds of sleepiness (global p-value = 0.02; see details in supplement; Table 7).

[0058] Additionally, apnea-hypopnea index quintiles 4-5, (compared to quintile 1; global p-value <0.01; Table 7) and only quintile 5 of hypoxic burden (compared to quintile 1 ;global p-value = 0.01; Table 7) can be associated with greater odds of sleepiness. Adjusted logistic regression models also indicated that ventilatory burden quintiles 3 and 5 (compared to quintile 1) can be associated with greater odds of hypertension (global p-value <0.001; see details in supplement; Table 8).

[0059] Likewise, compared to apnea-hypopnea index quintiles 4 and 5 (compared to quintile 1; global p-value <0.001) and hypoxic burden quintiles 3-5 (compared to quintile 1; global p-value <0.001) can be associated with greater odds of hypertension (Table 8).Relationship Between Ventilatory Burden and All-Cause Mortality and C VentilatoryDistribution Mortality

[0060] Of 5,804 subjects having available PSG data through the National Sleep Research Resource (NSRR), 4,784 subjects have valid PSG data and covariates for all-cause mortality (1039 events), whereas 4,155 have valid PSG data and covariates for C ventilatory distribution mortality (298 events). The breakdown of missing covariates and poor-quality PSG data is listed in Table 9 of the supplement. The results of the Cox regression models using quintiles of key predictors are shown in Table 10.

[0061] Unadjusted models reported in Table 10 show that all quintiles (2-5) of apnea- hypopnea index, hypoxic burden, and ventilatory burden, compared to the first quintile as the reference, are associated with greater risk of all-cause mortality. In addition, the association of all quintiles of ventilatory burden, relative to quintile 1 , to all-cause mortality significant even after adjusting for all quintiles of hypoxic burden. Model 4 having both ventilatory burden and hypoxic burden quintiles resulted in the highest concordance of 0.63 and is determined to be a significantly better fit than the models with quintiles of either apnea-hypopnea index, hypoxic burden, or ventilatory burden alone, (ventilatory burden+hypoxic burden vs. Apnea-hypopnea index: z=-5.3, p<0.001; ventilatory burden+hypoxic burden vs. Hypoxic burden: LR=68.3, p<0.001; ventilatory burden+hypoxic burden vs ventilatory burden: LR=108.8, p<0.001; LR=likelihood ratio). After adjusting for known covariates (Table 10), while the hazard ratios can be attenuated across models, apnea-hypopnea index quintiles 3-5 (compared to quintile 1;global p-value = 0.02), all quintiles of hypoxic burden (compared to quintile 1; global p-value <0.001), and ventilatory burden (compared to quintile 1 ; global p-value < 0.001; before and after adjusting for hypoxic burden) is determined to be significantly associated with greater risk of all-cause mortality. Partial likelihood tests (or Anova for nested models) indicate that Model 4 having quintiles of both ventilatory burden and hypoxic burden can be a significantly better fit than any of the other models with quintiles of either apnea-hypopnea index, hypoxic burden, or ventilatory burden alone (ventilatory burden+hypoxic burden vs. Apnea-hypopnea index: z=-3.2, p<0.001: ventilatory burden+hypoxic burden vs. Hypoxic burden: LR=17.2, p<0.005; ventilatory burden+hypoxic burden vs ventilatory burden: LR=23.1, p<0.001; LR=likelihood ratio). Hazard ratios for all covariates in the fully adjusted Model 4 are reported in Table 11 and the resulting survival curves are depicted in FIG. 8.

[0062] Unadjusted models reported in Table 10 indicated that all quintiles (2-5) apnea- hypopnea index, hypoxic burden, and ventilatory burden, compared to the first quintile as the reference, can be significantly associated with greater risk of Cventilatory distribution mortality. As in the case of all-cause mortality, Model 4 consisting of quintiles of both ventilatory burden and hypoxic burden have the highest concordance of 0.65 and can be a significantly better model fit than any of the other models with quintiles of either apnea- hypopnea index, hypoxic burden, or ventilatory burden alone (ventilatory burden+hypoxic burden vs. Apnea-hypopnea index: z=-2.6, p<0.001; ventilatory burden+hypoxic burden vs.Hypoxic burden: LR=36.6, p<0.001; ventilatory burden+hypoxic burden vs ventilatory burden: LR=31.1, p<0.001; LR=likelihood ratio). After adjusting for known covariates, onlyventilatory burden quintiles 3-5, relative to quintile 1, can be significantly associated with greater risk of Cventilatory distribution mortality (global p-value = 0.02; global p-value after adjusting for hypoxic burden = 0.04). In contrast, none of the quintiles of apnea-hypopnea index and hypoxic burden can be significantly associated with Cventilatory distribution mortality after adjusting for known covariates. Partial likelihood ratio tests (or anova for nested models) indicated that the adjusted Model 4 consisting of quintiles of both ventilatory burden and hypoxic burden with the highest concordance of 0.86, can be a significantly better fit than any of the other models with quintiles of either apnea-hypopnea index, hypoxic burden, or ventilatory burden alone (ventilatory burden-hypoxic burden vs. Apnea-hypopnea index: z=- 1.6, p<0.05; ventilatory burden-hypoxic burden vs. Hypoxic burden: LR=10.5, p<0.05; ventilatory burden+hypoxic burden vs ventilatory burden: LR=9.3, p<0.05; LR=likelihood ratio). Hazard ratios for all covariates in the fully adjusted Model 4 are reported in Table 11 in the supplement and the resulting survival curves are depicted in FIG. 9. Of note, when using continuous modelling of key predictors (apnea-hypopnea index, hypoxic burden, ventilatory burden: details in supplement), only ventilatory burden has shown to be significantly associated with both all-cause and Cventilatory distribution mortality.Excessive Daytime Sleepiness and All-Cause Mortality Prediction

[0063] The ventilatory distribution / HD / arousal distribution model for excessive daytime sleepiness prediction achieved an AUROC value of 0.86 ± 0.01 in SHHS dataset (see, for example, FIG. 10, at A), and an AUROC value of 0.59 in the external validation (MESA dataset) (see, for example, FIG. 11). In comparison, the single feature apnea-hypopnea index model achieved only an AUROC of 0.56 ± 0.02 in SHHS and 0.52 in MESA dataset. FIG. 10, at B, shows the 20 most important features in predicting excessive daytime sleepiness based on average impact on the model output magnitude in SHAP analysis. The HD is shown to contribute the most to the prediction of the excessive daytime sleepiness outcome in SHHS dataset.

[0064] The ventilatory distribution / HD / arousal distribution model for all-cause mortality prediction achieved an AUROC value of 0.93 ± 0.01, which is significantly higher than the apnea-hypopnea index model with an AUROC of 0.58 ± 0.01 (see, for example, FIG. 12, at A). Similar to the excessive daytime sleepiness model, again the HD is shown tocontribute the most to the prediction of the all-cause mortality outcome, followed by arousal distribution in the SHHS dataset (see, for example, FIG. 12, at B).

[0065] Moreover, models can be trained with single domain features of ventilatory distribution, HD, or arousal distribution on binary short- and long-term outcomes prediction. The results showed that all models using single domain features performed poorly compared to the combined ventilatory distribution / HD / arousal distribution model, with the highest performance observed in HD model (Table 12).Survival Model for All-Cause Mortality

[0066] The number of fatal events across time are shown in Table 14. The ventilatory distribution / HD / arousal distribution model for predicting survival trajectories achieved an AUROC value of 0.63 ± 0.08 with a C-index score of 0.63 ± 0.06 (higher is considered better) and an Integrated Brier score of 0.10 ± 0.01 (lower is considered better), while the apnea- hypopnea index model achieved an AUROC value of 0.54 with a C-index score of 0.55 and an Integrated Brier score of 0.16 in the prediction (see, for example, FIG. 13). The addition of demographic variables to the survival trajectory ventilatory distribution / HD / arousal distribution model, increased the AUROC value of 0.80 ± 0.03, a C-index score of 0.79 ± 0.02, and an Integrated Brier score of 0.08 ± 0.01.The mortality number across the time

[0067] Exemplary predicted survival trajectories are depicted in FIG. 14. As seen in FIG. 14, at A, the predicted survival trajectory matches the true outcome of an individual who has a fatal event at 1293 days better than apnea-hypopnea index. The survival curves between the two models almost coincide with each other up to approximately 500 days, followed by a clear deviation with the combined metrics survival model predicting a lower surv ival probability over time compared to the apnea-hypopnea index model. FIG. 14, at B shows the survival curves for an individual who survived with a right censored day of 4626. The combined metrics survival model predicts a higher survival probability than the apnea- hypopnea index model across time. Although these are examples that illustrate the differences in the two models, it should be noted that the overall performance of the physiological ML model has shown to be significantly better than the apnea-hypopnea index model, supporting the determination that the physiological ML model is more consistent and accurate than the apnea-hypopnea index model.DatasetsConsecutive night Polysomnography

[0068] Data from a total of N=205 subjects, analyzed to assess the night-to-night variability of a VB metric, can be derived, for example, by sampling community-dwelling healthy elderly volunteers to examine sleep, aging and Alzheimer’s disease biomarkers. Full in-lab data from a total of N=70 subjects are shown as assessed (Table 15). The data are associated with an EEG, EOG, EMG, airflow using nasal cannula / pressure transducer system, finger pulse oximetry for oxygen saturation, respiratory effort using respiratory inductance plethysmography, ECG, snoring (acoustic microphone) and body position measurements. Apneas can be detected.

[0069] At-home polysomnographic data from a total of N=135 subjects can be assessed and sleep monitoring for 2 nights of unsupervised home sleep testing (HST), using HST equipment.accessed. Airflow can be recorded from a nasal cannula / pressure transducer system, oxygen saturation from finger (e.g., EMBLETTA MPR SLEEP SYSTEM), and snoring recorded by acoustic microphone. Effort can be measured using respiratory inductance plethysmography. Respiratory events can be scored using AASM rules to describe subjects. ICC and Bland altman analysis can be conducted to assess the night-to-night variability in VB (Table 15; see, for example, FIG. 15).Signal Pre-processing

[0070] It is to be appreciated that a signal of interest can be the overnight airflow signal using a nasal cannula / pressure transducer system, and not necessarily segmented by sleep / wake. All studies have airflow monitored in-lab. Since the oscillatory frequency of respiration is typically less than 4 Hz, all analyzed airflow signals can be sampled, eliminating a need for performing sampling rate conversion. Further, the lowest sampling rate across all the datasets can be, for example, 32 Hz, which can be sufficient for analyses of an inspiratory breath shape. Airflow signals can be used “as-is,” e.g., raw without any software / hardware filters. Clipped signals can pre-processed.

[0071] A consequence of using unfiltered and raw signals can be the presence of baseline drift, commonly observed during NPSG, particularly in cases of airflow measured using nasal cannula. Low-pass filtering, though often recommended in such situations, may not eliminate drift, particularly where inspiratory and expiratory airflow are asymmetrical. Accordingly, baseline removal or detrending of the raw unfiltered airflow signal can be performed. Individual breaths overnight can be segmented and a start of inspiration detected automatically across the data. Changepoints (crossings) can be assessed and adjusted, for example, based on inspiration time. A given breath can be adjusted to zero baseline using a moving average of five prior breaths. This detrended signal can then be used for development of VB. Periods greater than 120 seconds with zero flow can be discarded and not used for analysis. For studies with true flow measurement (e.g., CPAP) the CPAP flow signal can be square transformed before calculating overnight ventilatory distribution, which improves consistency with the nasal cannula derived breaths.Derivation of Airflow from RIP effort belts

[0072] The SHHS unattended at-home polysomnograms may not acquire a nasal cannula / pressure transducer system airflow. Airflow can be measured using the ProTech thermistor (M325) and apnea / hypopnea can be scored using a combination of this airflow and thoracic and abdominal RIP efforts belts. While this setup suffices for the derivation of AHI and other routine respiratory indices, the present disclosure addresses airflow from a nasal cannula / pressure transducer system. More particularly, for a given cohort airflow can be derived from the volume signal, calculated as the sum of the thoracic and abdominal RIP effort belts. Note that ideally the volume measured using RIP effort belts requires calibration constants that depend on patient size and other factors. The present disclosure sufficiently uses uncalibrated RIP effort signals for assessing change in airflow amplitudes. With the sum of thoracic and abdominal RIP effort signals, the airflow (RIPFlow) signal can be derived using low-pass digital filter. The derived RIPF low signals can be also square transformed to resemble the airflow from nasal cannula pressure transducer system. An example tracing is shown in FIG. 16 from a subject in the Sau Paulo Epidemiologic Sleep Study (normals) cohort. The Sau Paulo Epidemiologic Sleep Study polysomnography data have both the sum of thoracic and abdominal RIP effort signals and the nasal cannula pressure transducer measured airflow. Traditionally the RIP effort signals are sampled at a lower frequency than the nasal cannula pressure transducer system. As a result, the RIP Flow generally appears smoother whencompared to the nasal cannula pressure transducer system airflow signal (see, for example, FIG. 16). Across the 100 subjects in a given cohort, sample-by-sample correlation between the derived RIP Flow and the nasal cannula pressure transducer system signal has been shown to be 0.83±0.01 (mean±std), whereas across the 18 subjects in a different respective cohort, the sample-by-sample correlation has shown to be 0.72±0.06. An example tracing in the SHHS study is shown in FIG. 17A which depicts the derived RIPFlow along with the other respiratory and SpO2 signals. In one or more implementations, evaluating VB can include utilizing the inspiratory part of the airflow (derived or as-is).

[0073] FIG. 17B illustrates airflow signal estimation using accelerometer patches. Tire present disclosure can include deriving airflow signals from thoraco-abdominal movement signals as depicted in FIG. 17B. The top figure represents the derived flow signal using the accelerometer patches, the middle figure indicates the derived flow using thoraco-abdominal RIP belts, while the bottom figure indicates the true airflow using nasal cannula / pressure transducer system. The derivation of the derived flow signals is undertaken using maximally lowpass filters. As can be seen, the shape of the breaths is identical to true flow (nasal cannula).Objective Hypoxic Burden

[0074] Hypoxic burden can be defined as the total area under the desaturation curve associated with >=3% desaturation events that are identified automatically. Prior to calculation of the total area under the desaturation curve, the SpO2 signal can be pre-processed as follows. Periods of wake identified using manually scored sleepAvakc can be discarded. Artifacts that indicated non-physiological values can be also discarded (e.g., values outside 56-100% oxygen saturation). A Savitzky-Golay filter can be then applied to the discontinuous segments to create an interpolated and continuous SpO2 signal. SpO2 nadirs can be identified automatically for this pre-processed signal using a peak-prominence value of 3%, i.e., decreases in SpO2 by more than 3% over 3 seconds or more can be deemed as a candidate “event”. The left- and right-peak associated with these candidate events can be identified automatically from the pre- processed SpO2 signal. The area bounded by these peaks and the identified nadir can be then evaluated and the cumulative sum of such desaturation events can be then defined as the hypoxic burden. An example tracing with the shaded desaturation area is shown in FIG. 11 .Association Between VB And All-Cause And CVD Mortality Using Continuous Modelling Of Exposures

[0075] To avoid assumptions regarding linearity as well as allow parsimonious interpretation, the present disclosure includes an assessment of the relationship between the exposure (AHI, HB, or VB) and all-cause and CVD mortality by dividing the exposures into quintiles. Exposures can be modeled as continuous variables (main + quadratic term) and assessed their relationships with all-cause and CVD mortality as reported in Table 16. Analysis can be carried out in a similar fashion as in those reported in Table 10. The comparisons between models can be carried out using the method described in the Statistical Analysis section. In these models, to meet the proportional hazard assumptions, the following variables can be transfonned: age (square root), BMI (squared). Across all the models, only VB has been shown to be significantly associated with both all-cause and CVD mortality and the models with VB (likewise VB+HB) are determined to be a significantly better fit than the models with AHI (likewise HB). The hazard ratios for all covariates in Model 4 (VB-HB) is reported in Table 17.Relationship Between VB and Daytime Sleepiness and Baseline Hypertension

[0076] Logistic regression models can be constructed to assess the relationship between VB and daytime sleepiness and baseline hypertension. For each of the models, global p-values can be evaluated using Wald chi-squared tests and the corresponding variance-covariance matrices. A cut-off of Epworth Sleepiness Scale (ESS) >=10 can be used to indicate whether a subject is sleepy or not. Baseline hypertension status (yes / no) can be derived using 2nd and 3rd blood pressure readings and / or use of hypertension medications. Known covariates can be also entered into the logistic regression models (age, BMI, gender, race, smoking status (only for hypertension) and time in bed). Wilcoxon rank-sum tests can be also conducted to assess the differences in VB among sleepy vs. fion-slccpy and hypertensive vs. normotensive subjects. Results of the logistic regression are shown in Table 7 and Table 8.Statistical Analyses

[0077] Statistical analyses can be performed using MATLAB (Mathworks, MA, USA) R2022a and R statistical software. Normality of data can be tested using the Shapiro-Wilk test. Non-parametric correlations can be used (Spearman’s), such as when data are not normally distributed. Mean and standard deviation are reported for normally distributed variables, whereas median and IQR are reported for non-normally distributed variables. Histogram comparison can be conducted by first calculating the chi-square distance29 between the two histograms and then converting the distance to a similarity metric defined as 1 / (1 + d(a,b)), where d(a,b) represents the chi-square distance between histograms a and b. Tire similarity metric ranges between 0 and 1, where values close to 0 indicate no similarity between the histograms. Two sample t-tests can be conducted to assess the difference between ventilatory burden across the Sau Paulo Epidemiologic Sleep Study and the DAYFUN (healthy normal cohorts). A linear mixed effects model can be used to assess the effect of CPAP (on / off) on ventilatory'' burden (vb ~ timepoint + (1 [subject)), with the reference timepoint as the baseline. Post-hoc comparisons can be carried out using estimated marginal means.

[0078] Unadjusted and adjusted Cox proportional models can be constructed to assess the relationship between ventilatory burden and C ventilatory distribution and all-cause mortality. Death from any cause in a respective cohort can be a primary outcome for the Cox proportional models. Moreover, a second set of models can be constructed using C ventilatory distribution mortality as a primary endpoint. C ventilatory distribution mortality in given cohorts, i.e., death due to coronary heart disease, can be determined using death certificates, autopsy results, subject’s physician and family or other proxy interviews. The mortality data as well as the polysomnography (PSG) data can be provided through the National Sleep Research Resource (NSRR). Covariates added can include age, gender, BMI, race, smoking status, time in bed, baseline hypertension, previous history' of congestive heart failure, and stroke based on previous literature and known associations with mortality. In total, four models can be constructed: apnea-hypopnea index (Model 1), hypoxic burden (Model 2), ventilatory burden (Model 3), and hypoxic burden+ventilatory burden (Model 4). The Concordance statistic can be used as a measure of goodness of fit for each of the models. Nested models can be compared by constructing an analysis of deviance (anova) table based on the log partial likelihood in R (anova.coxph). Partial likelihood tests can be also conducted for comparisons of non-nested models using a method (e.g., proposed by Fine), which assumes the null hypothesis that the non-nested models are equidistant in Kullback-Liebler metric applied to therank likelihood. The proportional hazards assumption can be validated using the Schoenfeld residual test. Variables failing to meet the proportional hazards assumption (p<0.05 using the Schoenfeld residual test) can be recursively binned into discrete groups of two, three, four, and so forth until the assumptions are met.

[0079] Apnea-hypopnea index, hypoxic burden, and ventilatory burden can be discretized into quintiles for help with interpretation without assuming linearity. Global p- values for each categorical variable can be created using Wald chi-squared test and the corresponding variance covariance matrices. Models 1-4 can be also analyzed using continuous (linear and / or quadratic) terms for all exposure variables (Table 16 and Table 17). Survival curves for ventilatory burden can be generated using the fully adjusted model for C ventilatory distribution and all-cause mortality (Model 4). Correlations between ventilatory burden, age, BMI, gender, apnea-hypopnea index, and hypoxic burden can be assessed using Spearman’s rho (non-normally distributed variables) or Pearson’s r (normally distributed variables).

[0080] The present disclosure accounts for non-linear interactions between the distinct pathophysiological domains of obstructive sleep apnea. The present disclosure goes beyond assigning equal weighting to all apnea and hypopnea events, irrespective of their physiological impact, including by providing a ML-based approach that combines the physiological measures across the most impacted domains in obstructive sleep apnea (ventilatory, hypoxic, and arousal), in a non-linear method to overcome the drawbacks of apnea-hypopnea index for predicting short- and long-term consequences in obstructive sleep apnea. Indeed, as evident here, across all models, the ML-based approach consistently outperforms the apnea-hypopnea index. Conceptually, the ventilatory distribution / HD / arousal distribution ML model uses the same components that are used in deriving the apnea-hypopnea index. Leveraging principles of explainable Al, the hypoxic domain has been shown to be particularly important in predicting short- and long-term outcomes. This aligns w’ith results on single domains of ventilatory distribution, HD, or arousal distribution prediction.

[0081] The apnea-hypopnea index (apnea-hypopnea index) can be a principal metric for identifying presence of obstructive sleep apnea and assessing severity. Unfortunately, the index has limitations, for example for capturing only the rate of respiratory events. The relationship between apnea-hypopnea index as an obstructive sleep apnea severity metric and long-term outcomes, generally, has been inconsistent. The present disclosure uses data from multiple epidemiological and clinical cohorts, for example, comprising over 5,000 subjects. One or more computing devices can be configured, for example, by executing programminginstructions stored on processor readable media, to determine that a measure of the ventilatory burden (ventilatory burden) in obstructive sleep apnea is strongly predictive of Cventilatory distribution and all-cause mortality and is associated with daytime sleepiness and hypertension. Accordingly, a normative range of ventilatory burden can be established using data from a plurality of cohorts of healthy subjects. The relationship of ventilatory burden to changes in the treatment of obstruction of the upper airway can be demonstrated to support a biological plausibility of using ventilatory burden as a metric to characterize obstructive sleep apnea independent of the immediate consequences of intermittent hypoxia or sleep fragmentation from each event.

[0082] The apnea-hypopnea index, particularly the apnea-hypopnea index 4%, which indicates the rate of hypopneas associated 4% or greater desaturations, implicitly is a combination of two immediate consequences of obstructive sleep apnea: ventilatory changes and associated intermittent hypoxia. Defining and characterizing obstructive sleep apnea accurately and independently across these domains, rather than combining them into a single metric, can provide greater flexibility to determine utility and possibly overcome limitations of the apnea-hypopnea index. The present disclosure addresses durations of apnea / hypopnea events, which can be used to evaluate a number of abnormally decreased breath amplitudes analogous to ventilatory burden. It is recognized herein that duration of events is not easily generalizable given significant heterogeneity in the duration of events across subjects. Along the hypoxia domain, metrics of desaturation severity, including hypoxic burden evaluated using manually marked respiratory events, are usable to contributing predictors of short- and long-term outcomes in obstructive sleep apnea. The model with combined ventilatory burden and hypoxic burden can have the highest concordance (goodness of fit), in fully adjusted models and improves the models with either apnea-hypopnea index or hypoxic burden alone. This supports using disentangled, objective, and generalizable metrics across the ventilatory and hypoxic domains as opposed to using the apnea-hypopnea index’s implicit, but fixed, combination.

[0083] The normative value observed for ventilatory burden (approximately 25% or lower) suggests that up to 25% of breaths during sleep may have reduced amplitude, a threshold below which adverse outcomes are unlikely: this threshold therefore can be used to exclude / define disease (obstructive sleep apnea). Moreover, the upper limit of the first quintile for ventilatory burden in the SHHS cohort can be shown to be 19.6%, and as a result the comparisons of hazard ratios in the Cox proportional models reported here are, in essence, acomparison of hazard ratios of abnormal to normal subjects. Further, in a respective cohort, the mean ventilatory' burden in subjects without obstructive sleep apnea (apuca-hypopnca index=0, N=86) can be 23.9%, which is below the 95th percentile of ventilatory burden observed in healthy subjects across other respective cohorts. This supports validity' of the normative range of ventilatory' burden derived in accordance with the teachings herein. The changes in ventilatory' burden observed with changes in ventilation using CPAP are expected and are similar to those usually observed with apnea-hypopnea index, but enhance the biological plausibility of ventilatory burden.

[0084] An apnea-hypopnea index has been reported to have a greater night-to-night variability' with ICC values ranging between 0.71 - 0.87 depending on the definition of apnea- hypopnea index (e.g., 3% vs. 4% desaturations, supine vs. non-supine, etc.). Furthermore, depending on the cut-offs as well as definitions used for apnea-hypopnea index, the designation of patients into severity categories can include night-to-night variability. In contrast, night-to- night variability in ventilatory' burden can be shown to be low across both in-lab and at-home polysomnography, further supporting its use in clinical management of obstructive sleep apnea.

[0085] In a respective cohort, REM-predominant obstructive sleep apnea (overall apnea-hypopnea index < 5) subjects show a significantly higher ventilatory' burden than those across healthy subjects, which suggests a possible utility of a single metric of ventilatory burden in identifying REM-predominant obstructive sleep apnea as opposed to the use of multiple metrics currently employed. In subjects with REM-predominant obstructive sleep apnea, the overall apnea-hypopnea index can be lower due to the lower proportion of REM sleep to total sleep. In contrast, the current derivation of ventilatory burden and the does not require sleep / wake scoring and breaths during R EM / NR EM are pooled together, thus providing a better description of the total underly'ing respiratory' disturbance during the night. The present disclosure recognizes the clinical significance of REM-predominant obstructive sleep apnea, and contributes to establishing an accurate description of the disorder, including to better assess the underly'ing ventilatory disturbance.

[0086] In operation and as noted herein, the present disclosure includes distributions of the ventilatory, hypoxic, and arousal domains for predicting short-term and long-term outcomes of obstructive sleep apnea, as opposed to a single aggregated value, such as hypoxic or ventilatory' burdens. The present disclosure includes technology to associate measurements across the three domains, whether separate or combined linearly, with long-term outcomes. This can generate a predictive power that outperforms apnea-hypopnea index. Indeed, thepresent disclosure reveals an increased predictive power when an input to an ML model is represented as a histogram rather than a single value, as more nuanced information can be captured that can reveal patterns not evident in a single summary measure. In other words, although two subjects may have the same value for hypoxemia, the underlying patterns of their hypoxemia could be quite different and not illustrated in a single hypoxic burden value, potentially leading to distinct clinical outcomes. Thus, by accounting for the variability within each domain, the present disclosure provides a more nuanced understanding of the multifactorial nature of obstructive sleep apnea and its impact on both short-term and longterm health outcomes.

[0087] Moreover, the most relevant features for predicting excessive daytime sleepiness are primarily related to the hypoxic domain, while long-term mortality prediction can be most significantly associated with both the hypoxic and arousal domains. These findings are consistent with previous research that highlights the major role of intermittent hypoxia in the pathophysiology of obstructive sleep apnea and its relationship to sleepiness and cardiovascular associated mortality. Hypoxia, due to repeated oxygen desaturation events during sleep, can lead to disruptions in sleep architecture and neurocognitive function, which may explain its strong association with excessive daytime sleepiness. On the other hand, the involvement of the arousal domain in long-term mortality outcomes aligns with the growing body of evidence suggesting that frequent arousals contribute to cardiovascular morbidity and mortality in obstructive sleep apnea patients. Nocturnal hypoxemia with arousals has been determined to increase ventricular repolarization lability, which might elevate the risk of arrhythmias and sudden cardiac death in patients with obstructive sleep apnea. Ventilatory domain features consistently are determined to be among the top features across all models, albeit not ranked highest. One explanation might be that the ML models prioritize the response (hypoxemia or arousal) to the stimulus (ventilatory). That is, for a given level of ventilatory features, hypoxia and / or arousal emerge as the most important features, suggesting a non-linear adjustment by the models. Future research should further explore the interactions between these domains to refine predictive models and better target therapeutic interventions.

[0088] The present disclosure further includes adding factors including age, sex, smoking, and BMI to a ML model containing predominantly physiological measures. This significantly improves the performance (increase in AUROC from 0.63 to 0.80). Age has long been recognized as a risk factor for obstructive sleep apnea , with older individuals typically experiencing more collapsible airway and a higher mortality risk. Sex differences can also playa critical role, with men more likely to develop obstructive sleep apnea at younger ages while the prevalence in women increases after post-menopause likely due to hormonal changes, and there are sex differences in both short- and long-term outcomes of obstructive sleep apnea. In addition, smoking and BMI are well-established modifiable risk factors for obstructive sleep apnea development and progression. The incorporation of these demographic features into the model provides a more comprehensive view of the factors influencing obstructive sleep apnea, enhancing the model’s ability to predict both short- and long-term outcomes. These results suggest that models that integrate both physiological data and demographic information may offer more accurate risk stratification and help guide personalized treatment strategies.

[0089] In one or more implementations of the present disclosure, one or more models such as transformers can analyze raw time series signals of an entire night’s data collection. Moreover, other PSG features relating to mortality, such as heart rate variability, can also be included in one or more models.

[0090] FIG. 18 illustrates a process and features in an accordance with an example implementation of the present disclosure, including clinical practice notes processed, including to de-identify patient information, such a name, medical record number, date of birth, or other information, including by using named entity recognition. Chunking processes can operate on the clinical notes to filter for relevant information, such as a patient reporting occasional snoring but no witnessed apneas or gasping episodes. Embeddings for each of a plurality of chunks can be generated and stored, for example, in data storage accessible by applications employing large language models. An embedding model can process symptom information associated with snoring, dry mouth, apneas, sleepiness, congestion, and other relevant data to generate vectors. Further operations can be performed, such via the FEISS library, for similarity searching and vector clustering, including as a function of cosine similarity vis-a-vis the symptom information. Moreover, the top-K embeddings for each symptom can be mapped to chunks and provided in a particular context, including for predictive measures. A large language model can be prompted to generate structured outputs. For example, a prompt can be generated for the LLM directing to clinical progress notes and determinations of existence of respective symptoms being reported. Boolean variables (e.g., True vs. False) can be assigned for each. Prompting the LLM can further include restrictions from subjective assumptions and a requirement only for explicitly stated information set forth in the clinical progress notes. Further, chunks can be provided to artificial intelligence models, such as LLAMA, formultimodal, reasoned, and contextual output. Moreover, the present disclosure can include post processing, such as to generate and extract structured JSON.

[0091] FIG. 19 illustrates an example medical clinical record and corresponding JSON extraction, generated in accordance with an example implementation of the present disclosure. Symptoms output that the LLM identified as present include snoring, sleepiness, feeling unrested during the day, and congested nose. In one or more implementations, output can be generated in accordance with the present disclosure that represents certain testing did not occur, such as for sleep apnea. The output can be usable to gauge accuracy and effectiveness of the LLM, such as to recognize that certain symptoms may not be recognized by the LLM nor represented in the output.

[0092] Referring to FIG. 20, a diagram is provided that shows an example hardware arrangement 2000 that is configured for providing the systems and methods disclosed herein and designated generally as system 2000. System 2000 can include one or more information processors 2002 that are at least communicatively coupled to one or more user computing devices 2004 across communication network 2006. Information processors 2002 and user computing devices 2004 can include, for example, servers, workstations (e.g., desktop and notebook computers), mobile computing devices (e.g., tablets and smartphones), networking devices (e.g., hubs, repeaters, bridges, switches and routers), and various controller devices configured with electronics (e.g., integrated circuits and packages). Further, one computing device may be configured as an information processor 2002 and a user computing device 2004, depending upon operations being executed at a particular time.

[0093] With continued reference to FIG. 20, information processor 2002 can be configured to access one or more databases for the present disclosure, including source code repositories and other information. However, it is contemplated that information processor 2002 can access any required databases via communication network 2006 or any other communication network to which information processor 2002 has access. Information processor 2002 can communicate with devices comprising databases using any known communication method, including a direct serial, parallel, universal serial bus (“USB”) interface, or via a local or wide area network.

[0094] User computing devices 2004 can communicate with information processors 2002 using data connections 2008, which are respectively coupled to communication network 2006. Communication network 2006 can be any data communication network. Dataconnections 2008 can be any known arrangement for accessing communication network 2006, such as the public internet, private Internet (e.g., VPN), dedicated Internet connection, or dialup serial line interface protocoL''point-to-point protocol (SLIPP / PPP), integrated services digital network (ISDN), dedicated leased-line sen-ice, broadband (cable) access, frame relay, digital subscriber line (DSL), asynchronous transfer mode (ATM) or other access techniques.

[0095] User computing devices 2004 have the ability to send and receive data and to provide received data and processed data across one or more networks. Arrangement 2000 can include software that provides functionality described in greater detail herein, and preferably resides on one or more information processors 2002 and / or user computing devices 2004. One of the functions performed by information processor 2002 is that of operating as a web server and / or a web site host. Information processors 2002 typically communicate with communication network 2006 across a permanent (i.e. un-switched) data connection 2008. Permanent connectivity ensures that access to information processors 2002 is always available.

[0096] FIG. 21 shows an example information processor 2002 and / or user computing device 2004 that can be used to implement the techniques described herein. The information processor 2002 and / or user computing device 2004 is intended to represent various forms of digital computers, such as laptops, desktops, workstations, personal digital assistants, seivers, blade servers, mainframes, and other appropriate computers. The components shown in FIG. 21, including connections and relationships, and their functions, are meant to be exemplary only, and are not meant to limit implementations of the inventions described and / or claimed in this document.

[0097] As shown in FIG. 21, the information processor 2002 and / or user computing device 2004 includes a processor 2102, a memory 2104, a storage device 2106, a high-speed interface 2108 connecting to the memory 2104 and multiple high-speed expansion ports 2110, and a low-speed interface 2112 connecting to a low-speed expansion port 2114 and the storage device 2106. Each of the processor 2102, the memory' 2104, the storage device 2106, the highspeed interface 2108, the high-speed expansion ports 2110, and the low-speed interface 2112, are interconnected using various busses, and can be mounted on a common motherboard or in other manners as appropriate. The processor 2102 can process instructions for execution, for example stored on processor-readable media (including, but not limited to non-transitory processor readable media) accessible by the information processor 2002 and / or user computing device 2004, such as instructions stored in the memory 2104 or on the storage device 2106 to display graphical information for a GUI on an external input / output device, such as a display2116 coupled to the high-speed interface 2108. In other implementations, multiple processors and / or multiple buses can be used, as appropriate, along with multiple memories and types of memory. Also, multiple computing devices can be connected, with each device providing portions of the necessary operations (e.g., as a server bank, a group of blade servers, or a multiprocessor system).

[0098] The memory 2104 stores information within the information processor 2002 and / or user computing device 2004. In some implementations, the memory 2104 is a volatile memory unit or units. In some implementations, the memory 2104 is a non-volatile memory unit or units. The memory 2104 can also be another form of computer-readable medium, such as a magnetic or optical disk.

[0099] The storage device 2106 is capable of providing mass storage for the information processor 2002 and / or user computing device 2004. In some implementations, the storage device 2106 can be or contain a computer-readable medium, e.g., a computer-readable storage medium such as a floppy disk device, a hard disk device, an optical disk device, or a tape device, a flash memory or other similar solid-state memory device, or an array of devices, including devices in a storage area network or other configurations. A computer program product can also be tangibly embodied in an information carrier. The computer program product can also contain instructions that, when executed, perform one or more methods, such as those described above. The computer program product can also be tangibly embodied in a computer- or machine-readable medium, such as the memory 2104, the storage device 2106, or memory on the processor 2102.

[0100] The high-speed interface 2108 can be configured to manage bandwidthintensive operations, while the low-speed interface 2112 can be configured to manage lower bandwidth-intensive operations. Of course, one of ordinary skill in the art will recognize that such allocation of functions is exemplary only. In some implementations, the high-speed interface 2108 is coupled to the memory 2104, the display 2116 (e.g., through a graphics processor or accelerator), and to the high-speed expansion ports 2110, which can accept various expansion cards (not shown). In an implementation, the low-speed interface 2112 is coupled to the storage device 2106 and the low-speed expansion port 2114. The low-speed expansion port 2114, which can include various communication ports (e.g., USB, Bluetooth, Ethernet, wireless Ethernet) can be coupled to one or more input / output devices, such as a keyboard, a pointing device, a scanner, or a networking device such as a switch or router, e.g., through a network adapter. Accordingly, the automated methods described herein can be implemented by invarious forms, including an electronic circuit configured (e.g., by code, such as programmed, by custom logic, as in configurable logic gates, or the like) to carry out steps of a method. Moreover, steps can be performed on or using programmed logic, such as custom or preprogrammed control logic devices, circuits, or processors. Examples include a programmable logic circuit (PLC), computer, software, or other circuit (e.g., ASIC, FPGA) configured by code or logic to carry out their assigned task. The devices, circuits, or processors can also be, for example, dedicated or shared hardware devices (such as laptops, single board computers (SBCs), workstations, tablets, smartphones, part of a server, or dedicated hardware circuits, as in FPGAs or ASICs, or the like), or computer servers, or a portion of a server or computer system. The devices, circuits, or processors can include a non-transitory computer readable medium (CRM, such as read-only memory (ROM), flash drive, or disk drive) storing instructions that, when executed on one or more processors, cause these methods to be carried out.

[0101] Accordingly, the present disclosure includes one more computing devices, including configured for providing a framework to harmonize ventilatory / hypoxic / arousal domains of obstructive sleep apnea for the future prediction of short- and long-term outcomes. Modeling reveals that combined metrics offered better predictive performance for both short- and long-term obstructive sleep apnea outcomes compared to apnea-hypopnea index. Additionally, relevant features for predicting excessive daytime sleepiness are determined to be related primarily to the hypoxic domain, while the most significant features for predicting long-tenn mortality are associated with both the hypoxic and arousal domains.

[0102] The terminology used herein is for the purpose of describing particular embodiments only and is not intended to be limiting of the invention. As used herein, the singular forms “a,” “an” and “the” are intended to include the plural forms as well, unless the context clearly indicates otherwise. It will be further understood that the terms “comprises” and / or “comprising,” when used in this disclosure, specify the presence of stated features, integers, steps, operations, elements, and / or components, but do not preclude the presence or addition of one or more other features, integers, steps, operations, elements, components, and / or groups thereof.

[0103] Particular embodiments of the subject matter described in this specification have been described. Other embodiments are within the scope of the following claims. For example, the actions recited in the claims can be performed in a different order andstill achieve desirable results. As one example, the processes depicted in the accompanying figures do not necessarily require the particular order shown, or sequential order, to achieve desirable results. In certain implementations, multitasking and parallel processing may be advantageous.

Claims

WHAT IS CLAIMED:1 . A computer-implemented method for predicting consequences of obstructive sleep apnea, the method comprising: accessing, by at least one computing device, information associated with overnight polysomnography: extracting, by the at least one computing device, features from the information, wherein the extracted features are across three physiological domains; representing, by the at least one computing de vice, each of the features across the three physiological domains as histograms; applying, by the at least one computing device, the histograms to two extreme gradient boosting (XGBoost) models; predicting, by the at least one computing device as a function of the two XGBoost models, short-term excessive daytime sleepiness (EDS) and long-term ail-cause mortality; applying, by the at least one computing device, the histograms to a gradient-boosting survival model: evaluating, by the at least one computing device as a function of the gradient-boosting survival model, survival trajectories; and evaluating, by the at least one computing device using artificial intelligence, feature importance of the information associated with overnight polysomnography for predicting a consequence of obstructive sleep apnea.

2. The method of claim I, further comprising: combining, by the at least one computing device, physiological measures across impacted domains in obstructive sleep apnea in a non-linear method.

3. The method of claim 1, wherein the impacted domains include ventilatory, hypoxic, and arousal.

4. The method of claim I, wherein predicting the consequence includes providing inputs represented as histograms to at least one machine learning model, and further comprising: processing output of the at least one machine learning model to reveal patterns.

5. The method of claim 4, wherein the patterns are not evident in a single summary measure.

6. The method of claim 1 , wherein predicting the consequence includes at least one distinct clinical outcome.

7. The method of claim 1 , wherein the at least one distinct clinical outcome is based on a function of a multifactorial nature of obstructive sleep apnea.

8. The method of claim 1, wherein the predicting includes at least one of shortterm outcome and long-term health outcome.

9. The method of claim 1, further comprising: predicting, by the at least one computing device, excessive daytime sleepiness as a function of a hypoxic domain.

10. The method of claim I, further comprising: predicting, by the at least one computing device, long-term mortality as a function of a prediction a hypoxic domain and an arousal domain.

11. A computer-implemented system for predicting consequences of obstructive sleep apnea, the system comprising: at least one computing device configured by executing instructions for: accessing information associated with overnight polysomnography; extracting features from the information, wherein the extracted features are across three physiological domains; representing each of the features across the three physiological domains as histograms; applying the histograms to two extreme gradient boosting (XGBoost) models; predicting, as a function of the two XGBoost models, short-term excessive daytime sleepiness (EDS) and long-term all-cause mortality; applying the histograms to a gradient-boosting survival model; evaluating, as a function of the gradient-boosting survival model, survival trajectories; andevaluating, using artificial intelligence, feature importance of the information associated with overnight polysomnography for predicting a consequence of obstructive sleep apnea.

12. The system of claim 11, wherein the at least one computing device is farther configured for: combining physiological measures across impacted domains in obstructive sleep apnea in a non-linear method.

13. The system of claim 11, wherein the impacted domains include ventilatory, hypoxic, and arousal.

14. The system of claim 11 , wherein predicting the consequence includes providing inputs represented as histograms to at least one machine learning model, and further wherein the at least one computing device is further configured for: processing output of the at least one machine learning model to reveal patterns.

15. The system of claim 14, wherein the patterns are not evident in a single summary measure.

16. The system of claim 11 , wherein predicting the consequence includes at least one distinct clinical outcome.

17. The system of claim 1 1 , wherein the at least one distinct clinical outcome is based on a function of a multifactorial nature of obstructive sleep apnea.

18. The system of claim 11 , wherein the predicting includes at least one of shortterm outcome and long-term health outcome.

19. The system of claim 11, wherein the at least one computing device is farther configured for: predicting excessive daytime sleepiness as a function of a hypoxic domain.

20. The system of claim 11, wherein the at least one computing device is further configured for: predicting long-term mortality as a function of a pred iction a hypoxic domain and an arousal domain.

21. A method of determining ventilatory burden in a subject comprising: determining the normalized ventilatory amplitude of said subject during a sleep period; determining if the median value of the proportion of breaths during said period is less than about 50% of the normalized amplitude of breaths during said period; and diagnosing said patient with ventilatory burden where said proportion falls below about 50%.

22. A method of determining ventilatory burden in a subject comprising: determining the normalized breath amplitude of a subject over a sleep period; and diagnosing said patient with ventilatory burden wherein more than about 25% of breaths during said period are less than about 50% of said normalized breath amplitude.

23. A method of treating obstructive sleep apnea in a subject in need thereof comprising: determining the normalized ventilatory amplitude of a subject during a sleep period; determining the median value of the normalized ventilatory amplitude during said period; and treating said subject for obstructive sleep apnea wherein said median value is less than about 50% of the normalized amplitude during said period.

24. A method of treating obstructive sleep apnea in a subject in need thereof comprising: determining the normalized breath amplitude of a subject over a sleep period; and treating said patient for obstructive sleep apnea wherein more than about 25% of breaths during said period are less than about 50% of said normalized breath amplitude.