System and method for encoding full-night, multichannel sleep study data via a foundational transformer

A patch-based transformer neural network model encodes PSG data for efficient and accurate sleep stage classification, addressing the limitations of manual annotation in PSG data analysis and enhancing the scalability and consistency of sleep disorder diagnoses.

WO2026035742A1PCT designated stage Publication Date: 2026-02-12MT SINAI SCHOOL OF MEDICINE
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
PCT/US2025/040738
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-08-05
Filing Date
2025-08-05
Publication Date
2026-02-12

AI Technical Summary

Technical Problem

Manual annotation and review of polysomnography (PSG) data for sleep disorder diagnosis is time-consuming, resource-intensive, and prone to inconsistency and inaccuracies, limiting the scalability and accuracy of sleep disorder diagnoses.

Method used

A computer-implemented method using a patch-based, self-supervised transformer neural network model to encode full-night, multichannel PSG data, followed by a supervised, bidirectional gated recurrent unit probing head for sleep stage classification, enabling efficient and accurate sleep stage prediction in a single step.

Benefits of technology

The method provides standardized, automated sleep stage determination and annotation, reducing variability and improving the scalability and consistency of sleep disorder diagnoses, while maintaining high accuracy and efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure US2025040738_12022026_PF_FP_ABST
    Figure US2025040738_12022026_PF_FP_ABST
Patent Text Reader

Abstract

A computer-implemented system and method for encoding full-night, multichannel sleep study data via a foundational transformer sleep stage classification, including: generating, by at least one computing device employing a patch-based, self-supervised transformer model, encodings, wherein the encodings are generated using full-night, multi-channel polysomnography data; inputting, by the at least one computing device, the encodings into a supervised, bidirectional, gated recurrent unit probing head; and classifying, by the computing device as a function of the encodings input to the supervised, bidirectional, gated recurrent unit probing head, sleep stages.
Need to check novelty before this filing date? Find Prior Art

Description

SYSTEM AND METHOD FOR ENCODING FULL-NIGHT, MULTICHANNEL SLEEP STUDY DATA VIA A FOUNDATIONAL TRANSFORMERCROSS-REFERENCE TO RELATED APPLICATIONS

[0001] This application claims the benefit of and priority to U.S. Provisional Patent Application Ser. No. 63 / 679,373, filed August 5, 2024, the entire contents of which are incorporated by reference herein.FIELD

[0002] The present disclosure generally relates to processing polysomnography data and, more specifically, to a computer-implemented method and system for processing polysomnography data using machine learning techniques to analyze sleep data, including classifying sleep stages.BACKGROUND

[0003] Sleep deprivation and sleep-related disorders are a public health epidemic affecting about 30% of the U.S. population, yet it is believed that only 5% have been properly diagnosed. Sleep deprivation and sleep-related disorders disrupt daily activities, mental health, and longevity, and can contribute to major health issues, such as cardiovascular disease, cancer, diabetes, and hypertension. Unfortunately, economic and social factors, as well as climate change, worsen sleep conditions for many.

[0004] Polysomnography (PSG) is a known method for diagnosing sleep disorders. Physiological signal data, including electroencephalogram (EEG), electrooculogram (EOG), electrocardiogram (ECG), and electromyogram (EMG) can be collected during a sleep study, in addition to vital signs, such as respiration rate and oxygen saturation, among others. These signals can be manually annotated and reviewed by a clinician where they identify sleep stages, for example, wake, non-rapid eye movement (NREM) stages 1 (N1), 2 (N2), and 3 (N3), and rapid eye movement (REM). The manual annotation and review process is time and resource intensive, which limits scalability of such sleep disorder diagnoses. Unfortunately, manual annotation and review processes can be prone to inconsistency and / or inaccuracies.

[0005] It is in view of these and other concerns that the present disclosure is made.SUMMARY

[0006] According to one or more example implementations of the present disclosure, a computer-implemented method for encoding full-night, multichannel sleep study data via a foundational transformer neural network comprises: generating, by at least one computing device employing a patch-based, self-supervised transformer neural network model, encodings, wherein the encodings are generated using full-night, multichannel polysomnography data; inputting, by at least one computing device, the encodings into a supervised, bidirectional, gated recurrent unit probing head; and classifying, by at least computing device as a function of the encodings input to the supervised, bidirectional, gated recurrent unit probing head, sleep stages.

[0007] In one or more example implementations, the sleep stage classification is implemented using a foundational transformer.

[0008] In one or more example implementations, the foundational transformer predicts sleep stages with a single step.

[0009] In one or more example implementations, the full-night, multi-channel polysomnography data include sleep signal data collected over at least eight hours.

[0010] In one or more example implementations, the sleep signal data are associated with at least one of brain, movement, cardiac, oxygen, and respiratory channels.

[0011] In one or more example implementations, the method further comprises providing, by at least one computing device, the classified sleep stages to at least one clinical application.

[0012] In one or more example implementations, the method further comprises annotating, by at least one computing device, a sleep study.

[0013] In one or more example implementations, the method further comprises preprocessing the full-night, multi-channel polysomnography data, wherein the preprocessing comprises at least one of resampling, instance normalization, and patching.

[0014] In one or more example implementations, the full-night, multi-channel polysomnography data is in patches of a predetermined duration that is less than 30 seconds, or about 4 seconds to about 8 seconds, or preferably 6 seconds.

[0015] In one or more example implementations, the full-night, multi-channel polysomnography data comprises random zero padding within the patches to form sleep signal data having a duration of about 6 hours to about 10 hours, or preferably 8 hours.

[0016] According to one or more example implementations of the present disclosure, a computer-implemented system for encoding full-night, multichannel sleep study data via a foundational transformer neural network comprises: at least one computing device configured by executing processor-readable instructions that configure the at least one computing device for: employing a patch-based, self-supervised transformer neural network model to generate encodings, wherein the encodings are generated using full-night, multi-channel polysomnography data; inputting the encodings into a supervised, bidirectional, gated recurrent unit probing head; and classifying, as a function of the encodings input to the supervised, bidirectional, gated recurrent unit probing head, sleep stages.

[0017] In one or more example implementations, the sleep stage classification is implemented using a foundational transformer.

[0018] In one or more example implementations, the foundational transformer predicts sleep stages with a single step.

[0019] In one or more example implementations, the full-night, multi-channel polysomnography data include sleep signal data collected over at least eight hours.

[0020] In one or more example implementations, the sleep signal data are associated with at least one of brain, movement, cardiac, oxygen, and respiratory channels.

[0021] In one or more example implementations, the at least one computing device is further configured for providing the classified sleep stages to at least one clinical application.

[0022] In one or more example implementations, the at least one computing device is further configured for annotating a sleep study.

[0023] In one or more example implementations, the processor-readable instructions further configure the at least one computing device for preprocessing the fullnight, multi-channel polysomnography data, wherein the preprocessing comprises at least one of resampling, instance normalization, and patching.

[0024] In one or more example implementations, the full-night, multi-channel polysomnography data is in patches of a predetermined duration that is less than 30 seconds, or about 4 seconds to about 8 seconds, or preferably 6 seconds.

[0025] In one or more example implementations, the full-night, multi-channel polysomnography data comprises random zero padding within the patches to form sleep signal data having a duration of about 6 hours to about 10 hours, or preferably 8 hours.

[0026] These and other aspects, features, and advantages can be appreciated from the accompanying description of certain embodiments of the invention and the accompanying drawing figures and claims.BRIEF DESCRIPTION OF THE DRAWINGS

[0027] Various features, aspects and advantages of the invention can be appreciated from the following detailed description and the accompanying drawing figures, in which:

[0028] FIG. 1 is a schematic diagram illustrating a sleep stage determination system including a foundational transformer in accordance with one or more example implementations of the present disclosure.

[0029] FIG. 2 is a flow diagram showing a sleep stage classification process of the system of FIG. 1 according to one or more example implementations of the present disclosure.

[0030] FIG. 3 is a flow diagram of a data preparation process for training the foundational transformer of FIG. 1 in accordance with one or more example implementations of the present disclosure.

[0031] FIG. 4A is a schematic diagram corresponding to FIG. 1 for illustrating the operations of training the system of FIG. 1 to process PSG data in accordance with one or more example implementations of the present disclosure.

[0032] FIG. 4B is a schematic diagram corresponding to FIG. 1 for illustrating the operations of finetuning the system of FIG. 1 to determine sleep stages in accordance with one or more example implementations of the present disclosure.

[0033] FIG. 5 is a schematic diagram illustrating a data processing apparatus and associated network and processing apparatuses for implementing the systems and methods of the present disclosure.

[0034] FIG. 6 is a schematic diagram illustrating the data processing apparatus ofFIG. 5 according to one or more example implementations of the present disclosure.

[0035] FIG. 7 shows plots of recreations of 10-second samples of signal data generated using a foundational transformer according to an example implementation of the present disclosure.

[0036] FIGS. 8A, 8B, and 8C show respective row normalized confusion matrices for sleep stage classifications on held-out test datasets by a foundational transformer model according to respective example implementations of the present disclosure.

[0037] FIGS. 8D, 8E, and 8F show respective row normalized confusion matrices for sleep stage classifications on independent test datasets by a foundational transformer model according to respective example implementations of the present disclosure.

[0038] FIG. 9 shows a graph plot comparing sleep staging results of respective foundational transformer models according to example implementations of the present disclosure.DETAILED DESCRIPTION

[0039] By way of overview and introduction, the present disclosure provides tools and techniques for monitoring and accurate sleep stage detection, event detection, and long-term monitoring, including to improve diagnosis and sleep habits. Such tools and techniques include customized artificial intelligence (Al) tools, which can process comprehensive PSG data and provide standardized, automated sleep stage determination and annotation systems and methods. The present disclosure can employ Al-based algorithms, including via transformer architecture, for analyzing time series data and in view of the need for more consistent and efficient sleep analyses, the present disclosure provides a customized, patch-based, self-supervised transformer model for standardizing PSG data encodings that reduces variability in sleep scoring and annotation of sleep signals.

[0040] The following one or more example implementations are described based on PSG data and sleep stage determination features of which may be incorporated into other types of time series physiological measurement data for health analysis classification tasks without departing from the spirit and the scope of the disclosure.

[0041] FIG. 1 is a schematic diagram illustrating a sleep stage determination system 100 in accordance with one or more example implementations of the presentdisclosure. In embodiments, elements of system 100 can be executed, and / or embodied, at one or more of data processing apparatus 502 and computing apparatuses 504 shown in FIG. 5.

[0042] As illustrated in FIG. 1, system 100 comprises a foundational transformer (hereinafter referred to as PFTSleep) 105, which, as will be described in further detail below, is trained to embody a patch-based, self-supervised transformer model. In one or more example implementations, data processing apparatus 502 (FIG. 5) executes PFTSleep 105 to comprise an embedding layer 110 adapted to convert input PSG signal data 112 — from, for example, electroencephalogram (EEG), electrooculography (EOG), electromyography (EMG), electrocardiogram (ECG), peripheral capillary oxygen saturation (SpO2), thoracic respiratory rate (Thor RR), and / or abdominal respiratory rate (Abdo RR) measurement signals, to name a few — into dense vector representations suitable for neural network processing. Each of the EEG, EOG, EMG, ECG, SpO2, Thor RR, and Abdo RR can be comprised in a respective channel of PSG signal data 112. At a next stage, PFTSleep 105 adds positional encoding 115 to the vector representations of embedding layer 110 via an adder 120 to provide sequence awareness for the subsequent stages of PFTSleep 105.

[0043] Next, PFTSleep 105 passes the sequenced representations through a multihead self-attention mechanism (Multi-Head Attention) 125 and a residual connection and layer normalization mechanism (“Add & Norm") 130. In embodiments, Multi-Head Attention 125 enables PFTSleep 105 to attend to different parts of the sequence simultaneously, capturing complex dependencies across time and channels, and Add & Norm 130 applies residual connections followed by layer normalization to stabilize and speed up training.

[0044] At a next stage, PFTSleep 105 comprises a feed forward layer 135 and another Add & Norm 140. In embodiments, feed forward layer 135 can be a fully connected layer applied independently to each position, allowing for non-linear transformations of the attended features, and Add & Norm 140 provides another residual connection and normalization step to maintain model stability for PFT Sleep 105.

[0045] In one or more example implementations, PFTSleep 105 projects the encoding output into a time embedding space 145, where each time point is represented as a vector. Accordingly, time embedding space 145 captures the temporal dynamics of sleep stages across a night. At a next stage, a recurrent neural network (RNN) or, more specifically, a gated recurrent unit (GRU) 150 processes the time-embedded vectors tomake stage-wise predictions 152 over an 8-hour sleep period. In one or more example implementations, the output 152 of the GRU 150 is a sequence of predicted sleep stages: a. W (Wake) b. N1, N2, N3 (Non-REM stages) c. REM (Rapid Eye Movement sleep)

[0046] FIG. 2 is a flow diagram showing a sleep stage classification process 200 of system 100 according to one or more example implementations of the present disclosure. In embodiments, process 200 can be executed at one or more of data processing apparatus 502 and computing apparatuses 504 shown in FIG. 5.

[0047] As shown in FIG. 2, process 200 initiates with step s201, where PFTSleep 105 generates encodings using PSG data 112. In one or more example implementations, PSG data 112 comprises full-night (e.g., 8 hours), multi-channel (e.g., 3 or 7 channels) PSG data.

[0048] Next, at steps s202, PFTSleep 105 projects the generated encodings to a time embedding space 145 and inputs the results to GRU 150, which embodies a supervised, bidirectional, gated recurrent unit probing head.

[0049] Then, process 200 concludes with step s203, where GRU 150 classifies sleep stages for the PSG data 1 12 as a function of the encodings input from PFTSleep 105 to GRU 150. In embodiments, the sleep stage classifications can be output to one or more display devices and / or stored in one or more data storage devices, for example, at one or more of data processing apparatus 502 and computing apparatuses 504 shown in FIG. 5.

[0050] In one or more example implementations, PFTSleep 105 is trained using customized PSG data for improved sleep stage prediction. FIG. 3 is a flow diagram of a data preparation process 300 for training PFTSleep 105 in accordance with one or more example implementations of the present disclosure. In embodiments, process 300 can be executed at one or more of data processing apparatus 502 and computing apparatuses 504 shown in FIG. 5.

[0051] As illustrated in FIG. 3, process 300 initiates with step s301, where data processing apparatus 502 extracts signal channels from one or more sleep study signal data files, for example, in European Data Format (EDF) and stores the extracted channels in an arrayed data loading format, for example, the Zarr format. In one or more example implementations, the sleep study data comprises seven (7) channels, including respectivechannels for EEG, EOG, EMG, ECG, SpO2, Thor RR, and Abdo RR. In certain embodiments, the sleep study data can include x channels (x ≥ 1 , or 2 ≤ x ≤ 10, or x = 3).

[0052] Next, at step s302, data processing apparatus 502 obtains labeled data with sleep stage identifications associated with the stored data of step s301. According to one or more example implementations, the labeled data is obtained from plural datasets of differing sleep studies. In certain embodiments, the labeled data can include hypnograms indicating REM and NREM (wake, N1, N2, N3, N4) sleep labeled, for example, every 30 seconds of sleep data. Accordingly, the average probabilities can be every 5 patches to get a prediction for each 30 seconds across the 8 hours of data. In some embodiments, sleep stage N4 can be truncated into stage N3 in the labeled sleep data.

[0053] Data processing apparatus 502 then, at step s303, resizes a unit for all sleep study signal data to a standard (“full night") sleep duration. In one or more example implementations, the standard (“full night”) sleep duration is set to 8 hours. In certain embodiments, the resizing includes trimming sleep study signal data to a standard sleep duration unit and / or zero padding sleep study signal data to extend to a standard sleep duration unit. In embodiments, the standard sleep duration (for a “full night” duration) can be set to another duration, for example, from about 6 hours to about 10 hours. In one or more example implementations, data processing apparatus 502 executes the zero padding by random zeroing out and augmentation of values within patches (or masked autoregression).

[0054] Next, at step s304, data processing apparatus 502 resamples all channels of the signal data to a standard sampling frequency. In one or more example implementations, the signal data is resampled to 125 Hz.

[0055] Data processing apparatus 502 then, at step s305, filters all signals to a predetermined window size. In one or more example implementations, the predetermined window size is 3. In some embodiments, bandpass filters (second order) can be applied to ECG (0.5—40 Hz), EEG (0.3-30 Hz), EMG (0.3-30 Hz), and EOG (0.3-30 Hz) signals, and lowpass filters can be applied to SpO2 (0.4 Hz) and respiration (0.5 Hz) signals. Accordingly, after resampling and padding at steps s3O3 and s304, the length of an input sample can be 7 x 3.6 million (7 channels, 125 Hz, 8-hour long signals) samples.

[0056] Next, at step s306, data processing apparatus 502 normalizes the resampled and resized sleep study data units and patches normalized data into patches of predetermined duration. In one or more example implementations, the predeterminedduration of the patches is six (6) seconds and the normalization is executed using a learned reversible instance normalization layer — for example, sample-channel-wise unit variance scaling with a learnable weight parameter — thus, effectively creating a 3D tensor with seven (7) (or x) channels. In embodiments, the predetermined duration of the patches can be less than thirty (30) seconds, or from about four (4) seconds to about eight (8) seconds.

[0057] Data processing apparatus 502 then, at step s307, passes the patches of each of the x (or 7) channels through an individual linear layer (input size: 750 samples per patch) to embed the 6-second patches into a vector — for example, of size 512 — for input into the transformer architecture to train and form PFTSleep 105. Thus, in certain embodiments, the patch input into the transformer for training PFTSleep 105 can be a 3D tensor of shape 7 x 4800 x 512 (7 channels, 4800 6-second patches in 8 hours, 512 patch feature representations).

[0058] Process 300 concludes with step s308, where data processing apparatus 502 inputs the vectors of step s307 to the transformer architecture of PFTSleep 105 for training PFTSleep 105.

[0059] In some embodiments, process 300 can incorporate elements of a PatchTST implementation with customized parameters. For example, PatchTST can usually implement 8 or 16 patch lengths. For the present disclosure, in one or more example implementations, the channel data is patched at step s206 into 6-second segments at 125 Hz, which has a length of 750 samples per patch. Advantageously, six- second patches can capture more specific sleep events throughout a sleep study duration, hi certain embodiments, the patches can be set to varying lengths, for example, less than 30 seconds, from about 4 seconds to about 8 seconds, or the like.

[0060] Additionally, based on the processing at steps s303-s306, zeroed out and augmented values of data for conforming the sleep study data to a uniform (or standard “full night”) sleep study duration, for example, 8 hours, are within patches, instead of zeroing out full patches of time series data during self-supervision.

[0061] In one or more example implementations, for providing PFTSleep 105 with an equal length input (8 hours) of training data across plural sleep studies given that not all sleep studies are at least 8 hours long, data processing apparatus 502 applies a padding mask at one or more of steps s302 and s303 to the labeled training data to indicate to PFTSleep 105 which parts of the input are actual sleep data and which parts are padded values. For example, in certain embodiments, sleep studies less than 8 hours can bepadded to 8 hours with the padded values adapted to be ignored by PFTSleep 105. Advantageously, during self-supervised training, PFTSleep 105 learns representations of 6-second patches of channel data and their relationship with each other over 8 hours (4800 patches) of sleep data. After self-supervised training, the representations generated at PFTSleep 105 can have a “foundational” understanding of the sleep study data, and these new features can be used to predict sleep stages. The trained PFTSleep 105, thus, embodies a patch-based, self-supervised transformer neural network model. In some embodiments, PFTSleep 105 can embody a larger encoder with additional layers and / or additional attention heads (125) to generate more complex representations for capturing more subtle features of the PSG data (112).

[0062] FIGS. 4 A and 4B are schematic diagrams corresponding to FIG. 1 for illustrating the operations of training system 100 to process PSG data and to determine sleep stages (or “Sleep Staging”), respectively. In embodiments, the operations shown in FIGS. 4A and 4B, including elements of system 100, can be executed, and / or embodied, at one or more of data processing apparatus 502 and computing apparatuses 504 shown in FIG. 5.

[0063] FIG. 4A illustrates self-supervised training of PFTSleep 105 according to one or more example implementations of the present disclosure. As shown in FIG. 4A, for training system 100, or PFTSleep transformer 105, data processing apparatus 502 inputs prepared data 402 to transformer 105 in a forward pass. Prepared data 402 corresponds to the data input to transformer 105 at step s308 of process 300. As described before, in one or more example implementations, prepared data 402 (or “Augmented Time Patch”) comprises PSG data that is augmented to a standard “full night” (or 8 hour) duration and patched into a predetermined duration, for example, 6 seconds.

[0064] In a forward pass, transformer 105 projects the encodings resulting from input of data 402 to time embedding space 145. The projected representations in time embedding space 145 are output to pretrain GRU, or “Pretrain Head,” 150a in a forward pass. As illustrated in FIG. 4A, GRU 150a generates a recreated time patch 405, based upon which a mean squared error (MSE) loss function 410 against a corresponding nonaugmented original time patch 407 is propagated in a backward pass (or “Backward Pass Pretaining”) through GRU 150a, time embedding space 145, and transformer 105. Accordingly, the backward pass pretraining updates model weights, for example, via gradient descent. In embodiments, the backward pass pretraining can include additional and / or alternative features for updating and training system 100.

[0065] FIG. 4B illustrates operations of finetuning GRU 150b after the training of system 100 shown in FIG. 4A. For sleep stage classification after transformer training, GRU 150 (or 150b) is trained to predict a sleep stage using the representations for each 6-second patch. In operation, RNN, or GRU, 150 recurrently calculates a final feature vector for each 6-second feature representation across all signals. In one or more example implementations, GRU 150 comprises a linear layer (not shown) through which the feature vector of time embedding space 145 is passed to output a class probability for each sleep stage. During training, RNN, or GRU, 150 learns the relationship among sleep stages of a current, a previous, and a subsequent patch (e.g., 6-second patches across x channels). In embodiments, various mechanisms can be implemented within RNN, or GRU, 150 to learn how much information should be passed from patch to patch during training. In one or more example implementations, GRU 150 includes two “gates” to determine how much information from a previous patch should be included in processing the feature vector (145) of a current patch and how much information from the current patch should be passed on to process the feature vector (145) of a next patch. As illustrated in FIG. 4B, for finetuning a trained GRU (or “Probing Head”) 150b to execute sleep stage determinations based on PSG data, a backward pass linear probing based on the sleep stage output 415 at GRU 150b is passed back through GRU 150b. In embodiments, the backward pass linear probing can include gradients from a classification loss (e.g., a focal loss function) against, for example, labeled data obtained at step s302 in process 300 (FIG. 3).

[0066] Advantageously, PFTSleep 105 is a foundational transformer that processes PSG data to provide encodings that are suitable for varied applications associated with sleep analysis, not only sleep staging. Furthermore, system 100 can be customized to specific sleep staging applications, which may involve different data characteristics, by finetuning GRU 150 without the need to retrain PFTSleep 105 or the overall system 100. In some embodiments, one or more Al tools with simpler architecture can be used in place of GRU 150 for faster finetuning to alternative datasets.

[0067] Referring now to FIG. 5, a block diagram is shown illustrating an example implementation of the present disclosure and that represents an association of a plurality of devices and the flow 508 of information associated with the devices. In the example shown in FIG. 5, various computing devices 502 and 504 are shown, each capable of executing desktop and / or mobile computing device web browser application(s) including MICROSOFT EDGE, INTERNET EXPLORER, CHROME, FIREFOX, and other (e.g.,SAFARI, OPERA). In addition to standard web browser application functionality, user information can be gathered via Push Notifications, and information can be retrieved from a computing device using a “REST” interface. Various mobile devices running different operating systems are shown, including IOS, ANDROID and other (e.g., PALM, WINDOWS or other mobile device) operating system.

[0068] In the example shown in FIG. 5, one or more data processing apparatuses 502 are operatively coupled to one or more user computing device(s) 504. Devices 502 / 504 can be respectively operated by one or more users skilled in the use of the proposed workflow, including, but not limited to, healthcare providers and associated staff, medical specialists, and / or biomechanical specialists. Healthcare providers can include, for example, physicians, physician assistants, nurses, therapists and / or other providers of healthcare services. Biomechanical specialists can include, for example, engineers specialized in biomechanics. Data processing apparatus 502 and / or user computing device 504 can be operable to access and / or store various information on database(s) including, for example, historic medical and procedure information patients, physicians, devices, or the like.

[0069] Continuing with reference to FIG. 5, network 506 is illustrated, which can be configured as a local area network (LAN), wide area network (WAN), Peer-to-Peer network (“P2P”), Multi-Peer network, the Internet, one or more telephony networks or a combination thereof, that is operable to connect data processing apparatus 502 and / or devices. Though many of the examples and implementations shown and described herein relate to product and / or service recommendations, many other forms of content can be provided and / or delivered by system 500.

[0070] FIG. 6 is a block diagram that illustrates functional elements of one or more of data processing apparatus 502 or computing device 504 and preferably include one or more central processing units (CPU) 602 used to execute software code in order to control operations, including of data processing apparatus 502, read only memory (ROM) 604, random access memory (RAM) 606, one or more network interfaces 608 to transmit and receive data to and from other computing devices across a communication network, storage devices 610 such as a hard disk drive, solid state drive, universal serial bus (USB) drive, floppy disk drive, tape drive, CD-ROM or DVD drive for storing program code, databases and application code, one or more input devices 612 such as a keyboard, mouse, track ball and the like, and a display 614.

[0071] The various components of devices 502 and / or 504 need not be physically contained within the same chassis or even located in a single location. For example, storage device 610 can be located at a site which is remote from the remaining elements of computing devices 502 and / or 504 and can even be connected to CPU 602 across communication network 506 via network interface 608.

[0072] The functional elements shown in FIG. 6 (designated by reference numbers 602-614) are preferably of the same categories of functional elements preferably present in computing device 502 and / or 504. However, not all elements need be present, for example, storage devices in the case of mobile computing devices (e.g., smartphones), and the capacities of the various elements are arranged to accommodate expected user demand. For example, CPU 602 in computing device 504 can be of a smaller capacity than CPU 602 as present in data processing apparatus 502. Similarly, it is likely that data processing apparatus 502 will include storage devices 610 of a much higher capacity than storage devices 610 present in computing device 504. Of course, one of ordinary skill in the art will understand that the capacities of the functional elements can be adjusted as needed. For example, one or more graphics processing units (GPU) can be utilized for processing and providing functionality shown and described herein. In addition, or in the alternative, a cluster of computing devices can work to provide functionality shown and described herein.

[0073] The nature of the present disclosure is such that one skilled in the art of writing computer executed code (software) can implement the described functions using one or more or a combination of a popular computer programming language including but not limited to C++, JAVA, ACTIVEX, HTML, XML, ASP, SOAP, IOS,OBJECTIVE C, ANDROID, TORR, PYTHON, MATLAB, and various web application development environments.

[0074] As used herein, references to displaying data on computing device 504 refer to the process of communicating data to the computing device 504 across communication network 506 and processing the data such that the data can be viewed on the user computing device 504 display 614 using a web browser, custom application or the like. The display screens on computing devices 502 / 504 present areas within system 500 such that a user can proceed from area to area within the system 500 by selecting a desired link. Therefore, each user’s experience with system 500 will be based on the order with which (s)he progresses through the display screens. In other words, because the system is not completely hierarchical in its arrangement of display screens, users canproceed from area to area without the need to “backtrack” through a series of display screens. For that reason and unless stated otherwise, the following discussion is not intended to represent any sequential operation steps, but rather the discussion of the components of system 500.

[0075] Although the present disclosure is described by way of example herein in terms of a web-based system using web browsers, custom applications and a web site server (data processing apparatus 502), and with mobile computing devices, system 500 is not limited to that particular configuration. It is contemplated that system 500 can be arranged such that computing device 504 can communicate with, and display data received from, data processing apparatus 502 using any known communication and display method, for example, using a non-Intemet browser Windows viewer coupled with a local area network protocol such as the Internetwork Packet Exchange (IPX). It is further contemplated that any suitable operating system can be used on computing device 504, for example, WINDOWS, MAC OS, OSX, LINUX, IOS, ANDROID and any suitable PDA or other computer operating system.

[0076] The present disclosure is further described below in connection with one or more example implementations.

[0077] Foundational transformers that can be included and configured according to PFTSleep 105 were trained and validated using deidentified, retrospective PSG data collected from multicenter cohort studies and made available through the National Sleep Research Resource (NSRR), including the Sleep Heart Health Study (SHHS), the Wisconsin Sleep Cohort (WSC), the Osteoporotic Fractures in Men (MrOS) Study, the Multi-Ethnic Study of Atherosclerosis (MESA) and the Apnea Positive Pressure Longterm Efficacy Study (APPLES). In total, the training and testing involved over one million hours of sleep study signal data from about 14,000 sleep studies acquired from the SHHS visit 1 PSG dataset collected from 1995 to 1998 (n = 5793), SHHS visit 2 PSG dataset collected from 2001 to 2003 (n = 2651), WSC PSG dataset collected from 2000 to 2015 (n = 2544), and MrOS Visit 1 PSG dataset collected in 2005 (n = 2900). TheMESA PSG dataset collected from 2010 to 2012 (n = 2055), the MrOS Visit 2 PSG dataset collected from 2009 to 2012 (n = 1025), and the APPLES PSG dataset collected from 2003 to 2008 (n = 1089) were used as independent test sets. SHHS, MESA, and MrOS were type II, unattended at-home sleep studies. WSC and APPLES were in-lab sleep studies.

[0078] Data extraction and preprocessing

[0079] The sleep study data from the published studies were extracted and preprocessed according to the present disclosure.

[0080] Signal channels were extracted from EDF files and stored in the Zarr format for more performant data loading. Specifically, a single EEG channel (C4-M1 or C3-M2), left EOG (E1 -M2), chin EMG, augmented lead II ECG, SpO2, thorax respiration, and abdomen respiration were extracted from the respective sleep database’ s EDF files. For each study, sleep stages were identified using Rechtshaffen and Kales criteria by a randomly assigned single scorer in a central reading center. Hypnograms indicating REM and NREM (wake, N1, N2, N3, N4) sleep were labeled for each 30 seconds of sleep. Thirty-second epochs that were labeled as movement or unknown were ignored during training. Sleep stage N4 was truncated into stage N3. Sleep studies were trimmed to 8 hours or extended to 8 hours with zero padding. All channels were resampled to 125 Hz. Wake stages were trimmed from the beginning or end of a sleep study to the second-largest sleep stage count if the wake stage count was larger than the second-largest sleep stage count. Wake stages removed were ignored during training. Median filters were applied to all signals with a window size of 3. Bandpass filters (second order) were applied to ECG (0.5-40 Hz), EEG (0.3-30 Hz), EMG (0.3-30 Hz), and EOG (0.3-30 Hz) signals, and lowpass filters were applied to SpO2 (0.4 Hz) and respiration (0.5 Hz) signals. After resampling and padding, the length of an input sample into our model was 7 x 3.6 million (7 channels, 125 Hz 8-hour long signals).

[0081] Input data preparation

[0082] Each 8-hour sleep study after resampling was normalized via a learned reversible instance normalization layer (sample-channel- wise unit variance scaling with a learnable weight parameter) and patched into 6-second patches, effectively creating a 3D tensor with 7 channels. Each channel’s patches were then passed through an individual linear layer (input size: 750) to embed the 6-second patches into a vector of size 512 for input into the transformer architecture. The patch input into the transformer was a 3D tensor of shape 7 x 4800 x 512 (7 channels, 4800 6-second patches in 8 hours, 512 patch feature representations).

[0083] Foundational Transformer Model Training

[0084] The initial learning rate was set to le-5 and a one cycle scheduler was used with a max learning rate of 0.01. The foundational transformer encoder (105) was trained with SHHS visits 1 and 2 with 80% training and 20% validation (without patient overlap) and a 10% augmentation mask, where 10% of values across all channels were augmentedwith random noise or set to zero. The encoder (105) comprised 2 layers with 2 attention heads each, a dimension of 512, and a feedforward size of 2048. The model was trained for 100 epochs with a batch size of 4 and gradient accumulation of 4 on two NVIDIA Al 0080GB Tensor Core GPUs. The training objective was to minimize the mean square error (MSE) between the original, non-augmented time series (407) and the recreated time series (405) (from the augmented data). An overview of the training process is shown in FIG. 4A. The signal data was recreated from the embeddings through a single linear layer to compute the MSE loss. The model with the lowest MSE loss on the validation set was saved over the 100 epochs of training and used for downstream tasks.

[0085] Gated Recurrent Unit (GRU) Probing Head Training

[0086] The GRU probing model (150) was trained using the transformer encoder (105) with the lowest mean squared error validation loss. A bidirectional GRU with a hidden size of 512 was trained to predict a sleep stage for every 6 second patch. Average pooling calculated a prediction for each 30 second sleep epoch (across 5 patches of data). The model was trained for 50 epochs with an initial learning rate of le-5 and a focal loss function, which identifies and focuses on harder to classify examples, compared to cross entropy. The focal gamma parameter was set to 2.0 and the alpha parameter was not set.

[0087] Three (3) foundational transformers configured according to PFTSleep 105 were trained and four (4) GRU classifier heads configured according to GRU 150 were trained using the three (3) foundational transformers to examine the effects of including the MESA dataset at different stages in training.

[0088] Example 1

[0089] For the first of five (5) example implementations of the present disclosure, a foundational transformer configured according to PFTSleep 105 was trained on 3 channels (EEG, left EOG, EMG) of sleep data only (non-MESA), which reached a minimum validation MSE loss after 67 epochs. A 3-channel GRU classifier configured according to GRU 150 (non-MESA) was trained with this foundational transformer.

[0090] Example 2

[0091] For the second of the five (5) example implementations, a foundational transformer configured according to PFTSleep 105 was trained with 7-channels (EEG, EOG, EMG, ECG, SpO2, Thor RR, and Abdo RR) of sleep data (non-MESA). A GRU classifier configured according to GRU 150 (non-MESA) was trained with this foundation transformer.

[0092] Example 3

[0093] For the third of the five (5) example implementations, an additional foundational transformer configured according to PFTSleep 105 was trained with the MESA dataset, which reached a minimum validation MSE loss after 66 epochs. A MESA-trained GRU classifier configured according to GRU 150 was trained with this foundational transformer.

[0094] Example 4

[0095] For the fourth of the five (5) example implementations, a non-MESA- trained transformer (configured according to PFTSleep 105) was paired with a MESA- trained GRU classifier (configured according to GRU 150).

[0096] Example 5

[0097] For the fifth of the five (5) example implementations, a MESA-trained transformer (configured according to PFTSleep 105) was paired with a non-MESA trained GRU classifier (configured according to GRU 150).

[0098] The hyperparameters of the GRU classifiers were identical, and they reached their minimum validation loss at epochs 12, 12, 13, and 10, respectively.

[0099] Statistical analysis and evaluation

[0100] The following includes a description of the statistical analysis and evaluation in connection with the five example implementations.

[0101] Dataset demographics were compared using independent t-tests and chi- squared tests depending on whether the variable was numeric (age) or categorical (gender). Sleep stage ratios per sample were compared using the Mann-Whitney U test.

[0102] Model performance was evaluated with micro accuracy, macro sensitivity, specificity, F1 scores, and Cohen’s Kappa for comparison to other models on the held- out test sets and independent datasets. A macro metric calculates the metric for each class separately and then takes the average across the metric calculated for each class. Meanwhile, a micro-metric aggregates all results together. Sensitivity, also known as recall or the true positive rate, indicates the ratio of true positives identified by the model overall positives in the dataset. Specificity, also known as true negative rate, indicates the ratio of true negatives identified to overall negatives in the dataset. F1 score is the harmonic mean between precision and sensitivity, where precision is the ratio of true positives identified over the total positives predicted (true positives and false positives). Cohen’ s Kappa is a measure of agreement between two raters that considers agreementby chance. The primary metrics used for comparison are Cohen’s Kappa and F1 scores. When comparing evaluation metrics of a machine learning model to another model, comparisons are not always direct, as training and validation data differ; however, we do our best to match validation sets for optimal comparison.

[0103] Further, sleep stage classification is inherently an imbalanced classification problem, and the metrics reported should account for this. Therefore, a macro area was included under the precision-recall curve (AUPRC) to examine precision and recall. The AUPRC metric summarizes the tradeoff between precision and recall for each class in a one-class versus rest approach by examining precision and recall at multiple decision thresholds (a cutoff probabi li ty that decides when the probability output from the model should be classified as positive or negative). A high precision value close to 1 indicates few false positives and a high recall value close to 1 indicates few false negatives. An AUPRC value close to 1 indicates that the model is good at identifying true positives while reducing false positives.

[0104] The macro area under the receiver operating characteristics curves (AUROC) and confusion matrices are also reported. AUROC is another common score in machine learning that examines the relationship between recall and specificity; however, it is less effective when there are few positive cases in the dataset.

[0105] Finally, model explainability was performed using the Integrated Gradients method with Captum, a model interpretability package for PyTorch (https: / / captum.ai / ). The Integrated Gradients method was chosen over perturbationbased methods because of its computational efficiency. It attributes an importance value to each time point in the input using the integral of the gradients. Details of the Integrated Gradients implementation are described in Supplementary Methods.

[0106] Example Results

[0107] Demographics of the studies, training, validation, and held-out test sets are presented in Table 1 below.

[0108] Table 1 (Demographics and Sleep Stage Proportions of SHHS Visits 1 and2, MESA, WSC, APPLES, and MrOS Visits 1 and 2, and the Training, Validation, and Held-Out Sets)

[0109] Age, gender, and sleep stages were balanced across the training, validation, and held-out test sets. There were no statistically discernible differences between the training dataset and the held-out testing dataset in age, gender, or sleep stages. Between the training and each individual independent testing dataset, there were statistically discernible differences for age, gender, and each sleep stage, except for APPLES REM sleep.

[0110] Hyperparameter tuning was performed on a subset of the training data, SHHS, to find optimal parameters. The transformer model reached a minimum validation set mean squared error (MSE) loss on epoch 67 (validation MSE loss: 0.0343). Training loss was stable until epoch 72, where a large jump occurred potentially indicating the model escaped from the minimum it was converging to. Validation loss decreased initially but was more variable. There were no signs of overfitting.

[0111] FIG. 7 shows plots of recreations of 10-second samples of signal data, which correspond to the recreated time patch 405 illustrated in FIG. 4A. In FIG. 7, recreated signals for ECG, EOG(L), EMG, EEG, SaO2, THOR RES, and ABDO RES (e.g., 405 in FIG. 4A) are shown in black and the respective original signals (e.g., 407 in FIG. 4A) are shown in indicated gray shades. As illustrated in FIG. 7, the recreated signals are closely aligned with the true signal data. It appears that signals that are periodic and lower frequency (respiration) are more easily recreated than higher frequency signals(EEG, EOG, EMG, ECG) or signals that lack periodicity (SpO2). For sleep stage classification, a GRU classifier configured according to GRU 150 reached a validation loss minimum after only 7 epochs and more effectively utilized the representations generated by the foundational transformer configured according to PFTSleep 105 than multi-layer perceptrons, transformer decoders, and convolutional heads.

[0112] Sleep stage classification results

[0113] Overall, the macro AUPRC scores for wake, N1, N2, N3, and REM sleep on the held-out test set, comprising data from SHHS, WSC, and MrOS visit 1 was 0.84, indicating PFTSleep (105) is good at identifying sleep stages correctly (few false positives) and recalling most of the correct sleep stage (few false negatives). For the independent test sets, MESA, MrOS visit 2, and APPLES, macro AUPRC scores were 0.68, 0.76, and 0.68.

[0114] PFTSleep (105) has strong performance in the wake, N2, and REM in both held-out and independent test sets with AUPRC scores of 0.98, 0.94, and 0.96 (aggregate held-out test set), 0.92, 0.83, and 0.94 (APPLES), 0.98, 0.88, and 0.92 (MrOS visit 2), and 0.92, 0.80, and 0.84 (MESA), respectively.

[0115] The Examples were compared to other sleep staging-specific models that were trained with similar datasets and PFTSleep (105) achieved competitive results. For SHHS, L-SeqSleepNet, SleepTransformer, and XSleepNet reached Cohen’s Kappa values of 0.83, meaning these models had a near “perfect agreement” with the human- scored sleep study. On an identical SHHS dataset split, PFTSleep (105) maintained strong performance as well (0.80 Cohen’s Kappa). Notably, performance is high using full night, sleep stage representations that were generated independently of the sleep stage classification task. U-Sleep, another multi-dataset trained sleep stage-specific classification model, reported a macro 80.0 F1 score on 100 held-out samples of SHHS, meaning that they had an optimal balance of high precision (fewer false positives) and high recall (fewer false negatives), compared to PFTSleep’ s (105) of 76.5 on 2533 held- out samples. For WSC, PFTSleep (105 achieved 0.78 Cohen’s Kappa indicating substantial agreement with human scorers, compared to Olesen et al’s two models with 0.72 and 0.75 Cohen’s Kappa values. For MrOS, visit 1 results (held-out test set) for PFTSleep (105) were compared to other models that used MrOS in training or as independent test sets (Zhang et al.). PFTSleep (105) achieved a Cohen’s Kappa of 0.82 with almost perfect agreement with human scorers, compared to Olesen’ s of 0.76 and Zhang’s of 0.70. Compared to USleep, PFTSleep (105) had comparable performance at72.6 macro F1 versus 77.0 macro F1 with small improvements in wake, N3, and REM sleep and a decrease in performance in N1. For MrOS Visit 2, an independent test set, PFTSleep 105 achieved a similar performance (66.7 to 74.1 macro F1) to RobustSleepNet, which used MrOS visit 2 in training. For APPLES, there were no comparable studies. Additional metrics and model comparisons can be found in Tables 2 and 3 below.Table 2 (Sleep Stage Evaluation Metrics in Comparison to State of the Art Models by Dataset)Table 3

[0116] The MESA dataset was used in combinations of including it in the training dataset used for the foundational transformer (105) and GRU classifier (150). WhenMESA was completely independent (Example 2), PFTSleep (105) achieved moderate agreement with human scorers at 0.60 Cohen’s Kappa on the entire MESA dataset of 2055 sleep studies compared to FullSleepNet’s 30% held-outtest set (-615 sleep studies) Cohen’s Kappa value of 0.76. USleep achieved a macro F1 score on a 100-held-out sample of 79.0, compared to PFTSleep’ s 56.7. PFTSleep still maintained high F1 scores for wake, N2, and REM sleep.

[0117] When PFTSleep was trained including the MESA dataset in the foundational transformer (105) only (Example 5), performance did not change substantially. This indicated that similar feature representations were being created by the foundational transformer (105), independent of whether the MESA dataset was included in the training process. When PFTSleep was trained with MESA data in the GRU classifier (150) only (Example 4), there was a strong performance increase. PFTSleep achieved substantial agreement with human scorers with a 0.76 Cohen’s Kappa value, identical to FullSleepNet and a macro F1 score of 73.1 compared to USleep’s score of 79.0. Performance increases were seen across all sleep stages, but most were in N1, with an AUPRC score increase of 0.28. Finally, when MESA was included in both the training of the foundational transformer (105) and GRU classifier (150) (Example 3), performanceremained the same as the GRU classifier-only model (Example 4). Additional metrics and comparisons can be found in Tables 4 and 5 below.Table 4Table 5

[0118] Overall, PFTSleep performed well in distinguishing wake, NREM, andREM sleep in independent test sets.

[0119] The majority of misclassifications, as seen in the confusion matrices inFIGS. 8A, 8B, 8C, 8D, 8E, and 8F, occurred among N1, N2, and N3. The data from the confusion matrices of FIGS. 8A-8F is reproduced in Tables 6-11 below.

[0120] FIG. 8A shows a row normalized confusion matrix for sleep stage classification (among Wake, N1, N2, N3, and REM stages) on a held-out test set forSHHS, with the rows for true labels and columns for predicted labels, or model outputs.

[0121] The data for FIG. 8 A is reproduced in Table 6 below.Table 6 (SHHS)

[0122] FIG. 8B shows a row normalized confusion matrix for sleep stage classification (among Wake, N1, N2, N3, and REM stages) on a held-out test set for WSC, with the rows for true labels and columns for predicted labels, or model outputs.

[0123] The data for FIG. 8B is reproduced in Table 7 below.Table 7 (WSC)

[0124] FIG. 8C shows a row normalized confusion matrix for sleep stage classification (among Wake, N1, N2, N3, and REM stages) on a held-out test set for MrOS visit 1, with the rows for true labels and columns for predicted labels, or model outputs.

[0125] The data for FIG. 8C is reproduced in Table 8 below.Table 8 (MrOS visit 1)

[0126] FIG. 8D shows a row normalized confusion matrix for sleep stage classification (among Wake, N1, N2, N3, and REM stages) on an independent test set forMESA, with the rows for true labels and columns for predicted labels, or model outputs.

[0127] The data for FIG. 8D is reproduced in Table 9 below.Table 9 (MESA)

[0128] FIG. 8E shows a row normalized confusion matrix for sleep stage classification (among Wake, N1, N2, N3, and REM stages) on an independent test set forMrOS Visit 2, with the rows for true labels and columns for predicted labels, or model outputs.

[0129] The data for FIG. 8E is reproduced in Table 10 below.Table 10 (MrOS Visit 2)

[0130] FIG. 8F shows a row normalized confusion matrix for sleep stage classifications (among Wake, N1, N2, N3, and REM stages) on an independent test set for APPLES, with the rows for true labels and columns for predicted labels, or model outputs.

[0131] The data for FIG. 8F is reproduced in Table 11 below.Table 9 (APPLES)

[0132] A 3-channel only (EEG, left EOG, EMG) foundational transformer model and GRU classifier (both the transformer and classifier were trained independently of MESA) (Example 1) was to compare performance to the 7-channel model (Example 2). FIG. 9 is a graph plot showing the results of the comparison. Cohen’s Kappa remained within 0.01 for all datasets, except for the MESA dataset, where it increased by 0.11, and decreased in APPLES by 0.04 when trained with the 7-channel model, as seen in FIG. 9. F1 scores remained similar except in N1 sleep where in all datasets except for MESA, F1 scores increased with the 3-channel model. AUPRC scores remained approximately the same, with small increases in N1 when training with the 3-channel model. In MESA, the opposite occurred and AUPRC scores were decreased across all stages when trained with the 3-channel model. REM sleep had a 0.15 AUPRC score increase in the 7-channel model. This was explored further in the model explainability section below.

[0133] Model explainabilitv

[0134] Model explainability was performed using the Integrated Gradients method. This method attributes the predictions to every single timepoint (in other words, every timepoint is treated as a feature). However, human-understandable features can include multiple timepoints. Additionally, calculating attributions for the full 8 hours is difficult to interpret as attributions are both positive and negative. Thus, there is a need to be able to effectively distinguish attributions that are contributing positively and negatively to the prediction. In certain embodiments, other time series explainable Al and attention-based relevance propagation for transformers can be used to provide further insights into interpreting these attributions of PSG signal data and the outcome.

[0135] Generally, the top 5% of the absolute value of attributions appeared to be within the EEG, EMG, EOG, and ECG channels, while SpO2 and respiration rates lacked attributions across all stages. Attributions for a particular sleep stage were notably present in prior or subsequent sleep epochs (independent of their sleep stage). In the examples, there were strong EEG attributions in N2 and attributions associated with prominent eye movements in REM.

[0136] Further, the 3 -channel and 7 -channel models were compared for a stage of REM sleep from MESA to try to further understand the higher performance in REM with the 7-channel model. The 3-channel model attributions for REM sleep in MESAhighlighted the EEG in the N2 stage before REM and the EOG during REM, but not in the REM stage after (this stage was misclassified as a wake). Consequently, the 7-channel model emphasized the EEG and EOG before, during, and after the REM stage, while also highlighting the ECG throughout (it classified all stages displayed correctly).

[0137] The foundational transformer, PFTSleep (105) of the present disclosure predicts sleep stages with comparable performance to other sleep stage-specific models. In contrast to the compared models, PFTSleep (105) was the only PSG-based model to input and encode an entire night’s (“full night”) sleep. In total, 8 hours of sleep signal data including brain, movement, cardiac, oxygen, and respiratory channels were encoded into a feature vector of size 7 x 4800 x 512. These features were used as input for the training of a supervised GRU classifier (150) that effectively extracted information from feature representations to predict sleep stages for the entire night. Notably, when training the GRU classifier (150), the foundational transformer (105) weights were frozen. Thus, this model used task-agnostic features to predict sleep stages.

[0138] In comparison to other sleep stage-specific Al models, performance was near state-of-the-art on SHHS, WSC, and MrOS visit 1 held-out test sets. While performance decreased on independent test sets MESA and APPLES, PFTSleep (105) (or system 100) still performed considerably well with moderate agreement with human scorers and Cohen’s Kappa values of 0.60 and 0.59 on a total of 3144 independent samples. Furthermore, wake, N2, and REM maintained high performance in these test sets, highlighting the model’s ability to distinguish wake, REM, and NREM sleep. Poorer performance in N1, N2, and N3 was expected, given that these stages had high variability among human scorers (0.24, 0.57, and 0.57 Cohen’s Kappa, respectively) and are more likely to be mislabeled. While MrOS visit 2 was treated as an independent test set, it was not fully independent as MrOS visit 1 was included in the training. Still, the model maintained performance on MrOS visit 2 patients years after the initial visit. MrOS visit 1 Cohen’s Kappa was near perfect agreement with human scorers with a value of 0.82, while MrOS visit 2 Cohen’s Kappa had substantial agreement at 0.75. This shows that the model remains consistent over time and could capture patient- specific features, emphasizing its potential use for longitudinal monitoring.

[0139] When PFTSleep’s GRU classifier (150) was trained with the addition of MESA, performance was near state-of-the-art for MESA with substantial agreement with human scorers and a Cohen’s Kappa of 0.76. Small performance increases were seen in other datasets as well, except for WSC (Cohen’ s Kappa decreased by 0.05). Furthermore,there were no identifiable performance increases when MESA was included in the foundational transformer’s (105) training, indicating that the features generated by PFTSleep (105) were similar, independent of whether MESA was included in the training process. Thus, to increase performance on unseen data, simply retraining the GRU classifier head (150 was adequate, which could be useful to adapt PFTSleep (105), or system 100, to new data quickly.

[0140] PFTSleep trained with only 3 channels (Example 1) performed well compared to the 7-channel PFTSleep (Example 2). In fact, it increased performance across N1 sleep in all datasets, except for MESA. Cohen’s Kappa was approximately the same across held-out test sets and MrOS visit 2 between the 7-channel and 3-channel models.However, the 7-channel model improved the performance on MESA, while slightly decreasing the performance on APPLES. It’s possible the addition of respiration rates and SpO2 added new, unseen features to the APPLES representations since nearly all APPLES subjects had at least mild obstructive sleep apnea. While the 3-channel model did perform well and an EEG, EMG, and EOG model is adequate for sleep staging, as demonstrated by other models, such as L-SeqSleepNet and USleep, the addition of other channels could provide additional feature representations to predict other outcomes in sleep beyond sleep stages.

[0141] While PFTSleep (105 and 150, or system 100) was compared to other models in sleep staging, PFTSleep of the present disclosure is fundamentally different and is directed to different goals. Rather than reaching the highest scores for sleep stage prediction across diverse datasets, such as USleep, PFTSleep is directed to learning robust feature representations of a full night of sleep. It was shown that these representations include information that can be extracted to predict sleep stages via another model (e.g., the GRU classifier 150) with high performance. According to some embodiments, there are additional features within these representations that can be used to predict even more outcomes related to sleep, not just sleep staging.

[0142] There are three other foundational models that also encode sleep study data. StagerNet developed by Banville et al. adapted self- supervised learning of EEG signals via a convolutional encoder to predict sleep stages. They encoded 30-second segments of multichannel EEG data and used the representations to predict sleep stages via logistic regression with 199 held-out samples from the Physionet Computing in Cardiology 2018 challenge. They also train another self-supervised convolutional encoder, ShallowNet, to predict abnormal EEGs. SleepFM built a foundationalconvolutional encoder and showed good performance on sleep stage classification and sleep-disordered breathing across multiple brain, cardiac, and respiratory channels. The input to their model was 30 seconds worth of multichannel sleep data. Their encoder learned relationships via a contrastive training mechanism among 30-second embeddings of brain, cardiac, and respiratory channels. During training, channels were separated into their own respective encoders and embeddings were compared pairwise and via a leave- one-out method. They assessed the model on an internal test set and an external test set of 100 samples from the Physionet Computing in Cardiology 2018 challenge. Both StagerNet and SleepFM did not validate their model on a larger, NSRR dataset for comparison. Finally, Ogg et al. built a foundational convolutional transformer using EEG signals and a sequence length of about 1 hour, demonstrating the utility of self-supervised models over fully supervised models and the encoder’s ability to translate to downstream tasks, such as sleep stage and brain age prediction. While they did utilize NSRR sleep studies, evaluation metrics were not reported for specific tasks, limiting comparison.

[0143] PFTSleep (or system 100) addresses several limitations in these studies. First, PFTSleep encodes a full night (8 hours) recording of sleep from channels that measure brain, movement, respiratory, oxygen, and cardiac activity. As demonstrated, these representations can be used to predict sleep stages for an entire night, in a single step. Additionally, PFTSleep learns representations of 6-second segments of sleep signal data, which could capture more specific sleep characteristics in comparison to encoders that model 30-second segments. Finally, PFTSleep was independently tested on 3144 samples, while StagerNet did not independently test, and SleepFM independently tested on only 100 samples from the Physionet challenge.

[0144] Advantages

[0145] The transformer architecture had recently gained notable popularity in the Al field through its use to train large language models (LLMs), such as bidirectional encoder representations from transformers (BERT) and Chat Generative Pretrained Transformers (ChatGPT). These have shown impressive performance in natural language understanding in many different domains, including medicine. Both language models are trained via a method known as self-supervision. In self-supervision, there are no outcomes or labels to predict. Instead, a model uses augmented or masked-out input data and attempts to predict or recreate the augmented values through the training procedure. For ChatGPT, the model is trained using sequences of words at varying lengths and it learns to predict the missing words at the end of the sequence, eventually learning to“understand” the language. BERT is similar but instead of predicting the missing words at the end of a sequence, it learns to predict missing words within sequences. During BERT’ s training procedure, as the name suggests, encoder “representations” are created as an intermediate step. These representations or feature vectors, which essentially summarize the original sequence of words that were passed into the model, can be used as input into other models to perform predictive tasks, such as classifying text or answering questions.

[0146] Notably, these feature representations can be learned from any data type, including images and time series data. Recent works have adapted similar training processes in medical imaging, clinical notes, and multi-modal health models. For example, a BERT-like model could be trained via self-supervision to learn to predict masked-out portions of images, effectively learning encoder representations that summarize the content in images. When training is complete, the model is hypothesized to have a “foundational” understanding of the input data, hence the term “foundational model.”

[0147] Recent advances in BERT-like models for time series data have expanded upon their proposed transformer representation framework to develop foundation models in various time series classification and forecasting tasks. However, the transformer’s performance was variable in comparison to simpler models until the release of the patch time series transformer (PatchTST), which increased performance on standardized datasets in the field in comparison to simpler models. The key idea behind the PatchTST architecture is that patches, or intervals of time series data (versus single time points), are more meaningful representations to encode because timepoints are meaningful in relation to nearby timepoints. Other transformer models with similar concepts have since been developed, such as time-frequency consistency (TFC) and MedFormer, and have made the case for utilizing the transformer in time series representation learning.

[0148] Previous work in automated sleep stage annotation have achieved reasonable accuracy and relatively high Cohen’s Kappa scores (0.75-0.83 depending on the dataset versus 0.76 for human scoring). Despite the better performance compared to humans, these models have two primary limitations. First, models are independently tested on small datasets, and adjusting to differences in sleep data acquisition and sleep staging annotation across clinics (or sleep datasets) has not been thoroughly explored. Second, models have heretofore been trained via full supervision to predict the sleep stage, potentially overfitting models to specific datasets they were trained on.

[0149] In addition to these primary limitations, model implementation limitations exist. First, existing models typically only utilize electroencephalograms (EEGs), electrooculography (EOG), and / or electromyography (EMG). While these signal channels may be adequate for sleep stage classification, they could limit the ability of the models to learn relationships among channels so that sleep stage algorithms can maintain performance on unseen datasets. Further, models are typically trained on 30-second segments or 30-second time-frequency images (sleep stages are labeled every 30 seconds per American Academy of Sleep Medicine guidelines), meaning that the input to a model is 30 seconds, and the output is a class indicating the sleep stage of that 30-second segment. Training models on 30-second segments of signal data, whether in the time or frequency domain, limits the ability of the model to classify sleep stages at intervals less than 30 seconds and identify subtle features that could help explain the sleep stage class, such as sleep spindles. In addition, models trained with 30-second inputs only correlate relationships between the input and output sleep stage label, without considering nearby or far away segments of the data, which could be meaningful. The longest interval of time considered (input data) in an existing model is 90 minutes while others typically only utilize 300 seconds or 30 seconds of sleep data.

[0150] In view of the foregoing, the present disclosure provides (1) a multichannel foundational transformer using 8-hour-long sleep study data, (2) that generates PSG representations suitable for sleep stage classification that performs equally or better than existing models as measured by Cohen’s Kappa, and (3) provides performance feedback on large independent test sets and elucidates whether differences in data collection or sleep stage annotation is a possible reason for decreased performance for unseen data.

[0151] Clinically, PFTSleep 105 of the present disclosure could help annotate sleep studies for wake, NREM, and REM and the GRU classifier 150 could be finetuned to specific clinic sleep data to classify all five sleep stages. Paired with the Integrated Gradients explainability method, PFTSleep 105 provides for discovering important aspects of PSG data.

[0152] Portions of the methods described herein can be performed by software or firmware in machine readable form on a tangible (e.g., non-transitory) storage medium. For example, the software or firmware can be in the form of a computer program including computer program code adapted to cause the system to perform various actions described herein when the program is run on a computer or suitable hardware device, andwhere the computer program can be embodied on a computer readable medium. Examples of tangible storage media include computer storage devices having computer- readable media such as disks, thumb drives, flash memory, and the like, and do not include propagated signals. Propagated signals can be present in a tangible storage media. The software can be suitable for execution on a parallel processor or a serial processor such that various actions described herein can be carried out in any suitable order, or simultaneously.

[0153] The headings used herein are for organizational purposes only and are not meant to be used to limit the scope of the description or the claims. As used throughout this application, the words "may" and "can" are used in a permissive sense (i.e., meaning having the potential to), rather than the mandatory sense (i.e., meaning must). To facilitate understanding, like reference numerals have been used, where possible, to designate like elements common to the figures. In certain instances, a letter suffix following a dash (. ..-b) denotes a specific example of an element marked by a particular reference numeral (e.g., 210-b). Description of elements with references to the base reference numerals (e.g., 210) also refer to all specific examples with such letter suffixes (e.g., 210-b), and vice versa.

[0154] It is to be further understood that like or similar numerals in the drawings represent like or similar elements through the several figures, and that not all components or steps described and illustrated with reference to the figures are required for all embodiments or arrangements.

[0155] The terminology used herein is for the purpose of describing particular embodiments only and is not intended to be limiting of the disclosure. As used herein, the singular forms "a," "an," and "the" are intended to include the plural forms as well, unless the context clearly indicates otherwise. It will be further understood that the terms “contains”, “containing”, “includes”, “including,” “comprises”, and / or “comprising,” and variations thereof, when used in this specification, specify the presence of stated features, integers, steps, operations, elements, and / or components, but do not preclude the presence or addition of one or more other features, integers, steps, operations, elements, components, and / or groups thereof, and are meant to encompass the items listed thereafter and equivalents thereof as well as additional items.

[0156] While the disclosure has described several example implementations, it will be understood by those skilled in the art that various changes can be made, and equivalents can be substituted for elements thereof without departing from the spirit andscope of the disclosure. In addition, many modifications will be appreciated by those skilled in the art to adapt a particular instrument, situation, or material to embodiments of the disclosure without departing from the essential scope thereof. Therefore, it is intended that the disclosure not be limited to the particular embodiments disclosed, or to the best mode contemplated for carrying out this disclosure, but that the disclosure will include all embodiments falling within the scope of the appended claims.

[0157] The subject matter described above is provided by way of illustration only and should not be construed as limiting. Various modifications and changes can be made to the subject matter described herein without following the example embodiments and applications illustrated and described, and without departing from the true spirit and scope encompassed by the present disclosure, which is defined by the set of recitations in the following claims and by structures and functions or steps which are equivalent to these recitations.

Claims

WHAT IS CLAIMED IS:

1. A computer-implemented method for encoding full-night, multichannel sleep study data via a foundational transformer neural network, the method comprising: generating, by at least one computing device employing a patch-based, selfsupervised transformer neural network model, encodings, wherein the encodings are generated using full-night, multi-channel polysomnography data; inputting, by at least one computing device, the encodings into a supervised, bidirectional, gated recurrent unit probing head; and classifying, by at least one computing device as a function of the encodings input to the supervised, bidirectional, gated recurrent unit probing head, sleep stages.

2. The method of claim 1, wherein the sleep stage classification is implemented using a foundational transformer.

3. The method of claim 2, wherein the foundational transformer predicts sleep stages with a single step.

4. The method of claim 1, wherein the full-night, multi-channel polysomnography data include sleep signal data collected over at least eight hours.

5. The method of claim 4, wherein the sleep signal data are associated with at least one of brain, movement, cardiac, oxygen, and respiratory channels.

6. The method of claim 1 , further comprising providing, by at least one computing device, the classified sleep stages to at least one clinical application.

7. The method of claim 1, further comprising annotating, by at least one computing device, a sleep study.

8. The method of claim 1, further comprising preprocessing the full-night, multichannel polysomnography data, wherein the preprocessing comprises at least one of resampling, instance normalization, and patching.

9. The method of claim 1, wherein the full-night, multi-channel polysomnography data is in patches of a predetermined duration that is less than 30 seconds, or about 4 seconds to about 8 seconds, or preferably 6 seconds.

10. The method of claim 9, wherein the full-night, multi-channel polysomnography data comprises random zero padding within the patches to form sleep signal data having a duration of about 6 hours to about 10 hours, or preferably 8 hours.

11. A computer-implemented system for encoding full-night, multichannel sleep study data via a foundational transformer neural network, the system comprising: at least one computing device configured by executing processor-readable instructions that configure the at least one computing device for: employing a patch-based, self-supervised transformer neural network model to generate encodings, wherein the encodings are generated using full-night, multi-channel polysomnography data; inputting the encodings into a supervised, bidirectional, gated recurrent unit probing head; and classifying, as a function of the encodings input to the supervised, bidirectional, gated recurrent unit probing head, sleep stages.

12. The system of claim 11, wherein the sleep stage classification is implemented using a foundational transformer.

13. The system of claim 11 , wherein the foundational transformer predicts sleep stages with a single step.

14. The system of claim 11, wherein the full-night, multi-channel polysomnography data include sleep signal data collected over at least eight hours.

15. The system of claim 14, wherein the sleep signal data are associated with at least one of brain, movement, cardiac, oxygen, and respiratory channels.

16. The system of claim 11, wherein the at least one computing device is further configured for providing the classified sleep stages to at least one cli nical application.

17. The system of claim 11, wherein the at least one computing device is further configured for annotating a sleep study.

18. The system of claim 1 1 , wherein the processor-readable instructions further configure the at least one computing device for preprocessing the full-night, multichannel polysomnography data, wherein the preprocessing comprises at least one of resampling, instance normalization, and patching.

19. The system of claim 11, wherein the full-night, multi-channel polysomnography data is in patches of a predetermined duration that is less than 30 seconds, or about 4 seconds to about 8 seconds, or preferably 6 seconds.

20. The system of claim 19, wherein the full-night, multi-channel polysomnography data comprises random zero padding within the patches to form sleep signal data having a duration of about 6 hours to about 10 hours, or preferably 8 hours.

Citation Information

Patent Citations

  • Methods for Explainability of Deep-Learning Models

    US20200151516A1

  • System and method for determining sleep stages based on non-cardiac body signals

    US20210085242A1

  • Learning representations of EEG signals with self-supervised learning

    US20230306267A1