Sleep diagnosis system based on sleep language model

By introducing SleepGPT and hierarchical Transformer networks, the problems of time-consuming and labor-intensive sleep diagnosis, prediction bias, and limitations of long-term series modeling are solved, achieving efficient and accurate sleep staging and disorder diagnosis, and improving the model's generalization ability and robustness.

CN121196483APending Publication Date: 2025-12-26SOUTH CHINA UNIV OF TECH
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202511531544.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-10-24
Publication Date
2025-12-26

AI Technical Summary

Technical Problem

Existing sleep diagnosis methods are time-consuming, labor-intensive, and highly subjective. They suffer from model prediction bias, limitations in long-term series modeling, and poor feature reuse capabilities, making them difficult to generalize in real-world clinical scenarios.

Method used

We use the SleepGPT sleep language model based on the GPT-2 architecture for context modeling, and combine it with a hierarchical Transformer network to extract local context features from short sleep segments. We then use a Transformer encoder to capture global context features and design a hierarchical Transformer network for sleep disorder diagnosis.

Benefits of technology

It improves the accuracy of sleep staging and the model's generalization ability in real clinical scenarios, reduces data annotation costs, overcomes the limitations of long-term sequence modeling, and enhances the efficiency and interpretability of sleep disorder diagnosis.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121196483A_ABST
    Figure CN121196483A_ABST
Patent Text Reader

Abstract

The invention discloses a sleep diagnosis system based on a sleep language model, and solves the problems of model prediction deviation, difficulty in long sequence modeling, poor feature reusability and the like in the existing sleep diagnosis method. The system comprises a data loading module, a data preprocessing module, a sleep stage staging module, a sleep stage staging calibration module and a sleep disorder diagnosis module. According to the system, preliminary sleep staging is carried out on electroencephalogram signals through a deep learning model; a sleep language model (called SleepGPT) based on a GPT-2 architecture is introduced, context semantic information of a sleep stage sequence is extracted, a preliminary staging result is calibrated in combination with a dynamic weighting mechanism, and the staging accuracy is improved; a hierarchical Transform network is designed, pre-trained SleepGPT is multiplexed to extract local features from a short sequence, global context features are integrated through a Transform encoder, and finally efficient sleep disorder auxiliary diagnosis is achieved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the technical field of sleep diagnosis and artificial intelligence, and in particular to a sleep diagnosis system based on a sleep language model. Background Technology

[0002] Sleep diagnostic technology is a comprehensive approach that involves collecting, recording, and analyzing sleep-related physiological signals and behavioral data to objectively assess sleep structure and identify sleep disorders. Its core lies in collecting, recording, and analyzing sleep-related physiological signals and behavioral data, transforming subjective sleep experiences into quantifiable objective indicators, namely different sleep stages, providing a scientific basis for the screening, diagnosis, classification, severity assessment, and treatment effectiveness evaluation of sleep disorders. Polysomnography (PSG) is a core examination technique that simultaneously records multiple physiological signals during sleep. The recorded data includes physiological activity signals such as electroencephalogram (EEG), electrooculogram (EOG), and electromyogram (EMG), possessing objectivity and completeness, and is widely used in sleep diagnosis.

[0003] Traditional sleep diagnosis relies on manual interpretation of sleep symptom recordings (PSGs). This involves dividing the entire night's recorded signal sequence into 30-second segments. Experts visually inspect each segment and label it with a sleep stage category: Wake, N1, N2, N3, REM, MOVE, and UNK. Wake represents wakefulness; N1, N2, and N3 represent stages I, II, and III of non-rapid eye movement (NREM) sleep, respectively; REM represents rapid eye movement (REM); MOVE represents motion artifacts; and UNK represents an unknown stage. This process generates a sleep stage sequence, characterizing the macroscopic structure of sleep. This manual, frame-by-frame labeling method is time-consuming, labor-intensive, and highly subjective, with a lack of consistency among experts. Therefore, developing efficient automated sleep staging and diagnosis technology is of great significance.

[0004] Although existing deep learning models can achieve automated staging, they still face the following limitations: model prediction bias, such as the imbalance of sleep stage categories and the scarcity of N1 stage samples, which leads to model prediction bias towards the majority class, restricting the model's generalization ability and robustness in real clinical scenarios; limitations in long-term sequence modeling, making it difficult to capture the global features of the entire night's sleep sequence; and the need to train independent models separately for sleep disorder diagnosis, resulting in poor feature reuse capabilities.

[0005] To address the aforementioned issues, this invention proposes a sleep diagnosis system based on a sleep language model. The system first extracts EEG data from the sleep apnea-serum (PSG), performing automated segmentation and signal cleaning to suppress noise and preserve valid physiological patterns. Subsequently, a deep learning model is used to perform preliminary sleep staging, and an innovative GPT-2-based sleep language model is introduced to treat sleep stage sequences as "sentences" for contextual modeling, calibrating the preliminary staging results and significantly improving staging accuracy. Furthermore, a hierarchical Transformer network is designed to extract temporal patterns of wakefulness and dynamic features of stage transitions from the sleep stage sequences, providing an interpretable and efficient computational tool for the early identification and clinical auxiliary diagnosis of sleep disorders.

[0006] The challenges of this invention lie in the construction and adaptation of a sleep language model, long-term sequence modeling, and hierarchical feature extraction. Transferring natural language processing model architectures like GPT-2 to sleep sequence analysis requires effectively encoding sleep stage sequences into "word vectors" and pre-training a sleep-specific language model on large-scale sleep data. This enables the model to understand the transition patterns of sleep stages, allowing for context-aware calibration of errors in the initial staging. Furthermore, sleep is a complex process encompassing both local and global structures, including local features of spindle waves and K-complex waves, as well as global features of the NREM-REM cycle. The number of time steps in a full night's sleep staging sequence far exceeds the sequence length limitations of sleep language models. A hierarchical Transformer network is designed to extract local contextual features from short sleep segments using the sleep language model, and to capture global contextual features from the output of the sleep language model using a Transformer encoder. This achieves the preservation of local waveform features, the capture of sleep stage phases and transition patterns, and demonstrates strong feature reuse capabilities, overcoming the limitations of long-term sequence modeling. Summary of the Invention

[0007] The purpose of this invention is to address the problems of current clinical sleep diagnosis methods, such as reliance on high-quality labeled data, limitations due to long-term sequence modeling, imbalanced sleep stage classification leading to majority class bias in predictions, and the need for separate training of independent models with poor feature reuse capabilities for sleep disorder diagnosis. This invention provides a sleep diagnosis system based on a sleep language model. This system achieves efficient and accurate sleep staging through automated and high-precision analysis of EEG data in PSG files, and extracts wakefulness and transition features using the sleep stage sequence structure, providing an efficient tool for the early identification and clinical auxiliary diagnosis of sleep disorders.

[0008] To achieve the above objectives, the technical solution provided by this invention is: a sleep diagnosis system based on a sleep language model, comprising: The data loading module is used to acquire and load multi-channel EEG data and sleep stage category labels from the PSG file. The sleep stage category labels include Wake, N1, N2, N3, REM, MOVE, and UNK, where Wake represents the wakefulness stage, N1, N2, and N3 represent stages I, II, and III of non-rapid eye movement sleep (NREM), respectively, REM represents the rapid eye movement sleep stage, MOVE represents motion artifacts, and UNK represents the unknown stage. The data preprocessing module is used to preprocess the multi-channel EEG data acquired and loaded by the data loading module, remove the EEG data labeled MOVE and UNK and their corresponding EEG data, filter out high-quality data segments, and output EEG segments that are divided according to a fixed duration and noise removed, i.e., epochs. The sleep stage segmentation module uses a pre-trained deep learning model to perform preliminary segmentation of the epochs output by the data preprocessing module, outputs the probability distribution sequence of different sleep stages corresponding to each epoch, and records the corresponding sleep stage category label predicted for each epoch, generating a preliminary segmentation of the sleep stage sequence. The sleep stage segmentation calibration module uses a pre-trained sleep language model to receive the probability distribution vector corresponding to different sleep stages for each epoch from the deep learning model. Based on this probability distribution vector, it predicts the probability distribution vector of the sleep stage category label corresponding to the current epoch, and dynamically weights it with the probability distribution vector of the sleep stage category label corresponding to the current epoch output by the deep learning model to generate a calibrated sleep stage sequence after initially segmenting the sleep stage sequence. The sleep language model is called SleepGPT, which adopts a lightweight GPT-2 architecture, including a token embedding layer, a position encoding layer, a three-layer Transformer decoder, and a linear layer. The token embedding layer maps the discrete sleep stage category labels (tokens) in the sleep stage sequence to a fixed-dimensional continuous vector representation. The position encoding layer maps the position index of the sleep stage sequence to a fixed-dimensional position vector, providing the model with the temporal information of the sleep stage in the sequence. The sequence position information enables the model to understand the temporal dependence and periodicity of sleep stages. The Transformer decoder includes a multi-head self-attention layer, a feedforward neural network, and a residual connection and normalization layer. The multi-head self-attention layer processes the input sequence in parallel through five independent attention heads and applies a causal attention mechanism in the time dimension, enabling the model to simultaneously pay attention to the contextual features of the input sequence. This ensures that the prediction at each time step depends only on the current and previous sequence information, preventing the leakage of future information. The feedforward neural network is a two-layer fully connected network that performs nonlinear transformations and feature enhancements on the output of the multi-head self-attention layer, learning the complex feature combinations and transformation rules in the sleep stage sequence. The residual connection and normalization layer maintains the original information flow through residual connections and stabilizes the feature distribution through layer normalization. The linear layer maps the output vector of the Transformer decoder to a probability distribution sequence of the predicted categories of the five sleep stages: Wake, N1, N2, N3, and REM. The sleep disorder diagnosis module predicts sleep disorder labels for subjects' long sleep stage sequences throughout the night based on a hierarchical Transformer network. The hierarchical Transformer network includes a pre-trained SleepGPT, a Transformer encoder, and a linear layer. The pre-trained SleepGPT acts as a feature extractor, extracting local contextual features from short sleep stage sequences. The Transformer encoder captures global contextual features from the output of SleepGPT, outputting a global feature vector sequence. Finally, the output global feature vector sequence is input into the linear layer to predict sleep disorder labels.

[0009] Furthermore, the data loading module includes a data source acquisition unit and an EEG reading unit, wherein: The data source acquisition unit is used to acquire PSG files and corresponding annotation files from local or wearable devices. The PSG file contains multi-channel physiological signals, including electroencephalogram (EEG), electrooculogram (EOG), and electromyogram (EMG). The EEG includes head information and data recordings. The head information includes the date and time of recording, the recording duration, the electrodes used and their locations, labels for each channel, and other recording device and patient information. The data recordings include physiological signal data for each channel acquired at a predetermined sampling frequency. The annotation file contains sleep stage category labels and records the start time and duration of each sleep stage. The EEG reading unit reads the EEG of a specified channel from the PSG file. The EEG is initially a one-dimensional array with a length equal to the total number of samples, where the total number of samples is determined by the file duration × sampling frequency.

[0010] Furthermore, the data preprocessing module performs a series of preprocessing operations on the EEG from the specified channel read by the EEG reading unit in the data loading module, obtaining EEG segments divided into fixed durations and noise-removed, forming a set of epochs, and saving them in NPZ format; wherein, the series of preprocessing operations include dividing the EEG into fixed durations, removing invalid labels, and parsing the sleep stage category labels to correspond with the segmented signal data; dividing the EEG into fixed durations means dividing the EEG into epochs according to a fixed duration, where the fixed duration is the duration of one epoch; removing invalid labels means removing the labels MOVE and UNK and their corresponding EEG data.

[0011] Furthermore, the deep learning model is XSleepNet, which adopts a dual-stream network architecture to process the original waveform signal and time-frequency features simultaneously. It also integrates multi-dimensional information through an adaptive fusion mechanism and finally outputs the classification result through a classification layer.

[0012] Furthermore, the sleep stage calibration module performs the following operations: 1) Receive the initial sleep stage sequence output by XSleepNet, and obtain the probability distribution sequence of each epoch corresponding to different sleep stages, defined as... Each element is a 5-dimensional probability distribution vector. If the total number of epochs is represented, then the 1st epoch... element ,satisfy , This represents the probability of an epoch corresponding to the awakening stage label. These represent the probabilities of the epoch corresponding to stage I, II, and III of the non-rapid eye movement sleep (NREM) sleep pattern, respectively. This represents the probability of an epoch corresponding to a REM sleep stage category label; 2) Change the sequence Save it as a two-dimensional array as input to SleepGPT, with shape as follows: , where 5 represents the predicted probability of the 5 sleep stage category labels corresponding to each epoch; 3) Use a fixed length of The context window, in steps In sequence Slide up; for the first... A 5-dimensional probability distribution vector ,in , cut The former There are 5-dimensional probability distribution vectors, that is, starting from the (t-k+1)th 5-dimensional probability distribution vector. Begin, take A continuous 5-dimensional probability distribution vector, denoted as ,in Reference The probability distribution subsequence is, The above information indicates that, as input to the model, the model will... Predict the sleep stage category label corresponding to the t-th epoch; 4) The input vector constructed in step 3) is processed through a token embedding layer and a position encoding layer. Mapped to a higher-dimensional space, and for Inject positional information into each of the 5-dimensional probability distribution vectors: ; In the formula, This represents the embedding vector output after processing through the token embedding layer and the position encoding layer. The token embedding matrix, which is the weight matrix of the token embedding layer, is then compared with... The dot product maps each sleep stage category to a unique 48-dimensional vector representation, forming a learnable embedding matrix with a shape of 5×48. The matrix contains five sleep stage category labels: Wake, N1, N2, N3, and REM. Each sleep stage category label corresponds to a unique vector representation. The position embedding matrix maps the position indices of the sleep stage sequence, i.e., integers from 0 to 89, to a 48-dimensional position vector matrix with a shape of 90×48. This provides the model with temporal position information of the sleep stage in the sequence, enabling the model to understand the temporal dependence and periodicity of the sleep stage. 5) The vector output in step 4) The input is processed by a 3-layer Transformer decoder. The processing procedure is as follows: ; In the formula, Indicates Transformer decoder, This represents the vector output by the l-th layer Transformer decoder. Indicates the first The vector output by the layer Transformer decoder, when hour for ; 6) Input the output vector obtained in step 5) into the linear layer to output the sleep stage prediction label corresponding to the t-th epoch. probability distribution , represented as: ; In the formula, SLM represents SleepGPT. This represents the predicted label sequence of sleep stages corresponding to the first t-1 epochs. Represents a linear layer. This represents the weight matrix of the linear layer in SleepGPT; 7) Calculate the probability distribution of XSleepNet for different sleep stage category labels corresponding to the current epoch. and Through weighting factors The probabilities of different sleep stages are then merged to obtain the final sleep stage probabilities. predict: ; In the formula, SSM represents XSleepNet, and the weighting factor is... By controlling the relative influence of the two models XSleepNet and SleepGPT, the sleep stage category label with the highest prediction probability is finally selected as the prediction result, and the calibrated sleep stage sequence is output.

[0013] Furthermore, the sleep disorder diagnosis module performs the following operations: 1) The subject's overnight long sleep stage sequence was segmented into short sleep stage sequences, and these sequences were padded into a batch to achieve a uniform length, represented as: ; In the formula, U represents the sequence of long sleep stages throughout the night for the subject. This represents the segmentation of the long sleep stage sequence into its first segment. Short sleep stage sequence, Indicates the number of segments in the long sleep phase sequence; 2) Using the pre-trained SleepGPT from the sleep stage staging calibration module, local contextual features are extracted from short sleep stage sequences, represented as: ; In the formula, The first generation generated for SleepGPT There are local feature vectors, and the sequence of all local feature vectors is . The input sequence that constitutes the Transformer encoder in a hierarchical Transformer network, where Corresponding to the CLS tag, in the Transformer architecture, the CLS tag is a special tag added to the beginning of the input sequence. Its final output can serve as a global representation of the entire sequence, which is suitable for classification tasks. 3) Use the Transformer encoder to capture global contextual features from the output of SleepGPT, represented as: ; In the formula, Represents the Transformer encoder. The output sequence of the Transformer encoder, i.e., the global feature vector sequence of the sleep phase sequence, is denoted as... , For the first All feature vectors, The global feature vector representing the output position of the CLS marker; 4) Input the global feature vector sequence obtained in step 3) into the linear layer to predict sleep disorder labels, as follows: ; In the formula, This represents the probability of predicting a sleep disorder label corresponding to a subject's sleep stage sequence. and These represent the weight matrix and parameter representation of the linear layer in a hierarchical Transformer network.

[0014] Compared with the prior art, the present invention has the following advantages and beneficial effects: 1. This invention designs a novel sleep language model called SleepGPT. The model's multi-head self-attention layer processes the input sequence in parallel through five independent attention heads and applies a causal attention mechanism in the temporal dimension. This allows the model to simultaneously focus on the contextual features of the input sequence, ensuring that predictions at each time step rely only on current and previous sequence information, preventing the leakage of future information and significantly improving the model's generalization ability on diverse datasets. Its "pre-training-fine-tuning" self-supervised learning paradigm helps solve the problem of supervised deep learning models heavily relying on high-quality labeled data for sleep staging, reducing data labeling costs.

[0015] 2. This invention introduces a weighting factor in the sleep stage staging calibration module. It controls the impact of the probability of the predicted sleep stage category label by XSleepNet and SleepGPT on the final prediction result, solves the problem of sleep stage staging bias in a single model, improves the overall staging accuracy, and enhances the model's generalization ability and robustness in real clinical scenarios.

[0016] 3. This invention designs a hierarchical Transformer network in the sleep disorder diagnosis module, reusing SleepGPT as a feature extractor to extract local contextual features from short sleep segments. It then uses a Transformer encoder to capture global contextual features from the output of the sleep language model, demonstrating strong feature reuse capabilities and eliminating the need for separately trained independent models for sleep disorder diagnosis. Furthermore, the method of segmenting, extracting features, and hierarchically predicting the subject's entire nightly long sleep sequence also addresses the limitations of long-term sequence modeling. Attached Figure Description

[0017] Figure 1 This is a schematic diagram showing the relationship between the various modules of the system of the present invention. Detailed Implementation

[0018] The present invention will be further described in detail below with reference to the embodiments and accompanying drawings, but the embodiments of the present invention are not limited thereto.

[0019] This embodiment discloses a sleep diagnosis system based on a sleep language model. It is a sleep staging and disorder diagnosis system developed using Python and capable of running on Windows devices. Figure 1 As shown, the system includes a data loading module, a data preprocessing module, a sleep stage staging module, a sleep stage staging calibration module, and a sleep disorder diagnosis module. This embodiment uses the overnight PSG data of subject ID 01 in the SS3 subset of the public dataset MASS as an example to provide a detailed description of the system implementation. The specific descriptions of each module are as follows: The data loading module is used to acquire and load multi-channel EEG data and sleep stage category labels from the PSG file. The sleep stage category labels include Wake, N1, N2, N3, REM, MOVE, and UNK, where Wake represents the wakefulness stage, N1, N2, and N3 represent stages I, II, and III of non-rapid eye movement sleep (NREM), respectively, REM represents the rapid eye movement sleep stage, MOVE represents motion artifacts, and UNK represents the unknown stage.

[0020] The data preprocessing module is used to preprocess the multi-channel EEG data acquired and loaded by the data loading module, remove the EEG data labeled MOVE and UNK and their corresponding labels, filter out high-quality data segments, and output EEG segments that are divided into fixed duration segments and noise removed, i.e., epochs.

[0021] The sleep stage segmentation module uses a pre-trained deep learning model to perform preliminary segmentation of the epochs output by the data preprocessing module. It outputs the probability distribution vector of each epoch corresponding to different sleep stages and records the corresponding sleep stage category label predicted for each epoch, generating a preliminary sleep stage sequence.

[0022] The sleep stage segmentation calibration module uses a pre-trained sleep language model to receive the probability distribution vector corresponding to different sleep stages for each epoch from the deep learning model. Based on this probability distribution vector, it predicts the probability distribution vector of the sleep stage category label corresponding to the current epoch, and dynamically weights it with the probability distribution vector of the sleep stage category label corresponding to the current epoch output by the deep learning model to generate a calibrated sleep stage sequence after initially segmenting the sleep stage sequence. The sleep language model is called SleepGPT, which adopts a lightweight GPT-2 architecture, including a token embedding layer, a position encoding layer, a three-layer Transformer decoder, and a linear layer. The token embedding layer maps the discrete sleep stage category labels (tokens) in the sleep stage sequence to a fixed-dimensional continuous vector representation. The position encoding layer maps the position index of the sleep stage sequence to a fixed-dimensional position vector, providing the model with the temporal information of the sleep stage in the sequence. The sequence position information enables the model to understand the temporal dependence and periodicity of sleep stages. The Transformer decoder includes a multi-head self-attention layer, a feedforward neural network, and a residual connection and normalization layer. The multi-head self-attention layer processes the input sequence in parallel through five independent attention heads and applies a causal attention mechanism in the time dimension, enabling the model to simultaneously pay attention to the contextual features of the input sequence. This ensures that the prediction at each time step depends only on the current and previous sequence information, preventing the leakage of future information. The feedforward neural network is a two-layer fully connected network that performs nonlinear transformations and feature enhancements on the output of the multi-head self-attention layer, learning the complex feature combinations and transformation rules in the sleep stage sequence. The residual connection and normalization layer maintains the original information flow through residual connections and stabilizes the feature distribution through layer normalization. The linear layer maps the output vector of the Transformer decoder to the probability distribution sequence of the predicted sleep stage category labels for Wake, N1, N2, N3, and REM.

[0023] The sleep disorder diagnosis module predicts sleep disorder labels for subjects' long sleep stage sequences throughout the night based on a hierarchical Transformer network. The hierarchical Transformer network includes a pre-trained SleepGPT, a Transformer encoder, and a linear layer. The pre-trained SleepGPT acts as a feature extractor, extracting local contextual features from short sleep stage sequences. The Transformer encoder captures global contextual features from the output of SleepGPT, outputting a global feature vector sequence. Finally, the output global feature vector sequence is input into the linear layer to predict sleep disorder labels.

[0024] Specifically, the data loading module includes a data source acquisition unit and an EEG reading unit. The data source acquisition unit acquires a PSG file and corresponding annotation file from a local device or wearable device. The PSG file contains multi-channel physiological signals, including electroencephalogram (EEG), electrooculogram (EOG), and electromyogram (EMG). The EEG includes head information and data recordings. Head information includes the recording date and time, recording duration, electrodes used and their locations, channel labels, and other recording device and patient information. Data recordings include physiological signal data for each channel acquired at a predetermined sampling frequency. The annotation file contains sleep stage category labels and records the start time and duration of each sleep stage. For example, it loads the PSG file 01_EEG.edf and its corresponding annotation file from the subject with ID 01 in the SS3 subset of MASS from a specified local path. The annotation file records the sleep stage category labels corresponding to each EEG segment, with a duration of 30 seconds, including Wake, N1, N2, N3, REM, MOVE, and UNK. The EEG reading unit reads the PSG file and selects the EEG signals from two differential channels, C3-A2 and C4-A1, based on clinical criteria. This signal is loaded as a one-dimensional array, the total length of which is determined by the file duration of 8 hours and the sampling frequency of 256 Hz, i.e., file duration × sampling frequency, approximately 7,372,800 data points.

[0025] Specifically, the data preprocessing module performs a series of preprocessing operations on the EEG from the specified channel read by the EEG reading unit in the data loading module, obtaining EEG segments divided into fixed-duration segments and noise-removed segments, forming a set of epochs, and saving them in NPZ format; wherein, the series of preprocessing operations include dividing the EEG into segments of fixed duration, removing invalid labels, and parsing the sleep stage category labels to correspond with the segmented signal data; dividing the EEG into segments of fixed duration means dividing the EEG into epochs of fixed duration, where fixed duration is one epoch. h The duration of the interval is specified; the removal of invalid tags refers to removing the tags MOVE and UNK and their corresponding EEG data. In this example, the continuous EEG signal is divided into 30-second intervals, resulting in 960 epochs. Each epoch contains 7680 data points, calculated based on 30 seconds × 256 Hz. According to the annotation file, epochs with the tags MOVE and UNK are identified and removed. In this embodiment, 12 invalid epochs are removed, leaving 948 valid epochs and their corresponding Wake, N1, N2, N3, and REM tags. The preprocessed EEG data, tag array, and sampling rate information are saved together as an NPZ format file for use by subsequent modules.

[0026] Specifically, the sleep stage segmentation module uses the XSleepNet deep learning model, which employs a two-stream network architecture to simultaneously process the original waveform signal and time-frequency features. It integrates multi-dimensional information through an adaptive fusion mechanism and finally outputs the classification result through a classification layer. In this example, EEG data from 948 epochs, with a shape of [948, 2, 7680] (representing 948 epochs, 2 channels, and 7680 data points), is input into the pre-trained XSleepNet model. The model outputs a 5-dimensional probability distribution vector for each epoch. The label with the highest probability in each vector is taken as the initial stage of that epoch, thus generating a preliminary sleep stage sequence.

[0027] Specifically, the sleep stage calibration module performs the following operations: 1) Receive the initial sleep stage sequence output by XSleepNet, and obtain the probability distribution sequence of each epoch corresponding to different sleep stages, defined as... Each element is a 5-dimensional probability distribution vector. If the total number of epochs is represented, then the 1st epoch... element ,satisfy , This represents the probability of an epoch corresponding to the awakening stage label. These represent the probabilities of the epoch corresponding to stage I, II, and III of the non-rapid eye movement sleep (NREM) sleep pattern, respectively. This represents the probability of an epoch corresponding to a REM sleep stage category label; 2) Change the sequence Save it as a two-dimensional array as input to SleepGPT, with shape as follows: , where 5 represents the predicted probability of the 5 sleep stage category labels corresponding to each epoch; 3) Use a fixed length of The context window, in steps In sequence Slide up; for the first... A 5-dimensional probability distribution vector ,in , cut The former There are 5-dimensional probability distribution vectors, that is, starting from the (t-k+1)th 5-dimensional probability distribution vector. Begin, take A continuous 5-dimensional probability distribution vector, denoted as ,in Reference The probability distribution subsequence is, The above information indicates that, as input to the model, the model will... Predict the sleep stage category label corresponding to the t-th epoch; 4) The input vector constructed in step 3) is processed through a token embedding layer and a position encoding layer. Mapped to a higher-dimensional space, and for Inject positional information into each of the 5-dimensional probability distribution vectors: ; In the formula, This represents the embedding vector output after processing through the token embedding layer and the position encoding layer. The token embedding matrix, which is the weight matrix of the token embedding layer, is then compared with... The dot product maps each sleep stage category to a unique 48-dimensional vector representation, forming a learnable embedding matrix with a shape of 5×48. The matrix contains five sleep stage category labels: Wake, N1, N2, N3, and REM. Each sleep stage category label corresponds to a unique vector representation. The position embedding matrix maps the position indices of the sleep stage sequence, i.e., integers from 0 to 89, to a 48-dimensional position vector matrix with a shape of 90×48. This provides the model with temporal position information of the sleep stage in the sequence, enabling the model to understand the temporal dependence and periodicity of the sleep stage. 5) The vector output in step 4) The input is processed by a 3-layer Transformer decoder. The processing procedure is as follows: ; In the formula, Indicates Transformer decoder, This represents the vector output by the l-th layer Transformer decoder. Indicates the first The vector output by the layer Transformer decoder, when hour for ; 6) Input the output vector obtained in step 5) into the linear layer to output the sleep stage prediction label corresponding to the t-th epoch. probability distribution , represented as: ; In the formula, SLM represents SleepGPT. This represents the predicted label sequence of sleep stages corresponding to the first t-1 epochs. Represents a linear layer. This represents the weight matrix of the linear layer in SleepGPT; 7) Calculate the probability distribution of XSleepNet for different sleep stage category labels corresponding to the current epoch. and Through weighting factors The probabilities of different sleep stages are then merged to obtain the final sleep stage probabilities. predict: ; In the formula, SSM represents XSleepNet, and the weighting factor is... By controlling the relative influence of the two models XSleepNet and SleepGPT, the sleep stage category label with the highest prediction probability is finally selected as the prediction result, and the calibrated sleep stage sequence is output.

[0028] Using a specific example, the probability distribution sequence of 948 epochs output by XSleepNet is taken as input. The probability distribution sequence has a shape of [948, 5], where 5 represents the 5 sleep stage category labels predicted based on the epoch. A length of... A sliding window, approximately 45 minutes of data, in steps. Slide along the sequence. For the t-th epoch, Extract the probability distribution subsequence of its first 90 epochs. SleepGPT is used as the context sequence input. The input sequence is processed through a token embedding layer and a position encoding layer. Mapped to a high-dimensional embedding vector . After processing by a 3-layer Transformer decoder, the output context encoding vector is obtained. Context encoding vector The probability distribution of the current epoch predicted by SleepGPT based on context is obtained through a linear layer. The output of XSleepNet With SleepGPT output According to the formula The fusion is performed. This embodiment sets a weighting factor. =0.6. Final value: The label with the highest probability is used as the calibrated staging result.

[0029] Specifically, the sleep disorder diagnosis module performs the following operations: 1) The subject's overnight long sleep stage sequence was segmented into short sleep stage sequences, and these sequences were padded into a batch to achieve a uniform length, represented as: ; In the formula, U represents the sequence of long sleep stages throughout the night for the subject. This represents the segmentation of the long sleep stage sequence into its first segment. Short sleep stage sequence, Indicates the number of segments in the long sleep phase sequence; 2) Using the pre-trained SleepGPT from the sleep stage staging calibration module, local contextual features are extracted from short sleep stage sequences, represented as: ; In the formula, The first generation generated for SleepGPT There are local feature vectors, and the sequence of all local feature vectors is . The input sequence that constitutes the Transformer encoder in a hierarchical Transformer network, where Corresponding to the CLS tag, in the Transformer architecture, the CLS tag is a special tag added to the beginning of the input sequence. Its final output can serve as a global representation of the entire sequence, which is suitable for tasks such as classification.

[0030] 3) Use the Transformer encoder to capture global contextual features from the output of SleepGPT, represented as: ; In the formula, Represents the Transformer encoder. The output sequence of the Transformer encoder, i.e., the global feature vector sequence of the sleep phase sequence, is denoted as... , For the first All feature vectors, The global feature vector representing the output position of the CLS marker; 4) Input the global feature vector sequence obtained in step 3) into the linear layer to predict sleep disorder labels, as follows: ; In the formula, This represents the probability of predicting a sleep disorder label corresponding to a subject's sleep stage sequence. and These represent the weight matrix and parameter representation of the linear layer in a hierarchical Transformer network.

[0031] In this example, the calibrated overnight long sleep stage sequence, containing 948 tags, is divided into 11 short sleep stage sequences of length 90, with zeros padded at the end if necessary, denoted as . The first Short sleep stage sequence Input into SleepGPT with freeze parameters, and extract the output of its last time step as the first time step. Local feature vectors 11 local feature vector sequences The input is fed into a Transformer encoder. This encoder captures the long-range dependencies between these local features and outputs a global feature vector sequence V. A representative vector from the global feature sequence is then selected. The input is fed into a linear layer, and the probability of a sleep disorder label is predicted using the softmax function.

[0032] The above embodiments are preferred embodiments of the present invention, but the embodiments of the present invention are not limited to the above embodiments. Any changes, modifications, substitutions, combinations, or simplifications made without departing from the spirit and principle of the present invention shall be considered equivalent substitutions and shall be included within the protection scope of the present invention.

Claims

1. A sleep diagnosis system based on a sleep language model, characterized in that, include: The data loading module is used to acquire and load multi-channel EEG data and sleep stage category labels from the PSG file. The sleep stage category labels include Wake, N1, N2, N3, REM, MOVE, and UNK, where Wake represents the wakefulness stage, N1, N2, and N3 represent stages I, II, and III of non-rapid eye movement sleep (NREM), respectively, REM represents the rapid eye movement sleep stage, MOVE represents motion artifacts, and UNK represents the unknown stage. The data preprocessing module is used to preprocess the multi-channel EEG data acquired and loaded by the data loading module, remove the EEG data labeled MOVE and UNK and their corresponding EEG data, filter out high-quality data segments, and output EEG segments that are divided according to a fixed duration and noise removed, i.e., epochs. The sleep stage segmentation module uses a pre-trained deep learning model to perform preliminary segmentation of the epochs output by the data preprocessing module, outputs the probability distribution sequence of different sleep stages corresponding to each epoch, and records the corresponding sleep stage category label predicted for each epoch, generating a preliminary segmentation of the sleep stage sequence. The sleep stage segmentation calibration module uses a pre-trained sleep language model to receive the probability distribution vector corresponding to different sleep stages for each epoch from the deep learning model. Based on this probability distribution vector, it predicts the probability distribution vector of the sleep stage category label corresponding to the current epoch, and dynamically weights it with the probability distribution vector of the sleep stage category label corresponding to the current epoch output by the deep learning model to generate a calibrated sleep stage sequence after initially segmenting the sleep stage sequence. The sleep language model is called SleepGPT, which adopts a lightweight GPT-2 architecture, including a token embedding layer, a position encoding layer, a three-layer Transformer decoder, and a linear layer. The token embedding layer maps the discrete sleep stage category labels (tokens) in the sleep stage sequence to a fixed-dimensional continuous vector representation. The position encoding layer maps the position index of the sleep stage sequence to a fixed-dimensional position vector, providing the model with the temporal information of the sleep stage in the sequence. The sequence position information enables the model to understand the temporal dependence and periodicity of sleep stages. The Transformer decoder includes a multi-head self-attention layer, a feedforward neural network, and a residual connection and normalization layer. The multi-head self-attention layer processes the input sequence in parallel through five independent attention heads and applies a causal attention mechanism in the time dimension, enabling the model to simultaneously pay attention to the contextual features of the input sequence. This ensures that the prediction at each time step depends only on the current and previous sequence information, preventing the leakage of future information. The feedforward neural network is a two-layer fully connected network that performs nonlinear transformations and feature enhancements on the output of the multi-head self-attention layer, learning the complex feature combinations and transformation rules in the sleep stage sequence. The residual connection and normalization layer maintains the original information flow through residual connections and stabilizes the feature distribution through layer normalization. The linear layer maps the output vector of the Transformer decoder to a probability distribution sequence of the predicted categories of the five sleep stages: Wake, N1, N2, N3, and REM. The sleep disorder diagnosis module predicts sleep disorder labels for subjects' long sleep stage sequences throughout the night based on a hierarchical Transformer network. The hierarchical Transformer network includes a pre-trained SleepGPT, a Transformer encoder, and a linear layer. The pre-trained SleepGPT acts as a feature extractor, extracting local contextual features from short sleep stage sequences. The Transformer encoder captures global contextual features from the output of SleepGPT, outputting a global feature vector sequence. Finally, the output global feature vector sequence is input into the linear layer to predict sleep disorder labels.

2. The sleep diagnosis system based on a sleep language model according to claim 1, characterized in that, The data loading module includes a data source acquisition unit and an EEG reading unit, wherein: The data source acquisition unit is used to acquire PSG files and corresponding annotation files from local or wearable devices. The PSG file contains multi-channel physiological signals, including electroencephalogram (EEG), electrooculogram (EOG), and electromyogram (EMG). The EEG includes head information and data recordings. The head information includes the date and time of recording, the recording duration, the electrodes used and their locations, labels for each channel, and other recording device and patient information. The data recordings include physiological signal data for each channel acquired at a predetermined sampling frequency. The annotation file contains sleep stage category labels and records the start time and duration of each sleep stage. The EEG reading unit reads the EEG of a specified channel from the PSG file. The EEG is initially a one-dimensional array with a length equal to the total number of samples, where the total number of samples is determined by the file duration × sampling frequency.

3. The sleep diagnosis system based on a sleep language model according to claim 2, characterized in that, The data preprocessing module performs a series of preprocessing operations on the EEG from the specified channel read by the EEG reading unit in the data loading module, resulting in EEG segments divided into fixed-duration segments and noise-removed segments, forming a set of epochs, and saving them in NPZ format. The series of preprocessing operations includes dividing the EEG into segments of fixed duration, removing invalid labels, and parsing the sleep stage category labels to correspond with the segmented signal data. Dividing the EEG into segments of fixed duration means dividing the EEG into epochs of fixed duration, where fixed duration is the length of one epoch. Removing invalid labels means removing the labels MOVE and UNK and their corresponding EEG data.

4. The sleep diagnosis system based on a sleep language model according to claim 3, characterized in that, The deep learning model is XSleepNet, which adopts a dual-stream network architecture to process the original waveform signal and time-frequency features simultaneously. It integrates multi-dimensional information through an adaptive fusion mechanism and finally outputs the classification result through a classification layer.

5. The sleep diagnosis system based on a sleep language model according to claim 4, characterized in that, The sleep stage calibration module performs the following operations: 1) Receive the initial sleep stage sequence output by XSleepNet, and obtain the probability distribution sequence of each epoch corresponding to different sleep stages, defined as... Each element is a 5-dimensional probability distribution vector. If the total number of epochs is represented, then the 1st epoch... element ,satisfy , This represents the probability of an epoch corresponding to the awakening stage label. These represent the probabilities of the epoch corresponding to stage I, II, and III of the non-rapid eye movement sleep (NREM) sleep pattern, respectively. This represents the probability of an epoch corresponding to a REM sleep stage category label; 2) Change the sequence Save it as a two-dimensional array as input to SleepGPT, with shape as follows: , where 5 represents the predicted probability of the 5 sleep stage category labels corresponding to each epoch; 3) Use a fixed length of The context window, in steps In sequence Slide up; for the first... A 5-dimensional probability distribution vector ,in , cut The former There are 5-dimensional probability distribution vectors, that is, starting from the (t-k+1)th 5-dimensional probability distribution vector. Begin, take A continuous 5-dimensional probability distribution vector, denoted as ,in Reference The probability distribution subsequence is, The above information indicates that, as input to the model, the model will... Predict the sleep stage category label corresponding to the t-th epoch; 4) The input vector constructed in step 3) is processed through a token embedding layer and a position encoding layer. Mapped to a higher-dimensional space, and for Inject positional information into each of the 5-dimensional probability distribution vectors: ; In the formula, This represents the embedding vector output after processing through the token embedding layer and the position encoding layer. The token embedding matrix, which is the weight matrix of the token embedding layer, is then compared with... The dot product maps each sleep stage category to a unique 48-dimensional vector representation, forming a learnable embedding matrix with a shape of 5×48. The matrix contains five sleep stage category labels: Wake, N1, N2, N3, and REM. Each sleep stage category label corresponds to a unique vector representation. The position embedding matrix maps the position indices of the sleep stage sequence, i.e., integers from 0 to 89, to a 48-dimensional position vector matrix with a shape of 90×48. This provides the model with temporal position information of the sleep stage in the sequence, enabling the model to understand the temporal dependence and periodicity of the sleep stage. 5) The vector output in step 4) The input is processed by a 3-layer Transformer decoder. The processing procedure is as follows: ; In the formula, Indicates Transformer decoder, This represents the vector output by the l-th layer Transformer decoder. Indicates the first The vector output by the layer Transformer decoder, when hour for ; 6) Input the output vector obtained in step 5) into the linear layer to output the sleep stage prediction label corresponding to the t-th epoch. probability distribution , represented as: ; In the formula, SLM represents SleepGPT. This represents the predicted label sequence of sleep stages corresponding to the first t-1 epochs. Represents a linear layer. This represents the weight matrix of the linear layer in SleepGPT; 7) Calculate the probability distribution of XSleepNet for different sleep stage category labels corresponding to the current epoch. and Through weighting factors The probabilities of different sleep stages are then merged to obtain the final sleep stage probabilities. predict: ; In the formula, SSM represents XSleepNet, and the weighting factor is... By controlling the relative influence of the two models XSleepNet and SleepGPT, the sleep stage category label with the highest prediction probability is finally selected as the prediction result, and the calibrated sleep stage sequence is output.

6. The sleep diagnosis system based on a sleep language model according to claim 5, characterized in that, The sleep disorder diagnosis module performs the following operations: 1) The subject's overnight long sleep stage sequence was segmented into short sleep stage sequences, and these sequences were padded into a batch to achieve a uniform length, represented as: ; In the formula, U represents the sequence of long sleep stages throughout the night for the subject. This represents the segmentation of the long sleep stage sequence into its first segment. Short sleep stage sequence, Indicates the number of segments in the long sleep phase sequence; 2) Using the pre-trained SleepGPT from the sleep stage staging calibration module, local contextual features are extracted from short sleep stage sequences, represented as: ; In the formula, The first generation generated for SleepGPT There are local feature vectors, and the sequence of all local feature vectors is . The input sequence that constitutes the Transformer encoder in a hierarchical Transformer network, where Corresponding to the CLS tag, in the Transformer architecture, the CLS tag is a special tag added to the beginning of the input sequence. Its final output can serve as a global representation of the entire sequence, which is suitable for classification tasks. 3) Use the Transformer encoder to capture global contextual features from the output of SleepGPT, represented as: ; In the formula, Represents the Transformer encoder. The output sequence of the Transformer encoder, i.e., the global feature vector sequence of the sleep phase sequence, is denoted as... , For the first All feature vectors, The global feature vector representing the output position of the CLS marker; 4) Input the global feature vector sequence obtained in step 3) into the linear layer to predict sleep disorder labels, as follows: ; In the formula, This represents the probability of predicting a sleep disorder label corresponding to a subject's sleep stage sequence. and These represent the weight matrix and parameter representation of the linear layer in a hierarchical Transformer network.

Citation Information

Patent Citations

  • Sleep staging-oriented electroencephalogram interpretability analysis method and related equipment

    CN115137374A

  • Self-attention sleep system and method based on polymorphic sleep data fusion

    CN119792762A

  • Sleep profiling system with feature generation and auto-mapping

    WO2016089313A1