Sleep decoding method and system based on generative pre-training transformer
Patent Information
- Application Number
- CN202610949777.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-06-29
- Publication Date
- 2026-09-25
AI Technical Summary
[0005]本发明的目的在于提出一种基于生成式预训练Transformer的睡眠解码方法及系统,以解决现有睡眠解码模型通道配置固定、任务泛化能力弱、时频信息割裂的技术问题,达到提升睡眠解码准确性、鲁棒性和临床适应性的效果
[0025]1. 本发明通过对多导睡眠监测信号进行标准化,并结合通道存在信息和掩码信息生成联合嵌入,使不同数据集、不同设备和不同通道配置下的PSG信号能够在统一通道框架下参与建模,降低固定通道配置对睡眠解码模型的限制,提高跨设备和跨数据集处理能力。
Smart Images

Figure CN122818010A_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of sleep medicine signal processing technology, specifically relating to a sleep decoding method and system based on generative pre-trained Transformer. Background Technology
[0002] Sleep is a vital physiological process for maintaining human health, significantly impacting cognitive regulation, emotional stability, and the prevention and control of various chronic diseases. With the accelerated pace of modern life, the incidence of sleep disorders is rising, and their association with neurological diseases such as Alzheimer's, epilepsy, and Parkinson's has attracted widespread attention. Therefore, developing accurate and efficient sleep state identification and analysis technologies is of great significance for early disease screening, individualized intervention strategy development, and brain function research.
[0003] In clinical practice, polysomnography (PSG) is widely used to record and analyze multimodal signals, including electroencephalography (EEG), eye movement (EMG), electromyography (EMG), respiration, heart rate, and blood oxygenation. Traditional sleep data analysis relies on manual annotation, which is time-consuming and highly subjective, limiting its application in large-scale populations and remote monitoring. In recent years, artificial intelligence, especially deep neural networks, has developed rapidly in sleep decoding tasks, enabling automatic completion of tasks such as sleep staging, event detection, and pathological identification, thus improving analysis efficiency and consistency. However, existing methods still face the following key challenges: (1) over-reliance on manually labeled data, which restricts the model's generalization ability; (2) dependence on fixed channel configurations, making it difficult for the model to adapt to real-world scenarios with multiple devices and datasets; and (3) failure to effectively integrate time-domain and frequency-domain information, which limits the model's decoding performance in multi-task environments.
[0004] To address the aforementioned issues, it is urgent to construct a unified model framework with universality, adaptability, and cross-domain modeling capabilities. As a new generation of AI learning paradigm, basic models have achieved significant results in fields such as image processing, language processing, and biomedicine. Their core advantage lies in obtaining general feature representation capabilities through self-supervised pre-training on large-scale unlabeled data, and achieving transfer and adaptation in various downstream tasks through minor fine-tuning. However, basic model research in the field of sleep is still in its early stages, mainly facing the following challenges: (1) lack of self-supervised pre-training methods adapted to PSG signals; (2) modeling difficulties caused by channel heterogeneity; and (3) immature deep fusion mechanisms of time-domain and frequency-domain information. Therefore, it is urgent to propose a novel generative pre-training model with multi-task expansion capabilities, channel structure adaptation capabilities, and time-frequency feature joint modeling capabilities to support the large-scale promotion and application of sleep decoding in clinical and scientific research. Summary of the Invention
[0005] The purpose of this invention is to propose a sleep decoding method and system based on generative pre-trained Transformer to solve the technical problems of fixed channel configuration, weak task generalization ability, and fragmented time-frequency information in existing sleep decoding models, thereby improving the accuracy, robustness and clinical adaptability of sleep decoding.
[0006] To achieve the above objectives, the present invention adopts the following technical solution.
[0007] A sleep decoding method based on generative pre-trained Transformer includes the following steps: The polysomnography signal was standardized to obtain a standardized multichannel sleep monitoring signal. Based on the standardized multi-channel sleep monitoring signal, a time-domain signal and a frequency-domain representation are generated; Block-marked embedding is performed on the time-domain signal and the frequency-domain representation respectively, and a joint embedding is generated based on channel presence information and mask information; The joint embedding is input into a unified time-frequency Transformer backbone network, and sleep feature representation is obtained through masked attention and domain-aware mapping; Based on the sleep feature representation, a multi-agent pre-training task is performed to train the unified time-frequency Transformer backbone network and obtain a pre-trained sleep decoding model. The multi-agent pre-training task includes a cross-domain contrastive learning task, a difficult negative sample matching task, and a mask reconstruction task. The pre-trained sleep decoding model is fine-tuned for downstream tasks based on labeled sleep monitoring data to obtain the target sleep decoding model. The sleep monitoring signal to be decoded is input into the target sleep decoding model to obtain the sleep decoding result.
[0008] Furthermore, the polysomnography signal is standardized to obtain a standardized multi-channel sleep monitoring signal, including: The format of the polysomnography signal is standardized and the sampling rate is aligned to obtain the preprocessed sleep monitoring signal. The preprocessed sleep monitoring signal is channel-mapped according to the standard channel set to obtain the channel-mapped signal; The presence information and missing channel markers are determined based on the channel mapping signal. The channel mapping signal is normalized to obtain a standardized multi-channel sleep monitoring signal.
[0009] Further, the preprocessed sleep monitoring signal is channel-mapped according to a standard channel set to obtain a channel-mapped signal, including: The original channels in the preprocessed sleep monitoring signal are matched with the standard channel set to obtain the channel matching result; Based on the channel matching results, the original channels in the preprocessed sleep monitoring signal are preserved, mapped, or fused to obtain a channel-mapped signal.
[0010] Further, based on the standardized multi-channel sleep monitoring signal, a time-domain signal and a frequency-domain representation are generated, including: The standardized multi-channel sleep monitoring signal is divided into sleep segments to obtain a time-domain signal; Perform time-frequency transformation on the time-domain signal to obtain an initial frequency-domain representation; The initial frequency domain representation is numerically stabilized to obtain the frequency domain representation.
[0011] Further, the time-domain signal is subjected to time-frequency transformation to obtain an initial frequency-domain representation, including: The time-domain signal is windowed according to the window function configuration and window overlap configuration to obtain a windowed time-domain signal; The windowed time-domain signal is subjected to spectral transformation to obtain an initial frequency-domain representation.
[0012] Further, block-marking embedding is performed on the time-domain signal and the frequency-domain representation respectively, and a joint embedding is generated based on channel presence information and mask information, including: Channel-independent convolution and embedding mapping are performed on the time-domain signal to obtain time-domain block labels; Channel-independent convolution and embedding mapping are performed on the frequency domain representation to obtain frequency domain block labels; The time domain block marker is fused with the classification marker, position embedding, channel embedding and block type embedding to obtain the time domain embedding, and the frequency domain block marker is fused with the classification marker, position embedding, channel embedding and block type embedding to obtain the frequency domain embedding; Based on the channel presence information and mask information, the time domain embedding and the frequency domain embedding are marked and constrained to obtain the restricted dual-domain embedding; The restricted dual-domain embedding is aligned by the number of blocks and then concatenated to obtain a joint embedding.
[0013] Furthermore, the restricted dual-domain embedding is aligned by the number of blocks and then concatenated to obtain a joint embedding, including: The block resolution configuration is determined based on the number of time domain blocks and the number of frequency domain blocks. The restricted two-domain embedding is reconstructed according to the block resolution configuration to obtain an aligned two-domain embedding; By concatenating the time-domain embedding and the frequency-domain embedding in the aligned dual-domain embedding, a joint embedding is obtained.
[0014] Furthermore, the jointly embedded input to the unified time-frequency Transformer backbone network is used to obtain a sleep feature representation through masked attention and domain-aware mapping, including: Layer normalization and masked attention calculation are performed on the joint embedding to obtain attention features; Based on the signal domain to which the attention feature belongs, the attention feature is input into the domain perception mapping path to obtain the domain perception feature; The domain-aware features are fused with the corresponding input features to obtain hierarchical time-frequency features; The time-frequency features of the aforementioned levels are passed and converged layer by layer to obtain a sleep feature representation.
[0015] Further, based on the signal domain to which the attention features belong, the attention features are input into the domain-aware mapping path to obtain domain-aware features, including: The temporal features in the attention features are input into the temporal perceptron to obtain the temporal mapping features; The frequency domain features in the attention features are input into the frequency domain perceptron to obtain the frequency domain mapping features; The time-domain mapping feature and the frequency-domain mapping feature are fused to obtain the intra-domain mapping feature; The domain-mapping features are input into a unified perceptron to obtain domain-aware features.
[0016] Further, based on the sleep feature representation, a multi-agent pre-training task is performed to train the unified time-frequency Transformer backbone network, resulting in a pre-trained sleep decoding model, including: Based on the sleep feature representation, determine the time domain representation and frequency domain representation of the same sleep segment; Cross-domain matching is performed on the time-domain representation and the frequency-domain representation to obtain the cross-domain contrast loss; Based on the cross-domain matching results, difficult negative sample pairs are determined, and the difficult negative sample pairs are matched and discriminated to obtain the difficult negative sample loss; The masked block is reconstructed based on the mask information, and the mask reconstruction loss is obtained. The unified time-frequency Transformer backbone network is trained based on the cross-domain contrast loss, the hard negative sample loss, and the mask reconstruction loss to obtain a pre-trained sleep decoding model.
[0017] Further, cross-domain matching is performed on the time-domain representation and the frequency-domain representation to obtain a cross-domain contrast loss, including: Extract the global representation of the time domain from the time domain representation and the global representation of the frequency domain from the frequency domain representation, respectively; Bidirectional cross-domain similarity is determined based on the time-domain global representation and the frequency-domain global representation; The cross-domain comparison loss is determined based on the bidirectional cross-domain similarity.
[0018] Further, based on the cross-domain matching results, difficult negative sample pairs are determined, and matching and discrimination are performed on the difficult negative sample pairs to obtain the difficult negative sample loss, including: Determine the cross-domain matching probability based on the cross-domain matching results; Eliminate correct matching pairs from the cross-domain matching probabilities to obtain candidate negative sample pairs; Difficult negative sample pairs are determined based on the matching probability of the candidate negative sample pairs; The difficult negative sample pairs are matched and discriminated to obtain the difficult negative sample loss.
[0019] Furthermore, the masked blocks are reconstructed based on the mask information to obtain the mask reconstruction loss, including: The masked and unmasked blocks in the joint embedding are determined based on the mask information; A reconstructed representation of the occluded block is generated based on the unoccluded block; The mask reconstruction loss is determined based on the reconstructed representation and the corresponding true representation.
[0020] Furthermore, the pre-trained sleep decoding model is fine-tuned for downstream tasks based on labeled sleep monitoring data to obtain the target sleep decoding model, including: Based on the downstream task type, a task decoder is configured after the pre-trained sleep decoding model to obtain the sleep decoding model to be fine-tuned. The labeled sleep monitoring data is input into the sleep decoding model to be fine-tuned to obtain the task prediction result; The sleep decoding model to be fine-tuned is trained based on the task prediction results and corresponding annotation information to obtain the target sleep decoding model.
[0021] Further, the sleep monitoring signal to be decoded is input into the target sleep decoding model to obtain the sleep decoding result, including: The sleep monitoring signal to be decoded is standardized to be consistent with the standardized multi-channel sleep monitoring signal to obtain the standardized signal to be decoded. The standardized signal to be decoded is input into the target sleep decoding model to obtain the sleep features to be decoded; The sleep features to be decoded are decoded by the task decoder in the target sleep decoding model to obtain the sleep decoding result.
[0022] Further, the sleep monitoring signal to be decoded is input into the target sleep decoding model to obtain the sleep decoding result. The method also includes performing channel-stage analysis on the sleep decoding result, the channel-stage analysis including: Based on the channel embedding of the target sleep decoding model and the sleep decoding results, the channel contribution information is determined; Based on the channel contribution information and the sleep stage information in the sleep decoding results, a channel-stage contribution relationship is generated.
[0023] A sleep decoding system based on generative pre-trained Transformer, comprising: An input embedding module is used to standardize the polysomnography signal to obtain a standardized multi-channel sleep monitoring signal; based on the standardized multi-channel sleep monitoring signal, a time-domain signal and a frequency-domain representation are generated, and block-marked embedding is performed on the time-domain signal and the frequency-domain representation respectively, and a joint embedding is generated based on channel presence information and mask information; The backbone network module is used to embed the joint input into a unified time-frequency Transformer backbone network and obtain a sleep feature representation through masked attention and domain-aware mapping; The model training module is used to perform a multi-agent pre-training task based on the sleep feature representation, train the unified time-frequency Transformer backbone network, and obtain a pre-trained sleep decoding model. The multi-agent pre-training task includes a cross-domain contrastive learning task, a difficult negative sample matching task, and a mask reconstruction task. The pre-trained sleep decoding model is then fine-tuned by downstream tasks based on labeled sleep monitoring data to obtain a target sleep decoding model. The multi-task output module is used to input the sleep monitoring signal to be decoded into the target sleep decoding model to obtain the sleep decoding result.
[0024] The present invention has achieved the following beneficial effects.
[0025] 1. This invention standardizes polysomnography (PSG) signals and generates joint embeddings by combining channel presence information and mask information. This enables PSG signals from different datasets, devices, and channel configurations to participate in modeling under a unified channel framework, reducing the limitations of fixed channel configurations on sleep decoding models and improving cross-device and cross-dataset processing capabilities.
[0026] 2. This invention generates time-domain signals and frequency-domain representations, and then performs block-marking embedding on both to generate joint embeddings. This allows the original waveform information and spectral dynamic information to participate in subsequent modeling in a unified data structure, reducing the insufficient utilization of features caused by the separation of time-domain and frequency-domain information.
[0027] 3. This invention uses a unified time-frequency Transformer backbone network to perform masked attention and domain-aware mapping on the joint embedding, enabling the model to perform hierarchical feature extraction between the time domain, frequency domain, and cross-domain joint representation, thereby enhancing the modeling ability of local changes, global patterns, and cross-domain complementary relationships of sleep signals.
[0028] 4. This invention trains a unified time-frequency Transformer backbone network through cross-domain contrastive learning tasks, difficult negative sample matching tasks, and mask reconstruction tasks, enabling the model to learn sleep feature representations using unlabeled PSG data, reducing dependence on large-scale manually labeled data, and improving the model's generalization ability under conditions of limited data and task transfer.
[0029] 5. This invention enables the target sleep decoding model to adapt to tasks such as sleep stage identification, pathological state classification, key sleep event detection, signal generation or signal reconstruction by fine-tuning the pre-trained sleep decoding model for downstream tasks, thereby enhancing the task scalability and practical deployment adaptability of the unified modeling framework.
[0030] 6. This invention performs channel-stage analysis on sleep decoding results, generates channel-stage contribution relationships based on channel embedding and sleep stage information, and enables the model output to correlate the feature contributions of different channels and different sleep stages, providing an interpretable technical basis for sleep decoding results. Attached Figure Description
[0031] Figure 1 This is a flowchart of a sleep decoding method based on generative pre-trained Transformer in an embodiment; Figure 2 This is a schematic diagram of the block tag embedding and joint embedding generation process of time-domain signal and frequency-domain representation in the embodiment; Figure 3 This is a schematic diagram of the unified time-frequency Transformer backbone network in the embodiment; Figure 4 This is a schematic diagram of the pre-training and downstream task fine-tuning process of the sleep decoding model in the embodiment. Detailed Implementation
[0032] To make the various technical features, advantages, or effects of the present invention more apparent and understandable, detailed descriptions are provided below through embodiments.
[0033] This invention provides a sleep decoding method based on generative pre-trained Transformer, such as... Figure 1As shown, this method targets polysomnography (PSG) signals. The PSG signals can originate from multiple publicly available PSG datasets, covering different subject groups, acquisition devices, and channel configurations, and can serve as pre-training, fine-tuning, and testing datasets for the model. The method employs a unified channel format and a unified time-frequency Transformer backbone network (also known as SleepGPT) to jointly model the time-domain signal and frequency-domain representation. This unified time-frequency Transformer backbone network includes an input embedding process, a backbone network feature extraction process, and a task decoding process, and adopts a two-stage training approach combining a pre-training stage and a downstream task fine-tuning stage.
[0034] Step S1: Standardize the polysomnography signal to obtain a standardized multichannel sleep monitoring signal.
[0035] Specifically, the standardization is used to unify the data format, sampling rate, units, and channel settings across different datasets. The standardization process may include format unification, sampling rate alignment, channel mapping, missing channel marking, artifact or outlier data marking, and normalization.
[0036] In an optional embodiment of the present invention, step S1 may include: Step S11: Standardize the format and align the sampling rate of the polysomnography signal to obtain a preprocessed sleep monitoring signal.
[0037] Specifically, the polysomnography (PSG) signals can come from multiple publicly available PSG datasets. By unifying the format and aligning the sampling rates, PSG signals from different sources can form a consistent data format and time sampling basis.
[0038] Step S12: Perform channel mapping on the preprocessed sleep monitoring signal according to the standard channel set to obtain the channel mapping signal.
[0039] Specifically, the standard channel set can be configured to use a unified standard channel input format according to the dataset or application scenario, preferably a five-channel input format.
[0040] In an optional embodiment of the present invention, step S12 may include: Step S121: Match the original channels in the preprocessed sleep monitoring signal with the standard channel set to obtain the channel matching result.
[0041] Specifically, the original channel can be a lead channel acquired from different PSG datasets, and the channel matching result is used to represent the correspondence between the original channel and the standard channel set.
[0042] Step S122: Based on the channel matching result, the original channels in the preprocessed sleep monitoring signal are preserved, mapped, or fused to obtain a channel-mapped signal.
[0043] Specifically, leads that exist in the original data but are not selected into the standard channel set can be discarded, mapped, or fused; leads that can match the standard channel set can be retained or mapped to the corresponding standard channel.
[0044] Step S13: Determine the channel presence information and missing channel markers based on the channel mapping signal.
[0045] Specifically, for leads missing in the standard channel set, a corresponding missing channel marker or channel presence mask is generated; the channel presence information is used to indicate whether each standard channel in the current sleep segment is available.
[0046] Step S14: Normalize the channel mapping signal to obtain a standardized multi-channel sleep monitoring signal.
[0047] Specifically, the normalization can be achieved by standardizing the channel mapping signal using the global mean and standard deviation of the training dataset.
[0048] Step S2: Generate a time-domain signal and a frequency-domain representation based on the standardized multi-channel sleep monitoring signal.
[0049] Specifically, the time-domain signal is used to represent the data of the standardized multi-channel sleep monitoring signal in the original waveform dimension, and the frequency-domain signal is used to represent the spectrum or time-spectrum data obtained by converting the time-domain signal. For signals containing... The training dataset for *n* sleep segments (epochs) can be represented as: , where each sleep segment This represents a PSG signal. This indicates the length of the sampling point in each channel within the sleep segment. It is determined by the product of the sampling rate and the duration of the sleep segment. For a multi-channel PSG signal, the time-domain input of a sleep segment can be expressed as: ,in, Indicates the number of channels; Indicates the time sampling length.
[0050] In an optional embodiment of the present invention, step S2 may include: Step S21: Divide the standardized multi-channel sleep monitoring signal into sleep segments to obtain a time-domain signal.
[0051] Specifically, each PSG record is divided into several sleep segments of fixed duration, preferably each sleep segment lasting 30 seconds.
[0052] Step S22: Perform time-frequency transformation on the time-domain signal to obtain an initial frequency domain representation.
[0053] Specifically, for the time-domain signal of each lead, a two-dimensional time-spectrum diagram can be generated using short-time Fourier transform.
[0054] In an optional embodiment of the present invention, step S22 may include: Step S221: Window the time domain signal according to the window function configuration and window overlap configuration to obtain a windowed time domain signal.
[0055] Specifically, the window function configuration can adopt a Hamming window, and the window overlap configuration can adopt a window function overlap rate of 50.0%.
[0056] Step S222: Perform spectral transformation on the windowed time domain signal to obtain an initial frequency domain representation.
[0057] Specifically, the spectrum transformation can be a short-time Fourier transform, and the number of FFT points can be 256.
[0058] Step S23: Numerically stabilize the initial frequency domain representation to obtain the frequency domain representation.
[0059] Specifically, the amplitude spectrum can be numerically stabilized by taking the logarithm, resulting in a frequency domain representation. Therefore, the model's input can form a set of two-domain data pairs. ,in, Indicates the first The time-domain signal of a sleep segment; Indicates the first Frequency domain representation or time-spectrum representation of each sleep segment.
[0060] Step S3: Block tag embedding is performed on the time domain signal and the frequency domain representation respectively, and joint embedding is generated based on channel presence information and mask information.
[0061] Specifically, see Figure 2 The time-domain signal and frequency-domain representation are mapped to high-dimensional block labels respectively, and the two types of block labels are combined into a unified joint embedding, which serves as the input to the unified time-frequency Transformer backbone network.
[0062] In an optional embodiment of the present invention, step S3 may include: Step S31: Perform channel-independent convolution and embedding mapping on the time-domain signal to obtain time-domain block labels.
[0063] Specifically, for one-dimensional time-domain signals By slicing and encoding it through channel-independent 1D convolution, we obtain... Block marker ,in, Indicates the signal length; Indicates the resolution of the time domain block; Indicates the number of time-domain blocks; Indicates the number of channels; express Embedded dimension.
[0064] Step S32: Perform channel-independent convolution and embedding mapping on the frequency domain representation to obtain frequency domain block labels.
[0065] Specifically, for two-dimensional frequency domain signals It is divided into channels by channel-independent 2D convolution. Block marker ,in, Indicates the frequency domain block resolution; Indicates the duration of a sleep segment; This represents the number of frequency points obtained by the FFT; Indicates the number of frequency domain blocks; Indicates the number of blocks in the frequency domain block set; Indicates the embedding dimension.
[0066] Step S33: The time domain block marker is fused with the classification marker, position embedding, channel embedding and block type embedding to obtain the time domain embedding, and the frequency domain block marker is fused with the classification marker, position embedding, channel embedding and block type embedding to obtain the frequency domain embedding.
[0067] Specifically, the time domain blocks are flattened in the first two dimensions, and special classification tags for representing global time domain information are added at the beginning of the sequence. Composed of time-domain embedded representation Learnable one-dimensional position embedding, channel embedding, and block type embedding are added to the time-domain block label, and the sum of these three with the block label constitutes the final time-domain input representation.
[0068] After flattening the frequency domain block labels, a classification label representing global frequency domain information is added at the beginning of the sequence. It is then added to the learnable position embedding, channel embedding, and block type embedding to form the frequency domain input representation. ,in, This indicates the number of frequency domain blocks in the frequency domain input representation.
[0069] Step S34: Mark and constrain the time domain embedding and the frequency domain embedding according to the channel existence information and the mask information to obtain the restricted dual-domain embedding.
[0070] Specifically, the masking information may include a channel presence mask, a data missing block mask, and a random block mask. The channel presence mask is used to identify whether each standard channel exists in the current sleep segment; the data missing block mask is used to identify blocks that are unusable due to sensor detachment, signal artifacts, acquisition failure, or insufficient data quality; the random block mask is used in the mask reconstruction task during the pre-training phase to randomly mask some blocks in the time domain block markers and frequency domain block markers.
[0071] Step S35: Align the number of blocks in the restricted dual-domain embedding and splice them together to obtain the joint embedding.
[0072] Specifically, the time-domain PSG signal and the frequency-domain PSG signal are respectively convolved through channels to form block labels. Some data missing blocks are marked as missing blocks, and some blocks are marked as masked blocks in pre-training. Then, the time-domain embedding and the frequency-domain embedding are connected to form a joint embedding.
[0073] In an optional embodiment of the present invention, step S35 may include: Step S351: Determine the block resolution configuration based on the number of time domain blocks and the number of frequency domain blocks.
[0074] Specifically, by setting the time domain block resolution and frequency domain block resolution This allows the number of time domain blocks to match the number of frequency domain blocks, and the number of blocks in the frequency direction can be further set as a constant.
[0075] Step S352: Reconstruct the restricted dual-domain embedding according to the block resolution configuration to obtain an aligned dual-domain embedding.
[0076] Specifically, the aligned dual-domain embedding includes time-domain embedding and frequency-domain embedding after block number alignment.
[0077] Step S353: The time domain embedding and frequency domain embedding in the aligned dual-domain embedding are concatenated to obtain the joint embedding.
[0078] Specifically, after alignment, the time-domain embedding and frequency-domain embedding are concatenated along the sequence dimension to form a joint embedding. , as the input to the 0th layer of the UTF Transformer, where, Indicates joint embedding; Indicates temporal embedding; Indicates frequency domain embedding; Indicates the number of channels; Indicates the number of blocks after synchronization, and equal or ; Indicates the embedding dimension.
[0079] Step S4: The joint embedding is input into the unified time-frequency Transformer backbone network, and sleep feature representation is obtained through masked attention and domain-aware mapping.
[0080] Specifically, see Figure 3 The unified time-frequency Transformer backbone network adopts a UTF Transformer structure and performs multi-layer feature extraction on the joint embedding. The UTF Transformer, through... A masked multi-head self-attention (MSA) mechanism is introduced into a stacked Transformer block to achieve unified modeling and interactive alignment of dual-domain signals.
[0081] In an optional embodiment of the present invention, step S4 may include: Step S41: Perform layer normalization and mask attention calculation on the joint embedding to obtain attention features.
[0082] Specifically, for the first The input to the layer UTF Transformer block is the previous layer, i.e., the first layer. Layer output The computation process of the mask multi-head self-attention module can be formally represented as: in, Indicates the first Attention characteristics of layers; This indicates a multi-head self-attention mask; Representation layer normalization; Indicates the first The output of the layer. By introducing masking operations to restrict information interaction between specific blocks, the masking operations can correspond to channels missing, signal artifacts, data missing blocks, or random masking blocks.
[0083] Step S42: Based on the signal domain to which the attention feature belongs, input the attention feature into the domain perception mapping path to obtain the domain perception feature.
[0084] Specifically, the UTF Transformer introduces three types of switched perceptrons in each block: the Time-domain Perceptron (T-PCN), the Frequency-domain Perceptron (F-PCN), and the Unified Perceptron (U-PCN). The T-PCN, F-PCN, and U-PCN are used to map embedded features from different domains to their respective semantic spaces.
[0085] In an optional embodiment of the present invention, step S42 may include: Step S421: Input the temporal features in the attention features into the temporal perceptron to obtain the temporal mapping features.
[0086] Step S422: Input the frequency domain features in the attention features into the frequency domain perceptron to obtain the frequency domain mapping features.
[0087] Step S423: The time domain mapping feature and the frequency domain mapping feature are fused to obtain the intra-domain mapping feature.
[0088] Step S424: Input the domain-mapping features into the unified perceptron to obtain domain-aware features.
[0089] In an optional embodiment of the present invention, the UTF Transformer... The output of the layer block is defined as follows: in, Indicates the first Layer output; Represents a time-domain perceptron; This represents a frequency domain perceptron; Represents a unified perceptron; Representation layer normalization; Indicates the first Attention characteristics of layers; and These represent the time-domain feature embedding and frequency-domain feature embedding from the masked MSA output, respectively.
[0090] Step S43: The domain-aware features are fused with the corresponding input features to obtain hierarchical time-frequency features.
[0091] Specifically, in the front of the UTF Transformer network In the lower-level blocks, T-PCN and F-PCN independently process the time-domain embedding and frequency-domain embedding of the mask MSA output, respectively; in the higher-level structure, especially at the top... In each UTF Transformer block, U-PCN is introduced to process the intermediate representation from the masked multi-head self-attention mechanism. Depending on the characteristics and modeling requirements of the downstream task, the high-level structure can adopt two processing methods: one is to input the representation output by the masked attention into T-PCN and F-PCN for independent modeling; the other is to directly input it into U-PCN for cross-domain information fusion.
[0092] Step S44: The hierarchical time-frequency features are passed and converged layer by layer to obtain a sleep feature representation.
[0093] Specifically, through multi-layer stacking and inter-domain modeling, a sleep fragment-level representation, Epoch embedding, is obtained. This Epoch embedding can be derived from the final layer... mark, The tags, along with pooled temporal and frequency domain block tags, are formed and used as sleep feature representations for decoding in pre-trained agent tasks and downstream tasks.
[0094] Step S5: Based on the sleep feature representation, perform a multi-agent pre-training task to train the unified time-frequency Transformer backbone network and obtain a pre-trained sleep decoding model. The multi-agent pre-training task includes a cross-domain contrastive learning task, a difficult negative sample matching task, and a mask reconstruction task.
[0095] Specifically, see Figure 4 The pre-training phase is used to train the main body of the model on large-scale unlabeled PSG data, while the downstream task fine-tuning phase is used to adapt the model to specific tasks on limited labeled data. The pre-training phase adopts a multi-agent task learning framework, integrating contrastive learning, hard negative sample mining, and masked autoencoder strategies.
[0096] In an optional embodiment of the present invention, step S5 may include: Step S51: Determine the time domain representation and frequency domain representation of the same sleep segment based on the sleep feature representation.
[0097] Specifically, in the comparative learning process, the batch size is set to... The model needs to be in The correct matching pair is identified from 10 possible time-frequency pairs. Time-domain embedding is encoded using stacked UTF Transformer blocks combined with T-PCN, while frequency-domain embedding is processed using a Transformer structure of the same depth combined with F-PCN. For data from batches... and batch The model uses the output from the final output. and The markers represent global time domain information and global frequency domain information, respectively, and the outputs are represented as follows: and .
[0098] Step S52: Perform cross-domain matching on the time domain representation and the frequency domain representation to obtain cross-domain contrast loss.
[0099] Specifically, the model constructs time-to-frequency (T2F) and frequency-to-time (F2T) comparison matrices by calculating the similarity between the time-domain and frequency-domain embeddings.
[0100] In an optional embodiment of the present invention, step S52 may include: Step S521: Extract the global representation of the time domain from the time domain representation and the global representation of the frequency domain from the frequency domain representation, respectively.
[0101] Specifically, the time-domain global representation can be derived from... The notation indicates that the global frequency domain representation can be derived from... The marker indicates.
[0102] Step S522: Determine the bidirectional cross-domain similarity based on the time-domain global representation and the frequency-domain global representation.
[0103] Specifically, the bidirectional cross-domain similarity includes time-to-frequency similarity and frequency-to-time similarity, which are defined as follows: in, Indicates the first The time-domain output of the first sleep segment and the first sleep segment Time-to-frequency similarity between the frequency domain outputs of individual sleep segments; Indicates the first The frequency domain output of the first sleep segment is related to the first Frequency-to-temporal similarity between the time-domain outputs of individual sleep segments; Representation layer normalization; Indicates the transpose operation; Indicates the first Temporal domain output of each sleep segment; Indicates the first Frequency domain output of each sleep segment; Indicates the first Frequency domain output of each sleep segment; Indicates the first Temporal domain output of each sleep segment; This represents a learnable temperature parameter.
[0104] Step S523: Determine the cross-domain comparison loss based on the bidirectional cross-domain similarity.
[0105] Specifically, the cross-domain contrast loss is used to bring matching pairs closer together within the same batch, and its definition is as follows: in, Indicates cross-domain comparison loss; Indicates batch size; Indicates the first The time-domain output of the first sleep segment and the first sleep segment Time-to-frequency similarity between the frequency domain outputs of individual sleep segments; Indicates the first The frequency domain output of the first sleep segment is related to the first Frequency-to-temporal similarity between the time-domain outputs of each sleep segment.
[0106] Step S53: Determine difficult negative sample pairs based on cross-domain matching results, and perform matching and discrimination on the difficult negative sample pairs to obtain the difficult negative sample loss.
[0107] Specifically, the difficult negative sample pairs are time-domain and frequency-domain representations of different sleep segments that have high similarity but are mismatched. These difficult negative sample pairs are re-inputted into the model and then... The labels are used for binary classification training to determine whether the time domain representation and frequency domain representation of the input match.
[0108] In an optional embodiment of the present invention, step S53 may include: Step S531: Determine the cross-domain matching probability based on the cross-domain matching results.
[0109] Specifically, the model uses the softmax function to normalize the T2F and F2T similarity matrices into probability distributions: in, Indicates the probability of a time-to-frequency match; This indicates the probability of frequency-to-time matching. Indicates time-to-frequency similarity; Indicates frequency-to-time similarity; Indicates the batch size.
[0110] Step S532: Eliminate correct matching pairs from the cross-domain matching probabilities to obtain candidate negative sample pairs.
[0111] Specifically, diagonal elements are masked to exclude correct matching pairs.
[0112] Step S533: Determine difficult negative sample pairs based on the matching probability of the candidate negative sample pairs.
[0113] Specifically, new mismatch pairs are generated from the remaining elements through random sampling to form training samples.
[0114] Step S534: Match and discriminate the difficult negative sample pairs to obtain the difficult negative sample loss.
[0115] The loss function for hard negative samples is: in, Represents the difficulty in negative sample matching loss; This indicates the number of samples in a batch; This represents the index of the sleep segment within the batch. ; Indicates the first The time domain representation of each sleep segment; Indicates the first Frequency domain representation of each sleep segment; Binary labels representing whether the time-domain representation matches the frequency-domain representation, where, Indicates a match. Indicates a mismatch; The model determines the first The time domain representation of sleep segments Frequency domain representation of the same sleep segment The probability of a match; Represented as the first Each time domain representation The selected frequency domain hard negative sample index satisfies ; Indicates and The frequency domain representation of the difficult negative sample pairs originates from a different source than the first. The first sleep segment A sleep segment; The model decision time domain representation With frequency domain difficult negative samples The probability of a mismatch; Represented as the first Each frequency domain represents The selected time-domain hard negative sample index satisfies ; Indicates and The time-domain representation of the difficult negative sample pairs originates from a different source than the first. The first sleep segment A sleep segment; This indicates that the model is determining the time-domain difficult negative samples. Frequency domain representation The probability of a mismatch; The natural logarithm can usually be used; This represents the total number of sample pairs used to calculate the loss in each batch, including... Positive sample pairs , Difficult negative sample pairs in the time-to-frequency direction ,as well as Difficult negative sample pairs from frequency to time direction .
[0116] Specifically, the training samples are re-input into the model, through... The labels are used for binary classification training, and the cross-entropy loss function is used to determine whether they match. Furthermore, in the time-frequency matching and contrastive learning stage, large-batch hard mining training of difficult negative samples can be achieved by merging batch data from different GPUs.
[0117] Step S54: Reconstruct the masked block based on the mask information to obtain the mask reconstruction loss.
[0118] Specifically, in the masked autoencoder process, the model is designed to simultaneously reconstruct the original signal in both the time and frequency domains. The model first embeds the input in the time domain. Frequency domain embedding A pair of randomly generated mask matrices are applied to identify the block regions that need to be reconstructed. These masks guide the model to focus on the unmasked blocks through a self-attention mechanism, thereby learning their contextual features for reconstruction.
[0119] In an optional embodiment of the present invention, step S54 may include: Step S541: Determine the masked and unmasked blocks in the joint embedding based on the mask information.
[0120] Step S542: Generate a reconstructed representation of the occluded block based on the unoccluded block.
[0121] Specifically, at the top In each UTF Transformer block, U-PCN can be introduced to integrate cross-modal features. All reconstruction outputs are fed into a linear layer to obtain time-domain reconstruction results and frequency-domain reconstruction results.
[0122] Step S543: Determine the mask reconstruction loss based on the reconstructed representation and the corresponding real representation.
[0123] Specifically, the mask reconstruction loss is calculated using absolute error to determine the reconstruction error, which is defined as follows: in, Indicates the loss during mask reconstruction; This represents the time-domain and frequency-domain signals obtained from the model reconstruction; This represents the corresponding real signal; loss calculations are performed independently in the time and frequency domains, respectively.
[0124] Step S55: Train the unified time-frequency Transformer backbone network based on the cross-domain contrast loss, the hard negative sample loss, and the mask reconstruction loss to obtain the pre-trained sleep decoding model.
[0125] Specifically, at each step of the pre-training phase, the model simultaneously performs three tasks: contrastive learning, hard mining of hard negative pairs, and mask reconstruction. These three tasks share the same Transformer encoder structure and collectively contribute to the final feature learning process. The total loss function is a weighted sum of the losses from these three tasks.
[0126] Step S6: Based on labeled sleep monitoring data, the pre-trained sleep decoding model is fine-tuned for downstream tasks to obtain the target sleep decoding model.
[0127] Specifically, after pre-training is completed, the process moves to the downstream task fine-tuning stage. In this stage, the pre-trained sleep decoding model is used as the backbone network, and task-specific layers or task decoders are configured after it based on the specific task type.
[0128] In an optional embodiment of the present invention, step S6 may include: Step S61: Configure the task decoder after the pre-trained sleep decoding model according to the downstream task type to obtain the sleep decoding model to be fine-tuned.
[0129] Specifically, the downstream task types may include sleep stage identification tasks, pathological state classification tasks, key sleep event detection tasks, signal generation tasks, or signal reconstruction tasks. For sleep stage identification tasks, the task decoder can output the corresponding sleep stage category based on the sleep segment-level representation; for pathological state classification tasks, the task decoder can output the sleep disorder, pathological state, or clinical risk category based on a single sleep segment, multiple sleep segment sequences, or an aggregated representation of overnight recordings; for key sleep event detection tasks, the task decoder can output the event category, event location, event duration, or event confidence level; for signal generation or signal reconstruction tasks, the task decoder can generate target channel signals, complete missing signals, or reconstruct PSG representations of specific domains based on fusion features.
[0130] Step S62: Input the labeled sleep monitoring data into the sleep decoding model to be fine-tuned to obtain the task prediction result.
[0131] Specifically, the labeled sleep monitoring data serves as a supervisory signal, used to provide corresponding labeled information to the sleep decoding model to be fine-tuned.
[0132] Step S63: Train the sleep decoding model to be fine-tuned based on the task prediction results and corresponding annotation information to obtain the target sleep decoding model.
[0133] Specifically, for different downstream tasks, some backbone network parameters can be frozen and the task decoder can be trained, or the unified time-frequency Transformer backbone network and task decoder can be fine-tuned end-to-end.
[0134] Step S7: Input the sleep monitoring signal to be decoded into the target sleep decoding model to obtain the sleep decoding result.
[0135] Specifically, the sleep monitoring signal to be decoded is the PSG signal that needs to be decoded for sleep, and the sleep decoding result is output by the target sleep decoding model based on the task decoder.
[0136] In an optional embodiment of the present invention, step S7 may include: Step S71: Standardize the sleep monitoring signal to be decoded to be consistent with the standardized multi-channel sleep monitoring signal to obtain a standardized signal to be decoded.
[0137] Step S72: Input the standardized signal to be decoded into the target sleep decoding model to obtain the sleep features to be decoded.
[0138] Step S73: The sleep features to be decoded are decoded by the task decoder in the target sleep decoding model to obtain the sleep decoding result.
[0139] Specifically, the sleep decoding results may include sleep stage identification results, pathological state classification results, key sleep event detection results, signal generation results, or signal reconstruction results.
[0140] Step S74: Determine channel contribution information based on the channel embedding of the target sleep decoding model and the sleep decoding result.
[0141] Specifically, when determining channel contribution information, the contributions of different channels in different sleep stages or different tasks can be statistically analyzed based on channel embedding, attention weights, mask perturbation results, gradient contributions, or changes in task output.
[0142] Step S75: Based on the channel contribution information and the sleep stage information in the sleep decoding result, generate a channel-stage contribution relationship.
[0143] Specifically, the channel-stage contribution relationship is used to represent the contribution relationship of specific brain region leads, specific PSG channels, or specific frequency band features in the wakefulness period, REM sleep, and non-REM sleep stages.
[0144] This invention also provides a sleep decoding system based on generative pre-trained Transformer, comprising: An input embedding module is used to standardize the polysomnography signal to obtain a standardized multi-channel sleep monitoring signal; based on the standardized multi-channel sleep monitoring signal, a time-domain signal and a frequency-domain representation are generated, and block-marked embedding is performed on the time-domain signal and the frequency-domain representation respectively, and a joint embedding is generated based on channel presence information and mask information; The backbone network module is used to embed the joint input into a unified time-frequency Transformer backbone network and obtain a sleep feature representation through masked attention and domain-aware mapping; The model training module is used to perform a multi-agent pre-training task based on the sleep feature representation, train the unified time-frequency Transformer backbone network, and obtain a pre-trained sleep decoding model. The multi-agent pre-training task includes a cross-domain contrastive learning task, a difficult negative sample matching task, and a mask reconstruction task. The pre-trained sleep decoding model is then fine-tuned by downstream tasks based on labeled sleep monitoring data to obtain a target sleep decoding model. The multi-task output module is used to input the sleep monitoring signal to be decoded into the target sleep decoding model to obtain the sleep decoding result.
[0145] To demonstrate the effectiveness and superiority of the method of this invention, the following experiments were conducted: In experimental verification, the method of this invention was applied to an automatic sleep staging task and compared with various existing sleep staging models, including SeqSleepNet, L-SeqSleepNet, XSleepNet, AttnSleep, DeepSleepNet, IITNet, SleepPyCo, U-Time, SleepTransformer, SleepXViT, and Cross-modalSleepTransformer. The experiments used publicly available datasets such as MASS, PhysioNet2018, SHHS-1, SleepEDF-20, and SleepEDF-78, and used accuracy, macro-average F1 score, and Kappa coefficient as evaluation metrics.
[0146] The results show that the method of this invention achieves or approaches the best performance on multiple benchmark datasets: 88.1% accuracy and 84.22% macro-average F1 score on the MASS dataset, with the macro-average F1 score increasing to 85.1% after adopting adaptive channel configuration; 89.14% accuracy, 82.4% macro-average F1 score, and 0.845 Kappa on the SHHS-1 dataset; and 87.8% accuracy, 81.6% macro-average F1 score, and 0.834 Kappa on the SleepEDF-20 dataset. These results demonstrate that the method of this invention exhibits good classification performance and generalization ability on sleep data from multiple sources, with multiple channel configurations, and of varying scales, and can more stably complete automatic sleep staging compared to existing methods.
[0147] Although the present invention has been disclosed above with reference to embodiments, it is not intended to limit the present invention. Appropriate modifications or equivalent substitutions made by those skilled in the art to the technical solutions of the present invention should be covered within the protection scope of the present invention, which is defined by the claims.
Claims
1. A sleep decoding method based on generative pre-trained Transformer, characterized in that, Includes the following steps: The polysomnography signal was standardized to obtain a standardized multichannel sleep monitoring signal. Based on the standardized multi-channel sleep monitoring signal, a time-domain signal and a frequency-domain representation are generated; Block-marked embedding is performed on the time-domain signal and the frequency-domain representation respectively, and a joint embedding is generated based on channel presence information and mask information; The joint embedding is input into a unified time-frequency Transformer backbone network, and sleep feature representation is obtained through masked attention and domain-aware mapping; Based on the sleep feature representation, a multi-agent pre-training task is performed to train the unified time-frequency Transformer backbone network and obtain a pre-trained sleep decoding model. The multi-agent pre-training task includes a cross-domain contrastive learning task, a difficult negative sample matching task, and a mask reconstruction task. The pre-trained sleep decoding model is fine-tuned for downstream tasks based on labeled sleep monitoring data to obtain the target sleep decoding model. The sleep monitoring signal to be decoded is input into the target sleep decoding model to obtain the sleep decoding result.
2. The method as described in claim 1, characterized in that, The polysomnography (PSG) signals were standardized to obtain standardized multichannel sleep monitoring signals, including: The format of the polysomnography signal is standardized and the sampling rate is aligned to obtain the preprocessed sleep monitoring signal. The preprocessed sleep monitoring signal is channel-mapped according to the standard channel set to obtain the channel-mapped signal; The presence information and missing channel markers are determined based on the channel mapping signal. The channel mapping signal is normalized to obtain a standardized multi-channel sleep monitoring signal.
3. The method as described in claim 1 or 2, characterized in that, Block-marked embedding is performed on the time-domain signal and the frequency-domain representation respectively, and a joint embedding is generated based on channel presence information and mask information, including: Channel-independent convolution and embedding mapping are performed on the time-domain signal to obtain time-domain block labels; Channel-independent convolution and embedding mapping are performed on the frequency domain representation to obtain frequency domain block labels; The time domain block marker is fused with the classification marker, position embedding, channel embedding and block type embedding to obtain the time domain embedding, and the frequency domain block marker is fused with the classification marker, position embedding, channel embedding and block type embedding to obtain the frequency domain embedding; Based on the channel presence information and mask information, the time domain embedding and the frequency domain embedding are marked and constrained to obtain the restricted dual-domain embedding; The restricted dual-domain embedding is aligned by the number of blocks and then concatenated to obtain a joint embedding.
4. The method as described in claim 3, characterized in that, The restricted dual-domain embedding is aligned by the number of blocks and then concatenated to obtain a joint embedding, including: The block resolution configuration is determined based on the number of time domain blocks and the number of frequency domain blocks. The restricted two-domain embedding is reconstructed according to the block resolution configuration to obtain an aligned two-domain embedding; By concatenating the time-domain embedding and the frequency-domain embedding in the aligned dual-domain embedding, a joint embedding is obtained.
5. The method as described in claim 1, characterized in that, The joint embedding is input into a unified time-frequency Transformer backbone network, and sleep feature representations are obtained through masked attention and domain-aware mapping, including: Layer normalization and masked attention calculation are performed on the joint embedding to obtain attention features; Based on the signal domain to which the attention feature belongs, the attention feature is input into the domain perception mapping path to obtain the domain perception feature; The domain-aware features are fused with the corresponding input features to obtain hierarchical time-frequency features; The time-frequency features of the aforementioned levels are passed and converged layer by layer to obtain a sleep feature representation.
6. The method as described in claim 5, characterized in that, Based on the signal domain to which the attention features belong, the attention features are input into the domain-aware mapping path to obtain domain-aware features, including: The temporal features in the attention features are input into the temporal perceptron to obtain the temporal mapping features; The frequency domain features in the attention features are input into the frequency domain perceptron to obtain the frequency domain mapping features; The time-domain mapping feature and the frequency-domain mapping feature are fused to obtain the intra-domain mapping feature; The domain-mapping features are input into a unified perceptron to obtain domain-aware features.
7. The method as described in claim 1, characterized in that, Based on the sleep feature representation, a multi-agent pre-training task is performed to train the unified time-frequency Transformer backbone network, resulting in a pre-trained sleep decoding model, including: Based on the sleep feature representation, determine the time domain representation and frequency domain representation of the same sleep segment; Cross-domain matching is performed on the time-domain representation and the frequency-domain representation to obtain the cross-domain contrast loss; Based on the cross-domain matching results, difficult negative sample pairs are determined, and the difficult negative sample pairs are matched and discriminated to obtain the difficult negative sample loss; The masked block is reconstructed based on the mask information, and the mask reconstruction loss is obtained. The unified time-frequency Transformer backbone network is trained based on the cross-domain contrast loss, the hard negative sample loss, and the mask reconstruction loss to obtain a pre-trained sleep decoding model.
8. The method as described in claim 7, characterized in that, Based on the cross-domain matching results, difficult negative sample pairs are identified, and matching and discrimination are performed on the difficult negative sample pairs to obtain the difficult negative sample loss, including: Determine the cross-domain matching probability based on the cross-domain matching results; Eliminate correct matching pairs from the cross-domain matching probabilities to obtain candidate negative sample pairs; Difficult negative sample pairs are determined based on the matching probability of the candidate negative sample pairs; The difficult negative sample pairs are matched and discriminated to obtain the difficult negative sample loss.
9. The method as described in claim 7, characterized in that, The masked blocks are reconstructed based on the mask information, resulting in a mask reconstruction loss, including: The masked and unmasked blocks in the joint embedding are determined based on the mask information; A reconstructed representation of the occluded block is generated based on the unoccluded block; The mask reconstruction loss is determined based on the reconstructed representation and the corresponding true representation.
10. A sleep decoding system based on generative pre-trained Transformer, characterized in that, include: An input embedding module is used to standardize the polysomnography signal to obtain a standardized multi-channel sleep monitoring signal; based on the standardized multi-channel sleep monitoring signal, a time-domain signal and a frequency-domain representation are generated, and block-marked embedding is performed on the time-domain signal and the frequency-domain representation respectively, and a joint embedding is generated based on channel presence information and mask information; The backbone network module is used to embed the joint input into a unified time-frequency Transformer backbone network and obtain a sleep feature representation through masked attention and domain-aware mapping; The model training module is used to perform a multi-agent pre-training task based on the sleep feature representation, train the unified time-frequency Transformer backbone network, and obtain a pre-trained sleep decoding model. The multi-agent pre-training task includes a cross-domain contrastive learning task, a difficult negative sample matching task, and a mask reconstruction task. The pre-trained sleep decoding model is then fine-tuned by downstream tasks based on labeled sleep monitoring data to obtain a target sleep decoding model. The multi-task output module is used to input the sleep monitoring signal to be decoded into the target sleep decoding model to obtain the sleep decoding result.