Sleep staging method based on sparse attention mechanism
By combining multiple segment sequence input and sparse attention mechanisms in sleep staging, the problems of high computational complexity and insufficient staging accuracy in the prior art are solved, and more efficient and accurate sleep staging is achieved.
Patent Information
- Application Number
- CN202510108123.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-23
- Publication Date
- 2025-06-06
- Estimated Expiration
- 2045-01-23
AI Technical Summary
Existing deep learning methods are difficult to effectively reduce the complexity of attention calculations in sleep staging, and at the same time use longer time-series context information, resulting in high computational complexity and insufficient staging accuracy.
The method of multiple segment sequence input is adopted, combining the sparse attention mechanism and the Transformer model, the time domain features are extracted through the convolution sliding window, and the Top-T sparse attention mechanism is applied in the Transformer encoding layer to reduce the computational complexity and capture long-range timing dependencies.
It significantly reduces the computational complexity, improves the accuracy and robustness of sleep staging, and can more accurately capture the long-term change information of sleep signals.
Smart Images

Figure CN120093214A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of biomedical signal processing, and in particular to a sleep staging method based on a sparse attention mechanism. Background Art
[0002] Sleep staging is an important method commonly used in scientific research to measure the quality and structure of individual sleep. At present, traditional sleep staging methods mainly rely on manual labeling, which is time-consuming and easily affected by the subjective experience of the labeler. The rise of automatic sleep staging technology makes it possible to automatically complete sleep staging using multi-channel signals such as EEG (electroencephalography), EOG (electro-oculography), EMG (electromyography) and combined with machine learning or deep learning models.
[0003] Most existing deep learning methods are based on models such as CNN, RNN or standard Transformer. CNN can extract local features, but it has certain limitations in modeling long-term correlations. Although RNN and LSTM can capture long-term dependencies, they process long sequences slowly and are prone to gradient vanishing or explosion problems. Since the standard Transformer uses full attention calculations, the computational complexity of long sequence data is extremely high and the memory consumption is huge. Sleep staging involves long time series, and it is usually necessary to capture the gradual changes in sleep stages within a large time range (several minutes to tens of minutes). At the same time, sleep signals have complex temporal dependencies and fine-grained feature differences, and efficient feature extraction and encoding methods are required to capture macroscopic time information and microscopic window features.
[0004] Therefore, a deep learning method that can reduce the computational complexity of attention and utilize longer temporal context information is needed to better adapt to the actual sleep staging task. The present invention proposes a new method to address the above problems. It fully utilizes the context information of sleep signals by inputting multiple segment sequences, and uses a sparse attention mechanism to effectively reduce the computational complexity. It uses window feature extraction and Transformer to extract features from segment sequences, and uses Transformer to extract features from the integrated features of each segment sequence, which significantly improves the staging accuracy. Summary of the invention
[0005] The purpose of the present invention is to overcome the shortcomings of the prior art and provide a sleep staging method based on a sparse attention mechanism, which can achieve efficient and accurate automatic sleep stage classification. The technology makes full use of sleep signal context information by using multiple segment sequence inputs, and uses a sparse attention mechanism to effectively reduce the computational complexity. It uses window feature extraction and Transformer to extract features from segment sequences, and uses Transformer to extract features from the integrated features of each segment sequence to enhance the ability to capture key features, thereby improving the accuracy and operating efficiency of sleep staging.
[0006] To achieve the above object, the technical solution provided by the present invention is: a sleep staging method based on sparse attention mechanism, comprising the following steps:
[0007] S1, obtaining a multi-channel sleep signal and preprocessing it to obtain a single-channel sleep signal, and dividing the single-channel sleep signal into a plurality of continuous segment sequences with a duration of t seconds;
[0008] S2, taking the target segment sequence and its preceding and following n segment sequence data, a total of m segment sequences, as input units, where m=2n+1, and n is a positive integer;
[0009] S3, extracting features from each segment sequence in the input unit, and using a convolution sliding window to extract time domain features of the segment sequence;
[0010] S4. For each segment sequence, the extracted time domain features are input into the Transformer encoding layer using the sparse attention mechanism. The sparse attention mechanism is used to select key feature connections within the segment to obtain feature-enhanced features.
[0011] S5, integrate the enhanced features of each segment sequence to form a comprehensive feature representation of the input unit, and then input the feature representation into the Transformer encoding layer to capture the temporal correlation between segments to obtain the enhanced input unit features;
[0012] S6. Input the feature-enhanced input unit features into the global average pooling layer, and then input the pooled features into the multi-layer fully connected layer for sleep stage classification output.
[0013] Further, in step S1, for the acquired multi-channel sleep signal, the preprocessing performed includes: filtering the multi-channel sleep signal with a bandpass filter to remove noise information; and intercepting the valid interval to retain only the interval of entering sleep, recording data from the first non-awake period and ending at the last non-awake period; selecting one channel from the preprocessed multi-channel sleep signal as a single-channel input.
[0014] Furthermore, in step S3, convolution sliding window feature extraction is used, and the step size is set to half of the window size to ensure that adjacent windows overlap each other and retain the smooth transition of local time features; the feature tensor obtained by sliding window feature extraction is input into the depthwise separable convolution layer, and then the channels are linearly combined using point convolution to obtain the time domain features of the segment sequence.
[0015] Further, in step S4, for each segment sequence, the extracted time domain features are input into the Transformer encoding layer using the sparse attention mechanism. The encoding layer is composed of Transformer Encoder and adopts the Top-T sparse attention mechanism to only calculate the attention weights between each query vector and its most relevant T key vectors, thereby significantly reducing the total attention calculation amount. The specific steps are as follows:
[0016] S41, feature input and linear transformation: Input the time domain features into the Transformer encoding layer, and generate query vectors, key vectors, and value vectors through linear transformation:
[0017] Q=XW Q
[0018] K=XW K
[0019] V=XW V
[0020] Where X is the input feature, which is the time domain feature after processing by the depth-separable convolutional layer, that is, the window feature matrix; Q is the query vector, which is used to calculate the similarity with the key vector and determine the focus of attention allocation; K is the key vector, which is matched with the query vector to determine the attention weight; V is the value vector, which is weighted and summed according to the attention weight to generate the final output feature; W Q , W K , W V is the weight matrix, a linear transformation matrix used to map input features to query vectors, key vectors, and value vectors;
[0021] S42. Similarity calculation: Calculate the similarity between the query vector and all key vectors:
[0022]
[0023]
[0024] In the formula, Q i is the i query vector; K j is the jth key vector; Score(Q i ,K j ) is the query vector Q iWith the key vector K j The similarity score of ; the symbol · is the dot product operation in the vector, which calculates the similarity of two vectors; d k is the dimension of the key vector; d model is the input feature dimension; num_heads is the number of attention heads;
[0025] S43, Top-T selection: For each query vector, select the top T key vectors with the highest similarity, and only calculate the attention weights of these key vectors to form a sparse attention connection:
[0026] Top_T(Q i )={K j |j∈Top_T)Score(Q i ,K j ))}
[0027] In the formula, Top_T(Q i ) is the same as the query vector Q i The index set of the first T key vectors with the highest similarity score; Score(Q i ,K j ) is the query vector Q i With the key vector K j Similarity score of Top_T(Score(Q i ,K j )) is the query vector Q i With all key vectors K j The scores of the top T in the similarity scores;
[0028] S44, attention weight calculation and weighted summation: only weighted summation is performed on the selected Top-T key vectors to generate output features:
[0029]
[0030] In the formula, α ij is the query vector Q i With the key vector K j Normalized attention weight of Top_T(Q i ) is the same as the query vector Q i The index set of the first T key vectors with the highest similarity scores; softmax(·) is the softmax function, which converts the input vector into a probability distribution so that the sum of all output values is 1; V j is the key vector K j Corresponding value vector; Attention i is the query vector Q i The final output features.
[0031] Furthermore, in step S5, the features of each segment sequence after feature enhancement are integrated, including concatenating the features of multiple segment sequences in the channel dimension and then passing them through a one-dimensional convolutional layer to form a comprehensive feature representation of the input unit; the feature representation of the input unit is input into the Transformer encoding layer to capture the temporal correlation between segments, and the encoding layer is composed of a Transformer Encoder.
[0032] Further, the specific steps of step S6 are as follows:
[0033] S61, inputting the feature-enhanced input unit features into the global average pooling layer to obtain pooled features;
[0034] S62, pass the pooled features through the first fully connected layer, introduce nonlinearity through the ReLU activation function, and apply Dropout regularization to reduce the risk of overfitting;
[0035] S63. After passing through the second fully connected layer, the features are further mapped to the output space of the sleep stage category, and the output is converted into the probability distribution of each category through the Softmax function, and the category with the largest probability is taken as the final prediction result.
[0036] Compared with the prior art, the present invention has the following advantages and beneficial effects:
[0037] 1. The present invention combines the target segment sequence and several segment sequence data before and after it into a group of inputs to form multiple segment sequence inputs, so that the present invention can capture the sleep signal change information over a longer time range. Compared with the traditional method of only analyzing the target segment sequence, the multiple segment sequence inputs significantly improve the method's understanding of the transition and continuity of physiological sleep stages, and achieve more accurate and stable staging results;
[0038] 2. The present invention adopts the Top-T sparse attention mechanism, which only calculates the first T key vectors with the highest similarity. This sparse attention mechanism not only effectively reduces the computational burden and storage overhead of the full attention, but also can highlight the most discriminative feature connections in the sleep signal, improve the efficiency of capturing key timing information, and compared with the traditional global attention model, it greatly reduces the computational complexity while maintaining high accuracy;
[0039] 3. The method of using window feature extraction and Transformer to extract features from segment sequences, and using Transformer to extract features from the integrated features of each segment sequence, can capture the potential long-range temporal dependencies between different segment sequences. The structure of extracting features from both segment sequences and the integrated features of each segment sequence makes up for the lack of continuity in the traditional single segment sequence analysis method.
[0040] 4. The present invention has strong adaptability and versatility, and can be extended to a variety of physiological signal analysis. Although the present invention focuses on sleep staging, its multiple segment sequence input, feature extraction method and sparse attention mechanism can also be extended to the long-term sequence analysis of other physiological signals, and has strong transferability and applicability;
[0041] In summary, the present invention cooperates with each other in terms of multiple segment sequence input, sparse attention mechanism, feature extraction, etc., which not only greatly reduces the amount of calculation, but also significantly improves the accuracy and robustness of sleep staging, and makes up for the shortcomings of the existing technology in insufficient utilization of long-term time series information or excessively high computational cost, and has obvious advantages and beneficial effects. BRIEF DESCRIPTION OF THE DRAWINGS
[0042] Figure 1 It is a framework diagram of the method of the present invention.
[0043] Figure 2 Schematic diagram of the sparse attention mechanism. DETAILED DESCRIPTION
[0044] The present invention is further described in detail below in conjunction with embodiments and drawings, but the embodiments of the present invention are not limited thereto.
[0045] like Figure 1 and Figure 2 As shown, this embodiment discloses a sleep staging method based on a sparse attention mechanism, the details of which are as follows:
[0046] S1, obtaining a multi-channel sleep signal and preprocessing it to obtain a single-channel sleep signal, and dividing the single-channel sleep signal into a plurality of continuous segment sequences with a duration of t seconds;
[0047] For the acquired multi-channel sleep signal, the preprocessing performed includes: filtering the multi-channel sleep signal with a bandpass filter to remove noise information; and intercepting the valid interval to retain only the interval of entering sleep, recording data from the first non-awake period and ending at the last non-awake period; selecting one channel from the preprocessed multi-channel sleep signal as a single-channel input.
[0048] S2, taking the target segment sequence and its preceding and following n segment sequence data, a total of m segment sequences, as input units, where m=2n+1, and n is a positive integer;
[0049] S3. Perform feature extraction on each segment sequence in the input unit and use a convolution sliding window to extract the time domain features of the segment sequence, as follows:
[0050] Convolutional sliding window feature extraction is used, and the step size is set to half of the window size to ensure that adjacent windows overlap each other and retain the smooth transition of local time features; the feature tensor obtained by sliding window feature extraction is input into the deep separable convolution layer, and then the point convolution is used to linearly combine each channel to obtain the time domain features of the segment sequence.
[0051] S4. For each segment sequence, the extracted time domain features are input into the Transformer encoding layer using the sparse attention mechanism. The sparse attention mechanism is used to select key feature connections within the segment to obtain feature-enhanced features.
[0052] For each segment sequence, the extracted time domain features are input into the Transformer encoding layer with sparse attention mechanism. The encoding layer consists of Transformer Encoder and adopts Top-T sparse attention mechanism to only calculate the attention weights between each query vector and its most relevant T key vectors, thereby significantly reducing the total amount of attention calculation. The specific steps are as follows:
[0053] S41, feature input and linear transformation: Input the time domain features into the Transformer encoding layer, and generate query vectors, key vectors, and value vectors through linear transformation:
[0054] Q=XW Q
[0055] K=XW K
[0056] V=XW V
[0057] Where X is the input feature, which is the time domain feature after processing by the depth-separable convolutional layer, that is, the window feature matrix; Q is the query vector, which is used to calculate the similarity with the key vector and determine the focus of attention allocation; K is the key vector, which is matched with the query vector to determine the attention weight; V is the value vector, which is weighted and summed according to the attention weight to generate the final output feature; W Q , W K , W V is the weight matrix, a linear transformation matrix used to map input features to query vectors, key vectors, and value vectors;
[0058] S42. Similarity calculation: Calculate the similarity between the query vector and all key vectors:
[0059]
[0060] In the formula, Q i is the i query vector; K jis the jth key vector; Score(Q i ,K j ) is the query vector Q i With the key vector K j The similarity score of ; the symbol · is the dot product operation in the vector, which calculates the similarity of two vectors; d k is the dimension of the key vector; d model is the input feature dimension; num_heads is the number of attention heads;
[0061] S43, Top-T selection: For each query vector, select the top T key vectors with the highest similarity, and only calculate the attention weights of these key vectors to form a sparse attention connection:
[0062] Top_T(Q i )={K j |j∈Top_T(Score(Q i ,K j (
[0063] In the formula, Top_T(Q i ) is the same as the query vector Q i The index set of the first T key vectors with the highest similarity score; Score(Q i ,K j ) is the query vector Q i With the key vector K j Similarity score of Top_T(Score(Q i ,K j )) is the query vector Q i With all key vectors K j The scores of the top T in the similarity scores;
[0064] S44, attention weight calculation and weighted summation: only weighted summation is performed on the selected Top-T key vectors to generate output features:
[0065]
[0066] In the formula, α ij is the query vector Q i With the key vector K j Normalized attention weight of Top_T(Q i ) is the same as the query vector Q i The index set of the first T key vectors with the highest similarity scores; softmax(·) is the softmax function, which converts the input vector into a probability distribution so that the sum of all output values is 1; V j is the key vector K j Corresponding value vector; Attentioni is the query vector Q i The final output features.
[0067] S5. Integrate the enhanced features of each segment sequence to form a comprehensive feature representation of the input unit, and then input the feature representation to the Transformer encoding layer to capture the temporal association between segments to obtain the feature of the input unit after feature enhancement; wherein, integrate the enhanced features of each segment sequence, including concatenating the features of multiple segment sequences in the channel dimension, and then passing through a one-dimensional convolutional layer to form a comprehensive feature representation of the input unit; input the feature representation of the input unit to the Transformer encoding layer to capture the temporal association between segments, and the encoding layer is composed of a Transformer Encoder.
[0068] S6: Input the feature-enhanced input unit features into the global average pooling layer, and then input the pooled features into the multi-layer fully connected layer for sleep stage classification output. The specific steps are as follows:
[0069] S61, inputting the feature-enhanced input unit features into the global average pooling layer to obtain pooled features;
[0070] S62, pass the pooled features through the first fully connected layer, introduce nonlinearity through the ReLU activation function, and apply Dropout regularization to reduce the risk of overfitting;
[0071] S63. After passing through the second fully connected layer, the features are further mapped to the output space of the sleep stage category, and the output is converted into the probability distribution of each category through the Softmax function, and the category with the largest probability is taken as the final prediction result.
[0072] The above embodiments are preferred implementation modes of the present invention, but the implementation modes of the present invention are not limited to the above embodiments. Any other changes, modifications, substitutions, combinations, and simplifications that do not deviate from the spirit and principles of the present invention should be equivalent replacement methods and are included in the protection scope of the present invention.
Claims
1. A sleep staging method based on sparse attention mechanism, characterized in that: The following steps are involved: S1, obtaining a multi-channel sleep signal and preprocessing it to obtain a single-channel sleep signal, and dividing the single-channel sleep signal into a plurality of continuous segment sequences with a duration of t seconds; S2, taking the target segment sequence and its preceding and following n segment sequence data, a total of m segment sequences, as input units, where m=2n+1, and n is a positive integer; S3, extracting features from each segment sequence in the input unit, and using a convolution sliding window to extract time domain features of the segment sequence; S4. For each segment sequence, the extracted time domain features are input into the Transformer encoding layer using the sparse attention mechanism. The sparse attention mechanism is used to select key feature connections within the segment to obtain feature-enhanced features. S5, integrate the enhanced features of each segment sequence to form a comprehensive feature representation of the input unit, and then input the feature representation into the Transformer encoding layer to capture the temporal correlation between segments to obtain the enhanced input unit features; S6. Input the feature-enhanced input unit features into the global average pooling layer, and then input the pooled features into the multi-layer fully connected layer for sleep stage classification output.
2. A sleep staging method based on sparse attention mechanism according to claim 1, characterized in that: In step S1, for the acquired multi-channel sleep signal, the preprocessing performed includes: filtering the multi-channel sleep signal with a bandpass filter to remove noise information; and intercepting the valid interval, retaining only the interval of entering sleep, recording data from the first non-awake period, and ending at the last non-awake period; selecting one channel from the preprocessed multi-channel sleep signal as a single channel input.
3. A sleep staging method based on sparse attention mechanism according to claim 2, characterized in that: In step S3, convolution sliding window feature extraction is used, and the step size is set to half of the window size to ensure that adjacent windows overlap each other and retain the smooth transition of local time features; the feature tensor obtained by sliding window feature extraction is input into the depthwise separable convolution layer, and then the channels are linearly combined using point convolution to obtain the time domain features of the segment sequence.
4. A sleep staging method based on sparse attention mechanism according to claim 3, characterized in that: In step S4, for each segment sequence, the extracted time domain features are input into the Transformer encoding layer using the sparse attention mechanism. The encoding layer consists of a Transformer Encoder and uses the Top-T sparse attention mechanism to only calculate the attention weights between each query vector and its most relevant T key vectors, thereby significantly reducing the total amount of attention calculations. The specific steps are as follows: S41, feature input and linear transformation: Input the time domain features into the Transformer encoding layer, and generate query vectors, key vectors, and value vectors through linear transformation: Q=XW Q K=XW K V=XW V Where X is the input feature, which is the time domain feature after processing by the depth-separable convolutional layer, that is, the window feature matrix; Q is the query vector, which is used to calculate the similarity with the key vector and determine the focus of attention allocation; K is the key vector, which is matched with the query vector to determine the attention weight; V is the value vector, which is weighted and summed according to the attention weights to generate the final output features; W Q , W K , W V is the weight matrix, a linear transformation matrix used to map input features to query vectors, key vectors, and value vectors; S42. Similarity calculation: Calculate the similarity between the query vector and all key vectors: In the formula, Q i is the i query vector; K j is the jth key vector; Score(Q i ,K j ) is the query vector Q i With the key vector K j The similarity score of ; the symbol · is the dot product operation in the vector, which calculates the similarity of two vectors; d k is the dimension of the key vector; d model is the input feature dimension; num_heads is the number of attention heads; S43, Top-T selection: For each query vector, select the top T key vectors with the highest similarity, and only calculate the attention weights of these key vectors to form a sparse attention connection: Top_T(Q i )={K j |j∈Top_T(Score(Q i ,K j ))} In the formula, Top_T(Q i ) is the same as the query vector Q i The index set of the first T key vectors with the highest similarity score; Score(Q i ,K j ) is the query vector Q i With the key vector K j Similarity score of Top_T(Score(Q i ,K j )) is the query vector Q i With all key vectors K j The scores of the top T in the similarity scores; S44, attention weight calculation and weighted summation: only weighted summation is performed on the selected Top-T key vectors to generate output features: In the formula, α ij is the query vector Q i With the key vector K j Normalized attention weight of Top_T(Q i ) is the same as the query vector Q i The index set of the first T key vectors with the highest similarity scores; softmax(·) is the softmax function, which converts the input vector into a probability distribution so that the sum of all output values is 1; V j is the key vector K j Corresponding value vector; Attention i is the query vector Q i The final output features.
5. A sleep staging method based on sparse attention mechanism according to claim 4, characterized in that: In step S5, the enhanced features of each segment sequence are integrated, including concatenating the features of multiple segment sequences in the channel dimension and then passing them through a one-dimensional convolutional layer to form a comprehensive feature representation of the input unit; the feature representation of the input unit is input into the Transformer encoding layer to capture the temporal correlation between segments, and the encoding layer is composed of a Transformer Encoder.
6. A sleep staging method based on sparse attention mechanism according to claim 5, characterized in that: The specific steps of step S6 are as follows: S61, inputting the feature-enhanced input unit features into the global average pooling layer to obtain pooled features; S62, pass the pooled features through the first fully connected layer, introduce nonlinearity through the ReLU activation function, and apply Dropout regularization to reduce the risk of overfitting; S63. After passing through the second fully connected layer, the features are further mapped to the output space of the sleep stage category, and the output is converted into the probability distribution of each category through the Softmax function, and the category with the largest probability is taken as the final prediction result.
Citation Information
Patent Citations
Long time series data prediction method based on bidirectional sparse mechanism Transform
CN116541435A
Sparse attention computation model and method, electronic device, and storage medium
WO2023221940A1
Cited By
Multi-source cooperative sensing information low-delay transmission method based on intelligent traffic networking
CN120785925A