Sleep staging method based on multi-head attention mechanism enhanced by dual-stream time and frequency

By introducing time-frequency dual-stream enhancement and multi-head attention mechanisms into the sleep staging method, combined with conditional random field optimization, the problem that existing methods cannot effectively capture sleep stage characteristics and transition rules is solved, and more accurate and efficient sleep staging results are achieved.

CN115399735BActive Publication Date: 2025-05-16NANJING UNIV OF AERONAUTICS & ASTRONAUTICS
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202210882992.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-07-26
Publication Date
2025-05-16
Estimated Expiration
2042-07-26

AI Technical Summary

Technical Problem

Existing machine learning methods cannot effectively capture the important characteristics of each stage of sleep, and ignore the transition rule information between the sleep stage and the previous and subsequent periods, resulting in inaccurate sleep staging results.

Method used

The multi-head attention mechanism based on time-frequency dual-stream enhancement is adopted to obtain the time-domain and frequency domain features through time-frequency transformation, and the multi-head self-attention mechanism is used to learn the correlation between the features, and the preliminary results are optimized in combination with conditional random fields to obtain the final sleep staging results.

Benefits of technology

It achieves more accurate and objective sleep staging results, and can effectively utilize different dimensions of EEG signals, reduce the burden of manual labeling by doctors, and improve the efficiency of sleep quality diagnosis.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115399735B_ABST
    Figure CN115399735B_ABST
Patent Text Reader

Abstract

The present invention discloses a sleep staging method based on a multi-head attention mechanism enhanced by time-frequency dual streams. It belongs to the field of engineering medicine; the steps are: obtaining sleep EEG signals and preprocessing them; using them as time domain information and deriving two branches, one of which is converted into frequency domain information after time-frequency transformation; extracting frequency domain features from the frequency domain information through a frequency domain feature extractor; extracting time domain features from the other through a time domain feature extractor; integrating the above two features to obtain time-frequency dual stream features; passing the time-frequency dual stream features through a feature context learning module to obtain preliminary results of sleep staging; inputting the preliminary results of sleep staging into a conditional random field for optimization to obtain the final results of sleep staging. The present invention not only utilizes the time domain information and frequency domain information of EEG signals, but also learns the association of feature contexts through a multi-head self-attention mechanism, and finally further optimizes the sleep staging results through a conditional random field, thereby obtaining accurate and objective sleep staging results.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of electroencephalogram signal processing, and relates to a sleep staging method based on a multi-head attention mechanism enhanced by a time-frequency dual stream; specifically, it relates to a sleep staging method based on a multi-head attention mechanism enhanced by a time-frequency dual stream. Background Art

[0002] Sleep is an integral part of human activities. We relax and rest during sleep. Sleep is also closely related to the immune system, metabolism, and memory. With the rapid development of modern society, stress, anxiety, and disease also come along. Many people face sleep health problems such as insomnia, sleep apnea syndrome, and drowsiness. They may also suffer from depression, cardiovascular and respiratory diseases.

[0003] Sleep staging is the basis for evaluating and diagnosing sleep quality. In order to evaluate sleep quality, doctors usually need to wear detection equipment for patients with sleep problems to detect sleep conditions and obtain a polysomnogram of the patient, which usually includes EEG, jaw electromyography, electrooculography, and electrocardiography. The doctor first needs to divide the polysomnogram into 30-second intervals, and then divide the 30-second polysomnogram into stages according to the evaluation criteria. The existing standards for judging sleep stages include the American Academy of Medicine (AASM) standard and the Rechtschaffe & Kales (R&K) standard. The R&K standard divides the sleep process into S1, S2, S3, and S4 stages of wakefulness, rapid eye movement, and non-rapid eye movement. The AASM standard merges S3 and S4, and the non-rapid eye movement period is divided into N1, N2, and N3 stages, while the rest remain unchanged, becoming five types of sleep stages.

[0004] Currently, problems such as sleep quality diagnosis and sleep staging require doctors to observe polysomnograms for manual annotation and diagnosis. This process requires experienced experts and is very time-consuming. In addition, different experts may have different evaluations of the same polysomnogram. In recent years, with the promotion of machine learning, some automatic sleep staging methods have also become popular.

[0005] Traditional machine learning sleep staging usually requires manual feature extraction. First, the data is preprocessed and filtered to obtain clean information without impurities, and then the information is extracted for features, and useful information is selected and input into the classifier to achieve the division of sleep stages. The key part here is the selection of features and the selection of classifiers. Commonly used features include time features, frequency features, and nonlinear features, such as power spectral density, differential entropy, sample entropy, etc. Some classification models include support vector machines, random forests, naive Bayes, etc.

[0006] However, manual feature selection is easily restricted by professional constraints and has certain limitations. Compared with traditional machine learning methods, deep learning methods can automatically extract information features from polysomnography and further achieve end-to-end sleep stage prediction, which has received more and more attention and use. Convolutional neural networks are often used for feature extraction, and recursive neural networks are used to learn the timing-related information between signals.

[0007] Existing deep learning-based sleep staging methods have complex inputs and cannot capture important information about each sleep stage. They also ignore the transformation rules between stages, making the transitions between some stages unnatural. There are also differences in the polysomnography acquisition equipment used in different hospitals, which requires the design of a model with common inputs to solve this problem. Therefore, how to better utilize the different characteristics of the model to solve these problems, use less information and calculations to better achieve sleep staging, and help doctors reduce stress is an issue that needs to be addressed at this stage. Summary of the invention

[0008] Purpose of the invention: The purpose of the invention is: sleep staging is the basis of clinical polysomnography staging. The current method is that professional doctors use manual methods to divide the various stages of sleep; this stage is quite time-consuming and boring, and some experts may also make mistakes in stage division due to personal bias and subjective factors; therefore, an attempt is made to use machine learning methods to complete automatic sleep stage division; the existing machine learning methods cannot effectively capture the important characteristics of each stage of sleep, and ignore the transition rule information between the sleep stage and the previous and next periods; therefore, a single-channel EEG automatic sleep staging method based on time-frequency domain information combined with a self-attention mechanism is proposed. From the perspectives of time domain and frequency domain, it can better utilize the information of different dimensions of EEG signals, and use the multi-head self-attention mechanism to learn time-related dependent information to obtain a preliminary sleep staging result. At the same time, considering the correlation between sleep stages, we use conditional random fields to further correct the preliminary results obtained to obtain the final prediction results. The whole process refers to the steps of manual staging by doctors. First, the sleep stage is staged according to the signal characteristics. When encountering an uncertain stage, the previous and next periods are considered to determine the current period. The sleep staging process based on this method is more efficient and objective.

[0009] The technical solution of the present invention is: the sleep staging method based on multi-head attention mechanism enhanced by time-frequency dual streams described in the present invention has the following specific operation steps:

[0010] Step (1.1), obtain sleep EEG signals from existing public datasets;

[0011] Step (1.2), preprocessing the acquired EEG signal to obtain a preprocessed EEG signal;

[0012] Step (1.3), using the preprocessed EEG signal as time domain information, the time domain information derives two branches, one of the branches is converted into frequency domain information after time-frequency transformation, and the converted frequency domain information is extracted with a frequency domain feature extractor;

[0013] The derived branch is passed through a time domain feature extractor to extract time domain features;

[0014] The extracted time domain features and frequency domain features are integrated to obtain the time-frequency dual-stream features;

[0015] Step (1.4), pass the time-frequency dual-stream features through the feature context learning module, use the multi-head self-attention mechanism to learn the correlation between features, and obtain the preliminary results of sleep staging;

[0016] Step (1.5), input the obtained preliminary result of sleep staging into the conditional random field for optimization, so as to obtain the final result of sleep staging.

[0017] Further, in step (1.1), the public data set refers to: a public sleep EEG data set;

[0018] The sleep EEG data set includes sleep EEG signals and sleep period labels annotated by professional doctors;

[0019] The sleep period label is specifically based on the existing sleep staging standard, with each 30 seconds as a window. Professional doctors then evaluate the 30-second EEG signal based on the waveform characteristics to determine the sleep period;

[0020] The sleep periods include wakefulness (W), non-rapid eye movement I (N1), non-rapid eye movement II (N2), non-rapid eye movement III (N3) and rapid eye movement (REM).

[0021] Furthermore, in step (1.2), the specific operation steps of preprocessing the acquired EEG signal are:

[0022] (1.2.1) Remove the labels of the exercise period and the sleep period that cannot be identified, and organize the data according to the five stages of sleep;

[0023] (1.2.2) Keep the EEG data from 30 minutes before the start of sleep to 30 minutes after the end of sleep, and discard the rest.

[0024] Furthermore, in step (1.3), the time-frequency transformation refers to: converting the time domain signal into the frequency domain signal by means of fast Fourier transform, and after the preprocessed EEG signal is subjected to fast Fourier transform, the data of the 0-25 Hz frequency band is intercepted as the frequency domain information.

[0025] Furthermore, in step (1.3), the time domain feature extractor refers to: a convolutional neural network composed of two branches, the convolution kernels of the convolution layers of the two branches have different sizes, and are used to explore feature information of different scales; wherein each branch is composed of a stacked combination of a convolutional layer, a batch normalization layer, a maximum pooling layer, a GELU activation function and a discard layer.

[0026] Furthermore, in step (1.3), the frequency domain feature extractor refers to a convolutional neural network composed of a stack of convolutional layers, batch normalization layers, maximum pooling layers, GELU activation functions, and discard layers.

[0027] Further, in step (1.4), the feature context learning module refers to: a neural network combined with a multi-head self-attention mechanism;

[0028] The module includes two sub-modules: multi-head self-attention and feedforward transmission, and the two sub-modules are repeated twice.

[0029] Furthermore, in step (1.5), the optimization in the conditional random field refers to: using the preliminary results of sleep staging output by the feature context learning module in step (1.4) as the input of the conditional random field, the method uses the preliminary results of sleep staging and true labels of the entire training set for training, and the preliminary results of sleep staging of the test set are optimized in the conditional random field to obtain the final results of the test set.

[0030] The beneficial effect of the present invention is that the sleep staging method proposed in the present invention, which is based on time-frequency dual-stream features and enhanced by a multi-head self-attention mechanism, not only utilizes the time domain information and frequency domain information of the EEG signal, but also learns the association of feature contexts through a multi-head self-attention mechanism, and finally further optimizes the sleep staging results through conditional random fields, thereby obtaining accurate and objective sleep staging results. BRIEF DESCRIPTION OF THE DRAWINGS

[0031] Figure 1 It is an operation flow chart of the present invention;

[0032] Figure 2 is a structural diagram of each module in the frequency domain feature extractor of an embodiment of the present invention;

[0033] Figure 3 It is a structural diagram of each module in the time domain feature extractor of an embodiment of the present invention. DETAILED DESCRIPTION

[0034] The present invention is further described in detail below in conjunction with the embodiments. It should be noted that the protection scope of the present invention is not limited to the following embodiments, and these examples are listed only for illustrative purposes and are not intended to limit the present invention in any way.

[0035] As shown in the figure, the present invention is aimed at single-channel EEG signals in polysomnograms collected by hospitals; according to the AASM sleep staging criteria, the single-channel EEG signals are first cut into time segments of 30s in length, and each 30s time segment is input into the proposed model to obtain the sleep staging result, which is one of the five stages of wakefulness W, non-rapid eye movement N1, non-rapid eye movement N2, non-rapid eye movement N3, and rapid eye movement REM; therefore, for a whole night of polysomnograms, only one of the EEG signal channels needs to be input to obtain the sleep staging result for the whole night, thereby realizing automatic, objective, and efficient sleep staging, helping doctors save precious time. In the follow-up, doctors can analyze sleep quality and make related diagnoses of sleep disorders based on the whole night of sleep staging results output by the model.

[0036] The process of single-channel EEG sleep staging method based on time-frequency domain information and attention mechanism is as follows Figure 1 As shown, in the first stage, the present invention performs a preliminary sleep staging on the signal of the 30s time segment, and the input is two parts of the single-channel EEG signal, one is the original time series signal, which is convenient for obtaining time domain information; the other is the frequency domain information after time-frequency conversion, which can be used to learn frequency domain features; the time domain and frequency domain information will enter the multi-core feature extractor module to learn and capture the features of the time domain and frequency domain respectively; after the learned time domain and frequency domain features are combined, they will enter the feature context learning module, in which the nested multi-head self-attention is used to learn the association and dependency information between features. This stage simulates the doctor's observation of the characteristic waveform of the 30s signal, so as to preliminarily determine the period to which the signal belongs; in the second stage, when imitating the doctor's final determination of the sleep stage, for the signal that cannot be determined, it is usually necessary to observe the stage of the signal in the periods before and after the period to finally determine the period to which the stage belongs; here, the conditional random field is used to learn the conversion transition rules between periods, which can effectively learn the final determination method of the doctor's context stage; the main modules of the method are as follows:

[0037] 1. Multi-core feature extractor module to obtain the time domain and frequency domain features of EEG signals:

[0038] For the input original 30s time-series EEG signal, fast Fourier transform is used for time-frequency conversion, and the frequency band of about 0-25Hz is intercepted as frequency domain information; the original data is used as time domain information; the time domain information and frequency domain information will enter the multi-core feature extractor to learn the time domain and frequency domain features from two perspectives respectively; the multi-core feature extractor is divided into two sub-modules, in which the time domain information is input into the multi-core time domain feature extractor, and the frequency domain information is input into the frequency domain feature extractor; the time and frequency domain features obtained by the two feature extractors are integrated and then input into the next module.

[0039] (1) Frequency domain feature extractor:

[0040] The frequency domain feature extractor is composed of a single-channel convolutional neural network layer stack, including two convolution modules and a maximum pooling module; each convolution layer in the convolution module is followed by a batch normalization and GELU activation function, the purpose is to normalize the data so that the model can have a better generalization effect; the maximum pooling layer is followed by a dropout layer with a certain probability to prevent overfitting of the model; the frequency domain feature extractor uses convolutional neural networks to capture the relevant important features of the EEG signal frequency domain band for frequency domain information. The specific module structure is as follows Figure 2 As shown;

[0041] (2) Time domain feature extractor:

[0042] The time domain feature processor consists of two branches. The convolution kernel sizes of the convolution layers of the two branches are different in order to explore the feature information of different frequencies. The size of the convolution kernel is related to the sampling rate of the EEG signal. Taking a sampling rate of 100Hz as an example, the convolution kernel sizes are set to 50 and 400, corresponding to time windows of 0.5s and 4s respectively. Taking a time window of 4s as an example, it can capture sinusoidal signals as low as 0.25Hz. The feature waveforms that can be captured are different for time windows of different sizes. Therefore, branches with different convolution kernel sizes can obtain signal features of two different scales. Similar to the frequency domain feature extractor, this module is also composed of a convolution layer, a batch normalization layer, a maximum pooling layer, a GELU activation function, and a discard layer. The specific model structure is as follows: Figure 3 shown.

[0043] 2. The feature context learning module learns the dependencies between features and obtains preliminary classification results;

[0044] The feature context learning module draws on the ideas of Transformer's multi-head self-attention and forward propagation. The purpose is to encode and learn the extracted time-frequency domain features. The multi-head self-attention idea is used to learn time-related dependency information. The multi-head idea can process features in parallel, which improves the parallel efficiency of the model. Finally, a preliminary sleep staging result for the 30s EEG information is output. The module consists of sub-modules composed of a multi-head self-attention module and a forward propagation module, which are stacked twice.

[0045] Multi-head self-attention can learn long-term dependencies. Compared with traditional self-attention methods, multi-head self-attention divides the input features into subspaces composed of multiple heads. Each subspace will learn the attention weights in the space, and the heads of different subspaces will also interact with each other to transfer attention information between different subspaces. Therefore, multi-head self-attention can improve the overall model's ability to pay attention to different positions; for the features output by the multi-core feature processor l is the feature length, d is the feature dimension; assuming that the number of heads of multi-head self-attention is H, which is actually 5 in this method, the input feature will be evenly divided into H subspaces, and the feature of each subspace is Where (1≤n≤H); for each subspace n, calculate its corresponding Q according to the learnable weight matrix n , K n 、V n :

[0046]

[0047]

[0048]

[0049] The self-attention A for each subspace n n By Q n , K n 、V n By performing a dot product operation, the specific calculation formula is:

[0050]

[0051] Multi-head self-attention will focus on the self-attention A of each subspace n To perform splicing operations:

[0052] MultiHeadAttention=Concat(A1...A n …A H )

[0053] The multi-head self-attention calculation result MHA will also perform a residual addition operation with the input features and then enter the forward propagation module; the forward propagation module will first perform layer normalization on the input M and then enter the two fully connected layers; the output result will perform a residual operation with the initial input again, and the output F will enter the fully connected layer, and then output the preliminary predicted sleep staging results.

[0054] 3. The conditional random field module is corrected to obtain the final prediction result:

[0055] For the preliminary results obtained in the previous module, since only the information characteristics of the 30-second EEG are considered, it is like a doctor will first have a preliminary judgment on the sleep stage in each 30-second time window. When the EEG information in the time window cannot fully determine the sleep stage, the doctor will also consider the stages before and after the 30-second time window to determine the current stage; based on this idea, the present invention proposes a method for correcting the sleep stage transition rules taking into account the period before and after, and based on the previous preliminary judgment of the sleep stage, the idea of ​​conditional random fields is used to perform sleep stage correction.

[0056] Conditional random field is a discriminant probability model based on an undirected graph that can consider the relationship between adjacent variables. This method is based on linear conditional random field; linear conditional random field defines two random sequences, one is the state sequence I = {i1, i2, …, i T}, one is the observation sequence O = {o1,o2,…,o T Here, the state sequence I is the final desired result, and the observation sequence O is the result of the preliminary prediction of the previous module, where i n , o n ∈{W,N1,N2,N3,REM}(1≤n≤T) represents the true label and the observed preliminary sleep staging result at time n; the final sleep staging prediction result needs to be determined based on the probability from the undirected graph composed of the observation sequence and the state sequence, and its conditional probability distribution is:

[0057]

[0058]

[0059]

[0060] f k (i n ,i n-1 ,o n ) is its characteristic function, which is specifically divided into transfer characteristic function t k (i n ,i n-1 ,o n ) and state characteristic function s l (i n ,i n-1 ,o n ),ω k is the weight of the feature function, K is the total number of feature functions;

[0061] For the constructed conditional probability distribution, the conditional likelihood function is maximized Find the optimal solution, where N is the length of the prediction sequence, I j and O j They represent the state value and observation value of the jth sample respectively; after the trained model is finally obtained, the Viterbi algorithm is used to solve the prediction value, that is, the preliminary sleep staging results are sequence optimized through the conditional random field to obtain the final sleep staging prediction result.

[0062] 4. Loss function setting:

[0063] Due to the imbalance of various categories in the sleep stage, a weighted cross entropy loss function is used:

[0064]

[0065] ω t It is an adjustable weight parameter for each category, M is the total number of samples, and T is the number of sample categories. is the true label of the mth sample, is the predicted label of the mth sample, which together constitute the training loss Loss of the model.

[0066] Example:

[0067] 1. Experimental Dataset

[0068] The public data set Sleep-edf-20 is taken from the version released by PhysioBank in 2013. There are a total of 20 healthy white subjects aged 25 to 101. The subjects were recorded at home for two consecutive days and nights. One day they took tranquilizers and the other day they did not take them. Each recording lasted about 20 hours. Except for subject No. 13 who only had one night of polysomnography data, all other subjects had two nights of polysomnography data, totaling 39 polysomnograms. Each polysomnogram contains two EEG channels (Fpz-Cz and Pz-Oz), one EOG channel, one chin EMG channel, respiration and body temperature, and event markers. The sampling rate of EOG and EEG signals is 100Hz.

[0069] 2. Experimental Setup

[0070] The experiment used the EEG signals of the Fpz-Cz channel, and only captured the data from 30 minutes before falling asleep to 30 minutes after waking up. The labels of the sleep stage data were annotated by experts and summarized into five categories according to the AASM standard, namely wakefulness, rapid eye movement and three non-rapid eye movement periods. In order to evaluate the reliability of the model, the experiment adopted twenty-fold cross-validation. For the twenty subjects, the data of nineteen subjects in each fold participated in the training, and the data of the remaining subject was used for verification. The average accuracy of the twenty-fold results was taken as the final result.

[0071] 3. Experimental results

[0072] Sleep-edf-20 W phase N1 N2 N3 REM period average Accuracy 93.1 30.25 88.35 88.25 91.2 86.2

[0073] Finally, it should be understood that the embodiments described in the present invention are only used to illustrate the principles of the embodiments of the present invention; other variations may also fall within the scope of the present invention; accordingly, the embodiments of the present invention are not limited to the embodiments explicitly introduced and described in the present invention.

Claims

1. A sleep staging method based on a multi-head attention mechanism enhanced by dual-stream time-frequency, characterized in that: The specific steps are as follows: Step (1.1), obtain sleep EEG signals from existing public datasets; Step (1.2), preprocessing the acquired EEG signal to obtain a preprocessed EEG signal; Step (1.3), using the preprocessed EEG signal as time domain information, the time domain information derives two branches, one of the branches is converted into frequency domain information after time-frequency transformation, and the converted frequency domain information is extracted with a frequency domain feature extractor; The derived branch is passed through a time domain feature extractor to extract time domain features; The extracted time domain features and frequency domain features are integrated to obtain the time-frequency dual-stream features; Step (1.4), pass the time-frequency dual-stream features through the feature context learning module, use the multi-head self-attention mechanism to learn the correlation between features, and obtain the preliminary results of sleep staging; Multi-head self-attention divides the input features into subspaces composed of multiple heads. Each subspace will learn the attention weights in the space. The heads of different subspaces will also interact with each other and transfer the attention information between different subspaces. Therefore, multi-head self-attention can improve the overall ability of the model to pay attention to different positions. For the features output by the multi-core feature processor l is the feature length, d is the feature dimension; assuming that the number of heads of multi-head self-attention is H, which is actually 5 in this method, the input feature will be evenly divided into H subspaces, and the feature of each subspace is Where: 1≤n≤H; for each subspace n, calculate its corresponding Q according to the learnable weight matrix n , K n 、V n : The self-attention A for each subspace n n By Q n , K n 、V n By performing a dot product operation, the specific calculation formula is: Multi-head self-attention will focus on the self-attention A of each subspace n To perform splicing operations: MultiHeadAttention=Concat(A1…A n …A H ) The multi-head self-attention calculation result MHA will also perform a residual sum operation with the input features and then enter the forward propagation module; the forward propagation module will first perform layer normalization on the input M and then enter the two fully connected layers; the output result will perform a residual operation with the initial input again, and the output F will enter the fully connected layer, and then output the preliminary predicted sleep stage results; Step (1.5), inputting the obtained preliminary result of sleep staging into the conditional random field for optimization, thereby obtaining the final result of sleep staging; For the preliminary results obtained in the previous module, a preliminary judgment will be made on the sleep stage of each 30-second time window. When the EEG information of the time window cannot fully determine the sleep stage, the stages before and after the 30-second time window will be considered to determine the current stage. Conditional random field is a discriminant probability model based on linear conditional random field; linear conditional random field defines two random sequences, one is the state sequence I = {i1, i2, …, i T }, one is the observation sequence O = {o1,o2,…,o T }; The observation sequence O is the result of the preliminary prediction of the previous module, where i n , o n ∈{W,N1,N2,N3,REM}(1≤n≤T) represents the true label and the observed preliminary sleep staging result at time n; the final sleep staging prediction result needs to be determined based on the probability from the undirected graph composed of the observation sequence and the state sequence, and its conditional probability distribution is: f k (i n ,i n-1 ,o n ) is its characteristic function, which is specifically divided into transfer characteristic function t k (i n ,i n-1 ,o n ) and state characteristic function s l (i n ,i n-1 ,o n ),ω k is the weight of the feature function, K is the total number of feature functions; For the constructed conditional probability distribution, the conditional likelihood function is maximized Find the optimal solution, where N is the length of the prediction sequence, I j and O j Represent the state value and observation value of the jth sample respectively; finally, after obtaining the trained model, the Viterbi algorithm is used to solve the predicted value; Due to the imbalance of various categories in the sleep stage, a weighted cross entropy loss function is used: ω t It is an adjustable weight parameter for each category, M is the total number of samples, and T is the number of sample categories; is the true label of the mth sample, is the predicted label of the mth sample, which together constitute the training loss Loss of the model.

2. The sleep staging method based on multi-head attention mechanism enhanced by time-frequency dual stream according to claim 1 is characterized in that: In step (1.1), the public data set refers to: a public sleep EEG data set; The sleep EEG data set includes sleep EEG signals and sleep period labels annotated by professional doctors; The sleep period label is specifically based on the existing sleep staging standard, with each 30 seconds as a window. Professional doctors then evaluate the 30-second EEG signal based on the waveform characteristics to determine the sleep period; The sleep periods include wakefulness (W), non-rapid eye movement I (N1), non-rapid eye movement II (N2), non-rapid eye movement III (N3) and rapid eye movement (REM).

3. The sleep staging method based on multi-head attention mechanism enhanced by time-frequency dual stream according to claim 1 is characterized in that: In step (1.2), the specific operation steps of preprocessing the acquired EEG signal are: (1.2.1) Remove the labels of the exercise period and the sleep period that cannot be identified, and organize the data according to the five stages of sleep; (1.2.2) Keep the EEG data from 30 minutes before the start of sleep to 30 minutes after the end of sleep, and discard the rest.

4. The sleep staging method based on multi-head attention mechanism enhanced by time-frequency dual stream according to claim 1 is characterized in that: In step (1.3), the time-frequency transformation refers to: converting the time domain signal into the frequency domain signal by means of fast Fourier transform, and after the preprocessed EEG signal is subjected to fast Fourier transform, the data of the 0-25 Hz frequency band is intercepted as the frequency domain information.

5. The sleep staging method based on multi-head attention mechanism enhanced by time-frequency dual stream according to claim 1 is characterized in that: In step (1.3), the time domain feature extractor refers to: a convolutional neural network composed of two branches, the convolution kernels of the convolution layers of the two branches have different sizes, and are used to explore feature information of different scales; wherein each branch is composed of a stacked combination of a convolutional layer, a batch normalization layer, a maximum pooling layer, a GELU activation function and a discard layer.

6. The sleep staging method based on multi-head attention mechanism enhanced by time-frequency dual stream according to claim 1 is characterized in that: In step (1.3), the frequency domain feature extractor refers to a convolutional neural network composed of a stack of convolutional layers, batch normalization layers, maximum pooling layers, GELU activation functions, and discard layers.

7. The sleep staging method based on multi-head attention mechanism enhanced by time-frequency dual stream according to claim 1 is characterized in that: In step (1.4), the feature context learning module refers to: a neural network combined with a multi-head self-attention mechanism; The module includes two sub-modules: multi-head self-attention and feedforward transmission, and the two sub-modules are repeated twice.

8. The sleep staging method based on multi-head attention mechanism enhanced by time-frequency dual stream according to claim 1 is characterized in that: In step (1.5), the optimization in the conditional random field refers to: using the preliminary results of sleep staging output by the feature context learning module in step (1.4) as the input of the conditional random field. This method uses the preliminary results of sleep staging and true labels of the entire training set for training, and the preliminary results of sleep staging of the test set are optimized in the conditional random field to obtain the final results of the test set.

Citation Information

Patent Citations

  • Sleep awakening analysis method based on deep learning

    CN110811558A

  • Algorithm for automatically staging sleep

    CN111631688A