A Multimodal and Multiscale Sleep Staging Method for Improving Category Confusion in N1 Stage
Through multimodal multi-scale feature extraction and comparison learning methods, the problem of low classification accuracy in N1 phase is solved, and higher classification accuracy in N1 phase and the accuracy of sleep stages is achieved.
Patent Information
- Application Number
- CN202310152184.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-02-22
- Publication Date
- 2025-07-01
- Estimated Expiration
- 2043-02-22
AI Technical Summary
Among the existing automated sleep staging methods, the N1 phase classification accuracy is low, and the existing methods do not fully utilize sleep data and fail to effectively extract representative features, resulting in the N1 phase, N2 phase and REM phase easily confusing.
The multimodal multi-scale feature extraction method is adopted to expand the number of samples in the N1 phase through data augmentation technology, and the EEG signal is extracted fine-grained feature, and more representative features are screened out through the residual SE attention module. The time dependence of the data is mined in combination with the Bi-LSTM method, and finally the staging boundary is expanded by using the comparison learning method to improve the differentiability of the N1 phase.
On the basis of ensuring the overall accuracy, significantly improve the classification accuracy of N1 phase, reduce the confusion rate between N1 phase and other stages, and improve the accuracy of sleep stages.
Smart Images

Figure CN116172515B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical fields of signal processing, pattern recognition, etc., and relates to a multi-modal and multi-scale sleep staging method for improving the category confusion in the N1 stage. Background Art
[0002] According to the American Academy of Sleep Medicine (AASM) standard and the sleep manual (Rechtschaffen and Kales, R&K), sleep staging is divided into the wakefulness (W) stage, rapid eye movement (REM) stage, non-rapid eye movement 1 (N1) stage, non-rapid eye movement 2 (N2) stage, and non-rapid eye movement 3 (N3) stage. The AASM standard divides the overnight polysomnogram (PSG) data into non-overlapping segments of 30 seconds each, and manually discriminates these data segments into different sleep stages one by one, which results in a large amount of time and effort being consumed for sleep staging. To solve this problem, numerous automated sleep staging methods have been developed for sleep staging. However, the existing automated sleep staging methods all have the problem of low accuracy in the N1 stage.
[0003] For sleep data, the duration of each sleep stage is not equal. Specifically, the N2 stage accounts for the majority, about 45%-55% of the total sleep time, while the N1 stage only accounts for 2%-5%. This problem exists in all available sleep datasets, such as the Montreal Archive of Sleep Studies (MASS) database and the PhysioNet's Sleep-EDF database. Some studies have used methods such as oversampling, class imbalance samplers, and GANs to supplement the minority classes, but these methods only expand the number of samples and do not guarantee the authenticity of the samples.
[0004] The existing models do not fully utilize sleep data to extract more representative feature representations. First, the modality is not fully utilized. Since different modality sleep signals can distinguish sleep states from different perspectives, more and more multi-modal deep learning models are used for sleep staging. However, most multi-modal automated sleep staging models treat different modalities in the same way and do not pay attention to the differences between different modalities. Second, the frequency is not fully utilized. The existing sleep staging methods pay less attention to different frequencies, and only a few studies distinguish between high frequency and low frequency. However, only coarsely using high frequency and low frequency cannot accurately extract the features of each sleep stage. Finally, the temporal structure is not fully utilized. Sleep data is a continuous time series, so it is very necessary to learn the temporal relationship of sleep staging.
[0005] Stage N1 is in the transitional stage between REM and N2 stages. The boundary of sleep staging is not clear, and it is easy to be confused with N2 and REM stages. According to the AASM rules, sleep data is divided into 30-second segments to define sleep staging. However, due to the limitations of data division, the divided sleep segments may contain sleep data of multiple stages, which causes certain difficulties for accurate staging. According to the distribution characteristics of sleep staging, the confusion phenomenon is particularly obvious in stage N1. Therefore, it is imperative to apply contrastive learning in sleep staging to expand the classification boundary.
[0006] In summary, the present invention aims at the problem of low classification accuracy of stage N1 in sleep staging, analyzes the reasons for this problem from the above three aspects, and proposes new ideas and thoughts. Therefore, the present invention proposes a multi-modal and multi-scale sleep staging method to improve the category confusion of stage N1. Summary of the Invention
[0007] In view of the above background, the present invention proposes a multi-modal and multi-scale sleep staging method to improve the category confusion of stage N1, which improves the classification accuracy of stage N1 on the basis of ensuring the overall accuracy. After preprocessing, this method first uses data augmentation technology to expand the number of samples in stage N1 and alleviate the impact of class imbalance on the classification accuracy of stage N1. Then, a multi-modal and multi-scale feature extraction method is used to fully extract sleep features, different processing is performed on different modal data, and a multi-scale feature extraction method is used to perform fine-grained feature extraction on the EEG modality to improve the effectiveness of features. The extracted multi-modal features are fused, and a residual SE attention module is used to screen out more representative features, and the Bi-LSTM method is used to fully exploit the temporal dependence of the data. Finally, aiming at the problem that stage N1 is easily confused with N2 and REM stages, a contrastive learning method is used to improve the similarity of sleep data features in the same stage and reduce the similarity of sleep data features in different stages, thereby further improving the distinguishability of stage N1.
[0008] To achieve the above object, the present invention adopts the following technical solutions:
[0009] A multi-modal and multi-scale sleep staging method for improving the category confusion of stage N1, comprising:
[0010] Step 1, preprocessing operation of the original signal
[0011] First, select the Fpz-Cz channel in the electroencephalography (EEG) signal, electrooculography (EOG) signal, and electromyography (EMG) signal as the experimental data. Divide the sleep data with a sampling frequency of T into segments of 30 seconds, delete the MOVEMENT and UNKNOWN tags and the corresponding data, and merge the N3 and N4 stages of data and collectively call them the N3 stage. Only retain the 30-minute W stage before and after the sleep time in the training set. After preprocessing, obtain N old samples, and each sample data x i is data with a length of 30×T, where i ∈ [1, 2, …, N old represents the serial number of the current sample. Assume the obtained sample set is X old , and the label set is Y old .
[0012] Step 2: Data augmentation
[0013] For the training set samples obtained in Step 1, use the stacking strategy. For two adjacent N1-stage samples, take the middle part of the two samples to form a new sample. Put the newly generated sample into the original data set X old to obtain a new sample set X, and correspondingly update the label set Y old to the new label set Y.
[0014] Step 3: Multi-modal multi-scale feature extraction
[0015] In order to make full use of the sleep data and extract more representative features, the present invention designs a multi-modal multi-scale feature extraction method. Perform feature extraction on the new sample data obtained in Step 2. Process the data obtained in Step 2 differently according to different modalities, and perform multi-scale feature extraction on the EEG signal by frequency. Then splice the features extracted from each modality, send them into the residual SE attention module, select the representative features and remove the redundant features. Finally, use a double-layer Bi-LSTM to learn the transition rules of sleep staging to obtain the final features.
[0016] Step 4: Comparison and classification
[0017] Through Step 3, representative features of sleep staging can be obtained. With the help of contrastive learning, the similarity between the features of synchronous sleep data is enhanced, and the similarity between the features of asynchronous sleep data is reduced, thereby expanding the staging boundary and making it easier to distinguish between each staging. At the same time, use a classifier to obtain the prediction result.
[0018] Compared with the prior art, the method proposed by the present invention has the following advantages:
[0019] In existing automated sleep staging methods, in the face of the class imbalance problem of sleep data, methods such as GAN are usually used to generate minority class data, and the authenticity of the generated samples cannot be guaranteed. General multimodal sleep staging methods do not pay attention to the differences between modalities, perform the same processing on data of different modalities, and pay less attention to frequency features. Individual methods only focus on high frequencies and low frequencies, and do not perform fine-grained feature extraction from the perspective of EEG signal characteristic waves. In addition, regarding the problem that the N1 stage in sleep staging is easily confused with other stages, especially the N2 stage and the REM stage, previous sleep staging methods have not paid much attention. In contrast, the present invention uses a data augmentation algorithm based on an overlapping strategy to generate N1 stage data, expanding the number of N1 stage samples while ensuring the authenticity of the samples; and uses a multimodal multi-scale feature extraction method to make full use of sleep data and extract more representative sleep features; finally, contrastive learning is used to expand the staging boundary to solve the problem of easy confusion in the N1 stage.
[0020] The present invention adopts an overlapping strategy, selects the middle part of adjacent two N1 stage data as the new N1 stage, which not only ensures the authenticity of the samples but also achieves the purpose of expanding the number of samples. Moreover, according to the different manual evaluation bases for different modalities, the present invention performs different feature extractions on data of different modalities, and at the same time uses a multi-scale feature extraction method to perform fine-grained feature extraction on EEG signal data. Contrastive learning is used to expand the staging boundary to solve the problem of easy confusion in the N1 stage, constraining the model to improve the similarity of sleep data in the same period and relatively reducing the similarity of sleep data in different periods, further improving the classification accuracy of the N1 stage. BRIEF DESCRIPTION OF THE DRAWINGS
[0021] Figure 1 is the flowchart of the method of the present invention;
[0022] Figure 2 is the overall structure diagram of the present invention;
[0023] Figure 3 is the data augmentation algorithm diagram based on the overlapping strategy;
[0024] Figure 4 is the multimodal multi-scale feature extraction module diagram;
[0025] Figure 5 is the contrastive learning effect diagram. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0026] The following further describes the present invention in combination with the drawings and specific implementation details.
[0027] A multimodal multi-scale sleep staging method for improving the category confusion in the N1 stage, the flowchart is as Figure 1As shown in the figure, the network is named the MultiModal and MultiScale Convolution Network (MMSC), and the overall network framework diagram is as follows Figure 2 as shown.
[0028] Step 1, original signal preprocessing
[0029] During the implementation of the present invention, the Sleep-EDF-20 dataset composed of the first 20 subjects in the publicly available dataset Sleep-EDF is used. In the experiment, the Fpz-Cz channel in the EEG signal, the EOG signal, and the EMG signal are selected as the experimental data, that is, the number of channels Q is 3. They are divided into segments at 30-second intervals, and the MOVEMENT and UNKNOWN tags and the data corresponding to these tags are deleted. The N3 and N4 stage data are merged and collectively referred to as the N3 stage. Assume that there is a set of training datasets X old and a set of test data X test . The data sampling frequency is T, and in the Sleep-EDF-20 dataset, T is 100 Hz. In X old , only the W stage within 30 minutes before and after the sleep time is retained. Finally, N old segments are obtained, and the length of each segment is 30×T, resulting in a sample data size of (Q, 30×T), that is, the dimension of X old is (N old , Q, (30×T)).
[0030] Step 2, data augmentation
[0031] Using the stacking strategy, two consecutive similar samples in X old are selected and . The data of the last v seconds of and the data of the first (30 - v) seconds of are spliced together to form a complete new sleep data with the label . In the present invention, v takes 15, denotes the i-th newly generated sleep data, and denotes the label corresponding to the i-th newly generated sleep data. As Figure 3 shown. In this way, all adjacent N1 stage samples are selected to generate new N1 stage samples, and the M newly generated samples are put into the original dataset X old to obtain a new sample set X, and accordingly update its label set Y and the total number of samples N. The generation algorithm formula based on the stacking strategy is as follows:
[0032]
[0033]
[0034] N = N old + M
[0035] wherein, represents the new samples generated for class c, c ∈ [0, 1, 2, 3, 4], and c takes 1 in the present invention. O(·) represents the stacking function of the above process. M is the number of new samples generated for class c. N is the number of original segments N old The number of samples in the training set X after data augmentation to generate M new N1-phase samples, and the dimension of X is (N, Q, (30 × T)).
[0036] Step 3, multi-modal multi-scale feature extraction
[0037] Take the new sample set X after data augmentation as the input data and send it into the feature extraction module for multi-modal multi-scale feature extraction. The structure diagram of this module is as Figure 4 shown. The feature extraction process is mainly divided into 4 steps:
[0038] (1) Multi-modal feature extraction. Extract sleep features from the sleep data X obtained in step 2 in different modalities. The sleep staging manual has different staging bases for data in different modalities. Specifically, judge the sleep stage according to the presence or absence of EMG signals, various characteristic waves in EEG signals, and the change speed of EOG signals. Send the different modality data of the input data into the corresponding modality feature extraction branches: directly send the EMG data into the convolutional module for feature extraction, send the EEG data into different convolutional kernel modules for multi-scale feature extraction, first perform short-time Fourier transform on the EOG data, and then send it into the convolutional module for feature extraction. The short-time Fourier transform uses a Hanning window with a window size of 100 to extract the time-frequency features of the EOG signal. The feature extraction formulas for EMG and EOG signals are as follows:
[0039] f EMG = E EMG (X EMG , s EMG )
[0040]
[0041] wherein, f EMG , f EOG are the features extracted from the EMG signal and the EOG signal respectively. X EEG and X EMG are the EEG signal and EMG signal data of the data set X respectively, is the EOG signal after short-time Fourier transform of the data set X. s EMG , s EOGThey are the convolution kernels for extracting EMG signal and EOG signal feature selection, where s EMG = 10, s EOG = (8, 8). E u (·) is a function for feature extraction using the convolution module ConvBlock, where u ∈ [EMG, EOG, EEG1, EEG2, EEG3], and ConvBlock includes a convolutional layer, a batch normalization layer, a GELU activation function layer, a max pooling layer, and a Dropout layer with a parameter of 0.5 to prevent overfitting. E EMG is the EMG convolution module, which includes three one-dimensional ConvBlocks. The convolutional kernel sizes in each convolution module are 10, 8, and 8 respectively, the convolutional strides are all 3, and the kernels of the max pooling layer are 6, 6, and 4 respectively, and the pooling strides are 3, 3, and 4. E EOG is the EOG convolution module, which includes three two-dimensional ConvBlocks. The convolutional kernel sizes in each convolution module are 8×8, 7×7, and 7×7 respectively, the convolutional strides are 3×3, 1×1, and 1×1 respectively, and the kernels of the max pooling layer are 6×6, 4×4, and 2×2 respectively, and the pooling strides are 3×3, 2×2, and 2×2.
[0042] (2) Multi-scale feature extraction. Using the EEG signal data X EEG in (1) as the input, select different-scale convolutional kernel sizes according to the characteristic waves to extract the multi-scale sleep features of the EEG signal, and output the EEG signal feature f EEG , where v ∈ [1, 2, 3]. The characteristic waves in the EEG signal mainly include α wave, β wave, δ wave, θ wave, K complex wave, and Spindle wave. Different sleep stages contain different frequencies of brain waves. The W stage contains α wave and β wave, the N1 stage contains θ wave, the N2 stage contains spindle wave and K complex wave, the N3 stage contains δ wave, and the REM stage contains α wave, β wave, and θ wave. Among them, the α wave corresponds to a frequency of 8 - 11 Hz, the β wave corresponds to a frequency of 11 - 30 Hz, the θ wave corresponds to a frequency of 4 - 8 Hz, the δ wave corresponds to a frequency of 0.5 - 3 Hz, the K complex wave corresponds to a frequency of 0.5 - 1 Hz, and the Spindle wave corresponds to a frequency of 12 - 14 Hz. The correspondence between sleep stages and waveforms, frequencies is shown in Table 1. The present invention selects 1 Hz to extract the features in the frequency range corresponding to the K complex wave to distinguish the N2 and N3 stages from other stages; selects 2 Hz to extract the features in the frequency range corresponding to the δ wave to distinguish the N2 stage and the N3 stage; selects 10 Hz to extract the features in the frequency range corresponding to the θ wave to distinguish the N1 stage from the W stage and the REM stage. The EEG signal feature extraction formula is as follows:
[0043]
[0044] Among them, concat(·) represents the concatenation function, which is the convolution kernel of different scales for extracting EEG signal features. In the present invention, the convolution kernel is obtained by dividing the sampling frequency T by the frequency, that is E EEG1 (·) is with as the convolution kernel of the EEG convolution module, which includes three one-dimensional ConvBlocks. The convolution kernel sizes in each convolution module are 100, 7, and 7 respectively, the convolution strides are 12, 1, and 1 respectively, the kernels of the max pooling layer are 4, 6, and 4 respectively, and the pooling strides are 2, 3, and 4. E EEG2 (·) is with as the convolution kernel of the EEG convolution module, which includes three one-dimensional ConvBlocks. The convolution kernel sizes in each convolution module are 50, 7, and 7 respectively, the convolution strides are 6, 2, and 1 respectively, the kernels of the max pooling layer are 6, 6, and 4 respectively, and the pooling strides are 3, 3, and 4. E EEG3 (·) is with as the convolution kernel of the EEG convolution module, which includes three one-dimensional ConvBlocks. The convolution kernel sizes in each convolution module are 10, 7, and 7 respectively, the convolution strides are 4, 3, and 3 respectively, the kernels of the max pooling layer are 6, 6, and 4 respectively, and the pooling strides are 3, 3, and 4.
[0045] Table 1 Corresponding table of sleep stages, waveforms, and frequencies
[0046]
[0047] (3) Attention-aware dimensionality reduction. After multi-modal and multi-scale feature extraction, the features f EMG , f EEG , f EOG obtained in (1) and (2) are used as inputs, and the output feature attn is obtained through the attention-aware dimensionality reduction module. First, dimensionality reduction is performed on f EOG extracted after short-time Fourier transform, and then feature concatenation is performed to obtain the fused feature feature multi , and the formula is as follows:
[0048] feature multi = concat(f EMG , f EEG , Reduction(f EOG ))
[0049] Among them, Reduction(·) is a dimensionality reduction function. As the number of features after splicing continuously increases, the training time of the model will gradually increase. To improve the model training speed, the present invention uses a residual SE attention module to reduce the feature dimension and remove redundant features. The formula is as follows:
[0050] feature attn = DownSample(feature multi )
[0051] Among them, DownSample(·) represents the residual SE attention function. The input fused feature feature multi , after passing through a one-dimensional convolutional layer with a convolution kernel of 1, batch normalization, and a ReLU activation function, and then passing through a one-dimensional convolutional layer with a convolution kernel of 1 and batch normalization. Then it passes through the SE layer. The SE layer sets a global average pooling (Global Average Pooling, GAP) to obtain the pointing information of different features, and sets two linear layers to learn the value weights of different features. The activation functions of the two linear layers are ReLU and Sigmoid respectively. Finally, through the downsampling layer, the downsampling layer takes the fused feature feature multi as the input, passes through a one-dimensional convolutional layer with a convolution kernel of 1 and a batch normalization layer, adds the input obtained in this step to the output obtained in the previous step, and finally passes through a ReLU activation function to obtain the key feature feature attn .
[0052] (4) Temporal feature learning module. The key feature feature attn obtained in (3) is input, and the final feature feature is output through the temporal feature learning module. Sleep data is a continuous time series. A double-layer Bi-LSTM is used to learn the transition rules of sleep staging. The formula is as follows:
[0053] feature = Temporal(feature attn )
[0054] Among them, Temporal(·) represents the double-layer Bi-LSTM function. The key feature is input, passes through a Bi-LSTM layer with 2 layers and a dropout of 0.5, and outputs the final feature feature. The dimension of feature is (N, 30, 56).
[0055] Step 4, comparison and classification
[0056] From step 3, the representative feature feature is obtained. Taking the feature feature as the input, to improve the problem of easy confusion in the staging of the transition phase, contrast learning is used to expand the sleep staging boundary, maximize the consistency between similar instances, and at the same time encourage the differences between dissimilar instances. Figure 5 It is the effect diagram of contrast learning. Figure 5 In the first figure, it is the distribution of the final feature feature without using contrast learning. Figure 5 In the second figure, it is the distribution of the final feature feature after using contrast learning. It can be seen that after using contrast learning, the features of the same stage are more concentrated, and the boundaries of features in different stages are clearer. The goal of contrast learning is to reduce the contrast loss, and the calculation formula of the contrast loss is as follows:
[0057]
[0058]
[0059]
[0060] Among them, N is the total number of samples, and feature i represents the feature of the i-th sample. is the number of samples with the same label as y i , and l ij is the similarity loss between the feature feature i and the feature feature j . represents the contrast loss of the feature feature i , and l con represents the contrast loss function. sim(feature i , feature k ) represents the cosine similarity between the sample feature i and the feature j . exp(·) represents the exponential function with the natural constant e as the base, bool(·) represents a function that returns a boolean value, and τ is the hyperparameter contrast temperature. In the present invention, τ = 0.05 is taken.
[0061] At the end of the entire model framework, the present invention sets a linear classification layer with 5 neurons according to the international standard AASM sleep staging criteria, and classifies the extracted features into 5 types of sleep stages, namely wakefulness (W), rapid eye movement (REM), non-rapid eye movement stage 1 (N1), non-rapid eye movement stage 2 (N2), and non-rapid eye movement stage 3 (N3).
[0062] Training and optimizing the network
[0063] During the experiment, the dataset was divided into 20 groups according to the subjects, and leave-one-out cross-validation was adopted. Each time, one group of subjects was selected as the test data X test , and the remaining 19 groups were used as the training data X. In the training data X, the data was divided into 10 groups, and one group was selected as the validation set X valid , and the remaining 9 groups were used as the training set X train to train the model. Finally, the prediction results of all 20 groups were combined to calculate various performance metrics.
[0064] During the training and optimization process of the network, a loss function combining classification loss and contrast loss was used. The total loss function is as follows:
[0065] L = l cls + w con * l con
[0066] where L is the overall loss function, l cls represents the classification loss function, and w con is the weight of the contrast loss function.
[0067] The weighted cross-entropy loss function was used as the classification loss function. Considering the large gap in the number of samples in each sleep stage and the different levels of difficulty in distinguishing them, a class weight was added to each class when calculating the cross-entropy loss. The classification loss function l cls is defined as follows:
[0068]
[0069] where N is the total number of samples, N c is the number of samples in class c. In the present invention, N c = 5. When the true class of sample x i is equal to c take 1, otherwise take 0, represents the predicted probability that sample x i belongs to class c. w c represents the weight of class c, which is given by the sample size of this class and the level of difficulty in distinguishing. Since the number of N1-stage samples is small and difficult to distinguish, w1 = 7 is set. The N3 stage is in the deep sleep stage without eye movement and muscle movement and is in the low-frequency stage, which is easy to distinguish, so w3 = 2 is set. The W stage, N2 stage, and REM stage have similar levels of difficulty in distinguishing, and w0, w2, and w4 are set to 3, 3, and 4 respectively according to the data volume.
[0070] The network is trained in batches. The batch size is set to 256, and the network is set to be trained for 100 epochs. The learning rate of the Adam optimizer starts from 1e-3 and is reduced to 1e-4 after 10 epochs. The parameter settings of Adam are as follows: the weight decay is set to 1e-3, betas (b1, b2) are (0.9, 0.999) respectively, and amsgrad is set to true. All convolutional layers are initialized using a Gaussian distribution with a mean of 0 and a variance of 0.02.
[0071] The parameter settings of the entire method framework are shown in Table 3. Among them, the convolutional module ConvBlock included in multiple modules is extracted, and its definition and parameters are shown in Table 2.
[0072] Table 2 Definition and parameters of ConvBlock
[0073]
[0074] Table 3 Parameter settings of the network framework of the method of the present invention
[0075]
[0076]
[0077] Performance evaluation of the network
[0078] The present invention uses four evaluation indicators: overall accuracy, N1-stage accuracy, macro F1-score, and Cohen's Kappa coefficient. These evaluation indicators are all calculated from true positives (TP), true negatives (TN), false positives (FP), and false negatives (FN).
[0079] Automated sleep staging aims to assist sleep staging, and accuracy is the most intuitive criterion for the credibility of the model.
[0080]
[0081] The main research task of the present invention is to improve the classification performance of the model for the N1 stage. Therefore, it is essential to observe the accuracy of the N1 stage.
[0082]
[0083] The sleep dataset is a class-imbalanced dataset, and the macro F1-score is an effective indicator commonly used to evaluate class-imbalanced data.
[0084]
[0085] On class-imbalanced datasets, the model is prone to favor the majority class and abandon the minority class. Therefore, Cohen's Kappa coefficient is used to penalize the "bias".
[0086]
[0087] Among them, P e refers to the sum of the "product of the actual label quantity and the predicted result quantity" for each class divided by the "square of the total number of samples". The two metrics of precision and recall are defined as
[0088]
[0089] Comparative analysis
[0090] Table 4 Comparison of the results of the method of the present invention and other methods on the Sleep-EDF-20 dataset
[0091]
[0092] In the comparative experiment, the publicly available dataset Sleep-EDF-20 was used to verify the performance of the model. The experimental results are shown in Table 4. The first column represents the data used, the second column represents the name of the sleep staging method, Accuracy represents the overall accuracy of sleep staging, Acc_N1 represents the classification accuracy of N1 stage, MF1-score represents the macro F1-socre, and Kappa represents Cohen's Kappa coefficient. The results of MMSC represent the average results of each evaluation index of the model on 20 subjects. The overall accuracy of MMSC is 92.9%, and the accuracy of N1 stage is 74.1%. Table 4 shows that compared with the multi-scale sleep staging methods DeepSleepNet, AttnSleep, SleepEEGNet using single modality and the sleep staging method ResnetnetLSTM that extracts multi-level features using a residual network while paying attention to transition rules, as well as the multi-task sleep staging method MultiTask CNN and the multi-granularity multi-channel sleep staging method Deep Fusion framework using multi-modality, the MSMC model of the present invention has better classification performance than several other methods due to its systematic problem processing process and powerful feature extraction module. In particular, the model of the present invention has achieved an improvement of about 19.6% - 49.2% in the classification of N1 stage, which indicates that MMSC has significant results in the classification of N1 stage and is suitable for quickly performing sleep staging.
Claims
1. A multi-modal and multi-scale sleep staging method for improving N1 stage category confusion, characterized in that, It includes the following steps: Step 1, preprocessing operation of the original signal First, select the Fpz-Cz channel in the electroencephalogram (EEG) signal, the electrooculogram (EOG) signal, and the electromyogram (EMG) signal as the experimental data, that is, the number of channels Q is taken as 3. Divide the sleep data with a sampling frequency of T into segments of 30 seconds, delete the MOVEMENT and UNKNOWN tags and the data corresponding to these tags, and merge the data of N3 and N4 stages and collectively call them the N3 stage; only retain the 30 minutes of W stage before and after the sleep time in the training set; after preprocessing, obtain N old samples, and each sample data x i is data with a length of 30×T, where i ∈ [1, 2, …, N old represents the serial number of the current sample. Assume that the obtained sample set is X old , and the label set is Y old ; Step 2, data augmentation For the training set samples obtained in Step 1, use the stacking strategy. For two adjacent N1-phase samples, take the middle part of the two samples to form a new sample; put the newly generated sample into the original data set X old to obtain a new sample set X, and accordingly update the label set Y old as the new label set Y; Step 3, multi-modal and multi-scale feature extraction A multi-modal and multi-scale feature extraction method is designed to make full use of sleep data and extract more representative features; for the new data set obtained in Step 2, feature extraction is performed; the data obtained in Step 2 is processed differently according to modalities, and multi-scale feature extraction is performed on the EEG signal by frequency; the features extracted from each modality are concatenated and fed into the SE attention module to select representative features and remove redundant features; finally, a double-layer Bi-LSTM is used to learn the transition rules of sleep staging to obtain the final features; The new sample set X after data augmentation is used as the input data and fed into the feature extraction module for multi-modal and multi-scale feature extraction; features The feature extraction process is divided into 4 steps: (1) Multi-modal feature extraction; sleep features are extracted from the sleep data X obtained in Step 2 according to modalities; human experts have different staging bases for different modalities of data. Specifically, sleep staging is judged according to the presence or absence of EMG signals, various characteristic waves in EEG signals, and the change speed of EOG signals; different modalities of the input data are fed into the corresponding modality feature extraction branches: EMG data is directly fed into the convolutional module for feature extraction, EEG data is fed into different convolutional kernel modules for multi-scale feature extraction, and EOG data is first subjected to short-time Fourier transform and then fed into the convolutional module for feature extraction; the short-time Fourier transform is performed using a Hanning window with a window size of 100 to extract the time-frequency features of the EOG signal; (2) Multi-scale feature extraction; using the EEG signal data X in (1) EEG as the input, select the convolutional kernel sizes of different scales according to the characteristic waves so as to extract the multi-scale sleep features of the EEG signal and output the EEG signal feature f EEG , (3) Attention-aware dimensionality reduction; after multi-modal multi-scale feature extraction, the features f EMG , f EEG , f EOG obtained in (1) and (2) are used as the input, and the output feature attn is obtained through the attention-aware dimensionality reduction module; first, dimensionality reduction is performed on the f EOG extracted after short-time Fourier transform, and then feature splicing is performed to obtain the fused feature feature multi , (4) Temporal feature learning module; The key feature obtained from the input (3) attn , and the final feature is output through the temporal feature learning module; the sleep data is a continuous time series, and a two-layer Bi-LSTM is used to learn the transition rules of sleep staging. Step 4, comparison and classification After Step 3, representative features of sleep staging can be obtained. By means of contrastive learning, the boundaries of sleep staging are expanded, the similarity between the features of synchronous sleep data is enhanced, and the similarity between the features of asynchronous sleep data is reduced, so that it is easier to distinguish between different stages; at the same time, a classifier is used to obtain the prediction result; according to the international standard AASM sleep staging criteria, 5 types of sleep staging are performed, namely wakefulness (W), rapid eye movement (REM), non-rapid eye movement stage 1 (N1), non-rapid eye movement stage 2 (N2), and non-rapid eye movement stage 3 (N3).
2. The multimodal multi-scale sleep staging method for improving N1 stage category confusion according to claim 1, wherein The data augmentation described in Step 2 specifically includes: Using the stacking strategy, select X old Two consecutive samples of the same type in and Combine The data of the last v seconds of and the data of the first (30 - v) seconds of to form a complete new sleep data labeled In, v takes 15, indicating the i-th newly generated sleep data, indicating the label corresponding to the i-th newly generated sleep data; select all adjacent N1-stage samples in this way, generate new N1-stage samples, and put the newly generated M samples into the original dataset X old to obtain a new sample set X, and accordingly update its label set Y and the total number of samples N; the generation algorithm formula based on the stacking strategy is as follows: N = N old + M Among them, represents the new samples generated for class c, where c takes 1; O(·) represents the stacking function of the above process; M is the number of new samples generated for class c; N is the number of original segments N old The number of samples in the training set X after data augmentation to generate M new N1-phase samples, and the dimension of X is (N, Q, (30×T)).
3. A multi-modal multi-scale sleep staging method for improving N1 stage category confusion according to claim 1, characterized in that, The multi-modal and multi-scale feature extraction method described in Step 3 specifically includes: The new sample set X after data augmentation is used as the input data and fed into the feature extraction module for multi-modal and multi-scale feature extraction; the feature extraction process is divided into 4 steps: (1) Multi-modal feature extraction: Extract sleep features from the sleep data X obtained in Step 2 in different modalities. Human experts have different staging bases for data of different modalities. Specifically, sleep staging is judged according to the presence or absence of EMG signals, various characteristic waves in EEG signals, and the speed of change of EOG signals. For different modality data of the input data, they are sent into the corresponding modality feature extraction branches: EMG data is directly sent into the convolutional module for feature extraction, EEG data is sent into different convolutional kernel modules for multi-scale feature extraction, and EOG data is first subjected to short-time Fourier transform and then sent into the convolutional module for feature extraction. The short-time Fourier transform uses a Hanning window with a window size of 100 to extract the time-frequency features of the EOG signal. The feature extraction formulas for EMG and EOG signals are as follows: f EMG = E EMG (X EMG , s EMG ) Among them, f EMG and f EOG are the features extracted from the EMG signal and the EOG signal respectively; X EEG and X EMG are the EEG signal and the EMG signal data of the dataset X respectively, is the EOG signal after the short-time Fourier transform of the dataset X; s EMG and s EOG are the convolutional kernels selected for extracting the features of the EMG signal and the EOG signal respectively, where s EMG = 10, s EOG = (8, 8); E u (·) is a function for feature extraction using the convolutional module ConvBlock, where u ∈ [EMG, EOG, EEG1, EEG2, EEG3], and ConvBlock includes a convolutional layer, a batch normalization layer, a GELU activation function layer, a max pooling layer, and a Dropout layer with a parameter of 0.5 added to prevent overfitting; E EMG is the EMG convolutional module, which includes three one-dimensional ConvBlocks. The convolutional kernel sizes in each convolutional module are 10, 8, and 8 respectively, the convolutional strides are all 3, and the kernels of the max pooling layers are 6, 6, and 4 respectively, and the pooling strides are 3, 3, and 4 respectively; E EOG is the EOG convolutional module, which includes three two-dimensional ConvBlocks. The convolutional kernel sizes in each convolutional module are 8×8, 7×7, and 7×7 respectively, the convolutional strides are 3×3, 1×1, and 1×1 respectively, and the kernels of the max pooling layers are 6×6, 4×4, and 2×2 respectively, and the pooling strides are 3×3, 2×2, and 2×2 respectively; (2) Multi-scale feature extraction; using the EEG signal data X in (1) EEG as the input, and selecting the convolutional kernel sizes of different scales according to the characteristic waves so as to extract the multi-scale sleep features of the EEG signal and output the EEG signal feature f EEG , where v ∈ [1, 2, 3]; the characteristic waves in the EEG signal mainly include α waves, β waves, δ waves, θ waves, K complex waves and Spindle waves; different sleep stages contain brain waves of different frequencies. The W stage contains α waves and β waves, the N1 stage contains θ waves, the N2 stage contains spindle waves and K complex waves, the N3 stage contains δ waves, and the REM stage contains α waves, β waves and θ waves; among them, the α wave corresponds to a frequency of 8 - 11 Hz, the β wave corresponds to a frequency of 11 - 30 Hz, the θ wave corresponds to a frequency of 4 - 8 Hz, the δ wave corresponds to a frequency of 0.5 - 3 Hz, the K complex wave corresponds to a frequency of 0.5 - 1 Hz, and the Spindle wave corresponds to a frequency of 12 - 14 Hz; the corresponding relationship between sleep stages and waveforms and frequencies is shown in Table 1; 1 Hz is selected to extract the features in the frequency range corresponding to the K complex wave to distinguish the N2 and N3 stages from other stages; 2 Hz is selected to extract the features in the frequency range corresponding to the δ wave to distinguish the N2 stage and the N3 stage; 10 Hz is selected to extract the features in the frequency range corresponding to the θ wave to distinguish the N1 stage from the W stage and the REM stage; the EEG signal feature extraction formula is as follows: Among them, concat(·) represents the concatenation function, is the convolution kernel of different scales for extracting EEG signal features; in, the convolution kernel is obtained by dividing the sampling frequency T by the frequency, that is E EEG1 (·) is the EEG convolution module with as the convolution kernel, which includes three one-dimensional ConvBlocks. The convolution kernel sizes in each convolution module are 100, 7, and 7 respectively, the convolution strides are 12, 1, and 1 respectively, the kernels of the max pooling layer are 4, 6, and 4 respectively, and the pooling strides are 2, 3, and 4; E EEG2 (·) is the EEG convolution module with as the convolution kernel, which includes three one-dimensional ConvBlocks. The convolution kernel sizes in each convolution module are 50, 7, and 7 respectively, the convolution strides are 6, 2, and 1 respectively, the kernels of the max pooling layer are 6, 6, and 4 respectively, and the pooling strides are 3, 3, and 4; E EEG3 (·) is the EEG convolution module with as the convolution kernel, which includes three one-dimensional ConvBlocks. The convolution kernel sizes in each convolution module are 10, 7, and 7 respectively, the convolution strides are 4, 3, and 3 respectively, the kernels of the max pooling layer are 6, 6, and 4 respectively, and the pooling strides are 3, 3, and 4; The correspondence between sleep staging and waveforms, frequencies is as follows: (3) Attention perception dimensionality reduction; After multi-modal and multi-scale feature extraction, the features f obtained from (1) and (2) EMG and f EEG and f EOG are used as inputs, and the output feature attn is obtained through the attention perception dimensionality reduction module; First, dimensionality reduction is performed on f EOG extracted after short-time Fourier transform, and then feature splicing is performed to obtain the fused feature feature multi . The formula is as follows: feature multi = concat(f EMG , f EEG , Reduction(f EOG )) Among them, Reduction(·) is a dimensionality reduction function; as the number of concatenated features increases, the training time of the model will gradually increase; the residual SE attention module is used to reduce the feature dimension, remove redundant features, and improve the model training speed. The formula is as follows: feature attn = DownSample(feature multi ) Among them, DownSample(·) represents the residual SE attention function; the input fusion feature feature multi , passes through a one-dimensional convolutional layer with a kernel size of 1, batch normalization, and the ReLU activation function, and then passes through a one-dimensional convolutional layer with a kernel size of 1 and batch normalization; then passes through the SE layer. The SE layer sets a global average pooling (Global Average Pooling, GAP) to obtain the pointing information of different features, and sets two linear layers to learn the value weights of different features. The activation functions of the two linear layers are ReLU and Sigmoid respectively; finally, passes through the downsampling layer. The downsampling layer uses the fusion feature feature multi as the input, passes through a one-dimensional convolutional layer with a kernel size of 1 and a batch normalization layer, adds the input obtained in this step to the output obtained in the previous step, and finally passes through the ReLU activation function to obtain the key feature feature attn ; (4) Temporal Feature Learning Module; the key feature "feature" obtained from the input (3) attn , and the final feature "feature" is output after passing through the temporal feature learning module; the sleep data is a continuous time series, and a two-layer Bi-LSTM is used to learn the transition rules of sleep staging. The formula is as follows: feature=Temporal(feature attn ) Among them, Temporal(·) represents a double-layer Bi-LSTM function. The key features are input, passed through a double-layer Bi-LSTM layer with 2 layers and a dropout of 0.5, and the final feature feature is output. The dimension of feature is (N, 30, 56).
4. A multi-modal and multi-scale sleep staging method for improving N1 stage category confusion according to claim 1, characterized in that, The comparison and classification method described in Step 4 includes the following steps: From Step 3, the representative feature feature is obtained. Taking the feature feature as the input, to improve the problem that the staging in the transition stage is easily confused, contrastive learning is used to expand the sleep staging boundary, maximize the consistency between similar instances, and at the same time encourage the differences between non-similar instances. The contrastive learning objective is to reduce the contrastive loss. The contrastive loss calculation formula is as follows: Among them, N is the total number of samples, and feature i represents the feature of the i-th sample, and N yi is the number of samples with the same label as y i , and l i,j is the similarity loss between the feature feature i and the feature feature j . represents the contrast loss of the feature feature i , and l con represents the contrast loss function; sim(feature i , feature k ) represents the cosine similarity between the sample feature i and feature j . exp(·) represents the exponential function with the natural constant e as the base, bool(·) represents a function that returns a boolean value, τ is the hyperparameter contrast temperature, and in this paper, τ = 0.05; At the end of the entire model framework, a linear classification layer with 5 neurons is set according to the international standard AASM sleep staging criteria, and the extracted features are subjected to 5-class sleep staging, namely wakefulness (W), rapid eye movement (REM), non-rapid eye movement stage 1 (N1), non-rapid eye movement stage 2 (N2), and non-rapid eye movement stage 3 (N3).