Sleep stage identification method and system based on multi-channel parallel time-frequency segmentation

Through the multi-channel parallel time-frequency segmentation method, combined with the multi-time frame self-attention mechanism and grouping convolution, multi-domain feature fusion is achieved, solving the problem of failure to effectively identify the sleep stage in the prior art, and significantly improving the accuracy of the recognition.

CN119942221APending Publication Date: 2025-05-06SHENYANG AEROSPACE UNIVERSITY
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510107289.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-23
Publication Date
2025-05-06

AI Technical Summary

Technical Problem

When identifying the sleep stage, the prior art fails to effectively take into account the synergy of multiple channels, and the feature fusion method only focuses on the channel dimension, resulting in insufficient prominence of key features and failure to take into account both time-domain and frequency-domain features.

Method used

The sleep stage recognition method based on multi-channel parallel time-frequency segmentation is adopted to generate brain maps through short-time Fourier transform, split time frames and frequency frames, and capture timing and spatial features using multi-time frame self-attention mechanism and grouping convolution, and realize multi-domain image feature fusion through channel and point attention mechanisms.

Benefits of technology

Effectively extracting time domain features, frequency domain features and spatial features significantly improves the ability to capture complex sleep signal features and improves the accuracy of sleep stage recognition.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119942221A_ABST
    Figure CN119942221A_ABST
Patent Text Reader

Abstract

The invention provides a sleep stage recognition method based on parallel time-frequency segmentation of multiple channels. The sleep stage recognition method comprises the following steps: S1, converting data of multiple channels into a brain map through short-time Fourier transform (STFT); s2, segmenting the time frames of the brain map into learning multi-time-frame inter-vector features; s3, capturing spatial features of sampling points of the brain map through grouping convolution; s4, segmenting the frequency frames of the brain map into learning multi-frequency frame inter-vector features; s5, focusing on channel dimensions and point-by-point dimensions of the time sequence vector features, the sampling point space features and the frequency vector features to realize multi-domain image feature fusion; and S6, classifying the images after feature fusion, and identifying different sleep stages. According to the method, the characteristics among the multiple time frame vectors, the characteristics among the multiple frequency frame vectors and the sampling point spatial characteristics are processed at the same time, so that the capability of capturing the complex sleep signal characteristics is remarkably improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of biological signal processing technology, and specifically relates to a sleep stage recognition method and system based on multi-channel parallel time-frequency segmentation. Background Art

[0002] Sleep is a basic physiological activity of human beings, which is characterized by a series of changes in brain, muscle, eye, heart and respiratory activities. This positive and standardized change plays a vital role in physical and mental health. Sleep disorders are prone to cause medical diseases such as obesity, cardiovascular disease, and neuropsychiatric diseases. They are also risk factors for cognitive decline, neurodegenerative diseases, mood changes, depression, and other neuropsychiatric diseases. The key to the diagnosis and treatment of sleep disorders is to accurately identify the sleep stage. In medical research, sleep is divided into five stages, namely wakefulness (W), first non-rapid eye movement sleep stage (N1), second non-rapid eye movement sleep stage (N2), third non-rapid eye movement sleep stage (N3) and rapid eye movement sleep (REM). Each stage is interconnected, alternates periodically during sleep, and is responsible for different physiological functions. The results of sleep stage analysis can be used to calculate macroscopic sleep structure parameters, such as sleep cycle, sleep latency, etc., which are of great significance for exploring various sleep disorders, from common ones such as obstructive sleep apnea to rare ones such as narcolepsy.

[0003] EEG signals and EOG signals are physiological signals directly generated by the human nervous system. They reflect a person's sleep condition by recording the electrophysiological activity of brain nerve cells in the cerebral cortex and the potential difference caused by eye movement. In clinical scenarios, physiological signals are usually collected during sleep. These signals can be fed back to experts or models to identify sleep stages. However, identification by experts relies on human subjectivity and is time-consuming and labor-intensive. Identification by models is not affected by human subjectivity and is more efficient. Therefore, how to use models to accurately identify sleep stages has become a technical problem that needs to be solved. Summary of the invention

[0004] In view of this, an object of the present invention is to provide a sleep stage recognition method and system based on multi-channel parallel time-frequency segmentation, so as to achieve accurate recognition of sleep stages by using hybrid model learning.

[0005] The technical solution of the present invention is: a sleep stage recognition method based on multi-channel parallel time-frequency segmentation, comprising:

[0006] S1: Convert the data of multiple channels into brain map through short-time Fourier transform (STFT);

[0007] S2: Divide the time frames of the brain map into learning features between vectors of multiple time frames;

[0008] S3: Capturing the spatial features of the sampling points of the brain map through grouped convolution;

[0009] S4: Split the frequency frames of the brain map into vector features for learning multiple frequency frames;

[0010] S5: Focus on the channel dimension and point-by-point dimension of the time series vector features, sampling point spatial features, and frequency vector features to achieve multi-domain image feature fusion;

[0011] S6: Classify the image after feature fusion to identify different sleep stages.

[0012] Preferably, S1 uses a short-time Fourier multimodal brain map generation module based on enhanced time-frequency image features to process EEG and EOG data into images, including:

[0013] For each 30-second period, use a 2-second Hamming window with a 50% overlap.

[0014] Select a 256-point short-time Fourier transform (STFT), take the first 128 points of the positive frequency component, calculate the absolute value of the component, obtain the amplitude information of 128 frequency channels, and exclude the DC component at 0 frequency;

[0015] Apply logarithmic transformation to the magnitude information;

[0016] The data were globally normalized to the mean and standard deviation of the entire dataset to adjust the data to zero mean and unit variance.

[0017] Preferably, in S2, the time frames of the time-frequency image are segmented, and the association between the time frame vectors is learned through a multi-time frame self-attention mechanism:

[0018] The time dimension of the time-frequency image is divided into multiple time frame vectors, and the time-frequency image is converted into multiple time frames, each of which represents the characteristics of multiple frequency segments contained in a specific time point;

[0019] After each time frame is transformed into a dimension through a fully connected layer, it is sent into a multi-time frame self-attention mechanism for multi-head attention operation. When all channels are calculated, the features are integrated through a fully connected layer to obtain time domain features that integrate multiple channels.

[0020] Preferably, grouped convolution is used in S3 to learn the spatial features of sleep time-frequency images:

[0021] Grouped convolution is applied to the spatial branch, combining the local pattern attention ability of convolution with spatial hierarchical information learning to extract cross-modal spatial features of sampling points from time-frequency images.

[0022] Preferably, in S4, the frequency branch is used to segment the frequency frames of the sleep time-frequency image, and the association between these frequency frames is learned through a multi-frequency frame self-attention mechanism:

[0023] The time-frequency image is segmented along the frequency dimension, and the time-frequency image is converted into multiple frequency frame vectors, each of which represents the characteristics of multiple time points contained in a specific frequency value;

[0024] Learn the relationship between frequency frames through the multi-head attention mechanism in the multi-frequency frame self-attention mechanism;

[0025] When all channels are calculated, the features are integrated through the fully connected layer to obtain the frequency vector features that combine multiple channels.

[0026] Preferably, the multi-dimensional image feature fusion method applied in S5 combines channel attention and point attention:

[0027] The channel attention expands the number of channels of the input feature B through a convolutional layer to obtain an intermediate result, then reduces the dimension of the result through an average pooling layer, extracts features through another convolutional layer, and finally compresses the channel back to the original dimension D through a fully connected layer to generate a channel weight matrix;

[0028] The point attention captures point-by-point features through a convolutional network and outputs a point weight matrix containing the weight value of each point:

[0029] Perform matrix multiplication on the channel weight matrix and the point weight matrix to obtain the final weight matrix:

[0030] W = Sigmoid(A c ×A p )

[0031] Split the weight matrix W and match each part with the corresponding feature map:

[0032] w1,w2,w3=Split(W)

[0033] out=w1×B t +w2×B f +w3×B g

[0034] Among them, B t is the output eigenvalue of the time branch, B f is the output eigenvalue of the frequency branch, B g is the output eigenvalue of the spatial branch.

[0035] Preferably, the different sleep stages in S6 are divided into five stages: wakefulness (W), first non-rapid eye movement sleep stage (N1), second non-rapid eye movement sleep stage (N2), third non-rapid eye movement sleep stage (N3) and rapid eye movement sleep (REM).

[0036] The present invention also provides a sleep stage recognition system based on multi-channel parallel time-frequency segmentation, comprising:

[0037] Short-time Fourier transform (SFT) multimodal brain map generation module based on enhanced time-frequency image features: short-time Fourier transform (SFT) is used to process EEG and EOG data into images;

[0038] Time series vector feature analysis module based on multi-time frame self-attention mechanism: divide the time frames of the time-frequency image and learn the association between these time frame vectors through the multi-time frame self-attention mechanism;

[0039] Sampling point spatial feature analysis module based on grouped convolution: grouped convolution is used to learn the spatial features of sleep time-frequency images;

[0040] Frequency vector feature analysis module based on multi-frequency frame self-attention mechanism: The frequency frames of the sleep time-frequency image are segmented through frequency branches, and the association between these frequency frames is learned through the multi-frequency frame self-attention mechanism;

[0041] Multi-dimensional image feature fusion module based on channel and point attention mechanism: multi-domain image feature fusion is performed on the channel dimension and point-by-point dimension of the time branch, frequency branch, and space branch features;

[0042] Classification module: The feature image integrated by the feature fusion module is classified by the classification module.

[0043] The present invention provides a sleep stage recognition method and system based on multi-channel parallel time-frequency segmentation. The method realizes the comprehensive extraction of multi-dimensional features. The method simultaneously processes the features between multi-time frame vectors, multi-frequency frame vectors, and sampling point spatial features. Compared with the traditional method that only focuses on time information, the method effectively makes up for the shortcomings of the prior art in the collaborative processing of multi-domain features, thereby significantly improving the ability to capture complex sleep signal features. The present invention also introduces channel attention and point attention mechanisms, focusing on channel level and point-by-point information respectively. The dynamic weight adjustment mechanism enables the model to adaptively assign weights according to the input data, highlighting key information and weakening irrelevant interference. This module significantly enhances the ability to integrate features and comprehensively optimizes the feature expression effect.

[0044] The network framework of the present invention maintains the consistency of the EEG channel and the EOG channel, and does not perform different processing for different channels. This design of the same processing method for any channel allows the model framework to receive data from any number of channels. In home scenarios, single-channel data is often used for comfort, while in clinical scenarios, multi-channel data is often used for accuracy. The network framework of the present invention is applicable to any of the above scenarios. BRIEF DESCRIPTION OF THE DRAWINGS

[0045] The accompanying drawings, which are incorporated in and constitute a part of this specification, illustrate embodiments consistent with the invention and, together with the description, serve to explain the principles of the invention.

[0046] Figure 1 The overall flow chart of sleep stage recognition based on multi-channel parallel time-frequency segmentation provided by the present invention;

[0047] Figure 2 A schematic diagram of the structure of a multi-dimensional image feature fusion module provided by the present invention;

[0048] Figure 3 A schematic diagram for comparing ablation experiment results provided by the present invention;

[0049] Figure 4 This is the sleep diagram of the subject numbered SC4061E0 in SleepEDF-20 provided by the present invention. DETAILED DESCRIPTION

[0050] Exemplary embodiments will be described in detail herein, examples of which are shown in the accompanying drawings. When the following description refers to the drawings, the same numbers in different drawings represent the same or similar elements unless otherwise indicated. The implementations described in the following exemplary embodiments do not represent all implementations consistent with the present invention. Instead, they are merely examples of systems consistent with some aspects of the present invention as detailed in the appended claims.

[0051] The English meanings involved in this implementation plan are as follows:

[0052] epoch: the number of iterations of the neural network

[0053] EEG: It is a medical technique used to measure the electrical activity of the brain. It is assessed by placing several electrodes (usually glued to the scalp) and recording the electrical signals generated from the brain.

[0054] EOG: is a medical technique used to measure eye movements. It is assessed by placing a pair of electrodes (usually on the eyelids and around the eye) and recording the electrical signals generated from the eye.

[0055] Python 3.8: A high-level programming language used to write computer programs. It is developed as an open source project led by Guido van Rossum and has broad community support.

[0056] NVIDIA RTX 4090 GPU: A high-end graphics processing unit (GPU) designed for demanding games and computer graphics applications.

[0057] Pytorch: An open source software library for data flow and differentiable programming, covering a wide range of tasks. It is mainly used for machine learning and deep learning, allowing researchers and developers to easily build and deploy models.

[0058] The automatic feature extraction and analysis of EEG signals will directly affect the accuracy of sleep stage recognition. The current automatic sleep stage recognition methods mainly have the following problems. First, they fail to take into account the synergy of multiple channels, so that the characteristics of different channels in the same sleep stage cannot be reflected. Secondly, most of them adopt a sequence-to-sequence mode, which puts a burden on the model. Then the feature fusion method only focuses on the channel dimension, resulting in insufficient prominence of key features. In addition, the current method fails to take into account time domain features and frequency domain features. Therefore, the present invention needs to use a fragment-to-fragment mode, considering which channels to coordinate and how to coordinate, how to take into account time domain features and frequency domain features, and how to effectively fuse different features. In addition, in order to ensure the integrity of the extracted features, the present invention also incorporates the spatial characteristics of the sampling points.

[0059] Since the present invention can effectively extract time domain features, frequency domain features and spatial features of sampling points, the constructed neural network itself also needs to have modules that are good at deep mining of information in time domain, frequency domain and spatial domain features. The present invention proposes combining multiple different types of network architectures, and finally adopts a fragment-to-fragment mode and uses a hybrid model learning method to achieve accurate recognition.

[0060] The sleep stage that the monitored person experiences in each 30-second time segment is automatically identified by monitoring and analyzing EEG and EOG signals. The sleep stages are determined by the American Academy of Sleep Medicine (AASM) and their characteristics are as follows:

[0061] Wake stage: The discharge activity of brain neurons occurs frequently, and the electroencephalogram shows rapid and variable characteristics.

[0062] Stage N1: Marking the transition from wakefulness to sleep, EEG changes begin to slow down. During this stage, the brain may produce high-amplitude alpha waves, with theta waves gradually appearing as the stage progresses.

[0063] Stage N2: An intermediate stage before deep sleep, characterized by prominent theta waves. In addition, this stage is characterized by rapid, rhythmic brain activity brought about by sleep spindles and K complexes, which are high-amplitude patterns that may be related to responses to environmental stimuli. EOG shows that eye movements completely stop in this stage.

[0064] Stage N3: is the last stage of the non-REM sleep stage and is characterized by delta waves, also known as slow wave activity, indicating the deepest sleep stage compared to stages one and two.

[0065] REM stage: A deeper sleep state than any non-REM sleep stage. During this stage, brain waves are fast and low amplitude, similar to the patterns seen during wakefulness, and are accompanied by rapid eye movements.

[0066] like Figure 1 As shown in FIG. 1 , the sleep stage recognition method based on multi-channel parallel time-frequency segmentation proposed in the present invention is as follows Figure 1 As shown, the system consists of six main parts: data processing, time branch, frequency branch, space branch, feature fusion and classification module. The first part converts the data of multiple channels into a brain map through short-time Fourier (STFT). The second part learns the features between multiple time frame vectors by dividing the time frame of the brain map. The third part captures the spatial features of the sampling points of the brain map through grouped convolution. The fourth part learns the features between multiple frequency frame vectors by dividing the frequency frame of the brain map. The fifth part realizes multi-domain image feature fusion by focusing on the channel dimension and point-by-point dimension of the three branch features. Finally, there is a classification module to identify different sleep stages. The present invention proposes to combine multiple different types of network architectures to achieve accurate recognition of sleep stages.

[0067] The present invention specifically comprises the following steps:

[0068] S1: Convert the data of multiple channels into brain map through short-time Fourier transform (STFT);

[0069] The EEG and EOG data were processed into images using short-time Fourier transform for subsequent segmentation into time and frequency frames. Short-time Fourier transform can map one-dimensional data to two dimensions, allowing the signal to unfold in the time-frequency dimension. It can show the characteristics of the signal more clearly. For each 30-second period, a Hamming window with a duration of 2 seconds was used, and a 50% overlap was set to capture local signal changes and reduce window effects. We chose a 256-point short-time Fourier transform (STFT), which provides sufficient frequency resolution for signal analysis. Considering that only the amplitude information of the frequency content is of interest, the first half of the fast Fourier transform (FFT) output is retained, which corresponds to the positive frequency component (for a 256-point FFT, the first 128 points are taken). Subsequently, we calculated the absolute values ​​of these components to obtain the amplitude information of 128 frequency channels, and intentionally excluded the direct current component (DC component) at 0 frequency. In order to enhance the frequency resolution and emphasize the low-frequency details, a logarithmic transformation was applied to the amplitude data. This step improves the visibility of low-frequency components in the spectrum. At the same time, the transformation also mitigates the impact of high-frequency noise, ensuring that the basic pattern of the signal is more clearly discernible. Finally, the data is globally normalized by the mean and standard deviation of the entire dataset to adjust the data to zero mean and unit variance.

[0070] S2: Divide the time frames of the brain map into learning features between vectors of multiple time frames;

[0071] The time frames of the time-frequency image are segmented, and the association between these time frame vectors is learned through the multi-time frame self-attention mechanism. Specifically, the time-frequency image is first segmented into multiple time frame vectors along the time dimension, and the time-frequency image is converted into multiple time frames, each of which represents the characteristics of multiple frequency bands contained in a specific time point. The multi-head attention mechanism requires that the dimension of each time frame is a multiple of the number of heads. Therefore, it is sent to the multi-time frame self-attention mechanism after the dimension is changed through the fully connected layer for the calculation of the multi-head attention mechanism. When all channels are calculated, their features are integrated through the fully connected layer to obtain the time domain features that integrate multiple channels.

[0072] S3: Capturing the spatial features of the sampling points of the brain map through grouped convolution;

[0073] Grouped convolution is used to learn the spatial features of sleep time-frequency images. Grouped convolution divides the convolution operation into multiple independent groups, each of which only performs convolution on some input channels instead of all input channels. This method can improve computational efficiency, enhance model expression capabilities, and parallelize computation. In the present invention, grouped convolution is used for spatial branches, combining the local pattern attention ability of convolution with spatial hierarchical information learning, to extract cross-modal sampling point spatial features from time-frequency images, so as to more comprehensively understand and learn complex sleep patterns.

[0074] S4: Split the frequency frames of the brain map into vector features for learning multiple frequency frames;

[0075] The frequency branch in the present invention divides the frequency frames of the sleep time-frequency image, and learns the association between these frequency frames through the multi-frequency frame self-attention mechanism. It is opposite to the division direction of the time branch. Specifically, the time-frequency image is first divided along the frequency dimension, and the time-frequency image is converted into multiple frequency frame vectors, each of which represents the features of multiple time points contained in a specific frequency value. Subsequently, the relationship between frequency frames is learned through the multi-head attention mechanism in the multi-frequency frame self-attention mechanism to capture frequency domain features. When all channels are calculated, their features are integrated through the fully connected layer to obtain frequency vector features that integrate multiple channels.

[0076] S5: Focus on the channel dimension and point-by-point dimension of the time series vector features, sampling point spatial features, and frequency vector features to achieve multi-domain image feature fusion;

[0077] The multi-dimensional image feature fusion method in this invention combines channel attention and point attention. Figure 2 As shown in the figure. Channel attention focuses on different channel parts, highlights important channels, and obtains a channel-dimensional weight matrix through the cooperation of pooling and convolution. Specifically, first, the number of channels of the input feature B is expanded through the convolution layer to obtain the intermediate result. Then, the result is reduced in dimension through the average pooling layer, and features are extracted through another convolution layer. Finally, the channel is compressed back to the original dimension D through the fully connected layer to generate the channel weight matrix.

[0078] Since focusing only on the channel level cannot fully capture key information. A pixel on the time-frequency image contains the features of a specific time-frequency point, and the importance of each pixel in identifying different sleep stages varies. Therefore, point attention is proposed to focus on each pixel, thereby effectively highlighting important features. Figure 2 As shown in (b), point attention captures point-by-point features through a convolutional network and outputs a point weight matrix containing the weight value of each point.

[0079] After obtaining the channel weight matrix and the point weight matrix, perform matrix multiplication on the two matrices to obtain the final weight matrix:

[0080] W = Sigmoid(A c ×A p )

[0081] Subsequently, the weight matrix W is split and each part is matched with the corresponding feature map to highlight the important features:

[0082] w1,w2,w3=Split(W)

[0083] out=w1×B t +w2×B f +w3×B g

[0084] Among them, B t is the output eigenvalue of the time branch, B f is the output eigenvalue of the frequency branch, B g is the output eigenvalue of the spatial branch.

[0085] S6: Classify the image after feature fusion to identify different sleep stages.

[0086] The classification module consists of four fully connected layers, which are designed to match the model output results with the number of sleep categories. Through the fully connected layers, the final sleep classification results are obtained.

[0087] The present invention realizes the processing of multi-time frame vector features, multi-frequency frame vector features, and sampling point space features at the same time. Figure 3 As shown in Figure 1, it is an ablation experiment on the SleepEDF-20 dataset. It can be observed that eliminating both the frequency branch and the spatial branch will significantly reduce the accuracy.

[0088] In order to enable a more thorough understanding of the present invention, the present invention will be further described in detail below in conjunction with the accompanying drawings and specific implementation methods.

[0089] Example 1

[0090] Step 1: Collect the original EEG signal and pre-process the EEG signal (this embodiment takes the small data set SleepEDF-20 data set as an example);

[0091] The preprocessing is: using a 2-second Hamming window and setting a 50% overlap, calculating a 256-point short-time Fourier transform (using a positive frequency quantity, namely taking the initial 128 points), and converting the original signal into a time-frequency image.

[0092] Step 2: To improve frequency resolution and emphasize low-frequency details, we log-transformed the amplitude data to ensure that the fundamental pattern of the signal is clearer.

[0093] Step 3: The data were globally normalized using the mean and standard deviation of the entire dataset to adjust the data to zero mean and unit variance.

[0094] Step 4: The processed time-frequency images of multiple channels are passed into the time branch, and divided into multiple time frames along the time dimension of the time-frequency images. The feature relationship between these time frame vectors is learned through the model in the time branch.

[0095] Step 5: The processed time-frequency images of multiple channels are passed into the frequency branch, and are divided into multiple frequency frames along the frequency dimension of the time-frequency images. The characteristic relationship between these frequency frame vectors is learned through the model in the frequency branch.

[0096] Step 6: The processed time-frequency images of multiple channels are passed to the spatial branch, and the spatial features of the sampling points of the signal data are extracted using grouped convolution.

[0097] Step 7: Pass the features learned by the time branch, frequency branch, and space branch into the feature fusion module to capture the importance of different channels in the channel attention and the importance of the feature points corresponding to each time-frequency image in the point attention.

[0098] Step 8: The feature image after integration by the feature fusion module is classified by the classification module, and the signal is classified into one of the five stages of sleep. The following is the experiment and evaluation:

[0099] The experiment uses the public dataset SleepEDF-20, which contains 39 sleep records from 20 participants over two nights, with the second night data of participant No. 13 missing. Each sleep record includes two EEG channels, one EOG channel, and other channels and event marker data. Both EEG and EOG signals are sampled at 100Hz. The specific sampling of this dataset is shown in Table 1.

[0100] Table 1: SleepEDF-20 dataset sampling

[0101]

[0102] The experimental environment is python 3.8, NVIDIA RTX 4090 GPU processor. The entire network is implemented using the Pytorch architecture. The dataset uses 10-fold cross validation, training for 200 epochs, the batch size is 64 samples, and the learning rate is 5e -6 .

[0103] In order to verify the effectiveness of the present invention, the sleep graph of the subject numbered SC4061E0 in the SleepEDF-20 dataset is taken as an example. Figure 4 , where (a) represents the predicted probability value of the model, from which we can see that the model has a high confidence in the category prediction. (b) represents the predicted value of the model, where the red area represents the prediction error, and (c) represents the true label of the data. From the comparison between (b) and (c), we can see that our model can accurately identify the sleep stage of an individual.

[0104] In order to verify the performance of the present invention, the model proposed by the present invention is compared with other existing different models. The comparison results are shown in Table 2 below.

[0105] Table 2: Performance comparison with other methods on SleepEDF-20

[0106]

[0107] The results show that the accuracy of the present invention on the SleepEDF-20 dataset, Cohen's kappa coefficient and MF1 evaluation index respectively achieved classification accuracy of 88.2%, 0.84 and 82.2. The present invention achieved high classification performance for different sleep stages, indicating that our method has a high degree of accuracy and is expected to be applied in clinical practice.

[0108] Table 3 Sampling of SHHS dataset

[0109]

[0110] Embodiment 2:

[0111] This example uses my proposed multi-task hybrid model based on parallel training on the large dataset SHHS:

[0112] Step 1: Resample the data of all channels to 100 Hz and then convert them into time-frequency images in the same way as SleepEDF-20.

[0113] Steps 2 to 8 are the same as those for SleepEDF-20.

[0114] The following experiments and evaluations are carried out:

[0115] The experiment used a public data set SHHS, and 329 subjects with normal sleep patterns were selected for analysis in the present invention. In each PSG record, two EEG channels and two EOG channels are included. The sampling rate of the EEG channel is 125Hz, and the sampling rate of the EOG channel is 50Hz. The specific sampling conditions of the data set are shown in Table 3.

[0116] In order to verify the performance of the present invention on the discrete data set SHHS, we compared it with the comparison method in Example 1. The comparison results are shown in Table 4 below.

[0117] Table 4 Performance comparison with other methods on SHHS

[0118]

[0119] The experimental results show that the accuracy, Cohen's kappa coefficient and MF1 evaluation indicators of the sleep network based on parallel time-frequency segmentation are 89.2, 0.85 and 81.4 respectively, which are significantly improved compared with other comparison methods, and effectively verify the effectiveness of identifying different sleep stages based on multi-branch sleep networks.

[0120] The above is only a preferred embodiment of the present invention. It should be pointed out that for ordinary technicians in this technical field, several improvements and modifications can be made without departing from the technical principles of the present invention. These changes and modifications should also be regarded as the scope of protection of the present invention.

Claims

1. A sleep stage recognition method based on multi-channel parallel time-frequency segmentation, characterized in that: include: S1: Convert the data of multiple channels into brain map through short-time Fourier transform (STFT); S2: Divide the time frames of the brain map into learning features between multi-time frame vectors; S3: Capturing the spatial features of the sampling points of the brain map through grouped convolution; S4: Split the frequency frames of the brain map into learning features between multi-frequency frame vectors; S5: Focus on the channel dimension and point-by-point dimension of the time series vector features, sampling point spatial features, and frequency vector features to achieve multi-domain image feature fusion; S6: Classify the image after feature fusion to identify different sleep stages.

2. The sleep stage recognition method based on multi-channel parallel time-frequency segmentation according to claim 1, characterized in that: S1 uses a short-time Fourier multimodal brain map generation module based on enhanced time-frequency image features to process EEG and EOG data into images, including: For each 30-second period, use a 2-second Hamming window with a 50% overlap. Select a 256-point short-time Fourier transform (STFT), take the first 128 points of the positive frequency component, calculate the absolute value of the component, obtain the amplitude information of 128 frequency channels, and exclude the DC component at 0 frequency; Apply logarithmic transformation to the magnitude information; The data were globally normalized to the mean and standard deviation of the entire dataset to adjust the data to zero mean and unit variance.

3. The sleep stage identification method based on multi-channel parallel time-frequency segmentation according to claim 1, characterized in that: In S2, the time frames of the time-frequency image are segmented, and the association between the time frame vectors is learned through the multi-time frame self-attention mechanism: The time dimension of the time-frequency image is divided into multiple time frame vectors, and the time-frequency image is converted into multiple time frames, each of which represents the characteristics of multiple frequency segments contained in a specific time point; After each time frame is transformed into a dimension through a fully connected layer, it is sent into a multi-time frame self-attention mechanism for multi-head attention operation. When all channels are calculated, the features are integrated through a fully connected layer to obtain time domain features that integrate multiple channels.

4. The sleep stage recognition method based on multi-channel parallel time-frequency segmentation according to claim 1, characterized in that: S3 uses grouped convolution to learn the spatial features of sleep time-frequency images: Grouped convolution is applied to the spatial branch, combining the local pattern attention ability of convolution with spatial hierarchical information learning to extract cross-modal spatial features of sampling points from time-frequency images.

5. The sleep stage identification method based on multi-channel parallel time-frequency segmentation according to claim 1, characterized in that: In S4, the frequency branch is used to segment the frequency frames of the sleep time-frequency image, and the association between these frequency frames is learned through the multi-frequency frame self-attention mechanism: The time-frequency image is segmented along the frequency dimension, and the time-frequency image is converted into multiple frequency frame vectors, each of which represents the characteristics of multiple time points contained in a specific frequency value; Learn the relationship between frequency frames through the multi-head attention mechanism in the multi-frequency frame self-attention mechanism; When all channels are calculated, the features are integrated through the fully connected layer to obtain the frequency vector features that combine multiple channels.

6. The sleep stage recognition method based on multi-channel parallel time-frequency segmentation according to claim 1, characterized in that, The multi-dimensional image feature fusion method used in S5 combines channel attention and point attention: The channel attention expands the number of channels of the input feature B through a convolutional layer to obtain an intermediate result, then reduces the dimension of the result through an average pooling layer, extracts features through another convolutional layer, and finally compresses the channel back to the original dimension D through a fully connected layer to generate a channel weight matrix; The point attention captures point-by-point features through a convolutional network and outputs a point weight matrix containing the weight value of each point: Perform matrix multiplication on the channel weight matrix and the point weight matrix to obtain the final weight matrix: W=Sigmoid(A c ×A p ) Split the weight matrix W and match each part with the corresponding feature map: w1,w2,w3=Split(W) out=w1×B t +w2×B f +w3×B g Among them, B t is the output eigenvalue of the time branch, B f is the output eigenvalue of the frequency branch, B g is the output eigenvalue of the spatial branch.

7. The sleep stage identification method based on multi-channel parallel time-frequency segmentation according to claim 1, characterized in that: The different sleep stages in S6 are divided into five stages: wakefulness (W), first non-rapid eye movement sleep stage (N1), second non-rapid eye movement sleep stage (N2), third non-rapid eye movement sleep stage (N3) and rapid eye movement sleep (REM).

8. A sleep stage recognition system based on multi-channel parallel time-frequency segmentation, characterized in that: include: Short-time Fourier transform (SFT) multimodal brain map generation module based on enhanced time-frequency image features: short-time Fourier transform (SFT) is used to process EEG and EOG data into images; Time series vector feature analysis module based on multi-time frame self-attention mechanism: divide the time frames of the time-frequency image and learn the association between these time frame vectors through the multi-time frame self-attention mechanism; Sampling point spatial feature analysis module based on grouped convolution: grouped convolution is used to learn the spatial features of sleep time-frequency images; Frequency vector feature analysis module based on multi-frequency frame self-attention mechanism: frequency frames of sleep time-frequency images are segmented through frequency branches, and the association between these frequency frames is learned through multi-frequency frame self-attention mechanism; Multi-dimensional image feature fusion module based on channel and point attention mechanism: multi-domain image feature fusion is performed on the channel dimension and point-by-point dimension of the time branch, frequency branch, and space branch features; Classification module: The feature image integrated by the feature fusion module is classified by the classification module.