EEG signal time-frequency domain denoising method based on noise attention mechanism
Through the time-frequency domain segmentation method and complex convolutional neural network based on the noise attention mechanism, the problem of poor denoising effect in the existing technology is solved, efficient artifact removal and signal retention are achieved, and the interpretability of the model is improved.
Patent Information
- Application Number
- CN202510306166.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-14
- Publication Date
- 2025-06-17
- Estimated Expiration
- 2045-03-14
AI Technical Summary
The existing EEG signal denoising method fails to effectively combine time-frequency domain information, resulting in poor denoising effect, poor interpretability of the model, and difficulty in explaining the distribution of artifacts in the frequency band.
The time-frequency domain segmentation method based on the noise attention mechanism is adopted, combined with the complex convolutional neural network, and the complex C3k2Unet model (DS-Natt-CCUNet model) of the two-stage-noise attention mechanism is used to accurately separate the artifacts in the time-frequency domain and retain useful signals.
It realizes the maximum removal of artifacts from the time-frequency domain perspective, while maximizing the original components of the EEG signal, improving the interpretability of the model, and significantly improving the noise removal effect in real scenarios.
Smart Images

Figure CN120162529A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of electroencephalogram signal denoising, and particularly relates to a time-frequency domain denoising method for electroencephalogram signals based on a noise attention mechanism. Background Technique
[0002] Electroencephalogram (EEG) signal is a non-invasive technique for recording the electric field of the brain through scalp electrodes, and its signal originates from the spatial summation of postsynaptic potentials of a large number of neuron populations. Due to its safety, high temporal resolution, and hypersensitivity to the dynamic changes of brain nerve signals, EEG signals have been widely used in fields such as medicine, psychology, neuroscience, and human-computer interaction. However, the low-amplitude characteristics of EEG signals make them vulnerable to interference from other physiological noises and environmental noises, and these interferences sometimes seriously affect. Artifacts can mask or distort the true electroencephalogram activities, making it difficult for researchers to accurately extract electroencephalogram wave features. For example, in the medical field, artifacts may interfere with doctors' diagnosis of epilepsy, brain injury, or other diseases; in EEG research based on machine learning or deep learning, artifacts will introduce noise and reduce the accuracy of the model; in real-time applications such as brain-computer interfaces, artifacts will cause system response delays or incorrect operations.
[0003] When collecting EEG signals, it is inevitably affected by various artifacts. The main artifacts in EEG signals are: electrocardiogram (ECG) signal, electromyogram (EMG) signal, electrooculogram (EOG) signal, and electrode motion (EM). Among them, ECG signal, EMG signal, and EOG signal are all physiological signals, which are physiological signals that are inevitably mixed in when patients collect EEG signals. EM signal is an artifact signal introduced due to the relative movement of scalp electrodes caused by the movement of the subject or the shaking of the lead wires during the collection process.
[0004] At present, there have also been some studies on artifact removal for EEG signals. In traditional methods, it mainly relies on filtering techniques, blind source separation techniques (Blind Source Separation, BSS), and signal decomposition techniques, or a combination of the three. Although these traditional methods show certain effects in specific scenarios, their inherent limitations significantly restrict practical applications. For example: The blind source separation (BSS) techniques (such as ICA, CCA) rely on the assumption of multi-channel data (number of channels ≥ number of sources), making it difficult to adapt to single-channel portable scenarios; Hybrid methods (such as EEMD-CCA, SSA-ICA) expand the single-channel applicability through signal decomposition, but parameter adjustment needs to be targeted at specific artifact types and lacks the ability to jointly process multiple artifacts; Filtering methods are prone to losing the effective EEG frequency band, while adaptive filtering (AF) and artifact subspace reconstruction (ASR) are limited by the difficulty of obtaining reference signals and parameter sensitivity. These defects have prompted researchers to turn to data-driven deep learning frameworks to achieve end-to-end adaptive denoising and improve the robustness in complex scenarios. In terms of deep learning techniques, some denoising neural network frameworks have also been proposed by scholars. More important studies in the past five years include: the NovelCL model based on CNN, the segmentation denoising network SDNet based on Resnet, the DuoCL model based on the combination of multi-scale CNN and LSTM, the EEGIFNet based on a dual-branch structure, and the EEGDNet based on the Transformer self-attention mechanism. For example, the patent with the publication number CN119157556A discloses an EEG signal denoising method based on a dual-path convolutional denoising network. It constructs a training set using EEG signal samples in a public dataset; constructs and trains a dual-path convolutional denoising network using the training set to obtain the trained dual-path convolutional denoising network; collects the original EEG signal and inputs it into the trained dual-path convolutional denoising network to obtain the denoised EEG signal. In the dual-path convolutional denoising network, the signal encoder encodes the original EEG signal into a feature vector; the mask generator converts the feature vector into a two-dimensional feature vector and performs local and global information extraction to generate a mask; the feature fuser multiplies the mask and the feature vector to obtain a denoised feature vector, and the signal decoder converts the denoised feature vector into a denoised EEG signal. Although it achieves a certain denoising effect in both single-artifact removal and mixed-artifact removal scenarios, there are still two major deficiencies: First, existing research focuses on processing the signal in the time domain and ignores the impact of artifacts on the frequency domain. Second, the existing deep learning model denoising research can be understood as using the powerful non-linear problem processing ability of the neural network to force the model of the EEG signal with artifacts to fit the clean EEG signal, and the interpretability of the model is poor. Due to the lack of comprehensive analysis of the time domain and frequency domain, the model is difficult to explain "which frequencies are corrected" or "how the artifacts are distributed in the frequency band", limiting the transparency and reliability of the denoising effect. Summary of the Invention
[0005] Aiming at the defects existing in the prior art, the purpose of the present invention is to provide a time-frequency domain denoising method for electroencephalogram (EEG) signals based on a noise attention mechanism, which combines the time-domain and frequency-domain information of EEG signals by a complex convolutional neural network, and innovatively proposes a noise attention mechanism based on time-frequency domain segmentation to more accurately separate artifacts in EEG signals and fully retain useful EEG signals, while improving the interpretability of the model, and solving the problem that the existing EEG signal denoising methods fail to consider denoising by combining the time-frequency domain.
[0006] To achieve the above object, the technical solution of the present invention is as follows:
[0007] A time-frequency domain denoising method for EEG signals based on a noise attention mechanism, comprising the following steps:
[0008] Step 1: Obtain an EEG signal with artifacts, and use the short-time Fourier transform to convert it into a time-frequency diagram for subsequent denoising from the perspective of combining the time-frequency domain;
[0009] Step 2: Use a one-dimensional Unet1D model to obtain a time mask of the EEG signal with artifacts, that is, the area of noise in the time domain; use the Fourier transform and a one-dimensional Unet1D model with the same structure to obtain a frequency mask of the EEG signal with artifacts, that is, the area of noise in the frequency domain; then combine the two to obtain a time-frequency mask of noise, that is, the approximate area of noise in the time-frequency domain;
[0010] Step 3: Input the time-frequency diagram of the EEG signal with artifacts obtained in Step 1 into the noise attention module of the dual-stage noise attention mechanism complex C3k2Unet model (DS-Natt-CCUNet model), and also input the time-frequency mask of noise obtained in Step 2 into the noise attention module to help the C3k2Unet model locate the area of noise. Finally, the C3k2Unet model outputs an accurate time-frequency diagram of noise;
[0011] Step 4: After obtaining the accurate time-frequency diagram of noise, input it into the noise attention module, and the input is still the time-frequency diagram of the EEG signal with artifacts. The DS-Natt-CCUNet model performs accurate denoising and outputs the time-frequency diagram of the clean EEG signal;
[0012] Step 5: Perform the inverse short-time Fourier transform on the time-frequency diagram of the clean EEG signal to finally obtain the clean EEG signal.
[0013] For the EEG signal with artifacts in Step 1, the window function used in the discrete short-time Fourier transform is the Hann window, and the specific formula of the discrete short-time Fourier transform is as follows:
[0014]
[0015] Among them, X[m, k] is the complex spectral coefficient at time frame m and frequency k, x[n] is the original discrete signal, w[n] is a window function of length L (such as Hann window), H is the frame shift (step size), N is the number of FFT points, j is the imaginary unit, and π is the pi; the original signal can be reconstructed through the inverse short-time Fourier transform (ISTFT), and its formula is:
[0016]
[0017] Among them, A is the normalization factor (usually A = Σ m ω 2 [n - mH]), w[n - mH] is the window function used in the inverse transform, and the position is determined by the offset of mH. Other symbols are consistent with the STFT formula.
[0018] The specific content of step two is: generate a time mask and a frequency mask through a parallel time-domain Unet1D network and frequency-domain Unet1D network respectively, and multiply the time mask and the frequency mask element by element to generate a time-frequency joint noise mask.
[0019] In the DS-Natt-CCUNet model in step three, the basic framework is Unet. Combine the complex-valued neural network (Complex-valued CNN) and the C3k2 feature extraction module in YOLOv11 to construct an encoder that can process the time-frequency diagram of EEG signals; combine CSPNet (Cross Stage Partial Network) and the spatial attention mechanism, and use the time-frequency mask / diagram of noise as the input to construct a noise attention mechanism.
[0020] The complex-valued neural network includes four parts:
[0021] (1) Complex Convolution layer (abbreviated as CConv)
[0022] The convolution kernel of the complex convolution layer consists of two sets of learnable parameters, the real part and the imaginary part. The input data is also represented in complex form. During the calculation process, the complex convolution follows the complex multiplication rule. The specific process is: the result of the convolution of the real part of the weight and the real part of the input minus the result of the convolution of the imaginary part of the weight and the imaginary part of the input is used as the real part of the output; the result of the convolution of the imaginary part of the weight and the real part of the input plus the result of the convolution of the imaginary part of the weight and the imaginary part of the input is used as the imaginary part of the output, so as to obtain the result of the complex convolution;
[0023] The detailed operation of the specific process is:
[0024] 1) Define the complex convolution kernel W = W real + iW imagand the complex input X = X real + X imag , where W real and W imag are the real and imaginary parts of the complex convolution kernel, and X real and X imag are the real and imaginary parts of the complex input;
[0025] 2) Calculate the real part of the output: where is the real part of the output, * is the convolution operation, and Conv(·) is the real convolution function;
[0026] 3) Calculate the imaginary part of the output: where is the imaginary part of the output;
[0027] (2) Complex Pooling Layer
[0028] The complex pooling layer is divided into two types: max pooling and average pooling. Similarly, it is applied to the complex domain. For max pooling: select the maximum value of the modulus of the complex numbers in the local area; for average pooling: calculate the average values of the real and imaginary parts of the complex numbers in the local area as the real and imaginary parts of the result. According to the actual performance of the two types of pooling, finally select complex max pooling as the method for feature dimensionality reduction;
[0029] (3) Complex Batch Normalization Layer (Complex BatchNormalization, abbreviated as CBN)
[0030] The complex batch normalization layer is an extension of batch normalization in the complex domain, aiming to handle the normalization problem of complex-valued data. By normalizing the statistical characteristics of small batches of data, the complex-valued feature distribution is kept stable, thereby accelerating training and improving model performance. Since complex data contains two dimensions: real and imaginary parts, the design of CBN needs to consider complex operations. Its specific calculation formula is as follows:
[0031]
[0032] where is the normalized complex value, x is the input complex-valued data, E[x] is the mean of the complex input x, V is the covariance matrix of the complex input, and its definition is shown in formula (6). To make V solvable, V must satisfy positive definiteness or semi-positive definiteness, and the condition for V to be semi-positive definite is the mean μ, covariance Γ, and pseudo-variance C, that is, as shown in formula (7):
[0033]
[0034] where V rris the covariance of the real part with the real part, V ri is the covariance of the real part with the imaginary part, V ir is the covariance of the imaginary part with the real part, V ii is the covariance of the imaginary part with the imaginary part, and Con(a, b) is the covariance calculation function.
[0035]
[0036] Finally, two learnable adjustment parameters γ and β are introduced to adjust the feature distribution, and their calculation formulas are as follows:
[0037]
[0038] Among them, is a positive semi - definite matrix containing four parameters, β is a complex vector, which is consistent with the dimension of the complex number. For the convenience of training and after batch normalization, satisfies the standard Gaussian distribution, and γ rr and γ ii are initialized to γ ri 、γ ir and the real and imaginary parts of β are all initialized to 0;
[0039] (4) Complex Activation Function Layer
[0040] The activation function used in the complex activation function layer is the ReLU activation function, which is defined as f(x) = max(0, x), that is, the positive part of the input value is retained, and the negative part is truncated to zero. Its calculation process in the complex domain is shown in formula (9), and the real part and the imaginary part of the input data are respectively processed using the ReLU activation function to obtain the real and imaginary parts of the result. Similarly, the complex Sigmoid activation function is shown in formula (10):
[0041]
[0042] Among them, CReLU(x) is the complex ReLU activation function, and ReLU(x) is the real ReLU activation function.
[0043]
[0044] Among them, CSigmoid(x) is the complex Sigmoid activation function, and Sigmoid(x) is the real Sigmoid activation function.
[0045] The method of using the noise attention module in the DS-Natt-CCUNet model in Step 3 or Step 4 to denoise and obtain a clean time-frequency map is specifically as follows:
[0046] The input complex time-frequency map first extracts preliminary complex features through CConv, and then is split (Split) along the channel dimension. Inheriting the architecture of CSPNet, the features are divided into two branches: Branch 1 directly retains the original complex features to maintain information integrity, and Branch 2 dynamically generates a weight map through a spatial attention mechanism. In the attention branch, the features undergo dual pooling operations of max pooling and average pooling to generate two types of spatial statistical descriptors; then the results of the dual pooling are fused through complex convolution (CConv), and combined with complex batch normalization (CBN) to dynamically adjust the feature distribution to adapt to different noise scenarios. Finally, the fused features are compressed into a spatial weight map within the range of [0,1] through complex Sigmoid (CSigmoid); this weight map is multiplied point by point with the original complex features of Branch 1, and finally the time-frequency map of the clean EEG signal is output.
[0047] The present invention also includes a system capable of running the above-mentioned method for denoising EEG signals in the time-frequency domain based on a noise attention mechanism.
[0048] The present invention also includes a device, comprising:
[0049] A memory: for storing a computer program for implementing the above-mentioned method for denoising EEG signals in the time-frequency domain based on a noise attention mechanism;
[0050] A processor: for implementing the above-mentioned method for denoising EEG signals in the time-frequency domain based on a noise attention mechanism when executing the computer program.
[0051] The present invention also includes a computer-readable storage medium storing a computer program, and the computer program implements the above-mentioned method for denoising EEG signals in the time-frequency domain based on a noise attention mechanism when executed by a processor.
[0052] Compared with the prior art, the beneficial effects of the present invention are:
[0053] 1. Steps 1 and 2 of the present invention are the preprocessing stage, adopting a time-domain - frequency-domain joint mask segmentation technology to achieve precise positioning of signal features through a multi-modal collaborative processing mechanism, and having the characteristics of dual-domain feature enhancement.
[0054] 2. Steps 3 and 4 of the present invention are the core processing stage, innovatively introducing the DS-Natt-CCUNet deep learning model to construct a noise attention architecture, which can realize noise feature extraction and pure EEG signal reconstruction, forming a unique dual-feature decoupling learning ability.
[0055] 3. Steps five and six of the present invention are the processing stages. Signal reconstruction is achieved through inverse time-frequency matrix transformation, and an acquisition parameter optimization scheme is generated based on the feature analysis results, having the dual application values of high signal restoration fidelity and strong clinical adaptability.
[0056] In summary, the present invention solves the problem that the existing EEG signal denoising methods fail to combine the time domain and the frequency domain simultaneously for comprehensive denoising of EEG signals, resulting in poor denoising effects and poor model interpretability. At the same time, it realizes that it can not only remove artifacts to the greatest extent from the time-frequency domain perspective, but also retain the original components of EEG signals to the greatest extent. In the real scenario (mixed myoelectric and electrocardiogram artifacts), from the perspective of the time-frequency diagram, the present invention is closer to the time-frequency diagram of clean EEG signals, and effective denoising in both the time domain and the frequency domain is achieved. In addition, the type of noise mixed into the original electroencephalogram signal can be analyzed based on the time-frequency diagram of the noise, thereby helping electroencephalogram signal acquisition personnel prevent the mixing of noise from the source. Therefore, the present invention can not only effectively remove various artifacts in EEG signals, but also put forward some specific suggestions for reducing the mixing of artifacts during the actual acquisition of electroencephalogram signals. Description of the Drawings
[0057] Figure 1 It is the overall flowchart of the time-frequency domain denoising method for electroencephalogram signals based on the noise attention mechanism.
[0058] Figure 2 It is the specific module diagram of the dual-stage - noise attention mechanism complex C3k2Unet model (DS-Natt-CCUNet).
[0059] Figure 3 It is the example diagram of a single signal and an EEG signal with artifacts.
[0060] Figure 4 It is the average experimental results of this method and other methods when removing a single artifact.
[0061] Figure 5 It is the comparison diagram of the experimental results of this method and other methods when removing a single artifact.
[0062] Figure 6 It is the average experimental results of this method and other methods in the real scenario (removing mixed myoelectric and electrocardiogram artifacts).
[0063] Figure 7 It is the time domain comparison diagram of EEG signals of this method and other methods in the real scenario (removing mixed myoelectric and electrocardiogram artifacts).
[0064] Figure 8 It is the time-frequency diagram comparison diagram of EEG signals of this method and other methods in the real scenario (removing mixed myoelectric and electrocardiogram artifacts) (for convenient comparison, the values are logarithmized).
[0065] Figure 9 It is the label of the time-frequency diagram of noise, the comparison diagram of the noise time-frequency mask and the noise time-frequency diagram output by the model.
[0066] Figure 10 It is the result diagram of the ablation experiment of the DS-Natt-CCUNet model. Specific implementation manner
[0067] Next, the process and advantages of the present invention will be described in detail with reference to the accompanying drawings.
[0068] As Figure 1 shown, a method for denoising EEG signals in the time-frequency domain based on a noise attention mechanism includes the following steps:
[0069] Step 1: Obtain the EEG signal with artifacts and convert it into a time-frequency diagram using the short-time Fourier transform, so as to perform denoising from the perspective of the combination of the time-frequency domain in the subsequent steps;
[0070] Step 2: Use a one-dimensional Unet1D model to obtain the time mask of the EEG signal with artifacts, that is, the area of noise in the time domain; use the Fourier transform and a one-dimensional Unet1D model with the same structure to obtain the frequency mask of the EEG signal with artifacts, that is, the area of noise in the frequency domain; then combine the two to obtain the time-frequency mask of noise, that is, the approximate area of noise in the time-frequency domain;
[0071] Step 3: Input the time-frequency diagram of the EEG signal with artifacts obtained in Step 1 into the noise attention module in the dual-stage - noise attention mechanism complex C3k2Unet model (DS-Natt-CCUNet model), and input the time-frequency mask of noise obtained in Step 2 into the noise attention module as well, to help the C3k2Unet model locate the area of noise. Finally, the C3k2Unet model outputs an accurate noise time-frequency diagram;
[0072] Step 4: After obtaining the accurate noise time-frequency diagram, input it into the noise attention module again to assist the model in performing accurate denoising. The input of the DS-Natt-CCUNet model is still the time-frequency diagram of the EEG signal with artifacts. At this time, the model outputs the time-frequency diagram of the clean EEG signal;
[0073] Step 5: Perform the inverse short-time Fourier transform on the time-frequency diagram of the clean EEG signal to finally obtain the clean EEG signal.
[0074] Combined with the noise time-frequency diagram obtained in Step 3, the artifacts mixed in the original EEG signal can be analyzed, so as to give corresponding optimization suggestions for the acquisition process of the EEG signal. The optimization suggestions include:
[0075] (1) When EMG artifacts are detected, it is recommended to reduce facial muscle activity;
[0076] (2) When ECG artifacts are detected, it is recommended to adjust the position of the electrocardiogram interference electrode;
[0077] (3) When EOG artifacts are detected, it is recommended to reduce eye movements.
[0078] For the EEG signal with artifacts in Step 1, the sampling rate of the EEG signal is 256 Hz, the duration is 2 s, the window function used in the discrete short-time Fourier transform is the Hann window, the window size is 126, and the step size is 8. The specific formula for the discrete short-time Fourier transform is as follows:
[0079]
[0080] where x[n] is the discrete signal, w[n] is the window function of length L (such as the Hann window), m is the time frame index, H is the frame shift (step size), k is the frequency index, and N is the number of FFT points; the original signal can be reconstructed through the inverse short-time Fourier transform (ISTFT), and its formula is:
[0081]
[0082] where A is the normalization factor (usually A = Σ m ω 2 [n - mH]), w[n - mH] is the window function used in the inverse transform, and its position is determined by the offset of mH. Other symbols are consistent with the STFT formula.
[0083] The Unet1D model used in Step 2. The Unet1D network includes an encoder composed of a convolutional layer, a batch normalization layer, and a ReLU activation function, aiming to roughly locate the region where the noise is located in the time domain and frequency domain. This article is based on the classic Unet model and adapts and adjusts its structural parameters in combination with the characteristics such as the length of the one-dimensional EEG signal. The encoding layer is mainly composed of a convolutional layer, a BN layer, and an activation function (ReLU) layer.
[0084] The specific Step 2 is as follows: Generate a time mask and a frequency mask through a parallel time-domain Unet1D network and frequency-domain Unet1D network respectively, and multiply the time mask and the frequency mask element by element to generate a time-frequency joint noise mask;
[0085] The structures of the time-domain Unet1D network and the frequency-domain Unet1D network include:
[0086] (1) The encoder is composed of 5 levels of downsampling modules, and each level includes a one-dimensional convolutional layer, a batch normalization layer, and a ReLU activation function;
[0087] (2) The decoder consists of a 4-level upsampling module and uses transposed convolution to restore features;
[0088] (3) The skip connection is achieved through cross-layer feature concatenation.
[0089] For the DS-Natt-CCUNet model in step three, its model structure diagram is as shown in Figure 1 and Figure 2 . The basic framework of this model is still Unet. In order to enable it to process time-frequency map features (complex numbers, including real and imaginary parts) and improve its feature mining performance, in this paper, a complex-valued convolutional neural network (Complex-valued CNN) is combined with the C3k2 feature extraction module in YOLOv11 to construct an encoder that can efficiently process the time-frequency maps of EEG signals; CSPNet (CrossStage Partial Network) is combined with a spatial attention mechanism, and with the time-frequency mask / map of noise as the input, a noise attention mechanism is constructed.
[0090] The complex neural network module includes four parts:
[0091] (1) Complex Convolution layer, abbreviated as CConv
[0092] The complex convolution layer is an extended form of the real convolution layer. Its core lies in extending the convolutional kernel parameters and input data from the real number domain to the complex number domain, so as to simultaneously model the amplitude and phase information of the signal during the feature extraction process. Specifically, the convolutional kernel of the complex convolution layer consists of two sets of learnable parameters, the real part and the imaginary part. The input data is also represented in complex form. During the calculation process, the complex convolution follows the complex multiplication rule. Assuming the convolutional kernel parameter W = W real + iW imag , and the input data X = X real + X imag , the convolution process is as shown in formula (3):
[0093] W * X = (W real * X real - W imag * X imag ) + i(W imag * X real + W real * X imag ) (3)
[0094] Among them, W real and W imag are the real and imaginary parts of the complex convolutional kernel, and X real and X imagare the real and imaginary parts of a complex input, and * represents the convolution operation.
[0095] The above formula is transformed into matrix form as shown in formula (4):
[0096]
[0097] where is the real part of the output, is the imaginary part of the output.
[0098] It can be seen from the above two formulas that the operation process is as follows:
[0099] 1) Define the complex convolution kernel W = W real + iW imag and the complex input X = X real + X imag ;
[0100] 2) Calculate the real part of the output:
[0101] 3) Calculate the imaginary part of the output:
[0102] where Conv(·) represents the real convolution function.
[0103] The process of complex convolution is split into four parts of real convolution operations and finally pieced together into a complex form. The specific process is: the result of the convolution of the real part of the weight and the real part of the input minus the result of the convolution of the imaginary part of the weight and the imaginary part of the input is used as the real part of the output; the result of the convolution of the imaginary part of the weight and the real part of the input plus the result of the convolution of the imaginary part of the weight and the imaginary part of the input is used as the imaginary part of the output. In this way, we can obtain the result of complex convolution.
[0104] (2) Complex Pooling Layer
[0105] The complex pooling layer is also divided into two types: max pooling and average pooling. The operations of maximization and averaging are applied to the complex domain. For max pooling: select the maximum value of the modulus of the complex numbers in the local area; for average pooling: calculate the average values of the real and imaginary parts of the complex numbers in the local area as the real and imaginary parts of the result. According to the actual performance of the two poolings, finally select complex max pooling as the method for feature dimensionality reduction.
[0106] (3) Complex Batch Normalization Layer (abbreviated as CBN)
[0107] The complex batch normalization layer is an extension of batch normalization in the complex domain, aiming to handle the normalization problem of complex-valued data. Its core idea is similar to that of real-valued BN, that is, by normalizing the statistical characteristics of mini-batch data, the distribution of complex-valued features is kept stable, thus accelerating training and improving model performance. However, since complex data contains two dimensions, the real part and the imaginary part, the design of CBN needs to consider the special properties of complex operations, that is, simply relying on translation and scaling in the real number domain cannot ensure that the means and variances of the real and imaginary parts of the complex numbers are consistent, and the normalized input may show a large deviation. Complex domain normalization can be achieved by utilizing the overall mean of the input data x and the covariance matrix in the complex domain to effectively standardize the input, thus overcoming the above problems and improving model performance. The specific calculation formula is as follows:
[0108]
[0109] where, is the normalized complex value, x is the input complex-valued data, E[x] is the mean of the complex input x, V is the covariance matrix of the complex input, V is the covariance matrix of the complex input, and its definition is shown in formula (6). It should be noted that in order for V to have a solution, V must satisfy positive definiteness or semi-positive definiteness, and the condition for V to be semi-positive definite is the mean μ, covariance Γ, and pseudo-variance C, that is, as shown in formula (7):
[0110]
[0111] where, V rr is the covariance of the real part with the real part, V ri is the covariance of the real part with the imaginary part, V ir is the covariance of the imaginary part with the real part, V ii is the covariance of the imaginary part with the imaginary part, and Con(a, b) is the covariance calculation function.
[0112]
[0113] Finally, two learnable adjustment parameters γ and β are introduced to adjust the feature distribution, and the calculation formula is as follows:
[0114]
[0115] where, is a semi-positive definite matrix containing four parameters, β is a complex vector, which is consistent with the dimension of the complex number. For the convenience of training and the normalized to satisfy the standard Gaussian distribution, γ rr and γ ii are initialized as γ ri and γ irBoth the real part and the imaginary part of β are initialized to 0.
[0116] Through the above operations, we can obtain the result of complex batch normalization.
[0117] (4) Complex Activation Function Layer
[0118] The activation function is an important component that introduces non-linearity in neural networks. Its role is to perform non-linear transformations on the outputs of neurons, thereby endowing the model with the ability to express complex mapping relationships. Among many activation functions, ReLU (Rectified Linear Unit) has become one of the most widely used activation functions in modern deep learning due to its simplicity and efficiency. Using ReLU as the activation function, the definition of ReLU is f(x) = max(0, x), that is, the positive part of the input value is retained, and the negative part is truncated to zero. This design not only has high computational efficiency but also effectively alleviates the vanishing gradient problem, especially performing excellently in deep networks. Therefore, based on the ReLU activation function, this paper will use its expression in the complex domain. Its calculation process is shown in formula (9), and the real part and the imaginary part of the input data are respectively processed using the ReLU activation function to obtain the real part and the imaginary part of the result. Similarly, the CSigmoid activation function is shown in formula (10):
[0119]
[0120] where CReLU(x) is the complex ReLU activation function and ReLU(x) is the real ReLU activation function.
[0121]
[0122] where CSigmoid(x) is the complex Sigmoid activation function and Sigmoid(x) is the real Sigmoid activation function.
[0123] The method for denoising to obtain a clean time-frequency map in the noise attention module of the DS-Natt-CCUNet model in step three or step four is specifically as follows:
[0124] As Figure 2As shown, the noise attention module in this paper processes the complex-form noise time-frequency map (generated by STFT, etc.) based on modules such as complex convolution. Its core design combines the spatial attention mechanism with the efficient feature processing idea of CSPNet. Specifically, the input complex time-frequency map first extracts preliminary complex features through CConv, and then is split (Split) along the channel dimension. Inheriting the architecture of CSPNet, the features are divided into two branches: Branch 1 directly retains the original complex features to maintain information integrity, and Branch 2 dynamically generates a weight map through the spatial attention mechanism. In the attention branch, the features undergo dual pooling operations of max pooling (highlighting significant noise or signal regions) and average pooling (capturing the global energy distribution) to generate two types of spatial statistical descriptors; then, the results of the dual pooling are fused through complex convolution (CConv), and combined with complex batch normalization (CBN) to dynamically adjust the feature distribution to adapt to different noise scenarios. Finally, the fused features are compressed into a spatial weight map within the range of [0, 1] through complex Sigmoid (CSigmoid). This weight map is multiplied element-wise with the original complex features of Branch 1 to achieve focused attention on the noise region, which can help the model better learn the time-frequency map of the noise in the second stage and better perform noise denoising in the third stage. The feature splitting strategy of CSPNet significantly reduces computational redundancy, and the spatial attention mechanism can more accurately focus on the key regions in the time-frequency map through the combination of dual pooling, finally outputting the time-frequency map of the clean EEG signal with features carrying adaptive weights.
[0125] The processing flow of the noise attention module includes:
[0126] (1) Split the input complex features into a main branch and an attention branch;
[0127] (2) In the attention branch:
[0128] 1) Perform complex max pooling and complex average pooling on the input features in parallel. The complex max pooling selects the maximum value of the complex modulus in the local region, and the complex average pooling calculates the average values of the real and imaginary parts respectively;
[0129] 2) Fuse the results of the dual pooling through a complex convolutional layer;
[0130] 3) Perform complex batch normalization processing on the fused features;
[0131] 4) Generate a spatial weight map through the complex Sigmoid activation function (CSigmoid);
[0132] (3) Perform complex element-wise multiplication of the spatial weight map and the main branch features.
[0133] The parameters of the inverse short-time Fourier transform in Step 5 are strictly matched with those in step (1), including:
[0134] (1) Use the same Hann window function;
[0135] (2) The window length is 126 sampling points;
[0136] (3) The overlap rate is 84.13% (calculated from a step size of 8 and a window length of 126).
[0137] In addition to achieving a denoising effect, this method can also identify the artifacts mixed in the EEG signals, such as:
[0138] (1) When the high-frequency components in the noise time-frequency diagram account for a relatively high proportion, it is determined as electromyogram artifact (EMG);
[0139] (2) When the noise time-frequency diagram shows a low-frequency periodic waveform, it is determined as electrocardiogram artifact (ECG);
[0140] (3) When the energy of the noise time-frequency diagram is concentrated in the δ / θ frequency band (0.5 - 8 Hz), it is determined as electrooculogram artifact (EOG).
[0141] The present invention also includes a system capable of running the above-mentioned time-frequency domain denoising method for EEG signals based on a noise attention mechanism.
[0142] The present invention also includes a device, including:
[0143] A memory: for storing a computer program for implementing the above-mentioned time-frequency domain denoising method for EEG signals based on a noise attention mechanism;
[0144] A processor: for implementing the above-mentioned time-frequency domain denoising method for EEG signals based on a noise attention mechanism when executing the computer program.
[0145] The present invention also includes a computer-readable storage medium, the computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, it implements the above-mentioned time-frequency domain denoising method for EEG signals based on a noise attention mechanism.
[0146] Experimental verification of the performance of the present invention
[0147] In order to fully verify the denoising performance of the present invention, in this paper, the denoising effects of single artifacts and mixed artifacts are verified under the conditions that the signal-to-noise ratios (SNRs) of EEG signals and artifacts in the EEG signals with artifacts are -5, -3, -1, 0, 1, 3, 5 respectively. The specific EEG, EMG, ECG, EOG, EM signals and EEG signals with artifacts are as Figure 3As shown. There are three indicators in total for the experiment: the signal-to-noise ratio SNR after denoising, the relative root mean square error RRMSE, and the correlation coefficient CC to quantify the denoising effect, and a full comparison is made with the mainstream denoising models in the past five years. Next, the calculation formulas of the three indicators will be introduced first:
[0148] RRMSE evaluates the deviation between the reconstructed EEG signal and the clean EEG signal. A lower RRMSE indicates better reconstruction performance. The calculation formula of RRMSE is as follows:
[0149]
[0150] Among them, y represents the clean EEG signal, represents the output of the model, that is, the denoised EEG signal, and N represents the number of signal sampling points.
[0151] To measure the similarity between two variables, the correlation coefficient (CC) is commonly used, and its calculation formula is as follows:
[0152]
[0153] Among them, represents the covariance of the true signal y and the model output signal σ y and are the standard deviations of the two respectively.
[0154] The signal-to-noise ratio (SNR) represents the ratio of signal power to noise power, and its calculation formula is as follows:
[0155]
[0156] For the denoising effect of a single artifact as Figure 4 and Figure 5 shown, Figure 4 The average performance of the model for four different artifacts (electrocardiogram signal, electromyogram signal, electrooculogram signal, and motion artifact) under 7 different experimental conditions shows that in terms of denoising of the four artifacts, the present invention is superior to other compared models, and there is a large improvement. Among them, the denoising effect of the electrocardiogram signal is the best, with an average RRMSE of 0.2042, an average CC of 0.9779, and an average SNR of 14.039. In contrast, the denoising effects of the electromyogram signal, electrooculogram signal, and motion artifact are not as good as that of the electrocardiogram signal, which is related to their frequency and amplitude characteristics. Especially for the electrooculogram signal, because its frequency band range overlaps greatly with the electroencephalogram signal, it is the most difficult to remove. Figure 5It is the performance of various models under specific different signal-to-noise ratio conditions. It can be seen that as the signal-to-noise ratio increases, the denoising effect of the model is better, which also conforms to the actual cognition, that is, the larger the signal-to-noise ratio, the fewer artifacts in the original EEG signal, so the better the denoising performance of the model.
[0157] For the real scenario (mixed artifacts of EMG and ECG), the denoising effect of the model is as Figure 6 , Figure 7 and Figure 8 shown. Figure 6 It is the average performance of the model under 7 different experimental conditions. It can be seen that in terms of removing mixed artifacts, the present invention is still far superior to other comparative methods. We can also see that removing mixed artifacts is more difficult than removing single artifacts because the components of the mixed artifacts are more complex and the frequency bands of the noise are also more complex. From Figure 7 it can be seen that the EEG signal after denoising by the present invention fits best with the clean EEG signal. Combining Figure 8 from the perspective of the time-frequency diagram, although the fitting effect of the DuoCL model is also good, its time-frequency diagram in the time-frequency domain is quite different from that of the clean EEG signal, which also reflects the advantage of the present invention: it can simultaneously perform denoising in the time-frequency domain, so as to achieve a better denoising effect and improve the interpretability of the model.
[0158] From Figure 9 it can also be seen that in addition to simply denoising, the present invention can also obtain the time-frequency regions of the artifacts mixed in the original EEG signal compared with other methods, which helps to solve the problem of artifact mixing at the source. For example: the present invention can obtain the time-frequency mask of the noise through two Unet1D models, so as to know the approximate region of the noise in the original signal; it can also input the time-frequency mask of the noise into the noise attention mechanism of the DS-Natt-CCUnet model to obtain a more accurate time-frequency diagram of the noise. This is very useful in practical applications. For example: assuming that there are many high-frequency components in the time-frequency diagram of the noise, it can be inferred that there may be EMG artifacts, and then the EEG signal acquisition process can be regulated accordingly, especially paying attention to keeping the facial muscles relaxed to prevent the mixing of EMG signals, which provides important targeted suggestions for the actual acquisition of EEG signals.
[0159] In addition, from Figure 10The ablation experiment results show that: with the gradual stacking of module functions, the model performance shows a significant improvement. Taking the ECG signal as an example, through the deep feature extraction ability of the CC3k2 module, CCUnet reduces the RRMSE from 0.2445 of CUnet to 0.2302, and at the same time the CC value is increased to 0.9705, verifying the enhancement effect of the CC3k2 module on the learning of time-frequency features; after further introducing the noise attention mechanism, that is, the noise time-frequency mask is used to guide the Natt-CCUnet model for denoising, the CC value of the EMG signal increases from 0.9410 to 0.9493, and the SNR of the mixed artifacts is increased by 0.65 dB, indicating that the noise attention mechanism can guide the model to better learn time-frequency features and suppress noise components. The complete DS-Natt-CCUnet model combines the above advantages and is guided by a more accurate noise time-frequency matrix for denoising. As a result, the CC value reaches 0.9113 in the mixed artifact scenario, a 5.1% improvement compared to the basic model, and the RRMSE for a single artifact (such as EOG) is the lowest (0.3108), and the SNR is the highest (10.1932). This result proves that the collaborative design of the CC3k2 module and the noise attention mechanism can not only strengthen the feature expression ability but also dynamically optimize the noise suppression process, thus achieving robust denoising in multiple noise scenarios.
[0160] Finally, when actually applying the present invention, only the involved model and its pre-trained parameters need to be deployed to the target device to achieve end-to-end automatic removal of EEG signal artifacts. This method is not only easy to operate but also can significantly improve the efficiency and accuracy of EEG signal preprocessing. Compared with traditional manual processing or much more complex multi-step preprocessing processes, this method greatly reduces the need for human intervention and at the same time reduces the possibility of introducing errors due to improper operation or inaccurate parameter adjustment. In addition, since the model is pre-trained, its performance shows strong robustness in multiple scenarios and can adapt to different types of EEG signal data, thus providing a more reliable basis for subsequent signal analysis and feature extraction. This efficient and automated processing method is especially suitable for application scenarios that require fast and real-time processing of large-scale EEG data, such as clinical diagnosis, brain-computer interface development, and neuroscience research.
[0161] In summary, the EEG signal time-frequency domain denoising method based on the noise attention mechanism proposed in the present invention has three major advantages: First, it can simultaneously remove various artifacts in the EEG from the time-frequency domain perspective, and the performance and interpretability of the model are strong; Second, the type of artifacts mixed in the original EEG signal can be analyzed by analyzing the noise time-frequency map obtained by the model, so as to put forward some specific suggestions for reducing the mixing of artifacts in the actual process of collecting EEG signals; Third, the model in the present invention can be easily deployed to the target device to achieve end-to-end automatic removal of EEG signal artifacts.
Claims
1. A method for denoising EEG signals in time and frequency domain based on noise attention mechanism, characterized in that: The following steps are involved: Step 1: Obtain the EEG signal with artifacts and convert it into a time-frequency diagram using short-time Fourier transform so that it can be denoised from the perspective of combining time and frequency domains; Step 2: Use the one-dimensional Unet1D model to obtain the temporal mask of the EEG signal with artifacts, that is, the area of noise in the time domain; The frequency mask of the EEG signal with artifacts is obtained using Fourier transform and the one-dimensional Unet1D model of the same structure, that is, the area of noise in the frequency domain; then the two are combined to obtain the time-frequency mask of the noise, that is, the approximate area of noise in the time-frequency domain; Step 3: Input the time-frequency diagram of the EEG signal with artifacts obtained in step 1 into the noise attention module in the complex C3k2Unet model of the dual-stage-noise attention mechanism, and also input the time-frequency mask of the noise obtained in step 2 into the noise attention module to help the C3k2Unet model locate the noise area. Finally, the C3k2Unet model outputs an accurate noise time-frequency diagram; Step 4: After obtaining the accurate noise time-frequency graph, it is input into the noise attention module. The input is still the time-frequency graph of the EEG signal with artifacts. The DS-Natt-CCUNet model performs accurate denoising and outputs the time-frequency graph of the clean EEG signal. Step 5: Perform inverse short-time Fourier transform on the time-frequency diagram of the clean EEG signal to finally obtain a clean EEG signal.
2. According to claim 1, a method for denoising an EEG signal in time and frequency domain based on a noise attention mechanism is characterized in that: The window function used in the discrete short-time Fourier transform of the artifact-bearing EEG signal in step 1 is the Hann window, and the specific formula of the discrete short-time Fourier transform is as follows: Where X[m,k] is the complex spectral coefficient of m time frame and k frequency, x[n] is the original discrete signal, w[n] is the window function of length L (such as Hann window), H is the frame shift (step length), N is the number of FFT points, j is the imaginary unit, and π is the circumference of a circle; the original signal can be reconstructed by inverse short-time Fourier transform (ISTFT), and the formula is: Where A is the normalization factor (usually A = Σ m ω 2 [n-mH]), w[n-mH] is the window function used in the inverse transform, the position is determined by the mH offset, and the rest is consistent with the STFT formula.
3. According to the method of claim 1, the method is characterized in that: The step 2 is specifically as follows: respectively generating a time mask and a frequency mask through a parallel time domain Unet1D network and a frequency domain Unet1D network, and multiplying the time mask and the frequency mask element by element to generate a time-frequency joint noise mask.
4. According to claim 1, a method for denoising an EEG signal in time and frequency domain based on a noise attention mechanism is characterized in that: The DS-Natt-CCUNet model mentioned in the step 3 has a basic framework of Unet, which combines the complex-valued CNN and the C3k2 feature extraction module in YOLOv11 to construct an encoder that can process the time-frequency graph of EEG signals; it combines CSPNet (Cross Stage Partial Network) with the spatial attention mechanism, and uses the time-frequency mask / graph of the noise as input to construct a noise attention mechanism.
5. According to claim 4, a method for denoising EEG signals in time and frequency domain based on noise attention mechanism is characterized in that: The complex neural network includes four parts: (1) Complex Convolution (CConv) The convolution kernel of the complex convolution layer is composed of two sets of learnable parameters, the real part and the imaginary part. The input data is also expressed in complex form. During the calculation process, the complex convolution follows the complex multiplication rule. The specific process is: the result of the convolution of the real part of the weight and the real part of the input minus the result of the convolution of the imaginary part of the weight and the imaginary part of the input is used as the real part of the output; the result of the convolution of the imaginary part of the weight and the real part of the input plus the result of the convolution of the imaginary part of the weight and the imaginary part of the input is used as the imaginary part of the output, thereby obtaining the result of the complex convolution; The detailed operation of the specific process is: 1) Define the complex convolution kernel W = W real +iW imag and complex input X = X real +X imag , where W real and W imag is the real and imaginary part of the complex convolution kernel, X real and X imag are the real and imaginary parts of the complex input; 2) Calculate the real part of the output: in is the real part of the output, * is the convolution operation, and Conv(·) is the real convolution function; 3) Calculate the output imaginary part: in is the imaginary part of the output; (2) Complex Pooling The complex pooling layer is divided into maximum pooling and average pooling. They are also applied to the complex domain. For maximum pooling, the maximum value of the modulus of the complex values in the local area is selected; for average pooling, the average of the real and imaginary parts of the complex numbers in the local area is calculated as the real and imaginary parts of the result. According to the actual performance of the two pooling methods, complex maximum pooling is finally selected as the method of feature dimensionality reduction. (3) Complex Batch Normalization (CBN) The complex batch normalization layer is an extension of batch normalization in the complex domain. It is designed to handle the normalization problem of complex-valued data. By normalizing the statistical characteristics of small batches of data, the distribution of complex-valued features remains stable, thereby accelerating training and improving model performance. Since complex data contains two dimensions, the real part and the imaginary part, the design of CBN needs to consider complex operations. The specific calculation formula is as follows: in, is the normalized complex value, x is the input complex value data, E[x] is the mean of the complex input x, V is the covariance matrix of the complex input, and its definition is shown in formula (6). In order for V to have a solution, V must satisfy positive definiteness or semi-positive definiteness, and the condition for V to be semi-positive definite is the mean μ, covariance Γ, and pseudo-variance C, which is shown in formula (7): Among them, V rr is the covariance between the real and real parts, V ri is the covariance of the real and imaginary parts, V ir is the covariance of the imaginary part and the real part, V ii is the covariance between the imaginary parts, Con(a,b) is the covariance calculation function; Finally, two learnable adjustment parameters γ and β are introduced to adjust the feature distribution, and the calculation formula is as follows: in, is a semi-positive definite matrix containing four parameters, β is a complex vector, which is consistent with the dimension of the complex number. For the convenience of training and batch normalization Satisfying the standard Gaussian distribution, γ rr and γ ii Initialize to γ ri , γ ir And the real and imaginary parts of β are initialized to 0; (4) Complex Activation Function Layer The activation function used in the complex activation function layer is the ReLU activation function, which is defined as f(x) = max(0, x), that is, the positive part of the input value is retained and the negative value is truncated to zero. The calculation process in the complex domain is shown in formula (9), where the real part of the input data is and the imaginary part The ReLU activation function is used to obtain the real and imaginary parts of the results. Similarly, the CSigmoid activation function is shown in formula (10): Among them, CReLU(x) is the complex ReLU activation function, and ReLU(x) is the real ReLU activation function; Among them, CSigmoid(x) is the complex Sigmoid activation function, and Sigmoid(x) is the real Sigmoid activation function.
6. According to the method of claim 1, the method is characterized in that: The method for denoising the noise attention module in the DS-Natt-CCUNet model in step 3 or step 4 to obtain a clean time-frequency graph is specifically as follows: The input complex time-frequency graph first extracts preliminary complex features through CConv, and then splits (Split) along the channel dimension. The architecture of CSPNet is inherited to divide the features into two branches: Branch 1 directly retains the original complex features to maintain information integrity, and Branch 2 dynamically generates a weight map through the spatial attention mechanism. In the attention branch, the features undergo double pooling operations of maximum pooling and average pooling to generate two types of spatial statistical descriptors; then the double pooling results are fused through complex convolution (CConv), and the feature distribution is dynamically adjusted in combination with complex batch normalization (CBN) to adapt to different noise scenarios. Finally, the fused features are compressed into a spatial weight map in the range of [0,1] through complex Sigmoid (CSigmoid); the weight map is multiplied point by point with the original complex features of branch 1, and finally the clean EEG signal time-frequency graph is output.
7. A system, characterized in that: A method for denoising an EEG signal in the time and frequency domain based on a noise attention mechanism can be run according to any of claims 1 to 6.
8. A device, characterized in that: include: Memory: used to store a computer program for implementing a method for denoising an EEG signal in the time and frequency domain based on a noise attention mechanism as described in any one of claims 1 to 6; Processor: used to implement the EEG signal time-frequency domain denoising method based on noise attention mechanism as described in any of claims 1-6 when executing the computer program.
9. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, it implements the method for denoising an EEG signal in the time and frequency domain based on a noise attention mechanism as described in any one of claims 1 to 6.
Citation Information
Patent Citations
Electroencephalogram signal denoising method based on one-dimensional residual convolutional neural network
CN109784242A
Electroencephalogram signal denoising method using time-frequency feature multi-scale dense fusion neural network
CN117332208A
Electroencephalogram signal denoising method based on double-path convolution denoising network
CN119157556A
Feature extraction method and apparatus based on time domain and frequency domain of speech signal, and echo cancellation method and apparatus
WO2023044962A1
Cited By
Seismic data denoising method of time-frequency domain mixed loss constraint convolutional network
CN120577873A
Multi-channel electroencephalogram signal noise reduction method based on U-shaped network and hierarchical attention
CN121350416A