Electroencephalogram signal denoising method based on frequency domain attention and time domain adaptive convolution
Through the frequency domain attention and time domain adaptive convolutional network framework, the problems of residual noise and insufficient signal reconstruction fidelity in EEG signal processing are solved, and efficient EEG signal denoising is achieved, which is suitable for clinical neuroscience and neurological disease diagnosis.
Patent Information
- Application Number
- CN202510828074.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-20
- Publication Date
- 2025-10-03
AI Technical Summary
The existing technology in EEG signal processing has problems such as insufficient utilization of frequency domain information, difficulty of fixed-size convolution kernels in adapting to multi-scale noise changes, and imperfect mechanism of fusion of time domain and spatial domain features, resulting in residual noise and insufficient signal reconstruction fidelity.
A framework based on frequency domain attention and time domain adaptive convolutional network is adopted to enhance noise suppression capability through the frequency domain attention mechanism, dynamically optimize feature extraction by combining the adaptive multi-scale convolution module, and improve signal reconstruction quality through the spatiotemporal feature fusion module to achieve high-fidelity EEG signal recovery.
Significantly reduces residual noise, improves peak signal-to-noise ratio and structural similarity, and achieves high-fidelity EEG signal recovery, suitable for clinical neuroscience, brain-computer interface and neurological disease diagnosis.
Smart Images

Figure CN120744313A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of electroencephalogram (EEG) signal processing technology, and in particular to an EEG signal denoising method based on frequency domain attention and time domain adaptive convolution, which is applicable to the fields of clinical neuroscience, brain-computer interface, and neurological disease diagnosis. Background Art
[0002] Electroencephalography (EEG) is a modern auxiliary examination method that uses an electroencephalogram to amplify the brain's own weak bioelectricity and record it as a curve graph to help diagnose diseases. It does not cause any trauma to the person being examined and is mainly used to examine organic intracranial lesions such as epilepsy, encephalitis, cerebrovascular disease and intracranial space-occupying lesions.
[0003] EEG signal acquisition involves placing electrodes on the scalp surface according to internationally specified points, detecting potential differences at each point. However, due to the influence of the skull and scalp on the transmission of electrical signals, the electrical signals collected from the scalp are often mixed with a lot of noise and contain very little useful information.
[0004] EEG signals are susceptible to noise contamination, including electromyography (EMG), electrooculography (EOG), power frequency interference, and sweating artifacts. Traditional methods such as independent component analysis (ICA) and wavelet denoising have limited effectiveness in complex noisy scenarios, making it difficult to effectively separate high-frequency noise from low-frequency artifacts. While existing deep learning models excel at time-domain denoising, they ignore the global dependencies of visual domain features, resulting in significant residual high-frequency noise.
[0005] Existing technologies have the following defects: First, traditional methods make insufficient use of frequency domain information and fail to fully explore the global correlation of frequency domain signals; second, fixed-size convolution kernels are difficult to adapt to the dynamic changes of multi-scale noise in EEG signals; third, the fusion mechanism of time domain and spatial domain features is imperfect, affecting the fidelity of signal reconstruction. Summary of the Invention
[0006] The purpose of this invention is to propose an EEG signal denoising method based on frequency-domain attention and time-domain adaptive convolutional networks. By combining an improved convolutional time-domain audio separation network (Conv-TasNet) framework with a multi-source noise simulation strategy, efficient EEG signal denoising is achieved. The core innovation of this invention lies in: enhancing the frequency-domain noise suppression capability through the frequency-domain attention mechanism, dynamically optimizing the time-domain feature extraction through an adaptive multi-scale convolution module, and improving the signal reconstruction quality through a spatiotemporal feature fusion module, ultimately achieving high-fidelity EEG signal recovery.
[0007] In terms of frequency-domain feature modeling, this invention overcomes the limitations of traditional time-domain models by introducing a frequency-domain attention mechanism. This method decomposes the signal into the frequency domain using a fast Fourier transform, analyzes global dependencies for both the real and imaginary parts, and uses multi-head self-attention to model the correlation between frequency-domain noise and the valid signal. Compared to traditional methods, this mechanism can accurately capture high-frequency noise components and significantly reduce residual noise.
[0008] In terms of time-domain feature extraction, the present invention optimizes the noise separation process through an adaptive multi-scale convolution module. By dynamically adjusting the convolution kernel size (3, 5, 7, 11), local and global features at different time scales are extracted in parallel. Batch normalization and random dropout techniques are combined to enhance the model's robustness to non-stationary EEG signals, effectively suppressing multiple sources of interference, such as myoelectric and oculoscopic noise.
[0009] In terms of spatiotemporal feature fusion, this paper improves signal reconstruction quality through a spatiotemporal feature collaborative optimization module. It integrates temporal dynamic features with spatial electrode distribution characteristics through a weighted strategy, and uses a cross-channel attention mechanism to align signal phases, avoiding loss of signal details and significantly improving peak signal-to-noise ratio and structural similarity indicators.
[0010] The specific implementation steps of the present invention are as follows: generate a standardized noisy signal through data preprocessing, extract time domain features through a multi-scale convolutional encoder, separate noise through a separation network combined with a frequency domain attention mechanism and an adaptive convolution module, optimize signal reconstruction through a spatiotemporal fusion module, and finally output a high-quality denoised signal through a decoder.
[0011] Data preprocessing is achieved through the following steps: noisy EEG signals and corresponding clean signals are collected from a public database, amplitude differences between channels are eliminated through a standardization formula, and a training dataset in a complex noise scenario is generated by weighted superposition of electromyographic noise, electrooculographic noise, power frequency interference, and sweating artifacts.
[0012] The model is constructed through the collaborative work of an encoder, a separation network, and a decoder. The encoder extracts time-domain features in parallel using multi-scale one-dimensional convolution kernels. The separation network achieves efficient separation of noise and signal through an adaptive convolution module and a frequency-domain attention mechanism. The decoder restores signal temporal continuity through a spatiotemporal fusion module and optimizes waveform details using adaptive convolution kernels.
[0013] The specific implementation process of the frequency domain attention mechanism is as follows: convert the time domain signal to the frequency domain through fast Fourier transform and decompose it into real and imaginary components; model the global dependency relationship of the real and imaginary components respectively through the multi-head self-attention mechanism; and restore the enhanced frequency domain features to the time domain through inverse Fourier transform to complete noise suppression.
[0014] Model training is implemented using the AdamW optimizer with a cosine annealing learning rate scheduling strategy. Mean squared error is used as the loss function, and noise weights are dynamically adjusted through end-to-end supervised learning to enhance model generalization. During training, noisy EEG signals are used as input and clean signals as labels, and network parameters are optimized through backpropagation.
[0015] The signal denoising process is completed through the following steps: normalizing the input signal, extracting features through the encoder, removing noise through the separation network, reconstructing the waveform through the decoder, and finally restoring the complete time series signal by splicing the denoised segments. The denoising effect is quantitatively evaluated using peak signal-to-noise ratio, mean absolute error, and structural similarity metrics.
[0016] Compared with existing technologies, the present invention has the following advantages: it fully exploits global frequency domain information through a frequency-domain attention mechanism, solving the problem of high-frequency noise residue in traditional methods; it dynamically captures the multi-scale characteristics of time-domain signals through an adaptive multi-scale convolution module, improving robustness to complex noise; and it collaboratively optimizes the signal reconstruction process through a spatiotemporal feature fusion module, ensuring the spatiotemporal fidelity of the denoised signal. Experiments have shown that this method significantly outperforms existing technologies in key indicators such as peak signal-to-noise ratio (27.41dB) and mean absolute error (0.183), providing efficient and reliable technical support for clinical EEG signal processing. BRIEF DESCRIPTION OF THE DRAWINGS
[0017] Figure 1 Schematic diagram of the overall process of the EEG signal denoising method of the present invention.
[0018] Figure 2 Schematic diagram of the structure based on frequency domain attention and time domain adaptive convolution module. DETAILED DESCRIPTION
[0019] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings, and will also provide detailed supplementary explanations of some details in the technical invention solutions.
[0020] This paper proposes an efficient EEG signal denoising method based on frequency-domain attention and time-domain adaptive convolution techniques. The experimental data is a publicly available EEG dataset, which contains single-channel raw EEG signals and is widely used for EEG signal processing and analysis.
[0021] like Figure 1 As shown in FIG, the overall process of the EEG signal denoising method of the present invention mainly includes five stages: data loading and normalization, noise simulation, model construction, training and verification, and evaluation and testing.
[0022] First, the raw EEG data is acquired and preprocessed through the data loading and normalization module. Then, the standardized signal is subjected to multi-source noise superposition using a noise simulation method. The noisy signal is denoised through the model structure. Finally, the model effect is evaluated and tested to obtain high-fidelity EEG signal output.
[0023] During the data loading and normalization phase, the original EEG signal must first be standardized, as shown below:
[0024] Normalization: Normalize the EEG signal and its corresponding noise signal to reduce the amplitude difference between different signals. Assuming the EEG signal is X(t), the normalization process is expressed as follows:
[0025]
[0026] Here μ X is the mean value of the EEG signal, σ X This ensures that the signals are within the same range, making it easier for the model to process.
[0027] In the noise simulation stage, multiple noise sources, such as electromyographic noise (EMG), electrooculographic noise (EOG), power frequency interference, Gaussian noise, and sweating artifacts, are superimposed to generate a noisy mixed signal that is closer to the clinical environment. The specific expression is as follows:
[0028] Noise simulation and enhancement: Add different types of noise to the EEG signal, including EMG noise, EOG noise, power frequency interference, and sweating artifacts. These noises are added to the EEG signal through weighting and superposition, as shown below:
[0029]
[0030] Here Y(t) is the EEG signal after adding noise, α and β are the weighting coefficients of EMG and EOG noise respectively, N gauss (t) is Gaussian noise, N powerline (t) is the power frequency noise, N sweat (t) is sweating artifact noise.
[0031] During the model construction phase, key modules such as multi-scale convolution, bidirectional attention, and spatiotemporal fusion are used to suppress multi-source noise in EEG signals and achieve high-fidelity signal reconstruction. Specifically, the noisy EEG signals are fed into a framework that combines a multi-layer convolutional network with an attention mechanism, and a more complete EEG feature representation is obtained through spatiotemporal feature fusion, as shown below:
[0032] Encoder: The encoder consists of an EEG-specific convolutional layer and two one-dimensional convolutional layers to extract high-dimensional features of the EEG signal. The EEG-specific convolutional layer is a convolutional module tailored for EEG signals. It uses multiple layers of one-dimensional convolution operations, combined with different convolution kernel sizes, to extract multi-scale features from the EEG signal. This convolutional layer performs the convolution operation using the following formula:
[0033]
[0034] Here y (l) (t) is the output after the lth convolution layer, x i (t) is the i-th input signal channel, is the weight of the convolution kernel of the lth layer, b (l) is the bias term, and ReLU is the activation function. By selecting convolution kernels of different sizes (such as 3, 5, 7, and 11), the encoder can simultaneously process information from different frequency bands, ensuring that the key signal features in the EEG signal are captured and noise is effectively filtered out.
[0035] The extracted multi-scale features are further processed through a standard one-dimensional convolutional layer.
[0036] The first convolution layer is as follows:
[0037]
[0038] Here W1 is the convolution kernel weight, L is the convolution kernel size, and the stride is L / 2.
[0039] The second convolution layer is as follows:
[0040]
[0041] Through these two layers of convolution, the signal is projected into a higher-dimensional feature space, further capturing local features. The convolutional output is normalized using batch normalization to enhance model stability. To prevent overfitting, a dropout layer is added to randomly drop some neurons to improve the model's generalization capabilities.
[0042] Separation network: The separation network uses a multi-layer stacking approach to separate signals from noise through depthwise separable convolution and frequency-domain self-attention mechanisms. It specifically includes the following parts:
[0043] Adaptive multi-scale convolution block: Each layer in the separation network first performs a multi-scale convolution operation to capture signal characteristics at different time scales. Convolution kernels of different sizes extract information from different frequency bands and fuse them together. The formula is as follows:
[0044] h multi=ReLU(BN(Conv1D3(h)+Conv1D5(h)+Conv1D7(h)+Conv1D 11 (h)))
[0045] The convolution kernel sizes here are 3, 5, 7, and 11 respectively.
[0046] Bidirectional self-attention mechanism: In order to capture frequency domain features, the time domain signal is converted to the frequency domain through fast Fourier transform (FFT), and then the self-attention mechanism is applied to the real and imaginary parts of the frequency domain signal respectively to enhance the frequency domain information. The specific steps are:
[0047] Fast Fourier Transform (FFT): H freq =FFT(h multi )
[0048] The attention calculation of the real and imaginary parts is as follows:
[0049]
[0050]
[0051] Finally, the frequency domain features are restored to the time domain through the inverse Fourier transform (IFFT):
[0052]
[0053] Spatiotemporal feature fusion: After processing each layer, the spatiotemporal feature fusion module is applied to further combine the temporal and spatial features to enhance the spatiotemporal perception capability of the model. The specific fusion formula is as follows:
[0054] h fused =LayerNorm(h+h temporal +h spatial )
[0055] Decoder: The decoder maps the features output by the separation network back to the time domain through linear layers, adaptive convolution layers, and spatiotemporal feature fusion modules to obtain the denoised EEG signal.
[0056] First, the features are restored to the time domain through a linear layer:
[0057] h time =Linear(h fused )
[0058] Then, an adaptive convolution operation is applied:
[0059] h conv =ReLU(BN(Conv1D(h time )))
[0060] Finally, combining the spatiotemporal features, the denoised EEG signal is output:
[0061]
[0062] The decoder is designed to ensure that important temporal and spatial information is preserved when recovering the signal.
[0063] During the training and validation phase, the model is optimized in an end-to-end manner, using AdamW or other advanced optimization algorithms, and combined with a learning rate scheduling strategy, so that the network can achieve stable convergence and excellent generalization capabilities in different noise environments. After training, the model is applied to actual EEG signal denoising and verified on a test set or clinical data. Training settings: The model is trained using the AdamW optimizer, with a learning rate of 0.0001, and the cosine annealing scheduler is used for learning rate adjustment. The training goal is to minimize the mean square error between the noisy EEG signal and the clean EEG signal, and the loss function is:
[0064]
[0065] Here T is the length of the signal, is the model output, and x(t) is the real clean signal.
[0066] During the evaluation and testing phase, the denoised EEG signals were compared with the true clean signals using multiple indicators such as peak signal-to-noise ratio, mean absolute error, structural similarity, and Pearson correlation coefficient to evaluate the denoising effect and fidelity of the model.
[0067] like Figure 2 The figure shows the structure diagram (i.e., model block diagram) of the frequency-domain attention and time-domain adaptive convolution module of the present invention. The model block diagram includes, from left to right, the input EEG signal, multi-layer one-dimensional convolution (Conv Layers), ReLU activation function, normalization layer (Layer Norm), multi-scale convolution (Multi-scale Conv), bidirectional attention (Bidirectional Attention Block), spatio-temporal feature fusion (Spatio-Temporal Fusion), linear superposition and segmentation (Linear+Overlap and Add), spatio-temporal fusion (Spatio-Temporal Fusion) and adaptive convolution block (AdaptiveConvBlock) modules, and finally outputs the denoised EEG signal.
[0068] exist Figure 2In the model structure shown, the features of EEG signals at different time scales are first captured through multi-layer one-dimensional convolution kernels; then the bidirectional attention module is entered to model the global dependency in the time domain and frequency domain respectively, effectively suppressing high-frequency noise and low-frequency artifacts; then, the spatiotemporal feature fusion module is used to integrate the time domain dynamic features and spatial domain distribution information to ensure phase consistency and amplitude accuracy during signal reconstruction.
[0069] At the output end, the model uses adaptive convolution blocks to fine-tune the time domain waveform, further removing residual noise and optimizing edge details. The final output EEG signal not only has a significant improvement in peak signal-to-noise ratio, but also performs well in indicators such as structural similarity and correlation coefficient.
[0070] The method of the present invention fully utilizes the advantages of multi-scale convolution and bidirectional attention mechanism, takes into account the time domain and frequency domain characteristics of EEG signals, and can achieve high-precision signal separation and reconstruction in complex noise environments. Figure 1 The overall process shown and Figure 2 The model block diagram shown works together to significantly improve the quality of the denoised EEG signal.
[0071] The specific implementation methods of the present invention can be flexibly adjusted according to actual needs at each stage. For example, the noise source type and weighting coefficient range can be changed in noise simulation, and the convolution kernel size or the number of attention heads can be adjusted in the model structure; all these improvements are within the scope of the technical concept of the present invention.
Claims
1. A method for denoising EEG signals based on frequency domain attention and time domain adaptive convolution, characterized in that: The following steps are involved: Perform data preprocessing and noise simulation on EEG signals to generate noisy mixed signals; The encoder extracts multi-scale features from the noisy signal to generate a high-dimensional spatiotemporal feature representation; By combining a separation network with adaptive multi-scale convolutional blocks and a bidirectional self-attention mechanism, it can distinguish signals from noise and enhance temporal dependencies. The spatiotemporal feature fusion module integrates the temporal dynamic features and the spatial distribution characteristics to generate fusion features; The fused features are mapped back to the time domain through the decoder, and the denoised EEG signal is output.
2. The method for denoising an EEG signal according to claim 1, wherein: The data preprocessing and noise simulation of the EEG signal specifically includes the following steps: The original EEG signal X(t) is obtained through the data acquisition unit and normalized to eliminate the amplitude difference between channels. The calculation formula is: Among them, X(t) is the original EEG signal, μ X is the mean, σ X is the standard deviation; By weighted superposition, multiple noise sources are simulated, including myoelectric noise, electrooculographic noise, Gaussian noise N gauss (t), power frequency interference N p Owerline (t) and sweating artifact N s weat(t), generate noisy EEG signal: Y(t)=X(t)+α·EMG(t)+β·EOG(t)+N gauss (t)+N powerline (t)+N sweat (t), where α and β are dynamically adjusted weighting coefficients with a value range of [0.1, 0.5], used to control the noise intensity; The Gaussian noise N gauss (t) obeys the normal distribution with mean 0 and variance 0.01, and the power frequency interference N powerline (t) is a sine wave signal with a frequency of 50 Hz or 60 Hz, and the sweating artifact N sweat (t) is a low-frequency (0.5-2Hz) random fluctuation signal.
3. The method for denoising an EEG signal according to claim 1, wherein: The multi-scale feature extraction of the noisy signal by the encoder specifically includes the following steps: Use multiple layers of one-dimensional convolution kernels (3, 5, 7, 11) to perform convolution operation on the input signal. The formula is: Among them, C in is the number of input channels, is the convolution kernel weight of the lth layer, b (l) is the bias term; The convolution output is normalized by batch normalization to reduce the internal covariate shift. The formula is: Among them, γ and β are learnable parameters, μ batch and σ batch is the mean and variance of the batch data, and ε is a small constant to prevent division by zero; Some neurons are randomly discarded through the Dropout layer, and the dropout probability is set to 0.2 to improve the generalization ability of the model.
4. The method for denoising an EEG signal according to claim 1, wherein: The separation network is combined with adaptive multi-scale convolution blocks and bidirectional self-attention mechanism to distinguish signals from noise and enhance temporal dependencies, specifically including the following steps: The features of different time scales are fused through adaptive multi-scale convolution blocks, and the formula is: h multi =ReLU(BN(Conv1D3(h)+Conv1D5(h)+Conv1D7(h)+Conv1D 11 (h))), where Conv1D k Represents a one-dimensional convolution operation with a convolution kernel size of k; Convert the time domain signal to the frequency domain through fast Fourier transform and decompose it into real part and the imaginary part H freq =FFT(h multi ), apply multi-head self-attention mechanism to the real part and imaginary part respectively to model the global dependency in frequency domain.
5. The method for denoising an EEG signal according to claim 1, wherein: The integration of temporal dynamic features and spatial distribution characteristics through the spatiotemporal feature fusion module specifically includes the following steps: The temporal dynamic features h_temporal and the spatial distribution features h_spatial are extracted through independent convolutional layers respectively; Adaptive weighting strategy is used to fuse the two types of features. The formula is: h fused =LayerNorm(λ·h temporal +(1-λ)·h spatial ), where λ is the dynamically adjusted fusion weight, ranging from [0.3, 0.7], which is calculated by the attention mechanism; Residual connections are used to preserve the original feature information and avoid gradient disappearance: h final =h+h fused .
6. The method for denoising an EEG signal according to claim 1, wherein: The decoder maps the fused features back to the time domain and outputs the denoised EEG signal, which specifically includes the following steps: Project the fused feature h_final to the time domain space through the linear layer: h time =Linear(h final ); The time domain waveform is further optimized through the adaptive convolution layer, the formula is: conv =ReLU(BN(Conv1D(h time ))); The decoding features are optimized twice by the spatiotemporal feature fusion module to ensure the continuity of signal timing: x = SpatioTemporalFusion(h conv ), wherein the spatiotemporal fusion module aligns the signal phases of different electrodes through a cross-channel attention mechanism.