Motor imagery electroencephalogram signal recognition technology fusing multi-band adaptive filtering and self-attention mechanism
By adopting multi-branch structure adaptive filtering and multi-scale feature extraction modules in motion imagination BCI technology, combined with the information fusion of self-attention mechanism, the problem of focusing only on single frequency features in the existing technology is solved, and the recognition accuracy of motion imagination signals and the generalization ability of the model are improved.
Patent Information
- Application Number
- CN202510218242.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-26
- Publication Date
- 2025-05-30
AI Technical Summary
Existing motion imagination BCI technology usually only focuses on a single frequency feature and ignores multi-band feature analysis, resulting in the introduction of interfering signals when extracting features, affecting the recognition effect.
The adaptive filtering module in the multi-branch structure extracts the frequency domain characteristics of multiple frequency bands, and the time domain and spatial domain characteristics are extracted through the multi-scale feature extraction module, and the branch information fusion module of the self-attention mechanism is combined to efficiently complete the motion imagination signal recognition task.
Through the integration of multi-band feature analysis and information fusion of self-attention mechanisms, the recognition accuracy of motion imagination tasks is improved, and the overall performance and generalization ability of the model are enhanced.
Smart Images

Figure CN120067814A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the fields of brain-computer interfaces and motor imagery, and particularly to a motor imagery electroencephalogram signal recognition technology that integrates multi-band adaptive filtering and self-attention mechanism. Background Art
[0002] Motor Imagery (MI) is a commonly used non-invasive electroencephalogram (EEG) signal in brain-computer interface technology, which refers to the electroencephalogram signal generated by imagining specific movements or muscle activities. The motor imagery-based brain-computer interface (MI-BCI) is a technology that records the neuronal activities in the brain, decodes the intention of the subject's motor imagery, and converts it into a control signal to control external devices. It is one of the most widely used brain-computer interface paradigms. Its basic principle comes from the activation of the brain regions related to limb movement in the sensorimotor area of the brain, that is, the phenomena of event-related synchronization and event-related desynchronization occur. Research shows that this imagined movement and the actual executed movement will show similar activation patterns in the brain sensorimotor area. In the field of rehabilitation medicine, patients with lost motor function can activate some damaged motor control regions through systematic motor imagery training to achieve the restoration of motor ability; in the field of entertainment games, users can control virtual characters or objects through motor imagery; in the field of intelligent driving, the driving behavior pattern can be obtained by analyzing the motor imagery electroencephalogram signal of the driver, and based on this, an electric vehicle control model can be built to achieve brain-computer interface-based autonomous driving.
[0003] Currently, the motor imagery BCI technology has been successfully applied to multiple fields. By identifying the electroencephalogram rhythm, the electroencephalogram signals from different imagined actions can be classified. In recent years, deep learning methods have developed greatly and achieved good results in fields such as natural language processing and computer vision. Some researchers have also applied deep learning methods to the field of motor imagery. Compared with traditional machine learning methods, deep learning methods can automatically extract features and have achieved better results in the field of motor imagery. However, many studies usually only focus on single-frequency features and lack comprehensive analysis of multi-band information, resulting in a large amount of irrelevant signals being mixed in when the model extracts frequency-domain features, which affects the recognition effect of motor imagery tasks. The motor imagery electroencephalogram signal recognition technology that integrates multi-band adaptive filtering and self-attention mechanism extracts the frequency-domain features of multiple bands through the adaptive filtering module in the multi-branch structure according to the characteristics of motor imagery signals, extracts time-domain and spatial-domain features using the multi-scale feature extraction module in each branch, and then uses the branch information fusion module based on the self-attention mechanism to efficiently complete the motor imagery signal recognition task. Summary of the Invention
[0004] The purpose of the present invention is to address the problem that the existing technology only focuses on a single frequency feature and ignores the analysis of multi-band features, resulting in the introduction of interference signals during feature extraction and thus affecting the recognition effect. A motion imagination electroencephalogram (EEG) signal recognition technology that combines multi-band adaptive filtering and self-attention mechanism is proposed.
[0005] To achieve the above technical objectives, the present invention adopts the following technical solutions:
[0006] A motion imagination EEG signal recognition technology that combines multi-band adaptive filtering and self-attention mechanism. The multi-branch model in this method uses a parallel branch structure to process the spectral information of different frequency bands respectively, and constructs a spatio-temporal hybrid feature extraction unit in each branch: capturing the temporal pattern of action potentials through temporal convolution, combining with spatial domain convolution to extract the spatial distribution characteristics of cortical potentials, and finally achieving high-precision parsing of the subject's motion intention through branch feature fusion based on the self-attention mechanism. Specifically, it includes the following steps:
[0007] Step 1: Preprocessing of EEG data; select key channels for the EEG signals in the dataset and preprocess the signals of the selected channels.
[0008] Step 2: Use an adaptive filtering module to screen the frequency bands of the preprocessed EEG data; screen the EEG data of different frequency bands as the input of the multi-scale feature extraction module through the method of online band-pass filtering.
[0009] Step 3: Use a multi-scale feature extraction module to extract the temporal and spatial domain features of the EEG signals;
[0010] Step 4: Complete the fusion of branch frequency band information based on the self-attention mechanism; the branch information fusion module calculates the similarity between the feature map channels of different branches through the self-attention mechanism and automatically assigns greater weights to the important channels among them, effectively improving the accuracy of motion imagination task recognition.
[0011] As a preferred technical solution in the present invention, in the above Step 1, the preprocessing operation of the motion imagination EEG signal is specifically to first screen specific electrodes to obtain effective signals, then use the exponential moving average algorithm (EMA) to smooth the signals, reduce the interference of noise on the effective information, ensure the effectiveness of the samples by accurately extracting the EEG segments strongly related to the motion imagination intention, and finally perform mean-variance normalization on each sample independently to eliminate the amplitude differences caused by attention fluctuations between different experiments and ensure the efficient extraction of subsequent effective features.
[0012] The exponential moving average algorithm is an algorithm used to smooth signals and perform normalization, and is often used to process time series data. This method adjusts the amplitude in the signal to a more standard range by means of weighted averaging, while preserving the changing trend of the signal and reducing mutations caused by noise or electrode movement. The formula for exponential sliding normalization is shown in Equation (1).
[0013]
[0014] In the formula, x t represents the original signal, x' t represents the normalized signal, u t and σ t 2 represent the current moving mean and variance respectively, and their calculation methods are shown in Equations (2) and (3).
[0015] μ t =βx t +(1-β)μ t-1 (2)
[0016]
[0017] In the formula, β represents the decay factor, which is set to 0.001.
[0018] As a preferred technical solution in the present invention, a Sinc convolutional layer is introduced in the second step for online band-pass filtering to screen data in different frequency bands and use it as the input of the multi-scale feature extraction module. Compared with the traditional band-pass filtering method, the Sinc convolutional layer is an adaptive band-pass filter that can be embedded in a deep learning model and can dynamically adjust the filtering parameters according to the needs of different branches. The Sinc convolutional layer optimizes the filtering process through two learnable parameters f L and f H When performing the convolution operation, the convolution kernel parameters are calculated from f L and f H through the time-domain form of the rectangular band-pass filter, as shown in Equation (4).
[0019] k(n)=g(n,f L ,f H )=2f H ×Sinc(2πnf H )-2f L ×Sinc(2πnf H ) (4)
[0020] In the formula, n represents the nth parameter in the Sinc convolution kernel, f L , f HThey represent the pass frequency and cut-off frequency of the rectangular band-pass filter respectively. The Sinc function is defined as in Equation (5).
[0021]
[0022] Since the above calculation process is completely differentiable, the filtering range parameter f of the Sinc convolution layer L and f H can be automatically optimized during the model training process, so as to achieve the effect of adaptively adjusting the filtering range.
[0023] In addition, in order to improve the performance of the rectangular band-pass filter, the Hamming window is usually used to window the convolution kernel parameters in the Sinc convolution, and the windowing process is shown in Equation (6).
[0024] k w (n) = k(n) × w(n) (6)
[0025] In the formula, k w (n) represents the convolution kernel parameters after windowing, and w(n) represents the Hamming window function, which is defined as in Equation (7).
[0026] w(n) = 0.54 - 0.46 × cos(2πn / L) (7)
[0027] where L represents the window length, which is the same as the length L sinc of the Sinc convolution kernel.
[0028] As a preferred technical solution in the present invention, in the third step, the multi-scale feature extraction module aims to fully mine the time-domain and space-domain features in the electroencephalogram signals of each frequency band. This module mainly consists of parallel time-domain convolution, space-domain convolution and pooling layers. The parallel time-domain convolution in the module is used to alleviate the strong individual differences of the motor imagery signals. Since there are significant differences in the time-domain features corresponding to different subjects under different tasks, this solution adopts diverse convolution kernel sizes to adapt to these changes. Specifically, each parallel time-domain convolution has the same number of output channels C, but adopts different convolution kernel sizes, which are (1, L1), (1, L2), (1, L3) respectively, to extract the time-domain features at different scales. In order to accelerate the model convergence, a batch normalization (BN) layer is added after each convolution layer. Finally, the feature maps generated by the three time-domain convolutions are concatenated in the channel dimension for fusion.
[0029] The spatial domain convolution in the module is mainly used to extract the spatial distribution features between each electrode. The size of its convolution kernel is set to (E, 1), where E represents the number of electrodes during signal acquisition. Since each channel of the feature map extracted by the time domain convolution represents different time domain features and has a certain degree of independence from each other. Therefore, when performing the spatial domain convolution, in order to maintain this independence, the number of convolution groups is set to the number of channels of the input feature map. At this time, the operation of the spatial domain convolution is decomposed into multiple independent convolution operations, that is, each input channel only uses its corresponding convolution kernel for calculation, thereby realizing the independent extraction of the spatial features of each channel.
[0030] After the extraction of the time domain and spatial domain features is completed, the exponential linear unit (ELU) shown in Equation (8) is used as the activation function.
[0031]
[0032] Finally, the size of the generated feature map is adjusted through average pooling to reduce the subsequent calculation amount of the model and improve the generalization ability of the model.
[0033] As a preferred technical solution in the present invention, a branch information fusion module based on the self-attention mechanism is adopted in the fourth step to weight the important channels of the feature map and enhance the feature expression ability. This module uses the significant position selection (SPS) algorithm to identify the key channels in the feature map and realizes cross-channel information fusion by constructing a self-attention matrix. Specifically, first, the input data is mapped into a query matrix and a value matrix through two two-dimensional convolutional layers, and then the significant position selection (SPS) algorithm is used to calculate the sum of squares of each channel of the query matrix and sum each channel to determine the significant position. According to the significant positions selected by SPS, an attention matrix is constructed and softmax normalization processing is performed. The value matrix is combined with the attention matrix through matrix operations to obtain an enhanced feature with the same size as the input feature. In order to further enhance the feature expression effect, the fused result will be transformed through a 1×1 convolution and added to the original input feature to obtain the fused feature. Finally, these features will be flattened and input into the linear layer for classification to generate the prediction result.
[0034] The advantages and beneficial effects of the technical solution adopted by the present invention are:
[0035] In view of the complex components and large individual differences of motor imagery signals, a multi-branch structure is introduced into the model of the present invention to extract frequency-domain information of different frequency bands respectively. An adaptive filtering module is designed, which can dynamically adjust the filtering range and effectively solve the problem of uncertain frequency ranges of different signal components; a multi-scale feature extraction module is proposed to cope with the individual differences of signals and integrate it into the multi-branch structure; a branch information fusion module is designed in combination with the self-attention mechanism, which is optimized for the temporal characteristics of the feature map, and improves the information fusion effect between branches while reducing the number of parameters of the module, thereby improving the overall performance and generalization ability of the model. BRIEF DESCRIPTION OF THE DRAWINGS
[0036] Figure 1 is a schematic flowchart of the present invention;
[0037] Figure 2 is a schematic structural diagram of the model framework of the present invention;
[0038] Figure 3 is a schematic structural diagram of the multi-scale feature extraction module of the present invention;
[0039] Figure 4 is a schematic structural diagram of the branch information fusion module of the present invention; DETAILED DESCRIPTION OF THE INVENTION
[0040] The following is a detailed description of the embodiments of the present invention in conjunction with the accompanying drawings of the specification:
[0041] A motor imagery electroencephalogram signal recognition technology that integrates multi-band adaptive filtering and self-attention mechanism, and the overall flowchart is as shown in the appendix Figure 1 shown. The designed model consists of an adaptive filtering module, a multi-scale feature extraction module and a branch information fusion module, and the overall structure is as shown in Figure 2 shown. The following is a specific description of the steps:
[0042] Step 1: Preprocessing of electroencephalogram data
[0043] The BCI Competition IV 2a dataset selected by the present invention is a publicly available brain-computer interface dataset provided by Graz University of Technology, which is widely used in the research of motor imagery (MI) tasks. This dataset contains electroencephalogram data of 9 subjects, and each subject performs 4 types of motor imagery tasks, namely left hand, right hand, tongue and both feet. A total of 25 electrodes are set up in the experiment to record electroencephalogram signals at a sampling rate of 250 Hz.
[0044] In the process of data preprocessing, the MNE-Python and Braindecode libraries are used to complete dataset reading and signal preprocessing. The specific steps are as follows:
[0045] (1) Channel screening: The electrooculogram channels were excluded, and only 22 EEG channels were selected as the source of key electrode signals.
[0046] (2) Sample truncation: Starting from the visual cue trigger, 4-second signal segments strongly related to the motor imagery task were extracted as the model input, avoiding irrelevant EEG fluctuations during task switching to ensure the validity of the data.
[0047] (3) Signal smoothing: The exponential moving average algorithm (EMA) was used to smooth the signals to reduce the interference of noise on the effective information.
[0048] (4) Normalization processing: Mean-variance normalization was independently performed on each sample to eliminate the signal amplitude differences caused by the fluctuations in the subjects' attention, ensuring the stability and efficiency of subsequent feature extraction.
[0049] After preprocessing, the sample size in the dataset is 22×1000, where 22 represents the number of electrode channels and 1000 represents the number of time points for each sample.
[0050] Step 2: Use an adaptive filtering module to perform frequency band screening on the preprocessed EEG data
[0051] Considering the high correlation between the Alpha and Beta frequency bands of EEG signals and the motor imagery task, the present invention selects these two frequency bands as the processing objects for each branch of the model. To achieve efficient frequency band separation, different initial pass frequencies f L and initial cut-off frequencies f H are used in the Sinc convolutional layers within each branch, enabling each branch to be specifically responsible for extracting information in the corresponding frequency band. The model consists of two branches, and the convolution kernel length of the Sinc convolutional layer in each branch is fixed at 60, and different frequency ranges are set respectively to accurately extract the features in the corresponding frequency bands: Alpha (8 - 13 Hz) and Beta (13 - 30 Hz). This screening process ensures that the data input into the multi-scale feature extraction module has clear spectral characteristics, thereby enhancing the model's ability to analyze motor imagery signals.
[0052] Step 3: Use a multi-scale feature extraction module to extract the time-domain and spatial-domain features of EEG signals
[0053] The structure of the multi-scale feature extraction module is as shown in the appendix Figure 3As shown, the module includes multi-scale time-domain convolution, spatial-domain convolution, and the ELU activation function. For the branches of different frequency band information, the optimal convolution kernel size of different time-domain convolutions is determined through experiments. Research shows that: the higher the signal frequency, the smaller the corresponding optimal convolution kernel size, that is, the optimal convolution kernel sizes of different frequency bands are not the same, and overall, they show an inverse relationship with the center frequency of the frequency band. Based on this rule, the present invention sets different combinations of convolution kernel sizes for the multi-scale time-domain convolution of each branch. Among them, the Alpha frequency band uses (30, 40, 50), and the Beta frequency band uses (20, 25, 30) to fully meet the requirements of time-domain feature extraction for signals in different frequency bands. The spatial-domain convolution consists of depthwise separable convolution, and its size parameter E is set according to the number of electrode channels of the samples in the dataset to ensure the effective capture of spatial features. After the extraction of time-domain and spatial-domain features, ELU is used as the activation function to enhance the feature expression ability and the nonlinear fitting ability of the model. Finally, the size of the feature map is adjusted through an average pooling layer with a window size of (1, P1) to optimize the feature representation and improve the generalization ability of the model.
[0054] Step 4: Complete the fusion of branch frequency band information based on the self-attention mechanism
[0055] The branch information fusion module uses the self-attention mechanism based on the significant position selection algorithm to calculate the similarity between the channels of the feature maps of different branches and adaptively enhance the weights of the key channels, thereby effectively improving the recognition accuracy of the motor imagery task. The structure of this module is as shown in the appendix Figure 4 As shown, the specific implementation process includes the following steps:
[0056] (1) Calculate the squared value of the query matrix: The input data is converted into a query matrix Q through a two-dimensional convolutional layer and reconstructed into a size of [h*w, c], and then the squared value is calculated in the channel dimension to form the Q 2 matrix.
[0057] (2) Obtain the saliency score matrix: Sum Q 2 in the channel dimension to form the saliency score matrix Q pow , which is used to measure the importance of each channel to screen out the most critical feature positions.
[0058] (3) Select the significant positions: Select the largest k values in Q pow and record their positions as index for subsequent matrix construction.
[0059] (4) Output the significant position matrix: Use the channel index of matrix Q pow and the selected significant positions index to construct a new matrix K with a size of [c, k], representing the extracted significant positions.
[0060] (5) Calculate the attention matrix and perform fusion: Based on the significant positions selected by the significant position selection algorithm, construct the attention matrix and perform softmax normalization on it. Combine the numerical matrix with the attention matrix through matrix multiplication to obtain enhanced features of the same size as the input feature map.
[0061] (6) Enhance feature representation: Transform the enhanced features through 1×1 convolution and add them to the original input features to obtain the final fused features.
[0062] (7) Classification prediction: Flatten the fused features and input them into the linear layer for classification to generate the final prediction results.
[0063] This module reduces the data dimension while retaining the attention to the most important information in the feature map, effectively improving the accuracy of motor imagery task recognition.
[0064] In the above model training process, the present invention adopts an experimental method independent of subjects, that is, only the data of a single subject is used for each training and testing. When training based on the BCI Competition IV 2a dataset, the two-stage early stopping method is used as the training method for the experiment, where Session1 is used as the training data and Session2 is used as the test data. To improve the stability and generalization ability of training, the training data is divided into a training set and a validation set according to a ratio of 8:2. Finally, the sample quantity ratio of each subject in the training set, validation set, and test set is 232:56:288.
[0065] During training, multi-class cross-entropy is used as the loss function, and the Adam optimizer is used to update the model parameters. The initial learning rate of the optimizer is set to 0.001, the L2 regularization factor is set to 0.0001, and the remaining parameters are kept at the default settings. To further alleviate the overfitting phenomenon, a maximum norm constraint is imposed on the model weights, where the maximum norm limit of the convolutional layer parameters is set to 1, and the maximum norm limit of the linear layer parameters is set to 0.25.
[0066] The experiments of the present invention are carried out under the Windows 10 operating system. Based on the PyTorch 2.0 deep learning framework, model construction, training, and testing are performed, and the corresponding Python version is 3.8. The experiment uses NVIDIA GeForce RTX 3090 as the computing device to provide efficient training performance.
Claims
1. A motor imagery EEG signal recognition technology that integrates multi-band adaptive filtering and self-attention mechanism. The multi-branch model in this method uses a parallel branch architecture to process the spectrum information of different frequency bands respectively, and constructs a time-space hybrid feature extraction unit in each branch: the action potential timing pattern is captured by time domain convolution, and the spatial distribution characteristics of the cortical potential are extracted by combining spatial domain convolution. Finally, the high-precision analysis of the subject's motor intention is achieved through branch feature fusion based on the self-attention mechanism. Specifically, the following steps are included: Step 1: Preprocessing of EEG data: Select key channels of EEG signals in the data set and preprocess the signals of the selected channels. Step 2: Use the adaptive filtering module to filter the frequency bands of the preprocessed EEG data; use the online bandpass filtering method to filter the EEG data of different frequency bands as the input of the multi-scale feature extraction module. Step 3: Use a multi-scale feature extraction module to extract the time domain and spatial domain features of the EEG signal; Step 4: Complete the fusion of branch frequency band information based on the self-attention mechanism; the branch information fusion module calculates the similarity between feature map channels of different branches through the self-attention mechanism, and automatically gives greater weights to important channels, effectively improving the accuracy of motor imagery task recognition.
2. According to claim 1, the motor imagery EEG signal recognition technology integrating multi-band adaptive filtering and self-attention mechanism is characterized in that: The step one preprocesses the motor imagery EEG signal by first selecting specific electrodes for obtaining effective signals, then smoothing the signals using an exponential moving average algorithm (EMA) to reduce the interference of noise on effective information, and ensuring the validity of the samples by accurately extracting EEG segments that are strongly related to motor imagery intentions. Finally, mean-variance normalization is performed independently on each sample to eliminate amplitude differences caused by attention fluctuations between different trials, thereby ensuring efficient extraction of subsequent effective features. The exponential sliding average algorithm is an algorithm used to smooth and normalize signals, and is often used to process time series data. This method uses weighted averaging to adjust the amplitude of the signal to a more standard range, while retaining the signal's changing trend and reducing mutations caused by noise or electrode movement. The formula for exponential sliding normalization is as follows: The x in the formula t represents the original signal, x' t represents the normalized signal, u t and σ t 2 They represent the current sliding mean and variance respectively, and their calculation methods are shown in formulas (2) and (3). m t =βx t +(1-β)μ t-1 (2) The β in the formula represents the attenuation factor and is set to 0.
001.
3. According to claim 1, the motor imagery EEG signal recognition technology integrating multi-band adaptive filtering and self-attention mechanism is characterized in that: In the step 2, a Sinc convolution layer is introduced to perform online bandpass filtering to filter data in different frequency bands, and the data is used as the input of the multi-scale feature extraction module. Compared with the traditional bandpass filtering method, the Sinc convolution layer is an adaptive bandpass filter that can be embedded in the deep learning model and can dynamically adjust the filtering parameters according to the needs of different branches. The Sinc convolution layer uses two learnable parameters f L and f H To optimize the filtering process, when performing the convolution operation, the convolution kernel parameters are determined by f L and f H It is calculated by the time domain form of the rectangular bandpass filter, as shown in equation (4). k(n)=g(n,f L ,f H )=2f H ×Sinc(2πnf H )-2f L ×Sinc(2πnf H ) (4) Where n represents the nth parameter in the Sinc convolution kernel, f L 、f H Respectively represent the pass frequency and cutoff frequency of the rectangular bandpass filter. The Sinc function is defined as in formula (5). Since the above calculation process is completely differentiable, the filter range parameter f of the Sinc convolution layer is L With f H It can be automatically optimized during the model training process, thereby achieving the effect of adaptively adjusting the filtering range. In addition, in order to improve the performance of the rectangular bandpass filter, a Hamming window is usually used in Sinc convolution to perform windowing on the convolution kernel parameters. The windowing process is shown in formula (6). k w (n)=k(n)×w(n) (6) In the formula, k w (n) represents the convolution kernel parameter after windowing, and w(n) represents the Hamming window function, which is defined as shown in formula (7). w(n)=0.54-0.46×cos (2πn / L) (7) Where L represents the window length, which is the same as the length of the Sinc convolution kernel L sinc same.
4. According to claim 1, the motor imagery EEG signal recognition technology integrating multi-band adaptive filtering and self-attention mechanism is characterized in that: In the step three, the multi-scale feature extraction module aims to fully exploit the time domain and spatial domain features in the EEG signals of each frequency band. The module is mainly composed of parallel time domain convolution, spatial domain convolution and pooling layer. The parallel time domain convolution in the module is used to alleviate the strong individual differences of motor imagery signals. Since there are significant differences in the time domain features corresponding to different subjects under different tasks, this scheme adopts a variety of convolution kernel sizes to adapt to these changes. Specifically, each parallel time domain convolution has the same number of output channels C, but uses different convolution kernel sizes, namely (1, L1), (1, L2), (1, L3), to extract time domain features at different scales. In order to accelerate model convergence, the module adds a batch normalization (BN) layer after each convolution layer. Finally, splicing is performed in the channel dimension to fuse the feature maps generated by the three time domain convolutions. The spatial domain convolution in the module is mainly used to extract the spatial distribution characteristics between various electrodes. The convolution kernel size is set to (E, 1), where E represents the number of electrodes during signal acquisition. Since each channel of the feature map extracted by time domain convolution represents different time domain features, they are independent of each other. Therefore, when performing spatial domain convolution, in order to maintain this independence, the number of convolution groups is set to the number of input feature map channels. At this point, the spatial domain convolution operation is decomposed into multiple independent convolution operations, that is, each input channel is calculated only using the convolution kernel corresponding to it, thereby achieving independent extraction of spatial features of each channel. After completing the extraction of time domain and spatial domain features, the exponential linear unit (ELU) shown in formula (8) is used as the activation function. Finally, the size of the generated feature map is adjusted through average pooling to reduce the subsequent calculation amount of the model and improve the generalization of the model.
5. According to claim 1, the motor imagery EEG signal recognition technology integrating multi-band adaptive filtering and self-attention mechanism is characterized in that: In the step 4, a branch information fusion module based on the self-attention mechanism is used to weight the important channels of the feature map and improve the feature expression ability. This module uses the salient position selection (SPS) algorithm to identify the key channels in the feature map, and realizes cross-channel information fusion by constructing a self-attention matrix. Specifically, the input data is first mapped into a query matrix and a numerical matrix through two two-dimensional convolutional layers, and then the salient position selection (SPS) algorithm is used to calculate the sum of the squares of each channel of the query matrix, and each channel is summed to determine the salient position. According to the salient positions selected by SPS, an attention matrix is constructed and softmax normalized. The numerical matrix is combined with the attention matrix through matrix operations to obtain enhanced features of the same size as the input features. In order to further improve the feature expression effect, the fused result is transformed by a 1×1 convolution and added to the original input features to obtain the fused features. Finally, these features are flattened and input into the linear layer for classification to generate prediction results.