Electroencephalogram motor imagery decoding method based on multi-branch convolutional network

CN122251034BActive Publication Date: 2026-09-29HUNAN UNIV OF SCI & TECH
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202610719758.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2026-05-25
Publication Date
2026-09-29
Estimated Expiration
2046-05-25

AI Technical Summary

Technical Problem

[0005]1)特征尺度缺乏生理学解释:现有多尺度方法的卷积核尺度选择大多基于数据调试与经验设定,难以准确匹配不同脑电生理节律的真实周期特性,从而导致模型往往难以稳定地捕捉到具有神经生理学意义的判别性模式;

Benefits of technology

[0028](1)本发明的基于多分支卷积网络的脑电运动想象解码方法,摒弃了传统方法中依赖经验试错确定卷积核尺度的做法,而是依据δ、θ、μ、β四个关键神经振荡节律的生理中心频率,直接计算各分支时间卷积核的长度。这一设计使得卷积核的感受野与对应频段的实际波形周期相匹配,能够定向、并行地提取不同频段的时空特征。该机制不仅增强了模型的神经生理学可解释性,还显著提高了对事件相关去同步/事件相关同步等关键判别性模式的捕捉稳定性和准确率。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122251034B_ABST
    Figure CN122251034B_ABST
Patent Text Reader

Abstract

The application discloses a kind of based on multi-branch convolution network's electroencephalogram motor imagination decoding method, it is related to electroencephalogram signal processing technical field, decoding method includes: the parallel input of preprocessed electroencephalogram signal respectively corresponding network branch of four physiological frequency bands of delta, theta, mu, beta;Each branch is according to the center frequency of corresponding frequency band to calculate time convolution kernel scale to extract time feature;Energy feature is extracted to the core mu frequency band using fully connected space reinforcement module, and sparse expression is realized to secondary delta, theta, beta frequency band using depth separable convolution module;After each branch feature is dynamically calibrated weight by independent self-attention mechanism and residual concatenation, the decoding result is output by global fusion module.The application based on multi-branch convolution network's electroencephalogram motor imagination decoding method realizes differentiating accurate feature extraction, effectively suppresses redundant noise, and significantly improves the generalization robustness of cross-subject and cross-period.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of electroencephalogram (EEG) signal processing technology, and in particular to a method for decoding EEG motor imagery based on multi-branch convolutional networks. Background Technology

[0002] Brain-computer interface (BCI) technology aims to directly establish communication and control pathways between the brain and the external world by decoding neural activity, providing a novel technological path for stroke rehabilitation, functional assistance, and other fields. The non-invasive EEG paradigm based on motor imagery has become one of the most mainstream research directions in the field of BCI due to its high safety, lack of external stimulation, and portable devices. The neural mechanism of the motor imagery paradigm lies in the fact that motor imagery can induce rhythmic activity changes in the brain's sensorimotor cortex similar to actual movement, mainly manifested as event-related desynchronization and event-related synchronization phenomena in specific frequency bands. Therefore, developing a method that can robustly and accurately decode weak event-related desynchronization / event-related synchronization features representing user intentions from complex EEG background noise with low signal-to-noise ratios, high individual variability, and high non-stationarity is a key technical problem that urgently needs to be solved in this field.

[0003] Currently, deep learning technology, as an end-to-end algorithm, can automatically learn high-dimensional feature representations from raw EEG data without requiring manual feature extraction. Among them, convolutional neural networks (CNNs) have become one of the most widely used methods in current motor imagery decoding research due to their characteristics such as local perception and hierarchical feature extraction. Given the limitations of single-scale CNNs in capturing multi-scale, multi-dimensional neural representations in EEG signals, some studies have shifted to using multi-scale convolutional architectures. These architectures extract multi-dimensional features through parallel convolutional pathways with different receptive fields, aiming to extract more comprehensive discriminative features related to motor intent.

[0004] Current deep learning-based methods for decoding motor imagery in EEG still face two major unresolved issues:

[0005] 1) Lack of physiological explanation for feature scale: The selection of convolution kernel scale in existing multi-scale methods is mostly based on data debugging and empirical settings, which makes it difficult to accurately match the true periodic characteristics of different EEG rhythms. As a result, the model often fails to capture discriminative patterns with neurophysiological significance.

[0006] 2) Lack of frequency-specific feature extraction strategies: The functional information and importance of EEG rhythms in motor imagery vary significantly across different frequency bands. Existing multi-scale models typically apply a uniform spatial feature extraction network to all frequency band branches, failing to perform targeted enhancement extraction of energy representations for core motor frequency bands. Furthermore, the indiscriminate extraction of secondary auxiliary frequency bands not only easily leads to excessive model parameters but also introduces a large amount of background noise. This feature extraction approach, lacking a distinction between primary and secondary bands, increases the risk of overfitting and severely limits its robustness and generalization ability across subjects and time periods. Summary of the Invention

[0007] To address the shortcomings of existing technologies, this invention provides a method for decoding EEG motor imagery based on multi-branch convolutional networks, which achieves differentiated and accurate feature extraction, effectively suppresses redundant noise, and improves the generalization robustness of the model across subjects and time periods.

[0008] To solve the above-mentioned technical problems, the technical solution proposed by this invention is as follows:

[0009] A method for decoding EEG motor imagery based on multi-branch convolutional networks includes the following steps:

[0010] Step S1: Obtain the raw EEG signal of the single test to be decoded and perform data preprocessing;

[0011] Step S2: The preprocessed raw EEG signal is input into a pre-trained attention-enhancing multi-branch convolutional network. The network contains four parallel processing branches corresponding to the δ band of the post-exercise fatigue state, the θ band of the exercise task preparation state, the μ band of the contralateral inhibition state, and the β band of the exercise maintenance and recovery state, respectively.

[0012] Step S3: In each of the parallel processing branches, the temporal features of the preprocessed raw EEG signal are extracted by a multi-scale temporal convolution module, wherein the temporal convolution kernel scale of each branch is calculated based on the center frequency of the corresponding physiological rhythm.

[0013] Step S4: For the extracted temporal features, a differential spatial feature extraction module is used to transform the feature dimension. Specifically, for the branch containing the μ frequency band, a fully connected spatial enhancement module is used to extract energy features. For the branches containing the δ, θ, and β frequency bands, a spatially depth-separable convolution module is used to perform feature sparsification.

[0014] Step S5: The spatial features output by each branch are sent to an independent self-attention mechanism module to dynamically calculate the attention weights at each time point, and the attention-weighted features are then concatenated with the original spatial features.

[0015] Step S6: The spliced ​​features output from the four branches are fed into the global feature fusion module along the feature map dimension for fusion, and the decoding and classification results of the motion imagination task are output by the classifier.

[0016] A further improvement to the above technical solution is as follows:

[0017] Preferably, the data preprocessing in step S1 specifically includes: using a bandpass filter to retain EEG neural activity signals in the 0.5 to 30 Hz frequency band, and performing Z-Score normalization on the filtered data to convert it into a standard normal distribution.

[0018] Preferably, the formula for the length scale of the temporal convolution kernel in step S3 is:

[0019]

[0020] in, To represent different processing branches; To cover the adjustment factor of the waveform period, The length scale of the convolution kernel corresponding to the branch time is... Sampling frequency, This is the approximate target center frequency.

[0021] Preferably, the fully connected spatial enhancement module for the branch containing the μ band in step S4 includes the following steps in sequence: extracting spatial correlation features across all EEG channels through a fully connected spatial convolutional layer; performing a square operation on the convolutional output to extract signal intensity; extracting spatiotemporal envelope information along the time dimension through an average pooling layer; and finally converting the features into a power spectral density representation through a logarithmic transformation.

[0022] Preferably, the spatial depth separable convolution module for the branches containing the δ, θ, and β frequency bands in step S4 includes the following processes in sequence: independently extracting spatial filtering features within each channel through depthwise convolution; performing linear combination and information integration across feature maps in the channel dimension through pointwise convolution; and finally outputting the processed feature map through an average pooling layer.

[0023] Preferably, the calculation process of the self-attention mechanism module in step S5 is as follows: the input spatial features are linearly transformed to generate a query matrix, a key matrix, and a value matrix respectively; then, the attention distribution weights at each time point are obtained through the dot product attention calculation formula to generate attention-weighted features.

[0024] Preferably, in step S5, the attention-weighted features and the original spatial features are concatenated along the channel dimension at the residual level to preserve local details and incorporate global contextual information.

[0025] Preferably, the global feature fusion module in step S6 is specifically used to: globally concatenate and stitch the enhanced features output from the four branches along the feature map dimension; use convolutional layers for cross-frequency band information interaction; reduce the dimensionality of each feature map into a single value through a global average pooling layer to form a feature vector; and finally output the classification probability by a fully connected classifier.

[0026] Preferably, the decoding and classification results of the motor imagery task include motor imagery categories for the left hand, right hand, foot, or tongue.

[0027] The EEG motor imagery decoding method based on multi-branch convolutional networks provided by this invention has the following advantages compared with existing technologies:

[0028] (1) The EEG motor imagery decoding method based on multi-branch convolutional networks of the present invention abandons the traditional approach of relying on trial and error to determine the scale of the convolutional kernel. Instead, it directly calculates the length of the temporal convolutional kernel of each branch based on the physiological center frequencies of the four key neural oscillation rhythms δ, θ, μ, and β. This design makes the receptive field of the convolutional kernel match the actual waveform period of the corresponding frequency band, enabling the directional and parallel extraction of spatiotemporal features of different frequency bands. This mechanism not only enhances the neurophysiological interpretability of the model, but also significantly improves the stability and accuracy of capturing key discriminative patterns such as event-related desynchronization / event-related synchronization.

[0029] (2) The EEG motor imagery decoding method based on multi-branch convolutional networks of the present invention designs an asymmetric spatial feature extraction module to address the differences in the functional contributions of different frequency bands in motor imagery tasks. For the μ frequency band, which carries the core motor imagery information, a fully connected spatial convolution combined with square and logarithmic transformations is used to transform the features into power spectral density representations with clear physical meaning, thereby strengthening the feature representation of event-related desynchronization / event-related synchronization. For the δ, θ, and β frequency bands, which play an auxiliary role, a parameter-efficient depthwise separable convolution is used for sparse representation, which significantly reduces the number of model parameters and suppresses background noise interference while retaining effective information.

[0030] (3) The EEG motor imagery decoding method based on multi-branch convolutional networks of the present invention introduces a self-attention mechanism independently in each frequency band branch, enabling the model to dynamically calibrate the importance weights of features at each time point, thereby accurately locating instantaneous discriminative neural activity in strong background noise and non-stationary EEG signals. At the same time, the self-attention weighted features are concatenated with the original spatial features at the residual level, preserving both local details and global contextual information. Attached Figure Description

[0031] Figure 1 This is a schematic diagram of the overall architecture of the multi-branch convolutional network of the present invention.

[0032] Figure 2 This is a schematic diagram of the self-attention mechanism in an embodiment of the present invention.

[0033] Figure 3 To visualize the feature space of different models on the BCIC-IV-2a dataset in the experiment, (a) is the feature distribution of model EEGNet, (b) is the feature distribution of model Deep ConvNet, (c) is the feature distribution of model Shallow ConvNet, (d) is the feature distribution of model EISATC-Fusion, (e) is the feature distribution of model ATCNet, and (f) is the feature distribution of model AEMBCNet of this invention. Detailed Implementation

[0034] The following provides a detailed description of specific embodiments of the present invention. It should be understood that the specific embodiments described herein are for illustrative and explanatory purposes only and are not intended to limit the scope of the invention.

[0035] like Figure 1 and Figure 2 As shown, the EEG motor imagery decoding method based on multi-branch convolutional networks of the present invention preprocesses the original motor imagery EEG signal using filters and standardization techniques during the EEG signal decoding process; extracts temporal features of different physiological frequency bands using multi-scale temporal convolutional blocks; strengthens core frequency band features and sparsifies secondary frequency band features using a differentiated spatial feature extraction strategy; and dynamically calibrates feature weights and performs residual splicing using a self-attention mechanism. Finally, the fused multidimensional hybrid features are used to complete the classification task of motor imagery. Figure 1 The C in the text represents Feature splicing.

[0036] The EEG motor imagery decoding method based on multi-branch convolutional networks of the present invention is performed in a decoding system, including:

[0037] The preprocessing module is used to acquire the raw EEG signals from a single trial to be decoded and to perform data preprocessing.

[0038] The attention-enhanced multi-branch convolutional network contains four parallel processing branches corresponding to the δ, θ, μ, and β frequency bands, respectively.

[0039] The global feature fusion module is used to fuse the spliced ​​features output from the four branches along the feature map dimension;

[0040] A classifier is used to output the decoding and classification results of the motion imagery task.

[0041] Each of the parallel processing branches includes:

[0042] The multi-scale temporal convolution module is used to calculate the temporal convolution kernel scale based on the center frequency of the corresponding physiological rhythm and to extract temporal features from the preprocessed EEG signal.

[0043] The differential spatial feature extraction module includes:

[0044] For the branch containing the μ band, a fully connected spatial enhancement module is used to extract energy features;

[0045] For the branches containing the δ, θ, and β frequency bands, spatially depth-separable convolutional modules are used for feature sparsification representation.

[0046] The self-attention mechanism module is used to dynamically calculate the attention weights at each time point and concatenate the attention-weighted features with the original spatial features.

[0047] The preprocessing module uses a bandpass filter to retain EEG neural activity signals in the 0.5 to 30 Hz frequency band, and performs Z-Score normalization on the filtered data to convert it into a standard normal distribution.

[0048] The global feature fusion module is used to globally cascade and concatenate the enhanced features output from the four branches along the feature map dimension; it uses convolutional layers to perform cross-frequency band information interaction; it uses global average pooling layers to reduce the dimensionality of each feature map into a single value to form a feature vector; and finally, a fully connected classifier outputs the classification probability.

[0049] The EEG motor imagery decoding method based on multi-branch convolutional networks in this embodiment specifically includes the following steps:

[0050] Step S1: Obtain the raw EEG signal of the single test to be decoded and perform data preprocessing;

[0051] In practical applications or model training, the specific process of acquisition and preprocessing is as follows:

[0052] S1-1 uses a non-invasive EEG acquisition device to acquire raw EEG data from a single trial when the subject performs a motor imagery task (such as imagining the left hand, right hand, foot, or tongue), and extracts signals from a specific time window (such as 0-3 seconds) during the motor imagery task.

[0053] S1-2, In order to suppress high-frequency muscle artifacts and low-frequency baseline drift, a bandpass filter is used to filter the intercepted EEG signal. In this embodiment, it is preferable to retain the EEG neural activity signal in the 0.5-30Hz frequency band.

[0054] S1-3, to reduce the high non-stationarity of EEG signals across time periods and subjects, Z-Score standardization was performed on the filtered EEG signal data to transform it into a standard normal distribution with zero mean and unit variance. The standardization calculation formula is as follows:

[0055] (1)

[0056] in, This represents the filtered EEG signal data. This represents standardized electroencephalogram (EEG) signal data. This represents the mean of the data segment. This represents the standard deviation of the data segment. After standardized preprocessing, the EEG signals are used to construct the network's input features. For a single trial, the signal is represented as a matrix. ,in For the feature space, This refers to the number of channels used for EEG acquisition. This represents the number of sampling points in the time dimension.

[0057] Step S2: The preprocessed raw EEG signal is input into a pre-trained attention-enhanced multi-branch convolutional network, which constructs four branches with clear physiological mappings in parallel.

[0058] The standardized EEG signals were divided into training and validation sets. On the training set, an attention-enhanced multi-branch convolutional network was pre-trained for 30 rounds using the Adam optimizer, a learning rate of 1e-3, and a batch size of 16. This network contains four parallel branches, corresponding to the four physiological frequency bands δ, θ, μ, and β. In this invention, each branch first extracts features through an independent convolutional module, then dynamically calibrates the feature weights via a self-attention mechanism, and finally integrates and classifies the multi-frequency features through residual concatenation and a global fusion module.

[0059] The four parallel branches correspond to four core EEG rhythms with clear neurophysiological significance in the human brain. The delta band corresponding to the post-exercise fatigue state contains 0.5-4 Hz frequency information, the theta band corresponding to the exercise task preparation state contains 4-8 Hz frequency information, the μ band corresponding to the contralateral inhibition state of exercise contains 8-13 Hz frequency information, and the β band corresponding to the exercise maintenance and recovery state contains 13-30 Hz frequency information.

[0060] Step S3: To map the EEG signal from the total frequency band feature space to the multi-branch sub-feature space, a multi-scale temporal convolution module is designed. In each of the parallel processing branches, temporal features are extracted from the standardized EEG signal through temporal convolution networks of different scales. The temporal convolution kernel scale of each branch is calculated based on the center frequency of the corresponding physiological rhythm.

[0061] In this embodiment, the multi-scale temporal convolution module enhances the accuracy of feature extraction by establishing a direct mapping between the receptive field of the convolution kernel and the neurophysiological rhythm cycle. The specific implementation process is as follows:

[0062] S3-1, Determine the approximate target center frequency of the physiological rhythm corresponding to the four branches. The target center frequencies for the δ, θ, μ, and β bands are set to 2Hz, 5Hz, 10Hz, and 25Hz, respectively.

[0063] S3-2, based on the sampling frequency of the EEG signal With approximate target center frequency Calculate the length scale of the temporal convolution kernel for each branch. The calculation formula is as follows:

[0064] (2)

[0065] in, To represent different processing branches; To cover the adjustment factor of the waveform period, this embodiment preferably uses... ;Regulatory factor The value used to control the number of waveform cycles covered by the convolution kernel can be adjusted between 0.5 and 2 according to the signal characteristics. In this embodiment, 1.0 is used to cover a single complete cycle.

[0066] In this embodiment, when the sampling frequency The calculated temporal convolution kernel lengths corresponding to the δ, θ, μ, and β frequency bands are 125, 50, 25, and 10, respectively (corresponding to physical time spans of 0.5 seconds, 0.2 seconds, 0.1 seconds, and 0.04 seconds, respectively).

[0067] S3-3, calculates the scale of multiple sets of convolution kernel pairs input matrices. Parallel temporal convolution operations are performed to initially extract spatiotemporal features across different frequency bands. This multi-scale temporal convolution process is represented as follows:

[0068] (3)

[0069] in, Indicates different processing branches, Indicates a temporal convolutional layer. These are the temporal convolutional layer weights for the corresponding branches. The time feature matrix output by the corresponding branch has its feature dimensions determined by the input. Transform into ,in The number of feature layers is determined by the number of convolutional kernels set.

[0070] Step S4: For the extracted temporal features, a differential spatial feature extraction module is used to transform the feature dimension. This module is a differential feature extraction strategy that balances discriminative power and lightweight. It adopts asymmetric processing for the physiological importance of different frequency bands. Since the branch containing the μ frequency band contains the most significant features of the motion imagery task, a fully connected spatial enhancement module is used to extract energy features. For the branches containing the δ, θ and β frequency bands, a spatial depth separable convolution module is used to perform feature sparsification representation.

[0071] To address the differences in the contribution of different physiological frequency bands to motion imagery decoding, an asymmetric spatial feature extraction strategy was designed. The specific implementation process is as follows:

[0072] S4-1, targeting the μ-band of contralateral motor inhibition, employs a fully connected spatial enhancement network. Specifically: first, fully connected spatial convolutional kernels are used to extract spatial correlation features of electrode locations across all EEG channels; then, signal intensity features of each feature map are extracted through squaring operations; next, average pooling layers are used to smooth the temporal dimension and extract spatiotemporal envelope information; finally, logarithmic transformation is used to compress the feature distribution, converting the signal features into a physiologically meaningful power spectral density representation. The calculation process is expressed in the following formula:

[0073] (4)

[0074] In the formula, It is a fully connected spatial convolutional layer. For average pooling layer, This is the time characteristic matrix of the μ-band branch. Spatial characteristics of μ-band output, These are the weights of the μ-band branch spatial convolutional layer. Its output feature dimension is determined by the input... Transform into ,in This represents the time length after pooling. This process significantly enhances the characterization ability of event-related desynchronization / event-related synchronization phenomena.

[0075] S4-2 employs a spatially depth-separable convolutional network to process the frequency bands for the post-exercise fatigue state (δ band), the exercise task preparation state (θ band), and the exercise maintenance and recovery state (β band), which play an auxiliary role. Specifically, spatial filtering features are first extracted independently within each channel using spatial depth-separate convolutions. Then, 1×1 pointwise convolutions are used to perform linear combination and information integration across feature maps along the channel dimension. Finally, the processed feature map is output after an average pooling layer. The feature output is represented as follows:

[0076] (5)

[0077] In the formula, and These represent depthwise convolutional and pointwise convolutional layers, respectively. This is the time feature matrix for the corresponding branch. and These are the weights of the depthwise convolutional layer and the pointwise convolutional layer, respectively. For the corresponding Output characteristics of the branch. Spatial characteristics of the output. Dimensions are determined by the input Transform into ,in This represents the time duration after pooling. This processing effectively captures rich multi-band features while significantly reducing the number of model parameters, thus suppressing background noise interference.

[0078] S4-3, Features after different spatial processing methods and Each branch is still output as an independent feature to the next level, and the feature dimension of each branch is 1. In addition, and Concatenating along the feature number dimension, the formula is expressed as:

[0079] (6)

[0080] in, The concatenated features are processed in a differential space, with the dimensionality changing from the input's four dimensions. Feature map transformation , This is a feature dimension concatenation operation. They are respectively The branch is the output feature after spatial feature extraction.

[0081] Step S5: In order to avoid interference from cross-frequency features of multiple branches and enable the model to accurately focus on key sample points within a specific rhythm, the spatial features output by each branch are sent to an independent self-attention mechanism module to dynamically calculate the attention weights at each time point and perform residual-level splicing.

[0082] The spatial features output in step S4 and After linear transformation, each branch generates a query matrix independently. Key matrix Sum matrix The linear transformation is achieved through a learnable weight matrix, which maps the input features to different representation spaces.

[0083] The attention distribution weights at each time point are obtained using the dot product attention calculation formula, which is expressed as follows:

[0084] (7)

[0085] in, The output features of the self-attention mechanism module have an output dimension of . , This is the scaling factor for the dimensions of the feature vectors (usually the dimensions of the key matrix). This is a self-attention mechanism operation. For the softmax function, This represents the time length after pooling. This process is achieved through calculation. and The similarity between them dynamically assigns higher weight to discriminative neural activities.

[0086] Step S6, the output features of the self-attention mechanism module Concatenating along the feature map dimensions, and concatenating features after differential space processing. They are fed together into the global feature fusion module for fusion, and the classifier outputs the decoding and classification results of the motion imagination task.

[0087] This step aims to integrate spatiotemporal information across the entire frequency band to complete the final classification task. The specific implementation process is as follows:

[0088] S6-1, the output features of the self-attention mechanism module are concatenated and stitched together along the feature map dimension. The stitched features... The formula is expressed as follows:

[0089] (8)

[0090] in, The features are concatenated along the feature map dimension, with the dimension size ranging from... Transform into , This is a feature dimension concatenation operation. They are respectively The output characteristics of the branch after the attention enhancement module.

[0091] S6-2 involves secondary splicing and fusion of spatial and attentional features, employing a key step of integrating multi-band and multi-dimensional information.

[0092] Specifically, what will be obtained and The feature maps are concatenated along their channel dimensions. Before concatenation, each feature has a dimension of 1. ,in The sum of the number of feature map channels representing the four frequency band branches. The concatenation operation expands this dimension to... This aggregates all information from all branches, after spatial filtering and attention weighting, along the channel dimension.

[0093] Subsequently, a one-dimensional convolutional layer is used to perform cross-channel fusion and compression of the concatenated features. This operation aims to achieve cross-frequency band feature interaction and generate an information-rich fused representation. This process is formally represented as:

[0094] (9)

[0095] in, This indicates the output characteristics of cross-channel fusion. This is a one-dimensional convolutional layer, and its output feature number is designed to be equal to the number of input features. Therefore, the output features Dimensions , This is a splicing operation along the channel dimension of the feature map.

[0096] Finally, the feature vector The data is fed into a fully connected classifier, and the softmax function is used to output the classification probability of the motion imagery task. Its formula is expressed as follows:

[0097] (10)

[0098] In the formula, It is a global average pooling layer. and These represent the weight matrix and bias term of the fully connected classifier. In this embodiment, the output is one of four motor imagery categories: left hand, right hand, foot, or tongue.

[0099] Experimental verification:

[0100] To further verify the effectiveness of the proposed method (AEMBCNet) and its advantages over existing technologies, tests were conducted on two motor imagery datasets with different characteristics. BCI CompetitionIV Dataset 2a (BCIC-IV-2a) is a commonly used four-class classification task dataset, while BCI CompetitionIV Dataset2b (BCIC-IV-2b) is a two-class classification dataset containing only a small number of channels. These are two classic publicly available brain-computer interface motor imagery datasets, widely used as benchmark datasets for algorithm validation and performance evaluation in the field of brain-computer interface research.

[0101] The BCIC-IV-2a dataset contains 22 channels of EEG signals from 9 participants, covering four imagery tasks: left hand, right hand, tongue, and foot. The experiment was conducted in two phases, with 72 trials per task per phase, totaling 288 trials. The signal sampling rate was 250 Hz, and the amplification resolution was 100 µV. Data from phase 1 was used for model training, and data from phase 2 was used for performance testing.

[0102] The BCIC-IV-2b dataset records EEG data from nine participants across five time slots. The first two slots are feedback-free experiments, while the latter three include feedback. Each slot contains 120 binary classification tasks, including left-hand and right-hand tasks, with each trial lasting 4 seconds. Recording electrodes include C3, Cz, and C4, with a sampling rate of 250 Hz. Typically, slots 1, 2, and 3 are used as training data, while slots 4 and 5 are used as test data.

[0103] All original channels of both datasets were retained (22 and 3 channels respectively), and the sampling rate was maintained at 250 Hz. In the motion visualization task, a 0-3s signal segment was extracted for processing, corresponding to 750 sampling points on the time axis. The data underwent a 1-40Hz bandpass filter to suppress high-frequency artifacts. The standardized data was then input into the model after splitting into training and test sets.

[0104] To objectively evaluate the performance of AEMBCNet, this validation selected five representative deep learning models as evaluation baselines:

[0105] EEGNet, a general-purpose, compact convolutional neural network, is based on the introduction of depthwise separable convolution, which effectively separates the temporal convolution from the spatial mapping between channels. While maintaining feature extraction capabilities, it improves the model's generalization performance on raw EEG data.

[0106] Deep ConvNet: This model references deep convolutional design in computer vision, consisting of multiple nonlinear transformations and progressive receptive field expansion modules. With sufficient samples, this model demonstrates strong feature modeling capabilities.

[0107] Shallow ConvNet: This architecture consists of large temporal and spatial convolutional layers, and uses squared operations, average pooling, and logarithmic activation to simulate power spectral density feature extraction. Due to its small parameter count and optimization for sensorimotor rhythms, this model exhibits good stability when processing small sample EEG data.

[0108] EISATC-Fusion: This is a multi-module fusion model that combines multi-scale receptive fields with attention mechanisms. Its design incorporates parallel convolutional kernels from the Inception module with multi-head self-attention mechanisms to address the uncertainties in the time-frequency domain of EEG signals. The model captures long-range dependencies through temporal convolutional networks, focusing on resolving signal non-stationarity and inter-subject variability.

[0109] ATCNet: This algorithm employs an attention-based temporal convolutional network, embedding self-attention modules within a temporal convolutional framework. Its mechanism allows the model to dynamically adjust the weights of different sampling points based on task relevance, thereby locating transient neural activities relevant to the motion task. ATCNet combines residual connections and sliding window techniques to balance local feature extraction with understanding of the global context.

[0110] To verify the model's generalization ability in the MI-EEG decoding task, in-subject intra-time and cross-time analyses were conducted on two datasets. In the intra-time analysis, 80% of the training time data was used for model training and 20% for validation. In the cross-time analysis, time period 1 of BCIC-IV-2a was used for training and time period 2 for testing; the first three time periods of BCIC-IV-2b were used for training and the last two time periods for testing. Five-fold cross-validation was used in the experiments, and the final performance was evaluated based on the average decoding accuracy across all participants.

[0111] In the comparative experiments, all models were trained for a uniform number of epochs of 1000, using the Adam optimizer and cross-entropy loss function. The batch size was set to 16, and the learning rate was 0.0005. All experiments were conducted on an Intel Core i5-12400F processor and an NVIDIA RTX 4060Ti GPU, using the PyTorch framework and Python 3.12 for both training and testing.

[0112] (a) Performance comparison

[0113] (1) Analysis of experimental results within the time period

[0114] The experimental results on the BCIC-IV-2a dataset are shown in Table 1. The AEMBCNet model proposed in this invention achieved the best average classification accuracy of 87.03% among all compared methods. Specifically, among the 9 subjects, AEMBCNet achieved the best or tied-best accuracy on subjects S03, S07, and S08, especially reaching a maximum of 96.53% on subject S07, outperforming other methods and demonstrating its ability to capture highly discriminative features. Compared with the second-best performing ATCNet (average 86.57%), AEMBCNet performed better on more than half of the subjects. Although its standard deviation of 8.00 is slightly higher than ATCNet, considering its higher average performance, this indicates that AEMBCNet improved overall accuracy without introducing excessive performance fluctuations. It is worth noting that AEMBCNet still maintained a stability of 74.32% and 74.97% on subjects S02 and S06, which are generally more difficult to identify, and is competitive with the baseline model. This indicates that it also has a certain robustness to EEG signals with low signal-to-noise ratio or special individual patterns.

[0115] Table 1. Comparison of classification accuracy within the same time period in the BCIC-IV-2a dataset.

[0116]

[0117] On the BCIC-IV-2b dataset, the AEMBCNet model achieved an average classification accuracy of 87.49%, the best result among all compared methods, as shown in Table 2. The model's advantages are specifically reflected in its superior individual performance and stability. Among the nine participants, AEMBCNet achieved the best accuracy on more than half (S01, S04, S05, S06, S09). Particularly on participants S01 and S05, its performance of 91.35% and 91.77% represents significant improvements of over 8% and 20% respectively compared to the worst-performing baseline model, demonstrating its effective decoding of difficult patterns with significant individual differences. More importantly, AEMBCNet achieved the lowest standard deviation of 5.59 among all models, indicating minimal performance fluctuations and the highest generalization stability across different participants. Although its accuracy was slightly lower than some baselines on participants S03 and S07, the differences were small, less than 2%, and did not constitute a significant disadvantage.

[0118] Table 2 Comparison of classification accuracy within the same time period in the BCIC-IV-2b dataset

[0119]

[0120] A comprehensive analysis of the results on both datasets demonstrates AEMBCNet's consistent lead in average classification accuracy, validating the effectiveness of its network architecture in extracting discriminative spatiotemporal features of motor imagery EEG. The performance improvements achieved by the model on most subjects, including those who performed poorly on some traditional methods, further prove its excellent generalization ability.

[0121] (2) Analysis of experimental results across time periods

[0122] In the more challenging cross-time period experimental setting, AEMBCNet achieved an average classification accuracy of 80.52% on the BCIC-IV-2a dataset, outperforming other comparative methods, as shown in Table 3. Specifically, its performance is approximately 4.5% higher than the second-ranked Deep ConvNet's 75.96%, highlighting the model's effectiveness in handling signal non-stationarity across time periods. Although AEMBCNet's standard deviation of 9.12 is higher than ATCNet, which focuses on time period adaptation, it is lower than Shallow ConvNet and DeepConvNet, indicating that its performance fluctuations are within an acceptable range and that it is more adaptable to unstable subjects. Looking at individual subjects, AEMBCNet performs particularly well on S03, S07, and S08, achieving accuracies of 94.1%, 88.89%, and 85.76%, respectively, maintaining a leading or near-optimal performance on most subjects. Especially on difficult subjects such as S02 and S06, its accuracy of 63.54% and 68.4% is better than the baseline. This indicates that the model's multi-scale feature fusion mechanism may better capture relatively stable neural patterns across time periods and reduce the dependence on instantaneous features of specific sessions, thus showing stronger robustness in the face of time period variations.

[0123] Table 3 Comparison of classification accuracy across time periods in the BCIC-IV-2a dataset

[0124]

[0125] On the BCIC-IV-2b dataset, AEMBCNet also achieved a top average accuracy of 87.26%, as shown in Table 4. Further analysis reveals its strengths in key individual breakthroughs and overall robustness. Among the nine participants, AEMBCNet achieved best performance on S01, S04, S07, and S09. Particularly on S01 participants, its accuracy of 78.24% significantly outperformed most baselines, such as EEGNet's 67.19%, representing an improvement of over 10%, demonstrating its effective adaptation to changes in feature distribution across time periods for some participants. Although its performance was slightly inferior to some best baselines on S03 and S05 participants, this is common in complex cross-participant EEG decoding, and its leading advantage on most other participants compensated for this minor fluctuation. In terms of stability, AEMBCNet's standard deviation was 8.45, placing it in the middle range among all models, lower than EEGNet's 10.12 but higher than Deep ConvNet's 7.86. Considering that it achieved the highest average accuracy, this level of fluctuation is within an acceptable range, indicating that the model significantly improves cross-time decoding performance without sacrificing stability.

[0126] Table 4. Comparison of classification accuracy across time periods in the BCIC-IV-2b dataset.

[0127]

[0128] (ii) Ablation test

[0129] To verify the effectiveness of each module in the attention-enhanced multi-branch convolutional network proposed in this invention, a systematic intra-subject cross-time ablation experiment was conducted on the BCIC-IV-2a dataset. The experiment evaluated the impact of different network configurations on the final decoding performance by sequentially removing or replacing key components in the model. The specific settings are as follows:

[0130] Multi-scale temporal convolution (MSTC): The different scale temporal convolution kernels of the four branches are uniformly replaced with a single scale to verify the necessity of multi-scale temporal feature extraction.

[0131] Spatial Convolution Weight Learning (SCWSL): This study removes the squared logarithmic nonlinear operation from this branch and explores its role in energy feature extraction.

[0132] Spatial Sparse Mapping (SSPC): Replace depthwise separable pointwise convolution with standard spatial convolution and evaluate its performance in feature selection.

[0133] Multi-branch self-attention: Removes the independent self-attention layers for each branch and directly uses spatial convolutional features for classification.

[0134] Attention Enhancement: Cancels the concatenation operation between attention-weighted features and original features, and evaluates the impact of feature fusion on global representation.

[0135] The experimental results are shown in Table 5. The complete model AEMBCNet exhibits the best average performance, indicating that each component positively contributes to the decoding task. Removing the SCWSL module reduces the accuracy to 72.96%, demonstrating that the squared logarithm operation plays a crucial role in enhancing the representation of energy features related to motion imagery. Removing the SSPC module leads to an increase in standard deviation, indicating that the deep separable structure is significant in filtering effective features and suppressing redundant information, thus contributing to improved model stability. Furthermore, the absence of the attention mechanism and its enhanced fusion strategy both result in varying degrees of accuracy reduction, proving the effectiveness of self-attention in focusing on key spatiotemporal features and enriching global contextual information.

[0136] Table 5. Results of cross-time ablation experiments using the AEMBCNet model on the BCIC-IV-2a dataset.

[0137]

[0138] (III) Visualization Comparison Results of Feature Space

[0139] To intuitively compare the discriminative power of the features learned by different models, the last layer features of each model on the test data of subject S07 in the BCIC-IV-2a dataset were extracted, and the t-SNE method was used to reduce the dimensionality to two-dimensional space for visualization. The results are as follows: Figure 3 As shown in (a) to (f).

[0140] The scatter dots in the diagram represent four types of motion visualization tasks: left hand (red), right hand (blue), both feet (green), and tongue (orange). In feature visualization, ideal classification features should satisfy intra-class compactness and inter-class separation.

[0141] from Figure 3 As can be observed in (a) to (f), the feature distributions of different models differ. In the AEMBCNet subgraph proposed in this invention ( Figure 3 In (f), the point clusters representing the four task classes show relatively clearer separation boundaries in two-dimensional space and higher intra-class clustering. In contrast, other baseline models ( Figure 3 In the feature maps (a)-(e), there is relatively more overlap between point clusters of different categories, especially in the region corresponding to the hand motion imagery task (red and blue dots). This visualization result suggests that the features extracted by the AEMBCNet model may have stronger discriminative power in distinguishing different motion imagery tasks.

[0142] The above embodiments are merely preferred examples of the present invention and are not intended to limit the present invention in any way. Although the present invention has been disclosed above with reference to preferred embodiments, it is not intended to limit the present invention. Therefore, any simple modifications, equivalent changes, and alterations made to the above embodiments based on the technical essence of the present invention without departing from the scope of the present invention should fall within the protection scope of the present invention.

Claims

1. A method for decoding EEG motor imagery based on multi-branch convolutional networks, characterized in that, Includes the following steps: Step S1: Obtain the raw EEG signal of the single test to be decoded and perform data preprocessing; Step S2: The preprocessed raw EEG signals are input into a pre-trained attention-enhancing multi-branch convolutional network. The network includes frequency bands corresponding to the post-exercise fatigue state and the exercise task preparation state, respectively. Frequency band, contralateral inhibition state of motion Frequency band and motion maintenance and recovery state The frequency band has four parallel processing branches; The frequency band includes frequency information from 0.5 to 4 Hz. The frequency band includes 4-8Hz frequency information, the The frequency band contains 8-13Hz frequency information, the The frequency band includes frequency information from 13 to 30 Hz; Step S3: In each of the parallel processing branches, the preprocessed raw EEG signal is subjected to temporal feature extraction using a multi-scale temporal convolution module. The temporal convolution kernel scale of each branch is calculated based on the center frequency of the corresponding physiological rhythm. The formula for the length scale of the temporal convolution kernel of each branch is: ; in, To represent different processing branches; To cover the adjustment factor of the waveform period, , The length scale of the convolution kernel corresponding to the branch time is... Sampling frequency, To approximate the target center frequency; when the sampling frequency At that time, the corresponding result was calculated. , , , The temporal convolution kernel lengths for the frequency bands are 125, 50, 25, and 10, respectively, corresponding to physical time spans of 0.5 seconds, 0.2 seconds, 0.1 seconds, and 0.04 seconds. Step S4: For the extracted temporal features, the feature dimension transformation is performed using the differential spatial feature extraction module. Specifically, for... In the branch containing the frequency band, a fully connected spatial enhancement module is used to extract energy features. The fully connected spatial enhancement network process is as follows: first, fully connected spatial convolutional kernels are used to extract the spatial correlation features of electrode locations across all EEG channels; then, the signal feature intensity of each feature map is extracted through squaring operations; next, average pooling layers are used to smooth the temporal dimension and extract spatiotemporal envelope information; finally, logarithmic transformation is used to compress the feature distribution, transforming the signal features into a physiologically meaningful power spectral density representation; for , and In the branch containing the frequency band, spatially depth-separable convolutional modules are used for sparse feature representation. Step S5: The spatial features output by each branch are sent to an independent self-attention mechanism module to dynamically calculate the attention weights at each time point. The attention-weighted features are then concatenated with the original spatial features along the channel dimension at the residual level to preserve local details and incorporate global contextual information. Step S6: The concatenated features output from the four branches are fed into the global feature fusion module along the feature map dimension for fusion. The classifier outputs the decoding classification result of the motion imagination task. The enhanced features output from the four branches are globally concatenated and concatenated along the feature map dimension. Convolutional layers are used for cross-frequency band information interaction. The global average pooling layer reduces the dimensionality of each feature map into a single value to form a feature vector. Finally, the fully connected classifier outputs the classification probability.

2. The EEG motor imagery decoding method based on multi-branch convolutional networks according to claim 1, characterized in that, The data preprocessing in step S1 specifically includes: using a bandpass filter to retain EEG neural activity signals in the 0.5 to 30 Hz frequency band, and performing Z-Score normalization on the filtered data to convert it into a standard normal distribution.

3. The EEG motor imagery decoding method based on multi-branch convolutional networks according to claim 1, characterized in that, In step S4, the target is , and The spatial depth separable convolutional module of the frequency band branch has the following processing steps: extracting spatial filtering features independently within each channel through depthwise convolution; performing linear combination and information integration across feature maps in the channel dimension through pointwise convolution; and finally outputting the processed feature map through an average pooling layer.

4. The EEG motor imagery decoding method based on multi-branch convolutional networks according to claim 3, characterized in that, The calculation process of the self-attention mechanism module in step S5 is as follows: the input spatial features are linearly transformed to generate a query matrix, a key matrix, and a value matrix, respectively; then the attention distribution weights at each time point are obtained through the dot product attention calculation formula to generate attention-weighted features.

5. The EEG motor imagery decoding method based on multi-branch convolutional networks according to claim 1, characterized in that, The decoding and classification results of the motor imagery task include motor imagery categories for the left hand, right hand, foot, or tongue.

Citation Information

Patent Citations

  • Brain-computer interface coding and decoding method and system and electronic equipment

    CN118069980A

  • Motor imagery electroencephalogram signal classification method based on multi-scale convolution and self-attention

    CN118349906A

  • Motor imagery electroencephalogram decoding method based on multi-view space-time convolution and attention mechanism

    CN121881085A