Motor imagery electroencephalogram decoding method and system
By combining signal enhancement based on the Transformer network architecture with graph convolutional networks, the problems of signal quality and cross-subject adaptability in motor imagery EEG decoding technology are solved, and high-precision motor imagery task classification and EEG signal feature representation are achieved.
Patent Information
- Application Number
- CN202511390031.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-09-26
- Publication Date
- 2026-01-09
AI Technical Summary
Existing motor imagery EEG decoding technologies face challenges such as low signal-to-noise ratio, complex spatiotemporal feature coupling, scarce training data, and large inter-subject differences, which limit decoding accuracy and model generalization ability.
A signal enhancement model based on the Transformer network architecture is used to process multi-channel EEG data to generate enhanced EEG signals. The signals are then classified by constructing an adjacency matrix and a graph convolutional network to represent brain function networks. Interpretable discriminative features are extracted by combining graph structure and self-attention mechanism.
It significantly improved the feature representation ability and model generalization ability of EEG signals, enhanced the classification accuracy of motor imagery tasks, strengthened cross-subject adaptability, and improved the scalability and flexibility of EEG paradigms.
Smart Images

Figure CN121300622A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of brain-computer interface technology, specifically to a method and system for decoding motor imagery EEG. Background Technology
[0002] Brain-computer interface (BCI) is a cutting-edge technology that enables human-computer interaction by decoding electroencephalogram (EEG) signals. It achieves communication and control between humans and external devices by directly analyzing the user's brain neural activity (e.g., EEG signals). Motor imagery (MI) is one of its core paradigms. Motor imagery-based BCIs are currently a hot research topic. Their basic principle is that by imagining the movement of specific limbs (e.g., left hand, right hand), users can trigger specific patterns of neural activity in the sensorimotor cortex of the brain. The system decodes these activity patterns to identify the user's motor intentions and control external devices.
[0003] However, EEG signals inherently possess limitations such as weak amplitude (μV level), strong non-stationarity, and low signal-to-noise ratio, and are susceptible to interference from electromyography, electrooculography, and environmental noise, thus restricting the decoding accuracy of motor imagery tasks. Traditional methods (such as CSP+LDA) rely on manual feature extraction, resulting in insufficient generalization ability; while deep learning methods can automatically learn features, they face the bottleneck of scarce training data—EEG data collection is costly, subject-to-subject variability is large, and small samples easily lead to model overfitting. Furthermore, existing data augmentation techniques (such as SMOTE and GAN) struggle to effectively preserve the physiological characteristics of EEG, and the lack of modeling of causal relationships between brain regions further limits the interpretability of features. Therefore, a decoding method that integrates high-quality data augmentation and causal feature mining is urgently needed to improve model robustness and cross-subject adaptability, promoting the practical application of BCI systems. Summary of the Invention
[0004] To address this, the present invention provides a method and system for decoding motor imagery EEG, aiming to solve the technical problem that existing motor imagery EEG decoding technologies face multiple challenges such as low signal-to-noise ratio, complex spatiotemporal feature coupling, scarce training data, and large differences among subjects. These challenges make it difficult to simultaneously and effectively extract interpretable discriminative features and overcome the overfitting problem caused by small sample learning, thus resulting in limited decoding accuracy and model generalization ability.
[0005] To achieve the above objectives, the present invention adopts the following technical solution:
[0006] According to a first aspect of the present invention, the present invention provides a method for decoding motor imagery EEG, the method comprising:
[0007] The raw multi-channel EEG data of the motor imagery task is acquired, and the multi-channel EEG data is processed using a signal enhancement model based on the Transformer network architecture to generate enhanced EEG signals.
[0008] The enhanced EEG signals are spliced together along the time dimension to form a global matrix, and an adjacency matrix representing the brain functional network is constructed based on the effective connectivity between different time series.
[0009] A graph structure is generated based on the adjacency matrix, and the derived features generated based on the enhanced EEG signals are used as node features and input into a graph convolutional network model for classification processing to obtain the classification result of the motor imagery task.
[0010] The signal enhancement model includes a position coding injection structure, an encoder structure consisting of a multi-head self-attention unit and a feedforward network unit, and an enhanced signal output structure connected in sequence.
[0011] Furthermore, before processing the multi-channel EEG data using a signal enhancement model based on a Transformer network architecture, the method further includes:
[0012] The multi-channel EEG data is preprocessed to obtain cleaner multi-channel EEG data;
[0013] The data preprocessing includes filtering out low-frequency drift and high-frequency interference using bandpass filtering; and / or removing physiological artifacts using independent component analysis.
[0014] Furthermore, the process of using a signal enhancement model based on a Transformer network architecture to process the multi-channel EEG data and generate enhanced EEG signals includes:
[0015] The multi-channel EEG data is divided into sequences according to the time step to form an input matrix;
[0016] For the input matrix, the positional coding injection structure is used to generate positional codes through a sine-cosine function. The positional codes and the signal features after linear projection of the multi-channel EEG data are added element by element to obtain the embedded representation.
[0017] The embedded representation is input into the multi-head self-attention unit to calculate the multi-head attention weights, and the multi-head outputs are concatenated and linearly fused to obtain the attention output representation;
[0018] The attention output representation and the embedding representation are added together to form a residual connection, and after layer normalization, they are input into the feedforward network unit. The model representation at each time point is then subjected to a nonlinear transformation to obtain the encoder output.
[0019] The encoder output is reconstructed to obtain an enhanced EEG signal;
[0020] The feedforward network unit includes two fully connected layers and a GELU activation function;
[0021] And / or,
[0022] The step of inputting the embedded representation into the encoder structure to calculate the multi-head attention weights includes:
[0023] The embedding representation is divided into multiple heads, and attention weights are calculated independently for each head, as shown in the following mathematical expression:
[0024]
[0025] Where Q, K, and V represent the query matrix, key matrix, and value matrix, respectively; d k This represents the scaling factor.
[0026] Further, the reconstruction of the encoder output to obtain the enhanced EEG signal includes:
[0027] The encoder output is flattened and reconstructed into an enhanced signal through a Dropout layer and a fully connected layer whose output dimension matches the temporal length of the original signal, thus serving as an enhanced EEG signal.
[0028] During data augmentation, the Adam optimizer is used, and the cross-entropy loss function is employed for model training.
[0029] Furthermore, the construction of an adjacency matrix representing the brain functional network based on the effective connectivity between different time series includes:
[0030] The autoregressive and cross-regressive expressions for different time series are constructed as follows:
[0031]
[0032]
[0033] Among them, X t Y t These represent the autoregressive expressions for different time series; X' t Y' t These represent the regression expressions for different time series; q represents the model parameters; α represents the regression expression for each time series. 1i β 1i α i δ i β i γ iAll are regression coefficients; ξ 1t ξ 2t ξ 3t ξ 4t All are residuals;
[0034] Based on the residuals, the corresponding variances are calculated, and the Granger causality relationships for different time series are obtained. The mathematical expressions are as follows:
[0035] F x→y =ln(T1 / T2)
[0036] F y→x =ln(T2 / T1)
[0037] Among them, F x→y F y→x This represents the causal relationship between different time series X and Y; T1 and T2 represent the variances of the corresponding time series, respectively.
[0038] Based on the Granger causality relationships of different time series, a Granger causality matrix containing all time series is formed, which serves as the adjacency matrix. The mathematical expression is as follows:
[0039]
[0040] Among them, F ij represents the Granger causality coefficients for electrodes i and j; n is the total number of electrodes.
[0041] Further, the generation of the graph structure based on the adjacency matrix includes:
[0042] The adjacency matrix is symmetrically normalized, and the mathematical expression is as follows:
[0043] L = ID -1 / 2 AD -1 / 2
[0044] Where A represents the adjacency matrix; D represents the degree matrix, D ii =∑ j A ij ;
[0045] For the symmetric normalized adjacency matrix, a multi-level pooling matrix is generated using the Graclus multi-level clustering algorithm;
[0046] In the multi-level pooling process, each iteration performs node weight calculation and greedy matching, merges the largest weight neighbor pairs to form supernodes, which serve as nodes in the coarsened graph, and generates the coarsened graph and the corresponding Laplacian matrix.
[0047] Furthermore, the step of using derived features generated based on the enhanced EEG signal as node features includes:
[0048] Complex Morlet wavelet transform is applied to the enhanced EEG signal of each channel to extract time-frequency energy features, as shown in the following mathematical expression:
[0049] CWT(c,τ,s)=∫x c (t)
[0050] Where, x c (t) represents the enhanced EEG signal of channel c; c, τ, and s represent the function parameter, translation parameter, and scale parameter, respectively.
[0051] The time-frequency energy features are concatenated with the enhanced timing signals of the corresponding channels to obtain node features.
[0052] Further, the input is fed into a graph convolutional network model for classification processing to obtain the classification result of the motion imagery task, including:
[0053] The node features and Laplacian matrix are input into the graph convolutional network model, and processed using Chebyshev polynomial approximation spectral convolution through graph convolutional layers. The mathematical expression is as follows:
[0054]
[0055] in, L represents the Laplace matrix; σ represents the ReLU activation; K represents the order; T k Z represents the k-th graph transformation operator; (l) Z represents the input feature matrix of the l-th convolutional layer; (l+1) This represents the output feature matrix of the (l+1)th convolutional layer; This represents the weight matrix of the k-th graph transformation operator;
[0056] After passing through the graph convolutional layer, the output is fed into a pooling layer and a classification layer to output the category probability, thus obtaining the classification result of the motion imagination task.
[0057] According to a second aspect of the present invention, the present invention provides a motor imagery EEG decoding system, the system comprising:
[0058] The data augmentation processing module is used to acquire the original multi-channel EEG data of the motor imagery task, and process the multi-channel EEG data using a signal augmentation model based on the Transformer network architecture to generate enhanced EEG signals.
[0059] The adjacency matrix construction module is used to splice the enhanced EEG signals along the time dimension into a global matrix and construct an adjacency matrix representing the brain functional network based on the effective connectivity between different time series.
[0060] The neural network classification module is used to generate a graph structure based on the adjacency matrix, and to input the derived features generated based on the enhanced EEG signals as node features into the graph convolutional network model for classification processing to obtain the classification result of the motor imagery task.
[0061] The signal enhancement model includes a position coding injection structure, an encoder structure consisting of a multi-head self-attention unit and a feedforward network unit, and an enhanced signal output structure connected in sequence.
[0062] Furthermore, the system also includes:
[0063] The data preprocessing module is used to preprocess the multi-channel EEG data to obtain cleaner multi-channel EEG data.
[0064] The data preprocessing includes filtering out low-frequency drift and high-frequency interference using bandpass filtering; and / or removing physiological artifacts using independent component analysis.
[0065] And / or,
[0066] An external device access module is used to convert the classification results into control signals and access them to external devices to drive the external devices to perform corresponding operations.
[0067] The present invention, by adopting the above technical solution, has at least the following beneficial effects:
[0068] This invention proposes a method for decoding EEG signals related to motor imagery. It acquires raw multi-channel EEG data from a motor imagery task, processes this data using a signal enhancement model based on a Transformer network architecture to generate enhanced EEG signals, and concatenates these signals along the time dimension into a global matrix. Based on the effective connectivity between different time series, an adjacency matrix representing the brain's functional network is constructed. A graph structure is generated based on this adjacency matrix, and the derived features generated from the enhanced EEG signals are used as node features and input into a graph convolutional network model for classification, yielding the classification result for the motor imagery task. This invention significantly improves the feature representation ability of EEG signals, enhances the model's generalization ability, and substantially improves classification accuracy, achieving high scalability and flexibility in the EEG paradigm.
[0069] It should be understood that the above general description and the following detailed description are exemplary and explanatory only, and are not intended to limit the invention. Attached Figure Description
[0070] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0071] Figure 1 A flowchart illustrating a motor imagery EEG decoding method according to an embodiment of the present invention is shown;
[0072] Figure 2 A schematic diagram of the structure of a signal enhancement model provided in an embodiment of the present invention is shown;
[0073] Figure 3 A schematic diagram of the structure of a motor imagery EEG decoding system according to an embodiment of the present invention is shown;
[0074] Figure 4 A schematic diagram of the structure of a motor imagery EEG decoding system provided in another embodiment of the present invention is shown. Detailed Implementation
[0075] Exemplary embodiments of the present disclosure will now be described in more detail with reference to the accompanying drawings. While exemplary embodiments of the present disclosure are shown in the drawings, it should be understood that the present disclosure may be implemented in various forms and should not be limited to the embodiments set forth herein. Rather, these embodiments are provided so that this disclosure will be thorough and complete, and will fully convey the scope of the disclosure to those skilled in the art.
[0076] It should be noted that, in this document, relational terms such as "first" and "second" are used merely to distinguish one entity or operation from another, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Furthermore, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitations, an element defined by the phrase "comprising..." does not exclude the presence of additional identical elements in the process, method, article, or apparatus that includes said element.
[0077] As mentioned earlier, traditional brain-computer interfaces face the following problems: 1. Limited data: High cost of collecting motor imagery data leads to insufficient training samples and the model is prone to overfitting; 2. Difficult feature extraction: The spatiotemporal features of EEG signals are strongly coupled, making it difficult for traditional methods to model the time and space dimensions simultaneously; 3. Poor cross-subject adaptability: The distribution of EEG signals varies greatly among different subjects, resulting in insufficient model generalization ability.
[0078] To address the aforementioned problems, embodiments of the present invention provide a method for decoding motor imagery using electroencephalography, such as... Figure 1 As shown, it may include at least the following steps S101 to S103:
[0079] Step S101: Obtain the original multi-channel EEG data of the motor imagery task, and process the multi-channel EEG data using a signal enhancement model based on the Transformer network architecture to generate enhanced EEG signals.
[0080] Acquiring raw, multi-channel EEG data is typically accomplished using specialized EEG acquisition equipment. For example, an EEG cap with 64 electrodes, conforming to international standards, can be worn on the subject's scalp. During the acquisition process, the subject performs different types of motor imagery tasks based on on-screen prompts, such as imagining movements of the left hand, right hand, both feet, or tongue. The EEG signals are recorded at a specific sampling rate (e.g., 160 Hz or higher), forming multi-channel time-series data, which constitutes the raw, multi-channel EEG data.
[0081] In one alternative embodiment, before processing the multi-channel EEG data using a signal enhancement model based on a Transformer network architecture, the multi-channel EEG data can be preprocessed to obtain cleaner multi-channel EEG data. Specifically, data preprocessing may include filtering out low-frequency drift and high-frequency interference using bandpass filtering techniques, and removing physiological artifacts using independent component analysis techniques.
[0082] Understandably, data preprocessing aims to extract effective features from complex multi-channel EEG data. In practice, the raw signal is first pre-processed using a bandpass filter ranging from 1 to 45 Hz to filter out low-frequency drift (such as baseline noise) and high-frequency interference (such as EMG artifacts and power supply noise), while retaining frequency band information relevant to the motor imagery task (such as mu and beta rhythms). For example, a fourth-order Butterworth bandpass filter can be used to constrain the signal frequency range to between 1 Hz and 45 Hz. The lower limit of 1 Hz effectively filters out low-frequency noise such as baseline drift caused by slow electrode drift and subject movement; while the upper limit of 45 Hz filters out most high-frequency noise, especially 50 Hz power supply interference. Subsequently, Independent Component Analysis (ICA) algorithms can be used to further remove physiological artifacts (such as EEG, ECG, and EMG signals). Understandably, Independent Component Analysis (ICA), as an effective blind source separation technique, can decompose multi-channel mixed signals into multiple statistically independent components. By analyzing the waveform, spectrum, and scalp topography features of these independent components, components related to typical physiological artifacts such as eye movements, heartbeats, and muscle activity can be identified, thereby effectively separating and removing artifacts and obtaining purer EEG signals with higher signal-to-noise ratios that better reflect real brain neural activity.
[0083] The preprocessing procedure in this embodiment of the invention not only significantly improves the signal-to-noise ratio of the signal, but also lays a reliable data foundation for subsequent feature extraction and decoding tasks, ensuring that the model can more accurately capture neural activity features related to motor imagery.
[0084] Furthermore, a signal enhancement model based on the Transformer network architecture is used to process multi-channel EEG data to generate enhanced EEG signals, aiming to address the problem of insufficient training samples caused by difficulties in EEG data acquisition. For example... Figure 2 As shown, the signal enhancement model in this embodiment of the invention includes a position coding injection structure, an encoder structure composed of a multi-head self-attention unit and a feedforward network unit connected in sequence, and an enhanced signal output structure.
[0085] Specifically, the data augmentation process may include the following steps S101-1 to S101-5:
[0086] Step S101-1: Divide the multi-channel EEG data into sequences according to the time step to form an input matrix.
[0087] The preprocessed MI-EEG signal (sampling rate 160Hz, 64 channels) is divided into sequences according to time steps. Each time step contains voltage values from multiple channels, forming an input matrix X∈R.T×C Where T is the number of time steps and C is the number of channels. The input matrix X is standardized and random Gaussian noise (standard deviation of 0.1) is added to enhance robustness.
[0088] Step S101-2: For the input matrix, a positional encoding injection structure is used to generate positional codes through a sine-cosine function. The positional codes and the signal features after linear projection of the multi-channel EEG data are added element by element to obtain the embedded representation.
[0089] In other words, for the input matrix X, the position code P∈R is generated using the sine-cosine function. T×dmodel Where dmodel is the dimension of the hidden layer of the model (e.g., 256). The positional encoding is then combined with the linearly projected signal features XWe (We∈R). C ×dmodel Adding each element together, we obtain the embedded representation E = XWe + P.
[0090] Step S101-3: Input the embedded representation into the multi-head self-attention unit to calculate the multi-head attention weights, and concatenate and linearly fuse the multi-head outputs to obtain the attention output representation.
[0091] The encoder structure in this embodiment of the invention includes a two-layer encoder, each layer comprising a multi-head self-attention unit, a layer normalization unit, and a feedforward network unit.
[0092] Specifically, the embedding representation E can be divided into multiple heads, and attention weights can be calculated independently for each head, as shown in the following mathematical expression:
[0093]
[0094] Where Q, K, and V represent the query matrix, key matrix, and value matrix, respectively; d k This represents the scaling factor.
[0095] The multi-head attention weight outputs are concatenated and then linearly fused before being output to the Cen normalization unit.
[0096] Step S101-4: Add the attention output representation and the embedding representation to form a residual connection, and after performing layer normalization (LayerNorm), input it into the feedforward network unit (FFN) to perform a nonlinear transformation on the model representation at each time point to obtain the encoder output.
[0097] The feedforward network unit consists of two fully connected layers and a GELU activation function. Residual connections and layer normalization are also applied to obtain the encoder output.
[0098] Step S101-5: Reconstruct the encoder output to obtain enhanced EEG signals.
[0099] In this step, the encoder output is flattened and then reconstructed into an enhanced signal using a Dropout layer and a fully connected layer whose output dimension matches the temporal length of the original signal. This enhanced signal serves as the augmented EEG signal. It should be noted that during data augmentation, the Adam optimizer is used, and the model is trained using the cross-entropy loss function.
[0100] In practical applications, after the encoder output undergoes a Flatten operation, it is reconstructed into an enhanced signal X'∈R through a Dropout layer (dropout rate of 0.5) and a fully connected layer (output dimension matches the time length of the original signal). T×C During the training phase, the cross-entropy loss function can be used, and the optimizer is Adam (learning rate 1). e-5 After training for 300 rounds, the encoder weights are retained as a pre-trained feature extractor to generate high-quality augmented data and alleviate the small sample size problem.
[0101] It should be noted that the enhanced EEG signals generated through steps S1 to S5 are adapted for downstream applications. X' retains the original spatiotemporal structure but contains more significant class discrimination features, which can be directly input into downstream models such as Graph Convolutional Networks (GCN) for topological feature extraction.
[0102] Step S102 involves splicing the enhanced EEG signals along the time dimension into a global matrix, and constructing an adjacency matrix representing the brain functional network based on the effective connectivity between different time series.
[0103] This step acquires the enhanced electroencephalogram (MI-EEG) signal X'∈R. N×T×C Where N is the number of trials, C is the number of channels, and T is the number of time points. Then, all experimental data are concatenated along the time dimension to form a global matrix X”∈R. (N×T)×C An adjacency matrix is generated based on the global matrix. Specifically, the adjacency matrix generation process may include the following steps S102-1 to S102-3:
[0104] Step S102-1: Construct autoregressive and cross-regressive expressions for different time series, as follows:
[0105]
[0106]
[0107] Among them, X t Y t These represent the autoregressive expressions for different time series; X' t Y' t These represent the regression expressions for different time series; q represents the model parameters; α represents the regression expression for each time series. 1i β1i α i δ i β i γ i All are regression coefficients; ξ 1t ξ 2t ξ 3t ξ 4t These are all residuals, i.e., the difference between the predicted value and the actual value. The variances of these four sets of residuals can be calculated as T1, T2, T3, and T4. Based on the Granger causality described above, if variance T3 is smaller than variance T1, it means that using information from past moments of time series X and Y to predict the value of X at time t is more accurate than using information from past moments of X alone. This indicates a causal relationship between time series X and Y, with Y being the cause of X. Therefore, the Granger causality relationship between X and Y can be derived from this logic.
[0108] Step S102-2: Calculate the corresponding variance based on the residuals to obtain the Granger causality relationship for different time series. The mathematical expression is as follows:
[0109] F x→y =ln(T1 / T2)
[0110] F y→x =ln(T2 / T1)
[0111] Among them, F x→y F y→x T1 and T2 represent the causal relationship between X and Y in different time series; T1 and T2 represent the variances of the corresponding time series, respectively.
[0112] It's understandable, F x→y This represents the linear directional effect pointing from X to Y, with X as the reference region; F y→x This represents the linear directional effect pointing from Y to X, with Y as the reference region. From this, the Granger causality matrix can be obtained.
[0113] Step S102-3: Based on the Granger causality relationships of different time series, form a Granger causality matrix containing all time series, which serves as the adjacency matrix. The mathematical expression is as follows:
[0114]
[0115] Among them, F ij Let represent the Granger causality coefficients for electrodes i and j; n is the total number of electrodes. Thus, by explicitly modeling causal relationships between brain regions using Granger causal networks, the interpretability of features is greatly improved.
[0116] Step S103: Generate a graph structure based on the adjacency matrix, and use the derived features generated based on the enhanced EEG signal as node features, input them into the graph convolutional network model for classification processing, and obtain the classification results of the motor imagery task.
[0117] In this embodiment of the invention, a graph structure is generated based on the adjacency matrix, namely, a coarsened graph and its corresponding Laplacian matrix. Specifically, the adjacency matrix can be symmetrically normalized, as shown in the following mathematical expression:
[0118] L = ID -1 / 2 AD -1 / 2
[0119] Where A represents the adjacency matrix; D represents the degree matrix, D ii =∑ j A ij .
[0120] For the symmetrically normalized adjacency matrix, a multi-level pooling matrix is generated using the Graclus multi-level clustering algorithm. In practical applications, a 5-level pooling matrix can be generated. During the multi-level pooling process, node weights are calculated at each iteration: W. ij =1 / (deg(i)+deg(j)); then perform greedy matching, merge the largest weight neighbor pairs to form supernodes, which serve as nodes in the coarsened graph, generating the coarsened graph and the corresponding Laplacian matrix L. (k) (k = 1, ..., 5).
[0121] Furthermore, derived features generated based on enhanced EEG signals are used as node features. Specifically, complex Morlet wavelet transform (center frequency 1Hz, bandwidth 3Hz) can be applied to the enhanced EEG signals of each channel to extract time-frequency energy features. The mathematical expression is as follows:
[0122] CWT(c,τ,s)=∫x c (t)
[0123] Where, x c (t) represents the enhanced EEG signal of channel c; c, τ, and s represent the function parameter, translation parameter, and scaling parameter (8-30Hz frequency band), respectively. The modulus |CWT|∈R is extracted from this. T×F (F is the number of frequency points), and then the time-frequency energy features are concatenated with the enhanced time-series signals of the corresponding channels to obtain the node features Vc∈R. T×(1+F) .
[0124] Furthermore, the node features V and the Laplacian matrix L are input into the graph convolutional network model, and processed through six graph convolutional layers. Each layer uses Chebyshev polynomial approximation spectral convolution, as shown in the following mathematical expression:
[0125]
[0126] in, For Chebyshev polynomials; L represents the Laplace matrix; σ represents ReLU activation; K represents the order, preferably 2; T k Z represents the k-th graph transformation operator; (l) Z represents the input feature matrix of the l-th convolutional layer; (l+1) This represents the output feature matrix of the (l+1)th convolutional layer; Let represent the weight matrix of the k-th graph transformation operator.
[0127] After the graph convolutional layer outputs, a pooling layer and a classification layer are applied to output the class probabilities, yielding the classification result for the motion visualization task. Specifically, each graph convolutional layer is followed by average pooling (pooling size 2), using L... (k) Hierarchical dimensionality reduction is performed. After global average pooling, two fully connected layers (512-dimensional hidden layer and 4-dimensional output layer) are applied, and SoftMax is used to output class probabilities, thus achieving classification of motion visualization tasks. This step, combined with a graph convolutional neural network, fully utilizes topological information and significantly improves classification accuracy.
[0128] This invention provides a method for decoding motor imagery in electroencephalogram (EEG) signals. It employs a hybrid model architecture combining Graph Convolutional Networks (GCNs) and Transformers for decoding motor imagery (MI) signals from EEG. This method significantly improves classification accuracy, enhances model generalization ability, and improves the feature representation capabilities of EEG signals. Specific beneficial effects include at least the following:
[0129] 1) Effectively modeling functional connectivity between brain regions and enhancing feature representation: GCN (Graph Convolutional Network) can model spatial dependencies between brain regions based on the topological structure of EEG channels (such as electrode locations or functional connectivity), capturing co-activation patterns between different brain regions. The Transformer's self-attention mechanism further optimizes spatiotemporal features, enhances the weights of key EEG components (such as mu / beta rhythms), and suppresses noise interference. This invention combines the advantages of both: GCN provides prior information on brain network structure, and Transformer adaptively learns dynamic spatiotemporal dependencies, enabling the model to extract neural representations related to motor imagery more accurately.
[0130] 2) Enhancing cross-subject generalization ability and alleviating the problem of overfitting in small samples: Traditional EEG decoding methods (such as CSP+LDA) rely on handcrafted features and have poor generalization ability; deep learning methods (such as CNN and LSTM) are easily affected by individual differences among subjects. The GCN-Transformer hybrid architecture proposed in this invention reduces the dependence on single-subject data through graph structure constraints (such as adjacency matrices based on anatomical or functional connectivity), thereby improving cross-subject adaptability. This invention improves data efficiency; the Transformer's attention mechanism can automatically focus on key time segments, reducing the need for large-scale training data and making it suitable for small-sample EEG scenarios.
[0131] 3) Supports interpretability analysis and assists in brain-computer interface optimization: The graph structure of GCN can reflect the causal or correlational nature of brain functional networks, such as the activation patterns of the sensorimotor cortex in motor imagery tasks. The attention weights of Transformer can visualize key time points and brain regions, helping researchers analyze the contribution of different frequency bands (such as 8-13Hz mu rhythm) to classification. This feature helps optimize the electrode layout and frequency band selection of the BCI system, improving the system's practicality.
[0132] 4) Applicable to multiple EEG paradigms and highly scalable: This invention is not only applicable to motor imagery (MI) tasks, but can also be extended to event-related potentials (ERPs), steady-state visual evoked potentials (SSVEPs), and other paradigms. Specifically, by adjusting the adjacency matrix of the GCN (e.g., based on functional connectivity or Granger causality), it can be adapted to different EEG experimental designs, exhibiting high flexibility and scalability.
[0133] Furthermore, as Figure 1 In specific implementation, embodiments of the present invention provide a motor imagery EEG decoding system, such as Figure 3 As shown, the device may include: a data augmentation processing module 310, an adjacency matrix construction module 320, and a neural network classification module 330.
[0134] The data augmentation processing module 310 can be used to acquire the original multi-channel EEG data of performing a motor imagery task, and process the multi-channel EEG data using a signal augmentation model based on the Transformer network architecture to generate enhanced EEG signals.
[0135] The adjacency matrix construction module 320 can be used to splice enhanced EEG signals into a global matrix along the time dimension and construct an adjacency matrix representing the brain functional network based on the effective connectivity between different time series.
[0136] The neural network classification module 330 can be used to generate a graph structure based on the adjacency matrix, and use the derived features generated based on the enhanced EEG signal as node features, inputting them into the graph convolutional network model for classification processing to obtain the classification results of the motor imagery task;
[0137] The signal enhancement model includes a position coding injection structure, an encoder structure consisting of a multi-head self-attention unit and a feedforward network unit, and an enhanced signal output structure connected in sequence.
[0138] Optionally, such as Figure 4 As shown, another embodiment of the present invention provides a motor imagery EEG decoding system, which further includes a data preprocessing module 340 and an external device access module 350.
[0139] The data preprocessing module 340 can be used to preprocess multi-channel EEG data to obtain cleaner multi-channel EEG data.
[0140] Data preprocessing includes filtering out low-frequency drift and high-frequency interference using bandpass filtering techniques; and / or removing physiological artifacts using independent component analysis techniques.
[0141] The external device access module 350 can be used to convert the classification results into control signals to access external devices, so as to drive the external devices to perform corresponding operations.
[0142] It should be noted that other corresponding descriptions of the functional modules involved in the motor imagery EEG decoding system provided in this embodiment of the invention can be found in the following references. Figure 1 The corresponding description of the method shown will not be repeated here.
[0143] Those skilled in the art will clearly understand that the specific working process of the systems, devices, modules and units described above can be referred to the corresponding process in the foregoing method embodiments. For the sake of brevity, it will not be repeated here.
[0144] Furthermore, the functional units in the various embodiments of the present invention can be physically independent of each other, or two or more functional units can be integrated together, or all functional units can be integrated into one processing unit. The integrated functional units described above can be implemented in hardware, or in software or firmware.
[0145] Those skilled in the art will understand that if the integrated functional unit is implemented in software and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, or all or part of it, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computing device (e.g., a personal computer, server, or network device) to execute all or part of the steps of the methods described in the embodiments of the present invention when running the instructions. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.
[0146] Alternatively, all or part of the steps of the foregoing method embodiments can be implemented by hardware (such as a computing device, personal computer, server, or network device) related to program instructions. The program instructions can be stored in a computer-readable storage medium. When the program instructions are executed by the processor of the computing device, the computing device executes all or part of the steps of the methods described in the various embodiments of the present invention.
[0147] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, and not to limit them. Although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that within the spirit and principles of the present invention, modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some or all of the technical features therein; and these modifications or substitutions do not cause the corresponding technical solutions to depart from the protection scope of the present invention.
Claims
1. A method for decoding motor imagery using electroencephalography, characterized in that, The method includes: The raw multi-channel EEG data of the motor imagery task is acquired, and the multi-channel EEG data is processed using a signal enhancement model based on the Transformer network architecture to generate enhanced EEG signals. The enhanced EEG signals are spliced together along the time dimension to form a global matrix, and an adjacency matrix representing the brain functional network is constructed based on the effective connectivity between different time series. A graph structure is generated based on the adjacency matrix, and the derived features generated based on the enhanced EEG signals are used as node features and input into a graph convolutional network model for classification processing to obtain the classification result of the motor imagery task. The signal enhancement model includes a position coding injection structure, an encoder structure consisting of a multi-head self-attention unit and a feedforward network unit, and an enhanced signal output structure connected in sequence.
2. The method according to claim 1, characterized in that, Before processing the multi-channel EEG data using a signal enhancement model based on a Transformer network architecture, the method further includes: The multi-channel EEG data is preprocessed to obtain cleaner multi-channel EEG data; The data preprocessing includes filtering out low-frequency drift and high-frequency interference using bandpass filtering; and / or removing physiological artifacts using independent component analysis.
3. The method according to claim 1, characterized in that, The process of using a signal enhancement model based on a Transformer network architecture to process the multi-channel EEG data and generate enhanced EEG signals includes: The multi-channel EEG data is divided into sequences according to the time step to form an input matrix; For the input matrix, the positional coding injection structure is used to generate positional codes through a sine-cosine function. The positional codes and the signal features after linear projection of the multi-channel EEG data are added element by element to obtain the embedded representation. The embedded representation is input into the multi-head self-attention unit to calculate the multi-head attention weights, and the multi-head outputs are concatenated and linearly fused to obtain the attention output representation; The attention output representation and the embedding representation are added together to form a residual connection, and after layer normalization, they are input into the feedforward network unit. The model representation at each time point is then subjected to a nonlinear transformation to obtain the encoder output. The encoder output is reconstructed to obtain an enhanced EEG signal; The feedforward network unit includes two fully connected layers and a GELU activation function; And / or, The step of inputting the embedded representation into the encoder structure to calculate the multi-head attention weights includes: The embedding representation is divided into multiple heads, and attention weights are calculated independently for each head, as shown in the following mathematical expression: Where Q, K, and V represent the query matrix, key matrix, and value matrix, respectively; d k This represents the scaling factor.
4. The method according to claim 3, characterized in that, The reconstruction of the encoder output to obtain enhanced EEG signals includes: The encoder output is flattened and reconstructed into an enhanced signal through a Dropout layer and a fully connected layer whose output dimension matches the temporal length of the original signal, thus serving as an enhanced EEG signal. During data augmentation, the Adam optimizer is used, and the cross-entropy loss function is employed for model training.
5. The method according to claim 1, characterized in that, The process of constructing an adjacency matrix representing the brain functional network based on the effective connectivity between different time series includes: The autoregressive and cross-regressive expressions for different time series are constructed as follows: Among them, X t Y t These represent the autoregressive expressions for different time series; X' t Y' t These represent the regression expressions for different time series; q represents the model parameters; α represents the regression expression for each time series. 1i β 1i α i δ i β i γ i All are regression coefficients; ξ 1t ξ 2t ξ 3t ξ 4t All are residuals; Based on the residuals, the corresponding variances are calculated, and the Granger causality relationships for different time series are obtained. The mathematical expressions are as follows: F x→y =ln(T1 / T2) F y→x =ln(T2 / T1) Among them, F x→y F y→x This represents the causal relationship between different time series X and Y; T1 and T2 represent the variances of the corresponding time series, respectively. Based on the Granger causality relationships of different time series, a Granger causality matrix containing all time series is formed, which serves as the adjacency matrix. The mathematical expression is as follows: Among them, F ij represents the Granger causality coefficients for electrodes i and j; n is the total number of electrodes.
6. The method according to claim 1, characterized in that, The graph structure generated based on the adjacency matrix includes: The adjacency matrix is symmetrically normalized, and the mathematical expression is as follows: L=I-D -1 / 2 AD -1 / 2 Where A represents the adjacency matrix; D represents the degree matrix, D ii =∑ j A ij ; For the symmetric normalized adjacency matrix, a multi-level pooling matrix is generated using the Graclus multi-level clustering algorithm; In the multi-level pooling process, each iteration performs node weight calculation and greedy matching, merges the largest weight neighbor pairs to form supernodes, which serve as nodes in the coarsened graph, and generates the coarsened graph and the corresponding Laplacian matrix.
7. The method according to claim 1, characterized in that, The step of using derived features generated based on the enhanced EEG signal as node features includes: Complex Morlet wavelet transform is applied to the enhanced EEG signal of each channel to extract time-frequency energy features, as shown in the following mathematical expression: CWT(c,τ,s)=∫x c (t) Where, x c (t) represents the enhanced EEG signal of channel c; c, τ, and s represent the function parameter, translation parameter, and scale parameter, respectively. The time-frequency energy features are concatenated with the enhanced timing signals of the corresponding channels to obtain node features.
8. The method according to claim 1, characterized in that, The input is fed into a graph convolutional network model for classification processing to obtain the classification result of the motion imagination task, including: The node features and Laplacian matrix are input into the graph convolutional network model, and processed using Chebyshev polynomial approximation spectral convolution through graph convolutional layers. The mathematical expression is as follows: in, L represents the Laplace matrix; σ represents the ReLU activation; K represents the order; T k Z represents the k-th graph transformation operator; (l) Z represents the input feature matrix of the l-th convolutional layer; (l+1) This represents the output feature matrix of the (l+1)th convolutional layer; This represents the weight matrix of the k-th graph transformation operator; After passing through the graph convolutional layer, the output is fed into a pooling layer and a classification layer to output the category probability, thus obtaining the classification result of the motion imagination task.
9. A motor imagery EEG decoding system, characterized in that, The system includes: The data augmentation processing module is used to acquire the original multi-channel EEG data of the motor imagery task, and process the multi-channel EEG data using a signal augmentation model based on the Transformer network architecture to generate enhanced EEG signals. The adjacency matrix construction module is used to splice the enhanced EEG signals along the time dimension into a global matrix and construct an adjacency matrix representing the brain functional network based on the effective connectivity between different time series. The neural network classification module is used to generate a graph structure based on the adjacency matrix, and to input the derived features generated based on the enhanced EEG signals as node features into the graph convolutional network model for classification processing to obtain the classification result of the motor imagery task. The signal enhancement model includes a position coding injection structure, an encoder structure consisting of a multi-head self-attention unit and a feedforward network unit, and an enhanced signal output structure connected in sequence.
10. The system according to claim 9, characterized in that, The system also includes: The data preprocessing module is used to preprocess the multi-channel EEG data to obtain cleaner multi-channel EEG data. The data preprocessing includes filtering out low-frequency drift and high-frequency interference using bandpass filtering technology; And / or, remove physiological artifacts using independent component analysis techniques; And / or, An external device access module is used to convert the classification results into control signals and access them to external devices to drive the external devices to perform corresponding operations.