Motor imagery electroencephalogram signal decoding method and system

By adaptively adjusting EEG signal features through the HFNet architecture, extracting multi-channel and multi-scale spatiotemporal features and performing deep time series modeling, problems such as low signal-to-noise ratio and large individual differences in motor imagery EEG signal decoding are solved, and decoding accuracy and robustness are improved.

CN120763703APending Publication Date: 2025-10-10NINGXIA UNIVERSITY
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510923963.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-04
Publication Date
2025-10-10

AI Technical Summary

Technical Problem

Existing technologies face problems in decoding motor imagery EEG signals, such as extremely low signal-to-noise ratio, large individual differences, severe noise pollution, high non-stationarity, and weak ERD/ERS signals, resulting in insufficient robustness and accuracy of the decoding model.

Method used

The coordinated fusion network (HFNet) architecture is adopted. Through the combination of the time series-statistical feature coordinator (CSFH), the heterogeneous fusion-efficient channel attention module (HF-ECA) and the time series weaving module (CW), the EEG signal features are adaptively adjusted, multi-channel and multi-scale spatiotemporal features are extracted, and the multi-head attention mechanism and the residual separable time series convolutional network are used for deep time series modeling.

Benefits of technology

It improves the decoding accuracy and robustness of motor imagery EEG signals, enhances the model's adaptability to individual differences and non-stationarity, effectively integrates multi-dimensional features, captures long-range dependencies, and improves classification accuracy and generalization ability.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120763703A_ABST
    Figure CN120763703A_ABST
Patent Text Reader

Abstract

The invention discloses a motor imagery electroencephalogram signal decoding method and system, and the method comprises the steps: obtaining the time sequence data of each channel of a motor imagery electroencephalogram signal, carrying out the averaging of the time sequence data along the time dimension, obtaining the features after the time average pooling, carrying out the standardization of the features in the channels, and obtaining the time sequence data of each channel of the motor imagery electroencephalogram signal. Gaussian weighting coefficients are generated through Gaussian weighting harmonic, the coefficients are broadcasted along the time dimension, weights matched with the time dimension are obtained, the weights and the motor imagery electroencephalogram signals are multiplied element by element, and the dynamically harmonic motor imagery electroencephalogram signals are obtained. Extracting multi-channel multi-scale spatial-temporal features of the signal, windowing a feature sequence, constructing a feature vector with a fixed length, and performing classification through a classifier to obtain a probability corresponding to each predefined motor imagery category; according to the method, the decoding performance can be effectively improved through cooperative work of adaptive coordination, feature heterogeneous fusion and channel attention optimization of the motor imagery electroencephalogram signals.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the fields of brain-computer interface and artificial intelligence technology, and more specifically to a method and system for decoding motor imagery EEG signals. Background Art

[0002] Brain-computer interface uses AI algorithms and wearable sensors to monitor and decode brain activity. It is a cutting-edge technology with great potential and is expected to revolutionize the fields of communication and control. It has been widely used in various fields including human-computer interaction, sports rehabilitation and disease treatment.

[0003] Electroencephalography (EEG), a technique for recording brain activity, has become a primary method for capturing the neural signatures of motor imagery due to its non-invasive, high temporal resolution, and cost-effectiveness. However, practical applications present three key challenges: First, the raw EEG signal has an extremely low signal-to-noise ratio and is susceptible to noise contamination, such as electromyographic artifacts and environmental interference; second, the signal exhibits significant non-stationarity and inter-individual variability, and EEG patterns may drift within the same subject over time; finally, the ERD / ERS phenomenon induced by motor imagery exhibits weak signal characteristics, placing stringent demands on feature extraction algorithms.

[0004] Although existing technologies have made some progress in decoding and classifying motor imagery electroencephalogram (EEG) signals, there are still several inherent challenges and methodological limitations. These problems limit the further improvement of the performance and widespread application of brain-computer interface systems.

[0005] The inherent characteristics of EEG signals limit decoding effectiveness: EEG signals have an extremely low signal-to-noise ratio and are susceptible to contamination by various noise and physiological artifacts. Their significant non-stationarity and high inter-individual variability make it extremely difficult to construct robust and generalizable decoding models. Furthermore, the neural signatures associated with motor imagery tasks (such as the ERD / ERS phenomenon) are inherently weak, making effective identification and extraction more challenging.

[0006] Traditional machine learning methods have significant limitations. These methods (e.g., those based on common spatial patterns (CSP) and subsequent classifiers like LDA and SVM) often rely on handcrafted features and struggle to fully capture the complex nonlinear dynamics and deep, abstract associations in EEG signals. These methods also have limited adaptability in handling the nonstationary nature of signals and individual differences, often limiting their classification accuracy and generalization across subjects.

[0007] Existing deep learning models still have common technical bottlenecks: Although deep learning has shown advantages in automatic feature extraction, existing models still have shortcomings in the following aspects: First, in terms of individualized adaptive calibration, that is, how to effectively deal with the distribution differences of EEG signals between different users and improve the personalized adaptability of the model; second, in the collaborative fusion and deep integration of multi-dimensional (time domain, spatial domain and frequency domain) and multi-scale features, the complementary advantages of various features have not been fully utilized; third, there is still much room for optimization in the fine modeling and dynamic calibration of complex long-range temporal dependencies in EEG signals; finally, the robustness of some models under small sample training conditions and the structural complexity and computational cost of high-performance models themselves are also important factors limiting their widespread application.

[0008] In summary, the above-mentioned existing technologies have extremely low signal-to-noise ratio and high variability of motor imagery EEG signals, are easily contaminated by various noises and physiological artifacts, have significant non-stationarity and poor individual adaptability, and are insufficient in feature extraction and fusion during time series modeling, resulting in poor robustness and decoding accuracy of the constructed decoding model. Summary of the Invention

[0009] In response to the problems existing in the above-mentioned fields, the present invention proposes a method and system for decoding motor imagery EEG signals. By adaptively coordinating motor imagery EEG signals, individual differences are overcome, multi-channel and multi-scale spatiotemporal features are extracted, and cross-scale spatiotemporal features are integrated. Through the fine weaving of deep temporal patterns and fine modeling of deep temporal dependencies in collaborative work, the input motor imagery EEG signals are gradually converted into highly discriminative feature vector representations, and finally the classifier outputs accurate classification results, thereby effectively improving the decoding performance.

[0010] To solve the above technical problems, the present invention discloses a method for decoding motor imagery EEG signals, comprising the following steps: Obtain the time series data of each channel of the motor imagery EEG signal; By averaging the time series data along the time dimension, we obtain the time-averaged pooled features, normalize the features within the channel, and generate Gaussian weighted coefficients through Gaussian weighted harmonization. The Gaussian weighted coefficients are broadcasted along the time dimension to obtain weights that match the time dimension of the motor imagery EEG signal. The weights are then multiplied element-by-element by the motor imagery EEG signal to obtain a dynamically harmonized motor imagery EEG signal. The multi-channel and multi-scale spatiotemporal features of the dynamically coordinated motor imagery EEG signals are extracted; the feature sequence of the multi-scale spatiotemporal features is windowed to construct a feature vector of fixed length, which is then classified by a classifier to obtain the probability corresponding to each predefined motor imagery category.

[0011] Preferably, the extraction of multi-channel and multi-scale spatiotemporal features of dynamically coordinated motor imagery EEG signals includes: The dynamically harmonized motor imagery EEG signals are spatially filtered channel by channel to capture the spatial distribution characteristics within each channel and perform dimensionality adjustment and feature fusion in sequence to obtain fused multi-channel and multi-scale spatiotemporal features. A channel attention mechanism is introduced to determine the adaptive weight according to the number of channels of the fused multi-channel and multi-scale spatiotemporal features, and the fused multi-channel and multi-scale spatiotemporal features are adaptively recalibrated by generating a channel attention weight vector.

[0012] Preferably, constructing a fixed-length feature vector includes: The feature sequence of adaptively recalibrated multi-channel and multi-scale spatiotemporal features is windowed, and the feature sequence of each sliding window is optimized through a multi-head attention mechanism; The optimized feature sequence of each sliding window is aggregated to form a feature vector of fixed length.

[0013] Preferably, the adaptive recalibration of the fused multi-channel multi-scale spatiotemporal features by generating a channel attention weight vector specifically includes: A channel attention mechanism is introduced to optimize the fused multi-channel and multi-scale spatiotemporal features. The spatial dimension information of each channel is compressed into a single scalar descriptor through a global average pooling operation. Local cross-channel interactions are modeled in the channel dimension through a one-dimensional convolution. The kernel size of this one-dimensional convolution is not fixed, but is adaptively determined by a function based on the number of channels after fusion. After performing one-dimensional convolution processing on the fused multi-channel and multi-scale spatiotemporal features, a channel attention weight vector ranging from 0 to 1 is generated through a Sigmoid activation function; the channel attention weight vector is multiplied by the spatiotemporal features of each channel to adaptively recalibrate the features of each channel.

[0014] Preferably, the optimization of the feature sequence of each sliding window by the multi-head attention mechanism specifically includes: The feature sequence of each sliding window is mapped to different representation subspaces through independent linear projection layers to obtain Q, K and V of the projected feature sequence; The projected feature sequences Q, K, and V are input into multiple parallel attention heads. Within each attention head, scaled dot product attention is performed. The dot product of the transposed Q and K is calculated, the dot product result is scaled, and the attention score is normalized into a probability distribution through the SoftMax function. The probability distribution is weighted and summed with V to obtain the output result of each attention head. The output results of all parallel attention heads are concatenated and integrated again through a final linear projection layer to obtain the optimized feature sequence of each sliding window.

[0015] Preferably, the probabilities corresponding to the various predefined motor imagery categories are obtained by inputting a feature vector of a fixed length into a classifier consisting of one or more fully connected layers and a Softmax activation function, and outputting the probabilities corresponding to the various predefined motor imagery categories through classification by the classifier.

[0016] Preferably, the fused multi-channel multi-scale spatiotemporal features are obtained by using a three-way parallel heterogeneous convolution architecture to transform the dynamic harmonic motor imagery EEG signal feature map at the beginning of each branch through a convolution kernel size of The depth convolution performs channel-by-channel spatial filtering, C is the number of input channels, corresponding to the EEG electrode dimension, capturing the spatial distribution characteristics within each channel; The spatial distribution features within each channel are subjected to standard intermediate processing, including batch normalization, nonlinear activation, pooling and dropout operations; one branch uses standard two-dimensional convolution acting on the time dimension to perform dense temporal pattern learning; the other two branches use separable convolution to decompose the standard convolution into depth-wise channel-by-channel temporal convolution and subsequent point convolution; The features extracted by the three parallel branches are dimensionally adjusted and then the feature maps are concatenated or element-by-element added to obtain fused multi-channel and multi-scale spatiotemporal features.

[0017] Preferably, a motor imagery EEG signal decoding system is also included, comprising: A signal acquisition module is used to obtain the time series data of each channel of the motor imagery EEG signal; A dynamic harmonization module is used to obtain time-averaged pooled features by averaging the time series data along the time dimension, normalize the features within the channel, and generate Gaussian weighted coefficients through Gaussian weighted harmonization; broadcast the Gaussian weighted coefficients along the time dimension to obtain weights that match the time dimension of the motor imagery EEG signal, and multiply the weights by the motor imagery EEG signal element by element to obtain a dynamically harmonized motor imagery EEG signal; The decoding module is used to extract the multi-channel and multi-scale spatiotemporal features of the dynamically coordinated motor imagery EEG signals; the feature sequence of the multi-scale spatiotemporal features is windowed to construct a feature vector of fixed length, and the feature vector is classified by a classifier to obtain the probability corresponding to each predefined motor imagery category.

[0018] Compared with the prior art, the present invention has the following beneficial effects: The motor imagery EEG signal decoding method proposed in the present invention adaptively adjusts and optimizes the time series data distribution and statistical characteristics of motor imagery EEG signals, thereby enhancing the adaptability and robustness of the model to signal changes in different individuals and different time periods, and reducing the negative impact of individual differences. The multi-channel and multi-scale spatiotemporal features of the dynamically coordinated motor imagery EEG signals are extracted, and heterogeneous feature information in multiple dimensions such as spatial domain, spectral domain (or time-frequency domain) and time domain is captured and effectively integrated from the EEG signals in parallel and at multiple scales. At the same time, an advanced attention mechanism is introduced to dynamically evaluate and strengthen the key features and channels that contribute more to the classification task, thereby achieving synergistic enhancement of features and effective compression of information. The input motor imagery EEG signals are gradually converted into a highly discriminative feature vector representation, and the model's ability to understand the signal evolution process is improved by capturing the long-range dependencies, multi-level temporal dynamics and subtle temporal patterns contained in the EEG signals. BRIEF DESCRIPTION OF THE DRAWINGS

[0019] Figure 1 This is a flow chart of the motor imagery EEG signal decoding method proposed by the present invention; Figure 2 The Harmonized Fusion Network (HFNet) architecture constructed for the embodiments of the present invention; Figure 3 A schematic diagram of the MSA module structure provided in an embodiment of the present invention. DETAILED DESCRIPTION

[0020] The following is a combination of the embodiments of the present invention Figure 1-Figure 3 , the technical solutions in the embodiments of the present invention are clearly and completely described. It should be understood that the terms used in the present invention are only used to describe specific implementation methods and are not intended to limit the present invention.

[0021] Example The embodiment of the present invention provides a coordinated fusion network (HFNet) and a motor imagery electroencephalogram (EEG) decoding method and system based on the network, which is used to overcome the core problems faced by the existing technology in processing EEG signals, such as low signal-to-noise ratio, high variability, insufficient feature extraction and fusion, thereby improving decoding accuracy and robustness. Figure 1 As shown, the method includes the following steps:

[0022] Obtain the time series data of each channel of motor imagery EEG signal; By averaging the time series data along the time dimension, we obtain the time-averaged pooled features, normalize the features within the channel, and generate Gaussian weighted coefficients through Gaussian weighted harmonization. The Gaussian weighted coefficients are broadcasted along the time dimension to obtain weights that match the time dimension of the motor imagery EEG signal. The weights are then multiplied element-by-element by the motor imagery EEG signal to obtain a dynamically harmonized motor imagery EEG signal. The multi-channel and multi-scale spatiotemporal features of the dynamically coordinated motor imagery EEG signals are extracted; the feature sequence of the multi-scale spatiotemporal features is windowed to construct a feature vector of fixed length, which is then classified by a classifier to obtain the probability corresponding to each predefined motor imagery category.

[0023] The core technology of the proposed HFNet and the motor imagery EEG decoding method based on it lies in its unique network architecture. Figure 2 As shown in FIG, the overall architecture of the HFNet of the present invention is shown.

[0024] The architecture connects three key modules in series: Chrono-Statistical Feature Harmonizer (CSFH), Heterogeneous Fusion-Efficient Channel Attention (HF-ECA), and Chronological Weaving (CW).

[0025] The acquired motor imagery EEG signal (MI-EEG signal) is used as the starting point of HFNet. It first enters the time series-statistical feature coordinator (CSFH) module. The time series data of each channel of the motor imagery EEG signal are averaged along the time dimension to obtain the time average pooled features. The features are then normalized within the channel and Gaussian weighted coefficients are generated through Gaussian weighting.

[0026] This module preprocesses and calibrates the input signal through a series of adaptive statistical adjustments and preliminary convolution operations, aiming to reduce noise and enhance signal stability and consistency.

[0027] The features processed by the CSFH module are sent to the heterogeneous fusion-efficient channel attention (HF-ECA) module, which broadcasts the Gaussian weighted coefficients along the time dimension to obtain weights that match the time dimension of the motor imagery EEG signal. The weights are then multiplied element-by-element with the motor imagery EEG signal to obtain a dynamically harmonized motor imagery EEG signal.

[0028] The HF-ECA module extracts rich multi-scale spatiotemporal features through its parallel multi-branch heterogeneous convolutional structure and uses an efficient channel attention mechanism to perform weighted optimization on these features to highlight channels with richer information.

[0029] The dynamically harmonized motor imagery EEG signals are spatially filtered channel by channel to capture the spatial distribution characteristics within each channel and perform dimensionality adjustment and feature fusion in sequence to obtain fused multi-channel and multi-scale spatiotemporal features. A channel attention mechanism is introduced to determine the adaptive weights based on the number of channels of the fused multi-channel, multi-scale spatiotemporal features. The fused multi-channel, multi-scale spatiotemporal features are adaptively recalibrated by generating a channel attention weight vector. The feature sequence of the adaptively recalibrated multi-channel, multi-scale spatiotemporal features is windowed, and the feature sequence of each sliding window is optimized using a multi-head attention mechanism.

[0030] The optimized feature maps are fed into the Temporal Weaving (CW) module, whose core is the Multi-Head Residual Attention Separable Temporal Convolutional Network (MHA-RSTCN), specifically designed to capture complex deep temporal dependencies and long-range patterns in feature sequences. The optimized feature sequences of each sliding window are aggregated to form a fixed-length feature vector.

[0031] Finally, the refined feature representation output by the CW module will pass through one or more fully connected layers (classifiers, not detailed in the figure but standard configuration) and the Softmax activation function. The classifier will classify the feature vector to obtain the probability corresponding to each predefined motor imagery category, and complete the probability prediction and classification judgment of different motor imagery task categories.

[0032] In terms of input, the motor imagery EEG signal processed by the method provided by the embodiment of the present invention is usually represented as a two-dimensional tensor ,in C The number of channels representing EEG signals (e.g. 22 channels), T Represents the number of time sampling points of each channel in the time window (for example, determined by the sampling rate and window duration, such as 1000 points for a 4-second window at a 250Hz sampling rate). Before inputting into HFNet, the raw EEG data is usually preprocessed, such as bandpass filtering to retain the frequency band of interest (such as 0.5-100Hz), and intercepting key time segments containing specific motor imagery tasks (such as left hand, right hand, feet, tongue imagination) from the continuous recording according to task prompts (such as visual or auditory instructions). Each time segment X i Corresponding to a category label y i, which is usually converted to a one-hot encoding during training. This paper adopts an end-to-end learning model, aiming to automatically learn and extract discriminative features directly from (pre-processed) multi-channel EEG time series signals, without the need for complex traditional manual feature engineering.

[0033] like Figure 2 The CSFH module shown in the upper left area is the first key component of the network. Its core function is to perform adaptive feature recalibration and preliminary spatiotemporal feature extraction on the input EEG signal to cope with the inherent noise, variability and non-stationarity of EEG signals. The specific technical implementation steps are as follows: The module input is a batch of EEG signal tensors, such as (in B is the batch size).

[0034] First, perform time average pooling: the time series data of each channel is averaged along the time dimension T This operation aims to aggregate the persistent and relatively stable neural activation patterns within each trial, thereby obtaining a feature representation that can represent the main trend of the trial. y , whose time dimension is compressed.

[0035] Secondly, channel normalization is performed: in order to enhance the robustness of the signal to amplitude fluctuations and unify the statistical scales between different channels, the features after time average pooling are normalized. y Perform normalization within the channel. Specifically, for each feature dimension (if there are multidimensional features) and each sample in the batch, calculate its mean across all channels and variance , and then apply the normalization formula to get the normalized features :

[0036] ; in, y bfc A feature of a specific sample in the channel c The value on . and are the mean and variance of the feature across all channels, respectively, which capture the global statistical distribution of the feature at that moment. The statistical deviation of each channel from the mean of the channel group is quantified. is a tiny constant set to prevent division by zero.

[0037] Implementing Gaussian weighted harmonization: using standardized features To generate a Gaussian weighting coefficient The weight The calculation formula is:

[0038] ; Among them, the weight The range of is (0, 1], which evaluates the reliability of the channel based on its statistical deviation. When the channel characteristics are highly consistent with the mean ( ), its weight When the value approaches 1, it indicates that the signal of the channel is considered stable and reliable. On the contrary, when the channel shows significant deviation, its weight will be attenuated, thereby suppressing potential abnormal signals. c It is a hyperparameter used to adjust the sensitivity to channel statistical deviations. Under the action of the Gaussian function, it gives lower weights to channels whose statistical characteristics deviate from the normal state and higher weights to channels close to the mean, so as to balance the suppression of abnormal signals and the smooth modulation of features. (whose time dimension is 1) will be broadcast along the time dimension so that its dimension is the same as the original input signal without time averaging. X Then, the original input signal X With the weight after broadcasting Perform element-wise multiplication ( Figure 1 Cross product symbol ), and obtain the dynamically harmonized output signal X '. This step aims to adaptively recalibrate the contribution of each channel based on the statistical properties of the channel.

[0039] Perform preliminary spatiotemporal feature extraction: signal after Gaussian weighted harmonization X ' is fed into a 2D convolutional layer that uses e.g. The CSFH module uses a convolution kernel of different sizes to perform preliminary spatiotemporal feature extraction on the signal. This step is usually accompanied by a batch normalization (BatchNorm2d) layer to stabilize the training process and accelerate convergence, as well as a nonlinear activation function (such as ReLU or ELU) to increase the nonlinear expression ability of the model. c The feature coordination method (when the feature is fixed) can effectively strengthen the temporally stable neural patterns, suppress channel-specific artifacts, and achieve adaptive contrast enhancement, thereby providing higher quality and more discriminative feature inputs for subsequent modules.

[0040] like Figure 2 The HF-ECA module, shown in the upper-right center area, follows the CSFH module and is dedicated to further enhancing the extraction of spatiotemporal features. It also introduces a channel-level adaptive feature calibration mechanism to capture richer feature interactions. Its technical solutions include: The core of this module is a three-way parallel heterogeneous convolution architecture. The input feature map (dynamically harmonized motor imagery EEG signal feature map) is firstly passed through a convolution kernel size of (1, C )( C The depthwise convolution (DepthwiseConv2D) with the input channels corresponding to the EEG electrode dimensions performs channel-by-channel spatial filtering, aiming to capture the spatial distribution characteristics within each channel. Subsequently, each branch performs standard intermediate processing, such as batch normalization (BN), nonlinear activation (such as ReLU), pooling (such as Average Pooling) and possible dropout operations. The key lies in the heterogeneous design of the temporal feature extraction stage: one branch uses a standard two-dimensional convolution ( Figure 2 Conv2D in the network (acts on the time dimension) for dense temporal pattern learning; while the other two branches use separable convolution (SeparableConv2D), which decomposes the standard convolution into a depth-wise channel-by-channel temporal convolution and subsequent point-wise convolution ( Convolution), thereby significantly reducing the number of calculation parameters and computational complexity while maintaining the ability to model dynamic time series information, achieving a balance between representation ability and computational efficiency. The features extracted by these parallel branches are adjusted in dimension (for example, by Convolution makes the number of channels uniform to F2) and then fusion (such as Figure 1 As shown in the convergence point of the three branches, it is usually the concatenation or element-by-element addition of feature maps to obtain fused multi-channel and multi-scale spatiotemporal features.

[0041] The fused multi-channel and multi-scale spatiotemporal features are then optimized through the efficient channel attention (ECA) mechanism. First, the spatial (or spatiotemporal fusion here) dimension information of each channel is compressed into a single scalar descriptor through the global average pooling (GAP) operation. Unlike the traditional channel attention mechanism, ECA innovatively uses a one-dimensional convolution (1D-Conv) to efficiently model local cross-channel interactions directly in the channel dimension without the need for dimensionality reduction. The kernel size of this 1D convolution is not fixed, but is based on the number of channels after fusion. C out Adaptively through the function k To determine, the interaction range can be dynamically adjusted with the number of channels:

[0042] ; in, 、 b is a hyperparameter, It means taking the nearest odd number. C out represents the number of channels of input features, It is a function that adaptively calculates the size of the one-dimensional convolution kernel based on the number of channels.

[0043] After 1D convolution, a Sigmoid activation function is used to generate a channel attention weight vector ranging from 0 to 1. These weights are finally multiplied back to the feature map at the input of the ECA module (or before GAP) channel by channel (e.g. Figure 2 The last As shown in Figure 3), the adaptive recalibration of the characteristics of each channel is achieved, that is, the channel characteristics that contribute greatly to the classification task are strengthened, while the characteristics of the channels with small contributions or noisy channels are suppressed.

[0044] like Figure 2 The CW module shown in the lower area is centered around the MHA-RSTCN, a key component for deep temporal modeling in HFNet. It is responsible for extracting hierarchical temporal patterns and capturing long-term dependencies from the feature sequences output by the HF-ECA module. The detailed technical solution is as follows:

[0045] Characteristic sequence input to the CW module Z i May first be split into multiple overlapping sliding windows ( T w is the time step within the window, d is the feature dimension) in order to perform local and detailed analysis while preserving the continuity of the sequence.

[0046] Each windowed feature sequence It is then optimized through the Multi-head Self-Attention (MSA) mechanism. Figure 3 The structure of MSA is described in detail: the input feature sequence (usually Q, K and V are derived from the same input , after layer normalization) are mapped to different representation subspaces through independent linear projection layers (Linear). The projected feature sequences Q, K and V are fed into multiple parallel attention heads ( Figure 3 In each attention head, scaled dot-product attention (Scaled Dot-Product Attention, as shown in the figure) is performed. Figure 3(See the enlarged detail on the left in the figure): The output of each attention head is obtained by calculating the dot product of the transpose of Q and K (MatMul(Q, K^T)), scaling the dot product result, normalizing the attention score to a probability distribution using the SoftMax function, and finally weighting the attention weight with V (MatMul(AttentionWeights, V)). The output results of all parallel attention heads are concatenated and integrated again through a final linear projection layer (Linear) to obtain an optimized feature sequence for each sliding window. The entire MSA module typically also includes residual connections and layer normalization to promote information flow and stabilize training, thereby obtaining contextual representations with stronger awareness of local dependencies.

[0047] The feature sequence obtained by the MSA module is then input into the Residual Separable Temporal Convolutional Network (RSTCN) module. The RSTCN module consists of multiple stacked residual modules, each of which uses separable temporal convolution to efficiently extract temporal features. To capture temporal dependencies at different scales and expand the receptive field, the temporal convolutions in the RSTCN module typically use dilated convolutions, whose dilation rate increases exponentially with the depth of the network. To ensure strict temporal order (i.e., the output at the current moment depends only on the input at the past and current moments), the convolution operation uses causal padding. Each convolutional unit also includes standard batch normalization (BN), nonlinear activation functions (such as ELU or ReLU), and dropout layers to prevent overfitting. Multi-level residual connections run through the RSTCN module, fusing shallow features with deep features (aligning them through projected convolutions when dimensionality mismatches). This helps to balance fine-grained and coarse-grained temporal cues and enhance the model's discriminative capabilities.

[0048] The final output of the CW module is typically an aggregation of the features obtained from each window, forming a fixed-length feature vector. This vector is then fed into the final classification part of the network, which typically consists of one or more fully connected layers (Dense layers) and a Softmax activation function, and outputs the probability corresponding to each predefined motor imagery category.

[0049] In summary, the HFNet of the present invention works together through the adaptive coordination of the original signal by the CSFH module, the heterogeneous fusion and channel attention optimization of the multi-scale spatial spectrum features by the HF-ECA module, and the fine weaving of the deep temporal patterns by the CW module, to gradually transform the input motor imagery EEG signals into highly discriminative feature representations, and finally the classifier outputs accurate classification results, thereby effectively improving the decoding performance.

[0050] The embodiments of the present invention provide a coordinated fusion network (HFNet) architecture and a method for decoding motor imagery EEG signals. The architecture consists of three core modules connected in series: a time series-statistical feature coordinator (CSFH), a heterogeneous fusion-efficient channel attention module (HF-ECA), and a time series weaving module (CW). These modules collaboratively process raw EEG signals to achieve motor imagery task classification, effectively addressing issues such as low signal-to-noise ratio (SNR), non-stationarity, individual differences, and weak motor imagery features in EEG signals. The key point is this specific three-module series structure and the corresponding decoding method.

[0051] The present invention also proposes a motor imagery EEG signal decoding system, comprising: A signal acquisition module is used to obtain the time series data of each channel of the motor imagery EEG signal; A dynamic harmonization module is used to obtain time-averaged pooled features by averaging the time series data along the time dimension, perform intra-channel standardization on the features, and generate Gaussian weighted coefficients through Gaussian weighted harmonization; broadcast the Gaussian weighted coefficients along the time dimension to obtain weights that match the time dimension of the motor imagery EEG signal, and multiply the weights by the motor imagery EEG signal element by element to obtain a dynamically harmonized motor imagery EEG signal; The decoding module is used to extract the multi-channel and multi-scale spatiotemporal features of the dynamically coordinated motor imagery EEG signals; the feature sequence of the multi-scale spatiotemporal features is windowed to construct a feature vector of fixed length, and the feature vector is classified by a classifier to obtain the probability corresponding to each predefined motor imagery category.

[0052] The decoding method and system proposed in the present invention can capture the long-range dependencies, multi-level temporal dynamics and subtle timing patterns contained in EEG signals, enhance the model's ability to understand the signal evolution process, and thus effectively improve decoding performance.

[0053] The present invention also provides a collaborative working method for each core module in the network, mainly including: CSFH module: adaptively calibrates time domain features through temporal statistical characteristics to overcome individual differences; HF-ECA module: optimizes cross-scale spatial spectral features through multi-branch heterogeneous convolutional fusion and efficient channel attention; CW module: utilizes multi-head self-attention and residual separable temporal convolutional network (MHA-RSTCN) to finely model deep temporal dependencies.

[0054] The decoding method provided by the embodiment of the present invention has the following advantages: Improving the adaptive calibration capability for individual differences and non-stationarity of EEG signals: This paper aims to introduce novel feature preprocessing or early feature coordination mechanisms (such as the temporal-statistical feature coordinator CSFH in HFNet) to adaptively adjust and optimize the temporal distribution and statistical characteristics of the input EEG signals in the initial stage of feature learning, thereby enhancing the model's adaptability and robustness to signal changes in different individuals and time periods, and reducing the negative impact of individual differences.

[0055] Achieve efficient extraction and collaborative deep fusion of multi-dimensional and multi-scale EEG features: This invention aims to design a unique feature extraction and fusion architecture (such as the heterogeneous fusion-efficient channel attention module HF-ECA in HFNet) that can capture and effectively integrate heterogeneous feature information in multiple dimensions, such as spatial domain, spectral domain (or time-frequency domain), and time domain, from EEG signals in parallel and at multiple scales. At the same time, by introducing an advanced attention mechanism, it dynamically evaluates and strengthens key features and channels that contribute more to the classification task, achieving collaborative feature enhancement and effective information compression.

[0056] Deepen and refine the modeling and characterization of dynamic temporal dependencies of EEG signals: This invention aims to adopt advanced temporal information processing modules (such as the temporal weaving module CW in HFNet) to more effectively capture and characterize the long-range dependencies, multi-level temporal dynamics, and subtle temporal patterns contained in EEG signals, thereby improving the model's ability to understand the signal evolution process.

[0057] Comprehensively improve the overall performance and application potential of motor imagery EEG signal decoding: Through the innovative design and optimization of the above aspects, the ultimate goal of this invention is to significantly improve the accuracy, stability, and generalization ability of motor imagery EEG signal classification across different subjects and conditions. This will not only help promote theoretical research on motor imagery brain-computer interface technology, but also provide more reliable and efficient technical support for the practical application of this technology in fields such as neurorehabilitation, disability assistance, intelligent control, education, and entertainment, promoting its development towards practicality and universalization.

[0058] The above description is only a preferred specific embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Any technician familiar with the technical field, within the technical scope disclosed by the present invention, who makes equivalent replacements or changes based on the technical solution and inventive concept of the present invention, should be covered by the scope of protection of the present invention.

[0059] In addition, unless otherwise specified, all technical and scientific terms used in the present invention have the same meaning as commonly understood by those skilled in the art to which the present invention belongs. All documents mentioned in this specification are incorporated by reference to disclose and describe the methods related to the documents. In the event of any conflict with any incorporated document, the content of this specification shall prevail.

Claims

1. A method for decoding motor imagery EEG signals, characterized in that: The following steps are involved: Obtain the time series data of each channel of motor imagery EEG signal; By averaging the time series data along the time dimension, we obtain the time-averaged pooled features, normalize the features within the channel, and generate Gaussian weighted coefficients through Gaussian weighted harmonization. The Gaussian weighted coefficients are broadcasted along the time dimension to obtain weights that match the time dimension of the motor imagery EEG signal. The weights are then multiplied element-by-element by the motor imagery EEG signal to obtain a dynamically harmonized motor imagery EEG signal. The multi-channel and multi-scale spatiotemporal features of the dynamically coordinated motor imagery EEG signals are extracted; the feature sequence of the multi-scale spatiotemporal features is windowed to construct a feature vector of fixed length, which is then classified by a classifier to obtain the probability corresponding to each predefined motor imagery category.

2. The method for decoding motor imagery EEG signals according to claim 1, wherein: The method of extracting the multi-channel and multi-scale spatiotemporal features of the dynamic and coordinated motor imagery EEG signal includes: The dynamically harmonized motor imagery EEG signals are spatially filtered channel by channel to capture the spatial distribution characteristics within each channel and perform dimensionality adjustment and feature fusion in sequence to obtain fused multi-channel and multi-scale spatiotemporal features. A channel attention mechanism is introduced to determine the adaptive weight according to the number of channels of the fused multi-channel and multi-scale spatiotemporal features, and the fused multi-channel and multi-scale spatiotemporal features are adaptively recalibrated by generating a channel attention weight vector.

3. The method for decoding motor imagery EEG signals according to claim 2, wherein: The step of constructing a fixed-length feature vector includes: The feature sequence of adaptively recalibrated multi-channel and multi-scale spatiotemporal features is windowed, and the feature sequence of each sliding window is optimized through a multi-head attention mechanism; The optimized feature sequence of each sliding window is aggregated to form a feature vector of fixed length.

4. The method for decoding motor imagery EEG signals according to claim 3, wherein: The adaptive recalibration of the fused multi-channel multi-scale spatiotemporal features by generating a channel attention weight vector specifically includes: A channel attention mechanism is introduced to optimize the fused multi-channel and multi-scale spatiotemporal features. The spatial dimension information of each channel is compressed into a single scalar descriptor through a global average pooling operation. Local cross-channel interactions are modeled in the channel dimension through a one-dimensional convolution. The kernel size of this one-dimensional convolution is not fixed, but is adaptively determined by a function based on the number of channels after fusion. After performing one-dimensional convolution processing on the fused multi-channel and multi-scale spatiotemporal features, a channel attention weight vector ranging from 0 to 1 is generated through a Sigmoid activation function; the channel attention weight vector is multiplied by the spatiotemporal features of each channel to adaptively recalibrate the features of each channel.

5. The method for decoding motor imagery EEG signals according to claim 4, wherein: The multi-head attention mechanism is used to optimize the feature sequence of each sliding window, specifically including: The feature sequence of each sliding window is mapped to different representation subspaces through independent linear projection layers to obtain Q, K and V of the projected feature sequence; The projected feature sequences Q, K, and V are input into multiple parallel attention heads. Within each attention head, scaled dot product attention is performed. The dot product of the transposed Q and K is calculated, the dot product result is scaled, and the attention score is normalized into a probability distribution through the SoftMax function. The probability distribution is weighted and summed with V to obtain the output result of each attention head. The output results of all parallel attention heads are concatenated and integrated again through a final linear projection layer to obtain the optimized feature sequence of each sliding window.

6. The method for decoding motor imagery EEG signals according to claim 5, wherein: The method of obtaining the probabilities corresponding to the predefined motor imagery categories is specifically to input a feature vector of a fixed length into a classifier composed of one or more fully connected layers and a Softmax activation function, and output the probabilities corresponding to the predefined motor imagery categories through classification by the classifier.

7. The method for decoding motor imagery EEG signals according to claim 6, wherein: The fused multi-channel multi-scale spatiotemporal features are obtained by using a three-way parallel heterogeneous convolution architecture to transform the dynamic harmonic motor imagery EEG signal feature map at the beginning of each branch through a convolution kernel size of The depth convolution performs channel-by-channel spatial filtering, C is the number of input channels, corresponding to the EEG electrode dimension, capturing the spatial distribution characteristics within each channel; The spatial distribution features within each channel are subjected to standard intermediate processing, including batch normalization, nonlinear activation, pooling and dropout operations; one branch uses standard two-dimensional convolution acting on the time dimension to perform dense temporal pattern learning; the other two branches use separable convolution to decompose the standard convolution into depth-wise channel-by-channel temporal convolution and subsequent point convolution; The features extracted by the three parallel branches are dimensionally adjusted and then the feature maps are concatenated or added element by element to obtain fused multi-channel and multi-scale spatiotemporal features.

8. A motor imagery EEG signal decoding system, characterized in that: include: A signal acquisition module is used to obtain the time series data of each channel of the motor imagery EEG signal; A dynamic harmonization module is used to obtain time-averaged pooled features by averaging the time series data along the time dimension, perform intra-channel standardization on the features, and generate Gaussian weighted coefficients through Gaussian weighted harmonization; broadcast the Gaussian weighted coefficients along the time dimension to obtain weights that match the time dimension of the motor imagery EEG signal, and multiply the weights by the motor imagery EEG signal element by element to obtain a dynamically harmonized motor imagery EEG signal; The decoding module is used to extract the multi-channel and multi-scale spatiotemporal features of the dynamically coordinated motor imagery EEG signals; the feature sequence of the multi-scale spatiotemporal features is windowed to construct a feature vector of fixed length, and the feature vector is classified by a classifier to obtain the probability corresponding to each predefined motor imagery category.

Citation Information

Cited By

  • Electroencephalogram signal decoding method based on time domain and frequency domain signal fusion

    CN122286669A