Motor imagery classification neural network method based on multi-scale spatial-temporal feature fusion

By introducing a neural network method of multi-scale spatiotemporal and spatial feature fusion into the motion imagination brain-computer interface, the problems of insufficient multi-scale dependence modeling in the time domain, lack of adaptability in the spatial feature selection and incomplete spatial feature fusion mechanism are solved, and more efficient EEG signal decoding and better system robustness are achieved.

CN120105201APending Publication Date: 2025-06-06NANJING UNIV OF AERONAUTICS & ASTRONAUTICS
View PDF 0 Cites 4 Cited by

Patent Information

Application Number
CN202510264972.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-06
Publication Date
2025-06-06

AI Technical Summary

Technical Problem

The existing motion imagination brain-computer interface decoding methods are insufficient in the time domain multi-scale dependence modeling, lack of adaptability in the selection of spatial domain feature and imperfect spatial and temporal feature fusion mechanism.

Method used

A motion imagination classification neural network method based on multi-scale spatial and temporal feature fusion is proposed. The short- and long-term features of EEG signals are captured through the multi-scale time domain convolution attention module, the adaptive channel weight module dynamically optimizes the spatial channel weight, and deeply fusions the spatial and temporal features through the improved Transformer architecture.

Benefits of technology

It significantly improves the decoding accuracy and system robustness of EEG signals, can better adapt to different input data and application scenarios, and improves the decoding accuracy and stability in motion imagination tasks.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120105201A_ABST
    Figure CN120105201A_ABST
Patent Text Reader

Abstract

The invention provides a motor imagery classification neural network method based on multi-scale spatial-temporal feature fusion. According to the method, the EEG signal decoding precision in a motor imagery task is effectively improved by designing a multi-scale spatial-temporal feature fusion network. The method comprises the following steps: firstly, processing an original EEG signal by adopting standardized preprocessing; secondly, extracting features of the signals under different time scales through a multi-scale time sequence attention module, and enhancing a long-term time sequence dependency relationship in combination with a self-attention mechanism; thirdly, a self-adaptive channel weighting module is introduced, the weight of each electrode signal is dynamically adjusted, and representation of spatial features is optimized; and finally, fusing the spatio-temporal characteristics through a Transform module, capturing complex interaction of the spatio-temporal dependency relationship, and finally outputting a decoding result through a classifier. The method can effectively overcome the defects of an existing method in the aspects of spatial-temporal feature fusion and dependent modeling, improves the accuracy of EEG signal classification and the system robustness, and has a wide application prospect.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of brain-computer interface technology, and is a neural network method for classifying motor imagery EEG signals. The method can efficiently decode the user's intention in motor imagery tasks through multi-scale spatiotemporal feature fusion, and is widely used in medical rehabilitation, human-computer interaction, intelligent assistance and other fields. Background Art

[0002] Brain-computer interface technology has developed rapidly in recent years. It enables direct interaction between the human brain and external devices by decoding brain signals, and has shown great application potential in the fields of medical rehabilitation, human-computer interaction, etc. Among them, motor imagery allows users to generate EEG signals by imagining body movements, thereby achieving control of devices or assisting rehabilitation treatment. However, due to the high nonlinearity, high noise and low signal-to-noise ratio characteristics of EEG signals, accurately decoding the user's movement intentions remains a very challenging problem.

[0003] Traditional convolutional neural network (CNN)-based models (such as Deep ConvNet, EEGNet, etc.) extract short-term features of EEG signals through local convolution kernels, but their inherent structure makes it difficult to capture long-term dependencies in time series. Although the improved model based on the attention mechanism improves the global time series modeling capability by integrating the Transformer module, existing methods mostly focus on dependency analysis at a single time scale and lack collaborative modeling of multi-scale long- and short-term time series patterns.

[0004] The spatial distribution characteristics of EEG signals are closely related to the activation patterns of brain functional areas, but existing spatial feature extraction methods have significant defects: methods based on traditional spatial filtering (such as CSP, SFBCSP, EM-CSP) rely on manual parameter adjustment and have high computational complexity, making it difficult to adapt to neurophysiological differences between individuals; automatic feature learning models based on deep learning (such as LightConvNet and MSFNet) avoid manual intervention, but they use a homogenization processing strategy and do not fully consider the heterogeneity of the contribution of electrodes in different brain regions to the task. For example, in the motor imagery task, the electrode signals of the sensorimotor cortex contain core discriminant information, while areas such as the prefrontal lobe may introduce redundant noise, but existing methods cannot dynamically identify key electrode channels, resulting in reduced feature expression efficiency.

[0005] In addition, the current mainstream architecture (whose spatiotemporal features are usually integrated through simple concatenation or weighted addition) fails to fully explore the deep interactive relationship of spatiotemporal features.

[0006] In summary, developing a new decoding architecture that can break through the limitations of existing technologies and combine multi-scale timing analysis, dynamic spatial optimization and deep feature fusion capabilities is the core challenge in this field. This is of great significance for improving the performance and reliability of brain-computer interface systems, and is also the key to promoting the practical application of brain-computer interface technology. Summary of the invention

[0007] In view of the problems of insufficient modeling of multi-scale dependence in the time domain, lack of adaptability in spatial feature selection and imperfect spatiotemporal feature fusion mechanism in existing motor imagery brain-computer interface decoding methods, the present invention proposes a motor imagery classification neural network method based on multi-scale spatiotemporal feature fusion, which aims to significantly improve the EEG signal decoding accuracy and system robustness by collaboratively modeling multi-scale temporal features, dynamically optimizing spatial channel weights and deeply fusing spatiotemporal information.

[0008] First, the original EEG signal is standardized and preprocessed to separate the task-related EEG components from the environmental noise. Secondly, a multi-scale time-domain convolutional attention module is designed to capture the short-term features of the EEG signal by deploying multiple different one-dimensional convolution kernels, and the long-term temporal dependencies are strengthened by combining the self-attention mechanism. Next, an adaptive channel weight module is introduced to dynamically generate a spatial weight matrix based on the variance of each electrode channel, suppressing redundant information in noise-sensitive areas such as the prefrontal lobe, while enhancing the feature expression of high-value brain areas such as the sensorimotor cortex. Finally, the improved Transformer architecture is used to deeply fuse spatiotemporal features, explore dynamic interaction patterns across brain regions and time points, and output the motion intention decoding results through a lightweight classifier. BRIEF DESCRIPTION OF THE DRAWINGS

[0009] Figure 1 It is the overall flow chart of the method framework of the present invention;

[0010] Figure 2 It is an overall framework diagram of the system to which the method described in the present invention is applied. DETAILED DESCRIPTION

[0011] The present invention proposes a multi-scale spatiotemporal feature fusion network method for motor imagery electroencephalogram (EEG) signal classification. Through the synergistic effect of three key modules, the network successfully solves the shortcomings of traditional methods in capturing short-term and long-term dependencies of EEG signals, and effectively improves the decoding accuracy in motor imagery tasks. In order to make the purpose, advantages and technical solutions of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be fully and clearly described in conjunction with the accompanying drawings.

[0012] S1, data preprocessing;

[0013] The input EEG signal data is preprocessed to ensure that the data input to the network has good quality. First, the EEG signal is segmented and data enhanced. The EEG signal is segmented into time windows of fixed length. These time windows are randomly spliced ​​to expand the diversity of the data set and avoid additional data collection. This method helps to improve the robustness of the model, especially when the number of samples is limited. The EEG signal is Z-score standardized to reduce the variability between subjects and ensure the balance of the signal, thereby accelerating the convergence of the model. Specifically, for each channel in each sample, its mean and standard deviation are calculated.

[0014]

[0015] In the formula, x i represents the original data, μ and σ are the mean and standard deviation of the channel respectively, x 0 is the standardized data.

[0016] S2, extracting multi-scale spatiotemporal features of EEG signals by designing a multi-scale temporal convolutional attention module;

[0017] Taking the preprocessed motor imagery EEG data as input, a multi-scale temporal convolution attention module is designed. This module extracts the local features of the signal at different time scales through six convolution layers with different convolution kernel sizes. First, the EEG signal is input into the multi-scale convolution module, and the signal is processed in parallel using different convolution kernel sizes to capture the features of different time scales. The formula of the convolution operation of the lth layer is as follows:

[0018]

[0019] Where x′ is the enhanced EEG signal, is the output of the l-th convolution operation.

[0020] After each convolutional layer, batch normalization is used to standardize the output of the convolutional layer, thereby reducing the internal covariance offset and accelerating the training process of the model. The six convolutional outputs are then concatenated to obtain

[0021] Furthermore, in order to capture global dependencies, the present invention also introduces a self-attention mechanism. The output x after multi-scale convolution is concatenated t Enter the self-attention network, and adjust the importance of each time point in the sequence by calculating the attention weights between time series data points, thereby enhancing the model's ability to learn time series information. The calculation formula for self-attention is as follows:

[0022]

[0023] Where Q, K, and V represent query, key, and value respectively. k The dimension of the key.

[0024] The attenuation parameter is introduced to effectively adjust the importance of each feature in the self-attention mechanism, thereby improving the learning process of the network. The formula is as follows:

[0025] x t =γ·Attention(Q,K,V)

[0026] Where γ is the attenuation parameter, x t It is the product of the decay parameter and the output of the self-attention network.

[0027] Next, the extracted timing features are further refined, and a timing encoder is designed to enhance the timing characteristics of the signal. The encoder formula is as follows:

[0028] z t =Encoder t (x t )

[0029] In the formula, Encoder t (x t ) consists of a convolutional layer, batch normalization, exponential linear unit activation, average pooling layer, and dropout layer. t It is the temporal feature representation obtained by the temporal encoder.

[0030] S3, evaluate the importance of different electrodes in the EEG signal by calculating the variance of each electrode;

[0031] Different electrodes in the EEG signal contribute differently to the EEG activity, so an adaptive channel weighting module is designed to determine the importance of different electrodes. This module calculates the variance of each channel and dynamically adjusts the channel weights, thereby emphasizing the electrodes that are more important to the classification task and suppressing redundant information. The specific implementation is to first calculate the variance of each channel:

[0032]

[0033] In the formula, x′ Ct is the signal value of the C channel at time point t, μ C is the mean value of channel C.

[0034] Using the calculated channel variance, we get the channel weight matrix, which reflects its relative importance. The learning formula of channel weight is as follows:

[0035] W=σ(W 2 ·Relu(W 1 Norm (Var)

[0036] Where σ is the Sigmoid activation function, W 1 and W 2 is the weight of the fully connected layer, and Norm(Var) is the normalized variance vector.

[0037] Next, multiply the weight matrix by the original data, the formula is as follows:

[0038] x s =x′⊙W

[0039] Design a spatial encoder to extract spatial features. The encoder formula is as follows:

[0040] z s =Encoder s (x s )

[0041] In the formula, Encoder s (x s ), which consists of three convolutional layers and one pooling layer. s is the spatial feature representation obtained by the spatial encoder

[0042] S4, using transformer to fuse features from time and space dimensions to capture complex spatiotemporal dependencies;

[0043] Concatenate the spatiotemporal features to obtain feature z h , input to the Transformer network for fusion processing. Transformer can capture the dependency between temporal and spatial information at the same time. The formula is:

[0044]

[0045] In the formula, h represents different heads, Q h ,K h , V h They represent the query, key, and value of the h-th header respectively.

[0046] The output of the self-attention mechanism is further transformed nonlinearly through the feedforward neural network to enhance the feature expression ability of the model. The calculation formula of the feedforward network is:

[0047] FFN(z h )=Norm(b 2 ·Relu(z h W 1 +b 1 )+z h W 2 )

[0048] Where b 1 and b 2 is the bias vector, W 1 and W 2 is the weight matrix.

[0049] S5, classification and output layer;

[0050] The features processed by the Transformer fusion module are input into the classifier for motion imagery task classification, and the predicted probability of the category is output. The cross entropy loss function is used for classification training. The formula of the classification loss function is as follows:

[0051]

[0052] In the formula, y i is the true label, is the predicted probability of the i-th class.

[0053] Through the above steps, the present invention provides a motor imagery classification neural network method based on multi-scale spatiotemporal feature fusion. The method can effectively capture the complex time dependency in EEG signals through a carefully designed multi-scale convolution mechanism, thereby improving the expression ability of signal features. By adaptively adjusting the weights of the channels, redundant information can be removed to ensure that the model focuses on the most discriminative signal features. On the basis of spatiotemporal feature extraction, combined with the Transformer fusion module, the method further optimizes the integration and representation of spatiotemporal features and enhances the model's learning ability for long-term and short-term temporal dependencies. In addition, the multi-level fusion and self-attention mechanism of the model significantly improve the accuracy and robustness in the classification task, so that the method can better adapt to different input data and application scenarios when processing complex EEG signals. Finally, the method of the present invention not only improves the accuracy of EEG signal classification, but also enhances its stability and reliability under different circumstances, and has significant practical application value.

Claims

1. A motor imagery classification neural network method based on multi-scale spatiotemporal feature fusion, the method comprising the following steps: S1. Standardize and preprocess the original EEG signal and expand the diversity of the data set through signal segmentation and data enhancement; S2. Design a multi-scale temporal convolutional attention module, which uses multiple convolutional layers with different kernel sizes and combines the self-attention mechanism to extract the features of EEG signals at different time scales and enhance the long-term temporal dependencies in the signals. S3, introduces an adaptive channel weighting module to dynamically generate a spatial weight matrix by calculating the variance of each electrode to optimize the feature expression of the electrode; S4, through the Transformer module, deep integration of spatiotemporal features is carried out to explore dynamic interaction patterns across brain regions and time points; S5. Use a lightweight classifier to classify the fused spatiotemporal features and output the decoding result of the motion intention.

2. The multi-scale temporal convolutional attention method described in claim 1 processes EEG signals in parallel through six convolutional layers with different convolution kernel sizes, captures temporal features at different time scales, and captures short-term and long-term dependencies at different time scales through a self-attention mechanism.

3. The adaptive channel weighting method described in claim 1 determines the importance of the electrode by calculating the variance of each electrode signal, and dynamically adjusts the weight of the spatial channel according to the variance of the electrode, thereby suppressing redundant information and enhancing the characteristic expression of key brain areas.

4. The Transformer-based method described in claim 1 fuses spatiotemporal features through a multi-head attention mechanism to capture the deep interaction of spatiotemporal dependencies in EEG signals.

5. The classification step described in claim 1 is to use a cross entropy loss function to train the spatiotemporal features and output the classification results of motion intention.

Citation Information

Cited By

  • Abnormal mode detection system for cardiac electrical ablation signals

    CN120788593A

  • Attention interpretable electroencephalogram decoding method based on adaptive fuzzy convolution and TSK guidance

    CN121167432A

  • Motor imagery eeg classification method based on multi-domain entropy heterogeneous gating and time difference

    CN122548446A

  • Motor imagery eeg classification method based on multi-domain entropy heterogeneous gating and time difference

    CN122548446B