Multi-scale feature-based motor imagery electroencephalogram signal classification method of 3D convolutional neural network
By using the 3D spatial attention module and multi-scale temporal attention module of 3D convolutional neural network in the MI EEG classification task, multi-scale features are extracted, and the problem of difficulty in decoding MI EEG signals in the prior art is solved, and higher classification accuracy and generalization performance are achieved.
Patent Information
- Application Number
- CN202510437136.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-09
- Publication Date
- 2025-05-09
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
The prior art is difficult to effectively decode low signal-to-noise ratio signals and individual variance in motor imagined electroencephalogram signals (MI EEG), resulting in poor generalization ability and poor classification performance of MI EEG classification tasks.
An end-to-end 3D convolutional neural network is proposed to automatically extract multi-scale spatial and temporal-related features of motor imagination EEG signals through 3D spatial attention module and multi-scale temporal attention module to improve classification performance.
By adaptively assigning weights, discriminant features related to motion imagination are extracted, the accuracy and generalization performance of MI EEG classification tasks are improved, and the impact of biological and environmental artifacts is reduced.
Smart Images

Figure CN119961789A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of motor imagery EEG signal classification, and specifically is a motor imagery EEG signal classification of a 3D convolutional neural network with multi-scale features. Background Art
[0002] A brain-computer interface (BCI) is a system that connects the human brain to computing devices by decoding neuronal activity. Electroencephalogram (EEG) signals have become one of the most widely used BCI signals due to their non-invasive, high temporal resolution, and low cost. Motor imagery electroencephalogram (MI EEG) signals are a BCI paradigm that reflects the user's voluntary and conscious movements without any external stimulation. When the subject actively imagines body movements, the sensorimotor cortical areas of the contralateral and ipsilateral hemispheres of the brain are activated. (8~12 Hz) and The power of the (16~26 Hz) rhythm will decrease or increase, which is called event-related desynchronization (ERD) and event-related synchronization (ERS), respectively. The core problem of MI EEG classification is to effectively decode the low signal-to-noise ratio (SNR) and significant individual variance of MI EEG signals into correct instructions. In this study, our goal is to accurately analyze brain activity to help post-stroke and paralyzed patients solve communication problems with the outside world.
[0003] According to the input format definition, there are currently two main research branches in deep learning-based multi-class MI EEG classification methods. One is to take the feature maps extracted from the original MI EEG signal as input, and the other is to focus directly on the input format as the original MI EEG signal. The former represents the MI EEG signal as a series of two-dimensional feature maps as the input format, and reduces noise and enhances low signal-to-noise ratio signals through manually selected feature extraction methods. However, the potential problem is that the extracted features must be manually designed by human experts. More importantly, MI EEG signals are non-stationary and can be easily corrupted by various biological fluctuations and events (such as blinking, muscle artifacts, fatigue, and attention level), which makes it difficult to manually select appropriate feature extraction methods across subjects, resulting in poor generalization ability and poor classification performance for multi-class classification tasks. The latter represents the original signal as a two-dimensional array, which uses the number of time sampling points as the array width and the number of electrodes as the array height. However, when the original MI EEG is represented as a 2D array input in the above way, deep learning-based methods usually ignore the spatial dependency of MI EEG data, which has been shown to be important for improving classification performance. The correlation between nearby sampled electrodes cannot be fully reflected in the 2D array either. Therefore, the final performance of the MI classification model will be affected. At the same time, some MI EEG classification methods that only use 3D convolutional networks have eliminated the above effects, but they do not perform well in classification tasks because they do not comprehensively extract multi-scale features in the spatial and temporal domains. Summary of the invention
[0004] In order to solve the above problems, the present invention proposes an end-to-end 3D convolutional neural network for extracting multi-scale spatial and temporal related features for motor imagery EEG signal classification tasks. The network is specifically composed of a 3D representation, a 3D spatial attention module, a multi-scale temporal attention module and a classification module. The goal of the 3D spatial attention module and the multi-scale temporal attention module is to adaptively assign higher weights to motion-related spatial channels and temporal sampling cues than to motion-irrelevant ones in all brain regions. They can define new compact feature representations of MI EEG in spatial and temporal domains and prevent the influence of biological and environmental artifacts to improve classification performance.
[0005] The method for classifying motor imagery EEG signals based on a 3D convolutional neural network with multi-scale features includes the following steps: S1: Collecting motor imagery EEG signal data; Collecting the motor imagery EEG signal data of the subject using an EEG signal acquisition device; S2: preprocessing the motor imagery EEG signal data; The motor imagery EEG signal data is represented by a three-dimensional tensor; the first dimension of the three-dimensional tensor represents the number of discrete time sampling points, and the second and third dimensions of the three-dimensional tensor represent the spatial positions of the electrode channels of the EEG signal acquisition device; S3: Extract spatial features from the three-dimensional tensor to obtain spatial features ; S4: For the spatial features Extract time features and obtain time features ; S5: The time feature To classify; The time feature Input to the fully connected layer and perform classification prediction through the Softmax function.
[0006] Preferably, step S3 comprises: The three-dimensional tensor is input into three separable 3D convolution blocks to generate 3D spatial feature blocks respectively. and ; The 3D spatial feature block and Reconstruct and transpose to obtain the transformed 3D space feature blocks and ; In the converted 3D space feature block First, matrix multiplication is performed, and then the spatial attention weight matrix is obtained through the Softmax function. ; The spatial attention weight matrix Used to represent the similarity weight between the electrode channels of the EEG signal acquisition device; In the converted 3D space feature block and the spatial attention weight matrix Perform matrix multiplication between to obtain a three-dimensional tensor space representation ; The three-dimensional tensor space is represented as Input to Convolution kernel, generating weighted spatial features ; Preset learnable parameters With the weighted spatial features Multiply and perform element-by-element summation on the three-dimensional tensor to obtain the spatial features .
[0007] Preferably, step S4 comprises: The spatial features Divided into three time slices; Input two of the time slices into two 2D convolution blocks respectively to obtain the initial time features and ; The initial time feature and the remaining time slices are reconstructed and transposed to obtain the conversion time features and ; The conversion time characteristics First, matrix multiplication is performed, and then the time attention weight matrix is obtained through the Softmax function. The temporal attention weight matrix Used to represent the similarity weight between the time slices; The conversion time characteristics and the temporal attention weight matrix Perform matrix multiplication between to obtain weighted time features ; Preset learnable parameters With the weighted time feature Multiply and perform element-by-element summation with the time slice to obtain the scale The time characteristics of the following .
[0008] Preferably, step S5 comprises: The temporal features are processed by an average pooling layer with a logarithmic nonlinear active function , reducing the time characteristics The time dimension of The low-level features in the ts are converted into high-level abstract features, and the high-level abstract features are connected into multi-scale spatiotemporal features. ; The multi-scale spatiotemporal features Input to the fully connected layer and perform classification prediction through the Softmax function.
[0009] Preferably, the 3D convolution block is The convolution kernel structure is the same.
[0010] Preferably, the step S1 comprises: the EEG signal acquisition device is an EEG helmet; the EEG helmet is calibrated; and the subject's scalp is cleaned with conductive paste or saline.
[0011] Preferably, during the calibration process, a standard resistance solution is used to test the contact impedance of the EEG helmet electrodes.
[0012] Preferably, the step S2 comprises: Performing bandpass filtering on the motor imagery EEG signal data; The motor imagery EEG signal data is normalized.
[0013] Preferably, the standardized motor imagery EEG signal data are divided, wherein 80% of the motor imagery EEG signal data are used for network training, and 20% of the motor imagery EEG signal data are used for testing.
[0014] The present invention has the following beneficial effects: This paper proposes an end-to-end 3D convolutional neural network to extract multi-scale spatial and temporal features to improve the accuracy performance of MI EEG classification tasks.
[0015] The 3D spatial attention module automatically assigns higher weights to most motion-related channels and lower weights to motion-irrelevant channels through the attention mechanism, extracting discriminative spatial features related to motor imagery and the correlation between any two electrode channels. Since it is independent of subjects, motor imagery tasks, and manually selected parameters, it can eliminate artifacts caused by manually selected channels and adaptively improve the accuracy of motor imagery classification tasks for different subjects.
[0016] The multi-scale temporal attention module uses the attention mechanism to assign adaptive weights to different time slices to meet the EEG analysis needs of different subjects in different time periods, eliminating the adverse effects of different subjects not being able to maintain concentration for a long time during the experiment. At the same time, since it does not rely on the different forms and contents of motor imagery tasks, it also improves the generalization performance of the network. BRIEF DESCRIPTION OF THE DRAWINGS
[0017] Figure 1 Schematic diagram of the process principle of the present invention Figure 2 Schematic diagram of the 3D spatial attention module structure Figure 3 Schematic diagram of the multi-scale temporal attention module structure DETAILED DESCRIPTION
[0018] In order to make the purpose, technical solution and advantages of the present invention more clearly understood, the present invention is further described in detail below in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not intended to limit the embodiments of the present invention.
[0019] like Figure 1 As shown, the classification of motor imagery EEG signals based on a 3D convolutional neural network with multi-scale features, the method flow is shown in the figure, and includes the following steps: Step S1: Use high-precision EEG signal acquisition equipment and EEG helmet to collect the subject's motor imagery EEG signal data. Ensure that the data acquisition environment is quiet and reduce interference signals.
[0020] In step S1, the EEG helmet should have the following characteristics: high channel number, high sampling rate, low noise, and good conductivity. In the preparation stage, the EEG helmet needs to be calibrated to ensure that the signals of all electrodes are in normal working condition. During the calibration process, the contact impedance of the electrode can be tested with a standard resistance solution. Next, the subject is prepared, including cleaning the scalp to reduce resistance and using conductive paste or saline to improve the contact effect between the electrode and the scalp.
[0021] Step S2: Data preprocessing: The method for preprocessing the EEG signal in step S2 is: Step S21: Using EEG signals with little preprocessing as input to the deep learning model can obtain more competitive results. Therefore, the present invention only performs bandpass filtering of 0.5-100 Hz on the MI-EEG dataset.
[0022] Step S22: The MI-EEG signals of the data set are standardized before being input into the model, and the formula is:
[0023] In the formula is the number of EEG signal channels.
[0024] Step S23: The definition of 3D representation is as follows: is the original MI EEG signal, where The first The subject's The trials are represented as a traditional 2D matrix, where the width of the matrix is the number of discrete time sampling points ( ), the height of the matrix is the number of electrode channels ( ). However, as mentioned above, the traditional 2D representation Therefore, we use the positions of EEG electrode channels to transform the traditional 2D matrix into Expanded to 3D tensor . This tensor can be expressed as , still represents the number of discrete time sampling points, while The plane is numbered k and corresponds to the electrode channel position in the standard 10 / 20 system of the International Electroencephalographic Society. The number k in indicates that it has the same relative position as electrode channel k, and its time sampling value is The data is formed into continuous data. k = 0 means there is no electrode channel and its time sampling value is zero. Add zero to The purpose is to Keep it as a 3D cube tensor to support the use of 3D convolution without introducing any noise.
[0025] This 3D representation not only uses the electrode distribution to explicitly preserve the relative spatial topological information between electrode channels, but also uses the sequential form of time sampling values to preserve the temporal information. Therefore, it is easy to use 3D convolution to extract spatiotemporal features. The data represented is divided into 80% of the data for network training and 20% for testing.
[0026] Step S3: Use the 3D spatial attention module to extract spatial features.
[0027] The 3D spatial attention module learns It automatically assigns higher weights to most motion-related channels and lower weights to channels that are not motion-related, such as Figure 2 The detailed steps are as follows: 3D Tensor It is first input into three separable 3D convolutional blocks , , Generate different 3D spatial feature blocks , and . , and belong and is 4, indicating the number of feature blocks. Next, , and will reconstruct and transpose ( , and ) are converted to different sizes, such as , and , used to perform matrix multiplication between them ( ). Then, the softmax function is applied to and To obtain the spatial attention weight matrix ( ).
[0028]
[0029] in ( )express The number of electrode channels in . The range is 0 to 1, indicating No. and The higher the electrode correlation characteristic, the higher the weight. The bigger it is.
[0030] implement and Another matrix multiplication between to obtain . yes The new attention-based spatial representation is obtained by ) adaptively aggregates the 3D spatial features of other electrode channels to update each electrode channel.
[0031] one Convolution kernel ( ) is because it has a strong ability to reduce the dimension of feature blocks, and The convolution operation only focuses on the feature block dimension, and the number of input data is constant. Therefore, it is suitable for efficiently processing a large number of input channels.
[0032] Among them, the 3D convolution block , , , The structure is the same.
[0033] Finally, by setting the learnable parameters Multiply , and the original EEG 3D representation By performing element-by-element summation, we obtain the attention-based adaptive spatial features as follows:
[0034] is a real number, and the optimistic value is automatically learned in the training of the proposed architecture. Compared with other spatial feature extractions (such as 2D convolution), the 3D spatial attention module focuses on the location of electrode channels in the real world. It enhances valuable motion-related features based on 3D spatial information and suppresses useless motion-irrelevant features, which is more consistent with the reality that different brain functional areas may have a certain impact on different motor imagery tasks for different subjects. Therefore, by innovatively introducing the attention mechanism into 3D convolution, this module complements the existing 3D neural network.
[0035] Step S4: extract the spatial features from step S3 is cut into three time slices ( , is the length of each time slice, is a scale along the time dimension The number of time slices) is fed into the designed multi-scale temporal attention module for temporal feature extraction.
[0036] To better extract time-invariant high-level features within each time slice, we use an attention mechanism to assign adaptive weights to different time slices to meet the requirements of EEG analysis, where different subjects focus on different time periods. Compared with previous methods, the multi-scale temporal attention module does not depend on subjects or tasks and is more robust to new subjects and tasks.
[0037] by Take the multi-scale temporal attention module processing as an example, Figure 3 As shown, is fed into two separate 2D convolutional blocks of the same structure ( and ) to generate the initial time features and , their sizes are Then, , and Reconstruct and transpose ( , and )generate , and , whose size is , and , used to perform matrix multiplication between them ( ). Then, the softmax function is applied to and To obtain the temporal attention weight matrix ( ).
[0038]
[0039] in It is a scale The number of time slices. It is under test and The similarity weight between time slices in . Focus on specific motion-related time slices that are easier to distinguish than other motion-independent slices. The bigger, and Then, the weighted sum of all EEG time slices is calculated ( ) to learn attention-based temporal representations ( ). Finally, another learnable parameter and Multiply and Perform element-wise sum operation to obtain the scale The attention-based adaptive temporal features under are as follows:
[0040]
[0041] Step S5: Input the obtained multi-scale features into the classification module for classification.
[0042] Next, the An attention-based temporally adaptive representation ( ) is fed into an average pooling layer with a logarithmic nonlinear active function (NA) to aggregate features in the time dimension in parallel. It further reduces the time dimension and transforms low-level features into high-level abstract features ( , and ), these features are concatenated into multi-scale spatiotemporal features .
[0043] Multi-scale spatiotemporal features The input is sent to the fully connected layer and the classification prediction is performed through the Softmax function. All the above convolutional layers accelerate network training through batch normalization and activate nonlinearity through the Exponential Linear Unit (ELU).
[0044] The following are the specific experimental results and instructions added: The public MI EEG dataset IV-2a is used to evaluate the proposed method. IV2a is a 25-channel dataset with 4 categories of MI tasks (left hand, right hand, foot, and tongue) from 9 healthy subjects. It contains 72 trials for each subject. The total number of trials is 5184. This experiment uses accuracy (Acc) as the evaluation indicator of model performance, in %, and the specific description is as follows:
[0045]
[0046] Where: is a positive sample, i.e. The number of correctly predicted samples in the class; For the The number of samples in the class; The number of classes.
[0047] As shown in Table 1, the present invention performs best in the four-category MIEEG classification task, with an average accuracy of 92.8%. The machine learning-based method TSGSP is limited by human expert knowledge and experience, resulting in the extracted features not being the most suitable for classification when different subjects show significant EEG dynamic features in different motor imagery tasks. Therefore, the average accuracy of the present invention is 11.5 percentage points higher than that of TSGSP (81.3%). Then, it is compared with two deep learning methods. HSCNN performs convolution on mixed scales to improve classification. Three kernel sizes with distance distribution are used to extract EEG information in time, space and frequency domains to adapt to different subjects. However, these methods use the original MIEEG as a 2D array input, omitting the latent spatial information in 3D space. In contrast, the present invention solves the above problems, and the average accuracy is 3.4% higher than that of HSCNN (89.4%). DJDAN proposes a dynamic joint domain adaptation convolutional neural network to learn discriminative features for MIEEG classification, while reducing the marginal and conditional distribution differences across domains through global and local discriminators. However, in the limited convolutional layers, the size of a single receptive kernel limits DJDAN from extracting high-level features to improve classification performance. The average accuracy of our method is 2.2% higher than that of DJDAN (90.6%).
[0048] Table 1 Performance comparison of the present invention and other methods on the IV-2a dataset
[0049] The above descriptions are only some specific implementation methods of the present invention, but the protection scope of the present invention is not limited thereto. Any changes or substitutions that can be easily thought of by any person familiar with the art within the technical scope disclosed in the present invention should be covered within the protection scope of the present invention.
Claims
1. A motor imagery EEG signal classification method based on a 3D convolutional neural network with multi-scale features, characterized in that: The following steps are involved: S1: Collecting motor imagery EEG signal data; Collecting the motor imagery EEG signal data of the subject using an EEG signal acquisition device; S2: preprocessing the motor imagery EEG signal data; The motor imagery EEG signal data is represented by a three-dimensional tensor; the first dimension of the three-dimensional tensor represents the number of discrete time sampling points, and the second and third dimensions of the three-dimensional tensor represent the spatial positions of the electrode channels of the EEG signal acquisition device; S3: Extract spatial features from the three-dimensional tensor to obtain spatial features ; S4: For the spatial features Extract time features and obtain time features ; S5: The time feature To classify; The time feature Input to the fully connected layer and perform classification prediction through the Softmax function.
2. The method for classifying motor imagery EEG signals based on a 3D convolutional neural network with multi-scale features according to claim 1, characterized in that: The step S3 comprises: The three-dimensional tensor is input into three separable 3D convolution blocks to generate 3D spatial feature blocks respectively. and ; The 3D spatial feature block and Reconstruct and transpose to obtain the transformed 3D space feature blocks and ; In the converted 3D space feature block First, matrix multiplication is performed, and then the spatial attention weight matrix is obtained through the Softmax function. ; The spatial attention weight matrix Used to represent the similarity weight between the electrode channels of the EEG signal acquisition device; In the converted 3D space feature block and the spatial attention weight matrix Perform matrix multiplication between to obtain a three-dimensional tensor space representation ; The three-dimensional tensor space is represented as Input to Convolution kernel, generating weighted spatial features ; Preset learnable parameters With the weighted spatial features Multiply and perform element-by-element summation on the three-dimensional tensor to obtain the spatial features .
3. The method for classifying motor imagery EEG signals based on a 3D convolutional neural network with multi-scale features according to claim 2, characterized in that: The step S4 comprises: The spatial features Divided into three time slices; Input two of the time slices into two 2D convolution blocks respectively to obtain the initial time features and ; The initial time feature and the remaining time slices are reconstructed and transposed to obtain the conversion time features and ; The conversion time characteristics First, matrix multiplication is performed, and then the time attention weight matrix is obtained through the Softmax function. The temporal attention weight matrix Used to represent the similarity weight between the time slices; The conversion time characteristics and the temporal attention weight matrix Perform matrix multiplication between to obtain weighted time features ; Preset learnable parameters With the weighted time feature Multiply and perform element-by-element summation with the time slice to obtain the scale The time characteristics of the following .
4. The method for classifying motor imagery EEG signals based on a 3D convolutional neural network with multi-scale features according to claim 1, characterized in that: The step S5 comprises: The temporal features are processed by an average pooling layer with a logarithmic nonlinear active function , reducing the time characteristics The time dimension of The low-level features in the ts are converted into high-level abstract features, and the high-level abstract features are connected into multi-scale spatiotemporal features. ; The multi-scale spatiotemporal features Input to the fully connected layer and perform classification prediction through the Softmax function.
5. The method for classifying motor imagery EEG signals based on a 3D convolutional neural network with multi-scale features according to claim 1, characterized in that: The 3D convolutional block is The convolution kernel structure is the same.
6. The method for classifying motor imagery EEG signals based on a 3D convolutional neural network with multi-scale features according to claim 1, characterized in that: The step S1 includes: the EEG signal acquisition device is an EEG helmet; the EEG helmet is calibrated; and the subject's scalp is cleaned with conductive paste or saline.
7. The method for classifying motor imagery EEG signals based on a 3D convolutional neural network with multi-scale features according to claim 6, characterized in that: During the calibration process, the contact impedance of the EEG helmet electrodes was tested using a standard resistance solution.
8. The method for classifying motor imagery EEG signals based on a 3D convolutional neural network with multi-scale features according to claim 1, characterized in that: The step S2 comprises: Performing bandpass filtering on the motor imagery EEG signal data; The motor imagery EEG signal data is normalized.
9. The method for classifying motor imagery EEG signals based on a 3D convolutional neural network with multi-scale features according to claim 8, characterized in that: The standardized motor imagery EEG signal data are divided, wherein 80% of the motor imagery EEG signal data are used for network training, and 20% of the motor imagery EEG signal data are used for testing.
Citation Information
Patent Citations
Electroencephalogram signal classification and identification method and system based on deep learning
CN116383696A
Fine motion motor imagery electroencephalogram signal classification method and device
CN117694907A
Cited By
Parameter adjusting method of motion intention recognition model and implantable neural rehabilitation equipment
CN121881163A