A method for identifying motor imagery electroencephalogram signals based on capsule networks
Through the 3D-CapsNet model combined with 3D convolution and capsule network, the problem of insufficient utilization of MI-EEG spatial information in the prior art is solved, the recognition accuracy is improved, individual differences is overcome, and more efficient feature extraction and classification is achieved.
Patent Information
- Application Number
- CN202211077974.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-09-05
- Publication Date
- 2025-07-29
- Estimated Expiration
- 2042-09-05
AI Technical Summary
The prior art is difficult to effectively retain and utilize the spatial information of MI-EEG in the recognition of motor imagination EEG signals, resulting in large individual differences, and deep learning methods are prone to overfitting when feature extraction and classification.
A three-dimensional capsule network (3D-CapsNet) model based on capsule network is adopted, combining 3D convolution and capsule network, feature extraction and classification are performed through dynamic routing connections, avoiding dimensionality reduction of pooled layers and retaining EEG detailed features.
It improves the recognition accuracy, overcomes individual differences, effectively extracts the temporal and spatial characteristics of MI-EEG, and reduces the risk of overfitting.
Smart Images

Figure CN115456016B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical fields of deep learning and brain-computer interface, and particularly relates to a method for identifying motor imagery electroencephalogram signals based on a capsule network. Background Art
[0002] For electroencephalogram (EEG) signal recognition, existing recognition technologies are mainly divided into two categories. One is to combine traditional manual feature extraction with machine learning algorithms for recognition; the other is to perform feature extraction and recognition based on a deep learning training model. The method of combining feature extraction with machine learning divides feature extraction and classification into two stages. The acquisition of the best features is subjective. If a suboptimal frequency band is selected during the feature extraction process, the classification performance will be affected. Due to the differences among different subjects, the method of selecting the best frequency band for each subject cannot be well applied to a larger population. The deep learning method embeds feature extraction and classification into an end-to-end network, minimizing the EEG signal preprocessing process, and is obviously more suitable for online BCI research.
[0003] However, in the above two major research methods, most only identify EEG signals represented in two-dimensional form and fail to fully reflect the spatial information contained in the EEG signals. MI-EEG is collected from the three-dimensional scalp surface and is a non-linear random time series with spatio-temporal information. Therefore, when processing EEG signals, it is more reasonable to consider both its temporal and spatial characteristics simultaneously. In addition, compared with the fields of image recognition and natural language processing, EEG signal recognition research needs to overcome the problems of small datasets and unclear EEG features. It has strict requirements for the network, which not only needs to fully extract the contained features but also avoid overfitting problems. At the same time, overcoming the differences among different subjects is also a problem that the network needs to solve.
[0004] A brain-computer interface (BCI) allows people to interact with the real world only through brain nerve activities. Motor imagery electroencephalogram (MI-EEG) signals are one of the widely used paradigms in BCI and are currently mainly applied to the field of motor rehabilitation. For MI-EEG-based rehabilitation training, on the one hand, by processing and transforming MI-EEG in a certain way, the movement intention is converted into instructions to control rehabilitation assistance devices such as wheelchairs and robotic arms, which to a certain extent solves the problem of communication between patients with muscle or nerve terminal damage and the environment; on the other hand, it promotes brain function remodeling to achieve functional compensation and ultimately restores some motor functions and improves the quality of life of patients.
[0005] MI-EEG recognition is the key to improving BCI performance. Based on the phenomena of event-related synchronization (ERS) and event-related desynchronization (ERD), a large number of MI-EEG classification methods have been proposed successively [1-6]. Among them, the method of combining feature extraction with machine learning has been successfully applied to MI-EEG classification. However, these methods divide feature extraction and classification into two stages, which makes the parameters of the feature extraction model and the classifier trained with different objective functions. In addition, the acquisition of the best features is subjective. If suboptimal frequency bands are selected during the feature extraction process, the classification performance will be affected. Most importantly, for a complex non-linear random time series, the method of manually determining the frequency band highly depends on experts' experience and their understanding of EEG. Due to the differences among different subjects, the method of selecting the best frequency band for each subject cannot be well applied to a larger population.
[0006] Recently, various deep learning methods have been applied to EEG classification, such as convolutional neural network (CNN) [7], recurrent neural network (RNN) [8], and capsule network (CapsNet) [9], etc. Using deep learning to directly recognize EEG signals does not require manual extraction of the features contained in the EEG signals. It embeds feature extraction and classification into an end-to-end network, and the parameters are jointly optimized, minimizing the EEG signal preprocessing process. Obviously, it is more suitable for online BCI research. When using deep learning for MI-EEG classification, the primary task is to represent MI-EEG in a form that can be processed by a deep model. In addition, EEG signal recognition research needs to overcome the problems of small datasets and unclear EEG features, which requires strict requirements for recognition methods. It is necessary to fully extract the contained features and avoid overfitting problems.
[0007] Currently, MI-EEG is usually represented in the form of a two-dimensional matrix, hereinafter referred to as 2DMI-EEG, that is, the number of sampling electrodes is used as the height and the sampling time step is used as the width. Another common method is to transform the EEG signal into a two-dimensional time-frequency image as the network input through methods such as short-time Fourier transform or wavelet transform. However, neither the representation method of the two-dimensional matrix nor the two-dimensional time-frequency image can retain the spatial information of MI-EEG, and the internal relationship between adjacent electrodes cannot be reflected in the two-dimensional matrix, which will affect the classification performance. In 2015, Bashivan et al.
[10] proposed a method that retains the original electroencephalogram spatial structure, spectral structure, and time structure. First, calculate the power spectrum of the EEG signal of each electrode, then calculate the sum of the absolute value squares of three selected frequency bands, and finally use the azimuth equal-area projection (AEP) method to map the electrode distribution map as the input image of the model. Based on this representation method, the recognition performance has been significantly improved, indicating that spatial features are extremely important for EEG-based classification tasks.
[0008] Existing technical solutions
[0009] In 2019, Zhao et al.
[11] proposed a 3D representation method for EEG signals, mapping the EEG time series into a three-dimensional array according to the electrode spatial distribution as the model input. This method can retain both time features and spatial features. And a multi-branch 3D convolutional neural network (3D CNN) was proposed to classify 3DMI-EEG. 3D CNN is composed of three branches with different receptive field sizes to extract MI-related features. The three branches are respectively named the small receptive field network (SRF), the medium receptive field network (MRF), and the large receptive field network (LRF). Finally, a fully connected layer combined with Softmax is used for classification. This is a relatively successful attempt in the classification of original EEG data. After that, Liu et al.
[12] conducted further research on this basis. Still using a three-branch structure, a dense connection method was introduced to improve the multi-branch 3D convolutional neural network for classifying 3DMI-EEG, deepening the network and overcoming overfitting to a certain extent, and the performance has been improved to a certain extent.
[0010] For the EEG signals in three-dimensional representation form, both
[11] and
[12] adopt the form of convolutional neural network. In order to retain more features of the EEG signal, no pooling layer is used for dimensionality reduction in the middle, and a multi-branch structure is adopted at the same time. Therefore, the number of network parameters is relatively large. In addition, although both time and space features are taken into account, the internal relationship between features cannot be expressed by the network, which affects the recognition performance to a certain extent.
[0011] References
[0012] [1] BOSTANOV V. BCI competition 2003 - data sets Ib and IIb: feature extraction from event - related brain potentials with the continuous wavelet transform and the t - value scalogram[J]. IEEE Transactions on Biomedical engineering, 2004, 51(6): 1057 - 1061.
[0013] [2] HSU W Y, SUN Y N. EEG - based motor imagery analysis using weighted wavelet transform features[J]. Journal of neuroscience methods, 2009, 176(2): 310 - 318.
[0014] [3] BURKE D P, KELLY S P, DE CHAZAL P, et al. A parametric feature extraction and classification strategy for brain - computer interfacing[J]. IEEE Transactions on Neural Systems and Rehabilitation Engineering, 2005, 13(1): 12 - 17.
[0015] [4] RAMOSER H, MULLER - GERKING J, PFURTSCHELLER G. Optimal spatial filtering of single trial EEG during imagined hand movement[J]. IEEE transactions on rehabilitation engineering, 2000, 8(4): 441 - 446.
[0016] [5]ANG KK,CHIN ZY,WANG C,et al.Filter bank commonspatial pattern algorithm on BCI competition IV datasets 2a and2b[J].Frontiers inneuroscience,2012,6:39.
[0017] [6]NOVI Q, GUAN C, DAT TH, et al. Sub-band common spatial pattern (SBCSP) for brain-computer inter-face [C] / / 20073rdInternational IEEE / EMBS Conference on Neural Engineering. IEEE, 2007: 204-207.
[0018] [7] LI MA, HAN JF, DUAN L JA novel MI-EEG imaging with the location information of electrodes[J]. IEEE Access, 2019,8:3197-3211.
[0019] [8]ABBASVANDI Z,NASRABADI A MA self-organized recurrent neural network for estimating the effective con-nectivity and its application to EEGdata[J].Computers in biology and medicine,2019,110:93-107.
[0020] [9] Chen Bo, Chen Lanlan, Jiang Runqiang. EEG emotion recognition based on integrated capsule network[J]. Computer Engineering and Applications, 2022, 58(08): 175-184.
[0021] CHEN QIN,CHEN LANLAN,JIANG RUNQIANG. Emo-tion recognition of EEG based Ensemble CapsNet[J].CEA,2022,58(08):175-184.
[0022]
[10] BASHIVAN P, RISHI I, YEASIN M, et al. Learning repre-sentations from EEG with deep recurrent-convolutional neural networks[J]. arXiv preprint arXiv:1511.06448, 2015.
[0023]
[11] ZHAO X, ZHANG H, ZHU G, et al. A multi-branch 3D convolutional neural network for EEG-based motor im-agery classification[J]. IEEE transactions on neural systems and rehabilitation engineering, 2019, 27(10): 2164-2177.
[0024]
[12] LIU T, YANG D. A Densely Connected Multi-Branch 3D Convolutional Neural Network for Motor Imagery EEG Decoding[J]. Brain Sciences, 2021, 11(2): 197. Summary of the Invention
[0025] To solve the above problems, the present invention proposes a method for identifying motor imagery EEG signals based on a capsule network, comprising the following steps:
[0026] S1. Map the motor imagery EEG signal time series into a three-dimensional array form according to the electrode spatial distribution;
[0027] S2. Combine a capsule network with 3D convolution to construct a three-dimensional capsule network 3D-CapsNet motor imagery EEG signal recognition model. Use the three-dimensional EEG signal described in S1 as the input of the recognition model. 3D-CapsNet consists of a 3D convolution module and a capsule network module. Among them, the 3D convolution module uses multiple layers of 3D convolution to extract features from the input EEG signal simultaneously in the time dimension and the inter-channel spatial dimension to obtain low-level features. The capsule network module has spatial detection ability. The low-level features output by the 3D convolution module are integrated by the capsule network to obtain high-level spatial vectors containing the relationships between features;
[0028] S3. The capsule network module is trained using a dynamic routing algorithm. The primary capsules and the motor capsules are connected by dynamic routing, and finally, the classification result is output through the non-linear activation function squash.
[0029] The beneficial effects of the present invention are as follows: By combining 3D convolution, a 3D-CapsNet motor imagery electroencephalogram (MI-EEG) signal recognition model is proposed, which improves the recognition accuracy and overcomes individual differences to a certain extent. 3D-CapsNet comprehensively considers the time dimension, channel space dimension, and the internal relationship between features of MI-EEG, maximizing the feature expression ability of the network. At the same time, the dynamic routing connection method between capsules in the capsule network replaces the traditional fully connected layer, avoiding the need for the network to use a pooling layer to reduce the feature dimension, so many EEG detail features are retained, thus ensuring effective feature extraction. BRIEF DESCRIPTION OF THE DRAWINGS
[0030] Figure 1 It is a schematic diagram of the 3D representation process of MI-EEG of the present invention; Left: International 10-20 system electrode montage; Middle: TP two-dimensional matrices; Right: Motor imagery electroencephalogram signal in 3D representation;
[0031] Figure 2 It is a structural diagram of the three-dimensional capsule network (3D-CapsNet) of the present invention;
[0032] Figure 3 It is a schematic diagram of different network structure diagrams involved in the experimental process of the present invention; (a) Three-dimensional convolutional capsule network 3D-CapsNet; (b) Three-dimensional convolutional network 3D-CNN; (c) Two-dimensional convolutional capsule network 2D-CapsNet;
[0033] Figure 4 It is the information transfer and routing process between capsules of the present invention. DETAILED DESCRIPTION OF THE INVENTION
[0034] Inspired by the dynamic routing connection method of the capsule network, a 3D-CapsNet motor imagery electroencephalogram signal recognition model is proposed by combining 3D convolution. The 3D convolution module uses multiple layers of 3D convolution to extract features from both the time dimension and the channel space dimension simultaneously, obtaining low-level features. The capsule network also has a certain spatial detection ability. The low-level features output by the 3D convolution module are integrated by the capsule network to obtain high-level spatial vectors containing the relationship between features, and finally, the classification result is output through the non-linear activation function squash. 3D-CapsNet comprehensively considers the time and spatial features of the original electroencephalogram signal, adopts the dynamic routing connection method, abandons the pooling layer to retain fine features, and maximizes the feature expression ability of the network. A method for recognizing motor imagery electroencephalogram signals based on a capsule network is proposed, and the specific implementation is as follows:
[0035] Example 1
[0036] A method for identifying motor imagery electroencephalogram signals based on a capsule network, comprising the following steps:
[0037] S1. Map the electroencephalogram (EEG) time series into a three-dimensional array form according to the electrode spatial distribution;
[0038] S2. Combine the capsule network with 3D convolution to construct a three-dimensional capsule network (3D-CapsNet) motor imagery EEG signal recognition model, and use the three-dimensional EEG signal described in S1 as the input of the recognition model. 3D-CapsNet consists of a 3D convolution module and a capsule network module. Among them, the 3D convolution module uses multiple layers of 3D convolution to extract features from the input EEG signal simultaneously from the time dimension and the inter-channel spatial dimension to obtain low-level features; the capsule network module has spatial detection capabilities, and the low-level features output by the 3D convolution module are integrated through the capsule network to obtain high-level spatial vectors containing the relationships between features;
[0039] S3. The capsule network module is trained using a dynamic routing algorithm, and the primary capsules and the motor capsules are connected by dynamic routing, and finally the classification result is output through the non-linear activation function squash.
[0040] Among them, the step S1 is executed as follows:
[0041] First, intercept the EEG signal by frame and obtain the numerical value of the current frame, and transform the numerical value of each frame into a two-dimensional matrix 2D-map of x×y according to the general spatial distribution of the sampling electrodes, and fill the unused electrode positions with 0 at the same time;
[0042] Then, use the time information of the EEG signal to expand the TP 2D-maps into a three-dimensional matrix of x×y×TP, where TP is the number of sampling points of each channel, and TP is a natural number.
[0043] Among them, the step S2 is executed as follows: The 3D convolution module consists of 5 3D convolution layers. The convolution layers are encapsulated into a convolution module, which can extract the basic features of data at multiple levels, provide local perception information for the primary capsule layer, and the number of convolution kernels gradually increases to ensure the correct extraction of increasingly rich features. Batch Normalization (BN) is implemented after each convolution to accelerate convergence and reduce overfitting. After passing through the convolution module, the input generates an output of 128 * 4 * 5 * 6, which is converted into a tensor of 128 * 4 * 5 * 6 and sent to the primary capsule layer. The primary capsule layer outputs 384 4D capsules. The primary capsules store different forms of spatial features of (Motorimagery EEG, MI-EEG). A dynamic routing connection is performed between the primary capsule layer and the motor capsule layer. The dynamic routing algorithm clusters capsules with similar predictions together, abstracts the motor capsules that can represent the differences between classes, and finally outputs the classification result through the non-linear activation function squash.
[0044] Among them, the step S3 is specifically as follows: The capsule network uses a dynamic routing algorithm for training. The information transfer and routing process between capsules only occurs between two consecutive capsule layers, that is, and s j adopt the dynamic routing algorithm; the specific process:
[0045]
[0046] First, u i (i = 1, 2,..., n) represents the detected low-level feature vector. Multiply the low-level feature vector u i with the corresponding weight matrix W ij to obtain the high-level output vector i represents the i-th low-level feature, and j represents the j-th primary capsule; as shown in Equation (1), the vector length encodes the probability of the corresponding feature, and the vector direction encodes the internal state of the feature; Also known as the primary capsule, the above steps encode the spatial relationship between the low-level feature and the high-level feature;
[0047] Secondly, weight the primary capsule , and the capsule uses the dynamic routing algorithm to learn the coupled sparse weight c ij . By adjusting c ij , the primary capsule sends the output to the appropriate motor capsule s j , s j is the result of weighted summation of the prediction vectors of multiple primary capsules. The prediction values that are close to each other will be clustered. The whole process is shown in Equation (2):
[0048]
[0049] Finally, s j After passing through the non-linear activation function squash, without changing the direction of the vector, the length is compressed to within 0 to 1, and the result is represented by the vector v j As shown in Equation (3), the vector length encodes the probability of the corresponding feature, and the vector direction encodes the internal state of the feature;
[0050]
[0051] The above three steps are the complete propagation process between capsules. Among them, the learning of the coupling coefficient c ij is the essence of the dynamic routing algorithm and is determined by Equation (4):
[0052]
[0053] In the formula, b ij is a temporary variable with an initial value of 0. After the first iteration, all coupling coefficients c ij are equal; as the iteration progresses, the value of b ij is updated, and the uniform distribution of c ij will change; the update formula of b ij is as shown in (5):
[0054]
[0055] 3D Representation of MI-EEG
[0056] Figure 1 shows the 3D MI-EEG mapping process of the EEG signals of a certain subject in the BCI Competition IV Dataset 2a. First, the EEG signals are intercepted frame by frame and the values of the current frame are obtained. According to the general spatial distribution of the sampling electrodes, each frame of values is transformed into a two-dimensional matrix 2D-map, and at the same time, the electrode positions not used are filled with 0; then, using the time information of the EEG signals, TP 2D-maps are extended into a three-dimensional matrix of x×y×TP, where TP is the number of sampling points for each channel. This way of representing MI-EEG in three dimensions according to the electrode distribution not only completely retains the time information existing in the EEG time series but also retains the spatial information existing in the electrode distribution while ensuring the processability of the EEG data.
[0057] 3D-CapsNet Hierarchical Structure
[0058] 3D-CapsNet is mainly composed of two parts: a 3D convolutional module and a CapsNet module, and its framework is as Figure 2As shown, first, multi-layer 3D convolution is used to extract features in the time dimension and the spatial dimension between channels, and primary features are abstracted; then, a convolutional capsule layer is used to detect the internal relationships between features; finally, a high-dimensional feature vector, called a motion capsule, is obtained through dynamic routing connection, and the squash function is combined to classify the motion capsule.
[0059] The specific parameters of the 3D-CapsNet hierarchical structure are as Figure 3 (a) shown, and the parameter settings are the optimal values obtained through continuous experimental attempts. The 3D convolution module can extract the basic features of data at multiple levels, providing local perception information for the primary capsule layer. The number of convolutional kernels gradually increases to ensure the correct extraction of increasingly rich features. Batch normalization (BN) is implemented after each convolution to accelerate convergence and reduce overfitting. After passing through the convolution module, the input generates 128 outputs of 4*5*6, which are converted into a tensor of 128*4*5*6 and sent to the primary capsule layer. The primary capsule layer outputs 384 4D capsules. The primary capsules store different forms of spatial features of MI-EEG. Dynamic routing connection is performed between the primary capsule layer and the motion capsule layer. The dynamic routing algorithm clusters capsules with similar predictions together, abstracting motion capsules that can represent the differences between classes, and finally outputs the classification result through the non-linear activation function squash.
[0060] Experimental environment
[0061] The 3D-CapsNet model is implemented in Python under the Pytorch framework. The experimental environment is the 11th Gen Intel(R) Core(TM) i5-11400H @ 2.70GHz 2.69GHz, 16GB of memory, NVIDIA GeForce RTX3050 graphics card, and 64-bit Windows 11 system.
[0062] Training algorithm and training strategy
[0063] Training algorithm: The capsule network uses the dynamic routing algorithm for training. As Figure 4 shown is the information transfer and routing process between capsules. The dynamic routing algorithm is only used between two consecutive capsule layers ( and s j ). First, u i (i = 1, 2, …, n) represents the detected low-level feature vector. The low-level feature vector u i is multiplied by the corresponding weight matrix W ij to obtain the high-level output vector As shown in Equation (1), the vector length encodes the probability of the corresponding feature, and the vector direction encodes the internal state of the feature. Also known as the primary capsule, this step encodes the spatial relationship and other important relationships between low-level features and high-level features.
[0064]
[0065] Where i represents the i-th low-level feature and j represents the j-th primary capsule.
[0066] Secondly, the primary capsules are weighted. This step is similar to scalar weighting in neurons, with the difference that neuron weights are learned through the backpropagation algorithm, while capsules use the dynamic routing algorithm to learn the coupled sparse weights c ij , by adjusting c ij , the primary capsules will send their outputs to the appropriate motor capsules s j , s j is the result of weighted summation of the prediction vectors of multiple primary capsules. Close prediction values will cluster, and the whole process is shown in Equation (2):
[0067]
[0068] Finally, s j passes through the non-linear activation function squash. Without changing the vector direction, the length is compressed to within 0 to 1, and the result is represented by the vector v j . As shown in Equation (3), the vector length encodes the probability of the corresponding feature, and the vector direction encodes the internal state of the feature.
[0069]
[0070] The above three steps are the complete propagation process between capsules. Among them, the learning of the coupling coefficient c ij is the essence of the dynamic routing algorithm and is determined by Equation (4):
[0071]
[0072] In the formula, b ij is a temporary variable with an initial value of 0. After the first iteration, all coupling coefficients c ij will be equal. As the iteration progresses, the value of b ij is updated, and the uniform distribution of c ij will change. The update formula for b ij is shown in (5):
[0073]
[0074] The capsule loss evaluation is performed using the Margin Loss function, denoted by L k . For each category k, there is Lk :
[0075] L k = T k max(0, m + - ||v k ||) 2 + λ(1 - T k )max(0, ||v k || - m - ) 2 (6)
[0076] When and only when there is a motor imagery of class k, T k = 1, m + = 0.9 and m - = 0.1, λ takes the empirical value of 0.5 to reduce the loss of some non - occurring classes, and the total loss is the sum of the losses of all motion capsules.
[0077] Training strategy: The present invention adopts a cropping training strategy. In cropping training, samples are generated by sliding a 3D window along the time dimension with a certain data step size. The window covers all electrodes, and the size of the window in the time dimension is set related to the EEG data sampling frequency and a specific task. The EEG signal cropping training strategy is a common method to enhance EEG signal training samples, similar to the cropping strategy in the field of image recognition. Multiple existing experiments have shown that compared with the training of complete samples, the training of cropped samples has better classification performance. The process conducts capsule network training by optimizing the margin loss function, and the number of training iterations is set to 80. The Adam stochastic optimization algorithm is used to dynamically adjust the learning rate, which can replace the classical Stochastic Gradient Descent (SGD) process to more effectively update the network weights and accelerate the convergence of the neural network.
[0078] Verification process
[0079] The experimental verification stage progresses layer by layer. First, the effectiveness of applying the capsule network to EEG signal recognition is verified in 3D - CapsNet. 3D - CapsNet is as shown in Figure 3 (a). Monitoring is carried out on the test data of 9 subjects for 80 training cycles. 3D - CapsNet performs excellently on the data of subjects 1, 3, 7, 8, and 9, reaching a relatively high accuracy after 40 iterations and generally tending to be stable. It also shows good performance on the data of subjects 4 and 6. For subjects 2 and 5, although the model performance is not as ideal as on the data of other subjects, it can still reach an accuracy of about 70%. Considering the overall performance of 3D - CapsNet on the data of all subjects, the model does not show a large deviation due to the change of subjects and has a certain robustness.
[0080] Secondly Figure 3 (b) further verifies the excellent performance of the capsule network on the 3D-CNN structure shown. In the experiment, the number of iterations when the loss value tends to be stable is observed as the number of iterations in the subsequent experiment, and the number of iterations is set to 80, and it is implemented using Pytorch. Except for subject 5, the remaining subjects all showed good performance under the dynamic routing connection; and a lower standard deviation was shown under the dynamic routing connection. It can be concluded from this that the capsule network is superior to the traditional convolutional neural network in MI-EEG recognition.
[0081] Finally, the performance comparison of the network improved based on the capsule network on 2D MI-EEG and 3D MI-EEG was verified. The 3D convolutions in 3D-CapsNet were all changed to use 2D convolutions, and the 2D-CapsNet framework structure is as Figure 3 (c) shown. An identification experiment was carried out on the 2D MI-EEG input and compared with the method of the present invention. For 9 subjects, the recognition accuracy using the 3D MI-EEG representation was higher than that using the 2D MI-EEG representation; from the perspective of the standard deviation, the standard deviation under the 3D MI-EEG representation was lower than that under the 2D MI-EEG representation, that is, the motor imagery EEG signals represented in 3D are more suitable for the case of decoding using a deep network.
[0082] 3D-CapsNet comprehensively considers the time dimension, channel space dimension and the internal relationship between features of MI-EEG, and maximizes the feature expression ability of the network; at the same time, the dynamic routing connection method between capsules in the capsule network replaces the traditional fully connected layer, avoiding the need for the network to use a pooling layer to reduce the feature dimension, so many EEG detail features are retained, thus ensuring effective feature extraction.
[0083] Experimental results
[0084] In an experimental way, the method in this paper was compared with similar studies. Among them, DeepNet, EEGNet, and ShallowNet are all based on 2D-EEG for electroencephalogram (EEG) signal decoding, and References
[11] and
[12] are based on 3DMI-EEG for EEG signal decoding. The classification accuracy results of different methods on the evaluation dataset of 9 subjects are shown in Table 1. It can be seen that References
[12] and this paper have certain advantages in recognition accuracy. The decoding accuracy of Reference
[11] is lower than that of EEGNet and ShallowNet, but the standard deviation of the accuracy is much smaller than the two. Generally speaking, the standard deviation in the 3D representation form is much smaller than that in the 2D representation. It can be inferred that the 3D representation form of the motor imagery EEG signal is more suitable for EEG signal decoding and can improve the recognition accuracy to a certain extent. Such a representation form is more conducive to retaining the common features of MI-EEG among different subjects, can overcome the individual differences to a certain extent, and has stronger interpretability. On the other hand, the method proposed in this paper shows better overall recognition accuracy than the leading literature. The accuracy on the dataset of 6 subjects is the highest among similar research methods, and the average accuracy is 2.805% higher than the second-best result.
[0085] Table 1 Comparison of classification accuracies of similar studies
[0086]
[0087] To further verify the performance of 3D-CapsNet, the Kappa value of the classification results was calculated and compared with References
[11] and
[12] . The results are shown in the table. The Kappa value is mainly used for consistency testing to measure the consistency between the model prediction results and the actual classification results. The value range of Kappa is -1.0 - 1.0, and the larger the value, the better the classification performance of the algorithm. The expression of the Kappa value is as follows:
[0088]
[0089] where P o is the overall classification accuracy of the total samples, and P e is used to evaluate the chance probability. Assuming that c is the total number of categories, T i (i = 1, 2,..., c) are the number of samples correctly classified in each category. The number of true samples in each category is a1, a2,... a c , and the number of samples predicted for each category is b1, b2,... b c , and the total number of samples is n. Then:
[0090]
[0091] The results are shown in Table 2. For all subjects except Subject 5, the Kappa values of the method proposed in this paper are better than those of the comparative literature. Therefore, 3D-CapsNet has good performance for 3DMI-EEG recognition.
[0092] Comparison of Kappa values in similar studies in Table 2
[0093]
[0094] As mentioned above, the above are only the preferred specific embodiments of the present invention, but the protection scope of the present invention is not limited thereto. Any person skilled in the art within the technical scope disclosed by the present invention, according to the technical solution and its concept of the present invention, making equivalent substitutions or changes should be covered within the protection scope of the present invention.
Claims
1. A method for identifying motor imagery electroencephalogram signals based on a capsule network, characterized in that, The steps include: S1. Map the EEG time series of motor imagery signals into a three-dimensional array according to the spatial distribution of electrodes; S2. The capsule network is combined with 3D convolution to construct a 3D capsule network 3D-CapsNet motor imagery EEG signal recognition model. The EEG signal in the three-dimensional array form described in S1 is used as the input of the recognition model. The 3D-CapsNet consists of a 3D convolution module and a capsule network module. The 3D convolution module uses multi-layer 3D convolution to extract features of the input EEG signal from both the time dimension and the inter-channel spatial dimension to obtain low-level features. The capsule network module has spatial detection capabilities. The low-level features output by the 3D convolution module are integrated through the capsule network to obtain a high-level spatial vector containing the relationship between features. S3 and the capsule network module are trained using a dynamic routing algorithm. The primary capsules and the motion capsules are connected using dynamic routing, and the classification results are finally output through a nonlinear activation function squash.
2. The method for recognizing motor imagery EEG signals based on capsule network according to claim 1, wherein The step S1 performs the following steps: First, the EEG signal is intercepted frame by frame and the value of the current frame is obtained. According to the general spatial distribution of the sampling electrodes, the value of each frame is transformed into an x×y two-dimensional matrix 2D-map, and the unused electrode positions are filled with 0. Then, the temporal information of the EEG signal is used to expand the TP 2D-map into a three-dimensional matrix of x×y×TP, where TP is the number of sampling points of each channel and TP is a natural number.
3. The method for identifying motor imagery EEG signals based on a capsule network according to claim 2, wherein The step S2 is performed as follows: the 3D convolution module is composed of 5 3D convolution layers, which are encapsulated into a convolution module, which can extract the basic features of the input EEG signal at multiple levels and provide local perception information to the main capsule layer. The number of convolution kernels gradually increases to ensure the correct extraction of more and more rich features; batch normalization BN is implemented after each convolution to accelerate convergence and reduce overfitting; the input generates 128 4*5*6 outputs after passing through the convolution module, which are converted into a 128*4*5*6 tensor and sent to the main capsule layer. The main capsule layer outputs 384 4-dimensional capsules. The main capsule stores spatial features of different forms of motor imagery EEG signals. Dynamic routing connections are performed between the main capsule layer and the motion capsule layer. The dynamic routing algorithm clusters capsules with similar predictions together, abstracts motion capsules that can represent inter-class differences, and finally outputs the classification results through the nonlinear activation function squash.
4. The method for identifying motor imagery EEG signals based on a capsule network according to claim 3, wherein The specific steps of step S3 are as follows: The capsule network uses a dynamic routing algorithm for training, and the information transfer and routing process between capsules only occurs between two consecutive capsule layers, that is, between and s j the dynamic routing algorithm is adopted; the specific process is as follows: First, u i (i = 1, 2, …, n ) represents the detected low-level feature vector. Multiply the low-level feature vector u i with the corresponding weight matrix W ij to obtain the high-level output vector i represents the i-th low-level feature, and j represents the j-th primary capsule; as shown in Equation (1), the vector length encodes the probability of the corresponding feature, and the vector direction encodes the internal state of the feature; also known as the primary capsule, the above steps encode the spatial relationship between the low-level feature and the high-level feature; Secondly, the primary capsules are weighted, and the capsules learn the coupled sparse weights c using the dynamic routing algorithm ij . By adjusting c ij , the primary capsules send the output to the appropriate motor capsules s j , where s j is the result of weighted summation of the prediction vectors of multiple primary capsules, and the prediction values that are close to each other will be aggregated. The whole process is shown in Equation (2): Finally, s j After passing through the non-linear activation function squash, without changing the direction of the vector, the length is compressed to within 0 to 1, and the result is represented by the vector v j As shown in Equation (3), the length of the vector encodes the probability of the corresponding feature, and the direction of the vector encodes the internal state of the feature; The above three steps are the complete propagation process between capsules. Among them, the learning of the coupling coefficient c ij is the essence of the dynamic routing algorithm and is determined by Equation (4): where b ij is a temporary variable with an initial value of 0. After the first iteration, all coupling coefficients c ij are equal; as the iteration progresses, the value of b ij is updated, and the uniform distribution of c ij will change; the update formula for b ij is shown in (5):