Motor imagery electroencephalogram detection method based on fusion of convolutional neural network and locust optimization algorithm

By fusing convolutional neural networks and group optimization algorithms in the motor imagination EEG signal classification algorithm, the shortcomings of existing algorithms in terms of classification accuracy and computational complexity are solved, and higher classification accuracy and lower computational complexity are achieved.

CN120180233APending Publication Date: 2025-06-20CHANGCHUN UNIV OF SCI & TECH
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510444586.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-10
Publication Date
2025-06-20

AI Technical Summary

Technical Problem

The existing classification algorithm for motor imagination EEG signal is insufficient in improving classification accuracy, especially due to the inter-test variability of EEG data and the computational complexity is high.

Method used

A model based on the fusion of convolutional neural networks and swarm optimization algorithms is designed, feature extraction is performed through multi-scale convolutional neural networks, and a locust optimization algorithm is combined to reduce the computational complexity and parameter requirements.

Benefits of technology

It improves the accuracy and performance of motion imagination feature classification, reduces the computational complexity, and allows the classifier to more accurately identify motion imagination information.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120180233A_ABST
    Figure CN120180233A_ABST
Patent Text Reader

Abstract

The invention discloses a motor imagery electroencephalogram detection method based on fusion of a convolutional neural network and a locust optimization algorithm, and belongs to the technical field of signal decoding. The method comprises the steps that an electroencephalogram plate collects motor imagery (MI) electroencephalogram signals, a convolutional neural network (CNN) is used for extracting time-frequency-space features, a ProbSparse self-attention mechanism is introduced into a convolutional layer, and the long sequence modeling complexity is reduced through a sparse attention matrix; globally searching and optimizing feature subsets and weights in combination with a locust optimization algorithm (GOA), training a model by using cross entropy loss and an Adam optimizer, and finally outputting a classification result and a visual feature contribution degree. The core innovation of the invention lies in the coupling of CNN and GOA, solves the problems of incomplete feature extraction and insufficient model generalization ability in the traditional method, and reduces the calculation complexity by using a ProbSparse mechanism. According to the method, the accuracy and robustness are remarkably improved in motor imagery classification, and the method is suitable for the fields of brain-computer interfaces, neural rehabilitation, intelligent control and the like.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of electroencephalogram signal feature classification, and relates to an optimization of an electroencephalogram detection algorithm for motor imagery based on a fusion model of a convolutional neural network and a swarm optimization algorithm. Background Art

[0002] A brain-computer interface (BCI) is a technology that directly connects the brain to an external device, aiming to control an external device or transmit information by reading and decoding neural signals. The development of intention decoding technology not only helps to improve the interaction ability between humans and mechanical devices, but also shows great potential in many fields such as medical treatment, rehabilitation, entertainment, and smart home. Electroencephalogram-based motor imagery is to use electroencephalogram (EEG) technology to detect and record the activities of the cerebral cortex, so as to identify the neural signals generated when an individual imagines a specific movement or action. This technology is usually used in the research and application of motor imagery and can explore the relationship between individual intention and brain activity. EEG based on motor imagery (MI) simulates various movement actions in the brain, such as imagining the movement of the hand or foot. To implement such a brain-computer interface system, accurate brain activity classification is crucial. Although previous studies have shown good performance, there is still room for improving the classification accuracy in constructing an efficient brain-computer interface application. The subject imagines limb movement in the MI experiment, and this experiment generates electroencephalogram (EEG) signals that can be collected by non-invasive methods. Due to the non-linearity, non-stationarity and low spatial resolution of EEG signals, the decoding of the collected EEG signals is complex and challenging. Existing MI classification studies can be divided into two categories: classical machine learning methods and deep learning-based methods. Traditional MI decoding methods first calculate predefined features, and then use established pattern classifiers, such as linear discriminant analysis (LDA) and support vector machine (SVM) to identify the user's intention. In particular, common spatial pattern (CSP) and filter bank common spatial pattern (FBCSP) are typical methods based on spatial filtering and are widely used in MI decoding. However, these classical methods are easily affected by the inter-trial variability of EEG data and highly dependent on handcrafted features; therefore, the obtained model may not be well generalized to unseen data. To improve the decoding accuracy, this paper proposes to combine the advantages of current swarm optimization algorithms and design an improved feature extraction algorithm. Summary of the Invention

[0003] To improve the accuracy and performance of MI decoding and combine the advantages and disadvantages of various existing algorithms, the present invention designs a model based on the fusion of a convolutional neural network and a swarm optimization algorithm. By designing a multi-scale convolutional neural network, using convolutional kernels of appropriate sizes to extract features from the input data, and fusing the swarm optimization algorithm and the feature classification algorithm, it is possible to achieve higher accuracy and lower complexity in motor imagery feature classification. The steps are as follows:

[0004] Use a self-developed EEG acquisition board (attached Figure 4 ) to collect the user's EEG signals. The EEG board converts the analog EEG signals into digital signals recognizable by a computer through an analog-to-digital conversion chip, and further transmits them to the host computer through serial communication for further data processing on the host computer.

[0005] The analog-to-digital conversion chip uses ADS1299 to receive the EEG signals processed by the analog front end, and sends the processed signals to the microcontroller STM32F103C8T6TR for signal processing and packaging, and then sends them to the host computer.

[0006] Further, on the host computer, feature extraction and classification of EEG signals are performed. A time average pool with a non-overlapping window size of W is used to extract robust time features. Calculate the power spectral density (PSD) of each channel, and use the fast Fourier transform (FFT) method to obtain frequency domain features. In the spatial domain, spatial filtering and time filtering techniques are used, combined with a convolutional neural network optimization algorithm for extraction, to extract discriminative and representative spatial features. Combine the frequency domain features and the spatial domain features to form a multi-dimensional feature vector for further feature classification and analysis tasks. In traditional CNNs, the convolutional layer is used to extract local features, and these features may lose some global information. Therefore, the present invention proposes to add a ProbSparse self-attention mechanism after the convolutional layer to capture these global dependencies. The ProbSparse self-attention mechanism optimizes long time series. This mechanism can sparsify the attention matrix, only calculate the most important attention scores, reduce the computational complexity, and make it more feasible to model long EEG signals. It helps to optimize the weights and structure of the network to improve the performance and generalization ability of the network. Combine the above fusion algorithm network with the locust optimization algorithm, and then use a fully connected layer to output the results.

[0007] Furthermore, when training the model, the cross-entropy loss function will be used to calculate the loss, and the Adam optimizer will be used to update the model parameters to improve the convergence speed and generalization ability of the model and enhance the performance of the deep learning model.

[0008] Further, analyze and compare the results. To improve the model, the model designed in this paper is compared with different models and different algorithms in existing research to improve the accuracy and training speed of the model.

[0009] Beneficial effects

[0010] Compared with the existing motor imagery EEG signal recognition technology, the present invention has the following beneficial effects:

[0011] The motor imagery classification based on EEG signals is completed through the multi-scale feature extraction method of EEG signals, the optimization of the model structure design, and the algorithm fusion design. To address the problem of incomplete EEG signal feature extraction, an algorithm fusion is proposed for improvement. Aiming at the problems of large parameter requirements and high computational complexity in the MI classification algorithm, the grasshopper optimization algorithm is fused to further reduce the computational complexity and reduce the parameter requirements. The spectral-spatiotemporal feature extraction based on CNN can study various interaction operators to explore the influence of the interaction between different frequency bands on motor imagery. An improvement is made on the basis of the traditional convolutional neural network, that is, a ProbSparse self-attention mechanism is added after the convolutional layer to optimize the feature extraction step. According to the fusion with the grasshopper optimization algorithm, parameter optimization is carried out to enhance the feature quality, thereby improving the performance of the classifier and enabling the classifier to more accurately identify MI information. Description of the drawings

[0012] Figure 1 is the algorithm structure flowchart of the present invention.

[0013] Figure 2 is the algorithm structure diagram of the CNN spectral-spatiotemporal feature extraction of the present invention.

[0014] Figure 3 is the flowchart of the grasshopper optimization algorithm of the present invention.

[0015] Figure 4 is the PCB diagram of the EEG acquisition board of the present invention. Detailed implementation manners

[0016] In order to make the purpose, technical solutions and advantages of the present invention clearer, the embodiments of the present invention will be further described in detail below with reference to the accompanying drawings. It should be understood that the specific implementations described herein are only used to explain the present invention and are not used to limit the present invention.

[0017] The EEG data set in the embodiment of the present invention is derived from 15 healthy volunteers (aged 24-35 years old), and the time domain, frequency domain, and spatial domain intention EEG features of each channel are extracted. Among them, 5 subjects are initial BCI users, and the others have experience in BCI experiments. All subjects have no history of neurological, psychiatric, and other related diseases. The subjects sat on a 60 (± 5) cm chair facing a 21-inch LCD monitor (refresh rate: 60Hz; resolution: 1600×1200). The approximate horizontal and vertical viewing angles are 37.7 degrees and 28.1 degrees, respectively. In this project, only the MI paradigm is involved, so only the process of the MI paradigm is described: for each subject, a black fixation cross appears in the center of the monitor for the first 3 seconds of each experiment, and then the left and right arrow visual prompts appear. The subject performs an imagination task of grasping with the appropriate hand for 4 seconds. After each task, the screen is blank for 6 seconds. The experiment is divided into two stages: train and test (note that this is not a training set and a test set); there are 100 experiments in each stage, and the image tasks of the left and right hands are balanced at the same time. After acquiring the data set, the EEG signals are preprocessed by downsampling, Butterworth bandpass filtering, etc. to eliminate the noise and artifacts generated during the data acquisition process and remove electrooculographic interference.

[0018] Figure 1 It is the algorithm structure flow chart of the present invention. The method comprises the following steps:

[0019] MI EEG data is obtained through EEG board sampling experiments. The host computer pre-processes the original EEG signals, including downsampling, filtering, baseline correction, and artifact removal, in order to improve data quality. Then, the signal is subjected to feature extraction and classification, and model training and optimization, and finally the relevant MI EEG information is output.

[0020] Furthermore, since EEG signals contain a lot of noise, it will make emotion recognition difficult and reduce the recognition accuracy, so the EEG signals must be preprocessed after data collection. The EEG signal preprocessing process usually includes the following contents:

[0021] Filtering: Since EEG signals are easily interfered by noise, regular noise is filtered, such as the power frequency interference around 50 Hz, and the higher frequency parts that do not belong to EEG signals are removed. In addition, EEG signals of different rhythms can be extracted for analysis according to needs. Downsampling: Usually, the sampling frequency of the collected EEG signals is as high as 1000 Hz. To reduce the amount of data and improve the calculation efficiency, the dataset can be downsampled. Segmentation and baseline correction: When collecting EEG signals, the entire process of the experiment is generally recorded. However, for EEG emotion recognition, only the part where the subject receives the stimulus or the part of a specific experiment is needed. Therefore, the EEG signals need to be segmented according to the markers in the experimental process. Baseline correction can highlight the changes in EEG signals before and after the stimulus to analyze the impact of the stimulus on the subject. At the same time, baseline correction can also prevent the impact caused by drift. Artifact removal: EEG signals are easily interfered by the subject's own physiological signals, and these interference filtering methods cannot remove them. Therefore, methods such as interpolating bad electrodes, manually removing bad segments, and independent component analysis are needed to remove these artifact interferences.

[0022] Appendix Figure 2 The following is a structural diagram of a spectrum-spatiotemporal feature extraction algorithm based on CNN provided in the implementation example of the present invention. The method includes the following content:

[0023] The convolutional neural network (CNN) mainly includes the following structures: input layer, convolutional layer, pooling layer, fully connected layer, and output layer. It is generally composed of alternately connected convolutional layers and pooling layers, plus the final fully connected layer for discrimination. First, a single preprocessed original EEG sample can be represented as a 2D map X ∈ R C×T, where C represents the number of EEG channels and T represents the time points. With the goal of explicit frequency band interaction, the EEG signal is divided into two feature frequency bands. Brain oscillations are usually classified into specific frequency bands (delta: <4 Hz, theta: 4 - 7 Hz, alpha: 8 - 12 Hz, beta: 12 - 30 Hz, gamma: >30 Hz). The EEG signal is first filtered into a low-frequency band (4 - 16 Hz) and a high-frequency band (16 - 40 Hz). The selection of these two frequency bands covers the mu rhythm and beta rhythm that are most relevant to the MI signal. We use X L ∈ R C×T to represent the EEG data filtered in the low-frequency band and X H ∈ R C×T to represent the EEG data filtered in the high-frequency band. We adopt a method combining spatial filtering and temporal filtering to learn spectrum-spatial resolution patterns. X is regarded as a one-dimensional image with multiple channels along the time dimension. The spectrum-spatial features are generated as follows:

[0024]

[0025] where F s and F tThey are one-dimensional spatial convolution and one-dimensional depth-time convolution respectively. In the sequel, U l ∈R f×T and U h ∈R f×T are the outputs representing the spectral-spatial features of each band, where F is the number of filters for each band. The operator for interaction between the same frequency bands:

[0026] ①Summation.I(U l ,U h ) = U l +U h

[0027] ②Concatenation.I(U l ,U h ) = [U l ,U h

[0028] ③Hadamard product.I(U l ,U h ) = U l ⊙U h

[0029] ④Linear projection I(U1,U h ) = W l U1 + W h U h

[0030] For the problem that correlations cannot be established for multiple related inputs, it is solved by the self-attention mechanism.

[0031] Furthermore, I(U l , U h ) represents the interaction function between multi-band features. The best performance for MI decoding from EEG signals can be obtained using element-wise addition operations. The spectral-spatial features generated after the convolutional layer are still high-dimensional temporal representations. For the generated long time series, ProbSparse self-attention processing is performed. The specific steps include:

[0032] S1 Key point selection: Calculate the dot product between each query vector and all key vectors to generate an attention score matrix. Instead of considering all points, a probability distribution is used to select a part of the key points to calculate the attention scores, which greatly reduces the computational amount.

[0033] S2 Probability distribution: Use a Gaussian distribution or other suitable probability distributions to select the key points for calculating the attention scores, and these points are considered to contribute the most to the final result of the model.

[0034] ​S3 Sparse Attention Matrix: Only the scores of these selected key points are retained in the attention matrix, and the scores at other positions are set to zero. This sparsification operation reduces the computational complexity from O(N 2 ) to O(N log N) or lower.

[0035] Figure 3 is the flow chart of the locust optimization algorithm of the present invention. The method includes the following contents:

[0036] Feature classification is performed in the fully connected layer of the CNN. Feature classification algorithm fusion usually involves the combination of multiple feature selection or feature extraction methods and classifiers, aiming to improve classification accuracy and robustness. The locust optimization algorithm can be used as a meta-heuristic algorithm to optimize the process of feature selection or feature combination. The locust optimization algorithm (Grasshopper Optimization Algorithm, GOA) is a new swarm intelligence optimization algorithm proposed by Saremi et al. based on the simulation of locust foraging behavior. According to the value of the objective function, the global optimal solution and the individual optimal solution are updated. When locusts gather for foraging, their positions change due to the interaction between locusts, the gravity of the locusts themselves, and the influence of the wind. Its mathematical model is expressed as: X i = S i + G i + A i , where X i represents the position of the i-th locust, S i represents social interaction, G i represents the gravity of the i-th locust, and A i represents wind advection.

[0037]

[0038] where is the distance between the i-th individual and the j-th individual in the d-th dimension, is the unit vector from the i-th individual to the j-th individual. s(r) represents the intensity of the social influence of the population, f is the attraction intensity, and l is the attraction range (in the simulation experiment of GOA, f = 0.5 and l = 1.5); g is the gravitational constant, and e g represents the unit vector pointing to the center of the earth. u is the constant drift, and e w is the unit vector in the wind direction. ub d and lb d represent the upper and lower bounds of the d-th dimension respectively. The behavior of locusts gathering for foraging can be simulated using Equation 6. However, this data model cannot be directly used to solve mathematical problems because locusts quickly reach the comfort zone and the locust swarm does not converge to the specified point. Therefore, the mathematical model is modified as follows:

[0039]

[0040] where \(c_{max}\) and \(c_{min}\) represent the maximum and minimum values of the parameter \(c\) respectively, \(l\) is the current iteration number, \(L\) is the maximum iteration number, and \(T\) d is the optimal position of the locust. It can be seen from Equation (8) that the update of the locust individual's position in the locust optimization algorithm is determined by the current position of the locust individual, the optimal position, and the positions of all other locust individuals. The steps and pseudocode of the basic locust optimization algorithm are as follows: Input: locust population size \(N\), search space dimension \(dim\), maximum iteration number \(L\), parameter \(c\); Step 1: Initialize the population \(X\) i , \(i = 1,\ldots,N\); Step 2: Evaluate the fitness value of each individual and store the optimal position in \(T\) d ; Step 3: Update the parameter \(c\) according to Equation (7); Step 4: Update the position of the locust individual according to Equation (8); Step 5: Evaluate the fitness value of each individual and update \(T\) preferentially d ; Step 6: Judge whether the termination condition is satisfied. If so, go to Step 7; otherwise, repeat Steps 3 - 5; Step 7: Stop the algorithm and output the optimal position \(T\) d and the optimal value.

[0041] The locust optimization algorithm can explore the feature subset or feature weight space through global search, which helps to find a better feature combination scheme. It is applicable to different types of feature selection and feature weight optimization problems and can be used in combination with various classifiers, such as support vector machines (SVM), neural networks, etc. The algorithm parameters and evaluation metrics can be adjusted according to specific problems to obtain the best performance. Generally speaking, as a meta - heuristic algorithm, the locust optimization algorithm has good application potential for feature selection and feature weight optimization in the feature classification step, which can improve the performance and generalization ability of the classifier.

[0042] Furthermore, the training model will calculate the loss using the cross - entropy loss function and update the model parameters using the Adam optimizer. The cross - entropy loss function is:

[0043]

[0044] Furthermore, to ensure that the classification performance is driven by task - specific features rather than noise or artifacts in the data. In the electroencephalogram decoding network, it can tell us which sampling points of the electroencephalogram signal are responsible for picking a certain label (task). Finally, the normalized spectral - spatial attributes are mapped to the corresponding electrode positions, thus generating an attribute pattern associated with the brain regions and frequency bands of each MI task. Judge whether the features are related to the well - known brain activation patterns related to MI.

[0045] In the MI decoding experiment, it is necessary to evaluate the performance of different methods and models. Accuracy is one of the most commonly used evaluation metrics, representing the ratio of the number of samples correctly identified by the model to the total number of samples. The average accuracy of identification is positively correlated with the model performance, that is, the higher the average accuracy, the better the model performance. The formula for calculating the average accuracy in a binary classification task is:

[0046]

[0047] Among them, TP represents the number of positive samples correctly identified as positive samples, TN represents the number of negative samples correctly identified as negative samples, FP represents the number of negative samples misjudged as positive samples, and FN represents the number of positive samples misjudged as negative samples. In addition to accuracy, the confusion matrix is also a commonly used evaluation metric for supervised learning tasks. The confusion matrix statistically analyzes the recognition results of each class of samples and then intuitively displays the number of each classification situation in the form of a matrix.

[0048] Finally, it should be noted that the above-described embodiments are only specific implementation manners of the present application, used to illustrate the technical solutions of the present application, rather than limiting it. The protection scope of the present application is not limited thereto. Although the present application has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that: any person skilled in the art within the technical scope disclosed by the present application can still modify the technical solutions described in the foregoing embodiments or easily conceive of changes, or perform equivalent replacements for some of the technical features; and these modifications, changes or replacements do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present application, and should all be covered by the protection scope of the present application. Therefore, the protection scope of the present application shall be subject to the protection scope of the claims.

Claims

1. A motor imagery EEG detection method based on the fusion of convolutional neural network and locust optimization algorithm, characterized in that: The following steps are involved: The EEG signal acquisition and preprocessing: use a self-developed EEG acquisition board to collect the user's motor imagery (MI) EEG signal, the EEG acquisition board includes an analog-to-digital conversion chip ADS1299 and a microcontroller STM32F103C8T6TR, and transmits the digital signal to the host computer through serial port communication; the original EEG signal is preprocessed, including downsampling, Butterworth bandpass filtering (0.5-45Hz), baseline correction, segmentation, and independent component analysis (ICA) to remove eye and muscle artifacts. The multi-scale spectrum-space-time feature extraction: the preprocessed EEG signal is divided into a low frequency band (4-16Hz) and a high frequency band (16-40Hz), and spatial filtering and temporal filtering are performed respectively; the frequency domain and space-time domain features are extracted by a multi-scale convolutional neural network (CNN), the CNN includes a one-dimensional spatial convolution layer and a deep temporal convolution layer, and the ProbSparse self-attention mechanism is introduced after the convolution layer to capture the global dependency of long time series; the spectrum-space features of the low frequency band and the high frequency band are fused by a frequency band interaction operator (including addition, concatenation, Hadamard product or linear projection) to generate a multi-dimensional joint feature vector. The feature selection and model optimization based on the locust optimization algorithm (GOA) is as follows: the multi-dimensional joint feature vector is input into the locust optimization algorithm, feature subsets or feature weights are optimized through global search, and features with discriminative power are screened; the classification model is trained using the optimized features, and the fully connected layer parameters of the classification model are updated through the cross entropy loss function and the Adam optimizer. The motor imagery classification and performance evaluation: outputs the classification results, calculates the classification accuracy, ROC curve and confusion matrix, and visualizes the contribution of key time-frequency region features through gradient weighted class activation mapping (Grad-CAM).

2. The electroencephalogram detection method according to claim 1, characterized in that: The specific steps of the ProbSparse self-attention mechanism include: calculating the dot product of the query vector and the key vector to generate an attention score matrix; selecting key points based on probability distribution for sparse processing, reducing the computational complexity of the attention matrix from O(N 2 ) is reduced to O(N log N); the attention scores of the selected key points are retained and the remaining positions are set to zero.

3. The electroencephalogram detection method according to claim 1, characterized in that: The mathematical model of the locust optimization algorithm (GOA) is:

4. The electroencephalogram detection method according to claim 1, characterized in that: The input of the multi-scale convolutional neural network is a two-dimensional EEG signal map X∈R C×T , where C is the number of channels and T is the time point; the network output is the fused spectrum-temporal feature U l and U h , and generate a joint feature vector through the frequency band interaction operator.

5. The electroencephalogram detection method according to claim 1, characterized in that: The classification model is a deep neural network, including a ResNet or LSTM structure. During the training process, a five-fold cross-validation method is used to divide the data set, and the hyperparameters are adjusted through Bayesian optimization.

6. The electroencephalogram detection method according to claim 1, characterized in that: The performance evaluation further includes calculating the correlation between the features and MI-related brain activation patterns and interpreting the contribution of the features to the classification results using Shapley value (SHAP).

7. The electroencephalogram detection method according to claim 1, characterized in that: The method comprises the hardware and software modules of any one of claims 1 to 6, wherein the hardware module is an independently developed EEG acquisition board, and the software module integrates signal preprocessing, feature extraction, optimization algorithm and classification model training functions.

8. The electroencephalogram detection method according to claim 1, characterized in that: A computer program is stored, and when the program is executed by a processor, the electroencephalogram detection method described in any one of claims 1 to 6 is implemented.

Citation Information

Cited By

  • Electroencephalogram channel and feature selection device and method based on proxy model

    CN121834325A