A Motion Imagery EEG Signal Classification Method Based on Parallel Multi-Scale Filter Bank Time-Domain Convolution of Transfer Learning

Through the improved DSAN-MSFBCNN model, transfer learning and field adaptive technology are used to solve the problem of low signal-to-noise ratio and large individual differences in the classification of motor imagination EEG signals, achieving higher classification accuracy and robustness.

CN116522106BActive Publication Date: 2025-07-25BEIJING UNIV OF TECH
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202310218152.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-03-08
Publication Date
2025-07-25
Estimated Expiration
2043-03-08

AI Technical Summary

Technical Problem

In the classification of motor imagination EEG signals, the existing technology has problems such as unstable EEG signals, low signal-to-noise ratio, large individual differences, and poor model generalization ability in the classification of EEG signals, especially the EEG signals between different subjects do not conform to the independent and same distribution, resulting in poor classification results.

Method used

The transfer learning method is adopted, combined with fine-tune fine-tune and domain adaptive network (DSAN) to improve the MSFBCNN model, freeze the highest-level full-connection layer, use DSAN to adapt the full-connection layer, and add classification and adaptive loss functions after the full-connection layer to generate companion objective functions, and adjust softmax layer parameters to improve feature distinction.

Benefits of technology

It significantly improves the classification accuracy of motor imagined EEG signals and the robustness of the model, improves the adaptability to different subjects, and enhances feature extraction and classification performance.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116522106B_ABST
    Figure CN116522106B_ABST
Patent Text Reader

Abstract

A method for classifying motor imagery electroencephalogram (EEG) signals based on parallel multi-scale filter bank time-domain convolution of transfer learning belongs to the field of computer software. Aiming at the problem that cross-subject EEG signals do not conform to independent and identically distributed, enhancing the time-domain information of the EEG signals therein, an adaptive layer fine-tuning feature extraction method based on deep transfer learning of the MSFBCNN network is proposed, abbreviated as "DSAN-MSFBCNN". First, the fine-tune method is adopted to fine-tune the network model of the pre-trained model MSFBCNN, and the network structure before the top fully connected layer of the model is frozen; secondly, the domain adaptation method DSAN is used to adapt the fully connected layer to provide features with higher discrimination, thereby improving the classification accuracy. The "DSAN-MSFBCNN" model can achieve a high accuracy in the motor imagery classification task. Compared with the MSFBCNN model, the present invention improves the feature extraction and classification performance of motor imagery EEG signals, and the fine-tuned model has higher generalization ability.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The invention discloses a motor imagery electroencephalogram signal classification method based on transfer learning parallel multi-scale filter group time domain convolution, which can be used to identify motor imagery limb parts and belongs to the field of computer software. Background Art

[0002] Brain-computer interfaces (BCI) establish a direct path between the human brain and the computer through brain signal recording and decoding technology. Motor imagery is one of the most representative paradigms. Electroencephalogram (EEG)-based BCI has developed rapidly in the field of motor imagery (MI). As a classic paradigm, MI has been studied and developed for decades. Its physiological basis is that body movement can produce mu (8-12 Hz) and beta (16-26 Hz) rhythms in the motor sensory area of the brain, accompanied by event-related synchronization and desynchronization (ERS / ERD). Event-related synchronization and desynchronization means that when people perform unilateral limb movements and motor imagery, the amplitude of the mu rhythm and beta rhythm of the contralateral primary sensorimotor cortex of the brain is significantly reduced, while the amplitude of the mu rhythm and beta rhythm of the ipsilateral primary sensorimotor cortex is significantly increased. According to the difference in the EEG rhythm of the sensorimotor area, the EEG signals of different motor imagery tasks can be identified and classified, which is of great help and research significance for stroke patients or patients with damaged nervous system to restore sensory and motor functions. Although the brain-computer interface technology based on the motor imagery paradigm has been widely used in rehabilitation and other fields, its decoding performance still cannot meet the needs of practical applications. Because EEG signals are unstable signals with a very low signal-to-noise ratio, the collected EEG signals are easily affected by noise such as electrooculography and electromyography. How to extract high-precision EEG signals from raw EEG signals is also one of the problems to be solved. At present, there are three main measures to reduce noise interference: prevention, elimination, and correction. Correction is the mainstream method, and commonly used methods include adaptive filtering, wavelet analysis, and blind source separation. In addition, for different subjects, the EEG signals generated by the same imagination task are different.

[0003] To solve the above problems, machine learning methods and deep learning methods are often used for feature extraction and classification at present. The traditional classification methods for motor imagery electroencephalogram (EEG) signals mainly consist of two steps: feature extraction and classification. First, a series of algorithms are used to extract features from the original EEG signals, and then the extracted features are input into the corresponding classifier for classification and discrimination to obtain the final result. The traditional feature extraction method is the common spatial pattern (CSP) and its variants. The idea of CSP is to find a set of spatial filters to optimally distinguish multiple EEG recordings. The Filter-Bank CSP (FBCSP) algorithm benefits from manual feature selection to select the optimal spatial filters for feature extraction. This method has the advantages of simplicity and accuracy. Other CSP-based methods also extract potentially valuable components after analyzing EEG signals. The commonly used classification methods include linear discriminant analysis (LDA), support vector machine (SVM), Bayesian classifier, etc. Unfortunately, EEG features vary over time and vary significantly among different individuals. Therefore, for new applications of MI, the need for more robust and general feature extraction techniques is gradually increasing. Due to the low signal-to-noise ratio and time-related covariates of EEG signals, a reliable feature extraction method is required. The effective application of deep learning in various fields has been applied to EEG signal processing and achieved better results compared with traditional methods. Lawhern et al. proposed EEGNet, a general and compact convolutional neural network specifically designed for EEG recognition tasks. Compared with the ordinary convolutional neural network structure, they used depthwise separable convolutions (depthwise convolution + pointwise convolution) to construct a specific EEG model. Dai et al. proposed a convolutional neural network with hybrid convolutional scales (HS-CNN). The method they proposed effectively solved the problem that using a single convolutional scale in CNN limits the classification effect and further improved the classification accuracy. Wu et al. proposed a parallel multi-scale filter bank convolutional neural network (MSFBCNN) for motor imagery EEG classification. The novelty of this model lies in using four parallel temporal convolutional layers to extract temporal features and fully utilizing the feature information in the end-to-end network. It can be found from the above research that lightweight networks and filter banks play a key role in deep learning-based motor imagery EEG classification.

[0004] Currently, aiming at the problems existing in traditional machine learning algorithms and deep learning algorithms, such as the dependence on a large amount of data, the electroencephalogram (EEG) signals of different motor imagery subjects not conforming to independent and identically distributed, and the small sample size, resulting in poor model classification effect and poor generalization ability, the idea of transfer learning can be combined. The purpose of transfer learning is to apply the knowledge or patterns learned from one task to other different but related tasks. By using the transfer method, two models that can share parameters in two domains are constructed, and then the transfer between different domains is realized by adjusting the parameters shared between the two domains, so as to shorten the differences existing in the EEG signals of different subjects. At present, the application of transfer learning in deep learning models mainly uses two methods, namely domain adaptation (DA) and model fine-tuning method (fine-tune). The domain adaptation method can solve the problem of the non-uniformity of the "domain" of EEG signals to a certain extent by reducing the differences between domains while maintaining the discrimination ability for tasks. The common domain adaptation metric, that is, the method for measuring the distance between domains, is the Maximum Mean Discrepancy (MMD). Syed Umar Amin et al. proposed a transfer method for multi-layer neural networks to fuse neural networks with different features and architectures to improve the accuracy of motor imagery classification of EEG signals. In this work, different convolutional features are used to capture spatio-temporal features from the original EEG signal data. In the domain adaptation method, the Deep Subdomain Adaption Network (DSAN) uses a Local Maximum Mean Discrepancy (LMMD) to align the relevant subdomain distributions of domain-specific layer activations in different domains to learn the transfer network, and has achieved the best classification effect among metric-based methods in recent years. A subdomain refers to a situation where, under a domain, similar samples can be divided into a subdomain according to some conditions. For example, using class labels as the division basis, the same class is placed in a subdomain. DSAN only uses one classification loss function and one adaptive loss function to generate an associated objective function for the hidden layer associated with the middle layer of the backbone network, that is, the measured output layer, to provide features with higher discrimination. In the fusion stage, the deconvolution technique is used to perform the upsampling operation to aggregate the multi-resolution features extracted from different levels. The main idea of fine-tuning is to perform model structure pruning and retraining on the given Pretrained model, so that while the original feature extraction ability of the pre-trained model can be fully released and utilized, it can better adapt to the new motor imagery EEG signals and further improve the model performance. The main fine-tuning strategies include two steps: model freezing and model retraining.

[0005] By discussing and analyzing the advantages and disadvantages of the above existing research methods, the present invention has obtained new inspirations and research ideas. Taking MSFBCNN as the backbone model, a new improved network model (DSAN-MSFBCNN) is proposed. First, the fine-tune method is adopted to fine-tune the network model of the pre-trained model MSFBCNN. The network structure before the top fully connected layer of the model is frozen, and the domain adaptation method DSAN is used to adapt the fully connected layer. After the fully connected layer, a classification loss function and an adaptive loss function are used to generate an associated objective function for the hidden layer associated with the middle layer of the backbone network, that is, the measured output layer, to provide features with higher discrimination, thereby improving the classification accuracy. Secondly, the relevant parameters of the softmax layer are adjusted according to the relevant information such as the classification categories of the new dataset, so as to further improve the robustness of the model. Compared with the MSFBCNN model, the classification method proposed by the present invention can more effectively improve the decoding performance of the motor imagery electroencephalogram signals. Summary of the Invention

[0006] The present invention proposes a classification method for motor imagery electroencephalogram signals based on transfer learning parallel multi-scale filter bank time domain convolution, which can effectively improve the feature extraction and classification performance of motor imagery electroencephalogram signals. First, the fine-tune method is adopted to fine-tune the network model of the pre-trained model MSFBCNN. The network structure before the top fully connected layer of the model is frozen, and the domain adaptation method DSAN is used to adapt the fully connected layer. After the fully connected layer, a classification loss function and an adaptive loss function are used to generate an associated objective function for the hidden layer associated with the middle layer of the backbone network, that is, the measured output layer, to provide features with higher discrimination, thereby improving the classification accuracy. Secondly, the relevant parameters of the softmax layer are adjusted according to the relevant information such as the classification categories of the new dataset, and the probability of each category is predicted according to the number of classification categories. The category with the highest probability is the final output category. Thereby further improving the robustness of the model. Compared with MSFBCNN, the method proposed by the present invention has a higher classification accuracy.

[0007] After research, discussion and repeated practice, the final solution of this method is determined as follows:

[0008] First, the original electroencephalogram dataset is preprocessed, and the dataset is divided into three parts: a training set, a validation set and a test set, which are input into the MSFBCNN model for training and testing, and the model generated after training is saved; secondly, based on the domain adaptation method DSAN and the fine-tune method, the pre-trained model is fine-tuned, trained and tested. Finally, the model classification results are obtained, and the classification results are evaluated to verify the effectiveness of the method.

[0009] The specific steps of the technical solution of the present invention are as follows:

[0010] Step 1. Data preprocessing: Use a band-pass filter to perform band-pass filtering on the motor imagery EEG signals, and then perform exponential moving average normalization on the filtered signals; divide the EEG signal dataset into a training set, a validation set, and a test set;

[0011] Step 2. Construct the MSFBCNN pre-trained model: Input the training set and validation set in Step 1 into the MSFBCNN model for training and validation, and finally generate and save the MSFBCNN pre-trained model;

[0012] Step 3. Fine-tuning of the pre-trained model and adding an adaptive layer: Adopt the fine-tune method to perform network model fine-tuning on the pre-trained MSFBCNN model, freeze the network structure before the top fully-connected layer of the model, use the domain adaptation method DSAN to adapt the fully-connected layer, use a classification loss function and an adaptive loss function after the fully-connected layer, and generate an associated objective function for the hidden layer associated with the intermediate layer of the backbone network, that is, the measured output layer, to provide features with higher discrimination. Finally, adjust the relevant parameters of the softmax layer according to the relevant information such as the classification categories of the new dataset;

[0013] Step 4. Input the test set in Step 1 into the model in Step 3 for training, validation, and classification, and evaluate the classification accuracy.

[0014] The present invention has the following advantages:

[0015] 1. Add an adaptive layer to the MSFBCNN model using the DSAN method, improve the model loss function, perform feature alignment on related sub-domains, and can extract features of different dimensions from EEG signals, thereby further improving the accuracy of the motor imagery classification task.

[0016] 2. Fine-tuning the model can not only maintain the network performance but also significantly reduce the complexity of the model and improve the model training speed. Brief description of the drawings

[0017] Figure 1 Overall flowchart of the method of the present invention

[0018] Figure 2 Detailed composition diagram of the DSAN-MSFBCNN network

[0019] Figure 3 Time schematic diagram of the motor imagery dataset Detailed implementation manners

[0020] In view of the problems that the low signal-to-noise ratio of electroencephalogram (EEG) signals leads to difficult feature extraction, hard classification, and uneven probability distribution of the EEG signals of the subjects, a classification method for motor imagery EEG signals based on transfer learning parallel multi-scale filter bank time-domain convolution is proposed. First, the finetune method is adopted to fine-tune the network model of the pre-trained model MSFBCNN. The network structure before the top fully connected layer of the model is frozen, and the domain adaptation method DSAN is used to adapt the fully connected layer. A classification loss function and an adaptive loss function are used in the fully connected layer to generate an associated objective function for the hidden layer associated with the middle layer of the backbone network, that is, the measured output layer, to provide features with higher discrimination, thereby improving the classification accuracy. Second, the relevant parameters of the softmax layer are adjusted according to the relevant information such as the classification categories of the new data set, so as to further improve the robustness of the model, providing an efficient and better-performing deep learning method for the classification of motor imagery EEG signals.

[0021] Figure 1 As shown in , the general flowchart of the method of the present invention can be decomposed into the following steps:

[0022] Step 1. Data preprocessing and data set division;

[0023] Step 2. Construct the DSAN-MSFBCNN model;

[0024] Step 3. Use the training set and the validation set to train the model;

[0025] Step 4. Test the model effect and evaluate the classification accuracy.

[0026] The specific details of each step are described in detail below:

[0027] Step 1:

[0028] (1) Band-pass filter the original motor imagery EEG signals with a 3rd-order Butterworth band-pass filter of 4 - 40 Hz to filter out the signals in the required frequency band;

[0029] (2) Perform exponential moving average normalization on the filtered signals, where the decay factor is set to 0.999 to reduce the influence of numerical differences on the model effect;

[0030] (3) Before starting the training, divide the preprocessed EEG data set. 80% of the data in the training samples is used as the training set, and the remaining 20% of the data is used as the validation set.

[0031] Step 2:

[0032] Aiming at the problems of difficult feature extraction and classification caused by low signal-to-noise ratio of electroencephalogram (EEG) signals, based on the MSFBCNN pre-trained model, the present invention proposes a new model improvement method, abbreviated as "DSAN-MSFBCNN". The network structure of this method is generally divided into four parts: Block1, Block2, fully connected layer, and adaptive layer. Here, the model is constructed using Pytorch. The following is a detailed description of each part:

[0033] (1) Block1

[0034] The feature extraction layer, based on the two-dimensional convolutional layer, adopts a parallel multi-scale filter bank convolutional neural network to fully extract time features and spatial features. The structure of this layer is successively four parallel temporal convolutional kernels, a batch normalization layer, a spatial convolutional kernel, and a batch normalization layer. The process of feature extraction is that the extracted temporal features are combined into the input of the spatial convolution after batch normalization, and then the spatial features are extracted through the processing of the spatial convolution and the dimension of the feature map is reduced; finally, after the spatial convolution, the dimension of the EEG channels is compressed to 1, and both the temporal and spatial convolutions are extended into three-dimensional feature matrices, which are normalized and sent to the next layer for feature dimensionality reduction. The weights of the batch normalization layer are initialized using the mean and a normal distribution with zero unit variance, the batch size is 64, and the weights are initialized to 1, the learning rate is 1e-3, and the decay weight is 1e-7. After multiple experiments, the final settings of the four temporal convolutional kernel sizes (kernel size) are (64,1), (40,1), (26,1), and (16,1), which have the best effect.

[0035] (2) Block2

[0036] The feature dimensionality reduction layer, in order to improve the non-linear expression ability of the network, uses the sum of squares and logarithmic non-linear functions to extract features related to the band power. And the average pooling layer is used to further reduce the temporal dimension and the third dimension. First, the output of the upper layer is subjected to feature extraction using a spatial convolutional layer with a convolutional kernel size of (C,1), a convolutional stride of 1, and a maximum norm weight constraint max_norm of 0.5, where C is the number of leads of the collected EEG signals. The activation function used in this layer is the square activation function Square and the logarithmic non-linear activation function Log, which can not only speed up the training speed but also improve the classification accuracy. After passing through the activation function, an average pooling layer with a size of 75x1 and a stride of 15x1 is used to process the features to reduce the number of parameters. Finally, through the Dropout method, a strategy to prevent overfitting, during the training process, the nodes in the corresponding layer are randomly discarded with a probability of 0.5 to alleviate the overfitting phenomenon of the network.

[0037] (3) Fully connected layer

[0038] Integrate all the obtained features through a fully connected layer, and at the same time add a maximum norm constraint to the fully connected layer for regularization. The maximum norm value is set to 0.5 to prevent overfitting and improve the generalization ability of the model.

[0039] (4) Adaptive layer

[0040] Based on the DSAN method, after the fully connected layer, an adaptive layer is added to the model. The LMMD adaptive loss metric is used to align the features. The first term uses cross-entropy as the classification loss, and the second term uses LMMD as the adaptive loss. The category division is used to define the sub-domains, and one sub-domain is one category. Therefore, in different data distributions p and q, the definition of LMMD is as follows:

[0041]

[0042] Represents the LMMD distance between different data distributions p and q. Represents using the Hilbert space as the mapping space for distance calculation; s and t represent the source domain and the target domain respectively. And Represent the source domain data and the target domain data respectively. And Represent the i-th instance data in the source domain and the j-th instance data in the target domain respectively; C represents the number of classification types of the data set. Represents the data under the classification category c (c ∈ C) Belonging to the weight of this category. Represents the data under the classification category c (c ∈ C) Belonging to the weight of this category. φ(·) is the Euler's formula. Used to calculate In The number of numbers relatively prime to Used to calculate In (The number of numbers relatively prime to

[0043] The general formula for calculating the Euler's totient function φ(x) is:

[0044]

[0045] Where p i , (i ∈ n) are the prime factors of x, and x is a positive integer.

[0046] Use multiple similar classes to define the sub-domains, and the definition of the weight w is as follows:

[0047]

[0048] ( denotes the weight of x belonging to the classification category c (c ∈ C); y i ; y ic represents the instance data x i 's c-th category label, y jc represents the instance data x j 's c-th category label.)

[0049] Since the source domain data uses real labels, while the target domain data uses the probability distribution predicted by the network. The expansion of LMMD is as follows:

[0050]

[0051] (C represents the number of classification types in the dataset, c represents the c-th (c ∈ C) classification category; n s represents the number of labeled data in the source domain, n t represents the number of unlabeled data in the target domain; sc, tc represent the data belonging to the c-th category in the source domain and target domain data for the c-th category, sl, tl respectively represent the data of the category in the source domain data and target domain data on the l-th layer of the network model; respectively represent the calculation results of the activation function for the i-th and j-th data on the l-th layer of the network model with respect to the sl category, respectively represent the calculation results of the activation function for the i-th and j-th layers of the network model with respect to the tl category; and respectively represent the weights of the i-th and j-th data in the source domain belonging to this category for the sc (sc ∈ C) classification category, and respectively represent the weights of the i-th and j-th data in the source domain belonging to this category for the tc (sc ∈ C) classification category; represents the LMMD distance between different data distributions p and q; ) is used to calculate the difference between the calculated values of the activation function for the i-th and j-th data of the source domain instance on the l-th layer of the network and .)

[0052] The entire network has L layers and is optimized using the following loss function:

[0053]

[0054] Among the above, is the probability of the final classification category result of the model for the i-th data in the source domain , is the probability of the actual classification category of the i-th data in the source domain, n sis the source domain data volume; the first item is the minimum cross-entropy as the classification loss, is to calculate and the entropy of, the second item is LMMD as the adaptive loss, and its calculation formula is shown in Equation (4). λ is the adaptive loss parameter, which will be iterated according to different datasets and model parameters during the experimental training process until the loss function reaches the minimum target or the maximum number of iterations, and then the training ends and the final λ value is determined. In this experiment, the maximum number of iterations is set to 800, and the minimum value of the loss function is set to 0.15.

[0055] The calculation formula of entropy is as follows:

[0056]

[0057] (where x i is the same sample point in the common sample space of the p distribution and the q distribution, and the size of the sample space is K.)

[0058] Step 3: Input the training set and validation set of EEG signals into the DSAN-MSFBCNN model for training. The training process is divided into two stages. The maximum number of iterations in the first stage is set to 800, and when the validation set loss function reaches the lowest, the training is terminated in advance to prevent overfitting and save training time. In the second stage, the validation set data is merged into the training set data for training. When the validation set loss value is less than the training set loss value in the first stage, the training is terminated in advance, and the maximum number of iterations is still 800. Record the model when the validation set loss value is the lowest during the second iteration process, use it to predict the test set samples, and record the test set accuracy. The above model training and testing are carried out for 9 subjects respectively to obtain 9 groups of test set accuracies, and record their average value as the final model accuracy.

[0059] In the experiment, the cross-entropy loss function is used for the training of all methods, the Adam method is used as the optimizer, the learning rate is set to 0.001, and the remaining parameters use the default values of the Adam method. The batch size (batchsize) of batch training is set to 64.

[0060] Step 4: Input the test set in Step 1 into the trained model in Step 4 for classification recognition, and evaluate the classification accuracy.

[0061] The dataset and experimental results used in the method of the present invention are described as follows:

[0062] 1. Dataset

[0063] The present invention uses two publicly available datasets, BCI Competition IV Dataset 2a and 2b, for experiments. The time schematic diagram of the datasets is as shown in Figure 2 All experimental data have been preprocessed using a band-pass filter with a frequency range of 0.5 - 100 Hz.

[0064] Dataset 2a contains electroencephalogram (EEG) signals of four types of motor imagery, namely left hand, right hand, both feet, and tongue, from 9 subjects. These EEG signals are collected from 22 electrodes at a sampling rate of 250 Hz, and a total of 576 trials (i.e., the number of samples is 576) are included. These 576 samples are collected over two days, and each day's experiment is recorded as 1 session. Each session contains samples of 4 categories, with 72 samples in each category. For Dataset 2a, data from 0.5 seconds to 2.5 seconds after the presentation of the cue is extracted as one sample. All samples have been labeled (i.e., marked which part of the motor imagery the sample corresponds to).

[0065] Dataset 2b contains electroencephalogram (EEG) signals of two types of motor imagery, namely left hand and right hand, from 9 subjects. These EEG signals are collected from 3 electrodes at the same sampling rate of 250 Hz. For each subject, the motor imagery task is divided into 5 sessions. Different from Dataset 2a, the first 2 sessions in Dataset 2b are conducted without feedback, that is, EEG imagery data without visual feedback, and the last 3 sessions are EEG imagery data with visual feedback. For Dataset 2b, data from 0.5 seconds to 4 seconds after the presentation of the cue is extracted as one sample.

[0066] Since the EEG characteristics of different subjects vary greatly, for the classification experiment of EEG signals, the classification accuracy needs to be calculated separately for each subject, and the average value of the classification accuracies of multiple subjects is used as the performance index of the model.

[0067] 2. Experimental Results and Discussion

[0068] To verify the effectiveness and generality of the method of the present invention, comparative experiments and ablation experiments are respectively conducted on the publicly available datasets 2a and 2b, and the experimental results are as follows:

[0069] (1) Cross-dataset Comparative Experiment

[0070] The new method proposed by the present invention and the MSFBCNN method are respectively compared using Datasets 2a and 2b, and the experimental results are shown in Table 2:

[0071] Table 2 Results of Cross-dataset Comparative Experiment

[0072]

[0073]

[0074] As can be seen from Table 2, on the 2a and 2b datasets, the accuracy of the method proposed by the present invention is higher than that of the MSFBCNN method, and the highest improvement can be about 8.5% on the 2a dataset.

[0075] (2) Ablation experiment

[0076] Two groups of ablation experiments were conducted on the 2a dataset for the method of the present invention. One group is DSAN-MSFBCNN without adding an attention mechanism, and the other group is DSAN-MSFBCNN without using a parallel multi-scale convolutional layer. The experimental results are shown in Table 3:

[0077] Table 3 Results of ablation experiments

[0078]

[0079]

[0080] As can be seen from Table 3, although the results of the two groups of ablation experiments are higher than those of the MSFBCNN method, they are both lower than the method proposed by the present invention. This shows that both of these two schemes are effective and essential, and when both are present, the method of the present invention has higher classification performance.

Claims

1. A method for classifying motor imagery electroencephalogram signals based on parallel multi-scale filter bank time-domain convolution of transfer learning, characterized in that, It includes the following steps: Step 1. Data preprocessing: Perform band-pass filtering on the motor imagery EEG signals using a band-pass filter, and then perform exponential moving average normalization on the filtered signals; divide the EEG signal dataset into a training set, a validation set, and a test set; Step 2. Construct the MSFBCNN pre-trained model: Input the training set and the validation set in Step 1 into the MSFBCNN model for training and validation, and finally generate and save the MSFBCNN pre-trained model; Step 3. Fine-tuning of the pre-trained model and adding an adaptive layer: Adopt the fine-tune method to fine-tune the network model of the pre-trained model MSFBCNN, freeze the network structure before the highest fully-connected layer of the model, use the domain adaptation method DSAN to adapt the fully-connected layer, use a classification loss function and an adaptive loss function after the fully-connected layer, and generate an associated objective function for the hidden layer associated with the middle layer of the backbone network, that is, the measurement output layer, to provide features with higher discrimination; finally, adjust the relevant parameters of the softmax classification layer according to the relevant information of the new dataset classification categories, and predict the probability of each category according to the number of classification categories. The category with the highest probability is the final output category; Step 4. Input the test set in Step 1 into the model in Step 3 for training, validation, and classification, and evaluate the classification accuracy.

2. The method for classifying motor imagery EEG signals based on transfer learning parallel multi-scale filter bank time domain convolution according to claim 1, wherein: Step 1 is specifically as follows: (1) Perform band-pass filtering on the original motor imagery EEG signals using a 3rd-order Butterworth band-pass filter with a frequency range of 4 - 40 Hz to filter out the signals in the required frequency band; (2) Perform exponential moving average normalization on the filtered signals, where the decay factor is set to 0.999 to reduce the impact of numerical differences on the model effect; (3) Before starting training, divide the preprocessed EEG dataset; use 80% of the data in the training samples as the training set, and the remaining 20% of the data as the validation set.

3. The method for classifying motor imagery EEG signals based on transfer learning parallel multi-scale filter bank time domain convolution according to claim 1, wherein: Step 2 is specifically to construct a DSAN-MSFBCNN model for subsequent model fine-tuning: The specific structure of the DSAN-MSFBCNN model is generally divided into four parts: Block1, Block2, fully-connected layer, and adaptive layer; the following is a detailed description of each part: (1) Block1 Feature extraction layer, based on a two-dimensional convolutional layer, adopts a parallel multi-scale filter bank convolutional neural network to fully extract temporal features and spatial features; The layer structure consists of four parallel temporal convolutional kernels, a batch normalization layer, a spatial convolutional kernel, and a batch normalization layer in sequence; the process of feature extraction is that the extracted temporal features are combined into the input of the spatial convolution after batch normalization, and then the spatial features are extracted through the processing of the spatial convolution and the dimension of the feature map is reduced; finally, after the spatial convolution, the dimension of the EEG channels is compressed to 1, and both the temporal and spatial convolutions are expanded into three-dimensional feature matrices, which are normalized and sent to the next layer for feature dimensionality reduction; The weights of the batch normalization layer are initialized using a normal distribution with a mean of zero and a unit variance, the batch size is 64, and the weights are initialized to 1, the learning rate is 1e-3, and the decay weight is 1e-7; four temporal convolutional kernels are set with sizes of (64,1), (40,1), (26,1), and (16,1) respectively; (2)Block2 The square sum logarithmic nonlinear function is used to extract features related to the band power; and the average pooling layer is used to further reduce the temporal dimension and the third dimension; first, the output of the upper layer is subjected to feature extraction using a spatial convolutional layer with a convolutional kernel size of (C,1), a convolutional stride of 1, and a maximum norm weight constraint max_norm of 0.5, where C is the number of leads of the collected EEG signals; the activation function of this layer uses the square activation function Square and the logarithmic nonlinear activation function Log, and after passing through the activation function, an average pooling layer with a size of 75x1 and a stride of 15x1 is used to process the features to reduce the number of parameters; finally, through the Dropout method, the nodes in the corresponding layer are randomly discarded with a probability of 0.5 during the training process to alleviate the overfitting phenomenon of the network; (3)Fully connected layer All the obtained features are integrated through the fully connected layer, and a maximum norm constraint is added to the fully connected layer for regularization processing, and the maximum norm value is set to 0.5; (4)Adaptive layer Based on the DSAN method, after the fully connected layer, an adaptive layer is added to the model, and the LMMD adaptive loss metric is used to align the features. The first term uses cross-entropy as the classification loss, and the second term uses LMMD as the adaptive loss; category partitioning is used to define sub-domains, and one sub-domain is one category, so in different data distributions p and q, the definition of LMMD is as follows: Denotes the LMMD distance between different data distributions p and q. Indicates that the Hilbert space is used as the mapping space for distance calculation; s and t represent the source domain and the target domain respectively. and respectively represent the source domain data and the target domain data. and respectively represent the i-th instance data in the source domain and the j-th instance data in the target domain; C represents the number of dataset classification types. Indicates the data belonging to this class under the classification category c (c ∈ C). Indicates the data belonging to this class under the classification category c (c ∈ C); φ(·) is the Euler's totient function. Used to calculate the number of numbers relatively prime to in ; (Used to calculate the number of numbers relatively prime to The general formula for the Euler's totient function φ(x) is calculated as: where p i , (i ∈ n) is a prime factor of x, and x is a positive integer; Multiple similar classes are used to define sub-domains, and the definition of the weight w is as follows: Indicates the weight of x under the classification category c (c ∈ C) i belonging to this category; y ic represents the instance data x i of the c-th class label, y jc represents the instance data x j of the c-th class label; Because the source domain data uses real labels, while the target domain data uses the probability distribution predicted by the network; the expansion formula of LMMD is as follows: Let \(C\) denote the number of dataset classification types, and \(c\) denote the \(c\)-th classification category (\(c\in C\)); \(n\) s denotes the number of labeled data in the source domain, \(n\) t denotes the number of unlabeled data in the target domain; \(s_c\) and \(t_c\) denote the data belonging to category \(c\) in the source domain and the target domain data respectively in category \(c\), and \(s_l\) and \(t_l\) denote the data of the category in the source domain data and the target domain data respectively on the \(l\)-th layer of the network model; respectively denote the calculation results of the activation function for the \(i\)-th and \(j\)-th data with respect to the \(s_l\) category on the \(l\)-th layer of the network model, respectively denote the calculation results of the activation function for the \(i\)-th and \(j\)-th layers of the network model with respect to the \(t_l\) category; and respectively denote the weights of the \(i\)-th and \(j\)-th data in the source domain belonging to this category under the classification category \(s_c\) (\(s_c\in C\)), and respectively denote the weights of the \(i\)-th and \(j\)-th data in the source domain belonging to this category under the classification category \(t_c\) (\(s_c\in C\)); denotes the LMMD distance between different data distributions \(p\) and \(q\); is used to calculate the numerical value of the activation function calculation of the \(i\)-th and \(j\)-th data of the source domain instances on the \(l\)-th layer of the network and the difference between them; The entire network has L layers, and the following loss function is used for optimization: Among the above, is the probability of the model for the i-th data in the source domain for the final classification category result, is the probability of the actual classification category of the i-th data in the source domain, n s is the source domain data volume; the first term is the minimum cross-entropy as the classification loss, is to calculate and of the entropy, the second term is the LMMD as the adaptive loss, and the calculation formula is shown in Equation (4). λ is the adaptive loss parameter, which will be iterated according to different data sets and model parameters during the experimental training process until the loss function reaches the minimum target or the maximum number of iterations, and then the training ends and the final λ value is determined. The maximum number of iterations is 800, and the minimum of the loss function is set to 0.15; The calculation formula for entropy is as follows: where x i is the same sample point in the common sample space of the p-distribution and the q-distribution, and the size of the sample space is K.

4. The method for classifying motor imagery EEG signals based on parallel multi-scale filter bank time-domain convolution with transfer learning according to claim 1, characterized in that: Step 3: Input the training set and validation set of EEG signals into the DSAN-MSFBCNN model for training. The training process is divided into two stages. The maximum number of iterations in the first stage is set to 800, and when the loss function of the validation set reaches the lowest, the training is terminated early to prevent overfitting and save training time. In the second stage, the validation set data is merged into the training set data for training. When the loss value of the validation set is less than the loss value of the training set in the first stage, the training is terminated early, and the maximum number of iterations is still 800. Record the model when the loss value of the validation set is the lowest during the second iteration process, use it to predict the test set samples, and record the test set accuracy. Conduct the above model training and testing on 9 subjects respectively to obtain 9 groups of test set accuracies, and record their average value as the final model accuracy. In the experiment, the training of all methods uses the cross-entropy loss function, uses the Adam method as the optimizer, sets the learning rate to 0.001, and uses the default values of the Adam method for the remaining parameters. The batch size of batch training is set to 64.

5. The motor imagery EEG signal classification method based on transfer learning parallel multi-scale filter bank time-domain convolution according to claim 1, characterized in that: Step 4: Input the test set in Step 1 into the trained model in Step 4 for classification and recognition, and evaluate the classification accuracy.

Citation Information

Patent Citations

  • Motor imagery electroencephalogram signal classification method based on deep learning and mixed noise data enhancement

    CN113269048A

  • Motor imagery electroencephalogram signal classification method based on channel attention and multi-scale time domain convolution

    CN114266276A