EEG Signal Classification Method Based on Decoupled Representation Learning for Continuous Fast Visual Demonstrations
Through decoupling representation learning and matching comparison learning, the problem of low classification accuracy of EEG signals caused by sample category imbalance is solved, and higher classification accuracy and more robust feature representation are achieved.
Patent Information
- Application Number
- CN202210001444.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-01-04
- Publication Date
- 2025-06-27
- Estimated Expiration
- 2042-01-04
AI Technical Summary
The prior art does not fully consider the problem of sample categories, resulting in low accuracy in classification of EEG signals.
By decoupling the learning process into a representation learning process and a classifier learning process, the classifier avoids the impact of the representation learning process, and the independent learning feature representation is adopted to solve the problem of category imbalance.
It improves the classification accuracy of EEG signals, obtains more robust feature representations, and enhances classification performance.
Smart Images

Figure CN114118176B_ABST
Abstract
Description
Technical Field
[0001] This application belongs to the technical field of signal processing, and relates to a method for classifying electroencephalogram (EEG) signals. Specifically, it relates to a method for classifying continuous rapid serial visual presentation (RSVP) EEG signals based on decoupled representation learning, which can be used for image classification. Background Art
[0002] Brain-computer interface (BCI) technology enables information transmission between the human brain and external devices. Analyzing and utilizing EEG signals can help disabled people control wheelchairs, typewriters, and artificial prosthetics, thereby improving their quality of life. In recent years, with the continuous progress of social information technology, the problem of information overload has become increasingly serious. Picture and video data repositories are growing at an exponential rate, and the scale, diversity, and potential sparsity of "targets of interest" in these data repositories hinder the effective retrieval of targets. RSVP, as a current popular BCI paradigm, realizes the efficient classification and retrieval of large-scale images by combining the human visual system with the event-related potential (ERP) of the cerebral cortex. BCI systems based on the RSVP paradigm are often used to help professionals, such as military photo reconnaissance personnel, effectively classify large-scale satellite images.
[0003] The accurate classification of EEG signals is the key to realizing a BCI system based on the RSVP paradigm. Current EEG signal classification methods are divided into traditional methods and deep learning methods. Traditional methods mainly rely on manual feature extraction, such as statistical features in the time domain, band power in the frequency domain, and discrete wavelet transform in the time-frequency domain. EEG signals have the characteristics of non-stationarity and low signal-to-noise ratio. The extraction of manual features heavily depends on expert-level experience and prior domain knowledge, which restricts the extraction of deep features to a certain extent.
[0004] Recently developed deep learning methods are taking the lead in improving RSVP EEG classification. Different from traditional methods, deep learning can automatically extract the deep spatio-temporal information of EEG signals. Convolutional neural networks are the most common and performant among deep learning methods. Specifically, by performing sliding convolution operations on the EEG data input to the network, the same convolutional kernel is used during a single sliding process. After the convolutional operation completes feature extraction, the features are sent to a fully connected layer for classification. Among them, the representative ones include DeepConvNet proposed by Schirrmeister et al. in the paper "Deep learning with convolutional neural networks for EEG decoding and visualization.", and EEGNet proposed by Lawhern Vernon et al. in the paper "EEGNet: a compact convolutional neural network for EEG-based brain-computer interfaces." and other deep learning methods. Both of these methods use time-domain convolution and spatial convolution. After obtaining features through convolutional operations, they are processed by a processing unit and then sent to a convolutional classifier for classification. Although these methods achieve basic classification, due to the lack of consideration of the class imbalance problem, it restricts the development of deep learning methods in RSVP EEG signal classification. For the RSVP paradigm, the number of non-target samples is much larger than that of target samples, and the main reasons are as follows: (1) During the visual presentation process, a sufficient time interval should be maintained between two targets, which is conducive to the induction of EEG signals. (2) A large number of non-target pictures need to be inserted into the fast sequence to maintain a high refresh rate. The problem of sample class imbalance is not fully considered in the prior art, resulting in a low accuracy rate of EEG signal classification. Summary of the Invention
[0005] The purpose of the present invention is to provide a method for classifying continuous rapid visual presentation EEG signals based on decoupled representation learning, aiming at the deficiencies in the above-mentioned prior art, so as to solve the problem of low accuracy rate of EEG signal classification caused by the insufficient consideration of the sample class imbalance problem in the prior art.
[0006] To achieve the above object, the technical concept adopted by the present invention is as follows: A dataset with imbalanced sample classes only affects the training of the classifier, but does not affect the learning of feature representation. The fundamental reason for the problem of imbalanced sample classes affecting classification performance is that the backpropagation mechanism of the neural network causes the collapse of the classifier, which will inevitably affect the learning of feature representation. By decoupling the learning process into a representation learning process and a classifier learning process, the influence of the classifier on the representation learning process is avoided. Specifically, preprocess the multi-channel EEG signals, and use decoupled representation learning to complete the classification of EEG signals, thereby solving the problem of low classification accuracy of EEG signals caused by insufficient consideration of the imbalanced sample classes. At the same time, make full use of all data information to learn more robust feature representations, and higher classification accuracy can be obtained.
[0007] The technical solution provided by the present invention is as follows:
[0008] This application provides a method for classifying continuous rapid visual demonstration EEG signals based on decoupled representation learning. The method includes the following steps: S1, acquisition and preprocessing of EEG signals; S2, construction of a decoupled representation learning network; S3, training the constructed decoupled representation learning network; S4, testing the trained decoupled representation learning network; S5, fine-tuning the tested decoupled representation learning network; S6, real-time detection of the fine-tuned decoupled representation learning network.
[0009] Furthermore, in step S1, the EEG signals are directly collected from the subjects.
[0010] Furthermore, the continuous rapid visual demonstration experiments participated by the subjects are successively in four states: preparation, viewing, interval, and waiting.
[0011] Furthermore, the preprocessing of the EEG signals is successively data segment selection, filtering, downsampling, and normalization processing.
[0012] Furthermore, the dataset composed of EEG signals is divided into a training set, a validation set, and a test set according to the ratio of 7:1.5:1.5.
[0013] Furthermore, the decoupled representation learning network constructed in step S2 is composed of a representation learning process and a classifier learning process.
[0014] Furthermore, the loss function of the representation learning process in step S3 is a contrastive loss function.
[0015] Furthermore, the loss function of the classifier learning process in step S3 is a cross-entropy loss function.
[0016] Further, if the classification accuracy of the decoupled representation learning network in the training set and the test set in step S3 differs by within 20%, the trained decoupled representation learning network is obtained.
[0017] Further, in step S6, the EEG signals of the subject are collected again and preprocessed, and the preprocessing method is the same as the above-mentioned preprocessing method.
[0018] Compared with the prior art, the beneficial effects of the present invention are as follows:
[0019] (1) The present invention decouples the learning process into a representation learning process and a classifier learning process, avoiding the influence of the classifier on the representation learning process, solving the problem of low classification accuracy caused by the class imbalance problem in continuous rapid visual demonstration classification, and thus improving the classification accuracy of EEG signals.
[0020] (2) The present invention uses matching contrast learning to independently learn feature representations. Matching contrast learning further normalizes the representation space and learns deep feature representations by using prior label information and cosine similarity, providing a basis for improving the classification accuracy of EEG signals, and thus improving the classification accuracy of EEG signals. BRIEF DESCRIPTION OF THE DRAWINGS
[0021] Figure 1 is a schematic diagram of the method for classifying EEG signals of continuous rapid visual demonstration based on decoupled representation learning provided by the present invention;
[0022] Figure 2 is a task timing diagram for collecting EEG signals in step S1 of the method for classifying EEG signals of continuous rapid visual demonstration based on decoupled representation learning provided by the present invention;
[0023] Figure 3 is a schematic diagram of the target image for performing the continuous rapid visual demonstration task in step S1 of the method for classifying EEG signals of continuous rapid visual demonstration based on decoupled representation learning provided by the present invention;
[0024] Figure 4 is a schematic diagram of the non-target image for performing the continuous rapid visual demonstration task in step S1 of the method for classifying EEG signals of continuous rapid visual demonstration based on decoupled representation learning provided by the present invention;
[0025] Figure 5 is a structural block diagram of the decoupled representation learning network constructed in step S2 of the method for classifying EEG signals of continuous rapid visual demonstration based on decoupled representation learning provided by the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0026] In order to make the implementation process of the present invention clearer, the following will be described in detail with reference to the accompanying drawings.
[0027] Example 1:
[0028] The present invention provides a method for classifying electroencephalogram (EEG) signals in continuous rapid visual demonstrations based on decoupled representation learning. The method provided by the present invention is an effective solution to the problem of class imbalance and can be used in different feature extraction networks, such as EEGNet, DeepConvNet, etc. As Figure 1 shown, the method includes the following steps:
[0029] S1. Acquisition and preprocessing of EEG signals;
[0030] S11. Acquisition of EEG signals;
[0031] The EEG signals of the present invention can be the EEG signals in existing EEG datasets or the EEG signals directly collected from subjects. Specifically, the present invention uses the EEG signals directly collected from subjects. Multiple subjects wear electrode caps. In this embodiment, there are 8 subjects. The continuous rapid visual demonstration experiment is completed through four states: preparation, viewing, intermittent, and waiting. During the continuous rapid visual demonstration, the EEG signals of the subjects are collected through the electrodes on the electrode cap. The specific process is as follows:
[0032] All subjects have normal vision or corrected vision. The age distribution of the subjects is from 19 to 27. Among the 8 subjects, 6 are male and 2 are female. Gender and age are not used as screening conditions for subjects, but only as information about the subjects in the embodiments of the present invention. All subjects have no history of nervous system diseases or other serious diseases, so that the experimental results obtained are more reliable. Before the experiment, each subject was clearly told the experimental precautions and all signed written consent forms. The subjects wear 64-channel EEG electrode caps with a sampling rate of 1024 Hz, and EEG paste is applied to keep the impedance of each electrode below 25 kΩ to ensure high-quality EEG signals.
[0033] As Figure 2 shown, each experiment has four states in chronological order, namely: preparation state, viewing state, intermittent state, and waiting state. In the preparation state, a crosshair will appear on the screen to facilitate the subject to concentrate and wait for continuous viewing of the picture sequence. After 2 s, the playback of the picture sequence will start. In the viewing state, in this embodiment, a total of 500 target images and 1000 non-target images are collected. Each time, 50 images are selected from them, among which 1 - 5 are target images and the rest are non-target images. They appear randomly in the center of the screen at a frequency of 10 Hz. The target images and non-target images in the picture sequence are set according to the purposes or characteristics of different experiments. In this embodiment, taking images of people and cars as target images and natural landscape images as non-target images as an example, respectively as Figure 3 and Figure 4As shown. The EEG signal when the subject views the target image is the target EEG signal, and the EEG signal when the subject views the non-target image is the non-target EEG signal. The method of the present invention can distinguish the target EEG signal and the non-target EEG signal to achieve the classification of EEG signals. During the viewing state, there will be an intermittent state after every 50 images are displayed to facilitate the subject to adjust the state; after 10 viewings, it will enter the waiting state for the subject to rest and adjust their state. After 4s, it will enter the next experiment; this process repeats 30 times, and a total of 15,000 single-trial samples are collected in this embodiment.
[0034] S12, preprocessing of EEG signals.
[0035] The EEG signals collected in step S11 are synthesized into a data set. After sequentially performing data segment selection, filtering, downsampling, and normalization on this data set, the preprocessed EEG signals are obtained. Data segment selection is to select 1 second of data in the interval [0s, 1s] from the collected EEG data, that is, to select the data from the start of the continuous rapid visual presentation of the target image or non-target image to 1 second after the start. The selected EEG signals contain event-related potential signals, which provides a basis for the classification of EEG signals. Filtering is to filter the selected time period data using a sixth-order Butterworth bandpass filter with a cut-off frequency of 0.1 - 48Hz, which can remove the interference of noise and obtain EEG signals with a higher signal-to-noise ratio. Downsampling is to reduce the sampling rate of the filtered data to 256Hz, which can effectively reduce the scale of EEG data, thereby reducing the computational amount and improving the classification efficiency of EEG signals. Normalization is to normalize the downsampled data using the Z-score method, which standardizes the data space and is beneficial to the training of deep learning.
[0036] The preprocessed EEG data is divided into a training set, a validation set, and a test set according to the ratio of 7:1.5:1.5, which can maximize the value of EEG data and is consistent with the division method in the prior art, facilitating the comparison of classification effects.
[0037] S2, construction of the decoupled representation learning network;
[0038] The present invention uses a matching contrast learning framework for independent representation learning. As Figure 5As shown in the figure, the decoupled representation learning network consists of a representation learning process and a classifier learning process, where the representation learning process is based on the matching contrast learning framework. The representation learning process and the classifier learning process share the same feature extractor through weight sharing, so that robust feature representations can be extracted and the classification accuracy can be improved. The feature extractor uses a convolutional neural network to extract the temporal and spatial features of the EEG signals successively. The specific network settings have been published in DeepConvNet proposed by Schirrmeister et al. in "Deep learning with convolutional neural networks for EEG decoding and visualization." The input of the representation learning process is two sample data randomly selected from the database, and the input of the classifier learning process is one sample data randomly selected from the database. First, the parameters of the feature extractor are learned through the representation learning process, and then the parameters of the classifier are learned through the classifier learning process.
[0039] The representation learning process is divided into four steps:
[0040] In the first step, sample data pairs are constructed. Two sample data are randomly selected from the sample database. If both sample data correspond to target EEG signals or non-target EEG signal images, they are defined as positive pairs; if one is a target EEG signal and the other is a non-target EEG signal, they are defined as negative pairs. In this way, it can adapt to the double-branch network structure;
[0041] In the second step, feature representations are extracted. The feature extractor uses a general feature extraction network (such as EEGNet, DeepConvNet, etc.) to extract the temporal and spatial features of the EEG signals successively, and obtains the general feature representations of the two sample data, that is, such as Figure 5 the h in 1 and h 2 shown;
[0042] In the third step, the projection head is used to map the feature representation to a low-dimensional space, which is convenient for calculating the contrast loss. The two branches of the representation learning process share a projection head through weight sharing. Specifically, it consists of two fully connected layers and an activation function;
[0043] In the fourth step, the cosine similarity is used to calculate the similarity between the feature representations of the two sample data mapped to the low-dimensional space in the third step. Through backpropagation, the contrast loss is minimized, and a more robust feature representation is learned, which is beneficial to improving the classification performance.
[0044] The classifier learning process first performs downsampling on the dataset in terms of sample dimensions to obtain a balanced dataset, which can promote the learning of the classifier, making the classifier not easily collapse and thus not easily affecting the representation learning process. Therefore, the method of the present invention can improve the classification performance of EEG signals. Then, freeze the weights of the representation learning process in the feature extractor. After freezing, the representation learning process has no effect on the feature extractor. Under the action of the information shared in the representation learning process, the feature extraction task of the classifier learning process is completed. The features extracted by the feature extraction task are more effective and are beneficial for the convolutional classifier to classify. The convolutional classifier is Figure 5 the classifier shown, which is used to classify the processed features.
[0045] S3. Train the constructed decoupled representation learning network;
[0046] Use the training set EEG data obtained in step S12 to iteratively train the decoupled representation learning network by the gradient descent method, and use the validation set EEG data obtained in step S12 to test the training results each time, and finally obtain the trained decoupled representation learning network. The representation learning process is first trained for 100 epochs, and the classifier learning process is trained for 20 epochs.
[0047] The relevant parameter settings during the training process are as follows:
[0048] Parameter settings and updates of the representation learning process. The loss function is the contrastive loss function, so that the same-class samples are closer in the representation space and the different-class samples are farther in the representation space, which is beneficial for classification. The expression of the ratio loss function is as follows:
[0049]
[0050] where L is the contrastive loss, M is the number of pairs of matching sample data, are the low-dimensional representations of the two samples in the sample data pair, q is the label, exp is the exponential function with as the base, and sim is the cosine function calculation method.
[0051] The optimizer adopts the adaptive moment estimation optimizer, and the initial learning rate is 0.001. Each time, 1024 pairs of single-trial samples are matched and sent into the representation learning process. First, the feature extraction of the sample data is performed, then it is mapped to the low-dimensional space by the projection head, and finally the cosine similarity of each pair of samples is calculated. The contrastive loss between the two samples is calculated according to the cosine similarity and the label, and then the adaptive moment estimation optimizer updates the parameters in the feature extractor and the projection head according to the contrastive loss, so that a more robust feature representation can be learned faster, and the above operations are iteratively executed 20 times.
[0052] Parameter setting and update of the classifier learning process. Set the number of training times to 150, the single-sample input volume to 4, and the loss function to the cross-entropy loss function:
[0053]
[0054] Among them, is the cross-entropy loss, N is the number of samples, y i is the one-hot encoding of the label, is the classification head or projection head, Graph Convolutional Neural Network, X i is the input data. The optimizer uses the Adaptive Moment Estimation optimizer, and the initial learning rate is 0.001.
[0055] First, freeze the parameters of the feature extractor. Each time, take 4 single-trial samples from the balanced dataset and send them into the classifier learning process. Use the trained feature extraction network to extract the features of the EEG data, and then send the features of the EEG data into the convolutional classifier for classification. Calculate the cross-entropy loss according to the classification result and the true label of the sample, and then the Adaptive Moment Estimation optimizer updates the parameters in the convolutional classifier of the disentangled representation learning network according to the cross-entropy loss to complete the classifier learning process. Traverse all the sample data in the training set to complete one training. Every 10 iterations of training, calculate the accuracy of the disentangled representation learning network on the training set and the validation set. Compare the accuracy of the disentangled representation learning network on the training set and the test set: If the accuracy of this network on the training set is more than 20% higher than that on the validation set, then overfitting has occurred. At this time, reduce the learning rate to 90% of the current value and retrain, and update the parameters of the convolutional classifier again; If the accuracy of this network on the training set and the test set differs by less than 20%, stop training to obtain the trained disentangled representation learning network.
[0056] S4. Test the trained disentangled representation learning network;
[0057] Directly send the EEG data in the test set into the trained disentangled representation learning network for classification to obtain the classification result, and count the classification result to obtain the classification accuracy of the trained disentangled representation learning network on the EEG data in the test set. The accuracy of the experimental result is 87.64%, which is higher than the 81.57% obtained by Schirrmeister et al. in the article "Deep learning with convolutional neural networks for EEG decoding and visualization."
[0058] S5. Fine-tune the tested disentangled representation learning network;
[0059] Adjust the learning rates of the feature extractor, projection head, and convolutional classifier in the decoupled representation learning network to 1 / 27, 1 / 9, and 1 / 3 of the original values respectively. Using the adjusted learning rates, fine-tune the tested decoupled representation learning network with the EEG data of the current subject to obtain a decoupled representation learning network suitable for the current subject to conduct online experiments. Similarly, for each subject, a decoupled representation learning network suitable for each subject to conduct online experiments can be obtained.
[0060] S6. Conduct real-time detection on the fine-tuned decoupled representation learning network.
[0061] Collect the EEG signals of each subject in step S1 again for online real-time detection. That is, first perform preprocessing on the collected EEG signals of each subject in sequence, including data segment selection, filtering, downsampling, and normalization; then send the preprocessed EEG signals into the fine-tuned ideal decoupled representation learning network to obtain the real-time classification results of the EEG data of each subject.
[0062] S61. Conduct the EEG signal detection tasks corresponding to the target images and non-target images according to the continuous rapid visual presentation experimental paradigm, and collect the EEG signals of the subject again in real time through the electrodes on the electrode cap. 64 scalp EEG channels are used for collection, and the sampling rate is 1024 Hz. Each subject collects 50 EEG data samples for each experiment, and a total of 10 experiments are conducted, collecting 500 real-time single-trial samples in total.
[0063] S62. Preprocess the EEG data of the subject obtained in real time according to the same method as in step S12, and send the preprocessed EEG data into the decoupled representation learning network that has been trained, tested, and fine-tuned in step S5 to obtain the real-time classification results of the EEG data.
[0064] Steps S5 and S6 make the method of the present invention easier to adapt to other subjects and improve the classification effect of other subjects.
[0065] The above are only the preferred embodiments of the present invention and are not used to limit the present invention. For those skilled in the art, the present invention can have various changes and modifications. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present invention shall be included within the protection scope of the present invention.
Claims
1. A method for classifying electroencephalogram signals in continuous fast visual demonstrations based on decoupled representation learning, characterized in that The method includes the following steps: S1, acquisition and preprocessing of electroencephalogram (EEG) signals; S2, construction of a decoupled representation learning network; S3, training the constructed decoupled representation learning network; S4, testing the trained decoupled representation learning network; S5, fine-tuning the tested decoupled representation learning network; S6, performing real-time detection on the fine-tuned decoupled representation learning network; The decoupled representation learning network constructed in step S2 consists of a representation learning process and a classifier learning process; The representation learning process includes: The first step, constructing sample data pairs; randomly selecting two sample data from the sample database, if both sample data correspond to target EEG signals or non-target EEG signal images, they are defined as positive pairs; if one is a target EEG signal and the other is a non-target EEG signal, they are defined as negative pairs; Step 2: Extract the feature representation to obtain the common feature representations h 1 and h 2 ; In the third step, use the projection head to project the feature representation h 1 and h 2 into a low-dimensional space, where the projection head consists of two fully connected layers and an activation function; The fourth step, using cosine similarity to calculate the similarity between the feature representations of the two sample data mapped to the low-dimensional space in the third step, and minimizing the contrastive loss through backpropagation; The loss function of the representation learning process in step S3 is a contrastive loss function; the expression of the contrastive loss function is: Among them, L is the contrastive loss, M is the logarithm of the number of matching sample data pairs, is the low-dimensional representation of the two samples in the sample data pair, q is the label, exp is the exponential function with as the base, and sim is the calculation method of the cosine function; The classifier learning process includes: first, performing downsampling operation on the dataset in the sample dimension to obtain a balanced dataset, and then freezing the weights of the representation learning process in the feature extractor, and completing the feature extraction task of the classifier learning process under the action of the information shared in the representation learning process.
2. The method for classifying electroencephalogram signals in continuous fast visual demonstrations based on decoupled representation learning according to claim 1, wherein The EEG signals in step S1 are directly collected from the subject.
3. The continuous fast visual demonstration EEG signal classification method based on decoupled representation learning according to claim 2, wherein The continuous rapid visual presentation experiments participated by the subject are successively in four states: preparation, viewing, interval, and waiting.
4. The method for classifying EEG signals of continuous fast visual demonstrations based on decoupled representation learning according to claim 3, characterized in that The preprocessing of the EEG signals is successively data segment selection, filtering, downsampling, and normalization processing.
5. The continuous fast visual demonstration EEG signal classification method based on decoupled representation learning according to claim 4, wherein The dataset composed of the EEG signals is divided into a training set, a validation set, and a test set according to the ratio of 7:1.5:1.
5.
6. The continuous fast visual demonstration EEG signal classification method based on decoupled representation learning according to claim 5, characterized in that, The loss function of the classifier learning process in step S3 is a cross-entropy loss function.
7. The method for classifying electroencephalogram signals in continuous fast visual demonstrations based on decoupled representation learning according to claim 6, wherein If the difference in the classification accuracy of the decoupled representation learning network in the training set and the test set in step S3 is within 20%, the trained decoupled representation learning network is obtained.
8. The method for classifying EEG signals of continuous fast visual demonstrations based on decoupled representation learning according to claim 7, wherein In step S6, the EEG signals of the subject are collected again and preprocessed, and the preprocessing method is the same as the preprocessing method described in claim 4.
Citation Information
Patent Citations
Motor imagery classification method based on convolutional neural network
CN110765920A
Cross-age face image recognition method based on de-entanglement representation learning
CN112766157A
Electroencephalogram emotion recognition method based on efficient convolutional neural network and comparative learning
CN113673434A