Electroencephalogram signal emotion detection method, system, equipment and medium
By combining time-frequency feature extraction, attention mechanism and self-supervised comparison learning methods, the problems of single features and unbalanced samples in EEG emotional recognition are solved, and efficient emotion detection is achieved.
Patent Information
- Application Number
- CN202510424761.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-07
- Publication Date
- 2025-07-04
- Estimated Expiration
- 2045-04-07
AI Technical Summary
In the prior art, the EEG signal emotion recognition method has problems such as single features, unbalanced samples and small samples, resulting in insufficient accuracy of emotion detection.
The time-frequency feature extraction module, attention mechanism model, self-supervised comparison learning module and fully connected prediction module are used to integrate the time-frequency features of EEG signals to improve detection accuracy through wavelet transformation, multi-scale convolution and self-supervised comparison learning.
It effectively overcomes the problems of single features and uneven samples, and significantly improves the accuracy of emotion detection.
Smart Images

Figure CN120241068A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of electroencephalogram (EEG) signal analysis, and particularly relates to an EEG signal emotion detection method, system, device and medium. Background Art
[0002] Emotion recognition is an important research direction in the fields of human-computer interaction, mental health monitoring, intelligent education, emotion computing, etc. Traditional emotion recognition methods mostly rely on inputs such as facial expressions, speech, and text. However, electroencephalogram (EEG) signals, as a physiological signal that directly reflects brain activities, have become a powerful emotion recognition tool with great potential. EEG signals have high timeliness, low invasiveness, and high spatial resolution, and can effectively reflect an individual's emotional state and psychological changes. Therefore, the research on emotion recognition based on EEG signals has attracted extensive attention.
[0003] In recent years, significant progress has been made in the field of EEG signal emotion recognition. Traditional methods mainly rely on classical machine learning algorithms, such as support vector machine (SVM), k-nearest neighbor (KNN), etc., combined with manually extracted features (such as spectral features, waveform features, etc.). With the rapid development of deep learning technology, emotion recognition methods based on deep learning models such as convolutional neural network (CNN), recurrent neural network (RNN), and Transformer have gradually become the mainstream of research. The method based on convolutional neural network (CNN) usually performs time-frequency transformation on EEG signals to generate spectrograms, and applies a convolutional neural network to the spectrograms for emotion recognition. The advantage of this method is that it can automatically extract features, avoid manual intervention, and has high classification accuracy; its disadvantage is that it cannot fully exploit the temporal features of EEG signals. The method based on recurrent neural network (RNN) can effectively process the time series information in EEG signals and identify long-term dependencies by capturing temporal dependencies, but its disadvantages are that the training process is long, the computational resources consumed are large, and it is sensitive to noise. The EEG signal classification method based on Transformer captures long-distance dependencies through the self-attention mechanism and can efficiently process the temporal features of EEG signals. Its advantages are that it has strong parallel computing capabilities and is suitable for long sequence data; however, this method has a complex model, a large demand for training data, and a high computational cost. Therefore, there is an urgent need for an emotion recognition method that comprehensively considers the overall features of EEG signals, has high detection capabilities, and can provide accurate classification results. Summary of the Invention
[0004] The purpose of the present invention is to provide an EEG signal emotion detection method, system, device and medium to solve the problems existing in the above-mentioned prior art.
[0005] To achieve the above purpose, the present invention provides an EEG signal emotion detection method, including:
[0006] Obtain electroencephalogram (EEG) signal data;
[0007] Input the EEG signal data into an emotion detection model for prediction and classification to obtain an emotion classification result; wherein, the emotion detection model includes a time-frequency feature extraction module, an attention mechanism model, a self-supervised contrast learning module, a pooling module, and a fully-connected prediction module that are connected in sequence.
[0008] Optionally, the training process of the emotion detection model specifically includes:
[0009] Obtain training data, where the training data includes EEG signal training data and corresponding emotion classification results;
[0010] Construct an initial emotion detection model, input the training data into the initial emotion detection model for prediction and classification, and perform training with the goal of minimizing the loss between the initial training result after prediction and classification and the emotion classification result corresponding to the EEG signal training data, to obtain a trained emotion detection model.
[0011] Optionally, the processing process of the emotion detection model specifically includes:
[0012] Input the EEG signal data into the time-frequency feature extraction module, perform wavelet transform on the EEG signal data to obtain the frequency-domain signal corresponding to the EEG signal data; use a parallel multi-scale one-dimensional convolutional layer to encode the EEG signal data to obtain a time-domain feature encoding; use a 3-layer one-dimensional convolutional layer to encode the frequency-domain signal to obtain a frequency-domain feature encoding; perform one-dimensional convolution on the EEG signal data and the frequency-domain signal respectively to obtain an initial time-domain feature and an initial frequency-domain feature;
[0013] Apply a self-attention mechanism to fuse the initial time-domain feature and the frequency-domain feature encoding to obtain a first time-frequency feature, and fuse the initial frequency-domain feature and the time-domain feature encoding to obtain a second time-frequency feature;
[0014] Connect and fuse the first time-frequency feature and the second time-frequency feature to obtain a fused feature vector;
[0015] Perform global average pooling on the feature vector to obtain a compressed feature vector;
[0016] Input the compressed feature vector into the fully-connected prediction module, and output a classification result through a softmax layer.
[0017] Optionally, the performing wavelet transform on the EEG signal data specifically includes:
[0018] Apply continuous wavelet transform to each channel of each sample in the electroencephalogram (EEG) signal data, convert the EEG signal from the time domain to the frequency domain, and generate a spectrogram; during the continuous wavelet transform process, use the Morlet wavelet as the mother wavelet.
[0019] Optionally, it further includes inputting the time-frequency features into a self-supervised contrast learning module, calculating the contrast loss according to the cosine similarity, and optimizing the emotion detection model by using the contrast loss.
[0020] An EEG signal emotion detection system includes:
[0021] A data acquisition module for acquiring EEG signal data;
[0022] An emotion detection module for inputting the EEG signal data into an emotion detection model for prediction and classification to obtain an emotion classification result; wherein, the emotion detection model includes a time-frequency feature extraction module, an attention mechanism model, a self-supervised contrast learning module, a pooling module, and a fully connected prediction module connected in sequence.
[0023] An electronic device includes a memory and a processor, the memory is used for storing a computer program, and the processor runs the computer program to enable the electronic device to execute the described EEG signal emotion detection method.
[0024] A computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, it implements the described EEG signal emotion detection method.
[0025] The technical effects of the present invention are:
[0026] The present invention proposes an EEG signal emotion detection method based on time-frequency feature fusion and self-supervised contrast learning. This method realizes the efficient detection of emotions by fusing the time-frequency features of EEG signals and self-supervised learning technology; this solution effectively overcomes the problems of single features, sample imbalance, and small samples existing in the prior art, and at the same time significantly improves the accuracy of emotion detection. Description of the Drawings
[0027] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for use in the embodiments. Obviously, the following described drawings are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on these drawings.
[0028] The drawings constituting a part of this application are used to provide a further understanding of this application. The schematic embodiments of this application and their descriptions are used to explain this application and do not constitute an improper limitation of this application. In the drawings:
[0029] Figure 1 This is a schematic diagram of the model structure in the embodiments of the present invention. Detailed implementation manners
[0030] Now, various exemplary implementation manners of the present invention will be described in detail. This detailed description should not be regarded as a limitation of the present invention, but rather as a more detailed description of certain aspects, features, and implementation schemes of the present invention.
[0031] It should be understood that the terms described in the present invention are only used to describe specific implementation manners and are not used to limit the present invention. Additionally, for the numerical ranges in the present invention, it should be understood that each intermediate value between the upper and lower limits of the range is also specifically disclosed. Each intermediate value within any stated value or stated range, as well as each smaller range between any other stated value or intermediate value within the stated range, is also included in the present invention. The upper and lower limits of these smaller ranges may be independently included or excluded from the range.
[0032] Without departing from the scope or spirit of the present invention, various improvements and changes can be made to the specific implementation manners of the description of the present invention, which are obvious to those skilled in the art. Other implementation manners obtained from the description of the present invention are obvious to those skilled in the art. The description and embodiments of this application are only exemplary.
[0033] Regarding the terms "comprising", "including", "having", "containing", etc. used herein, they are all open-ended terms, meaning including but not limited to.
[0034] It should be noted that, without conflict, the embodiments in this application and the features in the embodiments can be combined with each other. The following will refer to the drawings and combine with the embodiments to detail this application.
[0035] As Figure 1 shown, in this embodiment, an electroencephalogram (EEG) signal emotion detection method is provided, including: acquiring EEG signal data; inputting the EEG signal data into an emotion detection model for prediction and classification to obtain an emotion classification result; wherein, the emotion detection model includes a time-frequency feature extraction module, an attention mechanism model, a self-supervised contrast learning module, a pooling module, and a fully connected prediction module connected in sequence.
[0036] In this embodiment, first, the EEG signal is transformed from the time domain to the frequency domain through wavelet transform to generate a spectrogram; then, one-dimensional convolution (1DCNN) is used to encode the time domain and frequency domain of the EEG signal respectively; then, the fusion of time-frequency encoding is performed, and a time-frequency consistent EEG signal encoding is trained through self-supervised contrast learning; finally, the fused time-frequency encoding is connected and globally averaged pooled, and through a fully connected layer and a softmax layer, the emotion detection result is output.
[0037] This embodiment specifically includes the following steps:
[0038] Step 1: Use wavelet transform to convert the EEG signal from the time domain to the frequency domain and generate a spectrogram.
[0039] Step 2: Encode the EEG signal in the time domain using parallel multi-scale one-dimensional convolution (1DCNN) to extract time-domain features.
[0040] Step 3: Encode the EEG signal in the frequency domain through three layers of one-dimensional convolution (1DCNN) to extract frequency-domain features.
[0041] Step 4: After performing one-dimensional convolution (1DCNN) transformation on the EEG signal in the frequency domain, fuse it with the time-domain feature encoding in Step 2; after performing one-dimensional convolution (1DCNN) transformation on the EEG signal in the time domain, fuse it with the frequency-domain feature encoding in Step 3.
[0042] Step 5: Use self-supervised contrastive learning to train the feature vectors fused in Steps 4 and 5 to learn the time-frequency consistent EEG signal encoding.
[0043] Step 6: Concatenate and globally average pool the encoding in Step 5, then pass through a fully connected layer and a softmax layer to output the emotion detection result.
[0044] The present invention proposes an EEG signal emotion detection method based on time-frequency feature fusion and self-supervised contrastive learning. This method realizes efficient emotion detection by fusing the time-frequency features of EEG signals with self-supervised learning techniques. This solution effectively overcomes the problems of single features, sample imbalance, and small samples existing in the prior art, and at the same time significantly improves the accuracy of emotion detection.
[0045] The specific implementation process of this embodiment includes:
[0046] Step 1, assume that the size of the EEG signal (EEG) X is K×C×N, where K is the number of sampling points, C is the number of channels, and N is the number of samples. For the j-th channel of the i-th sample of X, perform continuous wavelet transform using formula (1) to obtain the spectrogram I(i,j).
[0047]
[0048] Where ψ(t) is the Morlet wavelet, a is the scale factor, and b is the translation factor. The EEG signal (EEG) X is transformed by wavelet transform to obtain a spectrogram I of S×C×N.
[0049] Step 2, Time-domain encoding of EEG signals. Since the original EEG signal data is a one-dimensional time series, the data is only correlated with time in the horizontal direction and has no correlation in the vertical direction. Therefore, a one-dimensional convolutional neural network (CNN) is used to extract features and encode the EEG signals in the time domain. The specific process is as follows:
[0050] (1) The EEG signals in the time domain go through 3 parallel groups of one-dimensional dilated convolution operations, which can capture the multi-scale time-domain features of the EEG signals without increasing the parameters. The structures of these 3 groups of one-dimensional dilated convolution are as follows:
[0051] The first group of one-dimensional dilated convolution has two layers: the first layer uses 16 filters, the kernel size is 3, the stride is 1, the dilation coefficient is 1, the activation function uses GeLU, and L2 regularization is used to prevent overfitting. The second layer is exactly the same as the first layer.
[0052] The second group of one-dimensional dilated convolution has two layers: the first layer uses 16 filters, the kernel size is 3, the stride is 1, the activation function uses GeLU, and L2 regularization is used to prevent overfitting. The difference between the second layer and the first layer is that the dilation coefficient is 2, and the others are the same.
[0053] The third group of one-dimensional dilated convolution has two layers: the first layer uses 16 filters, the kernel size is 3, the stride is 1, the activation function uses GeLU, and L2 regularization is used to prevent overfitting. The difference between the second layer and the first layer is that the dilation coefficient is 3, and the others are the same.
[0054] Then the results of the three groups of parallel convolutions are merged and connected together, and input into the following two one-dimensional convolutional layers.
[0055] (2) Apply two one-dimensional convolutional layer operations to the output result in (1): the first layer of convolution uses 32 filters, the kernel size is 3, the stride is 1, uses the GeLU activation function, and uses L2 regularization to prevent overfitting; the second layer of convolution uses 64 filters, the kernel size is 3, the stride is 1, also uses the ReLU activation function, and uses L2 regularization to prevent overfitting.
[0056] Step 3, Frequency-domain encoding of EEG signals: This is a network for frequency-domain feature extraction, and its purpose is to extract frequency-domain features from the spectrogram I in Step 1 using a 4-layer one-dimensional convolutional neural network (CNN). The specific convolution operations for each layer are as follows:
[0057] (1) The first layer of convolution uses 16 filters, the kernel size is 3, the stride is 1, the activation function uses GeLU, which can better capture the non-linear features of the frequency-domain signals. L2 regularization helps prevent overfitting, especially when dealing with complex frequency-domain data.
[0058] (2) The second-layer convolution uses 32 filters, with a kernel size of 3, a stride of 1, and the activation function uses GeLU, which can better capture the non-linear features of the frequency-domain signal. L2 regularization helps prevent overfitting, especially when dealing with complex frequency-domain data.
[0059] (3) The third-layer convolution uses 32 filters, with a kernel size of 3, a stride of 1, and the activation function uses GeLU, which can better capture the non-linear features of the frequency-domain signal. L2 regularization helps prevent overfitting, especially when dealing with complex frequency-domain data.
[0060] (4) The fourth-layer convolution uses 64 filters, with a kernel size of 3, a stride of 1, and the activation function uses GeLU, which can better capture the non-linear features of the frequency-domain signal. L2 regularization helps prevent overfitting, especially when dealing with complex frequency-domain data.
[0061] Step 4, Time-frequency feature fusion: Time-frequency feature fusion is the process of combining time-domain and frequency-domain information to enhance the model's ability to understand and model different features in the signal.
[0062] In electroencephalogram (EEG) signal analysis, time-frequency feature fusion can simultaneously capture the time-domain changes and frequency-domain patterns of the signal, which is very effective for tasks such as identifying complex brain activities or emotional states. The specific process is as follows:
[0063] (1) Perform one-dimensional convolution on the original input time-domain EEG signal to extract time-domain features. The convolution uses 64 filters, with a kernel size of 3, a stride of 1, and the activation function uses GeLU, which can better capture the non-linear features of the frequency-domain signal. L2 regularization helps prevent overfitting.
[0064] (2) Perform one-dimensional convolution on the frequency-domain EEG signal in Step 1 to extract frequency-domain features. This step performs further non-linear mapping on the frequency-domain features to help extract deeper features in the frequency domain. The convolution uses 64 filters, with a kernel size of 3, a stride of 1, and the activation function uses GeLU, which can better capture the non-linear features of the frequency-domain signal. L2 regularization helps prevent overfitting.
[0065] (3) Time-frequency feature fusion based on the attention mechanism. Time-Frequency Attention Fusion based on the attention mechanism automatically selects and weights different time and frequency features by introducing the self-attention mechanism, thereby enhancing the model's performance. The specific steps are as follows:
[0066] ① Let the time-domain encoding obtained in Step 2 be T X ∈R N×D, the frequency-domain encoding obtained in step 3 is F X ∈R N×D , the feature obtained in (1) of step 4 is T T ∈R N×D , the feature obtained in step 4(2) is F F ∈R N×D . Use the self-attention mechanism to fuse the features of T X and F F as well as F X and T T . The following takes the fusion of T X and F F as an example.
[0067] ② Perform linear transformation to obtain the time-frequency features Query (Q), Key (K), and Value (V). The specific formulas are as follows:
[0068]
[0069] Among them
[0070] ③ Calculate the attention scores in the time domain and frequency domain: According to the self-attention mechanism, the attention scores in the time domain and frequency domain are calculated through the similarity between Query and Key. The commonly used method is to calculate the dot product and then perform scaling processing:
[0071] Attention score in the time domain:
[0072]
[0073] Attention score in the frequency domain:
[0074]
[0075] Among them, d is the dimension of Query and Key, usually a preset hyperparameter.
[0076] ④ Use Softmax to calculate the probability distribution of the scores to obtain the attention weights:
[0077] Attention weight in the time domain:
[0078]
[0079] Attention weight in the frequency domain:
[0080]
[0081] ⑤ Use the attention weights to add Value:
[0082] Weighted Value output in the time domain:
[0083]
[0084] Weighted Value output in the frequency domain:
[0085]
[0086] ⑥ Weighted output fusing time domain and frequency domain:
[0087] FO TF = α T × Output T + α F × Output F (14)
[0088] Where, α T and α F are learned weighted coefficients.
[0089] Similarly, F X and T T are also fused through the above process, and finally the fused output is FO FT .
[0090] Step 5, calculate the self - contrast loss between FO TF and FO FT . The self - contrast loss generates positive sample pairs by different augmentations or perturbations of the same input, and at the same time generates negative sample pairs from different inputs. Then, through contrastive learning, it maximizes the similarity of positive sample pairs and minimizes the similarity of negative sample pairs. This can effectively improve the quality of feature representation and enable the model to achieve good results in unsupervised or semi - supervised tasks. FO TF and F O FT exactly belong to the sample pairs generated from different inputs, and their similarity is judged through contrastive learning. The specific calculation process is as follows:
[0091] (1) Perform L2 regularization on FO TF and FO FT respectively. L2 regularization scales each vector to unit length, that is, ensures its norm is 1.
[0092] (2) Calculate the cosine similarity between FO TF and FO FT :
[0093]
[0094] (3) Calculate the contrast loss:
[0095]
[0096] Step 6, emotion detection. Feed FOTF and FO FT After fusing with FO, the result is globally average pooled, so that the temporal information is compressed into a feature vector of a fixed length. Finally, a fully connected layer is used to generate the final output y_pred for the sentiment detection task. The specific process is as follows:
[0097] (1) Connect and fuse FO TF and FO FT directly, and the formula is as follows:
[0098] FO = concatenate(FO TF , O FT ) = [FO FT , O FT (17)
[0099] (2) Perform average pooling on the FO vector to obtain the vector AFO.
[0100] (3) Use the fully connected layer to output the prediction result y_pred, and the formula is as follows:
[0101] y_pred = σ(AFO * W + b) (18)
[0102] where σ is the Sigmoid function.
[0103] (4) Prediction loss function: The loss function between the predicted classification result y_pred and the true classification result y_true is calculated using cross-entropy, and the formula is as follows:
[0104]
[0105] where N is the number of samples, cnum is the number of sentiment categories, y_pred i,m is the predicted probability that the i-th sample belongs to the m-th class: y_true i,m is the one-hot encoded value (taking values of 0 or 1) of the true label of the i-th sample in the m-th class.
[0106] An electroencephalogram signal sentiment detection system, comprising:
[0107] A data acquisition module for acquiring electroencephalogram signal data;
[0108] A sentiment detection module for inputting the electroencephalogram signal data into a sentiment detection model for prediction classification to obtain a sentiment classification result; wherein, the sentiment detection model includes a time-frequency feature extraction module, an attention mechanism model, a self-supervised contrast learning module, a pooling module, and a fully connected prediction module connected in sequence.
[0109] An electronic device includes a memory and a processor. The memory is used to store a computer program, and the processor runs the computer program to enable the electronic device to execute the described electroencephalogram signal emotion detection method.
[0110] A computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, it implements the described electroencephalogram signal emotion detection method.
[0111] As described above, only the preferred specific implementation manners of the present application are provided, but the protection scope of the present application is not limited thereto. Any changes or substitutions that can be easily thought of by those skilled in the art within the technical scope disclosed by the present application should be covered by the protection scope of the present application. Therefore, the protection scope of the present application should be subject to the protection scope of the claims.
Claims
1. A method for emotion detection of electroencephalogram signals, characterized in that, Including: Obtaining electroencephalogram (EEG) signal data; Inputting the EEG signal data into an emotion detection model for prediction and classification to obtain an emotion classification result; wherein, the emotion detection model includes a time-frequency feature extraction module, an attention mechanism model, a self-supervised contrast learning module, a pooling module, and a fully-connected prediction module connected in sequence.
2. The method for detecting emotion from EEG signals according to claim 1, wherein, The training process of the emotion detection model specifically includes: Obtaining training data, where the training data includes EEG signal training data and corresponding emotion classification results; Constructing an initial emotion detection model, inputting the training data into the initial emotion detection model for prediction and classification, and training with the goal of minimizing the loss between the initial training result after prediction and classification and the emotion classification result corresponding to the EEG signal training data to obtain a trained emotion detection model.
3. The electroencephalogram signal emotion detection method according to claim 1, wherein The processing process of the emotion detection model specifically includes: Inputting the EEG signal data into the time-frequency feature extraction module, performing wavelet transform on the EEG signal data to obtain the frequency-domain signal corresponding to the EEG signal data; encoding the EEG signal data using a parallel multi-scale one-dimensional convolutional layer to obtain a time-domain feature encoding; encoding the frequency-domain signal using a 3-layer one-dimensional convolutional layer to obtain a frequency-domain feature encoding; performing one-dimensional convolution on the EEG signal data and the frequency-domain signal respectively to obtain an initial time-domain feature and an initial frequency-domain feature; Applying a self-attention mechanism to fuse the initial time-domain feature and the frequency-domain feature encoding to obtain a first time-frequency feature, and fusing the initial frequency-domain feature and the time-domain feature encoding to obtain a second time-frequency feature; Connecting and fusing the first time-frequency feature and the second time-frequency feature to obtain a fused feature vector; Performing global average pooling on the feature vector to obtain a compressed feature vector; Inputting the compressed feature vector into the fully-connected prediction module and outputting a classification result through a softmax layer.
4. A method for detecting emotional state from EEG signals according to claim 3, wherein, The performing wavelet transform on the EEG signal data specifically includes: Applying continuous wavelet transform to each channel of each sample in the EEG signal data to convert the EEG signal from the time domain to the frequency domain and generate a spectrogram; using a Morlet wavelet as the mother wavelet during the continuous wavelet transform process.
5. The method for detecting emotional state from EEG signals according to claim 3, characterized in that, It also includes inputting each time-frequency feature into the self-supervised contrast learning module, calculating a contrast loss according to the cosine similarity, and optimizing the emotion detection model using the contrast loss.
6. An electroencephalogram signal emotion detection system, characterized in that, Including: A data acquisition module for obtaining EEG signal data; An emotion detection module for inputting the EEG signal data into an emotion detection model for prediction and classification to obtain an emotion classification result; wherein, the emotion detection model includes a time-frequency feature extraction module, an attention mechanism model, a self-supervised contrast learning module, a pooling module, and a fully-connected prediction module connected in sequence.
7. An electronic device, characterized in that, Including a memory and a processor, the memory is used to store a computer program, and the processor runs the computer program to enable the electronic device to execute an EEG signal emotion detection method according to any one of claims 1-5.
8. A computer-readable storage medium, characterized in that, It stores a computer program, and when the computer program is executed by a processor, it implements an electroencephalogram signal emotion detection method according to any one of claims 1-5.
Citation Information
Patent Citations
Multi-feature emotion electroencephalogram recognition model establishment method and device based on graph convolution
CN116919422A
Emotional electroencephalogram classification method based on self-supervised comparative learning
CN117204864A
Emotion recognition method based on electroencephalogram signals, computer equipment and storage medium
CN117562542A
Physiological signal self-supervision representation learning method and system based on time-frequency reconstruction
CN119089378A
Fusion technology of image and EEG signal for real-time emotion recognition
KR102275436B1