Electroencephalogram emotion recognition method based on attention mechanism and multi-feature fusion
By adopting an attention mechanism and multi-feature fusion method in EEG emotion recognition technology, the time domain, frequency domain and spatial characteristics of EEG signals are extracted, and the problem of insufficient robustness and generalization ability in the existing technology is solved, and more accurate emotion classification is achieved.
Patent Information
- Application Number
- CN202510155814.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-12
- Publication Date
- 2025-06-03
AI Technical Summary
The existing EEG emotion recognition technology has insufficient robustness and generalization ability in cross-individual experiments. Single-dimensional feature extraction cannot fully reflect emotional state, and it is difficult to capture the dynamic changes in time, frequency and spatial of EEG signals.
The EEG emotion recognition method based on attention mechanism and multi-feature fusion is adopted, and the time domain, frequency domain and airspace features are extracted and fusion through the TSS-ENet network model. Combined with the improved Attention module and the sparse multi-head attention module, the global context features of the time series are extracted, and spatial features are extracted through the dynamic graph convolution network.
It improves the emotional classification performance, enhances the importance of spectrum characteristics, captures the dynamic changes of EEG signals over time, and dynamically extracts the interaction relationships between brain regions, improving the accuracy of emotional classification results.
Smart Images

Figure CN120086683A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of emotion recognition, and in particular to an electroencephalogram emotion recognition method based on an attention mechanism and multi-feature fusion. Background Art
[0002] In recent years, emotion recognition based on electroencephalogram (EEG) has attracted wide attention. Since emotion is an important mental state of human beings, it not only affects human daily behaviors and decisions, but also is closely related to mental health and brain diseases. As a non-invasive brain activity monitoring means, EEG can record the changes in electrical signals of different regions of the brain in real time. Its high temporal resolution gives it unique advantages in emotion recognition research, that is, it can directly reflect brain nerve activities, avoiding the interference of subjective factors and having the advantage of objectivity. EEG emotion recognition has been widely applied in fields such as mental health monitoring, human-computer interaction, intelligent medicine, and virtual reality.
[0003] EEG signals contain rich time-domain, frequency-domain, and spatial information. Feature extraction in a single dimension often makes it difficult to comprehensively reflect the emotional state; there are significant differences in EEG patterns among different individuals, resulting in poor robustness and generalization ability of the model in cross-individual experiments. Existing single-dimensional feature extraction, such as directly using time series features (such as mean, variance, waveform features, etc.), requires in-depth understanding of the data by designing to extract features. The extracted feature dimensions are limited and cannot adapt to complex tasks; extracting the spectral characteristics of signals based on Fourier transform or wavelet transform can reflect the frequency information related to emotions, but lacks the ability to model time and space dynamics; emotion recognition methods based on EEG spatial features mainly consider the spatial position relationship between electrodes and reconstruct electroencephalograms using electrode spatial information. The spatial features mainly reflect the dependence between channels, while the emotional state often requires frequency-domain information to characterize. Modeling the change of signals over time cannot be achieved only by spatial features, especially since EEG signals are strong time series. Summary of the Invention
[0004] In order to solve the above technical problems, the purpose of the present invention is to provide an electroencephalogram emotion recognition method based on an attention mechanism and multi-feature fusion, and the specific technical solution adopted is as follows:
[0005] An embodiment of the present invention provides an electroencephalogram emotion recognition method based on an attention mechanism and multi-feature fusion. The method includes the following steps:
[0006] Obtain electroencephalogram signal data, preprocess the electroencephalogram signal data to obtain preprocessed electroencephalogram signal data;
[0007] Taking the preprocessed electroencephalogram signal data as input samples, through the constructed TSS-ENet network model, extract and fuse the time-domain, frequency-domain, and spatial-domain features of the input samples, and output target high-dimensional features;
[0008] Send the target high-dimensional features into a classifier, and the classifier determines the emotion category of the target high-dimensional features and outputs an emotion classification result.
[0009] Further, the process of taking the preprocessed electroencephalogram signal data as input samples, through the constructed TSS-ENet network model, extracting and fusing the time-domain, frequency-domain, and spatial-domain features of the input samples, and outputting an emotion classification result includes:
[0010] Taking the preprocessed electroencephalogram signal data as input samples, and using a Transformer with an improved attention mechanism to extract the global context features of the time series of the input samples;
[0011] Extract the spectral features of the input samples, where the spectral features include power spectral density and differential entropy; enhance the power spectral density and differential entropy through weighted fusion, and then connect them with the global context features to obtain a comprehensive feature representation;
[0012] Input the comprehensive feature representation into a dynamic graph convolutional network to extract target high-dimensional features.
[0013] Further, the process of using a Transformer with an improved attention mechanism to extract the global context features of the time series of the input samples includes:
[0014] The input samples first pass through an improved Attention module to obtain the output features of the improved Attention module; the improved Attention module includes a channel attention module and a sparse multi-head attention module;
[0015] Add the output features of the Attention module and the input features through residual connection, and then perform normalization processing on the added features to obtain the normalized result;
[0016] Pass the normalized result to a feed-forward network for feature enhancement processing to obtain the output features of the feed-forward network; where the feed-forward network consists of two fully connected layers and a non-linear activation function;
[0017] Perform residual connection and normalization processing on the output features of the feed-forward network to obtain the global context features of the time series.
[0018] Further, the process of extracting the spectral features of the input samples includes:
[0019] Apply a band - pass filter to the input sample and extract signals in a preset number of target frequency bands;
[0020] Calculate the differential entropy and power spectral density of the signals in each target frequency band.
[0021] Further, the enhancement of the power spectral density and differential entropy through weighted fusion includes:
[0022] Use SE block fusion instead of concatenation for the differential entropy and power spectral density of the signals in all target frequency bands to form an output feature matrix; where the SE block consists of three basic steps: Squeeze, Excitation, and Scale.
[0023] Further, the implementation process of the channel attention module includes:
[0024] The input features go through two pooling operations, max - pooling and average - pooling, to generate two types of pooled features respectively;
[0025] Input the pooled features into a fully - connected layer network to generate channel attention weights;
[0026] Aggregate the two types of pooled features through element - wise addition, combine with the channel attention weights, and generate a channel attention map;
[0027] Multiply the input features and the channel attention map channel - by - channel to obtain the output features of the channel attention module.
[0028] Further, the implementation process of the sparse multi - head attention module includes:
[0029] Perform linear transformations on the input sample using three different weight matrices to obtain a query matrix, a key matrix, and a value matrix;
[0030] The query selection and sampling process downsamples the query matrix and the key matrix to generate a downsampled sparse matrix and a key matrix;
[0031] Perform dot - product operations on the downsampled sparse matrix and the key matrix, and perform normalization processing to generate a sparse attention map;
[0032] Combine the sparse attention map and the value matrix, and combine with the input sample through a residual connection to obtain the output features of the sparse multi - head attention module.
[0033] The present invention has the following beneficial effects:
[0034] The present invention provides an EEG emotion recognition method based on the attention mechanism and multi-feature fusion. This method effectively extracts features from EEG data, improving the emotion classification performance. A TSS-ENet (Temporal-Spectral-Spatial EmotionNet) model is constructed. This model effectively enhances the importance of spectral features, helps capture the dynamic changes of EEG signals over time, and is conducive to dynamically extracting the interaction relationships between brain regions to construct structured spatial features, thereby improving the accuracy of emotion classification results. BRIEF DESCRIPTION OF THE DRAWINGS
[0035] In order to more clearly illustrate the technical solutions and advantages in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the following drawings are only some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.
[0036] Figure 1 It is a flowchart of an EEG emotion recognition method based on the attention mechanism and multi-feature fusion according to an embodiment of the present invention;
[0037] Figure 2 It is a flowchart of the TSS-ENet network model in the embodiment of the present invention;
[0038] Figure 3 It is the implementation steps of step S2 in the embodiment of the present invention;
[0039] Figure 4 It is the structure diagram of the Att-Transformer model in the embodiment of the present invention;
[0040] Figure 5 It is the structure diagram of the improved Attention module in the embodiment of the present invention;
[0041] Figure 6 It is the structure diagram of the SE block in the embodiment of the present invention;
[0042] Figure 7 It is the structure diagram of the DGCN model and the classification module in the embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0043] In order to further explain the technical means and effects adopted by the present invention to achieve the predetermined invention purpose, the specific implementation methods, structures, features and effects of the technical solutions proposed by the present invention are described in detail below in conjunction with the accompanying drawings and preferred embodiments. In the following description, different "one embodiment" or "another embodiment" does not necessarily refer to the same embodiment. In addition, specific features, structures or characteristics in one or more embodiments may be combined in any suitable form.
[0044] Unless otherwise defined, all technical and scientific terms used in this article have the same meanings as those commonly understood by technicians in the technical field of the present invention. The EEG signal data involved in this article have been obtained with full consent and authorization, and the collection, use and processing of relevant information must comply with relevant laws, regulations and standards of relevant countries and regions.
[0045] The application scenarios targeted by the present invention may be:
[0046] Single-dimensional feature extraction methods have their own advantages and obvious limitations. Using only time domain or frequency domain features cannot reflect the full picture of EEG signals, while feature extraction methods that ignore spatial information are difficult to capture complex association patterns between brain regions. The neurophysiological mechanism of emotion involves complex interactions in the dimensions of time, frequency, and brain space. Therefore, it is of great significance to comprehensively extract time domain, frequency domain, and spatial features from EEG signals.
[0047] This embodiment provides an EEG emotion recognition method based on attention mechanism and multi-feature fusion. Figure 1 As shown, the following steps are included:
[0048] S1, acquiring EEG signal data, preprocessing the EEG signal data, and obtaining preprocessed EEG signal data.
[0049] As a specific implementation, for the preprocessed EEG signal data, the acquisition step may include:
[0050] Specifically, in the DEAP dataset, 32 channels of EEG signal data from 32 subjects were collected according to the international 10-20 system and downsampled to 128 Hz. The data was subjected to baseline removal and 1s sample segmentation to obtain N samples of data, denoted as S = {S 1 ,S 2 ,…,S N},in, Let \(m\) represent the number of channels and \(p\) represent the number of sampling points. Finally, the sample data is subjected to Z-Score standardization to obtain the final preprocessed sample, that is, the preprocessed electroencephalogram (EEG) signal data. Among them, the processes of baseline removal, sample segmentation, and standardization are all prior arts and not within the scope of protection of the present invention, so no detailed description will be given here. In addition, both the number of channels and the number of subjects of the EEG signal data can be set by the implementer according to the specific actual situation, and no specific limitation is made here.
[0051] S2. Use the preprocessed EEG signal data as the input sample, and through the constructed TSS-ENet network model, extract and fuse the time-domain, frequency-domain, and spatial-domain features of the input sample, and output the target high-dimensional features.
[0052] Here, the target high-dimensional features refer to multi-dimensional features that fuse time-domain, frequency-domain, and spatial-domain features, which can effectively avoid the fact that it is often difficult to comprehensively reflect the emotional state due to single-dimensional feature extraction; the TSS-ENet network model is used to extract and fuse features of multiple dimensions, and the flowchart of the TSS-ENet network model is as Figure 2 shown.
[0053] The above step S2 can be implemented through Figure 3 the steps S21 to S23 shown as follows:
[0054] S21. Use the preprocessed EEG signal data as the input sample, and use the Transformer with the improved attention mechanism to extract the global context features of the time series of the input sample.
[0055] In this embodiment, the preprocessed EEG signal data is used as the sample input, and the time features of the EEG data are extracted through Att-Transformer. The structural diagram of the Att-Transformer model is as Figure 4 shown, and the expression for describing the process of the entire time feature extractor can be:
[0056] In the formula, \(f\) Temporal represents the time feature, and \(E\) Temporal represents the time feature extraction function. Here, the global context features of the time series of the input sample are the time features, and the improved Attention module includes a channel attention module and a sparse multi-head attention module.
[0057] It should be noted that for the selection of the sparse multi-head attention module to improve the Attention module, when the standard Transformer model performs self-attention calculation, its complexity is \(O(N)\) 2), where N is the sequence length, which shows that as the sequence length increases, the computational and memory requirements increase sharply. The sparse attention mechanism is mainly used to reduce the computational complexity and memory consumption in long-sequence tasks, thereby improving the training and inference efficiency of the model. Although the sparse attention mechanism reduces the amount of computation, it still retains sufficient context information, enabling the model to effectively capture long-range dependencies in the sequence. EEG signals are usually time series with high-frequency sampling and high time resolution, which results in a large amount of data with time steps in a single experimental record, and the sparse attention mechanism is suitable for long EEG signals.
[0058] For selecting the channel attention module to improve the Attention module, EEG signals are collected from multiple channels, and each channel corresponds to the activity of different regions of the brain. In different tasks or states, the importance of different channels may change dynamically. Treating all channels equally directly may waste computational resources or introduce irrelevant information, affecting the model performance. However, adding channel attention in parallel can dynamically allocate weights, highlight key channels, suppress redundant information, and at the same time make up for the deficiency that temporal attention cannot model the relationship between channels, enhancing the model's comprehensive capture ability of the spatial and temporal features of EEG signals, thereby improving the overall performance and task adaptability.
[0059] As a specific implementation manner, the above step S21 can be implemented through steps S211 to S214 (not shown in the figure):
[0060] S211, the input sample first passes through the improved Attention module to obtain the output features of the improved Attention module.
[0061] In this embodiment, assuming the input sample X, it first passes through the improved Attention module. The structural diagram of the improved Attention module is as Figure 5 shown, and the descriptive expression of its process can be:
[0062] X ′ = Attention(X); where X ′ represents the output features of the improved Attention module.
[0063] As a specific implementation manner, the steps for obtaining the output features of the improved Attention module include:
[0064] First, for the channel attention module (CA module) in the improved Attention module, the input features undergo two pooling operations, max pooling (MaxPool) and average pooling (AvgPool), to generate two types of pooled features respectively. The pooled features are input into a fully connected layer network to generate channel attention weights. The two types of pooled features are aggregated by element-wise addition, combined with the channel attention weights, to generate a channel attention map. The input features are multiplied with the channel attention map channel by channel to obtain the output features of the channel attention module.
[0065] Second, for the sparse multi-head attention module (SMHA module) in the improved Attention module, three different weight matrices are used for linear transformation of the input samples to obtain the query matrix (Q), key matrix (K), and value matrix (V). The query selection and sampling process downsamples the query matrix and the key matrix to generate the downsampled sparse matrix (Q′) and key matrix (K′). A dot product operation is performed on the downsampled sparse matrix and the key matrix, followed by Softmax normalization to generate a sparse attention map. The sparse attention map is combined with the value matrix and combined with the input samples through a residual connection to obtain the output features of the sparse multi-head attention module.
[0066] Among them, the expressions for the query matrix (Q), key matrix (K), and value matrix (V) can be: Q = XW Q , K = XW K , V = XW V , where W Q represents the weight matrix of the query matrix, W K represents the weight matrix of the key matrix, and W V represents the weight matrix of the value matrix.
[0067] Finally, the output features of the channel attention module and the output features of the sparse multi-head attention module are fused by element-wise addition.
[0068] S212, the output features of the Attention module and the input features are added through a residual connection, and then the added features are normalized to obtain the normalized result.
[0069] In this embodiment, the output of the Attention module and the previous input are added through a residual connection (Add), and then normalized (Norm) to stabilize the training process. The descriptive expression of the process can be:
[0070] H′ = LayerNorm(X ′ + X); where H′ represents the result after normalization.
[0071] S213, Pass the result after normalization to the feed - forward network for feature enhancement processing to obtain the output features of the feed - forward network.
[0072] In this embodiment, the feed - forward network usually consists of two fully - connected layers and a non - linear activation function, which can extract deeper features. The calculation formula for enhancing the expression of the model can be:
[0073] FFN(H′) = ReLU(H′W 1 + b 1 )W 2 + b 2 ; where FFN(H′) represents the output features of the feed - forward network, ReLU represents the activation function, W 1 , b 1 , W 2 and b 2 represent learnable parameters.
[0074] S214, Perform residual connection and normalization processing on the output features of the feed - forward network to obtain the global context features of the time series.
[0075] It should be noted that the above - mentioned improved Attention module and Feed Forward parts, as an encoding module, can be stacked N times to enhance the expression ability of the model. N can take an empirical value of 2.
[0076] S22, Extract the spectral features of the input sample. The spectral features include power spectral density and differential entropy; enhance the power spectral density and differential entropy through weighted fusion, and then connect them with the global context features to obtain the comprehensive feature representation.
[0077] In this embodiment, the pre - processed EEG data is used as the sample input. After passing through the spectral module, the spectral features of the EEG data are extracted. The expression for describing the process of the entire spectral feature extractor can be:
[0078] where f Spectral represents the spectral features, and E Spectral represents the spectral feature extraction function.
[0079] The expression for concatenating the time feature f Temporal and the spectral feature f Spectral can be:
[0080] where f TSDenote the comprehensive feature representation, i.e., the time-frequency features of the EEG signals.
[0081] The above step S22 can be implemented through steps S221 to S222 (not shown in the figure):
[0082] S221, extract the spectral features of the input samples.
[0083] Here, the spectral features mainly include the calculation of band-pass filtering, differential entropy, and power spectral density.
[0084] First, apply a band-pass filter to the input samples to extract the signals in a preset number of target frequency bands.
[0085] Specifically, apply a band-pass filter to the input signal X to extract the signals in five frequency bands: Delta (0.5Hz - 3Hz), Theta (4Hz - 7Hz), Alpha (8Hz - 13Hz), Beta (14Hz - 30Hz), and Gamma (31Hz - 50Hz).
[0086] Second, calculate the differential entropy and power spectral density of the signals in each target frequency band.
[0087] Specifically, the differential entropy of the signal is used to quantify the distribution characteristics of the signal in a specific frequency band, and the calculation formula can be:
[0088] In the formula, σ 2 represents the variance of the target frequency band, log represents the logarithmic function, e represents the natural constant, and DE(X) represents the differential entropy of the target frequency band.
[0089] The power spectral density of the signal is used to quantify the power distribution of the signal in a specific frequency band, and the calculation formula can be:
[0090] In the formula, f end represents the end frequency of the target frequency band, f start represents the start frequency of the target frequency band, and PSD(X) represents the power spectral density of the target frequency band.
[0091] S222, enhance the power spectral density and differential entropy through weighted fusion, and then connect with the global context features to obtain the comprehensive feature representation.
[0092] First, use SE block fusion instead of concatenation for the differential entropy and power spectral density of the signals in all target frequency bands to form an output feature matrix.
[0093] Specifically, the SE block consists of three basic steps: Squeeze, Excitation, and Scale. The structure diagram of the SE block is as Figure 6As shown
[0094] In the first step, the Squeeze operation reduces each two-dimensional matrix to a single real number by averaging, thereby capturing a certain global receptive field.
[0095] Assume the input is where z k represents the global information obtained after the Squeeze operation on X K , C represents the number of channels, F represents the feature dimension, and X k (i, j) represents the feature value at the i-th channel and the j-th spatial position in the k-th sample, and K represents the number of samples.
[0096] In the second step, the Excitation operation generates different weights for each feature map using W.
[0097] s = F ex (z, W) = σ(W 2 δ(W 1 z)); where s represents the weight vector generated by the Excitation operation, F ex represents the Excitation operation function, z represents the global feature vector obtained by the Squeeze operation, W represents the set of weight parameters in the Excitation operation, σ represents the Sigmoid function, δ represents the ReLU function, W 2 represents the weight matrix of the second fully connected layer, and W 1 represents the weight matrix of the first fully connected layer.
[0098] In the third step, the Scale operation multiplies the weights obtained from the second step by the original feature map to achieve reweighting.
[0099] where represents the reweighted feature map, F scale represents the Scale operation function, X k represents the original feature map, and s k represents the weight of the original feature map.
[0100] Secondly, the output feature matrix is used as the input of the multi-band spectral characteristics and connected with the global context features to obtain the comprehensive feature representation.
[0101] S23, the comprehensive feature representation is input into the dynamic graph convolutional network to extract the target high-dimensional features.
[0102] In this embodiment, the comprehensive feature representation is input into the DGCN to extract spatial features, and the target high-dimensional features of time-frequency-space features are obtained. Its expression can be:
[0103] In the formula, f TSS represents the target high-dimensional features, that is, represents time-frequency-space features, and E Spatial represents the spatial feature extraction function.
[0104] Specifically, the comprehensive feature representation is input into the Dynamic Graph Convolutional Network (DGCN). By learning the dynamic associations between features, more discriminative features are extracted. The structural diagrams of the DGCN model and the classification module are as Figure 7 shown.
[0105] Here, the electroencephalogram (EEG) signals have high dynamics and non-stationarity. Under different times, tasks, or emotional states, the functional connectivity relationships of brain regions will change. However, a fixed adjacency matrix cannot capture this dynamic change. The dynamic adjacency matrix can reflect the changes in these connection patterns in real time, thereby improving the accuracy of emotion classification. Therefore, dynamic graph convolution is needed to capture spatial features.
[0106] The graph convolutional neural network is a neural network layer, and the expression of the propagation method between its layers can be:
[0107] In the formula, represents the adjacency matrix, I N represents the identity matrix, is the degree matrix of , represents the features of the l-th layer, and W (l) represents the parameters of the l-th layer of the graph convolutional neural network; for the input layer, H (0) =X, and σ is the non-linear activation function.
[0108] To optimize the graph convolutional layer, first, a random initial adjacency matrix is initialized. Subsequently, during the model training process, the dynamic adjacency matrix will be updated according to the current node features and learnable parameters in each layer or each round of iteration. The goal of the update is to optimize the loss function so that the generated adjacency matrix can better capture the interaction relationships between nodes and enhance the feature extraction ability of the graph convolution operation. The iteration process will continue until the model training is completed, and finally an optimal adjacency matrix that matches the input features and can effectively represent the graph structure is obtained.
[0109] So far, this embodiment has obtained the target high-dimensional features based on the attention mechanism and multi-feature fusion.
[0110] S3. Feed the target high-dimensional feature into a classifier, and the classifier determines the emotion category of the target high-dimensional feature and outputs an emotion classification result.
[0111] In this embodiment, f TSS After being flattened, it is input into a fully connected layer for binary classification to obtain a classification result. The process can be expressed as: That is, the target high-dimensional feature output by DGCN is fed into a classifier, and the classifier determines the emotion category of the feature and outputs an emotion classification result.
[0112] It should be noted that the model uses the Cross-Entropy Loss function as the optimization objective to measure the difference between the model prediction result and the actual label, and gradually minimizes the value of the loss function to optimize the model performance.
[0113] So far, this embodiment has obtained an emotion classification model with high accuracy.
[0114] To verify the advantages of the present invention, the method proposed by the present invention is compared with traditional methods. The comparison results of the experimental results on the DEAP dataset are shown in Table 1:
[0115] Table 1
[0116]
[0117] As can be seen from Table 1, the present invention can achieve relatively accurate emotion recognition results.
[0118] The Correlation-Aware Network (CAN) in Table 1 extends the attention-based recurrent neural network and integrates the correlation between electroencephalogram and eye movement extraction signals into its attention mechanism. Although CAN combines multi-modal information, it has weak modeling of the spatial and frequency domain features of electroencephalogram signals.
[0119] The Multi-modal Residual LSTM (MMResLSTM) constructs 4 LSTM layers for electroencephalogram signals and peripheral physiological signals respectively. By sharing the weights of each modality on each LSTM layer, it learns the temporal correlation between different modalities and enhances the ability to extract temporal features. However, it has insufficient processing of the spatial dependence relationship of electroencephalogram signals, and there is still room for performance improvement. Compared with the MMResLSTM model, the accuracy of the present invention is improved by about 5%.
[0120] The Parallel Convolutional Recurrent Neural Network (PCRNN) combines CNN and LSTM to extract spatio-temporal features for classification, but ignores the importance of the frequency domain information in electroencephalogram signals (such as the significant role of the energy distribution in different frequency bands in emotion recognition), resulting in insufficient representation ability of the model for emotion states.
[0121] The 4D-CRNN combines the frequency-domain, spatio-temporal information of electroencephalogram (EEG) signals, and transforms the differential entropy features of different channels into a 4D structure to train a deep model. Although certain achievements have been made, the model fails to effectively introduce an attention mechanism, resulting in insufficient attention to important features when processing complex EEG signals, which may reduce the recognition ability of the model.
[0122] The Regional Asymmetric Convolutional Neural Network (RACNN) can not only learn the regional information between adjacent channels but also learn the asymmetric differences between the two hemispheres, and it has achieved high performance in spatial feature modeling. However, due to the weak capture of time-domain features, there are still certain performance limitations.
[0123] The Dynamic Graph Convolutional Network (DGCNN) has significant advantages in spatial feature modeling, but its ability to fuse time series and frequency-domain information is limited, so its performance is slightly inferior. In addition to combining dynamic graph convolution, this invention also fuses time-domain and frequency-domain features, and the more comprehensive extraction of features improves the performance of the model.
[0124] CR-GCN uses graph convolution to capture local and global interactions between EEG channels. It focuses on capturing spatial interactions between channels but lacks in-depth modeling of time-dimensional features, which may lead to insufficient ability to capture dynamic emotional changes.
[0125] Generally speaking, compared with the above traditional methods, TSS-ENet not only fully integrates the time-domain, frequency-domain, and spatial features of EEG signals but also extracts key features through sparse attention and channel attention modules, significantly improving the performance of emotion recognition and providing a new direction for emotion recognition research.
[0126] To further verify the effect of this invention, ablation experiments on the DEAP dataset were conducted, as shown in Table 2:
[0127]
[0128] According to the adjustment methods of different modules in the ablation experiment, the changes in model performance can be explained as follows:
[0129] Sparse attention is replaced by the original multi-head attention:
[0130] In the experiment of sparse attention, the sparse self-attention mechanism in the current model is replaced by the traditional multi-head self-attention, and the efficiency drops to 96.79% and the arousal rate drops to 98.18%. This shows that sparse attention effectively reduces unnecessary information processing by focusing on key points in the time series and improves the efficiency of the model.
[0131] Remove channel attention:
[0132] In the experiment of channel attention, directly removing the channel attention mechanism in the model, the efficiency drops to 97.19%, and the arousal rate drops to 98.04%. This indicates that channel attention can dynamically adjust the importance of different EEG channels, which helps to strengthen the extraction of emotion-related features. Its removal weakens the model's ability to model the correlations between different channels, thus affecting the performance.
[0133] Replace the SE block with simple concatenation:
[0134] Replacing the mechanism of using the SE block for dynamically fusing frequency-domain features with simple feature concatenation, the efficiency drops to 96.69%, and the arousal rate drops to 97.50%. This shows that the SE block can effectively highlight the spectrum information related to emotions by adjusting the weights of frequency-domain features, improving the quality of feature fusion, while simple concatenation cannot effectively exploit the importance of frequency-domain information, resulting in insufficient utilization of spectral features by the model.
[0135] The best performance of the complete model:
[0136] The complete TSS-ENet model achieves the best performance with an efficiency of 97.62% and an arousal rate of 98.38% through the synergistic effect of sparse attention, channel attention, and the SE block.
[0137] The importance of each module is verified through ablation. The optimized designs of each module complement each other. By ensuring the integrity of each module, the effect of emotion recognition can be comprehensively improved.
[0138] So far, the above experimental results show that replacing or removing the key modules of the model will lead to a significant decline in performance, verifying the rationality and necessity of the EEG emotion recognition method based on attention mechanism and multi-feature fusion proposed in the present invention in multi-dimensional feature extraction and fusion.
[0139] The above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that: they can still modify the technical solutions recorded in the foregoing embodiments, or perform equivalent replacements on some of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the scope of the technical solutions of the embodiments of the present invention, and all should be included in the protection scope of the present invention.
Claims
1. A method for EEG emotion recognition based on attention mechanism and multi-feature fusion, characterized in that: The following steps are involved: Acquiring EEG signal data, and preprocessing the EEG signal data to obtain preprocessed EEG signal data; The preprocessed EEG signal data is used as an input sample, and the time domain, frequency domain and spatial domain features of the input sample are extracted and fused through the constructed TSS-ENet network model to output the target high-dimensional features; The target high-dimensional features are sent to a classifier, and the classifier determines the emotion category of the target high-dimensional features and outputs the emotion classification result.
2. According to claim 1, the method for EEG emotion recognition based on attention mechanism and multi-feature fusion is characterized in that: The preprocessed EEG signal data is used as an input sample, and the input sample is subjected to time domain, frequency domain and spatial domain feature extraction and fusion through the constructed TSS-ENet network model, and the emotion classification result is output, including: Taking the preprocessed EEG signal data as input samples, using the improved attention mechanism Transformer to extract global context features of the time series of the input samples; Extracting the frequency spectrum features of the input sample, the frequency spectrum features including power spectrum density and differential entropy; enhancing the power spectrum density and differential entropy by weighted fusion, and then connecting them with the global context features to obtain a comprehensive feature representation; The comprehensive feature representation is input into a dynamic graph convolutional network to extract target high-dimensional features.
3. The method for EEG emotion recognition based on attention mechanism and multi-feature fusion according to claim 2 is characterized in that: The Transformer using the improved attention mechanism extracts global context features of the time series of the input sample, including: The input sample first passes through the improved Attention module to obtain the output features of the improved Attention module; the improved Attention module includes a channel attention module and a sparse multi-head attention module; The output features of the Attention module and the input features are added through a residual connection, and the added features are normalized to obtain a normalized result; Passing the normalized result to a feedforward network for feature enhancement processing to obtain output features of the feedforward network; wherein the feedforward network is composed of two fully connected layers and a nonlinear activation function; The output features of the feedforward network are subjected to residual connection and normalization processing to obtain global context features of the time series.
4. The method for EEG emotion recognition based on attention mechanism and multi-feature fusion according to claim 2, characterized in that: The extracting the frequency spectrum feature of the input sample comprises: Applying a bandpass filter to the input samples to extract signals of a preset number of target frequency bands; Calculate the differential entropy and power spectral density of the signal in each target frequency band.
5. The method for EEG emotion recognition based on attention mechanism and multi-feature fusion according to claim 4, characterized in that: The step of enhancing the power spectrum density and differential entropy by weighted fusion includes: The differential entropy and power spectral density of the signals in all target frequency bands are fused using SE block instead of concatenated to form an output feature matrix; the SE block consists of three basic steps: Squeeze, Excitation and Scale.
6. The method for EEG emotion recognition based on attention mechanism and multi-feature fusion according to claim 2, characterized in that: The implementation process of the channel attention module includes: The input features are subjected to two pooling operations: maximum pooling and average pooling, generating two pooling features respectively; Input the pooled features into a fully connected layer network to generate channel attention weights; Aggregating the two pooling features by element-by-element addition and combining the channel attention weights to generate a channel attention map; The input feature is multiplied by the channel attention map channel by channel to obtain the output feature of the channel attention module.
7. The method for EEG emotion recognition based on attention mechanism and multi-feature fusion according to claim 2, characterized in that: The implementation process of the sparse multi-head attention module includes: For the input samples, three different weight matrices are used to perform linear transformation to obtain the query matrix, key matrix and value matrix; The query selection and sampling process downsamples the query matrix and the key matrix to generate a downsampled sparse matrix and a key matrix; Performing a dot product operation on the downsampled sparse matrix and the key matrix, and performing normalization processing to generate a sparse attention map; The sparse attention map is combined with the value matrix, and combined with the input sample through a residual connection to obtain the output features of the sparse multi-head attention module.
Citation Information
Cited By
CNN-Transform-based EEG emotion recognition method, device and system
CN120724310A
EEG signal processing method and device based on time-frequency fusion and multi-scale cross attention
CN121167639A
Electroencephalogram signal processing method and device based on time-frequency fusion and multi-scale cross attention
CN121167639B
Two-way emotion recognition method based on electroencephalogram micro-state features
CN121278476A
Motion imagination coupling system and method based on emotion prediction
CN121365240A