Brain electrical emotion recognition method and system based on multi-branch feature fusion
Through a multi-branch feature fusion network, combining channel attention weighting, dynamic graph convolution and adaptive Transformer technology, multiple features of EEG signals are extracted and fused, and adapted to different subjects through fine-tuning of adapters, solving the problems of low accuracy of EEG emotion recognition and difficulty in fusion of multiple features in the prior art, achieving more efficient emotion recognition.
Patent Information
- Application Number
- CN202510472790.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-16
- Publication Date
- 2025-05-13
- Estimated Expiration
- 2045-04-16
AI Technical Summary
Existing EEG emotion recognition technology has the risk of high computational complexity or overfitting when processing high-dimensional EEG signals, resulting in low accuracy of emotion recognition and difficulty in effectively integrating multiple features.
A multi-branch feature fusion network is adopted to extract and fuse the time, frequency and spatial features of EEG signals through channel attention-weighted spatial convolutional branches, dynamic graph convolution network branches and adaptive Transformer feature fusion networks, and quickly adapt to different subjects and data changes through fine-tuning by adapter.
The spatial and temporal feature representation ability and model generalization ability of EEG signals are improved, the emotional classification performance is enhanced, the risk of overfitting is reduced, and more accurate EEG emotional recognition across subjects is achieved.
Smart Images

Figure SMS_43 
Figure SMS_44 
Figure SMS_64
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of cross-subject EEG emotion recognition, and specifically, relates to an EEG emotion recognition method and system based on multi-branch feature fusion. Background Art
[0002] Emotions are the basis of human experience and play an important role in happiness, behavioral performance, social interaction, interpersonal communication, decision-making and cognitive functions. Affective computing technology enables computers to recognize, understand, represent, adapt and give feedback. Among them, emotional computing based on physiological signals has become a hot topic. EEG signals are widely used in the field of emotional computing due to their portability and high temporal resolution. Brain-computer interface (BCI) collects brain signals induced by emotional stimulation, performs pattern recognition, analyzes the type of brain emotion perception, and converts them into external instructions to achieve human-computer emotional interaction.
[0003] In recent years, recent advances in electrode technology and machine learning have improved electroencephalogram (EEG) analysis for emotion recognition. However, the inherent non-stationarity of EEG signals can vary from individual to individual or over time, which poses a challenge to developing a universally applicable model for all subjects. Simply transferring may lead to misclassification of the pre-trained model, resulting in poor classification performance for the target subject.
[0004] At the same time, EEG signals are characterized by high dimensionality and complexity, and it is not easy to extract effective emotional features from them. The extraction of time domain, frequency domain and spatial features requires complex algorithms and techniques, and the fusion of different features is also challenging. Some studies have tried to use deep learning methods, such as convolutional neural networks (CNN) and recurrent neural networks (RNN), to automatically extract and fuse different features. Although these methods have achieved good results, these methods have the risk of high computational complexity or overfitting when processing high-dimensional EEG signals, resulting in low accuracy in emotion recognition. Further research is still needed on the fusion of multiple features of EEG signals. Summary of the invention
[0005] In order to solve the problem of low emotion recognition accuracy in existing emotion feature extraction methods, the present invention provides an EEG emotion recognition method based on multi-branch feature fusion.
[0006] In order to achieve the above object, the present invention provides the following technical solutions: On the one hand, the present invention provides an EEG emotion recognition method based on multi-branch feature fusion, which specifically includes the following steps: Step 1, obtaining EEG data of multiple subjects; Step 2, preprocessing the EEG data of multiple subjects to obtain preprocessed EEG data of each subject; Step 3, extract the features of five frequency bands of the preprocessed EEG data of each subject, the five frequency bands are delta, theta, alpha, beta and gamma; Step 4: The features of the five frequency bands of the first subject are used as the target domain, and the features of the five frequency bands of the remaining subjects are used as the source domain; Step 5: Extract the time domain features and frequency domain features of the target domain, which are DE features and PSD features respectively; take 10% to 20% of the DE features and PSD features as adapter fine-tuning data, and use the remaining DE features and PSD features as test sets; extract the DE features and PSD features of the source domain as training sets; Step 6: Input the test set and the training set into a multi-branch feature fusion network together. Specifically, the multi-branch feature fusion network includes a channel attention weighted spatial convolution branch, a first dynamic graph convolution network branch and a second dynamic graph convolution network branch with the same structure; input the DE features in the test set and the training set into the first dynamic graph convolution network branch and the channel attention weighted spatial convolution branch respectively, and input the PSD features in the test set and the training set into the second dynamic graph convolution network branch; Step 7: The first dynamic graph convolutional network branch processes the input DE features to obtain time domain features. The second dynamic graph convolutional network branch processes the input PSD features to obtain frequency domain features ; The time domain features and frequency domain characteristics Connect them to get the time domain-frequency domain characteristics ; The channel attention weighted spatial convolution branch adopts the channel-adaptive weights of the channel attention weighted learning DE features; Finally, the weighted multi-channel EEG signal is subjected to spatial convolutional blocks. Processing is performed to finally obtain spatial features ; Step 8: The time domain-frequency domain features obtained in step 7 Input the adaptive Transformer feature fusion network with adapter fine-tuning to obtain high-level time-frequency domain features H A ; Step 9: Advanced time-frequency features and spatial characteristics Connect to get time-frequency-space features , and then the time domain-frequency domain-space features Input the adaptive Transformer feature fusion network with adapter fine-tuning to get the new output ; Step 10: Input the adapter fine-tuning data obtained in step 5 into the adapter fine-tuning module of the adaptive Transformer feature fusion network with adapter fine-tuning The result is used as the output of the adapter fine-tuning data H T ; Step 11: Output the new and adapters to fine-tune the output of data H T Through the fully connected layer and Softmax Function, respectively obtain The predicted sentiment probability and H T The predicted emotion probability of The maximum value of the predicted sentiment probability is taken as The classification result is The predicted sentiment label of H T The maximum value of the predicted emotion probability is taken as The classification result is The predicted sentiment label of Step 12, based on the predicted emotion probability, calculate the cross entropy loss function value of the adapter fine-tuning data and the training set; Step 13: Repeat steps 6 to 12 to obtain the trained model of the current subject by minimizing the cross entropy loss function value between the model prediction and the actual label; In step 14, the next subject is used as the target domain and the remaining subjects are used as source domains. Steps 5 to 13 are repeated until each subject is used as the target domain once, and finally a trained model for each subject is obtained.
[0007] On the other hand, the present invention provides an EEG emotion recognition system based on multi-branch feature fusion, which specifically includes the following modules: An EEG data acquisition module is used to acquire EEG data of multiple subjects; A preprocessing module, used for preprocessing the EEG data of multiple subjects to obtain the preprocessed EEG data of each subject; The frequency band feature extraction module is used to extract the features of five frequency bands of the preprocessed EEG data of each subject, and the five frequency bands are delta, theta, alpha, beta and gamma; A feature partitioning module, used to use the features of the five frequency bands of the first subject as target domains, and the features of the five frequency bands of the remaining subjects as source domains; The data set partitioning module is used to extract the time domain features and frequency domain features of the target domain, where the time domain features and frequency domain features are DE features and PSD features respectively; 10% to 20% of the DE features and PSD features are taken as adapter fine-tuning data, and the remaining DE features and PSD features are used as test sets; DE features and PSD features of the source domain are extracted as training sets; A data input module is used to input the test set and the training set into a multi-branch feature fusion network together, specifically: the multi-branch feature fusion network includes a channel attention weighted spatial convolution branch, a first dynamic graph convolution network branch and a second dynamic graph convolution network branch with the same structure; the DE features in the test set and the training set are respectively input into the first dynamic graph convolution network branch and the channel attention weighted spatial convolution branch, and the PSD features in the test set and the training set are input into the second dynamic graph convolution network branch; The feature fusion module is used to implement the following functions: the first dynamic graph convolutional network branch processes the input DE features to obtain time domain features The second dynamic graph convolutional network branch processes the input PSD features to obtain frequency domain features ; The time domain features and frequency domain characteristics Connect them to get the time domain-frequency domain features ; The channel attention weighted spatial convolution branch adopts the channel adaptive weights of the channel attention weighted learning DE features; finally, the spatial convolution block is used to weight the multi-channel EEG signal Processing is performed to finally obtain spatial features ; Adaptive feature fusion module 1 is used to integrate time domain and frequency domain features Input the adaptive Transformer feature fusion network with adapter fine-tuning to obtain high-level time-frequency domain features H A ; Adaptive feature fusion module 2 combines advanced time-domain and frequency-domain features and spatial characteristics Connect to get time-frequency-space features , and then the time domain-frequency domain-space features Input the adaptive Transformer feature fusion network with adapter fine-tuning to get the new output ; The fine-tuning module is used to input the adapter fine-tuning data obtained by the dataset partitioning module into the adapter fine-tuning module of the adaptive Transformer feature fusion network with adapter fine-tuning The result is used as the output of the adapter fine-tuning data H T ; Predicted emotion probability acquisition module, respectively, the new output and adapters to fine-tune the output of data H T Through the fully connected layer and Softmax Function, respectively obtain The predicted sentiment probability and H T The predicted emotion probability of The maximum value of the predicted emotion probability is taken as The classification result is The predicted sentiment label of H T The maximum value of the predicted emotion probability is taken as The classification result is The predicted sentiment label of The cross entropy loss function value calculation module calculates the cross entropy loss function value of the adapter fine-tuning data and the training set according to the predicted emotion probability; The training module is used to repeatedly call the data set partitioning module and the cross entropy loss function value calculation module in sequence to obtain the trained model of the current subject by minimizing the cross entropy loss function value between the model prediction and the actual label; The subject repetition module is used to take the next subject as the target domain and the remaining subjects as the source domains, and repeatedly call the dataset partitioning module~training module in sequence until each subject is taken as the target domain once, and finally obtain a trained model for each subject.
[0008] Compared with the prior art, the technical effects of the method of the present invention are as follows: (1) The present invention adopts a multi-feature extraction method, utilizing the time domain features (DE features) and frequency domain features (PSD features) of EEG signals. The channel attention weighted spatial convolution branch and two dynamic graph convolution network branches can fully extract and utilize the time, frequency, and spatial features of EEG signals.
[0009] (2) A channel-attention-weighted spatial convolution branch is designed in the multi-branch feature fusion network. This branch uses the channel-attention-weighted approach to learn the adaptive weights of each channel of the EEG signal, capturing the correlation of EEG features in the channel dimension while reducing redundant data. This improves the spatiotemporal feature representation capability and model generalization capability of the EEG signal, which is beneficial for subsequent spatial convolution to fully extract the spatial features of the EEG signal.
[0010] (3) Two dynamic graph convolutional network branches that rely on distance maps are used in the multi-branch feature fusion network. The EEG distance map can represent the natural geometric shape of the EEG electrodes and the correlation information between the electrodes, which improves the feature extraction capability of the graph convolutional network for EEG data and thus improves the recognition accuracy.
[0011] (4) Adopting an adaptive Transformer feature fusion network with adapter fine-tuning, we can extract and fuse high-level temporal, frequency, and spatial features, effectively associating the spatial distribution of EEG channels with deeply encoded emotion features, thereby improving the performance of emotion classification. The adapter fine-tuning module achieves fast cross-subject EEG emotion recognition by fine-tuning the adapter fine-tuning module, effectively avoiding the overfitting problem caused by subject-dependent methods and improving the recognition accuracy. DETAILED DESCRIPTION
[0012] For the sake of clarity and conciseness, not all features of an actual embodiment are described in the specification. However, it should be appreciated that in the process of developing any such actual embodiment, many implementation-specific decisions may be made to achieve the specific goals of the developers, and these decisions may vary from one implementation to another.
[0013] This embodiment provides a method for EEG emotion recognition based on multi-branch feature fusion, which specifically includes the following steps: Step 1: Obtain EEG data of multiple subjects.
[0014] Step 2: preprocessing the EEG data of multiple subjects to obtain preprocessed EEG data of each subject.
[0015] Preferably, the preprocessing includes filtering, removing noise and artifacts, and feature smoothing. Preprocessing can reduce interference and lay the foundation for accurate recognition in the later stage.
[0016] Step 3: Extract the features of the five frequency bands of each subject’s preprocessed EEG data. The five frequency bands are delta (0-4Hz), theta (4-8Hz), alpha (8-14Hz), beta (14-31Hz) and gamma (31-50Hz).
[0017] Step 4: The features of the five frequency bands of the first subject are used as the target domain, and the features of the five frequency bands of the remaining subjects are used as the source domain.
[0018] Step 5: Extract the time domain features and frequency domain features of the target domain. The time domain features and frequency domain features are DE features and PSD features respectively. Take 10% to 20% of the DE features and PSD features as adapter fine-tuning data (subsequently used in step 10), and use the remaining DE features and PSD features as the test set. Extract the DE features and PSD features of the source domain as the training set.
[0019] Specifically, the DE (differential entropy) feature can effectively distinguish low-frequency and high-frequency in EEG signals; the frequency spectrum of a fixed-length EEG segment is generally a Gaussian distribution N(µ,σ 2 ); PSD (power spectral density) features can effectively reflect the energy distribution of the signal in the frequency dimension and extract frequency characteristics related to emotions, cognition or pathology. The calculation formulas for DE features and PSD features are as follows:
[0020]
[0021] Where: —DE characteristics; -variance; —Natural constant, approximately equal to 2.71828; —The target domain or source domain in the DE feature calculation formula; —mean; —PSD features; —The target domain or source domain in the PSD feature calculation formula; —Window function, usually Hanning window or Gaussian window; -time; —Window length (time delay); — imaginary units; —angular frequency, which is used to describe the characteristics of EEG signals in the frequency domain, ω=2πf, f is the frequency.
[0022] Step six, input the test set and the training set into the multi-branch feature fusion network together. Specifically, the multi-branch feature fusion network includes a channel attention weighted spatial convolution branch, a first dynamic graph convolution network branch with the same structure, and a second dynamic graph convolution network branch; the DE features in the test set and the training set are respectively input into the first dynamic graph convolution network branch and the channel attention weighted spatial convolution branch, and the PSD features in the test set and the training set are input into the second dynamic graph convolution network branch.
[0023] The multi-branch feature fusion network can effectively capture the temporal, spectral and spatial characteristics of EEG simultaneously.
[0024] Step 7: The first dynamic graph convolutional network branch processes the input DE features to obtain time domain features. The second dynamic graph convolutional network branch processes the input PSD features to obtain frequency domain features ; The time domain features and frequency domain characteristics Connect them to get the time domain-frequency domain features .
[0025] Specifically, time domain features and frequency domain characteristics The calculation formula is as follows:
[0026]
[0027] Where: —Sparse adjacent matrix. The distance graph of the connections between nodes is composed of a sparse adjacent matrix The present embodiment adopts the natural geometric shape of 62 EEG electrodes under the international standard 10-20 EEG electrode placement conditions to represent , which is in line with the intuitiveness of the brain's physiological structure. The nodes of the graph correspond to EEG electrodes, and the edges represent the connections between electrodes. By integrating the connection information between EEG electrodes into the graph structure, various graph features can be easily extracted for data analysis.
[0028] D — , ∈ D , is a sparse adjacent matrix The degree matrix of ,in i and j are the rows and columns of the sparse adjacency matrix; —weight matrix;
[0029] Where: —According to the international standard 10-20 EEG electrode placement conditions, the electrodes and The Euclidean distance between — standard deviation of distance; —The sparsity threshold is set to 0.9 in this embodiment based on preliminary experiments and EEG domain knowledge.
[0030] Weight Matrix Used to perform linear transformations on DE (differential entropy) features and PSD (power spectral density) features. They play a role in adjusting feature representation in graph convolution operations, helping the model learn the relationship between different features and graph structures (represented by Ads).
[0031] , —ELU nonlinear function; —DE characteristics; —PSD features; ,in is the number of EEG channels, is the number of frequency bands.
[0032] The channel attention weighted spatial convolution branch uses channel attention weighted learning to learn the channel adaptive weights of DE features, which captures the correlation of input features in the channel dimension while reducing redundant data. Channel attention weighting can effectively improve the spatiotemporal feature representation ability and model generalization ability of EEG signals, which is beneficial to the spatial feature extraction of the spatial convolution module. Specifically, in this embodiment, the channel adaptive weight calculation formula of DE features is as follows:
[0033]
[0034] Where: — EEG signal The variance of the channels; —The variance of all channels of EEG signal; - Number of EEG channels; In this embodiment, C is 62; —The time point of the EEG signal; t -time; —The channel at time t c DE characteristics; —Channel adaptive weights of DE features; —DE features of all channels of EEG signals; — unit multiplication; (·)— Activation function; (·)— Activation function; (·) — normalization layer; , —The learning weight matrices of two fully connected layers (FC layers); is matrix multiplication.
[0035] Finally, the weighted multi-channel EEG signal is subjected to spatial convolutional blocks. Processing is performed to finally obtain spatial features , the calculation formula is as follows:
[0036] Where: —Spatial characteristics; —Spatial convolution block, whose convolution kernel size is ; (·) — Batch Normalization layer.
[0037] Step 8: The time domain-frequency domain features obtained in step 7 Input the adaptive Transformer feature fusion network with adapter fine-tuning to obtain high-level time-frequency domain features H A ; Specifically, the adaptive Transformer feature fusion network with adapter fine-tuning includes a multi-head self-attention mechanism module and adapter fine-tuning modules , and its calculation formula is as follows:
[0038]
[0039] Where: —Advanced time-frequency domain features; MHSA (∙) —Multi-head attention mechanism; Adapter (∙)—Adapter fine-tuning module. In this embodiment, the adapter fine-tuning module includes a linear layer, an ELU nonlinear function and a linear layer connected in sequence; LN (∙) — batch normalization layer; Preferably, multi-head self-attention mechanism module MHSA In the calculation formula of the self-attention mechanism:
[0040]
[0041] Where: - query; -key; -value; —Self-attention mechanism; —Query weight matrix, ; — key weight matrix, ; — value weight matrix, ; , —Hyperparameters; — sparse adjacency matrix; Softmax (∙)— Softmax function; In this step, the adaptive Transformer feature fusion network with adapter fine-tuning is used to fusion the input time-frequency features. H C Perform feature extraction, specifically learning time domain-frequency domain features in parallel Different representations of time domain-frequency domain features H C The information fusion of all elements in the can capture the time-frequency domain characteristics Internal dependencies to obtain high-level time-frequency features H A .
[0042] Among them, the adapter fine-tuning module is the part of the adaptive Transformer feature fusion network with adapter fine-tuning that is optimized for specific subjects or tasks. This module is a lightweight fine-tuning method that can quickly identify the characteristics of target subject samples without retraining the entire model, while maintaining model performance comparable to full fine-tuning. This module is implemented through adapter fine-tuning technology for transfer learning. The module includes linear layers, ELU nonlinear functions, and linear layers connected in sequence, and the adapter is also equipped with a skip connection to ensure parameter efficiency. The adaptive fine-tuning transfer learning method is applied to cross-subject emotion recognition, which proves that the method is parameter efficient when there are fewer target subject samples, while helping to improve the recognition accuracy of the target sample emotions.
[0043] Step 9: Advanced time-frequency features and spatial characteristics Connect to get time-frequency-space features , and then the time domain-frequency domain-space features Perform the time-frequency domain feature analysis in step 8 The same processing (that is, time-frequency-space features Input the adaptive Transformer feature fusion network with adapter fine-tuning), and get the new output .
[0044] Step 10: Network fine-tuning. Specifically, the adapter fine-tuning data obtained in step 5 is input into the adapter fine-tuning module of the adaptive Transformer feature fusion network with adapter fine-tuning. The result is used as the output of the adapter fine-tuning data H T .
[0045] This step can quickly identify the characteristics of the target subject sample without retraining the entire model, thereby improving the recognition accuracy of the target sample emotion.
[0046] Step 11: Output the new and the adapter fine-tunes the data output H T Through the fully connected layer and Softmax Function, respectively obtain The predicted sentiment probability and H T The predicted emotion probability of The maximum value of the predicted emotion probability is taken as The classification result is The predicted sentiment label of H T The maximum value of the predicted emotion probability is taken as The classification result is The predicted sentiment labels.
[0047] The predicted emotion label is the emotion state of the subject corresponding to the sample at the time of EEG acquisition.
[0048] Step 12, based on the predicted emotion probability, calculate the cross entropy loss function value of the adapter fine-tuning data and the training set; Specifically, the calculation formula of the cross entropy loss function is as follows:
[0049] Where: — loss function; —The actual label of the nth EEG data; —The nth piece of EEG data is compared with the j The predicted sentiment probability of the class; —Total number of samples; —Number of classes.
[0050] Step 13, repeating steps 6 to 12, and obtaining a trained model of the current subject by minimizing the cross entropy loss function value between the model prediction and the actual label. In this embodiment, the number of training iterations is 200.
[0051] In step 14, the next subject is used as the target domain and the remaining subjects are used as source domains. Steps 5 to 13 are repeated until each subject is used as the target domain once, and finally a trained model for each subject is obtained.
[0052] In order to analyze the feasibility and effectiveness of the method of the present invention, the following experiments were conducted.
[0053] 1. Dataset The core of the algorithm of the present invention lies in the fusion of multi-branch adaptive Transformer features in time domain, frequency domain and space, and through adapter fine-tuning, the model can quickly adapt to new subjects and data.
[0054] In order to verify the effectiveness of the algorithm, the present invention was evaluated on the Shanghai Jiao Tong University Emotional EEG Dataset (SEED) and (SEED-IV). The datasets were collected by 15 healthy subjects, and EEG data were obtained through a 62-channel ESI neural scanning system. Among them, SEED was collected by having the same participants watch 15 movie clips that aroused positive, neutral, and negative emotions; while the SEED-IV dataset was collected by having the same participants watch 24 movie clips that evoked neutral, sad, fearful, and happy emotions. The SEED signal was filtered through a 0-75 Hz bandpass filter and divided into non-overlapping intervals of 1 second each; the SEED-IV signal was filtered through a 1-75 Hz bandpass filter and divided into non-overlapping intervals of 4 seconds each. For the SEED and SEED-IV datasets, each subject and each clip was collected three times. By extracting and identifying features of these EEG signals, the performance of the method of the present invention on the EEG emotion recognition task can be comprehensively evaluated.
[0055] 2. Evaluation indicators The average accuracy (Acc) and variance (Std) are used as measurement indicators for sentiment judgment of model data sets, among which the average accuracy is the main indicator.
[0056] 3. Experimental results and analysis In order to verify the model's ability to perform emotion recognition on EEG datasets, the present invention selected some cross-subject emotion classification models in recent years as comparison models and conducted comparative experiments.
[0057] As shown in Table 1, after specific actual experiments, this experiment was conducted on a deep learning workstation with an Ubuntu system. The GPU of the workstation was NVIDIA GeForce RTX 2080Ti. The PyTorch neural network framework was used. The proportion of fine-tuning data in the target domain of the SEED dataset was set to 20%, and the proportion of fine-tuning data in the target domain of the SEED-IV dataset was set to 10%. The above method was run with a learning rate parameter of 0.001. The model was trained for 200 epochs with a learning rate of 1×10 −3 , the batch size is 32, and the experimental data that can be obtained are: for the three-classification and four-classification of EEG signals, the accuracy rates reached 99.68% and 96.13% on the SEED and SEED-IV datasets respectively.
[0058] It can be seen that in the cross-subject emotion recognition experiment of EEG signals, the present invention has achieved excellent results with less fine-tuning data, especially the accuracy of the three-classification has been close to 100%, effectively realizing the transfer learning of emotion.
[0059] 4. Quantitative analysis Table 1 shows the classification results of cross-subject emotion recognition on the SEED and SEED-IV datasets. The higher the classification accuracy, the better the model effect, and the lower the variance, the better the model can adapt to the characteristics of EEG signals of different individuals.
[0060] In Table 1, the average recognition accuracy of the present invention for the two data sets is higher than that of other algorithms, which proves the high efficiency of the present invention in the field of emotion transfer learning.
[0061] However, the variance of the present invention is not optimal. On the one hand, the reason may be that the average accuracy of the present invention is relatively high, and the difference in accuracy of different subjects has a greater impact on the overall variance. On the other hand, it may be that there is a focus on the optimization of hyperparameters. During the model training process, the optimization of hyperparameters focuses more on the goal of improving accuracy. For example, the selection of learning rate, regularization parameters, etc. is to prioritize the model to learn as accurate emotions as possible. However, because the goal of the present invention is to improve the accuracy of emotion recognition, non-optimal variance can be accepted under the premise of extremely high accuracy, which is also an aspect that we can further optimize in the future.
[0062] Table 1 Cross-subject experiments under different labeling conditions
[0063] Table 2 shows the experimental results of the model of the present invention under different proportions of fine-tuning data in the target domain. It can be seen that as the proportion of fine-tuning data increases, the classification effect of the model also shows a trend of continuous improvement. At first, the accuracy rate increases rapidly, but as the model accuracy approaches a bottleneck, the accuracy rate gradually decreases, and even increases the proportion of fine-tuning data but the average accuracy decreases slightly. However, as the proportion of fine-tuning data increases, the recognition accuracy of the model gradually increases until it reaches a bottleneck.
[0064] Our application can choose the proportion of fine-tuning data according to specific requirements. When the labeled data is sufficient, a higher average recognition accuracy can be obtained by increasing the proportion of fine-tuning data. When the labeled data is insufficient, the amount of fine-tuning data and the accuracy can be weighed to select the most appropriate proportion of fine-tuning data. In the SEED dataset, a good recognition accuracy can be obtained when the proportion of fine-tuning data is only 1%, and the best accuracy is achieved when the proportion of fine-tuning data is 20%; in the SEED-IV dataset, the trade-off when the proportion of fine-tuning data is 10% is more appropriate.
[0065] Table 2 Cross-subject experiments with different proportions of fine-tuning data in the target domain
[0066] In summary, in order to solve the problems of non-stationarity of EEG signals and difficulty in processing high-dimensional signals by multi-feature fusion in cross-subject EEG emotion recognition, the present invention proposes an EEG emotion recognition method based on multi-branch feature fusion. This method comprehensively extracts the time domain, frequency domain and spatial features of EEG signals, uses channel attention weighting to reduce redundant data, constructs a graph convolution branch network that depends on distance graphs to improve feature extraction capabilities, and uses an adaptive Transformer feature fusion network combined with adapter fine-tuning to adapt to different subjects and data changes, thereby improving the accuracy of emotion recognition, enabling it to better identify the emotional states of different subjects, and meeting the needs of actual applications.
[0067] In this way, the present invention successfully improves the effect of emotion recognition on EEG signals, making it more in line with people's needs. In order to verify the effectiveness of the algorithm of the present invention, the present invention is evaluated on the Shanghai Jiaotong University emotional EEG dataset SEED and SEED-IV. The experimental results show that the algorithm of the present invention achieves the best results in terms of average accuracy index and shows high efficiency in the field of emotion transfer learning. At the same time, through cross-subject experiments with different proportions of fine-tuning data, it can be seen that changes in the proportion of fine-tuning data will affect the classification effect of the model, and the appropriate proportion can be weighed and selected as needed to obtain a better recognition accuracy.
[0068] The above description is only an embodiment of the present invention, and the protection scope of the present invention is not limited to the enumerated contents. Any changes or substitutions that can be made by technicians within the technical scope disclosed by the present invention are within the protection scope of the present invention. The protection scope of the present invention shall be subject to the claims.
Claims
1. A method for EEG emotion recognition based on multi-branch feature fusion, characterized in that: The specific steps include: Step 1, obtaining EEG data of multiple subjects; Step 2, preprocessing the EEG data of multiple subjects to obtain preprocessed EEG data of each subject; Step 3, extract the features of five frequency bands of the preprocessed EEG data of each subject, the five frequency bands are delta, theta, alpha, beta and gamma; Step 4: The features of the five frequency bands of the first subject are used as the target domain, and the features of the five frequency bands of the remaining subjects are used as the source domain; Step 5: extract the time domain features and frequency domain features of the target domain, where the time domain features and frequency domain features are DE features and PSD features respectively; Take 10%~20% of DE features and PSD features as adapter fine-tuning data, and use the remaining DE features and PSD features as test sets; Extract DE features and PSD features of the source domain as training sets; Step 6: Input the test set and the training set into a multi-branch feature fusion network together. Specifically, the multi-branch feature fusion network includes a channel attention weighted spatial convolution branch, a first dynamic graph convolution network branch and a second dynamic graph convolution network branch with the same structure; input the DE features in the test set and the training set into the first dynamic graph convolution network branch and the channel attention weighted spatial convolution branch respectively, and input the PSD features in the test set and the training set into the second dynamic graph convolution network branch; Step 7: The first dynamic graph convolutional network branch processes the input DE features to obtain time domain features. The second dynamic graph convolutional network branch processes the input PSD features to obtain frequency domain features ; The time domain features and frequency domain characteristics Connect them to get the time domain-frequency domain features ; The channel attention weighted spatial convolution branch adopts the channel-adaptive weights of the channel attention weighted learning DE features; Finally, the weighted multi-channel EEG signal is subjected to spatial convolutional blocks. Processing is performed to finally obtain spatial features ; Step 8: The time domain-frequency domain features obtained in step 7 Input the adaptive Transformer feature fusion network with adapter fine-tuning to obtain high-level time-frequency domain features H A ; Step 9: Advanced time-frequency features and spatial characteristics Connect to get time-frequency-space features , and then the time domain-frequency domain-space features Input the adaptive Transformer feature fusion network with adapter fine-tuning to get the new output ; Step 10: Input the adapter fine-tuning data obtained in step 5 into the adapter fine-tuning module of the adaptive Transformer feature fusion network with adapter fine-tuning The result is used as the output of the adapter fine-tuning data H T ; Step 11: Output the new and adapters to fine-tune the output of data H T Through the fully connected layer and Softmax Function, respectively obtain The predicted sentiment probability and H T The predicted emotion probability of The maximum value of the predicted emotion probability is taken as The classification result is The predicted sentiment label of H T The maximum value of the predicted emotion probability is taken as The classification result is The predicted sentiment label of Step 12, based on the predicted emotion probability, calculate the cross entropy loss function value of the adapter fine-tuning data and the training set; Step 13: Repeat steps 6 to 12 to obtain the trained model of the current subject by minimizing the cross entropy loss function value between the model prediction and the actual label; In step 14, the next subject is used as the target domain and the remaining subjects are used as source domains. Steps 5 to 13 are repeated until each subject is used as the target domain once, and finally a trained model for each subject is obtained.
2. The EEG emotion recognition method based on multi-branch feature fusion as claimed in claim 1, characterized in that: In step 5, the calculation formulas of DE feature and PSD feature are as follows: Where: —DE characteristics; -variance; —Natural constant, approximately equal to 2.71828; —The target domain or source domain in the DE feature calculation formula; —mean; —PSD features; —The target domain or source domain in the PSD feature calculation formula; - window function; -time; - window length; — imaginary units; —Angular frequency.
3. The EEG emotion recognition method based on multi-branch feature fusion as claimed in claim 2, characterized in that: In step 7, time domain features and frequency domain characteristics The calculation formula is as follows: Where: — sparse adjacency matrix; D — , ∈ D , is a sparse adjacent matrix The degree matrix of ,in i and j are the rows and columns of the sparse adjacency matrix; —weight matrix; Where: —According to the international standard 10-20 EEG electrode placement conditions, the electrodes and The Euclidean distance between — standard deviation of distance; —The sparsity threshold is 0.9; , —ELU nonlinear function; —DE characteristics; —PSD features; ,in is the number of EEG channels, is the number of frequency bands.
4. The EEG emotion recognition method based on multi-branch feature fusion as claimed in claim 3, characterized in that: In step 7, the channel adaptive weight calculation formula of DE feature is as follows: Where: — EEG signal The variance of the channels; —The variance of all channels of EEG signal; —Number of EEG channels; —The time point of the EEG signal; t -time; —The channel at time t c DE characteristics; —Channel adaptive weights of DE features; —DE features of all channels of EEG signals; — unit multiplication; (·)— Activation function; (·)— Activation function; (·) — normalization layer; , —The learning weight matrices of the two fully connected layers; is matrix multiplication.
5. The EEG emotion recognition method based on multi-branch feature fusion as claimed in claim 4, characterized in that: In step 7, spatial features The calculation formula is as follows: Where: —Spatial characteristics; —Spatial convolution block, whose convolution kernel size is ; (·) — Batch Normalization layer.
6. The EEG emotion recognition method based on multi-branch feature fusion as claimed in claim 5, characterized in that: In step 8, the adaptive Transformer feature fusion network with adapter fine-tuning includes a multi-head self-attention mechanism module and adapter fine-tuning modules , and its calculation formula is as follows: Where: —Advanced time-frequency domain features; MHSA (∙) —Multi-head attention mechanism; Adapter (∙) —Adapter fine-tuning module, including a linear layer, an ELU nonlinear function, and a linear layer connected in sequence; LN (∙) — Batch Normalization layer.
7. The method for EEG emotion recognition based on multi-branch feature fusion as claimed in claim 6, characterized in that: The multi-head self-attention mechanism module MHSA In the calculation formula of the self-attention mechanism: Where: - query; -key; -value; —Self-attention mechanism; —Query weight matrix, ; — key weight matrix, ; — value weight matrix, ; , —Hyperparameters; — sparse adjacency matrix; Softmax (∙)— Softmax function.
8. The EEG emotion recognition method based on multi-branch feature fusion as claimed in claim 7, characterized in that: In step 12, the calculation formula of the cross entropy loss function is as follows: Where: — loss function; —The actual label of the nth EEG data; —The nth piece of EEG data is compared with the j The predicted sentiment probability of the class; —Total number of samples; —Number of classes.
9. An EEG emotion recognition system based on multi-branch feature fusion, characterized in that: The modules include: An EEG data acquisition module is used to acquire EEG data of multiple subjects; A preprocessing module, used for preprocessing the EEG data of multiple subjects to obtain the preprocessed EEG data of each subject; The frequency band feature extraction module is used to extract the features of five frequency bands of the preprocessed EEG data of each subject, and the five frequency bands are delta, theta, alpha, beta and gamma; A feature partitioning module, used to use the features of the five frequency bands of the first subject as target domains, and the features of the five frequency bands of the remaining subjects as source domains; The data set partitioning module is used to extract the time domain features and frequency domain features of the target domain. The time domain features and frequency domain features are DE features and PSD features respectively. Take 10%~20% of DE features and PSD features as adapter fine-tuning data, and use the remaining DE features and PSD features as test sets; Extract DE features and PSD features of the source domain as training sets; A data input module is used to input the test set and the training set into a multi-branch feature fusion network together, specifically: the multi-branch feature fusion network includes a channel attention weighted spatial convolution branch, a first dynamic graph convolution network branch and a second dynamic graph convolution network branch with the same structure; the DE features in the test set and the training set are respectively input into the first dynamic graph convolution network branch and the channel attention weighted spatial convolution branch, and the PSD features in the test set and the training set are input into the second dynamic graph convolution network branch; The feature fusion module is used to implement the following functions: the first dynamic graph convolutional network branch processes the input DE features to obtain time domain features The second dynamic graph convolutional network branch processes the input PSD features to obtain frequency domain features ; The time domain features and frequency domain characteristics Connect them to get the time domain-frequency domain features ; The channel attention weighted spatial convolution branch adopts the channel adaptive weights of the channel attention weighted learning DE features; finally, the spatial convolution block is used to weight the multi-channel EEG signal Processing is performed to finally obtain spatial features ; Adaptive feature fusion module 1 is used to integrate time domain and frequency domain features Input the adaptive Transformer feature fusion network with adapter fine-tuning to obtain high-level time-frequency domain features H A ; Adaptive feature fusion module 2 combines advanced time-domain and frequency-domain features and spatial characteristics Connect to get time-frequency-space features , and then the time domain-frequency domain-space features Input the adaptive Transformer feature fusion network with adapter fine-tuning to get the new output ; The fine-tuning module is used to input the adapter fine-tuning data obtained by the dataset partitioning module into the adapter fine-tuning module of the adaptive Transformer feature fusion network with adapter fine-tuning The result is used as the output of the adapter fine-tuning data H T ; Predicted emotion probability acquisition module, respectively, the new output and adapters to fine-tune the output of data H T Through the fully connected layer and Softmax Function, respectively obtain The predicted sentiment probability and H T The predicted emotion probability of The maximum value of the predicted emotion probability is taken as The classification result is The predicted sentiment label of H T The maximum value of the predicted emotion probability is taken as The classification result is The predicted sentiment label of The cross entropy loss function value calculation module calculates the cross entropy loss function value of the adapter fine-tuning data and the training set according to the predicted emotion probability; The training module is used to repeatedly call the data set partitioning module and the cross entropy loss function value calculation module in sequence to obtain the trained model of the current subject by minimizing the cross entropy loss function value between the model prediction and the actual label; The subject repetition module is used to take the next subject as the target domain and the remaining subjects as the source domains, and repeatedly call the dataset partitioning module~training module in sequence until each subject is taken as the target domain once, and finally obtain a trained model for each subject.
Citation Information
Patent Citations
Emotion recognition method and system fusing prior and automatic electroencephalogram characteristics
CN112836593A
Cross-domain electroencephalogram emotion recognition method and system based on space-time feature fusion model
CN114469137A
Brain electrical emotion recognition method based on attention mechanism space-time position coding
CN115736923A
Cross-subject electroencephalogram cognitive load assessment method based on cross attention network
CN117195038A
Multi-modal feature fusion emotion recognition method based on gating cross-attention mechanism
CN117370828A
Cited By
Electroencephalogram signal field adaptation method and system based on cross attention
CN120257027A