EEG Emotion Recognition Method and System Based on Multi-Branch Feature Fusion

Through the multi-branch feature fusion method, combined with channel attention and adaptive Transformer network, the problem of low accuracy in EEG-Electrophy emotional recognition across subjects is solved, and efficient emotion recognition effect is achieved, adapting to the changes in EEG data of different subjects.

CN119989160BActive Publication Date: 2025-07-04NORTHWEST UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510472790.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-04-16
Publication Date
2025-07-04
Estimated Expiration
2045-04-16

AI Technical Summary

Technical Problem

The prior art has the problem of low accuracy of emotion recognition in EEG-encephalopathy recognition across subjects, especially due to the difficulty of feature extraction and fusion caused by the non-stationarity of EEG signals and the high dimensional complexity, and there is a risk of high computational complexity or overfitting.

Method used

The multi-branch feature fusion method is adopted, including channel attention-weighted spatial convolutional branch, dynamic graph convolution network branch and adaptive Transformer feature fusion network. By extracting time domain, frequency domain and spatial features, and using the adapter fine-tuning module to quickly adapt, it realizes EEG emotional recognition across subjects.

Benefits of technology

It improves the accuracy of identification of EEG signals, reduces the problem of overfitting, improves the generalization ability and emotional classification performance of the model, and adapts to changes in EEG data of different subjects.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure SMS_43
    Figure SMS_43
  • Figure SMS_44
    Figure SMS_44
  • Figure SMS_64
    Figure SMS_64
Patent Text Reader

Abstract

The present invention belongs to the field of cross-subject electroencephalogram (EEG) emotion recognition, and discloses an EEG emotion recognition method and system based on multi-branch feature fusion. The method includes: Step 1, obtaining EEG data; Step 2, preprocessing the data; Step 3, extracting five-band features; Step 4, obtaining the target domain and the source domain; Step 5, dividing the data set; Step 6, inputting into the feature fusion network; Step 7, obtaining time-domain - frequency-domain features and spatial features; Step 8, obtaining high-level time-domain - frequency-domain features; Step 9, obtaining the output; Step 10, obtaining the adapter fine-tuning data output; Step 11, obtaining the predicted emotion probability; Step 12, calculating the cross-entropy loss function; Step 13, obtaining the trained model of the current subject being tested; Step 14, obtaining the trained models of all subjects being tested. The present invention more accurately achieves transfer learning between different samples, effectively improving the overall effect of cross-subject EEG emotion recognition.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of cross-subject EEG emotion recognition, and specifically, relates to an EEG emotion recognition method and system based on multi-branch feature fusion. Background Art

[0002] Emotions are the basis of human experience and play an important role in happiness, behavioral performance, social interaction, interpersonal communication, decision-making and cognitive functions. Affective computing technology enables computers to recognize, understand, represent, adapt and give feedback. Among them, emotional computing based on physiological signals has become a hot topic. EEG signals are widely used in the field of emotional computing due to their portability and high temporal resolution. Brain-computer interface (BCI) collects brain signals induced by emotional stimulation, performs pattern recognition, analyzes the type of brain emotion perception, and converts them into external instructions to achieve human-computer emotional interaction.

[0003] In recent years, recent advances in electrode technology and machine learning have improved electroencephalogram (EEG) analysis for emotion recognition. However, the inherent non-stationarity of EEG signals can vary from individual to individual or over time, which poses a challenge to developing a universally applicable model for all subjects. Simply transferring may lead to misclassification of the pre-trained model, resulting in poor classification performance for the target subject.

[0004] At the same time, EEG signals are characterized by high dimensionality and complexity, and it is not easy to extract effective emotional features from them. The extraction of time domain, frequency domain and spatial features requires complex algorithms and techniques, and the fusion of different features is also challenging. Some studies have tried to use deep learning methods, such as convolutional neural networks (CNN) and recurrent neural networks (RNN), to automatically extract and fuse different features. Although these methods have achieved good results, these methods have the risk of high computational complexity or overfitting when processing high-dimensional EEG signals, resulting in low accuracy in emotion recognition. Further research is still needed on the fusion of multiple features of EEG signals. Summary of the invention

[0005] In order to solve the problem of low emotion recognition accuracy in existing emotion feature extraction methods, the present invention provides an EEG emotion recognition method based on multi-branch feature fusion.

[0006] In order to achieve the above object, the present invention provides the following technical solutions:

[0007] On the one hand, the present invention provides an EEG emotion recognition method based on multi-branch feature fusion, which specifically includes the following steps:

[0008] Step 1, obtaining EEG data of multiple subjects;

[0009] Step 2: Preprocess the EEG data of multiple subjects to obtain the preprocessed EEG data of each subject;

[0010] Step 3: Extract the features of five frequency bands from the preprocessed EEG data of each subject. The five frequency bands are delta, theta, alpha, beta, and gamma;

[0011] Step 4: Use the features of the five frequency bands of the first subject as the target domain, and the features of the five frequency bands of the remaining subjects as the source domain;

[0012] Step 5: Extract the time-domain features and frequency-domain features of the target domain. The time-domain features and frequency-domain features are DE features and PSD features respectively; Take 10% - 20% from the DE features and PSD features as the adapter fine-tuning data, and use the remaining DE features and PSD features as the test set; Extract the DE features and PSD features of the source domain as the training set;

[0013] Step 6: Input the test set and the training set into the multi-branch feature fusion network. Specifically: The multi-branch feature fusion network includes a channel attention weighted spatial convolution branch, a first dynamic graph convolution network branch and a second dynamic graph convolution network branch with the same structure; Input the DE features in the test set and the training set into the first dynamic graph convolution network branch and the channel attention weighted spatial convolution branch respectively, and input the PSD features in the test set and the training set into the second dynamic graph convolution network branch;

[0014] Step 7: The first dynamic graph convolution network branch processes the input DE features to obtain time-domain features , and the second dynamic graph convolution network branch processes the input PSD features to obtain frequency-domain features ; Connect the time-domain features and the frequency-domain features to obtain time-domain - frequency-domain features ;

[0015] The channel attention weighted spatial convolution branch uses channel attention to weight and learn the channel adaptive weights of the DE features;

[0016] Finally, use the spatial convolution block to process the weighted multi-channel EEG signals to finally obtain spatial features ;

[0017] Step 8: Input the time-domain - frequency-domain features obtained in Step 7 into the adaptive Transformer feature fusion network with adapter fine-tuning to obtain high-level time-domain - frequency-domain features H A;

[0018] Step Nine, connect the high-level time-domain - frequency-domain features and spatial features to obtain time-domain - frequency-domain - spatial features , and then input the time-domain - frequency-domain - spatial features into the adaptive Transformer feature fusion network with adapter fine-tuning to obtain a new output ;

[0019] Step Ten, input the adapter fine-tuning data obtained in Step Five into the adapter fine-tuning module of the adaptive Transformer feature fusion network with adapter fine-tuning , and take the obtained result as the output of the adapter fine-tuning data H T ;

[0020] Step Eleven, respectively pass the new output and the output of the adapter fine-tuning data H T through the fully connected layer and Softmax function to respectively obtain 's predicted sentiment probability and H T 's predicted sentiment probability; among them, 's maximum predicted sentiment probability is regarded as 's classification result, and this classification result is 's predicted sentiment label; H T 's maximum predicted sentiment probability is regarded as 's classification result, and this classification result is 's predicted sentiment label;

[0021] Step Twelve, calculate the cross-entropy loss function value of the adapter fine-tuning data and the training set according to the predicted sentiment probability;

[0022] Step Thirteen, repeatedly execute Steps Six to Twelve, and obtain the trained model of the current subject by minimizing the cross-entropy loss function value between the model prediction and the actual label;

[0023] Step Fourteen, take the next subject as the target domain and the remaining subjects as the source domain, and repeat Steps Five to Thirteen until each subject has been used as the target domain once, and finally obtain the trained model of each subject.

[0024] On the other hand, the present invention provides an electroencephalogram emotion recognition system based on multi-branch feature fusion, specifically including the following modules:

[0025] An electroencephalogram (EEG) data acquisition module for acquiring EEG data of multiple subjects;

[0026] A preprocessing module for preprocessing the EEG data of multiple subjects to obtain the preprocessed EEG data of each subject;

[0027] A frequency band feature extraction module for extracting the features of five frequency bands of the preprocessed EEG data of each subject, and the five frequency bands are delta, theta, alpha, beta, and gamma respectively;

[0028] A feature division module for using the features of the five frequency bands of the first subject as the target domain and the features of the five frequency bands of the remaining subjects as the source domain;

[0029] A dataset division module for extracting the time-domain features and frequency-domain features of the target domain. The time-domain features and frequency-domain features are DE features and PSD features respectively; taking 10% - 20% from the DE features and PSD features as the adapter fine-tuning data, and using the remaining DE features and PSD features as the test set; extracting the DE features and PSD features of the source domain as the training set;

[0030] A data input module for jointly inputting the test set and the training set into a multi-branch feature fusion network. Specifically: the multi-branch feature fusion network includes a channel attention weighted spatial convolution branch, a first dynamic graph convolution network branch and a second dynamic graph convolution network branch with the same structure; inputting the DE features in the test set and the training set into the first dynamic graph convolution network branch and the channel attention weighted spatial convolution branch respectively, and inputting the PSD features in the test set and the training set into the second dynamic graph convolution network branch;

[0031] A feature fusion module for realizing the following functions: the first dynamic graph convolution network branch processes the input DE features to obtain time-domain features and the second dynamic graph convolution network branch processes the input PSD features to obtain frequency-domain features ; connecting the time-domain features and the frequency-domain features to obtain time-domain - frequency-domain features ; the channel attention weighted spatial convolution branch uses channel attention to weightedly learn the channel adaptive weights of the DE features; finally, using a spatial convolution block to process the weighted multi-channel EEG signals to finally obtain spatial features ;

[0032] An adaptive feature fusion module one for using the time-domain - frequency-domain features Input the adaptive Transformer feature fusion network with adapter fine-tuning to obtain high-level time-frequency domain features H A ;

[0033] Adaptive feature fusion module two, which connects the high-level time-frequency domain features and spatial features to obtain time-frequency domain-spatial features , and then inputs the time-frequency domain-spatial features into the adaptive Transformer feature fusion network with adapter fine-tuning to obtain a new output ;

[0034] Fine-tuning module, which is used to input the adapter fine-tuning data obtained by the dataset partitioning module into the adapter fine-tuning module of the adaptive Transformer feature fusion network with adapter fine-tuning , and the result obtained is used as the output of the adapter fine-tuning data H T ;

[0035] Prediction emotion probability acquisition module, which respectively passes the new output and the output of the adapter fine-tuning data H T through the fully connected layer and Softmax function to respectively obtain 's prediction emotion probability and H T 's prediction emotion probability; among them, 's maximum prediction emotion probability is regarded as 's classification result, and this classification result is 's prediction emotion label; H T 's maximum prediction emotion probability is regarded as 's classification result, and this classification result is 's prediction emotion label;

[0036] Cross-entropy loss function value calculation module, which calculates the cross-entropy loss function value of the adapter fine-tuning data and the training set according to the prediction emotion probability;

[0037] Training module, which is used to repeatedly call the dataset partitioning module~cross-entropy loss function value calculation module in sequence, and obtain the trained model of the current subject by minimizing the cross-entropy loss function value between the model prediction and the actual label;

[0038] The subject repetition module is used to take the next subject as the target domain and the remaining subjects as the source domain, and repeatedly call the dataset partitioning module ~ training module in sequence until each subject has been used as the target domain once, and finally obtain the trained models of each subject.

[0039] Compared with the prior art, the technical effects of the method of the present invention are as follows:

[0040] (1) The present invention adopts a multi-feature extraction method. By using the time-domain features (DE features) and frequency-domain features (PSD features) of EEG signals, the channel attention weighted spatial convolution branch and the two dynamic graph convolution network branches can fully extract and utilize the time, frequency, and spatial features of EEG signals.

[0041] (2) The channel attention weighted spatial convolution branch is designed in the multi-branch feature fusion network. This branch uses the channel attention weighted method to learn the adaptive weights of each channel of EEG signals, captures the correlation of EEG features in the channel dimension while reducing redundant data, improves the spatio-temporal feature representation ability of EEG signals and the model generalization ability, and is conducive to the subsequent spatial convolution to fully extract the spatial features of EEG signals.

[0042] (3) Two dynamic graph convolution network branches depending on the distance map are adopted in the multi-branch feature fusion network. The EEG distance map can represent the natural geometric shape of EEG electrodes and the correlation information between electrodes, improves the feature extraction ability of the graph convolution network for EEG data, and thus improves the recognition accuracy.

[0043] (4) The adaptive Transformer feature fusion network with adapter fine-tuning is adopted to achieve the extraction and fusion of high-level time, frequency, and spatial features, effectively correlates the spatial distribution of EEG channels with the emotion features encoded in depth, and thus improves the emotion classification performance. Among them, the adapter fine-tuning module realizes fast cross-subject EEG emotion recognition by fine-tuning the adapter fine-tuning module Adapter, effectively avoids the overfitting problem brought by the subject-dependent method, and improves the recognition accuracy. Specific embodiments

[0044] For clarity and conciseness, not all features of the actual embodiments are described in the specification. However, it should be understood that many embodiment-specific decisions may be made during the development of any such actual embodiment to achieve the specific goals of the developer, and these decisions may vary with different embodiments.

[0045] This embodiment provides an EEG emotion recognition method based on multi-branch feature fusion, which specifically includes the following steps:

[0046] Step 1: Obtain the EEG data of multiple subjects.

[0047] Step 2: Preprocess the EEG data of multiple subjects to obtain the preprocessed EEG data of each subject.

[0048] Preferably, the preprocessing includes filtering, removing noise and artifacts, and feature smoothing. Through preprocessing, interference can be reduced, laying a foundation for accurate later identification.

[0049] Step 3: Extract the features of five frequency bands from the preprocessed EEG data of each subject. The five frequency bands are delta (0 - 4Hz), theta (4 - 8Hz), alpha (8 - 14Hz), beta (14 - 31Hz), and gamma (31 - 50Hz).

[0050] Step 4: Use the features of the five frequency bands of the first subject as the target domain, and the features of the five frequency bands of the remaining subjects as the source domain.

[0051] Step 5: Extract the time-domain features and frequency-domain features of the target domain. The time-domain features and frequency-domain features are DE features and PSD features respectively; Take 10% - 20% from the DE features and PSD features as the adapter fine-tuning data (for subsequent Step 10), and use the remaining DE features and PSD features as the test set; Extract the DE features and PSD features of the source domain as the training set.

[0052] Specifically, DE (differential entropy) features can effectively distinguish between low-frequency and high-frequency in EEG signals; The spectrum of an EEG segment with a fixed length is generally Gaussian distributed N(µ,σ 2 ); PSD (power spectral density) features can effectively reflect the energy distribution of the signal in the frequency dimension and extract frequency characteristics related to emotions, cognition, or pathology. The calculation formulas for DE features and PSD features are as follows:

[0053]

[0054]

[0055] In the formula:

[0056] —DE feature;

[0057] —Variance;

[0058] —Natural constant, approximately equal to 2.71828;

[0059] —The target domain or source domain in the DE feature calculation formula;

[0060] — Mean;

[0061] — PSD feature;

[0062] — Target domain or source domain in the PSD feature calculation formula;

[0063] — Window function, usually using Hanning window or Gaussian window;

[0064] — Moment;

[0065] — Window length (time delay);

[0066] — Imaginary unit;

[0067] — Angular frequency, which is used to describe the characteristics of EEG signals in the frequency domain, ω = 2πf, f where f is the frequency.

[0068] Step 6: Input the test set and the training set into the multi-branch feature fusion network. Specifically, the multi-branch feature fusion network includes a channel attention weighted spatial convolution branch, a first dynamic graph convolution network branch and a second dynamic graph convolution network branch with the same structure; input the DE features in the test set and the training set into the first dynamic graph convolution network branch and the channel attention weighted spatial convolution branch respectively, and input the PSD features in the test set and the training set into the second dynamic graph convolution network branch.

[0069] The multi-branch feature fusion network can effectively capture the time, frequency and spatial features of EEG signals simultaneously.

[0070] Step 7: The first dynamic graph convolution network branch processes the input DE features to obtain time domain features , and the second dynamic graph convolution network branch processes the input PSD features to obtain frequency domain features ; connect the time domain features and the frequency domain features to obtain time-frequency domain features .

[0071] Specifically, the calculation formulas for the time domain features and the frequency domain features are as follows:

[0072]

[0073]

[0074] In the formula:

[0075] —Sparse adjacency matrix. The distance graph of the connections between nodes is determined by the sparse adjacency matrix and is represented by the natural geometric shape of 62-channel EEG electrodes under the international standard 10-20 EEG electrode placement conditions in this embodiment , which conforms to the intuitiveness of the brain physiological structure. The nodes of the graph correspond to EEG electrodes, and the edges represent the connections between electrodes. Integrating the connection information between EEG electrodes into the graph structure can conveniently extract various graph features for data analysis.

[0076] D — , ∈ D , is the degree matrix of the sparse adjacency matrix , where i and j are the rows and columns of the sparse adjacency matrix;

[0077] —Weight matrix;

[0078]

[0079] In the formula:

[0080] —The Euclidean distance between electrodes and under the international standard 10-20 EEG electrode placement conditions;

[0081] —The standard deviation of the distance;

[0082] —The threshold of sparsity, which is taken as 0.9 in this embodiment based on preliminary experiments and EEG field knowledge.

[0083] The weight matrix is used for linear transformation of DE (differential entropy) features and PSD (power spectral density) features. They play a role in adjusting the feature representation in the graph convolution operation, helping the model learn the relationship between different features and the graph structure (represented by Ads).

[0084] 、 —ELU non-linear function;

[0085] —DE feature;

[0086] — PSD feature; , where is the number of EEG channels, is the number of frequency bands.

[0087] The channel attention weighted spatial convolution branch adopts the channel adaptive weights of the channel attention weighted learning DE features, reducing redundant data while capturing the correlation of the input features in the channel dimension. Through channel attention weighting, the spatio-temporal feature representation ability and model generalization ability of EEG signals can be effectively improved, which is beneficial to the extraction of spatial features by the spatial convolution module. Specifically, in this embodiment, the calculation formula of the channel adaptive weights of the DE features is as follows:

[0088]

[0089]

[0090] In the formula:

[0091] — The variance of the th channel of the EEG signal;

[0092] — The variance of all channels of the EEG signal;

[0093] — The number of EEG channels; in this embodiment, C is 62;

[0094] — The time point of the EEG signal;

[0095] t — Moment;

[0096] — The DE feature of the c channel at time t;

[0097] — The channel adaptive weights of the DE feature;

[0098] — The DE features of all channels of the EEG signal;

[0099] — Element-wise multiplication;

[0100] (·) — Activation function;

[0101] (·) — Activation function;

[0102] (·) — Normalization layer;

[0103] , — Learning weight matrices of two fully connected layers (FC layers);

[0104] is matrix multiplication.

[0105] Finally, a spatial convolution block is used to process the weighted multi-channel EEG signals to finally obtain spatial features , and the calculation formula is as follows:

[0106]

[0107] In the formula:

[0108] — Spatial features;

[0109] — Spatial convolution block, whose convolution kernel size is ;

[0110] (·) — Batch normalization layer.

[0111] Step eight, input the time-domain and frequency-domain features obtained in step seven into the adaptive Transformer feature fusion network with adapter fine-tuning to obtain the advanced time-domain and frequency-domain features H A ;

[0112] Specifically, the adaptive Transformer feature fusion network with adapter fine-tuning includes a multi-head self-attention mechanism module and an adapter fine-tuning module , and its calculation formula is expressed as follows:

[0113]

[0114]

[0115] In the formula:

[0116] — Advanced time-domain and frequency-domain features;

[0117] MHSA (∙) — Multi-head attention mechanism;

[0118] Adapter(∙) - Adapter fine-tuning module. In this embodiment, the adapter fine-tuning module includes a linear layer, an ELU non-linear function, and a linear layer connected in sequence;

[0119] LN (∙) - Batch normalization layer;

[0120] Preferably, in the multi-head self-attention mechanism module MHSA The calculation formula of the self-attention mechanism is:

[0121]

[0122]

[0123] In the formula:

[0124] — Query;

[0125] — Key;

[0126] — Value;

[0127] — Self-attention mechanism;

[0128] — Query weight matrix, ;

[0129] — Key weight matrix, ;

[0130] — Value weight matrix, ;

[0131] 、 — Hyperparameters;

[0132] — Sparse adjacent matrix;

[0133] Softmax (∙) — Softmax Function;

[0134] In this step, the adaptive Transformer feature fusion network with adapter fine-tuning is used to extract features from the input time-frequency domain features H C Specifically, it first learns different representations of the time-frequency domain features in parallel, calculates the information fusion of all elements in the time-frequency domain features H C and can capture the time-frequency domain features Internal dependencies are obtained to get high-level time-domain and frequency-domain features H A 。

[0135] Among them, the adapter fine-tuning module is the part optimized for a specific subject or task in the adaptive Transformer feature fusion network with adapter fine-tuning. This module is a lightweight fine-tuning method that can quickly identify the target subject sample features without retraining the entire model, while maintaining model performance comparable to full-scale fine-tuning. This module is implemented through adapter fine-tuning technology for transfer learning. The module includes a linear layer, an ELU non-linear function, and a linear layer connected in sequence. At the same time, the adapter is also equipped with a skip connection to ensure parameter efficiency. Applying the adaptive fine-tuning transfer learning method to cross-subject emotion recognition proves that this method has parameter efficiency in the case of fewer target subject samples, and at the same time helps to improve the recognition accuracy of the emotions of target samples.

[0136] Step Nine: The high-level time-domain and frequency-domain features and spatial features are connected to obtain time-domain, frequency-domain, and spatial features , and then the time-domain, frequency-domain, and spatial features are processed in the same way as the time-domain and frequency-domain features in Step Eight (that is, the time-domain, frequency-domain, and spatial features are input into the adaptive Transformer feature fusion network with adapter fine-tuning), to obtain a new output 。

[0137] Step Ten: Network fine-tuning. Specifically: The adapter fine-tuning data obtained in Step Five is input into the adapter fine-tuning module of the adaptive Transformer feature fusion network with adapter fine-tuning T 。

[0138] This step can quickly identify the target subject sample features without retraining the entire model, and improve the recognition accuracy of the emotions of target samples.

[0139] Step Eleven: Respectively, the new output and the output H T of the adapter fine-tuning data are passed through a fully connected layer and Softmax function to respectively obtain the predicted emotion probability of H T and the predicted emotion probability of 。 Among them, the maximum value of the predicted emotion probability of is regarded as the classification result of The predicted sentiment label; H T The maximum value of the predicted sentiment probability is regarded as the classification result, and this classification result is the predicted sentiment label.

[0140] This predicted sentiment label is the emotional state of the subject corresponding to the sample at the moment of EEG acquisition.

[0141] Step Twelve, calculate the cross-entropy loss function value of the adapter fine-tuning data and the training set according to the predicted sentiment probability;

[0142] Specifically, the calculation formula of the cross-entropy loss function is as follows:

[0143]

[0144] In the formula:

[0145] —Loss function;

[0146] —The actual label of the nth EEG data;

[0147] —The predicted sentiment probability of the nth EEG data for the j th class;

[0148] —Total number of samples;

[0149] —Number of classes.

[0150] Step Thirteen, repeat Steps Six to Twelve, and obtain the trained model of the current subject by minimizing the cross-entropy loss function value between the model prediction and the actual label. In this embodiment, the number of training iterations is 200 times.

[0151] Step Fourteen, take the next subject as the target domain and the remaining subjects as the source domain, and repeat Steps Five to Thirteen until each subject has been used as the target domain once, and finally obtain the trained model of each subject.

[0152] In order to analyze the feasibility and effectiveness of the method of the present invention, the following experiments were carried out.

[0153] 1. Dataset

[0154] The core of the algorithm of the present invention lies in the multi-branch adaptive Transformer feature fusion in the time domain, frequency domain and space, and enables the model to quickly adapt to new subjects and data through adapter fine-tuning.

[0155] To verify the effectiveness of the algorithm, the present invention was evaluated on the Shanghai Jiao Tong University Emotion EEG Dataset (SEED) and (SEED-IV). The dataset was collected from 15 healthy subjects using a 62-channel ESI neuroscan system to obtain EEG data. Among them, SEED was collected by having the same participants watch 15 movie clips that evoked positive, neutral, and negative emotions; while the SEED-IV dataset was collected by having the same participants watch 24 movie clips that elicited neutral, sad, fearful, and happy emotions. The SEED signals were filtered through a band-pass filter of 0 - 75 Hz and segmented into non-overlapping intervals of 1 second each; the SEED-IV signals were filtered through a band-pass filter of 1 - 75 Hz and segmented into non-overlapping intervals of 4 seconds each. For both the SEED and SEED-IV datasets, signal collection was performed three times for each subject and each clip. By extracting and identifying features from these EEG signals, the performance of the method of the present invention in the EEG emotion recognition task can be comprehensively evaluated.

[0156] 2. Evaluation Metrics

[0157] The average accuracy (Acc) and variance (Std) were used as metrics for judging the emotion of the model dataset data, with the average accuracy as the main metric.

[0158] 3. Experimental Results and Analysis

[0159] To verify the ability of the model to perform emotion recognition on the EEG dataset, the present invention selected some cross-subject emotion classification models in recent years as comparison models for comparative experiments.

[0160] As shown in Table 1, through specific actual experiments, this experiment was conducted on a deep learning workstation with the Ubuntu system. The GPU of the workstation was NVIDIA GeForce RTX 2080Ti, using the PyTorch neural network framework. The proportion of fine-tuning data in the target domain of the SEED dataset was set to 20%, and the proportion of fine-tuning data in the target domain of the SEED-IV dataset was set to 10%. Running the above method with a learning rate parameter of 0.001, the model was trained for 200 epochs, the learning rate was 1×10 −3 , and the batch size was 32. The experimental data that could be obtained were: for the three-classification and four-classification of EEG signals, the accuracies reached 99.68% and 96.13% on the SEED and SEED-IV datasets respectively.

[0161] It can be seen that in the cross-subject emotion recognition experiment of the present invention on EEG signals, excellent results have been achieved with a small amount of fine-tuned data. In particular, the accuracy of three-class classification has approached 100%, effectively realizing emotion transfer learning.

[0162] 4. Quantitative analysis

[0163] Table 1 shows the classification experiment results of cross-subject emotion recognition on the SEED and SEED-IV datasets. Among them, the higher the classification accuracy, the better the model performance, and the lower the variance, the better the model can adapt to the EEG signal characteristics of different individuals.

[0164] In Table 1, the average recognition accuracy of the present invention for the two datasets is higher than that of other algorithms, which proves the efficiency of the present invention in the field of emotion transfer learning.

[0165] However, the variance of the present invention has not achieved the optimal value. On the one hand, the reason may be that the average accuracy of the present invention is relatively high, and the difference in accuracy among different subjects has a greater impact on the overall variance. On the other hand, it may be that there is a focus on the direction of hyperparameter optimization. During the model training process, the optimization of hyperparameters focuses more on the goal of improving accuracy. For example, the selection of learning rate, regularization parameter, etc. is to ensure that the model can learn accurate emotions as much as possible. However, because the goal of the present invention is to improve the emotion recognition accuracy, under the premise of extremely high accuracy, non-optimal variance can be accepted, which is also an aspect that can be further optimized in the future.

[0166] Table 1 Cross-subject experiments under different marking conditions

[0167]

[0168] Table 2 shows the cross-subject experiment results of the model of the present invention under different proportions of fine-tuned data in the target domain. It can be seen that as the proportion of fine-tuned data increases, the classification effect of the model also shows an increasing trend. At first, the accuracy increases rapidly, but as the model accuracy approaches the bottleneck, the improvement speed of the accuracy gradually decreases, and even the situation where the average accuracy slightly decreases when the proportion of fine-tuned data is increased occurs. However, generally, the recognition accuracy of the model gradually increases with the increase of the proportion of fine-tuned data until it reaches a bottleneck.

[0169] Our application can select the fine-tuning data proportion according to specific requirements. When there is enough labeled data, a higher average recognition accuracy can be obtained by increasing the fine-tuning data proportion. When there is a lack of labeled data, a trade-off can be made between the quantity and accuracy of the fine-tuning data to select the most suitable fine-tuning data proportion. In the SEED dataset, a good recognition accuracy can be obtained when the fine-tuning data proportion is only 1%, and the best accuracy is achieved when the fine-tuning data proportion is 20%. In the SEED-IV dataset, a trade-off of 10% for the fine-tuning data proportion is more appropriate.

[0170] Table 2 Cross-subject experiments with different fine-tuning data proportions in the target domain

[0171]

[0172] In summary, aiming at the problems existing in cross-subject EEG emotion recognition, such as the non-stationarity of EEG signals and the difficulty of multi-feature fusion for processing high-dimensional signals, the present invention proposes an EEG emotion recognition method based on multi-branch feature fusion. This method comprehensively extracts the time-domain, frequency-domain, and spatial features of EEG signals, uses channel attention weighting to reduce redundant data, constructs a graph convolutional network based on the distance-dependent graph to enhance the feature extraction ability, and applies an adaptive Transformer feature fusion network combined with adapter fine-tuning to adapt to different subjects and data changes, thereby improving the accuracy of emotion recognition and enabling it to better recognize the emotional states of different subjects, meeting the actual application requirements.

[0173] In this way, the present invention has successfully improved the effect of EEG signal emotion recognition, making it more in line with people's needs. To verify the effectiveness of the algorithm of the present invention, the present invention is evaluated on the Shanghai Jiao Tong University Emotion EEG Dataset SEED and SEED-IV. The experimental results show that the algorithm of the present invention has achieved the best results in terms of the average accuracy index and demonstrated high efficiency in the field of emotion transfer learning. At the same time, it can be seen from the cross-subject experiments with different fine-tuning data proportions that the change in the fine-tuning data proportion will affect the model classification effect, and the appropriate proportion can be selected through trade-off according to needs to obtain a better recognition accuracy.

[0174] The above are only the embodiments of the present invention, and the protection scope of the present invention is not limited to the listed content. Any changes or substitutions that can be made by any person skilled in the art within the technical scope disclosed by the present invention are within the protection scope of the present invention. The protection scope of the present invention is subject to the claims.

Claims

1. A method for electroencephalogram emotion recognition based on multi-branch feature fusion, characterized in that Specifically, it includes the following steps: Step 1: Obtain the electroencephalogram (EEG) data of multiple subjects. Step 2: Preprocess the EEG data of multiple subjects to obtain the preprocessed EEG data of each subject. Step 3: Extract the features of five frequency bands from the preprocessed EEG data of each subject. The five frequency bands are delta, theta, alpha, beta, and gamma respectively. Step 4: Use the features of the five frequency bands of the first subject as the target domain, and use the features of the five frequency bands of the remaining subjects as the source domain. Step 5: Extract the time-domain features and frequency-domain features of the target domain. The time-domain features and frequency-domain features are DE features and PSD features respectively. Take 10% - 20% from the DE features and PSD features as the adapter fine-tuning data, and use the remaining DE features and PSD features as the test set. Extract the DE features and PSD features of the source domain as the training set. Step 6: Input the test set and the training set into the multi-branch feature fusion network. Specifically, the multi-branch feature fusion network includes a channel attention weighted spatial convolution branch, a first dynamic graph convolution network branch and a second dynamic graph convolution network branch with the same structure. Input the DE features in the test set and the training set into the first dynamic graph convolution network branch and the channel attention weighted spatial convolution branch respectively, and input the PSD features in the test set and the training set into the second dynamic graph convolution network branch. Step 7: The first dynamic graph convolutional network branch processes the input DE features to obtain time-domain features , and the second dynamic graph convolutional network branch processes the input PSD features to obtain frequency-domain features ; Connect the time-domain features and the frequency-domain features to obtain time-frequency domain features ; The channel attention weighted spatial convolution branch uses channel attention to learn the channel adaptive weights of the DE features. Finally, a spatial convolution block is used to process the weighted multi-channel EEG signals to finally obtain spatial features ; Step eight, input the time-domain and frequency-domain features obtained in step seven into the adaptive Transformer feature fusion network with adapter fine-tuning to obtain high-level time-domain and frequency-domain features H A ; Step nine, combine the high-level time-domain and frequency-domain features and spatial features to obtain time-domain, frequency-domain, and spatial features . Then, input the time-domain, frequency-domain, and spatial features into an adaptive Transformer feature fusion network with adapter fine-tuning to obtain a new output ; Step 10: Input the adapter fine-tuning data obtained in Step 5 into the adapter fine-tuning module of the adaptive Transformer feature fusion network with adapter fine-tuning , and use the obtained result as the output of the adapter fine-tuning data H T ; Step Eleven, respectively take the new output and the output of the adapter fine-tuning data H T through the fully-connected layer and Softmax function to respectively obtain the predicted sentiment probability of H T and the predicted sentiment probability of the maximum value of the predicted sentiment probability of is regarded as the classification result of and this classification result is the predicted sentiment label of H T the maximum value of the predicted sentiment probability of is regarded as the classification result of and this classification result is the predicted sentiment label of Step 12: Calculate the cross-entropy loss function value of the adapter fine-tuning data and the training set according to the predicted emotion probability. Step 13: Repeat Steps 6 - 12. By minimizing the cross-entropy loss function value between the model prediction and the actual label, obtain the trained model of the current subject. Step 14: Use the next subject as the target domain and the remaining subjects as the source domain, and repeat Steps 5 - 13 until each subject has been used as the target domain once, and finally obtain the trained model of each subject.

2. The EEG emotion recognition method based on multi-branch feature fusion according to claim 1, characterized in that, In Step 5, the calculation formulas of the DE features and the PSD features are as follows: In the formula: — DE feature; — Variance; — The natural constant, approximately equal to 2.71828; — The target domain or source domain in the DE feature calculation formula; — mean; — PSD feature; — The target domain or source domain in the PSD feature calculation formula; — Window function; — moment; — Window length; — imaginary unit; — Angular frequency.

3. The EEG emotion recognition method based on multi-branch feature fusion according to claim 2, characterized in that In step seven, the time-domain features and the frequency-domain features are calculated as follows: In the formula: — Sparse adjacency matrix; D — , ∈ D , is a sparse adjacency matrix is the degree matrix of , where i and j are the rows and columns of the sparse adjacency matrix; — Weight matrix; In the formula: — Euclidean distance between electrodes under the 10 - 20 EEG electrode placement conditions according to international standards and ; — Standard deviation of distance; — Threshold of sparsity, taking 0.9; , , , — A pair of DE features and PSD features The weight matrices processed separately; , — ELU non-linear function —DE feature; — PSD feature; , where is the number of EEG channels, is the number of frequency bands.

4. The EEG emotion recognition method based on multi-branch feature fusion according to claim 3, wherein In Step 7, the calculation formula of the channel adaptive weight of the DE features is as follows: In the formula: — Variance of the th channel of the electroencephalogram signal; — Variance of all channels of electroencephalogram signals; — Number of EEG channels; —Time point of electroencephalogram signal; t — moment; —channel at time t c DE feature of; —Channel adaptive weights of DE features; — DE features of all channels of electroencephalogram signals; — Unit multiplication; (·)— Activation function; (·)— Activation function; (·) — Normalization layer; , — The learned weight matrices of two fully connected layers; is matrix multiplication.

5. The EEG emotion recognition method based on multi-branch feature fusion according to claim 4, wherein, In step seven, the spatial feature has the following calculation formula: In the formula: —Spatial features; — Spatial convolution block with a convolution kernel size of ; (·)—Batch Normalization layer.

6. The EEG emotion recognition method based on multi-branch feature fusion according to claim 5, characterized in that In Step 8, the adaptive Transformer feature fusion network with adapter fine-tuning includes a multi-head self-attention mechanism module and an adapter fine-tuning module , and its calculation formula is expressed as follows: In the formula: — Advanced time-domain and frequency-domain features; — Intermediate features in the processing of an adaptive Transformer feature fusion network with adapter fine-tuning; MHSA (∙) - Multi-head attention mechanism; Adapter (∙) - Adapter fine-tuning module, including a linear layer, an ELU non-linear function, and a linear layer connected in sequence; LN (∙) - Batch normalization layer.

7. The EEG emotion recognition method based on multi-branch feature fusion according to claim 6, characterized in that The multi-head self-attention mechanism module MHSA In it, the calculation formula of the self-attention mechanism is: In the formula: — Inquiry; — key; — value; — Self-attention mechanism; — Query weight matrix, ; — Key weight matrix, ; — value weight matrix, ; , — Hyperparameter; — Sparse adjacent matrix; Softmax (∙)— Softmax function.

8. The EEG emotion recognition method based on multi-branch feature fusion according to claim 7, characterized in that, In Step 12, the calculation formula of the cross-entropy loss function is as follows: In the formula: — Loss function; — The actual label of the nth electroencephalogram data; —The prediction emotion probability of the nth electroencephalogram data pair for the j nth class; — Total number of samples; - Class number.

9. An electroencephalogram emotion recognition system based on multi-branch feature fusion, characterized in that, Specifically, it includes the following modules: An EEG data acquisition module, which is used to acquire the EEG data of multiple subjects. A preprocessing module, which is used to preprocess the EEG data of multiple subjects to obtain the preprocessed EEG data of each subject. A frequency band feature extraction module, which is used to extract the features of five frequency bands from the preprocessed EEG data of each subject. The five frequency bands are delta, theta, alpha, beta, and gamma respectively. A feature division module, which is used to use the features of the five frequency bands of the first subject as the target domain and the features of the five frequency bands of the remaining subjects as the source domain. The dataset division module is used to extract the time-domain features and frequency-domain features of the target domain. The time-domain features and frequency-domain features are DE features and PSD features respectively; Take 10% - 20% from the DE features and PSD features as the adapter fine-tuning data, and use the remaining DE features and PSD features as the test set; Extract the DE features and PSD features of the source domain as the training set; The data input module is used to jointly input the test set and the training set into the multi-branch feature fusion network. Specifically: the multi-branch feature fusion network includes a channel attention weighted spatial convolution branch, a first dynamic graph convolution network branch and a second dynamic graph convolution network branch with the same structure; input the DE features in the test set and the training set into the first dynamic graph convolution network branch and the channel attention weighted spatial convolution branch respectively, and input the PSD features in the test set and the training set into the second dynamic graph convolution network branch; A feature fusion module for implementing the following functions: the first dynamic graph convolutional network branch processes the input DE features to obtain time-domain features , and the second dynamic graph convolutional network branch processes the input PSD features to obtain frequency-domain features ; connect the time-domain features and the frequency-domain features to obtain time-frequency domain features ; the channel attention weighted spatial convolution branch uses channel attention weighted learning to obtain the channel adaptive weights of the DE features; finally, use a spatial convolution block to process the weighted multi-channel EEG signals to finally obtain spatial features ; Adaptive Feature Fusion Module 1, used to fuse time-domain and frequency-domain features Input into the Adaptive Transformer Feature Fusion Network with adapter fine-tuning to obtain high-level time-domain and frequency-domain features H A ; Adaptive Feature Fusion Module 2, which combines high-level time-frequency domain features and spatial features to obtain time-frequency domain-spatial features , and then inputs the time-frequency domain-spatial features into an adaptive Transformer feature fusion network with adapter fine-tuning to obtain a new output ; A fine-tuning module for inputting the adapter fine-tuning data obtained by the dataset partitioning module into the adapter fine-tuning module of the adaptive Transformer feature fusion network with adapter fine-tuning , and using the obtained result as the output of the adapter fine-tuning data H T ; The predicted sentiment probability acquisition module respectively takes the new output and the output of the adapter fine-tuning data H T through the fully-connected layer and Softmax functions to respectively obtain the predicted sentiment probability of H T and the predicted sentiment probability of The maximum value of the predicted sentiment probability of is regarded as the classification result of and this classification result is the predicted sentiment label of H T The maximum value of the predicted sentiment probability of is regarded as the classification result of and this classification result is the predicted sentiment label of The cross-entropy loss function value calculation module calculates the cross-entropy loss function values of the adapter fine-tuning data and the training set according to the predicted sentiment probabilities; The training module is used to repeatedly call the dataset division module - the cross-entropy loss function value calculation module in sequence, and obtain the trained model of the current subject by minimizing the cross-entropy loss function value between the model prediction and the actual label; The subject repetition module is used to take the next subject as the target domain and the remaining subjects as the source domain, and repeatedly call the dataset division module - the training module in sequence until each subject has been used as the target domain once, and finally obtain the trained models of each subject.

Citation Information

Patent Citations

  • Emotion recognition method and system fusing prior and automatic electroencephalogram characteristics

    CN112836593A

  • Multi-feature-domain adaptive electroencephalogram emotion recognition method and device

    CN118568538A