Graph network electroencephalogram emotion recognition method based on generality and personality
By constructing an integrated adjacency matrix of commonality and personality graphs in the EEG emotion recognition model, and using bootstrap method and multi-task joint optimization strategy, the problem of difficulty in capturing individual differences and universal response interactions is solved, and the robustness and accuracy of the model are improved.
Patent Information
- Application Number
- CN202411883957.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-20
- Publication Date
- 2025-05-27
AI Technical Summary
Existing EEG emotion recognition models are difficult to effectively capture the complex interactions between individual differences and universal human responses, and the learned universal adjacency matrix may ignore individual-specific information.
A method for identifying EEG emotions based on commonality and personality is proposed. By constructing a seamless integration of commonality map and personality map, and updating commonality maps with bootstrap method, using graph diffusion convolution and multi-task joint optimization strategies, the robustness and accuracy of the model are enhanced.
It improves the robustness and accuracy of the model under different emotional states, can more effectively understand and distinguish emotional states, and adapt to the dynamic interaction between commonality and individuality.
Smart Images

Figure CN120045972A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of electroencephalogram (EEG) emotion recognition. Background Art
[0002] Emotions play a crucial role in human memory, decision-making, and various other cognitive behaviors. The task of emotion recognition from body language and physiological signals has received extensive attention, especially electroencephalogram (EEG) signals that directly reflect brain activities. A large number of studies have shown the advantages of high reliability and objectivity of emotion recognition based on EEG signals. In the field of deep learning, networks such as convolutional neural networks (CNNs), graph neural networks (GNNs), and transformers have been widely used to capture EEG information across spatial, temporal, and frequency domains. In addition, neuroscience research has also elucidated the general and specific spatial patterns in the human brain during emotion processing. Therefore, identifying similarities and differences between different emotions and individuals can effectively improve the performance of EEG emotion recognition models. However, to a large extent, the intricate interaction between individual differences and common human responses in EEG emotion recognition tasks has not been deeply explored. Inspired by the GNN model's ability to capture topological functional relationships between electrodes, this study aims to integrate commonalities and individualities in the graph structure to improve the performance of graph network-based EEG emotion recognition.
[0003] One of the main focuses of GNNs is to determine the relationships between nodes, especially by learning weighted edges to obtain an adjacency matrix, and the quality of the adjacency matrix has a great impact on the performance of GNNs. Usually, the adjacency matrix is constructed based on the distance between two nodes. However, this method is not suitable for EEG-based emotion recognition because electrodes or brain regions that are far apart spatially may still have close functional relationships. To solve this problem, some edges can be designed based on previous medical research. For example, it is generally believed that there are close interconnections between electrode pairs such as FP1 and FP2, O1 and O2, etc. Another method is to incorporate the entire adjacency matrix or the variable responsible for calculating the distance between nodes as a parameter into the model. However, in these models, the learned general adjacency matrix is often used for different people, which may ignore valuable individual-specific information. Thus, this method has the limitation of being unable to capture the subtle differences and unique characteristics existing within each individual.
[0004] Considering this issue, some studies have focused on developing individual-based adjacency matrices, that is, constructing adjacency matrices by manually calculating the correlation coefficients of each electrode for each sample. For example, the Pearson coefficient, cosine similarity, and feature distance between two nodes are calculated as the weights of the edges. Based on these methods, each input sample is processed individually during the training iteration without any information fusion with other samples. The coefficients are completely determined by the results observed by their respective electrodes without using any cross-sample insights. These methods have solved the previous problems to some extent, but often overemphasize individual differences and ignore the commonalities of brain activities among different individuals under the same emotional conditions. Summary of the Invention
[0005] In the present invention, in order to overcome these challenges and simultaneously consider common and individual electroencephalogram information, we propose a graph network-based electroencephalogram emotion recognition method for commonality and individuality. This graph network-based model innovatively constructs an adjacency matrix that seamlessly integrates two key elements, the common graph and the individual graph. These two graphs respectively encompass the common features and the unique individual features in EEG samples, thereby improving the model's ability to understand and distinguish emotional states. Specifically, to identify the common graph that transcends individual differences and applies to everyone under emotional stimuli, we adopt the bootstrap method in the model to update it based on the individual graph. This method can sample the information from the individual graph to the overall population, enabling the conformal graph to not only capture common basic features but also adapt to the dynamic interaction between commonality and individuality, thereby improving the robustness and accuracy of the model in different situations.
[0006] To capture the unique individual features in electroencephalogram samples, a graph learning module is designed in the individual graph step to calculate the individual graph for each input sample using TokenGT and graph diffusion convolution. Transformer-based models have been proven to have the advantage of learning global information and have shown highly competitive performance in various applications. In TokenGT, nodes, edges, and their types are integrated and encoded into a token stream, which then enters the multi-head attention structure, thereby improving the model's emotion recognition ability. To further enrich the high-dimensional feature space, we add two additional encoding components in TokenGT: encoding the original input data to retain its inherent features; position embedding to fuse the position information of the electrodes. At the same time, graph diffusion convolution can utilize the generalized graph diffusion method to discover the hidden connections between electrodes. This method first expands the number of edges connected to each node through a diffusion operation, and then prunes the edges according to the edge weights, thereby reducing the influence of noise and being able to learn more concealed connections.
[0007] Finally, to enhance the learning ability and accelerate the model convergence, we adopt a multi-task joint optimization strategy, combining contrastive learning tasks and self-supervised regression tasks into the emotion recognition task. Compared with single-task learning, using the additional information obtained from other tasks, the multi-task learning strategy can significantly enhance the generalization ability and robustness of the model. In this model, first, a contrastive learning task is designed to compare the learned graph structure with the correlation coefficient graph calculated from the input data, enabling the model to learn the rank correlation coefficient information between EEG signals. This not only learns statistical information but also strengthens the model's understanding of the internal structure of the data. Second, a regression task of reconstructing the input data based on the learned features is added. This method can ensure that the extracted features accurately reflect the representation of the original data while enhancing the model's ability to distinguish subtle differences in the data. BRIEF DESCRIPTION OF THE DRAWINGS
[0008] Figure 1 is the main structure diagram of the graph network EEG emotion recognition method based on commonality and individuality of the present invention DETAILED DESCRIPTION OF THE EMBODIMENTS
[0009] The following further describes in detail the specific embodiments of the present invention in conjunction with the drawings and embodiments.
[0010] Figure 1 is the main structure diagram of the graph network EEG emotion recognition method based on commonality and individuality of the present invention, and its process is as follows:
[0011] 1. Calculate differential entropy features: Differential entropy features can effectively extract the frequency domain information in the electrical signal.
[0012] First, use a Butterworth band-pass filter to obtain data in five frequency bands from the original data, namely the δ (1 - 4 Hz), θ band (4 - 8 Hz), α band (8 - 13 Hz), β band (13 - 30 Hz), and γ band (30 - 50 Hz). Calculate the differential entropy for the signals in each frequency band respectively according to the following formula:
[0013] x de = -∫p(x)log(p(x))dx (1)
[0014] where p(x) is the probability density function of the EEG signal approximated as a normal distribution.
[0015] 2. Encode the data:
[0016] Encode the obtained differential entropy features, common graph structure, and task information into the token stream to obtain the encoded result X E , and the specific calculation steps are shown in the following formula:
[0017]
[0018] In formula (2), MLP represents a linear layer, Emb represents the encoding layer of the pytorch module, and x de represents the input differential entropy data, and MLP 1 represents expanding the dimension of x de and using a linear layer to adjust it to MLP 2 represents expanding the dimension of n represents the number of electrodes, dim represents the number of feature dimensions, represents the real number field, and Emb node represents encoding according to the physical position of the electrodes, and Emb tpye represents encoding nodes as 0 and edges as 1, represents calculating the Laplacian matrix of the commonality graph, G com represents the commonality graph, I is the identity matrix, D is the degree matrix of graph G com and Λ c is a diagonal matrix composed of eigenvalues, U is the corresponding eigenvector matrix, and X E is the adaptive weighted sum of the four encoding results, represents using MLP 1 and MLP 2 to sum the results of encoding the differential entropy features, represents the result after encoding the task type and node position using the encoding module of pytorch, is encoding the commonality graph using the Laplace transform, represents encoding the types of nodes and edges. Encoding this information into the token stream can enable the subsequent multi-head attention module to achieve better performance without having to be designed too complexly, and these encodings also contain the general reaction characteristics of humans under the same emotional conditions represented by the commonality graph.
[0019] 3. Multi-head attention structure:
[0020] In this step, the encoded tokens are input into the multi-head self-attention structure to extract global information for learning the personality graph structure, and its calculation formula is as follows:
[0021]
[0022] Among them, H is the number of attention heads, h is each attention head, and d H is the size of each head. In formula (3), H = 8 and d H = 64, and X Eis the adaptive weighted sum of four encoding results, where QKV represents the Query, Key, and Value parts of the attention structure. represents the weights of these three parts for attention head h. is the weight matrix, which is initialized using the Kaiming method, and X t is the preliminary personality map obtained after passing through the multi-head attention module. The multi-head attention structure divides the encoded features into multiple heads through parallel attention structures, forming multiple sub-feature spaces, which allows the model to focus on feature information in different dimensions and thus capture key information more effectively.
[0023] 4. Graph diffusion convolution step:
[0024] In this step, the preliminary personality map X t is enhanced using graph diffusion convolution, and the specific calculation process is as follows:
[0025]
[0026] F in formula (4) d represents the GDC module, θ HK represents the coefficient using the Gaussian kernel, HK represents the Gaussian kernel, and L represents the generated transformation matrix, where L = AD -1 , u represents the diffusion step number, k represents each diffusion step, and A and D respectively represent the t adjacency matrix and degree matrix of X. Finally, the personality map G ind enters the next step.
[0027] 5. Contrast task step:
[0028] Considering the characteristics of electrical signals and DE features, we chose to use the Spearman coefficient to learn the contrast task of the graph structure to assist the emotion recognition classification task. This coefficient is a rank correlation coefficient that reflects the correlation between the direction and intensity of the trends of two random variables. Different from the Pearson coefficient, the Spearman coefficient aims to describe the monotonic relationship between variables, rather than the linear relationship. The calculation formula of the Spearman coefficient ρ p,q is as follows,
[0029]
[0030] R pl represents the rank of the differential entropy feature of the l-th frequency band of the p-th electrode of, and R ql represents the rank of the differential entropy feature of the l-th frequency band of the q-th electrode of, Denote the differential entropy feature \(x\) of the \(p\)-th electrode de,p as the average rank, denote the differential entropy feature \(x\) of the \(q\)-th electrode de,q as the average rank, \(b\) represents the feature dimension length \(5\) of the differential entropy feature \(x\), then the Spearman coefficient \(\rho\) de can be calculated simply using the following formula: p,q
[0031]
[0032] d l represents the difference in ranks of the \(l\)-th frequency band data pair between the \(p\)-th electrode and the \(q\)-th electrode in \(x\), \(b\) represents the feature dimension length \(5\) of the differential entropy feature \(x\). de <000020> de de p,q
[0033] After calculating the correlation coefficient \(\rho\) between each electrode and other electrodes in turn for the differential entropy feature \(x\) through formula (6) de the obtained correlation coefficient matrix is used as the adjacency matrix \(G\) p,q , and it is compared and calculated with the result obtained from formula (4) for the personality graph \(G\) coeff after passing through two layers of weight - shared graph convolutional networks respectively. The contrast loss calculation method is as follows: ind
[0034]
[0035] In the formula, \(s\) c,s represents the correlation between the results of the graph \(G\) ind and \(G\) coeff through the weight - shared graph convolutional layer. \(s\) c,m represents the correlation between the graph \(G\) ind and the graphs \(G\) ind and \(G\) coeff of other samples through the weight - shared graph convolutional layer. \(N\) represents the number of samples, \(2N\) means there are a total of \(2N\) sample pairs. \(\tau\) is the temperature coefficient, which is taken as \(0.1\) in this article. \(e\) c represents the feature vector output by the graph convolutional layer of sample \(c\) using \(G\) ind , \(e\) s represents the feature vector output by the graph convolutional layer of sample \(c\) using \(G\) coeff . When \(m\in[1,N]\), \(e\) m represents the feature vector output by the graph convolutional layer of sample \(m\) using \(G\) ind . When \(m\in(N,2N]\), \(e\) m represents the feature vector output by the graph convolutional layer of sample \(m - N\) using \(G\) coeff . Then the similarity is calculated in the following way:
[0036]
[0037] 6. Steps of the regression task:
[0038] This step mainly calculates the regression loss to fit the input information, and the differential entropy feature x of the input de successively passes through two adjacent matrices G ind The GCN layer with an output dimension of 32 and a linear layer with an output dimension of 5 attempt to fit the differential entropy feature x de , and the mean squared error loss MSE is used in the present invention to quantify the regression loss:
[0039]
[0040] where x de,i is the input differential entropy feature of the i-th sample, is the result of the model fitting of the i-th sample, and N represents the number of samples.
[0041] 7. Classification step:
[0042] In this step, first calculate the number of different graphs in the target classifier and perform a two-dimensional convolution operation on G ind using the same number of convolutional kernels of size 3*3, with a stride of 1 and a padding of 1. This operation obtains a graph structure suitable for different layers by focusing on local information. Then, use the convolution result to replace the original graph in the classifier in turn while keeping other structures unchanged, and use the modified classifier to perform emotion recognition classification prediction to obtain the prediction result. During the training process, the classification loss adopts the cross-entropy loss, and an adaptive learning method is adopted for the weight coefficients of the three losses in the loss calculation process. The calculation process is as follows:
[0043]
[0044] J in the formula represents the weighted sum of all losses, J mse and J ct As shown in steps 5 and 6, the third term of the polynomial is the calculation formula of the cross-entropy loss, p i represents the true classification of the i-th sample, y i represents the predicted classification of the i-th sample, N represents the number of samples, ||θ|| 2 represents L2 regularization, θ represents the weights and biases of each layer in formulas (2), (3), and (4), w 1 , w 2 , w 3 , w 4 is the adaptive weight of each item, initialized to 1, and w 1 , w 2 , w3 , w 4 and in formula (3) Both iterate in the direction of gradient descent with a step size of 0.005, record the J value obtained from each training. If J is greater than or equal to the J value of this training for 10 consecutive training results after a certain training, then record the J at this time as the minimum value, and record all the parameters in steps 2, 3, 4, and 7 when J takes the minimum value as the parameters used for final prediction. At this time, the training ends.
[0045] 8. Common graph update step
[0046] In this step, after t trainings, all the individual graphs G in this process ind are averaged, and the common graph G is updated according to the bootstrap method com , and the calculation method is as follows:
[0047]
[0048] N is the total number of samples in each training, G ind j,i represents the graph G generated by the i-th sample in the j-th training ind . After experimental tests with multiple changes to α and t, the coefficient α in this step is taken as 0.85, t is taken as 20 when using the SEED and DEAP datasets, and t is taken as 5 when using the SEED-IV dataset.
[0049] The present invention verifies the GAT, DGCNN, GCB - BLS, and MSFR - GCN classifiers on the SEED, SEED - IV, and DEAP datasets through the experimental method of leave - one - out cross - validation. The experiments confirm that the model proposed by the present invention improves the accuracy by 3%, 4%, and 3% respectively on the SEED, SEED - IV, and DEAP datasets, demonstrating its superiority in emotion recognition.
Claims
1. A graph network EEG emotion recognition method based on commonality and individuality, characterized by Include: Personality map step: Encode the input data and the common map and perform transformer operation to obtain a preliminary map, and then perform graph diffusion convolution enhancement on the preliminary map to obtain a personality map; Classification step: Through the comparison task and regression task, the personality graph learns the characteristics of both, and the personality graph is added to the graph convolution classifier through a linear layer to classify the emotions of the input data; Commonality map step: During the training process, the individuality maps of the previous batch are weighted averaged and the bootstrap method is used to update the conformal map for training the next batch of samples.
2. The graph network emotion recognition method based on commonality and individuality according to claim 1 is characterized in that: The personality graph step includes: designing a personality graph learning module composed of a token graph transformer (TokenGT) and a diffusion improved graph convolution (GDC) in sequence, wherein the former first undergoes a step of encoding into a token stream to obtain an encoded result X E , the specific calculation steps are shown in the following formula: In formula (1), W represents the adaptive parameter matrix, MLP represents the linear layer, and x de Represents the input differential entropy data, Emb represents the encoding layer of the pytorch module, represents the Laplacian matrix of the computational commonality graph, G com Represents the commonality graph, the encoded X E After a transformer layer, we get the initial learned personality graph X t , and then use GDC to enhance the personality map. The specific calculation process is as follows: F in formula (2) d represents the GDC module, θ HK represents the coefficients of the Gaussian kernel, and T represents the generated transformation matrix, where T=AD -1 , t represents the number of diffusion steps, A and D represent X t The adjacency matrix and degree matrix of ind Proceed to the next step.
3. The graph network emotion recognition method based on commonality and individuality according to claim 1 is characterized in that: The classification step comprises: Design a comparison task that calculates the Spearman coefficient graph based on the input differential entropy data and compares it with the personality graph to learn the Spearman coefficient information; design a regression task that performs graph convolution fitting on the input differential entropy data based on the personality graph to learn the direct features in the original data; and design the main goal of the emotion recognition task by inserting the personality graph into the graph convolution classifier through a linear layer.
4. The graph network emotion recognition method based on commonality and individuality according to claim 1 is characterized in that: The commonality graph step includes: The personality map generated according to each batch of samples is designed to use the bootstrap method for weighted sampling to calculate the conformal map, where the weights are adaptive weights and are iteratively optimized during the training process; for the calculation of each sample, the same conformal map is used for encoding.
5. The graph network emotion recognition method based on commonality and individuality according to claim 1 is characterized in that: The following steps are involved: 1) Calculate differential entropy features: Differential entropy features can effectively extract frequency domain information from electrical signals; First, a Butterworth bandpass filter is used to obtain data from five frequency bands from the original data, namely, δ (1-4 Hz), θ band (4-8 Hz), α band (8-13 Hz), β band (13-30 Hz) and γ band (30-50 Hz); the differential entropy of the signal in each frequency band is calculated according to the following formula: x de =-∫p(x)log(p(x))dx (3) Where p(x) is the probability density function of the EEG signal that is approximately normally distributed; 2) Encoded data: Encode the obtained differential entropy features, common graph structure, and task information into the token stream to obtain the encoded result X E , the specific calculation steps are shown in the following formula: In formula (2), MLP represents the linear layer, Emb represents the encoding layer of the pytorch module, and x de Represents the input differential entropy data, MLP1 represents x de Perform dimension expansion and use a linear layer to resize to MLP2 said it would Perform dimension expansion and use a linear layer to resize to n represents the number of electrodes, dim represents the number of feature dimensions, Represents the real number field, Emb node Indicates encoding based on the physical location of the electrode, Emb tpye It means that the node is encoded as 0 and the edge is encoded as 1. represents the Laplacian matrix of the computational commonality graph, G com represents the commonality graph, I is the identity matrix, and D is the graph G com The degree matrix of c is a diagonal matrix consisting of eigenvalues, U is the corresponding eigenvector matrix, X E is the adaptive weighted sum of the four encoding results, Represents the sum of the results of using MLP1 and MLP2 to encode the differential entropy features. It indicates the result of encoding the task type and node position using the encoding module of pytorch. is to encode the commonality graph using Laplace transform, Indicates the types of encoded nodes and edges; 3) Multi-head attention structure: The encoded token is input into the multi-head self-attention structure to extract global information and learn the personality graph structure. The calculation formula is as follows: Where H is the number of attention heads, h is the number of attention heads, and d H is the size of each head, H = 8 in formula (3), d H =64,X E It is the adaptive weighted sum of the four encoding results. QKV represents the Query, Key, and Value parts of the attention structure. represents the weights of the three parts of the attention head h, is the weight matrix, Initialize using Kaiming method, X t This is the preliminary personality map obtained after passing through the multi-head attention module; 4) Graph diffusion convolution step: Preliminary Personality Map X t Use graph diffusion convolution for enhancement. The specific calculation process is as follows: F in formula (4) d represents the GDC module, θ HK represents the coefficient of the Gaussian kernel, HK represents the Gaussian kernel, and L represents the generated transformation matrix, where L=AD -1 , u represents the number of diffusion steps, k represents each diffusion step, A and D represent X t The adjacency matrix and degree matrix of ind Go to the next step; 5) Compare task steps: We chose to use the Spearman coefficient to learn the contrast task of graph structure to assist the emotion recognition classification task; the Spearman coefficient ρ p,q The calculation formula is as follows, Represents the differential entropy characteristics of the lth frequency band of the pth electrode The ranking, represents the differential entropy characteristic of the lth frequency band of the qth electrode The ranking, represents the differential entropy characteristic x of the pth electrode de,p The average ranking of represents the differential entropy characteristic x of the qth electrode de,q The average rank of , b represents the differential entropy feature x de The feature dimension length is 5, then the Spearman coefficient ρ p,q The calculation can be simplified using the following formula: d l Represents x de The lth frequency band data pair of the pth electrode and the qth electrode in The difference in rank, b represents the differential entropy feature x de The feature dimension length is 5; The differential entropy feature x de The correlation coefficient ρ between each electrode and other electrodes is calculated in turn by formula (6): p,q The obtained correlation coefficient matrix is used as the adjacency matrix G coeff , and the personality graph G obtained by formula (4) ind The results obtained after passing through the two layers of weight-sharing graph convolutional networks are compared and calculated. The comparison loss is calculated as follows: Where s c,s Represents the same sample graph G ind and G coeff The correlation of the results of the graph convolutional layer through weight sharing, s c,m Representation graph G ind Graph G with other samples ind and G coeff The correlation of the results of the graph convolution layer through weight sharing, N represents the number of samples, 2N means there are 2N sample pairs in total, τ is the temperature coefficient, which is taken as 0.1 in this paper, e c Indicates that sample c uses G ind The feature vector output by the graph convolutional layer, e s Indicates that sample c uses G coeff The feature vector output by the graph convolution layer is e m Indicates that sample m uses G ind The feature vector output by the graph convolution layer is e m Indicates that sample mN uses G coeff The similarity is calculated as follows: 6) Steps of regression task: Calculate the regression loss to fit the input information, the differential entropy feature x of the input de Through two layers of adjacency matrix G ind A GCN layer with an output dimension of 32 and a linear layer with an output dimension of 5 try to fit the differential entropy feature x de , use the mean square error loss MSE to quantify the regression loss: where x de,i is the input differential entropy feature of the i-th sample, is the result of the i-th sample model fitting, and N represents the number of samples; 7) Classification steps: First, calculate the number of different graphs in the target classifier and ind A two-dimensional convolution operation with the same number of convolution kernels of size 3*3, a step size of 1, and a padding of 1 is used. This operation obtains a graph structure suitable for different layers by focusing on local information. The convolution results are then used to replace the graphs in the original classifier in turn while keeping other structures unchanged. The modified classifier is used for emotion recognition classification prediction to obtain the prediction results. During the training process, the classification loss adopts the cross entropy loss, and the weight coefficients of the three losses are adaptively learned in the loss calculation process. The calculation process is as follows: The J in the formula represents the weighted sum of all losses, J mse and J ct As shown in steps 5 and 6, the third term of the polynomial is the calculation formula for the cross entropy loss, p i represents the true classification of the i-th sample, y i represents the predicted classification of the i-th sample, N represents the number of samples, ||θ||2 represents L2 regularization, θ represents the weight and bias of each layer in formula (2), formula (3), and formula (4), w1, w2, w3, and w4 are the adaptive weights of each item, initialized to 1, and in each iteration process w1, w2, w3, and w4 are the weights of each item in formula (3). Iterate in the direction of gradient descent with a step size of 0.005, record the J value obtained in each training, if J is greater than or equal to the J value of the training for 10 consecutive trainings after a certain training, record J at this time as the minimum value, and record all the parameters in steps 2), 3), 4), and 7) when J takes the minimum value as the parameters used for the final prediction, and the training ends at this time; 8) Commonality graph update steps After t trainings, all personality graphs G in this process are ind Average and update the commonality graph G according to the bootstrap method com , the calculation method is as follows: N is the total number of samples in each training, G ind j,i Represents the graph G generated by the i-th sample in the j-th round of training ind , the coefficient α is taken as 0.85, t is taken as 20 when using the SEED and DEAP datasets, and is taken as 5 when using the SEED-IV dataset.