Graph diffusion convolution electroencephalogram emotion recognition method and system
By designing the frequency and time domain coding modules, the adjacency matrix and spatial relationship matrix are constructed, and diffusion convolution and bidirectional diffusion mechanisms are introduced, the existing EEG emotion recognition methods are solved, and more efficient EEG signal emotion classification is achieved.
Patent Information
- Application Number
- CN202510343065.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-21
- Publication Date
- 2025-07-04
AI Technical Summary
The existing EEG sentiment recognition methods have high computational complexity and are not very accurate enough, making it difficult to effectively integrate time-frequency domain features and capture the multi-hop interaction between EEG channels.
Design frequency and time domain coding modules, build adjacency matrix and spatial relationship matrix, introduce diffusion convolution and bidirectional diffusion mechanisms, simulate random walk process, and combine the full connection layer and the activation layer to perform multi-channel EEG signal emotional classification.
The accuracy and robustness of the model are improved, and the multi-hop interaction relationship between EEG channels can be better captured, the spatial characteristics of EEG signals can be fully considered, and the computational complexity can be reduced.
Smart Images

Figure CN120257091A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of electroencephalogram (EEG) signal processing and emotion classification, and in particular to a graph diffusion convolution EEG emotion recognition method and system. Background Art
[0002] As the macroscopic manifestation of the discharge activity of the brain's neuron groups, EEG signals contain rich time-frequency domain information. In the time domain, EEG signals have obvious dynamic characteristics and long-term dependence; in the frequency domain, EEG oscillations in different frequency bands reflect different neural activity patterns. For example, EEG activity in the δ band (1-4Hz) is usually associated with a deep sleep state; the α band (8-13Hz) is more obvious in a quiet, eyes-closed state; the β band (13-30Hz) is closely related to wakefulness, alertness, and cognitive activities; the γ band (30-50Hz) is prominent in advanced cognitive functions and when attention is focused. These changes in EEG signals in different frequency bands provide important clues for the analysis of emotional states.
[0003] Traditional methods for emotion classification of EEG signals often have many limitations. On the one hand, many methods process time domain or frequency domain features independently, and fail to fully utilize the complementarity of time and frequency information. For example, early studies may only focus on the signal amplitude changes in the time domain or the power spectrum density in the frequency domain, while ignoring the intrinsic connection between the two. Even if some studies try to simply superimpose or concatenate time and frequency features, this mechanical fusion method is difficult to explore the deep relationship between features and cannot effectively improve the accuracy of emotion classification. On the other hand, EEG signals have significant spatial characteristics, and there are complex functional connection relationships between different electrodes. At present, mainstream graph neural networks are mostly based on spectral convolution, which maps signals to the spectral domain through Fourier transform for convolution operations. However, spectral convolution has shortcomings in processing directed relationship graphs and dynamic interactions, and the computational complexity is also high. The interaction between EEG signal channels is directional and dynamic. For example, under emotional stimulation, the transmission of neural activity between different brain regions has a specific direction and time order, and spectral convolution is difficult to capture this complex relationship, which limits the effective modeling of the spatial characteristics of EEG signals.
[0004] With the continuous development of artificial intelligence technology, higher requirements are put forward for the accuracy and efficiency of electroencephalogram (EEG) signal emotion classification. How to effectively fuse time-frequency domain features and explore a graph convolutional mechanism more suitable for EEG signals has become a key issue in current research. Chinese Patent Publication No. CN117370828A discloses a multi-modal feature fusion emotion recognition method based on a gated cross-attention mechanism, which respectively designs a time-frequency domain dual-branch dynamic graph convolutional feature extraction network for EEG signals and a multi-dimensional feature encoding network for eye movement signals; constructs a multi-modal feature fusion module based on the gated cross-attention mechanism. The advantages of this method include: designing a time-frequency domain dual-branch dynamic graph convolutional EEG signal feature extraction network, which makes up for the problem of insufficient single-domain feature extraction in traditional methods; introducing the gated cross-attention mechanism to achieve effective fusion of EEG features and eye movement features, and improving the model recognition accuracy. However, the patent application uses graph convolution with high computational complexity and is difficult to capture multi-hop interaction relationships between EEG channels, thus not fully considering the spatial characteristics of EEG signals, resulting in a not-high-enough model accuracy. Summary of the Invention
[0005] The technical problem to be solved by the present invention lies in the problems of high computational complexity and insufficient accuracy of the existing EEG emotion recognition method.
[0006] The present invention solves the above technical problems through the following technical means: A graph diffusion convolutional EEG emotion recognition method, including:
[0007] S1. Design a frequency domain encoding module and a time domain encoding module to obtain frequency domain feature representations and time domain feature representations respectively, and fuse the two to obtain a fused feature;
[0008] S2. Construct an adjacency matrix representing the implicit connection pattern between different electrodes and a spatial relationship matrix representing the spatial relationship between electrodes, fuse the adjacency matrix and the spatial relationship matrix to obtain a final adjacency matrix, and use L1 regularization to constrain the final adjacency matrix;
[0009] S3. Based on the final adjacency matrix and the fused feature, introduce diffusion convolution and a bidirectional diffusion mechanism, simulate the random walk process on the graph, capture the multi-hop interaction relationships between nodes, and combine a fully connected layer and an activation layer to finally achieve emotion classification of multi-channel EEG signals.
[0010] Beneficial effects: The present invention adaptively integrates time domain and frequency domain features, fully utilizes the mutual relationship between the two, and thus obtains a more comprehensive EEG signal representation. This fusion strategy not only improves the complementarity of features, but also enhances the robustness and generalization ability of the model. By introducing diffusion convolution, the computational complexity is low, and by simulating the random walk process, it can better capture the multi-hop interaction relationships between EEG channels, fully consider the spatial characteristics of EEG signals, and the accuracy of the model is relatively high.
[0011] Further, S1 includes:
[0012] S11. The frequency-domain encoding module obtains the weighted features of each frequency band after weighting the input features and passing them through an activation function. The weighted features of each frequency band are fused with the input features to obtain a frequency-domain feature representation.
[0013] S12. The time-domain encoding module uses a GRU network to obtain the hidden state at the current moment and constructs a time-domain feature representation using the hidden state.
[0014] S13. The frequency-domain feature representation and the time-domain feature representation are fused with a gating coefficient to obtain a fused feature.
[0015] Even further, S11 includes:
[0016] The frequency-domain encoding module multiplies the input features element-wise with the features obtained by weighting the input features, passing them through a ReLU activation function, weighting them again, and then passing them through a Softmax activation function to obtain the weighted features of each frequency band. The features obtained by weighting the weighted features of each frequency band and passing them through another ReLU activation function are added to the weighted input features to obtain a frequency-domain feature representation.
[0017] Even further, S12 includes:
[0018] The time-domain encoding module inputs the input features at the current moment and the hidden state at the previous moment into the GRU network to obtain the hidden state at the current moment. The GRU network accumulates the outputs of the hidden states at each time step to obtain a time-domain feature representation, and this time-domain feature representation is processed by a fully connected layer and a normalization layer to make its dimension consistent with the target dimension of the frequency-domain feature representation.
[0019] Even further, S13 includes:
[0020] The sum of the frequency-domain feature representation and the time-domain feature representation is input into a one-dimensional convolutional layer, and then a gating coefficient is obtained through a Sigmoid activation function. The product of the gating coefficient and the frequency-domain feature representation plus the product of the time-domain feature representation and the gating parameter is used to obtain a fused feature, where the gating parameter is 1 minus the gating coefficient.
[0021] Further, S2 includes:
[0022] The product of the adjacency matrix and the first weighting coefficient is added to the product of the spatial relationship matrix and the second weighting coefficient to obtain the final adjacency matrix, where the second weighting coefficient is 1 minus the first weighting coefficient.
[0023] Further, the element in the i-th row and j-th column of the spatial relationship matrix d i,jis the distance between electrode i and electrode j calculated according to the standard 10-20 EEG electrode placement method, and δ is the calibration constant of the distance.
[0024] Further, S3 includes:
[0025] Introduce bidirectional diffusion convolution, that is, simultaneously simulate the forward and backward random walks to obtain the random walk result:
[0026]
[0027] Among them, D O represents the out-degree matrix, which is the sum of the out-edge weights of electrode i of the EEG signal. D O is a diagonal matrix, and its diagonal elements A ij represents the element in the i-th row and j-th column of the final adjacency matrix. D I represents the in-degree matrix, which is the sum of the in-edge weights of electrode i of the EEG signal. D I is a diagonal matrix, and its diagonal elements satisfy K is the order of the diffusion convolution, k is the k-th order diffusion convolution, and θ k,O is the learning parameter of the out-degree matrix, and θ k,I is the learning parameter of the in-degree matrix. A represents the final adjacency matrix, and X represents the fused feature;
[0028] Random walk result After passing through the alternately stacked fully connected layers and activation layers, the final emotion classification result is obtained.
[0029] The present invention also provides a graph diffusion convolution EEG emotion recognition system, including:
[0030] A node feature extraction module, which is used to design a frequency domain encoding module and a time domain encoding module to obtain the frequency domain feature representation and the time domain feature representation respectively, and fuse the two to obtain the fused feature;
[0031] An edge feature extraction module, which is used to construct an adjacency matrix representing the implicit connection pattern between different electrodes and a spatial relationship matrix representing the spatial relationship between electrodes, fuse the adjacency matrix and the spatial relationship matrix to obtain the final adjacency matrix, and use L1 regularization to constrain the final adjacency matrix;
[0032] A diffusion convolution module, which is used to introduce diffusion convolution and a bidirectional diffusion mechanism based on the final adjacency matrix and the fused feature, simulate the random walk process on the graph, capture the multi-hop interaction relationship between nodes, and combine fully connected layers and activation layers to finally realize the emotion classification of multi-channel EEG signals.
[0033] Further, the node feature extraction module is also used for:
[0034] S11. The frequency-domain encoding module obtains the weighted features of each frequency band after weighting the input features and passing them through an activation function. The weighted features of each frequency band are fused with the input features to obtain a frequency-domain feature representation;
[0035] S12. The time-domain encoding module uses a GRU network to obtain the hidden state at the current moment and constructs a time-domain feature representation using the hidden state;
[0036] S13. The frequency-domain feature representation and the time-domain feature representation are weighted and fused using a gating coefficient to obtain a fused feature.
[0037] Furthermore, S11 includes:
[0038] The frequency-domain encoding module multiplies the input features element-wise with the features output after weighting the input features by the ReLU activation function, then weighting again by the Softmax activation function to obtain the weighted features of each frequency band. The features obtained by weighting the weighted features of each frequency band by another ReLU activation function and adding them to the weighted input features result in a frequency-domain feature representation.
[0039] Furthermore, S12 includes:
[0040] The time-domain encoding module inputs the input features at the current moment and the hidden state at the previous moment into the GRU network to obtain the hidden state at the current moment. The GRU network accumulates the outputs of the hidden states at each time step to obtain a time-domain feature representation, and this time-domain feature representation is processed through a fully connected layer and a normalization layer to make its dimension consistent with the target dimension of the frequency-domain feature representation.
[0041] Furthermore, S13 includes:
[0042] The sum of the frequency-domain feature representation and the time-domain feature representation is input into a one-dimensional convolutional layer, and then a gating coefficient is obtained through a Sigmoid activation function. The product of the gating coefficient and the frequency-domain feature representation plus the product of the time-domain feature representation and the gating parameter results in a fused feature, where the gating parameter is 1 minus the gating coefficient.
[0043] Further, the edge feature extraction module is also used for:
[0044] Adding the product of the adjacency matrix and the first weighting coefficient to the product of the spatial relationship matrix and the second weighting coefficient to obtain the final adjacency matrix, where the second weighting coefficient is 1 minus the first weighting coefficient.
[0045] Further, the element in the i-th row and j-th column of the spatial relationship matrix d i,j is the distance between electrode i and electrode j calculated according to the standard 10 - 20 EEG electrode placement method, and δ is a calibration constant for the distance.
[0046] Furthermore, the diffusion convolution module is also used for:
[0047] Introducing bidirectional diffusion convolution, that is, simultaneously simulating forward and backward random walks to obtain the random walk result:
[0048]
[0049] where D O represents the out-degree matrix, which is the sum of the out-edge weights of electrode i of the EEG signal. D O is a diagonal matrix, and its diagonal elements A ij represents the element in the i-th row and j-th column of the final adjacency matrix. D I represents the in-degree matrix, which is the sum of the in-edge weights of electrode i of the EEG signal. D I is a diagonal matrix, and its diagonal elements satisfy K is the order of the diffusion convolution, k is the k-th order diffusion convolution, and θ k,O is the learning parameter of the out-degree matrix, and θ k,I is the learning parameter of the in-degree matrix. A represents the final adjacency matrix, and X represents the fused features;
[0050] The random walk result Passes through alternately stacked fully connected layers and activation layers to obtain the final emotion classification result.
[0051] The advantages of the present invention are as follows:
[0052] (1) The present invention adaptively integrates time-domain and frequency-domain features, fully utilizes the mutual relationship between the two, and thus obtains a more comprehensive EEG signal representation. This fusion strategy not only improves the complementarity of features but also enhances the robustness and generalization ability of the model. By introducing diffusion convolution, the computational complexity is low, and by simulating the random walk process, it can better capture the multi-hop interaction relationship between EEG channels, fully considering the spatial characteristics of EEG signals, and the accuracy of the model is relatively high.
[0053] (2) When constructing the graph structure, the present invention combines data-driven and spatial prior information. It not only captures the implicit connection pattern between electrodes through a trainable adjacency matrix but also models the spatial distribution of electrodes through physical distance. This hybrid strategy enhances the ability to model the spatial characteristics of EEG signals, enabling the model to better understand the functional association and physical layout characteristics between electrodes. BRIEF DESCRIPTION OF THE DRAWINGS
[0054] Figure 1 It is a flowchart of a graph diffusion convolution EEG emotion recognition method disclosed in an embodiment of the present invention;
[0055] Figure 2 Schematic diagram of the frequency-domain feature encoding process in a graph diffusion convolutional electroencephalogram emotion recognition method disclosed in an embodiment of the present invention;
[0056] Figure 3 Schematic diagram of the structure of the GRU network in step S12 of a graph diffusion convolutional electroencephalogram emotion recognition method disclosed in an embodiment of the present invention;
[0057] Figure 4 Schematic diagram of the diffusion convolution process in a graph diffusion convolutional electroencephalogram emotion recognition method disclosed in an embodiment of the present invention. Detailed implementation manners
[0058] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the embodiments of the present invention. Apparently, the described embodiments are some, but not all, of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.
[0059] Embodiment 1
[0060] As Figure 1 shown, Embodiment 1 of the present invention provides a graph diffusion convolutional electroencephalogram emotion recognition method. First, the electroencephalogram signal is preprocessed, including downsampling, removing electrooculogram artifacts, and band-pass filtering. Then, features are extracted from two dimensions of time domain and frequency domain. The frequency-domain features are obtained by calculating the differential entropy of different frequency bands, and the time-domain features are extracted by the GRU network. Next, a graph input is constructed. The node features are extracted by the frequency-domain encoding module and the time-domain encoding module respectively, and the time-frequency features are adaptively fused through a gating mechanism. The edge features are combined with data-driven and spatial prior information to construct an adjacency matrix. Subsequently, diffusion convolution is introduced to replace the traditional spectral graph convolution, and the multi-hop interaction relationship between electroencephalogram channels is captured by simulating the random walk process. The bidirectional diffusion convolution is combined to enhance the model's ability to capture complex interaction relationships. The emotion classification part transforms the output of the diffusion convolution layer through a fully connected layer and an activation layer, and finally outputs the classification result. The loss function consists of cross-entropy loss, L2 regularization, and L1 regularization. Specifically, it includes the following steps:
[0061] S1. By designing a frequency-domain encoding module and a time-domain encoding module, and fusing the outputs of the two, a unified node feature representation is obtained; the specific process is as follows:
[0062] S11. As Figure 2 shown, through the frequency-domain encoding module, using the attention mechanism and residual connection, the features of different frequency bands are extracted and fused, and finally a robust frequency-domain feature representation X is generatedF , for subsequent electroencephalogram (EEG) signal analysis tasks:
[0063] X F = ReLU(W3·X w ) + W r ·X f (1)
[0064] In Equation (1), and are both weight coefficients, D is the target feature dimension, and the input feature The final output frequency-domain feature representation Fuses band attention and deep features; among them, the feature X weighted by each band w is defined as:
[0065] X w = X f ⊙ Softmax(W2·ReLU(W1·X f ))(2)
[0066] In Equation (2), the input feature and are learnable parameters, and H is the hidden layer dimension. Softmax and ReLU are two activation functions. ReLU is used to set negative values to zero, and Softmax is used for normalization. ⊙ represents the element-wise multiplication operation, Characterizes the feature weighted by each band.
[0067] S12. Through the time-domain encoding module based on GRU, capture the dynamic characteristics and long-term dependencies of EEG signals, and after processing by the fully connected layer and the normalization layer, finally generate the time-domain feature representation X T , providing a basis for subsequent time-frequency feature fusion;
[0068] As Figure 3 shown, GRU (Gated Recurrent Unit) has high efficiency and stability in sequence modeling, and its operation process can be represented by the following formula:
[0069] r t = σ(W r x t + U r h t-1 + b r )(3)
[0070] Z t = σ(W z x t + U z h t-1 + b z )(4)
[0071]
[0072]
[0073] Among them, σ is the Sigmoid function, and tanh is the hyperbolic tangent function. Both are used for non-linear activation. r t represents the reset gate, whose function is to determine how much previous state information should be discarded when generating the candidate hidden state. Z t represents the update gate, which is used to control the update ratio of the final hidden state. W and U represent trainable weight matrices, and b represents the bias.
[0074] X T is the time-domain feature representation extracted and transformed by the GRU network. The operation of the GRU network can be expressed as:
[0075] h t = GRU(h t-1 , x t ) (7)
[0076] In equation (7), h t and h t-1 are the hidden layer states at the current time t and the previous time respectively, and x t represents the input at the current time. For the given electroencephalogram signal input where S represents the number of time instants, that is, the signal sampling frequency per unit time. 1 represents the extended time-domain feature dimension.
[0077] X T = [h1, h2,..., h t T
[0078] The GRU network accumulates the hidden layer outputs step by step in time to obtain the time-domain feature output of the electroencephalogram signal where d represents the hidden layer dimension of the GRU. Then, use the fully connected layer and the normalization layer to transform X T into the time-domain feature representation D represents the same target dimension as the frequency-domain coding module.
[0079] S13. Through the feature fusion strategy based on the gating mechanism, the contribution ratios of the time-domain and frequency-domain features are adaptively adjusted to achieve dynamic balance and effective fusion, and finally the fused feature X fused is generated, improving the stability and generalization ability of the feature representation;
[0080] For the time-domain feature representation and the frequency-domain feature representation A feature fusion strategy based on a gating mechanism is proposed. Specifically, first, two sets of features are added together, and then the addition result is input into a gating module composed of a one-dimensional convolutional layer and a Sigmoid activation function, where the convolutional kernel size is 1×1. This module aims to capture the local interaction relationships between different channels and generate a gating coefficient λ with a value range in [0,1]. Formally, the fusion process can be defined as:
[0081] λ = σ(Conv1d(X F +X T )) (8)
[0082] X fused = λ·X F +(1 - λ)·X T (9)
[0083] where σ(·) represents the Sigmoid activation function, Conv1d represents one-dimensional convolution, is the fused feature. This method adaptively adjusts the contribution ratio of time-domain and frequency-domain features on different channels through the learned gating coefficient λ, thereby achieving the dynamic balance and effective fusion of the two sets of features. Compared with the linear combination strategy with fixed fusion weights, this gating mechanism can make more full use of the complementarity between time-frequency features, and further improve the stability and generalization ability of feature fusion.
[0084] S2. By fusing a data-driven trainable adjacency matrix and a physical-distance-based spatial relationship matrix, and introducing L1 regularization constraints, an adjacency matrix A that combines functional association and physical layout characteristics is constructed F . That is, an adjacency matrix representing the implicit connection pattern between different electrodes and a spatial relationship matrix representing the spatial relationship between electrodes are constructed, the adjacency matrix and the spatial relationship matrix are fused to obtain the final adjacency matrix, and the final adjacency matrix is constrained by L1 regularization; the specific process is as follows:
[0085] A construction strategy of the adjacency matrix that combines data-driven and spatial prior information is designed, aiming to fully reflect the functional association and physical layout characteristics between electrodes. On the one hand, the present invention constructs a trainable adjacency matrix where C represents the number of channels. A l is initialized as the identity matrix I, that is, each electrode node is only connected to itself. During the model training process, A l participates in backpropagation as a learnable parameter, and its elements are dynamically adjusted for the adjacency matrix through backpropagation and the gradient descent algorithm. This construction method calculates the gradient information of A l through the L1 regularization loss function, and combines the Adam optimization algorithm to iteratively update the matrix elements, enabling it to adaptively capture the implicit functional connections between electrodes. Al The non-zero elements quantify the functional connection strength between electrodes, thus better uncovering the dynamic interaction relationships between EEG signal channels under emotional excitation, rather than relying on fixed physical distances or prior knowledge. A l During the training process, it can adaptively capture the implicit connection patterns between different electrodes. On the other hand, since the spatial relationship between electrodes is crucial for understanding brain activity patterns. Based on this understanding, the present invention adopts a spatial relationship modeling method based on physical distance to describe the spatial distribution of electrodes. For EEG signals containing C channels, a spatial relationship matrix is constructed whose element a i,j is defined as:
[0086]
[0087] where each row and each column of the spatial relationship matrix are the serial numbers of electrodes, the elements in the matrix are the spatial relationships between the corresponding electrodes, and d i,j is the distance between electrode i and electrode j calculated according to the standard 10-20 EEG electrode placement method, and δ is the calibration constant of the distance. The present invention sets δ to 5 to maintain about 20% effective connection. The two are fused into the final adjacency matrix in a weighted linear combination manner whose mathematical expression is:
[0088] A F = α * A l +(1 - α) * A s (11)
[0089] where A l represents the connection information obtained by data-driven learning, A s characterizes the prior structure based on physical distribution, and α ∈ [0, 1] is used to balance the contributions of these two parts of information. To avoid increasing the computational complexity and to alleviate the overfitting problem. The present invention uses the L1 regularization method to constrain the fused adjacency matrix A F
[0090]
[0091] where, ‖·‖1 represents the L1 norm, and the value of the regularization coefficient λ1 is set to 1e-3. It directly acts on A F to ensure that the fused adjacency matrix reflects both the internal pattern of the data and maintains sparsity.
[0092] S3. Based on the final adjacency matrix and fusion features, by introducing diffusion convolution and bidirectional diffusion mechanism, the random walk process on the graph is simulated to capture the multi-hop interaction relationships between nodes, and combined with the fully connected layer and activation layer, the emotional classification of multi-channel EEG signals is finally realized. The specific process is as follows:
[0093] As Figure 4 shown, the core idea of diffusion convolution is to model the propagation process of graph signals as a random walk process, and capture the multi-hop interaction relationships between nodes by simulating the diffusion of information in the graph. Different from spectral graph convolution which depends on the spectral decomposition of the graph, diffusion convolution directly propagates information between nodes through the number of diffusion steps, has higher flexibility, and is especially suitable for directed graph structures. The diffusion convolution operation can be expressed by Equation (10):
[0094]
[0095] where D O represents the out-degree matrix, which represents the sum of the out-edge weights of node i, and satisfies D ii = ∑ j A ij . θ k is a learnable parameter used to control the weight of each diffusion step. K represents the number of diffusion steps set to 2, which can capture the dependencies between channels within the multi-hop range and determines the propagation range of information in the graph. To further improve the flexibility of the model, the present invention introduces bidirectional diffusion convolution, that is, simultaneously considering the forward and backward random walks:
[0096]
[0097] where, D O represents the out-degree matrix, which represents the sum of the out-edge weights of electrode i of the EEG signal, D O is a diagonal matrix, and the elements on its diagonal A ij represents the element in the i-th row and j-th column of the final adjacency matrix. Each row and column of the final adjacency matrix is the serial number of the electrode, and the elements in this matrix represent the physical connection relationship between the corresponding electrodes. D I represents the in-degree matrix, which represents the sum of the in-edge weights of electrode i of the EEG signal, D I is a diagonal matrix, and the elements on its diagonal satisfy K is the order of the diffusion convolution, k is the k-th order diffusion convolution, θ k,O is the learning parameter of the out-degree matrix, θ k,Iis the learning parameter of the in-degree matrix, A represents the final adjacency matrix, and X represents the fused features. Specifically, this bidirectional diffusion mechanism enables the model to simultaneously capture the forward and backward influences between nodes, thereby more comprehensively understanding the complex interaction patterns between different brain regions in EEG signals. To convert the features into sentiment classification, the output features are transformed by alternately stacking fully connected layers and activation layers to obtain the final sentiment classification.
[0098] To verify the effectiveness of the method of the present invention, two experimental strategies, namely subject-dependent and subject-independent, are used to verify the effectiveness of the proposed TSGDN model in multiple aspects.
[0099] 1) Subject-dependent experiment: In the subject-dependent experiment, the data of each subject is used as the experimental data. Since the training and testing are only for a single subject and the dataset division is relatively balanced, the subject-dependent experiment often has a high accuracy rate. For each subject in the SEED dataset, in this section of the experiment, the first 9 trials of each subject are used as the training set, and the remaining 6 trials are used as the test set. For the DEAP dataset, in this section, a ten-fold cross-validation experimental strategy is adopted, that is, for the 40 trials of a subject, they are divided into 10 equal groups. In each experiment, 36 samples are selected as the training set for model training, and the remaining 4 samples are used as the test set to evaluate the performance of the model on the current group of data.
[0100] The TSGDN model achieved good performance in the subject-dependent experiment on the SEED dataset, with an average accuracy rate reaching 92.47%. And the model demonstrated good stability in experiments with different subjects, with a standard deviation of 6.90%. Among them, the accuracy rate of subject No. 15 was the highest, reaching 99.13%. The accuracy rates of subject No. 6 and subject No. 1 were also very high, being 98.99% and 98.48% respectively. However, the accuracy rates were relatively low on subject No. 11, subject No. 13, subject No. 10, subject No. 5, and subject No. 2, being 78.03%, 80.64%, 85.19%, 87.21%, and 87.93% respectively, which were lower than the average level. And the range between different subjects was relatively large, reaching 20.96%. This may be due to the natural differences in the response patterns of different subjects' brains to emotional stimuli, or there may be some individual-specific interference factors during the experiment, which affected the acquisition of EEG signals and the classification effect of the model. But overall, the accuracy rates of most subjects were concentrated between 90% and 95%. This shows that the TSGDN model can achieve good sentiment classification for most subjects.
[0101] Table 1 Comparison of test results of different models in the subject-dependent experiment on the SEED dataset
[0102]
[0103] As shown in Table 1, in the subject-dependent experiment of the SEED dataset, the TSGDN model achieved an accuracy of 92.47% and a standard deviation of 6.90%, outperforming other comparison models. The accuracies of the traditional machine learning methods SVM and DBN were 83.99% and 86.08% respectively, and the standard deviations were 9.72% and 8.34% respectively. This indicates that it is difficult to fully extract the emotional features in EEG signals simply by relying on feature engineering and shallow models. The deep learning methods based on graph structures demonstrated stronger feature extraction capabilities, where the accuracies of GAT+MLP and DGCNN reached 87.56% and 90.40% respectively. However, the standard deviations of these models were still relatively high, 8.22% and 8.49% respectively, indicating that there were significant fluctuations in their performance across different subjects. In contrast, TSGDN not only had a significant improvement in accuracy but also showed better stability. Specifically, TSGDN had an accuracy improvement of 2.07 percentage points compared to the closest TSGDN model, while reducing the standard deviation by 1.59 percentage points. This comprehensive performance improvement can be attributed to the following aspects: First, TSGDN can better model the multi-hop interaction relationships between EEG channels through diffusion convolution; second, the fusion of time-frequency domain features provides richer feature representations; finally, the graph structure design combined with spatial prior information strengthens the model's characterization of the spatial properties of EEG signals. The experimental results show that TSGDN has significant advantages in the EEG signal emotion classification task, providing new methodological ideas for related research.
[0104] Table 2 Experimental results of the TSGDN model and the baseline models in the subject-dependent experiment of the valence dimension (DEAP-V) and arousal dimension (DEAP-A) of the DEAP dataset
[0105]
[0106] Table 2 shows the results of the subject-dependent experiments of the TSGDN model and the baseline model in the valence dimension (DEAP-V) and arousal dimension (DEAP-A) of the DEAP dataset. In both dimensions, the TSGDN model achieved significantly better performance than the baseline model. In the valence dimension, the accuracies of the traditional machine learning methods SVM and DBN were only 51.60% and 51.58% respectively, and the standard deviations were 6.32% and 6.35% respectively. The deep learning methods based on graph structure showed obvious advantages. The accuracies of GAT+MLP and DGCNN were improved to 82.01% and 86.32% respectively, and the standard deviations were 8.77% and 6.04% respectively. The TSGDN model further improved the accuracy to 91.03%, while achieving a low standard deviation of 3.60%, demonstrating excellent performance and stability. In the arousal dimension, the performance of each model was similar but slightly different from that in the valence dimension. The accuracies of SVM and DBN were 55.80% and 62.43% respectively, with relatively large standard deviations of 9.49% and 11.87% respectively. The performance of GAT+MLP and DGCNN was significantly improved, and the accuracies reached 81.58% and 83.68% respectively, and the standard deviations were reduced to 9.77% and 5.68% respectively. The TSGDN model also performed well in this dimension, achieving an accuracy of 90.47% and a standard deviation of 4.70%.
[0107] 2) Subject-independent experiment: In the subject-independent experiment, the data of one subject is used as the test set, and the data of the remaining subjects is used as the training set. This method is also called Leave One Subject Out. Since the training and test data come from different subjects, the accuracy of the subject-independent experiment is often low.
[0108] Table 3 Experimental results of the subject-independent experiment
[0109]
[0110] Traditional machine learning methods perform poorly on this task. The accuracies of SVM and DBN are only 48.58% and 51.64% respectively, with standard deviations of 4.24% and 7.84% respectively. This indicates that traditional machine learning methods have obvious limitations in dealing with cross-subject emotion classification tasks. Deep learning methods based on graph structures demonstrate stronger performance. The GAT+MLP model achieves an accuracy of 77.56% with a standard deviation of 8.22%; the DGCNN model further improves the accuracy to 79.95%, but its standard deviation slightly increases to 9.02%. This shows that graph neural networks can better capture the spatial relationships and emotion features in EEG signals. Benefiting from the feature extraction ability of diffusion graph convolution, the TSGDN model performs best among all the compared methods, with an accuracy of 81.94% and a relatively low standard deviation of 6.60%. Compared with the DGCNN model with the closest performance, TSGDN not only improves the accuracy by 1.99 percentage points but also reduces the standard deviation by 2.42 percentage points. This result indicates that the TSGDN model has achieved significant improvements in both cross-subject generalization performance and model stability through innovative designs such as diffusion convolution and time-frequency feature fusion.
[0111] Table 4 Performance comparison between the TSGDN model and the baseline models in the valence dimension (DEAP-V) and arousal dimension (DEAP-A) in the subject-independent experiment on the DEAP dataset
[0112]
[0113] Table 4 shows the performance comparison between the TSGDN model and the baseline models in terms of valence dimension (DEAP-V) and arousal dimension (DEAP-A) in the subject-independent experiment of the DEAP dataset. In the valence dimension, traditional machine learning methods performed poorly. The accuracies of SVM and DBN were 48.58% and 51.64% respectively, and the standard deviations were 4.24% and 7.84% respectively. The performance of deep learning methods based on graph structure was improved. The accuracies of GAT+MLP and DGCNN reached 55.01% and 58.46% respectively, and the standard deviations were 8.77% and 7.85% respectively. The TSGDN model performed best, with an accuracy of 64.60% and a standard deviation of 7.50%, which was 6.14 percentage points higher than that of DGCNN. In the arousal dimension, the performance of each model was generally higher than that in the valence dimension. The accuracies of SVM and DBN were 50.75% and 56.68% respectively, and the standard deviations were 4.87% and 13.28% respectively. The graph neural network models showed better performance. GAT+MLP and DGCNN reached accuracies of 60.58% and 61.65% respectively, and the standard deviations were 9.77% and 13.34% respectively. The TSGDN model also performed best in this dimension, with an accuracy of 63.09% and the standard deviation dropped to 7.61%, showing better stability. Overall, the TSGDN model outperformed other baseline methods in both dimensions, and achieved a lower standard deviation while maintaining a high accuracy, indicating its stronger generalization ability and stability in the cross-subject emotion recognition task. It is worth noting that the performance of all models in the arousal dimension was generally better than that in the valence dimension, which may be related to the fact that the arousal state has more significant characteristic patterns in EEG signals.
[0114] In summary, the present invention can extract and fuse EEG signal features from multiple perspectives, improve the understanding and extraction ability of the global features of EEG signals. This multi-dimensional information integration improves the stability and accuracy of emotion classification, enabling the present invention to better identify emotional states.
[0115] With the above technical solutions, traditional methods of the present invention usually process time-domain or frequency-domain features independently, ignoring the complementarity between time-frequency information. By designing a gating fusion mechanism, the present invention adaptively integrates time-domain and frequency-domain features, fully utilizes the mutual relationship between the two, and thus obtains a more comprehensive EEG signal representation. This fusion strategy not only enhances the complementarity of features, but also improves the robustness and generalization ability of the model. Traditional spectrogram convolution has limitations in processing directed graphs and dynamic interactions, and has a high computational complexity. The present invention introduces diffusion convolution, which can better capture the multi-hop interaction relationships between EEG channels by simulating the random walk process. Diffusion convolution is not only applicable to directed graph structures, but also can more naturally process dynamic interactions, thereby improving the model's ability to model the spatial characteristics of EEG signals. In the subject-independent experiment, the TSGDN model demonstrated strong cross-subject generalization ability. Although the accuracy of the subject-independent experiment decreased compared to the subject-dependent experiment, TSGDN still maintained a relatively stable performance, indicating its strong adaptability and generalization ability in processing data of different individuals.
[0116] Embodiment 2
[0117] Based on Embodiment 1, Embodiment 2 of the present invention further provides a graph diffusion convolution EEG emotion recognition system, including:
[0118] A node feature extraction module, configured to design a frequency-domain encoding module and a time-domain encoding module to obtain a frequency-domain feature representation and a time-domain feature representation respectively, and fuse the two to obtain a fused feature;
[0119] An edge feature extraction module, configured to construct an adjacency matrix representing the implicit connection pattern between different electrodes and a spatial relationship matrix representing the spatial relationship between electrodes, fuse the adjacency matrix and the spatial relationship matrix to obtain a final adjacency matrix, and use L1 regularization to constrain the final adjacency matrix;
[0120] A diffusion convolution module, configured to introduce diffusion convolution and a bidirectional diffusion mechanism based on the final adjacency matrix and the fused feature, simulate the random walk process on the graph, capture the multi-hop interaction relationships between nodes, and combine a fully connected layer and an activation layer to finally realize the emotion classification of multi-channel EEG signals.
[0121] Specifically, the node feature extraction module is further configured to:
[0122] S11. The frequency-domain encoding module obtains the weighted features of each frequency band after weighting the input features with a weight and an activation function, and fuses the weighted features of each frequency band with the input features to obtain a frequency-domain feature representation;
[0123] S12. The time-domain encoding module uses a GRU network to obtain the hidden layer state at the current moment, and constructs a time-domain feature representation using the hidden layer state;
[0124] S13. Weightedly fuse the frequency-domain feature representation and the time-domain feature representation using a gating coefficient to obtain a fused feature.
[0125] More specifically, S11 includes:
[0126] The frequency-domain encoding module multiplies the feature output by the ReLU activation function after weighting the input feature and then weighting it again by the Softmax activation function element-wise with the input feature to obtain the weighted feature for each frequency band. The weighted feature for each frequency band is input into another ReLU activation function after weighting, and the obtained feature is added to the weighted input feature to obtain the frequency-domain feature representation.
[0127] More specifically, S12 includes:
[0128] The time-domain encoding module inputs the input feature at the current time and the hidden state at the previous time into the GRU network to obtain the hidden state at the current time. The GRU network accumulates and outputs the hidden state at each time step to obtain the time-domain feature representation, and this time-domain feature representation is processed by a fully connected layer and a normalization layer to make its dimension consistent with the target dimension of the frequency-domain feature representation.
[0129] More specifically, S13 includes:
[0130] Add the frequency-domain feature representation and the time-domain feature representation and input them into a one-dimensional convolutional layer, then obtain a gating coefficient through a Sigmoid activation function. Add the product of the gating coefficient and the frequency-domain feature representation to the product of the time-domain feature representation and the gating parameter to obtain the fused feature, where the gating parameter is 1 minus the gating coefficient.
[0131] Specifically, the edge feature extraction module is also used for:
[0132] Add the product of the adjacency matrix and the first weighting coefficient to the product of the spatial relationship matrix and the second weighting coefficient to obtain the final adjacency matrix, where the second weighting coefficient is 1 minus the first weighting coefficient.
[0133] Specifically, the element in the i-th row and j-th column of the spatial relationship matrix d i,j is the distance between electrode i and electrode j calculated according to the standard 10-20 EEG electrode placement method, and δ is the calibration constant of the distance.
[0134] Specifically, the diffusion convolution module is also used for:
[0135] Introduce bidirectional diffusion convolution, that is, simultaneously simulate forward and backward random walks to obtain the random walk result:
[0136]
[0137] where, D ODenotes the out-degree matrix, representing the sum of the out-edge weights of electrode i of the EEG signal, D O is a diagonal matrix, and its diagonal elements A ij denotes the element at the i-th row and j-th column of the final adjacency matrix, D I Denotes the in-degree matrix, representing the sum of the in-edge weights of electrode i of the EEG signal, D I is a diagonal matrix, and its diagonal elements satisfy K is the order of the diffusion convolution, k is the k-th order diffusion convolution, θ k,O is the learning parameter of the out-degree matrix, θ k,I is the learning parameter of the in-degree matrix, A represents the final adjacency matrix, and X represents the fused features;
[0138] Random walk result Through the alternately stacked fully connected layers and activation layers, the final sentiment classification result is obtained.
[0139] The above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions recorded in the foregoing embodiments, or perform equivalent replacements for some of the technical features; and these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.
Claims
1. A graph diffusion convolutional electroencephalogram emotion recognition method, characterized in that, Including: S1. The frequency-domain encoding module and the time-domain encoding module respectively obtain the frequency-domain feature representation and the time-domain feature representation, and fuse the two to obtain the fused feature; S2. Construct an adjacency matrix representing the implicit connection pattern between different electrodes and a spatial relationship matrix representing the spatial relationship between electrodes, fuse the adjacency matrix and the spatial relationship matrix to obtain the final adjacency matrix, and use L1 regularization to constrain the final adjacency matrix; S3. Based on the final adjacency matrix and the fused feature, introduce diffusion convolution and a bidirectional diffusion mechanism to simulate the random walk process on the graph, capture the multi-hop interaction relationship between nodes, and combine a fully connected layer and an activation layer to finally realize the emotion classification of multi-channel EEG signals.
2. The method for EEG emotion recognition based on graph diffusion convolution according to claim 1, wherein, S1 includes: S11. The frequency-domain encoding module obtains the weighted features of each frequency band after weighting the input features and passing them through an activation function, and fuses the weighted features of each frequency band with the input features to obtain the frequency-domain feature representation; S12. The time-domain encoding module uses the GRU network to obtain the hidden layer state at the current moment, and constructs the time-domain feature representation using the hidden layer state; S13. The frequency-domain feature representation and the time-domain feature representation are weighted and fused using a gating coefficient to obtain the fused feature.
3. The method for EEG emotion recognition based on graph diffusion convolution according to claim 2, wherein S11 includes: The features output by the frequency-domain encoding module after weighting the input features, passing them through the ReLU activation function, weighting them again, and passing them through the Softmax activation function are multiplied element-wise with the input features to obtain the weighted features of each frequency band. The features obtained by weighting the weighted features of each frequency band and inputting them into another ReLU activation function are added to the weighted input features to obtain the frequency-domain feature representation.
4. A graph diffusion convolutional EEG emotion recognition method according to claim 2, characterized in that, S12 includes: The time-domain encoding module inputs the input feature at the current moment and the hidden layer state at the previous moment into the GRU network to obtain the hidden layer state at the current moment. The GRU network accumulates the output of the hidden layer state at each time step to obtain the time-domain feature representation, and the time-domain feature representation is processed by a fully connected layer and a normalization layer to make its dimension consistent with the target dimension of the frequency-domain feature representation.
5. The graph diffusion convolutional EEG emotion recognition method according to claim 2, characterized in that, S13 includes: The sum of the frequency-domain feature representation and the time-domain feature representation is input into a one-dimensional convolutional layer, and then a gating coefficient is obtained through a Sigmoid activation function. The product of the gating coefficient and the frequency-domain feature representation plus the product of the time-domain feature representation and the gating parameter is used to obtain the fused feature, where the gating parameter is 1 minus the gating coefficient.
6. The method for EEG emotion recognition based on graph diffusion convolution according to claim 1, wherein S2 Including: The product of the adjacency matrix and the first weighting coefficient plus the product of the spatial relationship matrix and the second weighting coefficient is used to obtain the final adjacency matrix, where the second weighting coefficient is 1 minus the first weighting coefficient.
7. A graph diffusion convolutional electroencephalogram emotion recognition method according to claim 1, characterized in that The element in the i-th row and j-th column of the spatial relationship matrix d i,j is the distance between electrode i and electrode j calculated according to the standard 10-20 EEG electrode placement method, and δ is the calibration constant of the distance.
8. A graph diffusion convolutional EEG emotion recognition method according to claim 1, characterized in that S3 Including: Introduce bidirectional diffusion convolution, that is, simulate the forward and backward random walks simultaneously to obtain the random walk result: Among them, D O represents the out-degree matrix, which represents the sum of the out-edge weights of electrode i of the EEG signal. D O is a diagonal matrix, and its diagonal elements A ij represents the element in the i-th row and j-th column of the final adjacency matrix. D I represents the in-degree matrix, which represents the sum of the in-edge weights of electrode i of the EEG signal. D I is a diagonal matrix, and its diagonal elements satisfy K is the order of the diffusion convolution, k is the k-th order diffusion convolution, and θ k,O is the learning parameter of the out-degree matrix, and θ k,I is the learning parameter of the in-degree matrix. A represents the final adjacency matrix, and X represents the fused features; Random walk result After alternately stacking fully connected layers and activation layers, the final sentiment classification result is obtained.
9. A graph diffusion convolutional EEG emotion recognition system, characterized in that, Including: A node feature extraction module for designing a frequency-domain encoding module and a time-domain encoding module to respectively obtain a frequency-domain feature representation and a time-domain feature representation, and fusing the two to obtain a fused feature; An edge feature extraction module for constructing an adjacency matrix representing the implicit connection pattern between different electrodes and a spatial relationship matrix representing the spatial relationship between electrodes, fusing the adjacency matrix and the spatial relationship matrix to obtain the final adjacency matrix, and using L1 regularization to constrain the final adjacency matrix; The diffusion convolution module is used to introduce the diffusion convolution and bidirectional diffusion mechanism based on the final adjacency matrix and the fused features, simulate the random walk process on the graph, capture the multi-hop interaction relationships between nodes, and combine the fully connected layer and the activation layer to finally achieve the emotion classification of multi-channel EEG signals.
10. A graph diffusion convolutional EEG emotion recognition system according to claim 9, wherein, The node feature extraction module is also used for: S11. The frequency domain encoding module obtains the weighted features of each frequency band after weighting the input features and passing them through the activation function, and fuses the weighted features of each frequency band with the input features to obtain the frequency domain feature representation; S12. The time domain encoding module uses the GRU network to obtain the hidden layer state at the current moment, and constructs the time domain feature representation using the hidden layer state; S13. The frequency domain feature representation and the time domain feature representation are weighted and fused using the gating coefficient to obtain the fused features.
Citation Information
Patent Citations
Multi-modal feature fusion emotion recognition method based on gating cross-attention mechanism
CN117370828A
Cited By
A method for fault diagnosis of axle box bearings under different operating conditions based on random walk graph kernel
CN122571245A