Electroencephalogram signal fine-grained emotion recognition method and system
By using the common fluctuation edge center analysis method and the CoEdge-STNet model in EEG signal recognition, the problems of fewer emotions recognition types and neglect of dynamic interaction in the prior art are solved, and higher recognition accuracy and efficiency are achieved.
Patent Information
- Application Number
- CN202510320494.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-18
- Publication Date
- 2025-06-20
- Estimated Expiration
- 2045-03-18
AI Technical Summary
The prior art has fewer types of emotional recognition in fine-grained emotional recognition of EEG signals, ignoring the dynamic interaction and information complementarity between brain regions, resulting in low recognition accuracy and efficiency.
The co-fluctuation edge time series and functional connection matrix are constructed by the edge center analysis method based on co-fluctuation, and combined with the CoEdge-STNet model, higher-order spatiotemporal features are extracted for fine-grained emotion recognition.
Through information complementarity and dynamic analysis, the accuracy and generalization of fine-grained emotion recognition are improved, information loss caused by a single feature is overcome, and a more comprehensive and systematic perspective is provided to reveal the brain's emotional perception mechanism.
Smart Images

Figure CN120180306A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to a method and system for fine-grained emotion recognition of electroencephalogram signals. Background Art
[0002] The importance of emotions to human physical and mental health cannot be ignored. With the continuous development of emotion recognition technology, people have begun to pay attention to more fine-grained emotion recognition, such as the recognition of specific emotion types such as joy, anger, fear, sadness, etc. EEG signal technology, as a non-invasive, fast, cost-effective and user-friendly emotion recognition tool, can objectively and accurately map the individual's internal emotional state by measuring the voltage fluctuations caused by the flow of ion currents in brain neurons.
[0003] At present, emotion recognition of EEG signals is limited to a few simple emotions. Research on fine-grained emotion recognition can promote the development of the field of affective computing, improve the accuracy and efficiency of fine-grained emotion recognition, and provide stronger support for the research and application of emotional brain-computer interfaces.
[0004] The existing technology of emotion recognition based on EEG signals is as follows: extracting features in the time domain, frequency domain, etc. on the preprocessed EEG signals, and then using machine learning or deep learning models for classification and recognition. Feature extraction is a key factor in the entire pattern recognition. However, EEG signals are high-dimensional and complex time series signals that change dynamically, and the existing technology does not fully consider the dynamic interactions between brain regions. How to efficiently obtain the high-dimensional nonlinear dynamic features of EEG signals and effectively identify them is a difficult problem that needs to be solved in fine-grained emotion recognition of EEG signals.
[0005] The disadvantages of the prior art include at least:
[0006] (1) Few types of emotion recognition
[0007] In the field of EEG emotion recognition, although there are common datasets such as DEAP, DREAMER, and SEED series, these datasets still need to be improved in terms of the number of subjects, the types of emotions covered, and the diversity of situations, making it difficult to fully evaluate the stability and generalization ability of the model in different practical scenarios. Especially in the application of brain-computer interfaces, the requirements of real-time and generalization further increase the complexity and challenges of research.
[0008] (2) Ignoring high-order interaction information between edges
[0009] Feature extraction is a key factor in the entire pattern recognition. The features used for emotion recognition mostly come from time domain, frequency domain and other features. EEG signals are high-dimensional and complex time series signals that change dynamically. Existing technologies often use node-based functional connections. Such methods may ignore the high-order interaction information between edges and do not fully consider the dynamic interaction between brain regions.
[0010] (3) Ignoring information complementarity
[0011] The existing technologies generally focus on using the extracted features for classification and recognition, while ignoring the time dependence contained in the original time series, which is crucial for revealing the dynamic change information of EEG signals. Summary of the Invention
[0012] In view of this, it is necessary to provide a method and system for fine-grained emotion recognition of EEG signals.
[0013] The present invention provides a method for fine-grained emotion recognition of EEG signals, which includes the following steps: S1, selecting the EEG signals of the subjects from the FACED dataset, and performing frequency division and segmentation on the preprocessed EEG signals of the subjects; S2, for all the preprocessed EEG signals after frequency division and segmentation, using the co-fluctuation-based edge central analysis method to construct a co-fluctuation edge time series ETS and a co-fluctuation edge functional connection matrix EFC; S3, constructing a CoEdge-STNet model according to the high-order spatio-temporal characteristics of the co-fluctuation edge time series ETS and the co-fluctuation edge functional connection matrix EFC, and the CoEdge-STNet model is a fine-grained emotion classification model; S4, using the CoEdge-STNet model to perform fine-grained emotion recognition under different frequency bands within the subjects; S5, using the CoEdge-STNet model to perform fine-grained emotion recognition under different frequency bands across subjects.
[0014] Preferably, the selecting the EEG signals of the subjects from the FACED dataset includes: randomly selecting n subjects, or selecting n fixed consecutive subjects according to the subject numbers, where n is any natural number not exceeding the total number of subjects in the FACED dataset.
[0015] Preferably, the step S1 includes:
[0016] For each selected subject, 28 independent trials are participated in, and the EEG signals of 32 channels with a duration of 30 s are collected; obtaining the EEG signals in the Delta band of 1 - 4 Hz, Theta band of 4 - 8 Hz, Alpha band of 8 - 14 Hz, Beta band of 14 - 30 Hz, and Gamma band of 30 - 47 Hz through band-pass filtering; using a 1 s length as a non-overlapping window to perform equal-length segmentation on the preprocessed EEG signals.
[0017] Preferably, the step S3 includes:
[0018] The CoEdge-STNet model includes: two spatio-temporal collaborative convolution modules ST-Conv Black1, ST-Conv Black2, and a fully connected layer fc. Among them, each spatio-temporal collaborative convolution module extracts spatio-temporal features through the combination of two temporal convolution layers and a graph convolution layer.
[0019] Preferably, step S3 includes:
[0020] Step S31, the first spatio-temporal collaborative convolution module ST-Conv Black1: The input shape is (batch_size, num_nodes, num_features);
[0021] The output of the first spatio-temporal collaborative convolution module ST-Conv Black1 is (batch_size, num_nodes, 64), where 64 represents the feature dimension of each node;
[0022] Step S33, the second spatio-temporal collaborative convolution module ST-Conv Black2: The input shape is (batch_size, num_nodes, 64);
[0023] The output of the second spatio-temporal collaborative convolution module ST-Conv Black2 is (32, 496, 128);
[0024] Step S35, flattening and Dropout: After passing through the two spatio-temporal collaborative convolution modules, a feature map with the shape of (batch_size, num_nodes, 128) is obtained. Flatten (batch_size, num_nodes, 128) into (batch_size, num_nodes * 128), and apply the Dropout layer to randomly discard some neurons to prevent overfitting;
[0025] Step S36, the data is processed through the fully connected layer fc. The flattened features are input into the fully connected layer to obtain the final output.
[0026] Preferably, step S31 includes:
[0027] a. The first temporal convolution layer Convld-1: nn.Conv1d(in_channels, spatial_channels, kernel_size = 3, padding = 1) performs one-dimensional convolution. The kernel size is 3, the stride is defaulted to 1, and the padding is 1 to ensure that the lengths of the input and output are the same;
[0028] Convolution operation: x = self.Convld-1(x), which converts the input data of (batch_size, num_nodes, num_features) into (batch_size, num_nodes, 32) through convolution, and the number of features per node is mapped to 32;
[0029] b. ReLU activation: x = F.relu(x), applying the ReLU activation function to the output after convolution to enhance non-linearity;
[0030] c. Transpose dimensions: x = x.transpose(1, 2), converting the shape from (batch_size, num_nodes, 32) to (batch_size, 32, num_nodes) to prepare for graph convolution;
[0031] d. Graph Convolution: x = self.spatial(x, efc);
[0032] e. ReLU activation: x = F.relu(x), applying the ReLU activation function to the output after graph convolution again;
[0033] f. To send the data to the next temporal convolution layer Convld-2, transpose it again, Transpose dimensions: x = x.transpose(1, 2), and transpose it back to (batch_size, num_nodes, 32) again;
[0034] g. Second temporal convolution layer (Convld-2): The input shape is (batch_size, num_nodes, in_channels), where in_channels is 32, and the output shape is (batch_size, num_nodes, TimeTrack_channels), and TimeTrack_channels is 64.
[0035] Preferably, the step S33 includes:
[0036] a. First temporal convolution layer Convld-1, the input shape is (batch_size, num_nodes, spatial_channels), and this convolution operation performs convolution along the node dimension, and its output shape is (batch_size, num_nodes, spatial_channels);
[0037] b. Apply the ReLU activation function to the result after convolution;
[0038] c. Before applying graph convolution, transpose the data to make it conform to the input format of graph convolution;
[0039] d. The data enters the graph convolution layer GraphConvolution: The graph convolution layer propagates the features of each node and combines them with the features of the co-fluctuation edge functional connection matrix EFC constructed;
[0040] e. Apply the ReLU activation function to the result after graph convolution again;
[0041] f. Transpose the data back so that it can be fed into the next temporal convolution layer Convld-2;
[0042] g. The second temporal convolution layer Convld-2 performs convolution along the node dimension, and the output shape is (batch_size, num_nodes, temporal_channels).
[0043] Preferably, the step S4 includes:
[0044] Within-subject, for all 28 trials representing 9 different emotion categories of each subject, three-fold cross-validation is performed, where each fold includes complete trials representing nine emotion categories to ensure balanced distribution of emotion categories and sample sizes within each fold; The experimental results of each subject are reported as the average result of three-fold cross-validation.
[0045] Preferably, the step S5 includes:
[0046] Between-subject, for all 10 subjects, leave-one-out method is adopted for emotion category recognition. Each time, all trial data of one subject are taken as the validation set, and all trial data of the remaining subjects are taken as the training set. The above steps are repeated 10 times, and finally the average result of 10 times is reported.
[0047] The present invention also provides an electroencephalogram signal fine-grained emotion recognition system, which includes a frequency division and segmentation module, an analysis module, a model construction module, a within-subject fine-grained emotion recognition module, and a between-subject fine-grained emotion recognition module, where:
[0048] The frequency division and segmentation module is used to select the electroencephalogram signals of the subjects from the FACED dataset and perform frequency division and segmentation on the preprocessed electroencephalogram signals of the subjects;
[0049] The analysis module is used to construct the co-fluctuation edge time series ETS and the co-fluctuation edge functional connection matrix EFC for all the electroencephalogram signals after frequency division and segmentation by using the co-fluctuation-based edge central analysis method;
[0050] The model construction module is used to construct the CoEdge-STNet model according to the high-order spatio-temporal characteristics of the co-fluctuating edge time series (ETS) and the co-fluctuating edge functional connectivity matrix (EFC). The CoEdge-STNet model is a fine-grained emotion classification model;
[0051] The intra-subject fine-grained emotion recognition module is used to perform fine-grained emotion recognition using the CoEdge-STNet model at different frequency bands within a subject;
[0052] The cross-subject fine-grained emotion recognition module is used to perform fine-grained emotion recognition using the CoEdge-STNet model at different frequency bands across subjects.
[0053] The present invention can effectively recognize fine-grained emotions. Starting from the edge-centric perspective, by combining the co-fluctuating edge functional connectivity matrix (EFC) and the co-fluctuating edge time series (ETS) to form a richer feature space, it provides a new idea for extracting electroencephalogram emotion recognition features. This model can analyze the spatio-temporal characteristics of brain activities in multiple emotional states, providing a new perspective and method for brain function research, and helping to reveal the working principles and mechanisms of the brain in complex emotional states. The beneficial effects of the present invention specifically include:
[0054] (1) Overcoming information loss caused by a single feature
[0055] The present invention combines the co-fluctuating edge time series (ETS) constructed from the edge-centric perspective with the features of the high-dimensional co-fluctuating edge functional connectivity matrix (EFC) to achieve information complementarity and dynamic analysis, thereby providing a more comprehensive and systematic perspective to reveal the mechanism of the brain's perception of complex emotions.
[0056] (2) Improving the accuracy and generalization of fine-grained emotion recognition
[0057] The present invention can ensure that the model can exhibit good performance in different populations and situations. In particular, it can improve the current situation of low recognition accuracy in the prior art.
[0058] (3) Promoting the development of the field of affective computing
[0059] The present invention enables intelligent machines to understand human emotions, providing stronger support for the research and application of affective brain-computer interfaces. Clinically, it can help treat some severe mental diseases, such as treatment-resistant depression. In daily life, it can help ordinary people relieve anxiety and improve their mood. In the game and entertainment industry, it can adjust the game difficulty or music recommendations according to the player's emotions, providing a more personalized experience. Description of the Drawings
[0060] Figure 1It is a flowchart of the method for fine-grained emotion recognition of electroencephalogram signals in the present invention;
[0061] Figure 2 It is a schematic diagram of the technical route provided by an embodiment of the present invention;
[0062] Figure 3 It is a within-subject schematic diagram of the classification report of the FACED dataset - 10 subjects - 5 frequency bands - fine-grained emotion nine-classification provided by an embodiment of the present invention;
[0063] Figure 4 It is a cross-subject schematic diagram of the classification report of the FACED dataset - 10 subjects - 5 frequency bands - fine-grained emotion nine-classification provided by an embodiment of the present invention;
[0064] Figure 5 It is a hardware architecture diagram of the fine-grained emotion recognition system for electroencephalogram signals in the present invention. Detailed implementation manners
[0065] The present invention will be further described in detail below with reference to the accompanying drawings and specific embodiments.
[0066] Refer to Figure 1 As shown, it is an operation flowchart of a preferred embodiment of the method for fine-grained emotion recognition of electroencephalogram signals in the present invention. Please also refer to Figure 2 :
[0067] Step S1, select the electroencephalogram signals of the subjects from the FACED (Finer-grained Affective Computing EEG Dataset) dataset, and perform frequency division and segmentation on the preprocessed electroencephalogram signals of the subjects. Specifically:
[0068] The selection of the electroencephalogram signals of the subjects from the FACED dataset: n subjects can be randomly selected, or n consecutive fixed subjects can be selected according to the subject numbers. n is any natural number not exceeding the total number of subjects in the FACED dataset. For example, n can be 10, 20, or 30.
[0069] In this embodiment, the first 10 subjects are selected from the FACED dataset. For each selected subject, they all participated in 28 independent trials and the electroencephalogram signals of 32 channels with a duration of 30 s were collected; the electroencephalogram signals in the Delta, Theta, Alpha, Beta, and Gamma frequency bands were obtained through band-pass filtering, which are 1 - 4 Hz, 4 - 8 Hz, 8 - 14 Hz, 14 - 30 Hz, and 30 - 47 Hz respectively; with 1 s as the length as a non-overlapping window, the preprocessed electroencephalogram signals were segmented into equal lengths. Among them, 28 trials represent a total of 9 different emotion categories, and the labels of the emotion categories are digital labels from 0 to 8. The digital labels from 0 to 8 represent the nine emotion categories to be recognized.
[0070] Step S2: For all the EEG signals of the subjects after frequency division and segmentation, use the edge center analysis method based on co-fluctuation to construct the co-fluctuated edge time series (ETS) and the co-fluctuated edge functional connectivity matrix (EFC). Specifically:
[0071] For all the EEG signals of the subjects after frequency division and segmentation, use the edge center analysis method to construct the co-fluctuated edge time series ETS and the co-fluctuated edge functional connectivity matrix EFC, obtaining an edge time series matrix with a dimension of 496*250*30 for each trial sample (where: 496 represents the number of edges, 250 represents the number of sampling points, and 30 represents the number of segments), and an edge functional connectivity matrix with a dimension of 496*496*30 for each trial sample (where: 496 represents the number of edges, and 30 represents the number of segments).
[0072] Step S3: According to the high-order spatio-temporal characteristics of the co-fluctuated edge time series (ETS) and the co-fluctuated edge functional connectivity matrix (EFC), construct the CoEdge-STNet (Co-fluctuated Edge Spatio-Temporal Network) model. The CoEdge-STNet model is a fine-grained emotion classification model. Specifically:
[0073] The complete CoEdge-STNet model includes: 2 spatio-temporal collaborative convolution modules (ST-Conv Black) and a fully connected layer (fc). Among them, each spatio-temporal collaborative convolution module extracts spatio-temporal features through the combination of two temporal convolution layers (Convld-1 and Convld-2) and a graph convolution layer (Graph Convolution).
[0074] The structure of the CoEdge-STNet model is:
[0075] · The first spatio-temporal collaborative convolution module (ST-Conv Black1): Extract spatio-temporal features through the temporal convolution layer, graph convolution layer, and activation function:
[0076] · Convld-1
[0077] · Graph Convolution
[0078] · Convld-2
[0079] · The second spatio-temporal collaborative convolution module (ST-Conv Black2): Further extract more complex spatio-temporal features. The specific structure is the same as that of ST-Conv Black1.
[0080] · Flatten and Dropout: Flatten the features and apply Dropout to prevent overfitting.
[0081] · Fully connected layer (fc): Map the extracted spatio-temporal features to the output class space.
[0082] The above two ST-Conv Blacks are processed sequentially in the network, and the output of ST-Conv Black1 is used as the input of ST-Conv Black2.
[0083] Use temporal convolutional layers to extract edge time series features; use graph convolutional layers to handle the complexity of edge time series while effectively capturing the spatial dependencies contained in the functional connectivity matrix, thereby further enriching the feature representation.
[0084] The CoEdge-STNet model introduces the ReLU activation function after multiple convolutional layers to enhance the expressive power and deep learning ability. The above non-linear activation mechanism can significantly enhance the model's recognition ability for complex data patterns, enabling the model to learn higher-level features. The features output by the temporal convolutional layer are flattened and then passed through the fully connected layer to generate the final prediction result. Dropout technology is applied before the fully connected layer. As a regularization method, Dropout can randomly discard the outputs of some neurons, thereby effectively reducing the overfitting phenomenon and improving the generalization ability of the model.
[0085] Normalize the co-fluctuating edge time series ETS before feeding it into the model. The following are the parameter settings of the CoEdge-STNet model:
[0086] Parameter settings:
[0087] batch_size = 32, lr = 0.0001, weight_decay = 0.01, num_classes = 9
[0088] self.ST-Conv Black1 = ST-Conv Black(num_features, 32, 64, num_nodes)
[0089] self.ST-Conv Black2 = ST-Conv Black(64, 32, 128, num_nodes)
[0090] Model initialization settings:
[0091] model = NeuralNetwork(num_nodes, num_features, num_classes, dropout_rate = 0.5)
[0092] For the FACED dataset used, the shape of the input data is (32, 496, 250), where: 32 is the batch size, 496 is the number of co-fluctuating edge time series (num_nodes) constructed using the edge center view, and 250 is the number of features (num_features) of each co-fluctuating edge time series. The size of the edge functional connectivity network constructed from the co-fluctuating edge time series is 496 * 496.
[0093] Step S31, the first spatio-temporal co-convolution module ST-Conv Black1: The input shape is (batch_size, num_nodes, num_features), which is (32, 496, 250) in this embodiment, and includes the following processing steps:
[0094] a. The first temporal convolution layer (Convld-1): nn.Conv1d(in_channels, spatial_channels, kernel_size = 3, padding = 1) performs one-dimensional convolution. The kernel size is 3, the stride is defaulted to 1, and the padding is 1 to ensure that the lengths of the input and output are the same.
[0095] Convolution operation: x = self.Convld-1(x), which converts the input data of (batch_size, num_nodes, num_features) into (batch_size, num_nodes, 32) through convolution. The number of features of each node is mapped to 32.
[0096] b. ReLU activation: x = F.relu(x), which applies the ReLU activation function to the output after convolution to enhance non-linearity.
[0097] c. Transpose dimensions: x = x.transpose(1, 2), which converts the shape from (batch_size, num_nodes, 32) to (batch_size, 32, num_nodes) to prepare for graph convolution.
[0098] d. Graph convolution (GraphConvolution): x = self.spatial(x, efc).
[0099] The first step of the graph convolution operation: support = torch.matmul(x, self.weight), which is a matrix multiplication step that transforms the input features through the weight matrix weight to obtain the linearly transformed features support.
[0100] The second step of graph convolution operation: The graph convolution propagates features through the adjacency matrix: output = torch.matmul(efc, support), that is, performs matrix multiplication on the features support after linear transformation by the co-fluctuating edge functional connection matrix (EFC).
[0101] The third step of graph convolution operation: Return output + self.bias, that is, add the bias term.
[0102] e. ReLU activation: x = F.relu(x), and then apply the ReLU activation function to the output after graph convolution.
[0103] f. To send the data to the next temporal convolution layer t2, transpose it again, with the transposed dimensions: x = x.transpose(1, 2), and transpose it back to (batch_size, num_nodes, 32).
[0104] g. The second temporal convolution layer (Convld-2):
[0105] self.Convld-2 = nn.Conv1d(spatial_channels, TimeTrack_channels, kernel_size = 3, padding = 1), the second temporal convolution layer (Convld-2): The input shape is (batch_size, num_nodes, in_channels), where: in_channels is 32, and the output shape is (batch_size, num_nodes, TimeTrack_channels), and TimeTrack_channels is 64.
[0106] Step S32, the output of the first spatio-temporal collaborative convolution module ST-Conv Black1 is (batch_size, num_nodes, 64), where the feature dimension of each node is 64.
[0107] Step S33, the second spatio-temporal collaborative convolution module (ST-Conv Black2):
[0108] ST-Conv Black2 is the second spatio-temporal collaborative convolution module, with an input shape of (batch_size, num_nodes, 64), and includes the following processing steps:
[0109] (Convld-2): nn.Conv1d(spatial_channels, Convld-_channels, kernel_size = 3, padding = 1), spatial_channels = 32, TimeTrack_channels = 128.
[0110] a. First temporal convolutional layer (Convld-1): Convld-1 is a 1D convolutional layer. The input shape is (batch_size, num_nodes, spatial_channels), i.e., (32, 496, 32). This convolutional operation performs convolution along the node dimension (i.e., convolves the features of each node), and its output shape is (batch_size, num_nodes, spatial_channels), i.e., (32, 496, 64).
[0111] b. Apply the ReLU activation function to the result of the convolution: The shape of x remains (32, 496, 64) because the ReLU activation function does not change the shape of the tensor.
[0112] c. Before applying the graph convolution, transpose the data to conform to the input format of the graph convolution: The shape of x is transposed from (32, 496, 64) to (32, 64, 496).
[0113] d. The data enters the graph convolutional layer (GraphConvolution):
[0114] The graph convolutional layer propagates the features of each node, combines them with the features of the constructed co-fluctuating edge functional connection matrix EFC, and the output shape remains (32, 64, 496).
[0115] e. Apply the ReLU activation function to the result of the graph convolution again: The shape of x remains (32, 64, 496).
[0116] f. Transpose the data back to feed it into the next temporal convolutional layer: The shape of x is transposed from (32, 64, 496) to (32, 496, 64), restoring to (batch_size, num_nodes, spatial_channels).
[0117] g. Second temporal convolutional layer (Convld-2):
[0118] Convld-2 is a 1D convolutional layer. Convld-2 performs convolution along the node dimension, and the output shape is (batch_size, num_nodes, temporal_channels), that is, (32, 496, 128).
[0119] Step S34, the output shape of the second spatio-temporal collaborative convolution module ST-Conv Black2 is (32, 496, 128).
[0120] Step S35, Flatten and Dropout:
[0121] After passing through two spatio-temporal collaborative convolution modules, a feature map with the shape of (batch_size, num_nodes, 128) is obtained.
[0122] Next, Flatten: x = x.view(x.size(0), -1), which flattens (batch_size, num_nodes, 128) into (batch_size, num_nodes * 128), that is, combines the num_nodes * 128 dimensions for passing to the fully connected layer. For the current embodiment, the flattened shape is (batch_size, 496 * 128).
[0123] Dropout: x = self.dropout(x), applies the Dropout layer to randomly discard some neurons to prevent overfitting.
[0124] Step S36, Fully connected layer (fc):
[0125] The data is processed through the fully connected layer fc. Fully connected: x = self.fc(x), inputs the flattened features into the fully connected layer to obtain the final output. The shape of the fully connected layer is (batch_size, 9), that is, each sample is classified into one of the 9 emotion categories. Here, the weight matrix shape of the fc layer is (num_nodes * 128, num_classes), that is, (496 * 128, 9). Multiply the flattened feature vector by the weight matrix and add the bias to finally obtain a 9-dimensional output vector, with each dimension corresponding to the score of a category.
[0126] Step S4, use the CoEdge-STNet model for fine-grained emotion recognition under different frequency bands within the subject. That is, in this embodiment, within-subject emotion recognition is performed separately under different frequency bands (a total of 5), and it needs to be performed under each frequency band. Finally, 1 fine-grained emotion recognition result is reported for each of the 5 frequency bands. Please refer to Figure 3 , specifically:
[0127] In this embodiment, within-subject, for each subject, 28 trials representing 9 different emotion categories are subjected to three-fold cross-validation. Each fold includes complete trials representing nine emotion categories to ensure balanced distribution of emotion categories and sample sizes within each fold. The experimental results of each subject are reported as the average result of the three-fold cross-validation.
[0128] Step S5, across subjects and at different frequency bands, use the CoEdge-STNet model for fine-grained emotion recognition. That is, in this embodiment, cross-subject emotion recognition is performed separately at different frequency bands (a total of 5). It is required for each frequency band. Finally, 1 fine-grained emotion recognition result is reported for each of the 5 frequency bands. Please refer to Figure 4 Specifically:
[0129] In this embodiment, across subjects, for all 10 subjects, the leave-one-out method is adopted for emotion category recognition. Each time, all the trial data of one subject are used as the validation set, and all the trial data of the remaining subjects are used as the training set. The above steps are repeated 10 times, and finally the average result of the 10 times is reported.
[0130] Refer to Figure 5 As shown, it is the hardware architecture diagram of the electroencephalogram signal fine-grained emotion recognition system 10 of the present invention. Please also refer to Figure 2 This system includes: a frequency division and segmentation module 101, an analysis module 102, a model construction module 103, an within-subject fine-grained emotion recognition module 104, and a cross-subject fine-grained emotion recognition module 105. Among them:
[0131] The frequency division and segmentation module 101 is used to select the electroencephalogram signals of the subjects from the FACED dataset and perform frequency division and segmentation on the preprocessed electroencephalogram signals of the subjects. Specifically:
[0132] The selection of the electroencephalogram signals of the subjects from the FACED dataset: n subjects can be randomly selected, or n consecutive fixed subjects can be selected according to the subject numbers. n is any natural number not exceeding the total number of subjects in the FACED dataset. For example, n can be 10, 20, or 30.
[0133] In this embodiment, the first 10 subjects are selected from the FACED dataset. For each selected subject, 28 independent trials are participated in, and electroencephalogram (EEG) signals of 32 channels with a duration of 30 s are collected. The EEG signals in the Delta, Theta, Alpha, Beta, and Gamma frequency bands are obtained through band-pass filtering, which are 1 - 4 Hz, 4 - 8 Hz, 8 - 14 Hz, 14 - 30 Hz, and 30 - 47 Hz respectively. Taking 1 s as the length as a non-overlapping window, the preprocessed EEG signals are segmented into equal lengths. Among them, the 28 trials represent 9 different emotion categories, and the labels of the emotion categories are digital labels from 0 to 8, and the digital labels 0 - 8 represent the nine emotion categories to be recognized.
[0134] The analysis module 102 is used to construct a co-fluctuation edge time series (ETS) and a co-fluctuation edge functional connectivity matrix (EFC) for all the subjects' EEG signals after frequency division and segmentation using the edge central analysis method based on co-fluctuation. Specifically:
[0135] The analysis module 102 uses the edge central analysis method for all the subjects' EEG signals after frequency division and segmentation to construct a co-fluctuation edge time series ETS and a co-fluctuation edge functional connectivity matrix EFC, obtaining an edge time series matrix with a sample dimension of 496 * 250 * 30 for each trial (where: 496 represents the number of edges, 250 represents the number of sampling points, and 30 represents the number of segments), and an edge functional connectivity matrix with a sample dimension of 496 * 496 * 30 for each trial (where: 496 represents the number of edges, and 30 represents the number of segments).
[0136] The model construction module 103 is used to construct a CoEdge-STNet model according to the high-order spatio-temporal characteristics of the co-fluctuation edge time series (ETS) and the co-fluctuation edge functional connectivity matrix (EFC). The CoEdge-STNet model is a fine-grained emotion classification model. Specifically:
[0137] The complete CoEdge-STNet model includes: 2 spatio-temporal collaborative convolution modules (ST-Conv Black) and a fully connected layer (fc). Among them, each spatio-temporal collaborative convolution module extracts spatio-temporal features through the combination of two temporal convolution layers (Convld-1 and Convld-2) and a graph convolution layer (Graph Convolution).
[0138] The CoEdge-STNet model structure is:
[0139] · The first spatio-temporal collaborative convolution module (ST-Conv Black1): Extract spatio-temporal features through a temporal convolution layer, a graph convolution layer, and an activation function:
[0140] · Convld-1
[0141] · Graph Convolution
[0142] · Convld-2
[0143] · Second Spatiotemporal Collaborative Convolution Module (ST-Conv Black2): Further extract more complex spatiotemporal features. The specific structure is the same as ST-Conv Black1
[0144] · Flatten and Dropout: Flatten the features and apply Dropout to prevent overfitting.
[0145] · Fully connected layer (fc): Map the extracted spatiotemporal features to the output class space.
[0146] The above two ST-Conv Blacks are processed sequentially in the network, and the output of ST-Conv Black1 is used as the input of ST-Conv Black2.
[0147] Use the temporal convolutional layer to extract edge time series features; use the graph convolutional layer to handle the complexity of the edge time series while effectively capturing the spatial dependencies contained in the functional connectivity matrix, thereby further enriching the feature representation.
[0148] The CoEdge-STNet model introduces the ReLU activation function after multiple convolutional layers to enhance the expression ability and deep learning ability. The above non-linear activation mechanism can significantly enhance the model's ability to recognize complex data patterns, enabling the model to learn higher-level features. The features output by the temporal convolutional layer are flattened and then passed through the fully connected layer to generate the final prediction result. The Dropout technique is applied before the fully connected layer. As a regularization method, Dropout can randomly discard the outputs of some neurons, thereby effectively reducing the overfitting phenomenon and improving the generalization ability of the model.
[0149] Normalize the co-fluctuating edge time series ETS before feeding it into the model. The following are the parameter settings of the CoEdge-STNet model:
[0150] Parameter settings:
[0151] batch_size = 32, lr = 0.0001, weight_decay = 0.01, num_classes = 9
[0152] self.ST-Conv Black1 = ST-Conv Black(num_features, 32, 64, num_nodes)
[0153] self.ST-Conv Black2 = ST-Conv Black(64, 32, 128, num_nodes)
[0154] Model initialization settings:
[0155] model = NeuralNetwork(num_nodes, num_features, num_classes, dropout_rate = 0.5)
[0156] For the FACED dataset used, the shape of the input data is (32, 496, 250), where: 32 is the batch size, 496 is the number of co-fluctuating edge time series (num_nodes) constructed using the edge center view, and 250 is the number of features (num_features) for each co-fluctuating edge time series. The size of the edge functional connectivity network constructed by the co-fluctuating edge time series is 496 * 496.
[0157] First, the first spatio-temporal co-convolution module ST-Conv Black1: The input shape is (batch_size, num_nodes, num_features), which is (32, 496, 250) in this embodiment, and includes:
[0158] a. The first temporal convolutional layer (Convld-1): nn.Conv1d(in_channels, spatial_channels, kernel_size = 3, padding = 1) performs one-dimensional convolution. The kernel size is 3, the stride is defaulted to 1, and the padding is 1 to ensure that the lengths of the input and output are the same.
[0159] Convolution operation: x = self.Convld-1(x), which converts the input data of (batch_size, num_nodes, num_features) to (batch_size, num_nodes, 32) through convolution. The number of features for each node is mapped to 32.
[0160] b. ReLU activation: x = F.relu(x), which applies the ReLU activation function to the output after convolution to enhance non-linearity.
[0161] c. Transpose dimensions: x = x.transpose(1, 2), which converts the shape from (batch_size, num_nodes, 32) to (batch_size, 32, num_nodes) to prepare for graph convolution.
[0162] d. Graph Convolution: x = self.spatial(x, efc).
[0163] The first step of the graph convolution operation: support = torch.matmul(x, self.weight), which is a matrix multiplication that transforms the input features through the weight matrix weight to obtain the linearly transformed features support.
[0164] The second step of the graph convolution operation: the graph convolution propagates the features through the adjacency matrix: output = torch.matmul(efc, support), that is, performs matrix multiplication on the features support linearly transformed by the co-fluctuating edge function connection matrix (EFC).
[0165] The third step of the graph convolution operation: return output + self.bias, that is, add the bias term.
[0166] e. ReLU activation: x = F.relu(x), and then apply the ReLU activation function to the output after graph convolution.
[0167] f. To send the data to the next temporal convolutional layer t2, transpose it again, with the transposed dimensions: x = x.transpose(1, 2), and transpose it back to (batch_size, num_nodes, 32) again.
[0168] g. The second temporal convolutional layer (Convld-2):
[0169] self.Convld-2 = nn.Conv1d(spatial_channels, TimeTrack_channels, kernel_size = 3, padding = 1), the second temporal convolutional layer (Convld-2): the input shape is (batch_size, num_nodes, in_channels), where: in_channels is 32, and the output shape is (batch_size, num_nodes, TimeTrack_channels), and TimeTrack_channels is 64.
[0170] Secondly, the output of the first spatio-temporal collaborative convolutional module ST-Conv Black1 is (batch_size, num_nodes, 64), where the feature dimension of each node is 64.
[0171] Then, the second spatio-temporal collaborative convolutional module (ST-Conv Black2):
[0172] ST-Conv Black2 is the second spatio-temporal collaborative convolution module with an input shape of (batch_size, num_nodes, 64), including:
[0173] (Convld-2): nn.Conv1d(spatial_channels, Convld-_channels, kernel_size = 3, padding = 1), where spatial_channels = 32 and TimeTrack_channels = 128.
[0174] a. The first temporal convolutional layer (Convld-1): Convld-1 is a 1D convolutional layer with an input shape of (batch_size, num_nodes, spatial_channels), i.e., (32, 496, 32). This convolutional operation performs convolution along the node dimension (i.e., convolving the features of each node), and its output shape is (batch_size, num_nodes, spatial_channels), i.e., (32, 496, 64).
[0175] b. Apply the ReLU activation function to the result of the convolution: The shape of x remains (32, 496, 64) because the ReLU activation function does not change the shape of the tensor.
[0176] c. Before applying the graph convolution, transpose the data to conform to the input format of the graph convolution: The shape of x is transposed from (32, 496, 64) to (32, 64, 496).
[0177] d. The data enters the graph convolutional layer (GraphConvolution):
[0178] The graph convolutional layer propagates the features of each node, combines them with the features of the constructed co-varying edge functional connection matrix EFC, and the output shape remains (32, 64, 496).
[0179] e. Apply the ReLU activation function to the result of the graph convolution again: The shape of x remains (32, 64, 496).
[0180] f. Transpose the data back to feed it into the next temporal convolutional layer. That is, the shape of x is transposed from (32, 64, 496) to (32, 496, 64), restoring to (batch_size, num_nodes, spatial_channels).
[0181] g. The second temporal convolutional layer (Convld-2):
[0182] Convld-2 is a 1D convolutional layer. Convld-2 performs convolution along the node dimension, and the output shape is (batch_size, num_nodes, temporal_channels), that is, (32, 496, 128).
[0183] Subsequently, the output shape of the second spatio-temporal collaborative convolution module ST-Conv Black2 is (32, 496, 128).
[0184] Then, flattening and Dropout:
[0185] After passing through two spatio-temporal collaborative convolution modules, a feature map with the shape of (batch_size, num_nodes, 128) is obtained.
[0186] Flattening: x = x.view(x.size(0), -1), flatten (batch_size, num_nodes, 128) into (batch_size, num_nodes * 128), that is, merge the num_nodes * 128 dimensions for passing to the fully connected layer. For the current embodiment, the flattened shape is (batch_size, 496 * 128).
[0187] Dropout: x = self.dropout(x), apply the Dropout layer to randomly discard some neurons to prevent overfitting.
[0188] Finally, the fully connected layer (fc):
[0189] The data is processed through the fully connected layer fc. Fully connected: x = self.fc(x), input the flattened features into the fully connected layer to obtain the final output. The shape of the fully connected layer is (batch_size, 9), that is, each sample is classified into one of 9 emotion categories. Here, the weight matrix shape of the fc layer is (num_nodes * 128, num_classes), that is, (496 * 128, 9). Multiply the flattened feature vector by the weight matrix and add the bias to finally obtain a 9-dimensional output vector, with each dimension corresponding to the score of a category.
[0190] The within-subject fine-grained emotion recognition module 104 is used to perform fine-grained emotion recognition using the CoEdge-STNet model at different frequency bands within the subject. That is, in this embodiment, within-subject emotion recognition is performed separately at different frequency bands (a total of 5), and it is required for each frequency band. Finally, 1 fine-grained emotion recognition result is reported for each of the 5 frequency bands. Please refer to Figure 3, specifically:
[0191] In this embodiment, within-subject, for each subject, 28 trials representing 9 different emotion categories are subjected to three-fold cross-validation. Each fold includes complete trials representing nine emotion categories to ensure balanced distribution of emotion categories and sample sizes within each fold. The experimental results of each subject are reported as the average result of three-fold cross-validation.
[0192] The cross-subject fine-grained emotion recognition module 105 is used to perform fine-grained emotion recognition using the CoEdge-STNet model under different frequency bands across subjects. That is, in this embodiment, cross-subject emotion recognition is performed separately under different frequency bands (a total of 5), and it is required for each frequency band. Finally, 1 fine-grained emotion recognition result is reported for each of the 5 frequency bands. Please refer to Figure 4 , specifically:
[0193] In this embodiment, across subjects, for all 10 subjects, the leave-one-out method is adopted for emotion category recognition. Each time, all the trial data of one subject are used as the validation set, and all the trial data of the remaining subjects are used as the training set. The above steps are repeated 10 times, and finally the average result of 10 times is reported.
[0194] The present invention combines the co-fluctuating edge time series (ETS) constructed from the edge-center perspective with the high-dimensional co-fluctuating edge functional connectivity matrix (EFC) features, achieving information complementarity and dynamic analysis, enabling the CoEdge-STNet model to capture complex emotional states that are difficult to discover only through time-domain or frequency-domain features, thereby providing a more comprehensive and systematic perspective to reveal the internal connections of emotional states.
[0195] The co-fluctuating edge functional connectivity matrix (EFC) and the co-fluctuating edge time series (ETS) have different feature representation forms. Combining the two can form a richer feature space. The co-fluctuating edge time series (ETS) can reflect the dynamic change process of EEG signals, helping to understand the dynamic process of brain activities in emotional states; combining with the co-fluctuating edge functional connectivity matrix (EFC) can more deeply analyze the spatio-temporal features of brain activities in emotional states.
[0196] The present invention can promote the development of the field of affective computing, enabling intelligent machines to understand human emotions, providing more powerful support for the research and application of affective brain-computer interfaces. Clinically, it can help treat some severe mental diseases, such as treatment-resistant depression. In daily life, it can help ordinary people relieve anxiety and improve their emotions. In addition, affective brain-computer interfaces also have great commercial value, bringing innovation to the game and entertainment industries, being able to adjust game difficulty or music recommendations according to players' emotions, providing a more personalized experience.
[0197] Although the present invention has been described with reference to the current preferred embodiments, those skilled in the art should understand that the above preferred embodiments are only used to illustrate the present invention and are not used to limit the protection scope of the present invention. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle scope of the present invention should be included within the scope of the present invention's rights protection.
Claims
1. A fine-grained emotion recognition method for EEG signals, characterized in that: The method comprises the following steps: S1, select the EEG signals of the subjects from the FACED data set, and perform frequency segmentation on the preprocessed EEG signals of the subjects; S2, for all the EEG signals of the subjects after frequency segmentation, the co-fluctuation edge time series ETS and co-fluctuation edge functional connection matrix EFC were constructed using the co-fluctuation edge center analysis method; S3, constructing a CoEdge-STNet model according to the high-order spatiotemporal characteristics of the co-fluctuation edge time series ETS and the co-fluctuation edge functional connection matrix EFC, wherein the CoEdge-STNet model is a fine-grained emotion classification model; S4, fine-grained emotion recognition using the CoEdge-STNet model at different frequency bands within the subject; S5, fine-grained emotion recognition using the CoEdge-STNet model in different frequency bands across subjects.
2. The method according to claim 1, characterized in that The selecting of the subject's EEG signals from the FACED data set includes: randomly selecting n subjects, or selecting fixed and continuous n subjects according to the subject numbers, where n is any natural number not exceeding the total number of subjects in the FACED data set.
3. The method according to claim 2, characterized in that The step S1 comprises: Each selected subject participated in 28 independent trials and collected 32 channels of EEG signals with a duration of 30 seconds. The EEG signals in the Delta frequency bands of 1-4 Hz, Theta frequency bands of 4-8 Hz, Alpha frequency bands of 8-14 Hz, Beta frequency bands of 14-30 Hz, and Gamma frequency bands of 30-47 Hz were obtained through band-pass filtering. The preprocessed EEG signals were segmented into equal lengths with a non-overlapping window length of 1 second.
4. The method according to claim 3, characterized in that The step S3 comprises: The CoEdge-STNet model includes: two spatiotemporal co-convolution modules ST-Conv Black1 and ST-Conv Black2 and a fully connected layer fc, wherein each spatiotemporal co-convolution module extracts spatiotemporal features through a combination of two temporal convolution layers and one graph convolution layer.
5. The method according to claim 4, characterized in that The step S3 comprises: Step S31, the first spatiotemporal collaborative convolution module ST-Conv Black1: the input shape is (batch_size, num_nodes, num_features); Step S32, the output of the first spatiotemporal co-convolution module ST-Conv Black1 is (batch_size, num_nodes, 64), where 64 represents the feature dimension of each node; Step S33, the second spatiotemporal collaborative convolution module ST-Conv Black2: the input shape is (batch_size, num_nodes, 64); Step S34, the output of the second spatiotemporal collaborative convolution module ST-Conv Black2 is (32, 496, 128); Step S35, flattening and Dropout: After passing through the two spatiotemporal co-convolution modules, a feature map with a shape of (batch_size, num_nodes, 128) is obtained. (batch_size, num_nodes, 128) is flattened into (batch_size, num_nodes*128), and a Dropout layer is applied to randomly discard some neurons to prevent overfitting; Step S36, the data is processed by the fully connected layer fc, and the flattened features are input into the fully connected layer to obtain the final output.
6. The method according to claim 5, characterized in that The step S31 comprises: a. The first time convolution layer Convld-1: nn.Conv1d(in_channels, spatial_channels, kernel_size = 3, padding = 1) performs a one-dimensional convolution with a kernel size of 3, a default step size of 1, and a padding of 1 to ensure that the input and output lengths are consistent; Convolution operation: x = self.Convld-1(x), the input data of (batch_size, num_nodes, num_features) is converted to (batch_size, num_nodes, 32) through convolution, and the number of features of each node is mapped to 32; b. ReLU activation: x = F.relu(x), apply the ReLU activation function to the convolution output to enhance nonlinearity; c. Transpose dimension: x = x.transpose(1,2), convert the shape from (batch_size, num_nodes, 32) to (batch_size, 32, num_nodes) in preparation for graph convolution; d. GraphConvolution: x = self.spatial(x,efc); e. ReLU activation: x = F.relu(x), then apply the ReLU activation function to the output of the graph convolution; f. In order to feed the data into the next time convolution layer Convld-2, transpose it again, transpose the dimension: x = x.transpose(1,2), and transpose it back to (batch_size, num_nodes, 32); g. Second temporal convolutional layer (Convld-2): The input shape is (batch_size, num_nodes, in_channels), where in_channels is 32, and the output shape is (batch_size, num_nodes, TimeTrack_channels), where TimeTrack_channels is 64.
7. The method according to claim 6, characterized in that The step S33 comprises: a. The first time convolution layer Convld-1, the input shape is (batch_size, num_nodes, spatial_channels), the convolution operation is performed along the node dimension, and its output shape is (batch_size, num_nodes, spatial_channels); b. Apply the ReLU activation function to the convolution result; c. Before applying graph convolution, transpose the data to make it conform to the input format of graph convolution; d. The data enters the graph convolution layer GraphConvolution: The graph convolution layer propagates the features of each node and combines them with the features of the constructed co-fluctuation edge functional connection matrix EFC; e. Apply the ReLU activation function to the result of graph convolution again; f. Transpose the data back so that it can be fed into the next time convolution layer Convld-2; g. The second temporal convolution layer Convld-2 performs convolution along the node dimension, and the output shape is (batch_size, num_nodes, temporal_channels).
8. The method according to claim 7, characterized in that The step S4 comprises: Within the subjects, a three-fold cross-validation was performed on all 28 trials representing 9 different emotion categories for each subject, where each fold included complete trials representing nine emotion categories to ensure a balanced distribution of emotion categories and sample numbers within each fold; the experimental results for each subject were reported as the average results of the three-fold cross-validation.
9. The method according to claim 8, characterized in that The step S5 comprises: Across subjects, for all 10 subjects, the leave-one-out method was used to perform emotion category recognition. Each time, all trial data of one subject was selected as the validation set, and all trial data of the remaining subjects were used as the training set. The above steps were repeated 10 times, and finally the average result of 10 times was reported.
10. A fine-grained emotion recognition system for EEG signals, characterized in that: The system includes a frequency segmentation module, an analysis module, a model building module, an intra-subject fine-grained emotion recognition module, and an inter-subject fine-grained emotion recognition module, wherein: The frequency division and segmentation module is used to select the EEG signals of the subjects from the FACED data set and perform frequency division and segmentation on the preprocessed EEG signals of the subjects; The analysis module is used to construct the co-fluctuation edge time series ETS and the co-fluctuation edge functional connection matrix EFC for all the EEG signals of the subjects after frequency segmentation using the co-fluctuation-based edge center analysis method; The model building module is used to build a CoEdge-STNet model according to the high-order spatiotemporal characteristics of the co-fluctuation edge time series ETS and the co-fluctuation edge functional connection matrix EFC, and the CoEdge-STNet model is a fine-grained emotion classification model; The intra-subject fine-grained emotion recognition module is used to perform fine-grained emotion recognition using the CoEdge-STNet model at different frequency bands within the subject; The cross-subject fine-grained emotion recognition module is used to perform fine-grained emotion recognition using the CoEdge-STNet model in different frequency bands across subjects.
Citation Information
Patent Citations
Multiple fuzzy representation and anomaly detection method for space-time dynamics of gas turbine equipment
CN116702041A
Electroencephalogram signal identification method and device, medium and equipment
CN118436358A
Emotion and character parameters for diffusion model content generation systems and applications
US20240304177A1