An Emotion Recognition Method Based on Spatiotemporal Fusion and Multiple Attention Mechanisms
Through the emotion recognition method of space-time fusion and multiple attention mechanisms, the problems of insufficient temporal information fusion and low cross-subject recognition accuracy in EEG emotion recognition are solved, and efficient emotion recognition and cross-subject generalization ability are achieved.
Patent Information
- Application Number
- CN202510644984.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-20
- Publication Date
- 2025-07-18
- Estimated Expiration
- 2045-05-20
AI Technical Summary
The existing EEG emotion recognition model cannot effectively capture changes between channels and time points when processing non-Euclidean data, ignores spatiotemporal information fusion, insufficient feature extraction, and insufficient recognition accuracy across subjects.
Adopting an emotion recognition method based on space-time fusion and multiple attention mechanisms, by constructing a spatial feature extractor and a temporal feature extraction module, combining a multi-head attention mechanism and a graph convolution network, using a triple-tuple center loss function and a width learning system, adversarial domain adaptation technology is introduced to enhance the cross-domain adaptability of the model.
The accuracy and generalization ability of emotion recognition have been improved, especially in cross-subject scenarios, and the recognition accuracy rate of up to 93.71% has been achieved.
Smart Images

Figure CN120162682B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of emotion recognition, and particularly to an emotion recognition method based on spatio-temporal fusion and multiple attention mechanisms. Background Art
[0002] The importance of machines in various fields has been continuously increasing, promoting the development of multiple aspects such as daily life, industry, education, and entertainment, and affecting the global employment pattern. With the progress of technology, the ability of computers to recognize and respond to human intentions in real time is essential for human-computer interaction (HCI). The goal of human-computer interaction is to enable machines to understand and adapt to human needs, thereby improving the efficiency and quality of interaction. Emotion recognition (ER), as a field that studies how computers perceive and interpret human emotions, has extensive research value and application prospects, especially in fields such as intelligent assistants, healthcare, smart homes, and education. It combines multiple disciplines such as artificial intelligence, cognitive neuroscience, psychology, and computer science.
[0003] In the application of emotion recognition, a computer analyzes information in multiple modalities such as speech, facial expressions, gestures, and physiological signals to judge a person's emotional state. Physiological signals are widely used in the field of emotion recognition because they directly reflect the physical changes of an individual during emotional experience. These changes are usually not affected by an individual's subjective consciousness and can therefore objectively reflect the emotional state. Physiological signals used in emotion recognition include electroencephalogram (EEG), electrocardiogram, etc., and each signal has its unique characteristics.
[0004] Although some progress has been made in emotion recognition tasks based on electroencephalogram (EEG), there are still several challenges to be solved as follows:
[0005] First, some methods rely solely on convolutional neural networks (CNNs) for learning. However, since EEG signals and their derived features are non-Euclidean data, CNN-based methods cannot effectively capture the changes between different channels and time points. This limitation greatly affects the expressive power of the model and makes it unable to accurately reflect the functional connections between channels during emotion arousal;
[0006] Second, existing emotion recognition models often only focus on the extraction of temporal information or spatial information, separately processing dynamic changes or static features, and fail to fully consider the comprehensive fusion and utilization of spatio-temporal information. However, the expression of emotions is a spatio-temporal intertwined process, and relying solely on features in a single dimension often cannot comprehensively capture the complex changes of emotions. Therefore, ignoring the importance of spatio-temporal fusion may lead to limitations in recognition performance;
[0007] Third, although there have been many studies on emotion recognition based on electroencephalogram (EEG), EEG is a complex nonlinear signal. Its non-stationarity easily leads to distribution deviations between different subjects. At the same time, EEG signals usually contain stronger noise than images. How to extract more discriminative and reliable features from EEG signals is still a challenge;
[0008] Fourth, most feature fusion and classification networks lack effective feature enhancement in the final emotion recognition stage, which results in the model failing to fully mine and enhance the important emotion features in the data, thus affecting the recognition accuracy and generalization ability;
[0009] Fifth, due to inter-individual differences, cross-subject EEG emotion recognition is a challenging task. In cross-subject tasks, the training data comes from multiple subjects, while the test data comes from new subjects that the model has not seen. Since there are significant differences in the EEG patterns of different individuals (such as signal amplitude, frequency distribution, noise level, etc.), the features learned by the model on the training data may not be directly transferred to new subjects. Therefore, how to make the model learn to adapt to the unique signal patterns of different individuals so that it can still maintain high emotion recognition accuracy on unseen subject data to achieve stronger generalization ability is still a key issue in current research. Summary of the invention
[0010] In order to solve the above technical problems, the present invention provides an emotion recognition method based on spatiotemporal fusion and multiple attention mechanisms with simple algorithm and high recognition accuracy.
[0011] The technical solution of the present invention to solve the above technical problems is: an emotion recognition method based on spatiotemporal fusion and multiple attention mechanisms, comprising the following steps:
[0012] S1: Construct a spatial feature extractor, which extracts features from EEG signals to obtain feature M1;
[0013] S2: Feature M1 generates spatial feature F2 through the graph convolutional network GAT based on the multi-head attention mechanism;
[0014] S3: construct a time feature extraction module, which extracts features from the EEG signal to obtain a time series feature F3;
[0015] S4: Concatenate the spatial feature F2 and the temporal feature F3, and use the triple center loss function to bring the feature representation of similar samples closer and push the feature representation of heterogeneous samples farther away to obtain the spatiotemporal feature M st ;
[0016] S5: The spatiotemporal features M stAfter being flattened, it is input into the Broad Learning System (BLS) for feature mapping, and the output result is F B ;
[0017] S6: A classifier is constructed using a fully connected layer and a focal loss function to achieve the recognition and classification of emotions; while training the emotion classification model, the gradient reversal layer of the adversarial domain adaptor and the loss function of the domain discriminator are used to directly align the feature distributions of the source domain and the target domain, continuously enhancing the cross-domain adaptation ability of the model.
[0018] In the above emotion recognition method based on spatio-temporal fusion and multiple attention mechanisms, in step S1, the spatial feature extractor includes a Convolutional Block Attention Module (CBAM), a first convolutional layer Conv1, and a second convolutional layer Conv2. The first convolutional layer and the second convolutional layer extract depth features, and CBAM extracts features from channels related to key emotions;
[0019] For the i-th electroencephalogram signal X i , the process of obtaining the fused feature M1 through the spatial feature extractor is as follows:
[0020]
[0021]
[0022]
[0023]
[0024] Among them, is the output of CBAM, is the convolutional attention mapping function, T is the number of electroencephalogram signals, is the output of the first convolutional layer, , both represent convolutional operations, BN represents normalization operation, is the output of the second convolutional layer, is the concatenation operation.
[0025] In the above emotion recognition method based on spatio-temporal fusion and multiple attention mechanisms, the process of step S2 is as follows:
[0026]
[0027] Among them, C represents the number of channels in GAT, k represents the k-th input node, represents the feature M1 input to the k-th node, represents the weight related to each node and its adjacent nodes, that is, the self-attention coefficient of the node; Represents the weight of each node itself, and N represents the number of multi - heads.
[0028] In the above - mentioned emotion recognition method based on spatio - temporal fusion and multi - attention mechanism, in step S3, the time - feature extraction module includes a long short - term memory network LSTM, an attention mechanism AM, and a convolutional layer Conv. In the time - feature extraction module, the i - th electroencephalogram signal X i Is respectively input into LSTM and Conv. The output result of LSTM is input into the attention mechanism AM for weighting, and the weighted result is added to the output of Conv. The added result is then fused and concatenated with X i The whole process is as follows:
[0029]
[0030] Among them, Conv represents the convolution operation, Is the output of the convolutional layer, Represents the self - attention function, Is the output of LSTM, Is the output of the attention mechanism AM, and LSTM( ) is the function for constructing the long short - term memory network.
[0031] In the above - mentioned emotion recognition method based on spatio - temporal fusion and multi - attention mechanism, the process of step S4 is:
[0032]
[0033] Among them, m represents the margin, representing the minimum interval between positive and negative samples; 、 、 Are respectively the concatenated feature vectors of the anchor point, positive sample, and negative sample, Represents the triplet center loss function, Represents the squared Euclidean distance function.
[0034] In the above - mentioned emotion recognition method based on spatio - temporal fusion and multi - attention mechanism, in step S5, the spatio - temporal feature M st After being flattened, it is input into BLS for feature mapping. The BLS architecture consists of two layers. The first layer contains P linear layers integrated with the Sigmoid operation. The multiple features obtained in the first layer are concatenated and input into the second layer for linear transformation. The output features of the first and second layers are concatenated, and the result F is obtained through a linear layer B The whole process is as follows:
[0035]
[0036] Among them, Represents the j-th linear transformation of the first-layer integration of BLS, is the weight matrix corresponding to the j-th linear layer of the first-layer integration of BLS, is the weight matrix corresponding to the linear layer of the second layer of BLS, is the bias term corresponding to the j-th linear layer of the first-layer integration of BLS, is the bias term corresponding to the linear layer of the second layer of BLS, Represents the linear transformation of the second-layer integration of BLS, is the output feature of the j-th linear layer of the first-layer integration of BLS, is the output feature of the second layer of BLS, Flatten( ) represents the flattening operation, is the flattened spatio-temporal feature, and Sigmoid ( ) is the Sigmoid activation function.
[0037] In the above emotion recognition method based on spatio-temporal fusion and multi-attention mechanism, in step S6, F is mapped to emotion category scores through a fully connected layer, and the training process is optimized by combining the focal loss function. The process is as follows: B And the training process is optimized by combining the focal loss function. The process is as follows:
[0038]
[0039] where W is the weight matrix, b is the bias vector, S is the number of samples, is the true class label of the t-th sample, is the probability that the model predicts the t-th sample belongs to of, is 's weight coefficient, which is used to handle the class imbalance problem, and r is the focusing parameter of the focal loss, is the focal loss function.
[0040] In the above emotion recognition method based on spatio-temporal fusion and multi-attention mechanism, in step S6, the leave-one-subject cross-validation method is used for cross-subject emotion recognition evaluation. First, the data of one subject in the DEAP dataset is used as the target domain, and the data of other subjects is used as the source domain. The adversarial domain adaptation operation is performed iteratively until all subjects have been used as the target domain once. A subject refers to an individual participant in the experiment.
[0041] In the above emotion recognition method based on spatio-temporal fusion and multi-attention mechanism, in step S6, the adversarial domain adaptor includes a gradient reversal layer and a domain discriminator. The domain discriminator is a standard binary classifier, 0 represents the source domain, and 1 represents the target domain. The goal of the domain discriminator itself is to distinguish between the source domain and the target domain. By introducing the gradient reversal layer, the domain discriminator gradually becomes unable to distinguish whether the data features come from the source domain or the target domain, thus achieving domain adaptation;
[0042] Input F B into the gradient reversal layer to obtain feature G, and then input the obtained feature G into the domain discriminator. The domain discriminator is a simple fully connected neural network. The input layer inputs feature G, and the hidden layer has 3 fully connected layers, which are used to extract domain-related features. Through linear transformation, the input feature is mapped to the output feature for the domain discrimination task. The process is as follows:
[0043] ;
[0044] ;
[0045] ;
[0046] Among them, represents the gradient reversal layer, represents the output of the th fully connected layer, is the weight matrix, is the bias vector, f( ) represents the linear transformation, represents the feature input into the fully connected layer.
[0047] In the above emotion recognition method based on spatio-temporal fusion and multiple attention mechanisms, in step S6, the goal of the domain discriminator is to minimize the binary cross-entropy loss, and its formula is defined as:
[0048]
[0049] Among them, is the output of the domain discriminator for the cntth sample, that is, the probability of the target domain; is the domain label of the sample, takes the value of 0 to represent the source domain, and takes the value of 1 to represent the target domain, is the loss function of the domain discriminator.
[0050] The beneficial effects of the present invention are as follows:
[0051] 1. The present invention discloses an emotion recognition method based on spatio-temporal fusion and multiple attention mechanisms. First, it fully excavates temporal and spatial information, and combines multiple attention mechanisms to effectively capture spatial topology and temporal features in electroencephalogram data. Subsequently, the spatio-temporal features are fused, and a triplet center loss is introduced as an auxiliary loss for feature learning to make the fused latent feature vectors more discriminative. Then, the extracted features are input into a width learning system to further amplify the output features of the attention neural network, thereby enhancing the expressive ability of the features. Finally, a classifier composed of a fully connected layer and a focal loss function is used for emotion recognition, which can effectively improve the accuracy of emotion analysis and provide accurate data support for fields such as emotion regulation, mental health monitoring, and human-computer interaction.
[0052] 2. The present invention introduces an adversarial domain adaptation technology. Adversarial domain adaptation is a transfer learning technology that makes the features generated by the feature extractor indistinguishable by the domain discriminator through adversarial training, thereby reducing the distribution difference between the source domain and the target domain. By introducing an adversarial domain adapter, the present invention effectively improves the generalization ability of the model in cross-subject scenarios. BRIEF DESCRIPTION OF THE DRAWINGS
[0053] Figure 1 is the overall flowchart of the present invention.
[0054] Figure 2 is the overall model architecture diagram of the present invention.
[0055] Figure 3 is the specific architecture diagram of the convolutional block attention module CBAM of the present invention.
[0056] Figure 4 is the structure diagram of the adversarial domain adapter of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0057] The following further describes the present invention with reference to the drawings and embodiments.
[0058] As Figures 1 - 4 shown, an emotion recognition method based on spatio-temporal fusion and multiple attention mechanisms includes the following steps:
[0059] S1: Construct a spatial feature extractor. The spatial feature extractor extracts a feature M1 from the electroencephalogram signal. This fusion aims to enhance the feature information related to emotions.
[0060] The spatial feature extractor includes a convolutional block attention module CBAM, a first convolutional layer Conv1, and a second convolutional layer Conv2. The first convolutional layer and the second convolutional layer extract deep features, and CBAM extracts features from the channels related to key emotions;
[0061] For the i-th electroencephalogram signal Xi The process of obtaining feature M1 through feature extraction by the spatial feature extractor is as follows:
[0062]
[0063]
[0064]
[0065]
[0066] Among them, is the output of CBAM, is the convolutional attention mapping function, T is the number of electroencephalogram signals, is the output of the first convolutional layer, , both represent convolutional operations, BN represents the normalization operation, is the output of the second convolutional layer, is the concatenation operation.
[0067] With the help of CBAM, the convolutional neural network can better focus on the important topological region feature extraction stage. In addition, compared with more complex attention mechanisms, CBAM introduces relatively fewer parameters. It seamlessly integrates the channel attention (CA) and spatial attention (SA) components. CA dynamically adjusts the importance of different channels, while SA focuses on aggregating the spatial information at different positions in the input features.
[0068] S2: Feature M1 generates spatial feature F2 through the graph convolutional network GAT based on the multi-head attention mechanism.
[0069] The convolutional operation in the graph convolutional network GCN may lead to over-smoothing of graph information, thus reducing the difference between nodes. This may cause GCN to perform poorly when facing electroencephalogram data with fine-grained features. In contrast, GAT proposes an attention mechanism that can adaptively focus on node relationships and channel importance.
[0070] The process of step S2 is as follows:
[0071]
[0072] Among them, C represents the number of channels in GAT, k represents the k-th input node, represents the feature M1 input to the k-th node, represents the weight related to each node and its adjacent nodes, that is, the self-attention coefficient of the node; represents the weight of each node itself, N represents the number of multiple heads, which is configured as 5 in the present invention.
[0073] S3: Construct a time feature extraction module, which extracts features from the EEG signal to obtain a time series feature F3.
[0074] In step S3, the temporal feature extraction module includes a long short-term memory network LSTM, an attention mechanism AM and a convolutional layer Conv. In the temporal feature extraction module, the i-th EEG signal X i The output of LSTM is input into the attention mechanism AM for weighting, and the weighted result is added to the output of Conv, and the added result is added to X. i Perform fusion splicing; the whole process is as follows:
[0075]
[0076] Among them, Conv represents the convolution operation, is the output of the convolutional layer, represents the self-attention function, which can adaptively focus on node relationships and channel importance; is the output of LSTM, is the output of the attention mechanism AM, and LSTM( ) is a function for constructing a long short-term memory network.
[0077] Through this multi-layer feature fusion method that combines LSTM, attention mechanism AM and convolutional layer Conv, the model can better understand the temporal pattern of EEG signals, thereby improving the performance in emotion recognition tasks.
[0078] S4: Concatenate the spatial feature F2 and the temporal feature F3, and use the triple center loss function to bring the feature representation of similar samples closer and push the feature representation of heterogeneous samples farther away to obtain the spatiotemporal feature M st , to enhance the discriminability of features.
[0079] The process of step S4 is as follows:
[0080]
[0081] Among them, m represents the boundary, which represents the minimum interval between positive and negative samples; , , are the concatenated feature vectors of anchor points, positive samples, and negative samples, respectively. represents the triple center loss function, Represents the squared Euclidean distance function.
[0082] The triplet center loss function adjusts the parameters of the feature extraction model through the optimization process. The backpropagation of the triplet center loss function and the update of the model parameters optimize the feature extraction model, making the extracted feature vectors more qualified in the feature space.
[0083] S5: Flatten the spatio-temporal feature M st and input it into the broad learning system BLS for feature mapping, and output the result F B .
[0084] In the step S5, flatten the spatio-temporal feature M st and input it into BLS for feature mapping. The BLS architecture consists of two layers. The first layer contains P linear layers integrated with the Sigmoid operation. The multiple features obtained by the first layer are concatenated and input into the second layer for linear transformation. The output features of the first and second layers are concatenated, and the result F is obtained through a linear layer B ; The whole process is as follows:
[0085]
[0086] Among them, represents the j-th linear transformation integrated by the first layer of BLS, is the weight matrix corresponding to the j-th linear layer integrated by the first layer of BLS, is the weight matrix corresponding to the linear layer of the second layer of BLS, is the bias term corresponding to the j-th linear layer integrated by the first layer of BLS, is the bias term corresponding to the linear layer of the second layer of BLS, represents the linear transformation integrated by the second layer of BLS, is the output feature of the j-th linear layer integrated by the first layer of BLS, is the output feature of the second layer of BLS, Flatten( ) represents the flattening operation, is the flattened spatio-temporal feature, and Sigmoid ( ) is the Sigmoid activation function.
[0087] BLS aims to adopt non-linear operations to enhance the expression ability of input features, mainly by expanding the network width rather than increasing the network depth, thereby improving the expressiveness and training efficiency of the model.
[0088] S6: Construct a classifier using a fully connected layer and a focal loss function to realize the recognition and classification of emotions; while training the emotion classification model, directly align the feature distributions of the source domain and the target domain through the gradient reversal layer of the adversarial domain adaptor and the loss function of the domain discriminator, and continuously enhance the cross-domain adaptation ability of the model.
[0089] In step S6, F is mapped to emotion category scores through a fully connected layer, and the training process is optimized by combining the focal loss function, effectively solving the problem of class imbalance in emotion data, thereby improving the model's recognition ability for minority class emotions. The entire process is as follows: B It is mapped to emotion category scores through a fully connected layer, and the training process is optimized by combining the focal loss function, effectively solving the problem of class imbalance in emotion data, thereby improving the model's recognition ability for minority class emotions. The whole process is as follows:
[0090]
[0091] Among them, W is the weight matrix, b is the bias vector, S is the number of samples, is the true class label of the t-th sample, is the probability that the model predicts the t-th sample belongs to , is 's weight coefficient, used to handle the class imbalance problem, r is the focusing parameter of the focal loss, is the focal loss function.
[0092] Adversarial domain adaptation refers to the ability to transfer the model of the source domain to the target domain by learning the differences between the source domain and the target domain. The leave-one-subject cross-validation method is used to evaluate cross-subject emotion recognition. First, the data of one subject in the DEAP dataset is used as the target domain, and the data of other subjects is used as the source domain. The adversarial domain adaptation operation is performed iteratively until all subjects have been used as the target domain once. A subject refers to an individual participant in the experiment.
[0093] As Figure 4 shown, the adversarial domain adaptor includes a gradient reversal layer and a domain discriminator. The domain discriminator is a standard binary classifier. 0 represents the source domain, and 1 represents the target domain. The goal of the domain discriminator itself is to distinguish between the source domain and the target domain. By introducing the gradient reversal layer, the domain discriminator gradually becomes unable to distinguish whether the data features come from the source domain or the target domain, thereby achieving domain adaptation. Specifically, the gradient reversal layer has no functional effect in the forward propagation, and assigns a negative value to the gradient in the reverse propagation process, so that the parameters of the feature extractor are optimized in the direction that the domain discriminator cannot distinguish. In this way, the feature extractor can learn domain-invariant features, thereby reducing the distribution difference between the source domain and the target domain and achieving the purpose of domain adaptation.
[0094] Input F B into the gradient reversal layer to obtain the feature G, and then input the obtained feature G into the domain discriminator. The domain discriminator is a simple fully connected neural network. The input layer inputs the feature G, and the hidden layer has 3 fully connected layers, which are used to extract domain-related features and map the input features to output features through linear transformation for the domain discrimination task. The process is as follows:
[0095] ;
[0096] ;
[0097] ;
[0098] Among them, represents the gradient reversal layer, represents the output of the -th fully connected layer, is the weight matrix, is the bias vector, and f( ) represents the linear transformation, represents the features input to the fully connected layer.
[0099] The objective of the domain discriminator is to minimize the binary cross-entropy loss, and its formula is defined as:
[0100]
[0101] Among them, is the output of the domain discriminator for the cnt-th sample, that is, the probability of the target domain; is the domain label of the sample, taking the value of 0 represents the source domain, and taking the value of 1 represents the target domain, is the loss function of the domain discriminator.
[0102] The present invention fully extracts the temporal and spatial information in the EEG signals through a multi-attention mechanism, and at the same time fuses and enhances the feature maps to comprehensively learn the feature representation of the emotional state. In addition, by introducing an adversarial domain adaptor, the generalization ability of the model is effectively improved, thereby significantly improving the performance of emotion classification. In the subject-related and subject-independent experiments on the DEAP dataset, the proposed model shows excellent performance compared with other state-of-the-art methods, and the average recognition accuracy based on the subject experiment reaches 93.71%.
Claims
1. A method for emotion recognition based on spatio-temporal fusion and multiple attention mechanisms, characterized in that, It includes the following steps: S1: Construct a spatial feature extractor, which extracts features from the EEG signals to obtain feature M1; S2: Feature M1 generates spatial feature F2 through a graph convolutional network GAT based on the multi-head attention mechanism; S3: Construct a temporal feature extraction module, which extracts features from the EEG signals to obtain temporal feature F3; S4: Concatenate the spatial feature F2 and the temporal feature F3, and use the triple center loss function to bring the feature representation of similar samples closer and push the feature representation of heterogeneous samples farther away to obtain the spatiotemporal feature M st ; S5: Flatten the spatio-temporal feature M st and input it into the Broad Learning System (BLS) for feature mapping, and output the result F B ; S6: Construct a classifier using a fully connected layer and a focal loss function to realize the recognition and classification of emotions; while training the emotion classification model, the feature distributions of the source domain and the target domain are directly aligned through the gradient reversal layer of the adversarial domain adaptor and the loss function of the domain discriminator, continuously enhancing the cross-domain adaptation ability of the model.
2. The emotion recognition method based on spatio-temporal fusion and multiple attention mechanisms according to claim 1, characterized in that, In step S1, the spatial feature extractor includes a convolutional block attention module CBAM, a first convolutional layer Conv1, and a second convolutional layer Conv2. The first convolutional layer and the second convolutional layer extract depth features, and CBAM extracts features from the channels related to key emotions; For the i-th electroencephalogram signal X i , the process of obtaining feature M1 by feature extraction through a spatial feature extractor is as follows: Among them, is the output of CBAM, is the convolutional attention mapping function, T is the number of EEG signals, is the output of the first convolutional layer, and both represent convolutional operations, BN represents normalization operation, is the output of the second convolutional layer, is the concatenation operation.
3. The emotion recognition method based on spatio-temporal fusion and multiple attention mechanisms according to claim 2, wherein The process of step S2 is as follows: Among them, C represents the number of channels in the GAT, k represents the k-th node of the input, represents the feature M1 input to the k-th node, represents the weight associated with each node and its adjacent nodes, that is, the self-attention coefficient of the node; represents the weight of each node itself, and N represents the number of multi-heads.
4. The emotion recognition method based on spatio-temporal fusion and multiple attention mechanisms according to claim 3, wherein In the step S3, the time feature extraction module includes a long short-term memory network (LSTM), an attention mechanism (AM), and a convolutional layer (Conv). In the time feature extraction module, the i-th electroencephalogram signal X i is respectively input into the LSTM and the Conv. The output result of the LSTM is input into the attention mechanism AM for weighting, and the weighted result is added to the output of the Conv. The added result is then fused and concatenated with X i The whole process is as follows: Among them, Conv represents the convolution operation, is the output of the convolutional layer, represents the self-attention function, is the output of the LSTM, is the output of the attention mechanism AM, and LSTM( ) is a function for constructing a long short-term memory network.
5. The emotion recognition method based on spatio-temporal fusion and multi-attention mechanism according to claim 4, characterized in that, The process of step S4 is as follows: Among them, m represents the margin, which represents the minimum interval between positive and negative samples; , , are the concatenated feature vectors of the anchor point, positive sample, and negative sample respectively, represents the triplet center loss function, represents the squared Euclidean distance function.
6. The emotion recognition method based on spatio-temporal fusion and multiple attention mechanisms according to claim 5, characterized in that, In the step S5, the spatio-temporal feature M st is flattened and then input into the BLS for feature mapping. The BLS architecture consists of two layers. The first layer contains P linear layers integrated with the Sigmoid operation. The multiple features obtained by the first layer are concatenated and input into the second layer for linear transformation. The output features of the first and second layers are concatenated, and the result F is obtained through a linear layer B ; the whole process is as follows: Among them, represents the j-th linear transformation of the first-layer integration of BLS, is the weight matrix corresponding to the j-th linear layer of the first-layer integration of BLS, is the weight matrix corresponding to the linear layer of the second layer of BLS, is the bias term corresponding to the j-th linear layer of the first-layer integration of BLS, is the bias term corresponding to the linear layer of the second layer of BLS, represents the linear transformation of the second-layer integration of BLS, is the output feature of the j-th linear layer of the first-layer integration of BLS, is the output feature of the second layer of BLS, Flatten( ) represents the flattening operation, is the flattened spatio-temporal feature, and Sigmoid ( ) is the Sigmoid activation function.
7. The emotion recognition method based on spatio-temporal fusion and multi-attention mechanism according to claim 6, characterized in that, In the step S6, F is mapped to the emotion category score through a fully connected layer, and the training process is optimized by combining the focal loss function, and the process is as follows: B where \(W\) is the weight matrix, \(b\) is the bias vector, \(S\) is the number of samples, is the true class label of the \(t\)-th sample, is the probability that the model predicts the \(t\)-th sample belongs to and, is the weight coefficient of, which is used to handle the class imbalance problem, \(r\) is the focusing parameter of the focal loss, is the focal loss function.
8. The emotion recognition method based on spatio-temporal fusion and multiple attention mechanisms according to claim 7, characterized in that, In step S6, the leave-one-subject cross-validation method is used for cross-subject emotion recognition evaluation. First, the data of one subject in the DEAP dataset is used as the target domain, and the data of other subjects is used as the source domain. The adversarial domain adaptation operation is cyclically executed until all subjects have been used as the target domain once. A subject refers to an individual participant in the experiment.
9. The emotion recognition method based on spatio-temporal fusion and multiple attention mechanisms according to claim 8, wherein In step S6, the adversarial domain adaptor includes a gradient reversal layer and a domain discriminator. The domain discriminator is a standard binary classifier, where 0 is the source domain and 1 is the target domain. The goal of the domain discriminator itself is to distinguish the source domain from the target domain. By introducing the gradient reversal layer, the domain discriminator gradually becomes unable to distinguish whether the data features come from the source domain or the target domain, thus realizing domain adaptation; Input F B Input F into the gradient reversal layer to obtain feature G, and then input the obtained feature G into the domain discriminator. The domain discriminator is a simple fully connected neural network. The input layer inputs feature G, and the hidden layer has three fully connected layers, which are used to extract domain-related features and map the input features to output features through linear transformation for the domain discrimination task. The process is as follows: ; ; ; Among them, represents a gradient reversal layer, represents the output of the -th fully connected layer, is the weight matrix, is the bias vector, and f( ) represents a linear transformation, represents the features input to the fully connected layer.
10. The emotion recognition method based on spatio-temporal fusion and multiple attention mechanisms according to claim 9, wherein In step S6, the goal of the domain discriminator is to minimize the binary cross-entropy loss, and its formula is defined as: Among them, is the output of the domain discriminator for the cnt-th sample, that is, the probability of the target domain; is the domain label of the sample, taking the value of 0 indicates the source domain, and taking the value of 1 indicates the target domain, is the loss function of the domain discriminator.
Citation Information
Patent Citations
Multi-core width learning-based electroencephalogram emotion classification method and device, medium and equipment
CN113011493A
RSVP type unbalanced electroencephalogram signal classification method based on adaptive channel mixed attention mechanism and decoupling learning
CN119385578A