Self-learning multi-modal emotion recognition method based on multi-scale cavity attention
Through the multi-scale hollow attention convolution module and the time-frequency space three-dimensional feature extraction network combined with the self-learning weight module, the recognition accuracy problem in the fusion of EEG signals and facial expression modes is solved, and a more efficient emotion recognition effect is achieved.
Patent Information
- Application Number
- CN202510475209.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-16
- Publication Date
- 2025-07-29
AI Technical Summary
Existing emotion recognition technologies are difficult to effectively integrate information from two modalities: EEG signal and facial expression, especially the problems of different importance of each basic action unit on the face, different distances between key action units, and low recognition accuracy, and different confidence levels of each modal during decision-level fusion.
The multi-scale hollow attention convolution module is used to extract facial expressions, and the EEG signal is processed by combining the time-frequency and space three-dimensional feature extraction network. The dynamic weighting fusion of modal information is performed at the decision level through the self-learning weight module.
It improves the accuracy and robustness of emotion recognition, can more accurately capture the multi-scale features of facial expressions and the multi-dimensional features of EEG signals, and dynamically adjusts the modal weights to improve overall recognition performance.
Smart Images

Figure CN120387093A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of computer vision and pattern recognition, and specifically relates to a self-learning multi-modal emotion recognition method based on multi-scale dilated attention. Background Art
[0002] Emotion not only synthesizes the physiological states of various human sensations, thoughts, and behaviors, but also is the psychological and physiological reaction generated by various external stimuli. Accurately recognizing emotions is of great significance in fields such as the medical field, smart home, autonomous driving, commerce, and nursing companionship.
[0003] Existing emotion recognition research is divided into two categories according to the type of collected signals: emotion recognition based on behavioral performance and emotion recognition based on neurophysiological signals. Signals at the behavioral performance level are easy to collect. For example, facial expressions can intuitively show a person's emotional state, but they are subjective. The person being collected can disguise and hide their true emotional feelings, affecting the effect of emotion recognition. For neurophysiological signals, although the collection conditions of electroencephalogram (EEG) signals are extremely strict and are vulnerable to noise interference, because they are difficult to disguise, they have reliable recognition results. Since emotion is a complex psychophysiological phenomenon, it is difficult to establish a sound emotion recognition model relying solely on a single signal. Research on multi-modal emotion recognition has gradually become a research hotspot and an important development direction in this field. However, the importance of different modalities such as EEG signals and facial expressions changes with the fluctuation of emotions, the expressiveness of each modality signal in the task is different, and there are problems with different confidence levels of each modality in decision-level fusion.
[0004] In addition, human facial expressions are composed of multiple different Facial Action Units (AUs). The importance of these AUs varies. For example, the AUs around the mouth and eyes are usually more important than those around the nose. Moreover, there are dependency relationships at different distances in expressions. For instance, the distance between the eyebrows and eyes is relatively close, while the distance between the eyebrows and mouth is relatively far. Ordinary neural networks cannot accurately capture the key AU features and the correlations between AUs at different distances, and will introduce some useless information. Some studies have used multi-scale networks to mine information at different distances. Zhao et al. proposed a Global Multi-scale Local Attention Network (MA-Net) for FER in the wild environment. By fusing global multi-scale features and local attention features, it effectively solved the problems of occlusion and non-frontal poses. Li et al. proposed a face expression recognition method combining a lightweight Swin Transformer and a multi-scale feature fusion module, reducing the number of model parameters and achieving model lightweighting. Ni Jinyuan et al. used feature maps at different levels in the pyramid convolution structure to obtain information at different scales. In the above methods for obtaining multi-scale information, multi-scale features of continuous regions are extracted. However, there is also some discontinuous information in expressions. If the above methods are used to extract multi-scale features, some useless information will be introduced. Currently, the research on multi-scale information of expressions mainly considers regions of different sizes, but rarely considers the interaction of discontinuous regions. Summary of the Invention
[0005] The present invention provides a self-learning multi-modal emotion recognition method based on multi-scale dilated attention. By using two parallel network branches to perform emotion recognition on facial expressions and electroencephalogram signals respectively, and then fusing the outputs of the two modalities at the decision level to obtain the final classification result, it can solve the problems of low recognition accuracy caused by different importance of each basic action unit of the face and different distances between key action units, as well as the problem of different confidence levels of each modality during decision-level fusion.
[0006] The technical solution adopted by the present invention to solve its technical problems is as follows:
[0007] A self-learning multi-modal emotion recognition method based on multi-scale dilated attention, the steps include:
[0008] Preprocess the obtained facial expression images and then input them into the multi-scale dilated attention convolution module. The multi-scale dilated attention convolution module extracts features at different scales through a parallel three-branch convolution structure. The three-branch convolution structure is set such that the convolution kernel sizes of each branch are the same, and the dilation rates are different from each other. After the features output by the three-branch convolution structure are concatenated in the channel dimension, they are calibrated by the attention mechanism to obtain an enhanced facial expression feature map, which is then sent to the fully connected layer for emotion recognition;
[0009] When the original EEG signal is input into the time-frequency-space three-dimensional feature extraction network, the time-frequency-space three-dimensional feature extraction network decomposes the original EEG signal and then calculates the differential entropy features. The calculation results of the differential entropy features are processed by a global attention module including a spectral attention module, a spatial attention module, and a temporal attention module, and then a time-frequency-space multi-dimensional feature representation is output, and emotion recognition is performed through a fully connected layer;
[0010] The facial expression emotion recognition result and the EEG signal emotion recognition result are input into the self-learning weight module, and the final emotion recognition result is generated through dynamic weighted fusion.
[0011] Further, the specific operations performed within the multi-scale dilated attention convolution module are as follows: The preprocessed facial expression image data passes through two convolutional groups, each convolutional group containing two convolutional layers, to extract the global low-level features in the image data; The global low-level features are input into a three-branch convolutional structure, and the three branches extract features of different scales, which are concatenated in the channel dimension and merged into a comprehensive feature map; The merged comprehensive feature map is calibrated in parallel through a channel attention mechanism and a spatial attention mechanism; Finally, the output feature maps of the channel attention mechanism and the spatial attention mechanism are fused to obtain the final enhanced feature map; The enhanced feature map passes through a convolutional group, a max pooling layer, and a batch normalization layer and then is sent to a fully connected layer for emotion recognition.
[0012] Further, the specific settings of the three-branch convolutional structure are as follows: The first branch uses a 3×3 convolutional kernel with a dilation rate of 1; The second branch uses a 3×3 convolutional kernel with a dilation rate of 2; The third branch uses a 3×3 convolutional kernel with a dilation rate of 3.
[0013] Further, the channel attention mechanism first performs global average pooling on the comprehensive feature map to obtain the global features of each channel, learns the importance weights of each channel through two fully connected layers, and then multiplies these weights with the original comprehensive feature map channel by channel to achieve channel weighting; The spatial attention mechanism performs global pooling on the comprehensive feature map in the channel dimension to obtain a spatial feature map. The spatial feature map passes through a 1×1 convolution and a Sigmoid activation function to learn the importance weights of each spatial position, and multiplies these weights with the original comprehensive feature map element by element to achieve spatial weighting.
[0014] Further, the specific operations performed within the time-frequency-space three-dimensional feature extraction network are as follows: First, the original EEG signal is divided into N segments of equal length, each segment being T seconds long. Then, each segment of the signal is decomposed into four frequency bands using band-pass filtering, and the differential entropy features for each segment are calculated separately for each frequency band over a period of 0.5 seconds. Based on the relative positions of the 32 electrodes, the differential entropy feature vectors are transformed into 2D maps, constructing an 8×9 two-dimensional matrix, where the positions without electrodes are filled with zeros. Then, the 2D differential entropy feature maps of different frequency bands are stacked together to obtain a 4×8×9 three-dimensional feature matrix. The three-dimensional feature matrix is processed by a spectral attention module, a spatial attention module, and a temporal attention module to output a multi-dimensional feature representation of time-frequency-space, and the final emotion recognition result is obtained through a fully connected layer.
[0015] Further, the specific operations performed within the self-learning weight module are as follows: The facial expression emotion recognition result and the EEG signal emotion recognition result are concatenated along the feature dimension to obtain a fused feature representation. The fused feature representation is subjected to a non-linear transformation through a multi-layer fully connected network, and the non-linearly transformed features are reshaped into a sequence form and input into a multi-head self-attention layer for processing. This layer simultaneously performs weighted summation on the feature sequence through multiple attention heads, and each attention head independently calculates the linear transformations of the query, key, and value to dynamically learn the contribution weights of the facial expression modality and the EEG signal modality to emotion classification. After the outputs of all attention heads are concatenated, they undergo a linear transformation to obtain the final weighting coefficients. Finally, according to the obtained weighting coefficients, the information of each modality is adaptively weighted under different emotions, and the features are mapped to the target classification dimension through a series of fully connected layers, and the classification recognition probability is output through an activation function.
[0016] The beneficial effects of the present invention include:
[0017] The present invention designs a multi-scale dilated attention convolution module for facial expression emotion recognition. By applying different dilation rates, the receptive field size is increased without adding extra parameters, realizing the joint extraction of image detail information and global features. The channel attention mechanism and spatial attention mechanism are introduced to enhance the expression of facial emotion features, enabling the learning of discontinuous and cross-regional facial emotion information and improving the model performance. For the emotion recognition of electroencephalogram (EEG) signals, the present invention adopts a time-frequency-space three-dimensional feature extraction network, which can simultaneously retain the time-domain, spatial, and frequency-domain features in EEG signals. Combining with the global attention mechanism, global feature extraction is performed on EEG signals from multiple dimensions to make the most of the emotional information in EEG signals. Aiming at the problems that the importance of different modalities of EEG signals and facial expressions changes with the fluctuation of emotions and the expressive power of each modality signal in the task is different, the present invention designs a self-learning weight module for decision-level fusion, enabling the network to dynamically adjust the modality weights, self-learn the decision distributions of the two modalities, pay attention to more favorable information, prevent low-confidence modality from interfering with the recognition results, and improve the accuracy and robustness of multi-modal emotion recognition. Brief Description of the Drawings
[0018] Figure 1 is the multi-modal emotion recognition model diagram of the present application;
[0019] Figure 2 is the network structure diagram of the multi-scale dilated attention convolution module MSDAC;
[0020] Figure 3 is the EEG electrode position diagram;
[0021] Figure 4 is the network structure diagram of the self-learning weight module;
[0022] Figure 5 is the confusion matrix on the DEAP dataset. Detailed Embodiments
[0023] Next, the technical solutions of the present invention will be clearly and completely described in conjunction with the accompanying drawings. Obviously, the described embodiments are some, but not all, of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.
[0024] In addition, the technical features involved in different embodiments of the present invention described below can be combined with each other as long as they do not conflict with each other.
[0025] Compared with single-modal emotion recognition, multi-modal emotion recognition can explore the consistency and complementary features of emotional data in different modalities. Fully integrating this information enables machines to better understand human emotions. Facial expressions are one of the most natural and intuitive signals for humans to convey emotional states and intentions. Analyzing the information contained therein can provide a good understanding of human consciousness and psychological activities. Electroencephalogram (EEG) is controlled by the central nervous system, directly reflecting the activity state of brain neurons. It is not easily controlled by subjective consciousness and can better reflect the objective and real emotional state, with a higher recognition rate than other physiological signals. Therefore, combining EEG signals controlled by the central nervous system with the most intuitive and effective facial expressions can improve the emotion recognition rate to a certain extent. Based on this design idea, this application proposes a self-learning multi-modal emotion recognition method based on multi-scale dilated attention. By using two parallel networks to perform emotion recognition on two modalities respectively and fusing the recognition results of the two modalities at the decision level, the overall final emotion recognition accuracy is improved.
[0026] First, a multi-scale dilated attention convolutional module MSDAC is designed for facial expressions to learn the information in discontinuous regions of the face and introduce channel-spatial attention to enhance facial features, achieving average accuracies of 74.1%, 99.69%, 98.05% (valence) and 96.15% (arousal) on the Fer2013, CK+ and DEAP datasets respectively. To retain the rich features in multiple dimensions of time domain, space domain and frequency domain in EEG signals, the original EEG signals are converted into three-dimensional time-spectrum-space feature representations for emotion recognition. Finally, the two modalities are fused with self-learning weights at the decision level. The model designed in this invention is experimented on the DEAP dataset, achieving accuracies of 98.66% and 97.49% in the classification of valence and arousal dimensions respectively. Compared with existing methods, the extraction of key facial emotion features is more accurate, and the self-learning weight fusion method more efficiently selects the optimal fusion strategy, further improving the overall performance of the model.
[0027] Example 1: This example proposes a self-learning multi-modal emotion recognition method based on multi-scale dilated attention, as Figure 1As shown, emotion recognition is performed on facial expressions and EEG signals through two parallel network branches, and then the outputs of the two modalities are fused at the decision level to obtain the final classification result. For the problem of low recognition accuracy caused by different importance of each basic action unit in the face and different distances between key action units in the face recognition branch, a multi-scale dilated attention convolutional module (MSDAC) is designed; the EEG signal recognition branch uses an emotion recognition network based on three-dimensional time-frequency-spatial features to capture more comprehensive emotion features; for the problem of different confidence levels of each modality during decision-level fusion, a self-learning weight module is designed to dynamically weight and fuse the results of the two modalities, automatically learning the optimal fusion strategy in the classification of different emotions to improve the final recognition performance.
[0028] Multi-scale dilated attention convolutional module:
[0029] Facial expressions are usually composed of multiple discontinuous action units. A surprised expression may simultaneously include the dilation of the levator palpebrae superioris and the retraction of the orbicularis oris. And there are differences in scale among different AUs. For example, the corrugator supercilii may involve a larger range of muscle movements, while the frontalis muscle involves a smaller range of movements. For the problem that ordinary neural networks cannot accurately capture the key expression action units due to different distances and discontinuities between AUs in facial expressions, the present invention designs a multi-scale dilated attention convolutional module to solve the above problems. This module is divided into two parts. The multi-scale dilated convolution part extracts AU information at different scales and distances; the fusion of channel and spatial attention further enhances the expression ability of key features. The module structure diagram is as Figure 2 shown.
[0030] First, preprocess the facial expression image data to identify the face region and remove the redundant background. The preprocessed facial expression image data passes through two convolutional groups, each containing two convolutional layers. The main purpose is to extract the global low-level features in the image data, such as basic visual elements like edges and textures, laying the foundation for subsequent high-level semantic feature extraction. Next, a parallel three-branch convolutional structure is designed to achieve feature extraction at different scales. Different dilation rates are set for each branch to expand the receptive field without increasing unnecessary parameters, capturing spatial emotional information in different ranges, thus adapting to the changes and distributions of facial expression features at different scales. The first branch uses a 3×3 convolutional kernel with a dilation rate of 1 to moderately expand the receptive field and capture relatively fine facial expression changes. The second branch uses a 3×3 convolutional kernel with a dilation rate of 2 to further expand the receptive field to capture a wider range of local context information. The third branch also uses a 3×3 convolutional kernel with a dilation rate of 3 to provide the widest receptive field that can cover a larger facial area, effectively capturing feature information such as the overall facial contour and large-scale muscle movements. Traditional convolutional networks rely on stacking multiple convolutional layers to increase the receptive field, resulting in an increase in computational complexity and information dilution. Dilated convolution expands the receptive field while maintaining the resolution of the feature map by setting different dilation rates, which is crucial for capturing global expression information and key features. Through multiple parallel branches, the network can simultaneously capture features at different scales, enhancing the model's adaptability and recognition ability for various scale features in the image.
[0031] The three branches extract features at different scales and are concatenated in the channel dimension to synthesize a comprehensive feature map. The merged comprehensive feature map passes through the calibration of two attention mechanisms in parallel: the channel attention mechanism and the spatial attention mechanism. The channel attention mechanism first performs global average pooling on the comprehensive feature map to obtain the global features of each channel, and learns the importance weights of each channel through two fully connected layers. Then, these weights are multiplied element-wise with the original comprehensive feature map to achieve channel weighting. The spatial attention mechanism performs global pooling on the comprehensive feature map in the channel dimension to obtain a spatial feature map. The spatial feature map passes through a 1×1 convolution and a Sigmoid activation function to learn the importance weights of each spatial position, and these weights are multiplied element-wise with the original comprehensive feature map to achieve spatial weighting. Finally, the output feature maps of the channel attention and spatial attention mechanisms are fused to obtain the final enhanced feature map. The enhanced facial features are fed into a fully connected layer for emotion recognition after passing through the last convolutional group, the max-pooling layer, and the batch normalization layer. The Dropout layer randomly discards the outputs of some neurons to prevent overfitting.
[0032] Time-frequency-spatial three-dimensional feature extraction network:
[0033] To make full use of the features in the time, frequency, and spatio dimensions of electroencephalogram (EEG) signals, this application constructs the EEG into a three-dimensional feature structure. First, the original EEG signals are divided into N segments of equal length, each segment being T seconds long. Then, each segment of the signal is decomposed into four frequency bands, namely θ [4 - 8 Hz], α [8 - 14 Hz], β [14 - 31 Hz], and γ [31 - 45 Hz], using band-pass filtering. Differential Entropy (DE) can reflect the rate of change and trend of EEG signals and is the most representative feature in the frequency domain characteristics. Therefore, the DE features of a 0.5 s time window are calculated for each segment of the signal in each frequency band. The differential entropy expression is as follows:
[0034]
[0035] where h(x) is the differential entropy of the signal and f(x) is the probability density function of continuous information.
[0036] For EEG signals that follow a Gaussian distribution where is the variance and μ is the signal mean. Its differential entropy can be expressed as:
[0037]
[0038] First, according to the relative positions of 32 electrodes, the differential entropy feature vectors are transformed into a compact 2D map, constructing an 8×9 two-dimensional matrix, where the positions without electrodes are filled with zeros, as Figure 3 shown. Then, the two-dimensional differential entropy feature maps of different frequency bands are stacked together to obtain a 4×8×9 three-dimensional feature matrix.
[0039] The shape of the input data is (batch_size, 6, 4, 8, 9), which includes 6 time steps, 4 channels, and an 8×9 spatial dimension. To adapt to subsequent processing, the data is reshaped and input into a two-dimensional convolutional layer. The spectral attention module and the spatial attention module are used to process spectral features and spatial features respectively. Both adopt the multi-head attention mechanism, which can extract information from multiple subspaces simultaneously and dynamically allocate weights through the self-attention mechanism to more accurately capture the key features in the input data. Finally, the temporal attention generates weights for each time step through linear transformation and the ReLU activation function, focusing on the features of important time steps, so as to better capture the dynamic changes in the time dimension. The EEG signal outputs a multi-dimensional feature representation of time, frequency, and space through the global attention modules of spectrum, space, and time, and obtains the final emotion classification result through a fully connected layer.
[0040] Self-learning weight module:
[0041] To effectively fuse modal data with different importance and complementarity and prevent low-confidence modalities from affecting the final accuracy, this application proposes a self-learning weight module for achieving decision-level dynamic weighted fusion of electroencephalogram (EEG) signals and facial expressions. The core advantage of this module is that it can automatically learn and adjust the weights of each modality according to the input different modal data, thereby adaptively optimizing the decision fusion strategy in different emotion classification scenarios. The self-learning weight module not only effectively fuses multi-modal data but also automatically learns the optimal fusion strategy during the training process, avoiding the limitations of traditional fixed-weight fusion methods.
[0042] The probability distribution outputs of facial expressions and EEG signals are used as inputs for decision-level fusion and are weighted and fused through the designed self-learning weight module. The specific schematic diagram is as Figure 4 shown. In this process, the probability distribution of each modality is concatenated into a large feature vector, and the feature dimension is gradually expanded through a fully connected layer and the nonlinear activation function GELU is introduced.
[0043] Specifically: Assume that the probability distribution output X E of the EEG signal, and the probability distribution output X F of the facial expression, with the dimension of [N, D], where N is the number of samples and D is the feature dimension. The features of the two modalities are concatenated along the feature dimension to obtain the fused feature representation X fused :
[0044]
[0045] The fused features are nonlinearly transformed through a multi-layer fully connected network. Taking the first layer as an example:
[0046] H1 = GELU(X fused W1 + b1)
[0047] where and are the weights and biases of the i-th layer (i = 1, 2, 3, 4), GELU is the activation function, and H i is the output of the i-th layer (i = 1, 2, 3, 4).
[0048] The features after the nonlinear transformation are reshaped into a sequence form suitable for the attention mechanism:
[0049] H1' = reshape(H1, [N, L, E])
[0050] where L is the sequence length, E is the embedding dimension, and H1' is the output after reshaping.
[0051] Subsequently, it is processed through a multi-head self-attention layer, which simultaneously performs weighted summation on the feature sequence through multiple attention heads. Each attention head independently calculates the linear transformation of queries, keys, and values, thereby dynamically learning the contribution weights of each modality to emotion classification. For the input feature sequence H1', each attention head h calculates its query matrix Q h , key matrix K h , and value matrix V h :
[0052]
[0053] where is the learnable parameter matrix of the h-th attention head; the value matrix is weighted and summed using the attention score matrix to obtain the output H of the h-th attention head h :
[0054]
[0055] d k is the dimension of the key;
[0056] The outputs of all attention heads are concatenated and linearly transformed to obtain the final multi-head self-attention output H a :
[0057] H a = W O (H h1 , H h2 ,..., H h )
[0058] where W O is the final output linear transformation matrix. According to the learned weighted coefficients, the module adaptively weights the information of each modality under different emotions. After passing through a series of fully connected layers again, the features are mapped to the target classification dimension, and the classification probability is output through the Softmax activation function. W out and b out are the weight matrix and bias vector of the output layer:
[0059] Y = Softmax(H4W out + b out )
[0060] The weights are automatically learned during the dynamic weighting process. During the training process, the module adjusts its internal parameters according to the contribution degree of each modality to the emotion classification result. The module does not simply concatenate the two modalities together, but decides how to combine these modalities according to the characteristics of each situation. The self-learning weight module can dynamically adjust the optimal decision fusion strategy according to different emotions.
[0061] Example 2: This example shows the experimental results of method effectiveness verification.
[0062] Data set:
[0063] (1) DEAP: The DEAP data set records the electroencephalogram (EEG) signals of 32 subjects while watching 40 different video clips, as well as the facial videos of the first 22 subjects while watching the videos. The length of each video clip is approximately one minute. For each clip, the subjects provided ratings of valence and arousal levels from 1 to 9 as labels. In this example, these labels were divided into two categories with a threshold of 5. To ensure the consistency of multimodal data, only the data of the first 22 subjects with facial videos were used for analysis in this example. The experimental data of each subject was 63 seconds. After deducting the 3-second preparation time at the beginning, the data length during the actual task was 60 seconds. The EEG signals were downsampled to 128 Hz. The 3-second baseline signal was divided into 6 segments of 0.5 seconds each, and DE features in four frequency bands were extracted from each segment. To ensure the consistency of the facial expression data of the subjects with the EEG signals, the facial videos were captured every 0.5 seconds. Therefore, each subject generated 40 * 60 * 2 = 4800 facial expression images, and a total of 105,600 images were generated for 22 subjects. To improve the recognition accuracy of the facial images, precise cutting of the images was adopted to effectively remove the irrelevant background, thereby highlighting the facial features and enhancing the usability of the images.
[0064] (2) FER-2013: The FER-2013 data set consists of 35,887 grayscale images, each with a resolution of 48×48 pixels and in a single-channel (grayscale) format. These images are classified into seven basic expressions: angry, disgusted, fearful, happy, sad, surprised, and neutral, thus providing a broad expression classification system for training and evaluating facial expression recognition algorithms. The data set is divided into three parts: the training set contains approximately 28,709 images, the public test set contains approximately 3,589 images, and the private test set also contains approximately 3,589 images, facilitating systematic training and performance evaluation.
[0065] (3) CK+: The CK+ is a facial expression dataset extended based on Cohn-Kanade (CK). It adds facial expression data of 123 adults to the original 97 subjects. The participants' ages range from 18 to 50 years old, 69% are female, 81% are European Americans, 13% are African Americans, and 6% are from other groups. The subjects performed a series of 23 facial displays, including single action units and combinations of action units. Each person slowly changed from a natural state to a given expression, including seven emotions: neutral, angry, contemptuous, disgusted, fearful, happy, sad, and surprised. Each expression is a set of sequences, with a total of 593 image sequences, and 327 of these sequences have emotional sequence labels. The last three frames closest to the peak emotion were extracted from each sequence as experimental data.
[0066] Result analysis:
[0067] (1) Effectiveness of the multi-scale dilated attention convolution module: To verify the effectiveness of the multi-scale dilated attention convolution module in the facial recognition task, it was compared with existing methods on the Fer2013 and CK+ datasets, as shown in Table 1. The final experimental results demonstrate the excellent classification and recognition performance of the module proposed in this application.
[0068] Table 1 Comparison of the accuracy of the multi-scale dilated attention convolution module and other facial emotion recognition algorithms
[0069]
[0070] Among them, Document 1: Yang F. AttenVGG: Integrating Attention Mechanisms with VGGNet for Facial Expression Recognition [C] / / 2024 5th International Conference on Electronic Communication and Artificial Intelligence (ICECAI). IEEE, 2024: 332 - 337.
[0071] Document 2: Yang L, Zhang H, Li D, et al. Facial expression recognition based on transfer learning and SVM [C] / / Journal of Physics: Conference Series. IOP Publishing, 2021, 2025(1): 012015.
[0072] Document 3: Minaee S, MinaeiM, AbdolrashidiA. Deep-emotion: Facial expression recognition using attentional convolutional network [J]. Sensors, 2021, 21 (9): 3046.]
[0073] Document 4: Zu F, Zhou CJ, Wang
[0074] The multi-scale atrous attention convolution module proposed in this application achieved an average accuracy of 74.1% on the Fer2013 dataset and 99.69% on the CK+ dataset. The average accuracy of all models on the Fer2013 dataset was significantly lower than that on the CK+ dataset. This may be due to the low data quality of the Fer2013 dataset, which consists of 48×48 grayscale images and contains a large number of side profile and occluded facial images, which affected the experimental results. Experimental results demonstrate that the multi-scale atrous attention convolution module demonstrates excellent classification and recognition performance in the facial emotion recognition task. By introducing multi-scale atrous convolution and the channel-spatial attention mechanism, it can more effectively capture the multi-scale features and key regions of facial action units, significantly improving classification accuracy.
[0075] To verify the parameter advantage of multi-scale hole convolution over ordinary multi-scale convolution, Table 2 lists the parameter amount required for the MSDAC designed in this application and the parameter amount of the multi-scale attention convolution (MSAC) using ordinary convolution. From the data in Table 2, it can be seen that compared with the MSAC network, while maintaining the same input size, the parameter size (Paramssize) of MSDAC is reduced by about 0.13MB, and the total storage requirement (Estimated Total Size) is reduced by 6.33MB. The accuracy of MSDAC in the Valence and Arousal dimensions on the DEAP dataset reached 98.05% and 96.15% respectively, which are better than the results of 97.62% (Valence) and 95.94% (Arousal) of the MSAC network. This shows that MSDAC can reduce the computational complexity while still maintaining good recognition results.
[0076] Table 2 Comparison of the Parameter Quantities of MSDAC and MSAC
[0077]
[0078] To further verify the effectiveness of each module in the multi-scale dilated attention convolution module, an ablation experiment was designed. The channel-spatial attention module (C-SA) and the multi-scale dilated convolution module (MSDC) were removed in turn, and the performance was compared with that of the complete model on the DEAP dataset. The results are shown in Table 3.
[0079] Table 3 Ablation Experiment on the DEAP Dataset
[0080]
[0081] Experiments show that on the DEAP dataset, the valence accuracy rate is 92.47% and the arousal accuracy rate is 91.53% without the C-SA and MSDC modules; after adding the multi-scale dilated module, the facial emotion recognition accuracy rate is improved to 96.27% (valence) and 94.31% (arousal); when both the C-SA and MSDC modules are added, the accuracy rate reaches 98.05% (valence) and 96.15% (arousal). The above results show that the channel-spatial attention enhances the merged feature map, effectively improving the accuracy of the overall model. The combination of the multi-scale dilated convolution and the channel-spatial attention module can accurately capture the key facial expression features, significantly improving the overall performance of the model. The two modules play an important role in the model. Especially when dealing with complex facial expression data, their synergistic effect is crucial for facial emotion recognition.
[0082] To verify the effectiveness of the self-learning weight module, the accuracy rates of electroencephalogram (EEG) unimodal, facial unimodal, ordinary weighted multi-modal fusion, and self-learning weight multi-modal fusion were compared. Ordinary weighting assigns fixed weights to the decision results of different modalities and fuses them in the way of weighted summation. The formula is expressed as:
[0083] X fused = kX E +(1 - k)X F
[0084] where k ∈ [0, 1] is the fixed fusion weight.
[0085] Table 4 Verification of the Effectiveness of the Self-learning Weight Module
[0086]
[0087] The experimental results show that the combination of electroencephalogram (EEG) signals and facial expressions can effectively improve the accuracy of emotion recognition, showing a significant improvement compared to the accuracy of single-modal emotion recognition. Multimodal fusion can integrate information from different modalities and reduce the uncertainty brought by a single modality. The ordinary weighted multimodal fusion method achieved relatively high accuracy in both valence and arousal classification tasks, but it was still lower than the self-learning weight multimodal fusion method. The self-learning weight module demonstrated the best performance in multimodal fusion, being 1.41% higher than the ordinary weighted fusion method in the valence dimension and 1.02% higher in the arousal classification task. Through the self-learning weight module, the information of EEG and facial modalities can be more effectively fused, thus significantly improving the classification accuracy.
[0088] To demonstrate the recognition performance of the multimodal emotion recognition model of this application in the valence and arousal dimensions, confusion matrices on the DEAP dataset were plotted. Figure 5 (a) is the confusion matrix of emotion recognition in the valence dimension of the DEAP dataset, Figure 5 (b) is the confusion matrix of emotion recognition in the arousal dimension.
[0089] To further verify the effectiveness of the multimodal emotion recognition proposed in this application, this method was compared with other multimodal emotion recognition methods, and the results are shown in Table 5.
[0090] Table 5 Comparison between the multimodal method in this paper and other multimodal methods
[0091]
[0092] Among them, Reference 5: Zhang Y, Liu H, Wang D, et al. Cross-modal credibility modelling for EEG-based multimodal emotion recognition[J]. 2024.
[0093] Reference 6: Liao J, Zhong Q, Zhu Y, et al. Multimodal physiological signal emotion recognition based on convolutional recurrent neural network[C] / / IOP conference series: materials science and engineering. IOP Publishing, 2020, 782(3): 032005.
[0094] Document 7: Zhu Q, Zheng C, Zhang Z, et al. Dynamic Confidence-Aware Multi-ModalEmotion Recognition[J]. IEEE Transactions onAffective Computing, 2023.
[0095] Document 8: WuY, LiJ.Multi-modalemotionidentificationfusingfacialexpressionandEEG[J].MultimediaTools andApplications, 2023,82(7):10901-10919.
[0096] Document 9: SalamaE S, El-KhoribiRA, ShomanME, etal. A3D-convolutional neural network framework with ensemble learning techniques for multi-modal emotion recognition [J]. Egyptian Informatics Journal, 2021, 22(2): 167-176.
[0097] Document 10: Wang S, Qu J, Zhang Y, et al. Multimodal emotion recognition from EEG signals and facial expressions[J]. IEEEAccess, 2023, 11: 33061-33068.
[0098] Experimental results demonstrate that the combination of EEG and facial expressions performs exceptionally well among multimodal fusion methods. References 5 to 10 all employed EEG signals and facial expressions, achieving slightly higher results than EEG combined with peripheral physiological signals or ECG signals. This suggests that combining EEG and facial expression features can more accurately identify emotional states. This method achieved the highest accuracy on the DEAP dataset, with a valence accuracy of 98.66% and an arousal accuracy of 97.49%, validating its effectiveness.
[0099] In summary, this application proposes a multi-modal emotion recognition method based on the fusion of electroencephalogram (EEG) signals and facial expressions. Experimental results show that by designing a multi-scale dilated attention convolutional module and introducing an EEG signal recognition network based on time-frequency-spatial three-dimensional features, the model's ability to extract emotion features and recognition accuracy can be significantly improved. The multi-scale dilated attention convolutional module for facial emotion recognition has an average accuracy of 74.1% on FER2013, 99.69% on CK+, and on the DEAP dataset, the valence and arousal accuracies reach 98.05% and 96.15% respectively, which is better than existing methods. By combining EEG and facial modalities and making full use of their respective advantages, the designed decision-level fusion method with self-learning weights further enhances the model's attention to key information and dynamic fusion ability of multi-modal information. On the DEAP dataset, the emotion recognition accuracies in the valence dimension and arousal dimension reach 98.66% and 97.49% respectively. The above results prove the effectiveness and superiority of the model in the field of emotion recognition. EEG signals provide internal physiological information of emotions, while facial expressions provide external behavioral information of emotions. The combination of the two can more comprehensively capture the multi-dimensional features of emotions.
[0100] Obviously, the above embodiments are merely examples for clear illustration and not limitations on the implementation manners. For those of ordinary skill in the art, other different forms of changes or modifications can be made based on the above description. It is not necessary and impossible to enumerate all implementation manners here. The obvious changes or modifications derived therefrom still fall within the protection scope of the present invention.
Claims
1. A self-learning multi-modal emotion recognition method based on multi-scale hollow attention, characterized by the following steps Including: The obtained facial expression image is preprocessed and then input into a multi-scale dilated attention convolutional module. The multi-scale dilated attention convolutional module extracts features of different scales through a parallel three-branch convolutional structure. The three-branch convolutional structure is set such that the convolutional kernel sizes of each branch are the same, and the dilation rates are different from each other. After the features output by the three-branch convolutional structure are concatenated in the channel dimension, they are calibrated by an attention mechanism to obtain an enhanced facial expression feature map, which is sent to a fully connected layer for emotion recognition. The original electroencephalogram (EEG) signal is input into a spatio-temporal-frequency three-dimensional feature extraction network. The spatio-temporal-frequency three-dimensional feature extraction network decomposes the original EEG signal and then calculates differential entropy features. The calculation results of the differential entropy features are processed by a global attention module including a spectral attention module, a spatial attention module, and a temporal attention module, and then a spatio-temporal-frequency multi-dimensional feature representation is output, and emotion recognition is performed through a fully connected layer. The facial expression emotion recognition result and the EEG signal emotion recognition result are input into a self-learning weight module, and the final emotion recognition result is generated through dynamic weighted fusion.
2. The self-learning multi-modal emotion recognition method with multi-scale hollow attention according to claim 1, characterized in that, The specific operations performed within the multi-scale dilated attention convolutional module are as follows: The preprocessed facial expression image data passes through two convolutional groups, each convolutional group containing two convolutional layers, to extract global low-level features in the image data. The global low-level features are input into the three-branch convolutional structure, and the three branches extract features of different scales, which are concatenated in the channel dimension to form a comprehensive feature map. The combined comprehensive feature map is calibrated in parallel by a channel attention mechanism and a spatial attention mechanism. Finally, the output feature maps of the channel attention mechanism and the spatial attention mechanism are fused to obtain the final enhanced feature map. The enhanced feature map passes through a convolutional group, a max-pooling layer, and a batch normalization layer, and then is sent to a fully connected layer for emotion recognition.
3. A self-learning multi-modal emotion recognition method based on multi-scale hollow attention according to claim 2, characterized in that, The specific settings of the three-branch convolutional structure are as follows: The first branch uses a 3×3 convolutional kernel with a dilation rate of 1; the second branch uses a 3×3 convolutional kernel with a dilation rate of 2; the third branch uses a 3×3 convolutional kernel with a dilation rate of 3.
4. A self-learning multi-modal emotion recognition method based on multi-scale hole attention according to claim 3, characterized in that The channel attention mechanism first performs global average pooling on the comprehensive feature map to obtain the global features of each channel, learns the importance weights of each channel through two fully connected layers, and then multiplies these weights with the original comprehensive feature map channel by channel to achieve channel weighting. The spatial attention mechanism performs global pooling on the comprehensive feature map in the channel dimension to obtain a spatial feature map. The spatial feature map passes through a 1×1 convolution and a Sigmoid activation function to learn the importance weights of each spatial position, and multiplies these weights with the original comprehensive feature map element by element to achieve spatial weighting.
5. A self-learning multi-modal emotion recognition method based on multi-scale dilated attention, characterized in that The specific operations performed within the spatio-temporal-frequency three-dimensional feature extraction network are as follows: First, the original EEG signal is divided into N segments of equal length, each segment being T seconds long. Then, each segment of the signal is decomposed into four frequency bands using a band-pass filter, and the differential entropy features of 0.5 s are calculated for each segment of the signal on each frequency band. The differential entropy feature vector is transformed into a 2D map according to the relative positions of the 32 electrodes, constructing an 8×9 two-dimensional matrix, where the positions without electrodes are filled with zeros. Then, the 2D differential entropy feature maps of different frequency bands are stacked together to obtain a 4×8×9 three-dimensional feature matrix; The three-dimensional feature matrix is processed by the frequency spectrum attention module, spatial attention module, and temporal attention module to output a multi-dimensional feature representation of time-frequency-space, and the final emotion recognition result is obtained through a fully connected layer.
6. A self-learning multi-modal emotion recognition method based on multi-scale dilated attention according to claim 1, characterized in that The specific operations performed within the self-learning weight module are as follows: The facial expression emotion recognition result and the electroencephalogram signal emotion recognition result are concatenated along the feature dimension to obtain a fused feature representation. The fused feature representation is subjected to a non-linear transformation through a multi-layer fully connected network, and the non-linearly transformed features are reshaped into a sequence form and input into a multi-head self-attention layer for processing. This layer simultaneously performs weighted summation on the feature sequence through multiple attention heads. Each attention head independently calculates the linear transformation of the query, key, and value, dynamically learning the contribution weights of the facial expression modality and the electroencephalogram signal modality to emotion classification. The outputs of all attention heads are concatenated and then subjected to a linear transformation to obtain the final weighted coefficients; Finally, according to the obtained weighted coefficients, the information of each modality is adaptively weighted under different emotions, and the features are mapped to the target classification dimension through a series of fully connected layers again, and the classification recognition probability is output through an activation function.
Citation Information
Cited By
Depression state prediction method and system based on voice multi-scale time domain perception
CN120783803A
Neurological disease diagnosis system based on space-time attention and dynamic domain self-adaption
CN120932823A
Neurological disease diagnosis system based on spatio-temporal attention and dynamic domain adaptation
CN120932823B
Two-way emotion recognition method based on electroencephalogram micro-state features
CN121278476A
Emotion recognition method based on multi-modal signal fusion
CN121287145A