A facial expression recognition method based on multi-scale feature cross fusion and contrast multi-head attention

By employing multi-scale feature cross-fusion and contrast-separated multi-head attention methods, the problems of local feature focusing and semantic consistency loss in facial expression recognition were solved, achieving high-accuracy facial expression recognition with a recognition rate of 89.93%.

CN117079328BActive Publication Date: 2025-11-21BEIHANG UNIV
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202311049241.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-08-20
Publication Date
2025-11-21
Estimated Expiration
2043-08-20

AI Technical Summary

Technical Problem

Existing deep learning-based facial expression recognition methods are prone to problems such as local feature focusing and loss of semantic consistency before and after horizontal flipping, making it difficult to effectively handle the diversity and complexity of expression recognition.

Method used

We employ a multi-scale feature cross-fusion and contrastive separation multi-head attention approach. Through a multi-scale feature cross-fusion module with a reward and punishment mechanism, a separation multi-head attention module, and an attention consistency balancing framework based on contrastive learning, we improve sample separability and attention consistency, thus addressing the diversity and complexity of facial expression features.

Benefits of technology

It significantly improved the accuracy of facial expression recognition to 89.93%, and improved it by 3.03% to 5.15% compared with existing methods on the RAF-DB dataset. It effectively reduced the error rate of personalized expressions and enhanced the semantic consistency of images before and after flipping.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN117079328B_ABST
    Figure CN117079328B_ABST
Patent Text Reader

Abstract

The application provides a facial expression recognition method based on multi-scale feature cross fusion and contrast separation multi-head attention. The effectiveness is verified by using the public facial expression dataset RAF-DB: first, the training set and the test set are established; second, the training set is data enhanced, and a multi-scale feature cross fusion and contrast separation multi-head attention network is designed, and the training set is input into the model for training; finally, the test set is input into the trained network for performance test. The main innovation of the application is that a multi-scale feature cross fusion module based on a reward and punishment mechanism is designed, so that the model learns more rich features, and the separability of the sample is improved; a separation multi-head attention module is proposed to model feature information from multiple regions; an attention consistency balancing framework based on contrast learning is designed to calculate the consistency loss of the attention heat map, improve the learning of the latent knowledge, and the recognition accuracy of the model on the RAF-DB dataset is 89.93%, which is better than the existing optimal method.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention provides a facial expression recognition method based on multi-scale feature cross-fusion and contrast separation multi-head attention, which offers a new approach for the application of computer technology in the field of facial expression recognition. Background Technology

[0002] Facial expressions, as a crucial means of communication, can convey rich emotions, cognition, and intentions. Facial Expression Recognition (FER) is the process of converting images into facial expression annotation categories. In the medical field, facial expressions can assist doctors in diagnosing numerous mental illnesses such as depression, schizophrenia, anxiety disorders, and bipolar disorder. In the field of human-computer interaction, given my country's increasingly severe aging population, establishing human-computer emotional interaction systems can enable barrier-free communication between humans and robots. In the transportation sector, to reduce road traffic accidents caused by fatigued driving, in-vehicle systems for driver fatigue detection can be designed to promptly alert drivers based on their fatigue levels.

[0003] Facial expression feature extraction is crucial for facial expression recognition. Currently, feature extraction methods are generally divided into two categories: geometric feature-based and texture feature-based. Geometric feature-based methods utilize changes in facial shape to describe facial expressions, typically requiring manual annotation of facial feature points or automatic localization using feature point detection algorithms, which demands high accuracy and cost. Texture feature-based methods, on the other hand, utilize changes in facial texture and color, such as local binary patterns, Gaussian mixture models, and deep learning models. These methods usually do not require manual annotation or detection of facial feature points, but directly extract features from the entire facial image, capturing more facial expression details and variations, and adapting to different face shapes and expression intensities. However, research on facial expression recognition remains highly challenging. First, facial expressions are complex information conveyed through various non-verbal forms, such as eye contact, based on physiological changes in facial muscles. Second, different people express the same emotion in different ways, and the same person may express the same emotion in different states. Therefore, effectively handling the diversity and complexity of expression recognition is a significant challenge. Summary of the Invention

[0004] To address the issues of existing deep learning-based facial expression recognition methods easily falling into local feature focusing and loss of semantic consistency before and after horizontal flipping, this invention provides a facial expression recognition method based on multi-scale feature cross-fusion and contrastive separation multi-head attention. The proposed method is validated on the internationally publicly available facial expression recognition dataset RAF-DB through four steps: data processing, model construction, model training, and model testing. First, the obtained RAF-DB dataset is preprocessed to establish training and testing sets. Second, a multi-scale feature cross-fusion and contrastive separation multi-head attention network model is constructed. Then, the preprocessed training set is input into the constructed model for training. Finally, the preprocessed test set is input into the trained model for performance testing, and the accuracy of facial expression recognition is output.

[0005] A facial expression recognition method based on multi-scale feature cross-fusion and contrastive separation multi-head attention according to an embodiment of the present invention includes the following steps:

[0006] Step 1: Obtain the parameters of the internationally publicly available color facial expression recognition dataset RAF-DB and the ResNet-50 pre-trained model;

[0007] Step 2: Select images from the dataset for training and testing respectively, standardize the size of the training images, design a random data augmentation step, establish the training dataset, standardize the size of the test images, and establish the test dataset;

[0008] Step 3: Use the PyTorch deep learning framework and combine it with the ResNet-50 network to build a multi-scale feature cross-fusion and contrast separation multi-head attention network;

[0009] Step 4: Load the pre-trained model parameters obtained in Step 1 into the multi-scale feature cross-fusion and contrast separation multi-head attention network established in Step 3, and then input the training dataset established in Step 2 into the network to train the facial expression recognition model. After the training is completed, save the model parameters.

[0010] Step 5: Load the model parameters saved in Step 4 to obtain the trained facial expression recognition model. Input the test set established in Step 2 into the model to obtain the facial expression recognition results.

[0011] in:

[0012] In step 2, data preprocessing includes random rotation, random cropping, and random erasure.

[0013] In step 3, the constructed multi-scale feature cross-fusion and contrastive separation multi-head attention network includes: a multi-scale feature cross-fusion module based on a reward and punishment mechanism, a separation multi-head attention module, and an attention consistency balancing framework based on contrastive learning. The multi-scale feature cross-fusion module based on a reward and punishment mechanism is used to extract multi-level features from abstract to concrete and improves the separability of samples with the help of the reward and punishment mechanism. The separation multi-head attention module is used to focus on multiple local feature locations and uses a dispersion loss function to guide multiple attention heads to focus on different local feature locations. The attention consistency balancing framework based on contrastive learning is used to improve the attention consistency of images before and after flipping and can learn potential semantic consistency.

[0014] In step 4:

[0015] When training the multi-scale feature cross-fusion and contrast separation multi-head attention network model, the parameters of the multi-scale feature cross-fusion and contrast separation multi-head attention network are updated iteratively through error backpropagation and stochastic gradient descent. The training data batch size is set to 256, the model learning rate is set to 0.0001, the Adam optimizer is used to optimize the model parameters, and the cross-entropy loss function is used as the loss function. After 500 iterations of training, the model parameters are saved.

[0016] The main advantages of the facial expression recognition method based on multi-scale feature cross-fusion and contrastive separation multi-head attention proposed in this invention include:

[0017] 1. To address the inherent problems of inter-class similarity and intra-class differences, a multi-scale feature cross-fusion module based on a reward and punishment mechanism is proposed. Due to the diversity and complexity of facial expression features, sample categories are difficult to distinguish; therefore, combining a loss function with a reward and punishment mechanism improves the separability of samples. The multi-scale feature cross-fusion module is used to extract multi-level features;

[0018] 2. To address the problem of models easily getting stuck in local feature focus, a separate multi-head attention module is proposed. This method guides multiple attention heads to model facial expression features at different locations, which can significantly reduce the error rate of personalized expressions;

[0019] 3. To address the issue of semantic consistency loss after image horizontal flipping, a contrastive learning-based attention consistency balancing framework is proposed. This framework eliminates attentional bias in the model by comparing attention heatmaps before and after image flipping. Attached Figure Description

[0020] Figure 1 This is a flowchart of a facial expression recognition method according to an embodiment of the present invention.

[0021] Figure 2This is a diagram of a multi-head attention network framework based on multi-scale feature cross-fusion and contrast separation according to an embodiment of the present invention.

[0022] Figure 3 This is a diagram of a multi-scale feature cross-fusion module based on a reward and punishment mechanism according to an embodiment of the present invention.

[0023] Figure 4 This is a diagram of a split multi-head attention module according to an embodiment of the present invention.

[0024] Figure 5 This is a structural diagram of an attention head unit according to an embodiment of the present invention.

[0025] Figure 6 This is a diagram of an attention consistency balancing framework based on contrastive learning according to an embodiment of the present invention. Detailed Implementation

[0026] The overall process of a facial expression recognition method based on multi-scale feature cross-fusion and contrast separation multi-head attention according to an embodiment of the present invention is as follows: Figure 1 As shown, it includes:

[0027] Step S1: Obtain the internationally publicly available RAF-DB dataset for color facial expression recognition and the parameters of the ResNet-50 pre-trained model;

[0028] Step S2: Select the images from RAF-DB for training and testing respectively, including:

[0029] Standardize the size of training images, design random data augmentation, and establish a training dataset.

[0030] Standardize the size of the test images and create a test dataset.

[0031] Specifically, it includes:

[0032] Step S2.1: First, define the training set D = {(X i ,y i The training set images are resized to 224×224 pixels, i = 1, ..., N.

[0033] Step S2.2: Then, use three data augmentation methods: random rotation, random cropping, and random erasure.

[0034] Step S2.3: Randomly rotate the image between (-20, 20) and randomly select an angle of 0.2 to rotate it around the center.

[0035] Step S2.4: Randomly crop the image. First, fill the top, bottom, left and right sides of the image with 32 pixels of 0, and then randomly crop 224×224 pixels from 288×288 pixels with a probability of 0.2.

[0036] Step S2.5: Randomly erase the image, select the occlusion ratio between (0.02, 0.25) to crop a square area, and fill it with 0;

[0037] Step S3: Using the PyTorch deep learning framework, a multi-scale feature cross-fusion and contrast separation multi-head attention network is built using the ResNet-50 network. This multi-scale feature cross-fusion and contrast separation multi-head attention network (e.g., Figure 2 (As shown) includes:

[0038] Specifically, it includes:

[0039] A) A multi-scale feature cross-fusion module based on a reward and punishment mechanism is used to extract multi-level features from abstract to concrete, and then improve the separability of samples through clustering. The multi-scale feature cross-fusion module based on a reward and punishment mechanism is as follows: Figure 3 As shown, it includes:

[0040] A1) First, the ResNet-50 pre-trained model is used as the backbone network to extract the feature pyramid of the image. The feature pyramid is composed of abstract to concrete feature maps as shown in Equation 14:

[0041]

[0042] in, Let G represent the features output by the i-th stage of the ResNet-50 pre-trained model, where G represents the number of layers in the feature pyramid, and a 1×1 convolution is used to unify the number of channels to the same level. same;

[0043] A2) Secondly, the multi-scale feature cross-module based on the reward and punishment mechanism contains a total of three feature pyramids, namely F in F t and F out The first pyramid F in To the middle pyramid F t This is used to diffuse concrete features to abstract features. The feature H×W is expanded to 2H×2W through bilinear interpolation upsampling. After being added to the more abstract features, a 3×3 convolution is used for further fusion. The intermediate pyramid F... t To the end of the pyramid F out This method is used to diffuse abstract features to concrete features. It halves the H×W size through pooling, adds it to more specific features, and then further fuses them using a 3×3 convolution. This achieves cross-fusion of multi-scale features and unifies the feature scale of the final pyramid using pooling layers. Add up items of the same size;

[0044] A3) Then, for the fused feature map, this invention improves the center loss function and proposes a reward-penalty mechanism loss function that not only rewards the movement of sample features towards their class center but also penalizes the movement of sample features towards non-class centers. From C ′ M class centers C = {C} are randomly selected from a Gaussian distribution. i Let |i=1,…,M} be set as learnable parameters. The loss function of the reward and punishment mechanism is shown in Equation 15:

[0045]

[0046] In the formula Let C' represent the loss function of the reward and punishment mechanism. i The number of channels, ε is 10. -6 λ is used to avoid the denominator being zero, and λ is used to control the weights of the two parts.

[0047] B) A separate multi-head attention module, used to focus on different local locations of the feature map, wherein the separate multi-head attention module is as follows: Figure 4 As shown, it includes:

[0048] B1) First, the multi-head attention module contains multiple parallel attention heads in the horizontal direction. Each head consists of a separate structure of spatial attention units and channel attention units. The input of each attention head is the aggregated feature map output. The attention heads are as follows: Figure 5 As shown;

[0049] B2) Secondly, the spatial attention unit uses Inception and residual structures. Each spatial attention unit contains V Inception-residual blocks, each block consisting of three layers of convolutional operations with 1×3, 3×3, and 3×1 kernels, a ReLU activation function, and a residual structure. V is defined as a hyperparameter. The Inception structure is used to make the model focus on different feature map scales. The role of the spatial attention unit is shown in Equation 16:

[0050]

[0051] In the formula, W represents the spatial attention unit. sa Indicated parameter, This represents the output set of the multi-head spatial attention unit. This represents the output of the j-th head;

[0052] B3) Then, the channel attention unit contains W linear-residual blocks, each block consisting of a linear transformation layer, a ReLU activation function layer, and a residual structure. The function of the channel attention unit is shown in Equation 17:

[0053]

[0054] In the formula, W represents the channel attention unit. ca express The parameters, This represents the output set of a multi-channel attention unit. This represents the output of the j-th header;

[0055] B4) Finally, to overcome the limitation that the multi-head attention module simply stacks H attention heads horizontally without guaranteeing that each head focuses on different positions of the feature map, thus failing to achieve the goal of focusing on global information, a dispersion loss function is used to distribute the attention of multiple attention heads. The definition of the dispersion loss function is as follows:

[0056] As shown in Equation 18:

[0057]

[0058] In the formula, H represents the number of attention heads. Let represent the variance of the j-th channel at point (k, l) in the feature map of the i-th sample.

[0059] C) A contrastive learning-based attention consistency balancing framework first involves horizontally flipping the original image during image preprocessing to obtain a flipped image. Then, the original and flipped images are sequentially input into a multi-scale feature cross-fusion module based on a reward-penalty mechanism and a multi-head attention module. Finally, the two feature maps are used to calculate the contrastive loss function. This framework adds constraints to the feature learning process, improving the consistency of the images before and after flipping. The specific details of the contrastive learning-based attention consistency balancing framework are as follows: Figure 6 As shown, it includes:

[0060] C1) First, the original image is augmented to obtain a flipped image. The image pair consisting of the original image and the flipped image is simultaneously input into an encoder consisting of a multi-scale feature cross-fusion module based on a reward and punishment mechanism and a multi-head attention module.

[0061] C2) Then, the contrastive loss is calculated on the two feature maps obtained. The definition of the contrastive loss function is as follows:

[0062] As shown in equations 19 and 20:

[0063]

[0064]

[0065] In the formula, Defined as the features output from the spatial attention unit of the original image. This represents the features of the image flipped from the spatial attention unit's output. Flip means flipping the image. Flip to Can and Pixel positions correspond;

[0066] Step S4: Load the pre-trained model parameters obtained in Step S1 into the multi-scale feature cross-fusion and contrast separation multi-head attention network established in Step S3, and then input the training dataset established in Step S2 into this network to train the facial expression recognition model. After training is complete, save the model parameters, including:

[0067] The parameters of the multi-scale feature cross-fusion and contrast separation multi-head attention network are updated iteratively by backpropagation of error and stochastic gradient descent. The training data batch size is set to 256, the model learning rate is set to 0.0001, the Adam optimizer is used to optimize the model parameters, and the cross-entropy loss function is used as the loss function. The model parameters are saved after 500 iterations of training.

[0068] Step S4.1: Load the ResNet-50 pre-trained model parameters obtained in step S1 into the ResNet-50 backbone network of the multi-scale feature cross-fusion and contrast separation multi-head attention network established in step S3;

[0069] Step S4.2 For the preprocessed training set D = {(X i ,y i First, the training set samples (i = 1, ..., N) are horizontally flipped to obtain D. ori and D filp Secondly, it is input into a multi-scale feature cross-fusion and contrastive separation multi-head attention network loaded with ResNet-50 pre-trained model parameters, to obtain feature X output by the multi-scale feature cross-fusion module based on a reward and punishment mechanism. i ", calculate rewards and penalties for losses;

[0070] Step S4.3: The feature X output by the multi-scale feature cross-fusion module based on the reward and punishment mechanism is... i "In the input separation multi-head attention module, the outputs of multiple attention heads are obtained." Calculate the dispersion loss;

[0071] Step S4.4: Analyze the features output by the spatial attention units of the two augmented datasets. and Calculate the contrast loss;

[0072] Step S4.5: Then calculate the probability output of these two samples belonging to each category, i.e., out. ori and out flip The error between the model output and the label is calculated based on the cross-entropy loss function, which is shown in Equation 21:

[0073]

[0074] In the formula, Let p(y=j|x) represent the cross-entropy loss function. i ) represents the probability value calculated using the softmax function.

[0075] Step S4.6: Finally, combining the above four losses, the final training loss function is obtained as shown in Equation 22:

[0076]

[0077] In the formula, λ1, λ2 and λ3 are respectively and The weights are used for the parameters of the multi-scale feature cross-fusion module and the multi-head attention module. Backpropagation, updating and iterating 500 times;

[0078] Step S5: Load the model parameters saved in Step S4 to obtain the trained facial expression recognition model. Input the test set established in Step S2 into the model to obtain the facial expression recognition results, including:

[0079] Step S5.1: Place the test sample x test The input is fed into a pre-trained multi-scale feature cross-fusion and contrastive separation multi-head attention network to obtain the probability output of the sample belonging to each category, i.e., out. test out test The category with the highest probability is the category of the test sample;

[0080] Step S5.2: The classification performance of the facial expression recognition method with multi-scale feature cross-fusion and contrast separation multi-head attention is described using the recognition accuracy Acc, which is calculated as shown in Formula 23 below:

[0081]

[0082] In the formula N true N represents the number of correctly classified items. total This represents the total number of samples.

[0083] Through experiments, the model proposed in this invention achieved an accuracy of 89.93% on the RAF-DB dataset, reaching state-of-the-art performance. Quantitative performance comparisons with other methods are shown in Table 1. Compared to RAN, which also utilizes a region attention mechanism, this method captures more features from different regions and improves accuracy by 3.03%. The center loss function is used for DACL; compared to DACL, this invention improves the center loss function, achieving a 2.15% improvement.

[0084] Table 1. Performance comparison of the proposed method and the existing best model on the RAF-DB dataset.

[0085]

[0086] To test the impact of model settings on model performance, this section conducted extensive ablation experiments, demonstrating the contributions of the multi-scale feature cross-fusion module based on a reward-penalty mechanism, the reward-penalty loss function, the dispersion loss function, the contrastive loss function, and the split multi-head attention module to the model, as shown in Table 2. Here, CE represents the cross-entropy loss function, RPLowss represents the reward-penalty loss function, MHA is the split multi-head attention module, CLoss is the contrastive loss function, ZLoss is the dispersion loss function, and MSCF is the multi-scale feature cross-fusion module. When the model uses only the backbone network, the test result is only 86.46%. Adding a clustering loss function increases the separability of samples, improving the model's accuracy to 87.34%. Notably, adding the split multi-head attention module significantly improves the recognition result, with the final model recognition rate reaching 89.93%.

[0087] Table 2 Performance Comparison of Different Scheme Models

[0088]

[0089] The foregoing has provided a detailed description of the facial expression recognition method based on multi-scale feature cross-fusion and contrastive separation multi-head attention provided by this invention. However, it is clear that the scope of this invention is not limited thereto. Various modifications to the above embodiments without departing from the scope of protection defined by the appended claims are within the scope of this invention.

Claims

1. A facial expression recognition method based on multi-scale feature cross-fusion and contrast separation multi-head attention, characterized in that... include: Step S1: Obtain the parameters of the internationally publicly available color facial expression recognition dataset RAF-DB and the ResNet-50 pre-trained model; Step S2: Select images from the dataset for training and testing respectively, unify the size of the training images and design a random data augmentation step to establish the training dataset, unify the size of the test images and establish the test dataset; Step S3: Use the PyTorch deep learning framework and combine it with the ResNet-50 network to build a multi-scale feature cross-fusion and contrast separation multi-head attention network; Step S4: Load the pre-trained model parameters obtained in step S1 into the multi-scale feature cross-fusion and contrast separation multi-head attention network, and then input the training dataset established in step S2 into the network to train the facial expression recognition model. After the training is completed, save the model parameters. Step S5: Load the model parameters saved in step S4 to obtain the trained facial expression recognition model. Input the test set established in step S2 into the model to obtain the facial expression recognition results. in: Step S2 includes random rotation, random cutting, and random erasure. Let the training dataset be D = {(X} i ,y i )|i=1,…,N}, where: N represents the number of labeled training data; This represents an input image containing C channels, with a width and height of H and W respectively. y i These are labels for each image: 1 indicates surprise, 2 indicates fear, 3 indicates disgust, 4 indicates happiness, 5 indicates sadness, 6 indicates anger, and 7 indicates neutrality. The step S3 includes: A) Construct a multi-scale feature cross-fusion module based on a reward and punishment mechanism to extract multi-level features from abstract to concrete, and then improve the separability of samples through clustering. The operation of the multi-scale feature cross-fusion module based on the reward and punishment mechanism includes: A1) The ResNet-50 pre-trained model is used as the backbone network to extract the feature pyramid of the image. The feature pyramid consists of abstract to concrete feature maps and is represented as follows: in, Let G represent the features output by the i-th stage of the ResNet-50 network, where G represents the number of layers in the feature pyramid, and a 1×1 convolution is used to unify the number of channels to the same level. same; A2) The multi-scale feature cross-fusion module based on a reward and punishment mechanism contains three feature pyramids, namely F in F t and F out The first pyramid F in To the middle pyramid F t This is used to diffuse concrete features to abstract features. The feature H×W is expanded to 2H×2W through bilinear interpolation upsampling. After being added to the more abstract features, a 3×3 convolution is used for further fusion. The intermediate pyramid F... t To the end of the pyramid F out This method is used to diffuse abstract features to concrete features. It halves the H×W size through pooling, adds it to more specific features, and then further fuses them using a 3×3 convolution. This achieves cross-fusion of multi-scale features and unifies the feature scale of the final pyramid using pooling layers. Adding elements of the same size together yields the feature map X′. i ; A3) For the fused feature map, a reward-penalty loss function is used, which not only rewards the movement of sample features towards their class center but also penalizes the movement of sample features towards non-class centers. M class centers C = {C'} are randomly selected from a C'-dimensional Gaussian distribution. i Let |i=1,M…,M} be set as learnable parameters, and the loss function of the reward and punishment mechanism be: In the formula Let C' represent the loss function of the reward and punishment mechanism. i The number of channels, ε is 10. -6 λ is used to avoid the denominator being zero, and λ is used to control the weights of the two parts. B) Construct a separate multi-head attention module to focus on different local locations in the feature map, including: B1) The split multi-head attention module contains multiple parallel attention heads in the horizontal direction. Each attention head consists of a split structure of spatial attention units and channel attention units. The input of each attention head is the aggregated feature map output. B2) Spatial attention units use Inception and residual structures. Each spatial attention unit contains V Inception-residual blocks. Each block consists of three layers of convolutional operations with 1×3, 3×3, and 3×1 kernels, a ReLU activation function, and a residual structure, where V is defined as a hyperparameter. The Inception structure is used to make the model focus on different feature map scales. The role of the spatial attention unit is: In the formula, W represents the spatial attention unit. sa express The parameters, This represents the output set of the multi-head spatial attention unit. This represents the output of the j-th attention head; B3) The channel attention unit contains W linear residual blocks, each block consisting of a linear transform layer, a ReLU activation function layer, and a residual structure. The function of the channel attention unit is: In the formula, W represents the channel attention unit. ca express The parameters, This represents the output set of a multi-channel attention unit. This represents the output of the j-th attention head; B4) To overcome the problem that simply stacking H attention heads horizontally in a multi-head attention module does not guarantee that each head focuses on different positions of the feature map, a dispersion loss function is used to distribute the attention of multiple attention heads. for: In the formula, Let represent the variance of the j-th channel at point (k, l) in the feature map of the i-th sample. C) Establish an attention consistency balancing framework based on contrastive learning. First, during image preprocessing, the original image is horizontally flipped to obtain a flipped image. Then, the original image and the flipped image are sequentially input into a multi-scale feature cross-fusion module based on a reward-penalty mechanism and a multi-head attention module. Finally, the two feature maps are used to calculate the contrastive loss using a contrastive loss function. This framework is used to add constraints to the feature learning process and improve the consistency of the images before and after flipping, including: C1) First, the original image is augmented to obtain a flipped image. The image pair consisting of the original image and the flipped image is simultaneously input into an encoder consisting of a multi-scale feature cross-fusion module based on a reward and punishment mechanism and a multi-head attention module. C2) Then, the contrastive loss is calculated on the two feature maps obtained. The contrastive loss function is... for: In the formula, Defined as the features output from the spatial attention unit of the original image. This represents the features output from the spatial attention unit that indicate the image being flipped; "flip" stands for flip. and The pixel positions correspond.

2. The facial expression recognition method based on multi-scale feature cross-fusion and contrast separation multi-head attention as described in claim 1, characterized in that... Step S4 includes: Step S4.1: Load the obtained ResNet-50 pre-trained model parameters into the ResNet-50 backbone network of the multi-scale feature cross-fusion and contrast separation multi-head attention network; Step S4.2: For the preprocessed training set D = {(X i ,y i First, the training set samples (i = 1, ..., N) are horizontally flipped to obtain D. ori and D filp Secondly, it is input into a multi-scale feature cross-fusion and contrastive separation multi-head attention network loaded with ResNet-50 pre-trained model parameters, to obtain the feature X″ output by the multi-scale feature cross-fusion module based on a reward and punishment mechanism. i Calculate rewards, punishments, and losses; Step S4.3: The feature X″ output by the multi-scale feature cross-fusion module based on the reward and punishment mechanism is... i The input is separated into multiple attention modules, resulting in the outputs of multiple attention heads. Calculate the dispersion loss; Step S4.4: Analyze the features output by the spatial attention units of the two augmented datasets. and Calculate the contrast loss; Step S4.5: Then calculate the probability output of these two samples belonging to each category, i.e., out. ori and out flip The error between the model output and the label is calculated based on the cross-entropy loss function. for: In the formula, Let p(y=j|x) represent the cross-entropy loss function. i () represents the probability value calculated using the softmax function. Step S4.6: Finally, combining the above four losses, the final training loss function is obtained as follows: In the formula, λ1, λ2 and λ3 are respectively and The weights are used for backpropagation to update the parameters of the multi-scale feature cross-fusion module and the multi-head attention module for 500 iterations. Step S5 includes: Step S5.1: Place the test sample x test The input is fed into a pre-trained multi-scale feature cross-fusion and contrastive separation multi-head attention network to obtain the probability output of the sample belonging to each category, i.e., out. test out test The category with the highest probability is the category of the test sample; Step S5.2: The classification performance of the facial expression recognition method with multi-scale feature cross-fusion and contrast separation multi-head attention is described using the recognition accuracy Acc, which is calculated as follows: In the formula N true N represents the number of correctly classified items. total This represents the total number of samples.

3. The facial expression recognition method based on multi-scale feature cross-fusion and contrast separation multi-head attention as described in claim 2, characterized in that: Step S2 includes: S21) First, define the training set D = {(X i ,y i The training set image size is adjusted to 224×224 pixels if |i=1,…,N}. S22) Then, three data augmentation methods are used: random rotation, random cropping, and random erasure. For random rotation, the image is rotated around the center at a random degree with a probability of 0.2 between (-20, 20). For random cropping, the image is first filled with 32 pixels of 0 on all sides, and then 224×224 pixels are randomly cropped from 288×288 pixels at a probability of 0.

2. For random erasure, the image is cropped with a square area between (0.02, 0.25) by selecting the occlusion ratio, and then filled with 0.

4. The facial expression recognition method based on multi-scale feature cross-fusion and contrast separation multi-head attention according to claim 2, characterized in that: The parameters of the multi-scale feature cross-fusion and contrast separation multi-head attention network were updated iteratively using error backpropagation and stochastic gradient descent. The training data batch size was set to 256, the model learning rate was set to 0.0001, the Adam optimizer was used to optimize the model parameters, and the cross-entropy loss function was used as the loss function. The model parameters were saved after 500 iterations of training.

5. A computer-readable storage medium storing a computer-executable program that enables a processor to perform the method according to any one of claims 1-4.

Citation Information

Patent Citations

  • Expression recognition method for partially shielded face based on multi-scale attention mechanism

    CN115457641A