Facial emotion recognition method and system based on generative adversarial separation and expression exchange

By adopting the generation framework of visual transformers and generative adversarial networks in facial expression recognition technology, the problems of low recognition accuracy of facial expression recognition in real environments and high cost of model pre-training in the prior art are solved, and higher quality expression generation and better generalization capabilities are achieved.

CN120148085APending Publication Date: 2025-06-13HEFEI UNIV OF TECH
View PDF 0 Cites 3 Cited by

Patent Information

Application Number
CN202510250900.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-04
Publication Date
2025-06-13

AI Technical Summary

Technical Problem

The existing facial expression recognition technology has low recognition accuracy in real environments. Due to traditional manual features or shallow learning, it is difficult to adapt to different angles and spontaneous facial images. The large model leads to high pre-training costs, and training of unpaired data sets is challenging.

Method used

Using a generation framework based on vision transformer (ViT) and generative adversarial network (GAN), a self-supervised masked autoencoder pre-trained model is input, and two unpaired face expression images are generated and exchanged face features are generated and exchanged, facial key points features are fused, and model parameters are optimized using generators and discriminators.

Benefits of technology

It improves the diversity and realism of generated images, expands the scale of the data set, reduces training costs, enhances the generalization ability of the model, can better process new data, and maintain identity consistency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120148085A_ABST
    Figure CN120148085A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of medical electronic equipment, and discloses a face emotion recognition method and system based on generative adversarial separation and expression exchange. The training process of the adopted facial expression recognition model comprises the following steps: pre-training a visual converter based on a self-supervised mask auto-encoder; extracting face feature vectors corresponding to the two face expression images through an image encoder in a visual converter; performing feature decoupling on the face feature vector to separate emotion-related features and emotion-unrelated features, and performing feature exchange to generate exchanged face features; fusing the face key point features with the exchange face features to generate fused features; generating a reconstructed facial expression image and a synthesized exchange facial expression image by using a generator based on the fusion features, and constraining identity consistency of the reconstructed image and the original facial expression image; a discriminator is used for discriminating the authenticity of a reconstructed image, and meanwhile, a classifier is used for performing facial emotion classification on the basis of emotion related features, so that parameters of a facial expression recognition model are optimized. The core problems of poor model generalization, low generation quality and difficulty in non-paired data training in the prior art are effectively solved, and the training cost is reduced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of medical electronic devices, and particularly relates to a facial emotion recognition method and system based on generating adversarial separation and exchanging expressions. Background Art

[0002] The core of the facial expression recognition (FER) task is to classify human facial expressions and map complex facial expressions to basic emotion categories, such as happiness, sadness, anger, fear, disgust, and surprise. With the increasing depth of human-computer interaction, facial expression recognition has shown broad application prospects in fields such as social robots, medical diagnosis, and driving safety. Facial expression recognition technology encompasses key technologies and methods, such as feature extraction, datasets and annotations, and classification algorithms. Traditional facial expression recognition (FER) methods are mainly based on handcrafted features or shallow learning, resulting in limited recognition accuracy and weak adaptability. These methods show certain performance on datasets in a controlled laboratory environment. However, in reality, facial images often come from real environments, are taken from different angles, and are spontaneous. Factors existing in the real environment, such as occlusion, background illumination, and head pose changes, will all affect the performance of facial expression recognition.

[0003] The rise of deep learning has brought revolutionary changes to facial expression recognition (FER). Convolutional neural networks (CNNs), with their powerful local feature extraction capabilities, have become the mainstream method in the field of facial expression recognition. In recent years, vision transformers (ViTs) have performed well in global feature modeling through self-attention mechanisms, providing new ideas for facial expression recognition. In addition, with the rapid development of generative adversarial networks (GANs), many innovative frameworks based on GANs have emerged. By generating high-quality facial expression data, they help the model learn decoupled facial expression features. Despite some progress, the accuracy of facial expression recognition in the prior art still has room for improvement. The model is relatively large, resulting in high pre-training costs. At the same time, there are still some negative impacts on personal identity. It is challenging to train the model on unpaired real-world datasets and difficult to generate high-quality facial expressions for unknown subjects. Summary of the Invention

[0004] To solve the above technical problems, the present invention provides a facial emotion recognition method and system based on generating adversarial separation and exchanging expressions.

[0005] To solve the above technical problems, the present invention adopts the following technical solutions:

[0006] A facial emotion recognition method based on generating adversarial separation and exchanging expressions, wherein the training process of the facial expression recognition model adopted includes:

[0007] Pre-train the vision transformer based on self-supervised masked autoencoders;

[0008] Construct a facial expression recognition model based on a vision transformer and a generative adversarial network. Input two unpaired facial expression images, and extract the corresponding facial feature vectors of the two facial expression images through the image encoder in the vision transformer;

[0009] By decoupling the features of the facial feature vectors, separate the emotion-related features and emotion-unrelated features, and perform feature exchange to generate exchanged facial features;

[0010] Fuse the facial key-point features of the facial expression image with the exchanged facial features to generate fused features;

[0011] Based on the fused features, use the generator to generate reconstructed facial expression images and synthesized exchanged facial expression images, and constrain the identity consistency between the reconstructed images and the original facial expression images;

[0012] Use the discriminator to perform expression classification and authenticity discrimination on the reconstructed images, perform facial emotion classification based on the emotion-related features through the classifier, and optimize the parameters of the facial expression recognition model by combining the discriminator loss, reconstruction loss, orthogonal constraint, and identity constraint loss.

[0013] In one embodiment, the step of inputting two unpaired facial expression images and extracting the corresponding facial feature vectors of the two facial expression images through the image encoder in the vision transformer specifically includes:

[0014] Take the two facial expression images as the input of the facial expression recognition model, then use the image encoder to divide the input facial expression images into N image patches, map each image patch to a vector through a linear transformation, and then map it to a facial feature vector containing all the information of the facial expression image after adding a trainable class token and position encoding.

[0015] In one embodiment, the step of decoupling the features of the facial feature vectors to separate the emotion-related features and emotion-unrelated features specifically includes:

[0016] Input the facial feature vectors corresponding to the two facial expression images into a two-layer feature decoupler for expression separation; in the two-layer feature decoupler, add a dense block on the basis of the weight learning module, decouple the emotion-related features through the decoupling block, and then use the residual connection to obtain the emotion-unrelated features.

[0017] In one embodiment, the face feature vectors corresponding to the two face expression images are input into a double-layer feature decoupler for expression separation; in the double-layer feature decoupler, a dense block is added on the basis of the weight learning module, and the emotion-related features are decoupled through the decoupling block, and then the emotion-unrelated features are obtained by using residual connection, which specifically includes:

[0018] The weight learning module extracts global information through global context attention and performs interaction of multiple tokens. First, the attention weight g of each token is calculated through 1D convolution and softmax:

[0019] The face feature vector is weighted and summed using the attention weight to obtain the weighted feature z g :

[0020] For z g A channel transformation is performed to obtain the transformed feature z g_t ; The element-wise addition is performed using the broadcast mechanism to obtain the feature z se ;

[0021] Then the feature is further enhanced through the dense block, and the output of each layer is continuously connected with the previous features to obtain the enhanced feature z dense ;

[0022] Then the enhanced feature z dense is subjected to feature extraction through the decoupling block to obtain the emotion-related feature v e :

[0023] The orthogonal residual is calculated, that is, the emotion-unrelated feature v id is obtained: v id = z - v e .

[0024] In one embodiment, the feature exchange to generate the exchanged face features specifically includes:

[0025] The face feature vectors z a , z b corresponding to the two face expression images are respectively input into the same feature decoupler FD, and the vector pairs are respectively the emotion-related feature and the emotion-unrelated feature corresponding to z a , are respectively the emotion-related feature and the emotion-unrelated feature corresponding to z b ;

[0026] A feature exchange operation is performed on the two vector pairs: the emotion-related features in each vector pair are exchanged to obtain two new vector pairs The two vectors in each new vector pair are added to obtain the face feature vectors z a , zb Corresponding swapped face features

[0027] In one embodiment, the fusion of the facial key-point features of the facial expression image with the swapped face features to generate fused features specifically includes:

[0028] Through a pre-trained facial key-point detection network, for the two input facial expression images x a , x b , respectively extract the facial key-point features For the facial key-point features Perform image block division and linear transformation, and map them into vectors f a , f b , and fuse the swapped face features with the vector f b , and the swapped face features with the vector f a respectively to generate fused features; the swapped face features respectively represent the swapped face features corresponding to the two facial expression images x a , x b .

[0029] In one embodiment, based on the fused features and using a generator to generate a reconstructed facial expression image and a synthesized swapped facial expression image, and constraining the identity consistency between the reconstructed image and the original facial expression image specifically includes:

[0030] After receiving the fused features, as well as the emotion-related features and emotion-unrelated features corresponding to the two facial expression images, the generator adds mask vectors of the same shape, and based on the generator and using the cross-attention mechanism, generates a reconstructed facial expression image and a synthesized swapped facial expression image; the synthesized swapped facial expression image contains the identity of the original facial expression image and the expression of another facial expression image;

[0031] Taking the original facial expression image, as well as the corresponding reconstructed facial expression image and synthesized swapped facial expression image as a group of image pairs, input them into the arcface model to extract feature vectors, and calculate the cosine similarity between the feature vectors to keep the identity of each group of image pairs consistent.

[0032] In one embodiment, using a discriminator to perform expression classification and authenticity discrimination on the reconstructed image specifically includes:

[0033] By adding an auxiliary classifier in the discriminator to classify the expression of the reconstructed image and judge the authenticity of the reconstructed image, to compete with the generator that generates the reconstructed image;

[0034] The facial emotion classification by the classifier based on emotion-related features specifically includes:

[0035] Input the emotion-related features into a multi-layer perceptron, perform a mean operation on the sequence, then add the class label representing the emotion, and after being processed by multiple fully connected layers, output the predicted facial expression category.

[0036] In one embodiment, optimizing the parameters of the facial expression recognition model by combining discriminative loss, reconstruction loss, orthogonal constraint, identity constraint loss, and emotion classification loss specifically includes:

[0037] After extracting the input image features through a discriminator for auxiliary classification, calculate the discriminative loss using the BCEWithLogits loss;

[0038] Obtain the reconstruction loss by calculating the pixel-level difference between the original facial expression image and the reconstructed facial expression image;

[0039] Calculate the orthogonal constraint through the inner product of emotion-related features and emotion-unrelated features;

[0040] Generate the identity constraint loss when constraining the identity consistency between the reconstructed image and the original facial expression image;

[0041] Calculate the emotion classification loss through the logical values of the true label categories corresponding to the training samples and the logical values of each category of the training samples.

[0042] A facial emotion recognition system based on generating adversarial separation and swapping expressions includes:

[0043] A pre-training module that pre-trains a vision transformer based on a self-supervised masked autoencoder;

[0044] A face feature vector extraction module that constructs a facial expression recognition model based on a vision transformer and a generative adversarial network, inputs two unpaired facial expression images, and extracts the face feature vectors corresponding to the two facial expression images through the image encoder in the vision transformer;

[0045] An expression separation and swapping module that decouples the features of the face feature vectors, separates emotion-related features and emotion-unrelated features, and performs feature swapping to generate swapped face features;

[0046] A modality fusion module that fuses the facial key point features of the facial expression image with the swapped face features to generate fused features;

[0047] A reconstruction supervision module that generates a reconstructed facial expression image and a synthesized swapped facial expression image based on the fused features using a generator, and constrains the identity consistency between the reconstructed image and the original facial expression image;

[0048] The training module uses a discriminator to classify the expressions of the reconstructed images and determine their authenticity, and uses a classifier to classify facial emotions based on emotion-related features, and optimizes the parameters of the facial expression recognition model by combining the discriminative loss, reconstruction loss, orthogonal constraint, and identity constraint loss.

[0049] Compared with the prior art, the beneficial technical effects of the present invention are as follows:

[0050] 1. The present invention proposes a generation framework based on a Vision Transformer (ViT) and a Generative Adversarial Network (GAN). By using the auxiliary task of separating and swapping expressions, by inputting two facial expression images, it is possible to generate high-quality facial expression images with identity-invariant expression swapping. This innovative method not only improves the diversity of the generated images, but also actually expands the scale of the dataset on the basis of the original dataset, provides more training and test samples for subsequent research, and reduces the training cost.

[0051] 2. The present invention retrains the model with the expressions of different identities generated by expression swapping, further expanding the number of the dataset, enabling the model to better learn and understand the data, reducing the overfitting of the model to a specific dataset, making it more generalizable, and being able to better process new and unseen data.

[0052] 3. The adversarial task of facial expression swapping proposed by the present invention can perform good emotional entanglement and learn more facial details, so that the generated images have higher diversity and realism, and the model can better understand the fine-grained features of facial expressions.

[0053] 4. The information of facial key points (landmarks) introduced by the model of the present invention can guide the model to focus on emotion-related regions, and at the same time add a supervision signal of identity ID testing to the generated images, which helps to keep the identity unchanged. The two pulling each other can help the Generative Adversarial Network (GAN) generate more real and high-quality images. BRIEF DESCRIPTION OF THE DRAWINGS

[0054] Figure 1 It is a schematic flowchart of the method in the embodiment of the present invention;

[0055] Figure 2 It is a schematic framework diagram of the facial expression recognition model in the embodiment of the present invention;

[0056] Figure 3 It is a schematic diagram of the expression separation and swapping module in the embodiment of the present invention;

[0057] Figure 4 It is a schematic diagram of the feature decoupler in the embodiment of the present invention;

[0058] Figure 5 Schematic diagram of the facial key-point modality fusion module in the embodiment of the present invention;

[0059] Figure 6 Schematic diagram of the identity supervision module in the embodiment of the present invention. Detailed implementation manners

[0060] A preferred implementation manner of the present invention will be described in detail below with reference to the accompanying drawings.

[0061] In the present invention, a general vision transformer (ViT) model with powerful facial expression representation is pre-trained based on a self-supervised manner using a facial expression image dataset. Then, by adding an auxiliary task of generating adversarial separation and swapping expressions, a landmark modality fusion module (LMF module) is used to provide more expression-related details, and an identity supervision module (IS module) is used to keep the identity unchanged. Through a generative adversarial network (GAN) framework, high-quality facial expression images after swapping expressions are generated to achieve better facial expression recognition. At the same time, the facial expression recognition datasets used in the present invention all come from real environments, and all facial expression images are unpaired.

[0062] As Figure 2 shown, in a vision transformer (ViT)-generative adversarial network (GAN) framework, in order to obtain higher-precision facial expression recognition, the present invention uses an expression separation and swapping module (ESE module) to separate pure expression features and remove the influence of other factors as much as possible. At the same time, the adversarial generation of the generative adversarial network (GAN) framework is used to promote facial expression recognition. In order to guide the model to pay more attention to expression-related regions, the present invention adds a landmark modality fusion module to provide more expression-related details and better perform expression swapping. At the same time, in order to maintain the identity consistency of the facial expression images in the model, the present invention introduces an identity supervision module, which keeps the identity unchanged during the expression swapping or reconstruction of the original image of the facial expression image, prompting the model to generate higher-quality images. Finally, the facial expression recognition model ESEVG of the present invention is obtained.

[0063] As Figure 1 shown, a facial emotion recognition method based on generating adversarial separation and swapping expressions, and the training process of the facial expression recognition model adopted includes the following steps:

[0064] S1, pre-training the vision transformer based on a self-supervised masked autoencoder;

[0065] S2, constructing a facial expression recognition model based on the vision transformer and the generative adversarial network, inputting two unpaired facial expression images, and extracting the corresponding facial feature vectors of the two facial expression images through the image encoder in the vision transformer;

[0066] S3. By performing feature decoupling on the face feature vector, separating the emotion-related features and emotion-unrelated features, and performing feature exchange to generate an exchanged face feature;

[0067] S4. Fusing the face key-point features with the exchanged face feature to generate a fused feature;

[0068] S5. Based on the fused feature and using a generator to generate a reconstructed face expression image and a synthesized exchanged face expression image, and constraining the identity consistency between the reconstructed image and the original face expression image;

[0069] S6. Using a discriminator to perform expression classification and authenticity discrimination on the reconstructed image, performing facial emotion classification based on the emotion-related features through a classifier, and optimizing the parameters of the facial expression recognition model by combining the discriminator loss, reconstruction loss, orthogonal constraint, identity constraint loss, and emotion classification loss.

[0070] Based on self-supervised Masked Autoencoder (MAE), this invention pre-trains the Vision Transformer (ViT), constructs a ViT-GAN joint framework by combining the Generative Adversarial Network (GAN), uses unpaired face expression images as input, separates the emotion-related features and identity features (emotion-unrelated features) through the Emotion Separation and Exchange module (ESE), introduces the Landmark Modal Fusion (LMF) to enhance the expression details, combines the Identity Supervision module (IS) to constrain the identity consistency of the generated images, and simultaneously through adversarial training, orthogonal constraint, and multi-task optimization, achieves high-precision facial emotion recognition, low pre-training cost (reducing annotation dependence and computing resources), and high-quality expression generation (identity unchanged, details realistic), effectively solving the core problems of poor model generalization, low generation quality, and difficulty in training with unpaired data in the prior art.

[0071] In one embodiment, step S1 specifically includes: pre-training a vanilla Vision Transformer (ViT) using a self-supervised Masked Autoencoder (MAE) method. Here, the present invention initializes the Vision Transformer with the weights of the ImageNet dataset, and then uses the existing largest facial expression recognition dataset AffectNet as the pre-training dataset. The ImageNet 1K dataset contains approximately 1000 facial expression categories, with each category containing about 1000 images, totaling millions of images. AffectNet has approximately 400,000 labeled facial expression images. After initialization with the weights of the ImageNet 1K dataset, the pre-trained model has learned general features such as textures and shapes. Then, by pre-training with the AffectNet dataset, the Vision Transformer can learn more specific expression-related features to adapt to the facial expression recognition task. Based on Masked Autoencoder (MAE) pre-training, the image encoder in the Vision Transformer learns local and global facial structures by reconstructing the missing pixels of randomly masked image patches, enabling a stronger feature representation of facial expression images.

[0072] In one embodiment, for the input of two unpaired facial expression images in step S2, extracting the corresponding facial feature vectors of the two facial expression images through the image encoder in the Vision Transformer specifically includes:

[0073] Taking the two facial expression images as the input of the facial expression recognition model, and then using the image encoder to divide the input facial expression images into N image patches, and mapping each image patch to a vector through a linear transformation. After adding a trainable class token and position encoding, it is mapped to a facial feature vector that contains all the information of the facial expression images.

[0074] Based on the pre-training, the present invention combines the Vision Transformer (ViT) and the Generative Adversarial Network (GAN) framework to obtain a new network structure, and at the same time separates pure expression features through an auxiliary task of expression swapping to complete a specific image-to-image conversion task. Through this auxiliary task of separating and swapping expressions based on generative adversarial, it is possible to highly purify the emotion-related features and emotion-unrelated features, enabling the facial expression recognition model to capture finer-grained features of facial muscle movements, thereby better understanding facial expressions and achieving better expression recognition.

[0075] First, two facial expression images are used as the input of the facial expression recognition model. Here, each facial expression image x belongs to one of K basic facial expression categories, and W, H, and C represent width, height, and number of channels respectively. Then, the image encoder E in the Vision Transformer is used to process the input facial expression image x ∈ {x a , x bIt is divided into N small blocks of p×p, and each block is mapped to a vector of a fixed length through a linear transformation After adding trainable [class] tokens and positional encoding, it is mapped to a high-dimensional face feature vector z∈{z a ,z b}, which contains all the information of the facial expression image. Here N + 1 represents the number of tokens, and L represents the feature dimension. Different from the image encoder E in pre-training, there is no masking operation on the input facial expression image here. The input of the image encoder is all the facial feature tokens, and the same image encoder E is used for the two input facial expression images.

[0076] In one embodiment, in step S3, by performing feature decoupling on the face feature vector, the emotion-related features and emotion-unrelated features are separated, specifically including:

[0077] The face feature vectors corresponding to the two facial expression images are input into a two-layer feature decoupler for expression separation; in the two-layer feature decoupler, a dense block is added on the basis of the weight learning module, the emotion-related features are decoupled through the decoupling block, and then the emotion-unrelated features are obtained by using residual connection.

[0078] In one embodiment, the step of inputting the face feature vectors corresponding to the two facial expression images into a two-layer feature decoupler for expression separation; in the two-layer feature decoupler, a dense block is added on the basis of the weight learning module, the emotion-related features are decoupled through the decoupling block, and then the emotion-unrelated features are obtained by using residual connection, specifically including:

[0079] The weight learning module extracts global information through global context attention and performs the interaction of multiple tokens. First, the attention weight g of each token is calculated through 1D convolution and softmax:

[0080] The weighted sum of the face feature vectors is obtained using the attention weights to get the weighted feature z g :

[0081] For z g A channel transformation is performed to obtain the transformed feature z g_t ; The element-wise addition is performed using the broadcast mechanism to obtain the feature z se ;

[0082] Then the feature is further enhanced through the dense block, and the output of each layer is continuously connected with the previous feature to obtain the enhanced feature z dense ;

[0083] Then the enhanced feature z denseFeature extraction is performed to obtain emotion-related feature v e :

[0084] Calculate the orthogonal residual, that is, obtain the emotion-unrelated feature v id : v id = z - v e .

[0085] As Figure 3 shown, specifically, after obtaining the feature vector z ∈ {z a , z b}, the separation of expressions and the exchange of features will be performed in the Expression Separation and Exchange Module (ESE). For the separation of features, as Figure 4 shown, the present invention uses a double-layer feature decoupler FD. In FD, first, a DenseBlock is added on the basis of the Weight Learning Module (WL), and finally, the emotion-related feature and the residual are decoupled through the Decoupling Module (DM). Specifically, WL extracts global information through global context attention, performs the interaction of multiple tokens, and first calculates the attention weight of each token through 1D convolution and softmax

[0086] g = softmax(conv(z i )); (1)

[0087] Next, use the attention weight to perform weighted summation on the feature to obtain

[0088]

[0089] Perform channel transformation on z g :

[0090]

[0091] Among them, FC represents the fully connected layer.

[0092] Use the broadcast mechanism to perform element-wise addition to obtain the feature

[0093] v se = z + z g_t ; (4)

[0094] Then further enhance the feature through DenseBlock, continuously connect the output of each layer with the previous feature, and obtain the enhanced feature z dense :

[0095] z dense = DenseBlock(z se ) =

[0096] [z se , DenseLayer 1 (z se ), DenseLayer 2 ([z se , DenseLayer 1 (z se )]),... ; (5)

[0097] Then, the enhanced feature z dense is subjected to feature extraction by DM to obtain the sentiment vector

[0098] v e = FC2(Swish(FC1(z dense ))) ; (6)

[0099] The orthogonal residual is calculated

[0100] v id = z - v e . (7)

[0101] To ensure that the sentiment feature and the residual are orthogonal, the present invention realizes the orthogonality constraint by calculating the inner product between them. The present invention hopes that the inner product between the sentiment feature and the residual is close to zero, and the orthogonality constraint is described as follows:

[0102]

[0103] Here z a , z b Using the same feature decoupler FD, we can obtain two vector pairs.

[0104] In one embodiment, the feature swapping in step S3 to generate the swapped face feature specifically includes:

[0105] The face feature vectors z a , z b corresponding to two face expression images are respectively input into the same feature decoupler FD, and the vector pairs are respectively the sentiment-related feature and the sentiment-unrelated feature corresponding to z a , are respectively the sentiment-related feature and the sentiment-unrelated feature corresponding to z b ;

[0106] Perform a feature swapping operation on the two vector pairs: swap the sentiment-related features in each vector pair to obtain two new vector pairs Add the two vectors in each new vector pair to obtain the face feature vector z a ,z b The corresponding swapped face feature

[0107] After decoupling the features to separate the emotional features, perform a feature swapping operation on the two vector pairs. Specifically, swap the emotion-related features in each pair to obtain two new vector pairs Only need to add the two vectors in each vector pair to obtain the swapped face feature Here

[0108]

[0109] In one embodiment, the fusion of the face key point feature and the swapped face feature in step S4 to generate a fusion feature specifically includes:

[0110] Through a pre-trained face key point detection network for the two input face expression images x a ,x b , respectively extract the face key point features For the face key point features Perform image block division and linear transformation, and map them into vectors f a ,f b , and fuse the swapped face feature and f b , the swapped face feature and f a respectively perform feature fusion to generate a fusion feature; the swapped face features respectively represent the swapped face features corresponding to the two face expression images.

[0111] As Figure 5 shown, in order to make the swapped expression more natural and realistic, the present invention adds face key point (landmark) information and adds a landmark modality fusion module (LMF), which can guide the model to pay more attention to the expression-related areas and help generate high-quality swapped face expression images. First, through an existing face key point (landmark) detection network, extract the face key point (landmark) features from the two input original face expression images {x a ,x b}, Perform operations such as image block division and linear transformation on the feature , and map it into a vector f∈{f a ,f b}, and then the LMF receives the swapped face feature vector z * ​and the face key point (landmark) feature vector f, where The present invention will and f b , and f a are respectively subjected to feature fusion, aiming to provide more information about relevant expressions for the swapped image through the face key points (landmarks), and guiding the change of the expression of the swapped image.

[0112] For and f b when performing feature fusion, for all tokens except the [class] tag and f b , LMF first performs dimensional transformation and then projects to a low-dimensional latent embedding using a 1×1 convolutional layer to obtain the feature maps P z , where H is equal to W, and their product is N. Then, Softmax is used to weight the feature map for P z to highlight the features of certain regions. Finally, P f is added to the weighted feature map to obtain the feature map

[0113] P e = P f + Softmax(P z ) ⊙ P z ; (10)

[0114] where ⊙ represents element-wise multiplication, and then the fused feature is obtained through convolution and dimensional transformation which is briefly described as follows:

[0115] P a = Reshape(conv(P e )); (11)

[0116] Adding the [class] token, and through self-attention operation, information interaction is performed between the [class] token and other feature tokens to obtain the fused feature

[0117]

[0118] At this time contains the emotion-irrelevant features of x a and the emotion-related features of the swapped x b , and under the condition that the personal identity remains unchanged, it is shown as x a having an expression category. Similarly, for b and f and fa , the fused features can also be obtained through the above process

[0119] In one embodiment, in step S5, based on the fused features and using a generator to generate a reconstructed facial expression image and a synthesized swapped facial expression image, and constraining the identity consistency between the reconstructed image and the original facial expression image, specifically including:

[0120] After receiving the fused features, as well as the emotion-related features and emotion-unrelated features corresponding to two facial expression images, the generator adds mask vectors of the same shape, and based on the generator and using the cross-attention mechanism, generates a reconstructed facial expression image and a synthesized swapped facial expression image; the synthesized swapped facial expression image contains the identity of the original facial expression image and the expression of another facial expression image;

[0121] Taking the original facial expression image, as well as the corresponding reconstructed facial expression image and synthesized swapped facial expression image as a group of image pairs, input them into the arcface model to extract feature vectors, and calculate the cosine similarity between the feature vectors to keep the identity of each group of image pairs consistent.

[0122] Specifically, as Figure 6 shown, after completing the expression swapping, in order to supervise that the identity of the generated facial expression image remains unchanged, the present invention adds an identity supervision module (IS) to perform identity constraint on the generated swapped image. First, the generator G receives the fused features from the LMF and the decoupled from the separator S. At the same time, in order to keep consistent with the input of the upstream pre-trained decoder decoder, the present invention adds mask vectors of the same shape, and uses the cross-attention mechanism to obtain a reconstructed facial expression image through the generator G and a synthesized swapped facial expression image Here, the weights of the generator G are the decoder pre-trained by the masked autoencoder (MAE), and the synthesized swapped facial expression image contains the a original identity and the b expression of x. Similarly, contains the b original identity and the a expression of x. Then, two groups of image pairs and are input into the arcface network model to perform identity constraint on the original image and the generated image, extract the feature vectors of these images through the model Then calculate the cosine similarity between them to keep the identity of each group of image pairs consistent. For The identity constraint loss L can be obtained id_a :

[0123]

[0124] Similarly, for the identity constraint loss L can be obtained id_b .

[0125] Since the images in the dataset under real - world conditions are all unpaired, in order to generate high - quality facial expression images, the L2 loss is used to calculate the pixel - level difference between the original image and the reconstructed image, and the reconstruction loss L rec :

[0126]

[0127] Meanwhile, for the original image and the image synthesized by swapping expressions, the pixel difference in the non - expression component region is small, and the L1 loss is used for constraint:

[0128]

[0129] In one embodiment, in step S6, the discriminator is used to classify the expression of the reconstructed image and judge its authenticity, and the classifier is used to perform facial emotion classification based on emotion - related features, specifically including:

[0130] The expression of the reconstructed image is classified by adding an auxiliary classifier in the discriminator, and the authenticity of the reconstructed image is judged to compete with the generator that generates the reconstructed image; the emotion - related features are input into a multi - layer perceptron, and a mean operation is performed on the sequence, and then added to the class label representing the emotion. After being processed by multiple fully - connected layers, the predicted facial expression category is output.

[0131] To obtain higher - quality generated images, the discriminator D is used to judge whether the generated image is real or fake to compete with the generator, and at the same time an auxiliary classifier is added in the discriminator to classify the expression of the image. Here, the discriminator weight is the decoder decoder pre - trained by the masked auto - encoder (MAE), and the auxiliary classifier is a multi - layer perceptron (MLP). By processing the input image the feature z d is obtained, and only the feature z d z with [class] cls is finally output through the classifier for classification

[0132]

[0133] Among them the first K units are used for class emotion discrimination, and the last unit is responsible for distinguishing real and fake images.

[0134] Calculate the loss in the multi-label classification task through the BCEWithLogits loss. The specific discrimination loss is described as follows:

[0135]

[0136] where y i is the true value of the i-th label. A true value of 1 indicates the presence of the label, and a true value of 0 indicates the absence of the label. It represents the true label of x. At the same time, if y logit is y, it means the input image is a real image, y k+1 is 1, otherwise it is 0.

[0137] Specifically, the classifier C is a multi-layer perceptron (MLP), and it predicts the facial expression category according to the emotion-related features obtained by the feature decoupler FD During the training process, the emotion feature representation is input into the classifier, and a mean operation is performed on the sequence, and then added to the [class] token v representing the emotion cls After being processed by multiple fully connected layers, the predicted emotion category is output Ensure that v e contains potential emotion factors:

[0138] y e = MLP(v cls + mean(v e ))). (18)

[0139] In one embodiment, optimizing the parameters of the facial expression recognition model by combining the discrimination loss, reconstruction loss, orthogonal constraint, identity constraint loss, and emotion classification loss specifically includes:

[0140] After extracting the input image features through the discriminator, perform auxiliary classification, and calculate the discrimination loss using the BCEWithLogits loss;

[0141] Obtain the reconstruction loss by calculating the pixel-level difference between the original facial expression image and the reconstructed facial expression image;

[0142] Calculate the orthogonal constraint through the inner product of the emotion-related features and the emotion-unrelated features;

[0143] Generate the identity constraint loss when constraining the identity consistency between the reconstructed image and the original facial expression image;

[0144] Calculate the emotion classification loss through the logical values of the true label categories corresponding to the training samples and the logical values of each category of the training samples.

[0145] The identity constraint loss, orthogonal constraint, reconstruction loss, and discriminative loss have been introduced in detail above.

[0146] The present invention uses cross-entropy loss to supervise the K-class sentiment classification task, and the sentiment classification loss L emo is as follows:

[0147]

[0148] where B is the batch size, and y e [i, label[i]] is the logit value of the true label class corresponding to the i-th training sample, and y e [i, j] is the logit value of the j-th class of the i-th training sample.

[0149] In addition, during the testing process, the present invention only needs the encoder E, the feature decoupler FD, and the classifier C to well complete the facial expression recognition task.

[0150] It should be understood that although the steps in the flowchart of the accompanying drawings of the specification are shown in sequence according to the indication of the arrows, these steps are not necessarily executed in the order indicated by the arrows. Unless otherwise clearly stated in this article, there is no strict order restriction for the execution of these steps, and these steps can be executed in other orders. Moreover, at least a part of the steps in the flowchart of the accompanying drawings of the specification may include multiple steps or multiple stages. These steps or stages are not necessarily executed at the same moment, but can be executed at different moments, and the execution order of these steps or stages is not necessarily sequential, but can be executed alternately or in turn with at least a part of other steps or steps or stages in other steps.

[0151] The present disclosure also provides a facial emotion recognition system based on generative adversarial separation and exchanged expressions. The system may include a system (including a distributed system), software (application), module, component, server, client, etc. that uses the method described in the embodiments of this specification and combines the necessary implementation hardware. Based on the same innovative concept, the systems in one or more embodiments provided by the embodiments of the present disclosure are as described in the following embodiments. Since the implementation solutions for the system to solve problems are similar to the method, the implementation of the specific system in the embodiments of this specification can refer to the implementation of the foregoing method, and the repeated parts will not be elaborated. As used hereinafter, the term "module" or "modular" may be a combination of software and / or hardware that can achieve a predetermined function. Although the systems described in the following embodiments are preferably implemented in software, implementation in hardware, or a combination of software and hardware is also possible and contemplated.

[0152] The system includes:

[0153] A pre-training module that pre-trains a vision transformer based on a self-supervised masked auto-encoder;

[0154] A face feature vector extraction module that constructs a facial expression recognition model based on a vision transformer and a generative adversarial network. Input two unpaired facial expression images, and extract the face feature vectors corresponding to the two facial expression images through the image encoder in the vision transformer;

[0155] An expression separation and exchange module that decouples the features of the face feature vectors, separates the emotion-related features and emotion-unrelated features, and performs feature exchange to generate exchanged face features;

[0156] A modality fusion module that fuses the face key-point features with the exchanged face features to generate fused features;

[0157] A reconstruction supervision module that generates a reconstructed facial expression image and a synthesized exchanged facial expression image based on the fused features using a generator, and constrains the identity consistency between the reconstructed image and the original facial expression image;

[0158] A training module that uses a discriminator to discriminate the authenticity of the reconstructed image, and at the same time performs facial emotion classification based on the emotion-related features through a classifier, and optimizes the parameters of the facial expression recognition model by combining the discriminant loss, reconstruction loss, orthogonal constraint, and identity constraint loss.

[0159] The system of the present invention corresponds to the method, and the embodiments applicable to the method are equally applicable to the system.

[0160] For those skilled in the art, it is obvious that the present invention is not limited to the details of the above-described exemplary embodiments, and without departing from the spirit or basic characteristics of the present invention, the present invention can be implemented in other specific forms. Therefore, in any aspect, the embodiments should be regarded as exemplary and non-limiting. The scope of the present invention is defined by the appended claims rather than the above description. Therefore, it is intended to embrace all changes falling within the meaning and scope of the equivalent elements of the claims in the present invention, and any reference signs in the claims should not be regarded as limiting the claims involved.

[0161] In addition, it should be understood that although this specification is described according to the embodiments, not every embodiment only contains an independent technical solution. This narrative way of the specification is only for clarity. Those skilled in the art should regard the specification as a whole, and the technical solutions in each embodiment can also be appropriately combined to form other embodiments that can be understood by those skilled in the art.

Claims

1. A facial emotion recognition method based on generative adversarial separation and expression exchange, characterized in that: The training process of the adopted facial expression recognition model includes: Pre-training of visual transformers based on self-supervised mask autoencoders; Construct a facial expression recognition model based on visual transformer and generative adversarial network, input two unpaired facial expression images, and extract the facial feature vectors corresponding to the two facial expression images through the image encoder in the visual transformer; By decoupling the facial feature vector, emotion-related features and emotion-irrelevant features are separated, and feature exchange is performed to generate exchanged facial features; The facial key point features of the facial expression image are fused with the exchanged face features to generate fused features; Based on the fusion features, a generator is used to generate a reconstructed facial expression image and a synthesized exchanged facial expression image, and the identity consistency between the reconstructed image and the original facial expression image is constrained; The discriminator is used to classify the expression of the reconstructed image and judge its authenticity. The classifier is used to classify facial emotions based on emotion-related features. The parameters of the facial expression recognition model are optimized by combining the discrimination loss, reconstruction loss, orthogonal constraint and identity constraint loss.

2. The facial emotion recognition method based on generative adversarial separation and expression exchange according to claim 1, characterized in that: The input of two unpaired facial expression images and the extraction of facial feature vectors corresponding to the two facial expression images through an image encoder in a visual transformer specifically include: Two facial expression images are used as the input of the facial expression recognition model. The input facial expression images are then divided into N image blocks using an image encoder. Each image block is mapped into a vector through a linear transformation. After adding trainable category labels and position encoding, it is mapped into a facial feature vector that contains all the information of the facial expression image.

3. The facial emotion recognition method based on generative adversarial separation and expression exchange according to claim 1, characterized in that: The feature decoupling of the facial feature vector is performed to separate the emotion-related features and the emotion-irrelevant features, specifically including: The facial feature vectors corresponding to the two facial expression images are input into a two-layer feature decoupler for expression separation. In the two-layer feature decoupler, dense blocks are added on the basis of the weight learning module, and the emotion-related features are decoupled by the decoupling blocks, and then the emotion-irrelevant features are obtained by using residual connections.

4. The facial emotion recognition method based on generative adversarial separation and expression exchange according to claim 3, characterized in that: The facial feature vectors corresponding to the two facial expression images are input into a double-layer feature decoupler to separate the facial expressions; in the double-layer feature decoupler, a dense block is added on the basis of the weight learning module, the emotion-related features are decoupled by the decoupling block, and the emotion-irrelevant features are obtained by using the residual connection, which specifically includes: The weight learning module extracts global information through global context attention and interacts with multiple tags. First, the attention weight g of each tag is calculated through 1D convolution and softmax: Use the attention weight to perform weighted summation on the face feature vector to obtain the weighted feature z g : Right g Perform channel transformation to obtain the transformation feature z g_t ; Use the broadcast mechanism to perform element-by-element addition to obtain feature z se ; Then, the features are further enhanced through dense blocks, and the output of each layer is continuously connected with the previous features to obtain the enhanced features z dense ; Then the enhanced feature z is processed by the decoupling block. dense Perform feature extraction to obtain emotion-related features v e : Calculate the orthogonal residual, that is, get the emotion-independent feature v id :v id =zv e .

5. The facial emotion recognition method based on generative adversarial separation and expression exchange according to claim 1, characterized in that: The feature exchange to generate exchanged facial features specifically includes: The facial feature vector z corresponding to the two facial expression images a ,z b Input the same characteristic decoupler FD respectively, and we can get the vector pair They are z a The corresponding emotion-related features and emotion-irrelevant features, They are z b Corresponding emotion-related features and emotion-irrelevant features; Perform feature exchange operation on two vector pairs: exchange the emotion-related features in each vector pair to obtain two new vector pairs Add the two vectors in each new vector pair to get the facial feature vector z a ,z b Corresponding exchange facial features 6. The facial emotion recognition method based on generative adversarial separation and exchange of expressions according to claim 1, characterized in that: The step of fusing the facial key point features of the facial expression image with the exchanged facial features to generate fused features specifically includes: The pre-trained facial key point detection network is used to input two facial expression images x a ,x b , respectively extract the key features of the face Key point features of the face Perform image block division and linear transformation, respectively mapped into vector f a ,f b , will exchange facial features and vector f b , exchange facial features and vector f a Perform feature fusion separately to generate fusion features; exchange facial features Represents two facial expression images x respectively a ,x b The corresponding exchange facial features.

7. The facial emotion recognition method based on generative adversarial separation and expression exchange according to claim 1, characterized in that: The method of generating a reconstructed facial expression image and a synthesized exchanged facial expression image based on fusion features and using a generator, and constraining the identity consistency between the reconstructed image and the original facial expression image, specifically includes: The generator receives the fusion features and the emotion-related features and emotion-irrelevant features corresponding to the two facial expression images, adds a mask vector of the same shape, and generates a reconstructed facial expression image and a synthesized exchanged facial expression image based on the generator and using a cross-attention mechanism; the synthesized exchanged facial expression image contains the identity of the original facial expression image and the expression of the other facial expression image; The original facial expression image, the corresponding reconstructed facial expression image and the synthesized exchanged facial expression image are taken as a group of image pairs, input into the arcface model to extract feature vectors, and the cosine similarity between the feature vectors is calculated to keep the identity of each group of image pairs consistent.

8. The facial emotion recognition method based on generative adversarial separation and exchange of expressions according to claim 1, characterized in that: The method of using the discriminator to classify the expression of the reconstructed image and to judge the authenticity specifically includes: By adding an auxiliary classifier to the discriminator to classify the expression of the reconstructed image and judge the authenticity of the reconstructed image, it can compete with the generator that generates the reconstructed image; The facial emotion classification based on emotion-related features by a classifier specifically includes: The emotion-related features are input into a multi-layer perceptron, and the mean operation is performed on the sequence. Then, they are added to the category label representing the emotion, and processed through multiple fully connected layers to output the predicted facial expression category.

9. The facial emotion recognition method based on generative adversarial separation and exchange of expressions according to claim 1, characterized in that: The combination of identification loss, reconstruction loss, orthogonal constraint, identity constraint loss and emotion classification loss optimizes the parameters of the facial expression recognition model, specifically including: After extracting the input image features through the discriminator, auxiliary classification is performed, and the identification loss is calculated using the BCEWithLogits loss; The reconstruction loss is obtained by calculating the pixel-level difference between the original facial expression image and the reconstructed facial expression image; Orthogonality constraints are calculated by the inner product of emotion-related features and emotion-irrelevant features; The identity constraint loss is generated when constraining the identity consistency between the reconstructed image and the original facial expression image; The sentiment classification loss is calculated by the logical value of the true label category corresponding to the training sample and the logical value of each category of the training sample.

10. A facial emotion recognition system based on generative adversarial separation and expression exchange, characterized in that: include: Pre-training module, which pre-trains the visual transformer based on a self-supervised mask autoencoder; The facial feature vector extraction module builds a facial expression recognition model based on the visual transformer and the generative adversarial network. Two unpaired facial expression images are input and the facial feature vectors corresponding to the two facial expression images are extracted through the image encoder in the visual transformer. The expression separation and exchange module decouples the facial feature vector, separates the emotion-related features from the emotion-irrelevant features, and performs feature exchange to generate exchanged facial features; The modality fusion module fuses the facial key point features of the facial expression image with the exchanged face features to generate fused features; A reconstruction supervision module generates a reconstructed facial expression image and a synthesized exchanged facial expression image based on the fused features and using a generator, and constrains the identity consistency between the reconstructed image and the original facial expression image; In the training module, the discriminator is used to classify the expression of the reconstructed image and make authenticity judgment. The classifier is used to classify facial emotions based on emotion-related features, and the parameters of the facial expression recognition model are optimized by combining the discrimination loss, reconstruction loss, orthogonal constraint and identity constraint loss.

Citation Information

Cited By

  • Facial expression migration method and device based on adversarial auto-encoder, equipment and medium

    CN120495073A

  • Emotion analysis method and system

    CN120526807A

  • AI digital human expression and facial feature migration method and system

    CN120997351A