An expression recognition method, device, electronic device and storage medium

Through the self-attention mechanism, the global and local features of expressions are integrated and fine-grained adversarial learning is carried out, which solves the problem of distribution differences in expression recognition across data sets, and improves the recognition accuracy and robustness.

CN115273175BActive Publication Date: 2025-06-27SOUTH CHINA NORMAL UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202210710861.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-06-22
Publication Date
2025-06-27
Estimated Expiration
2042-06-22

AI Technical Summary

Technical Problem

The prior art has a problem of distribution differences in expression recognition across data sets, resulting in poor recognition effect of model in practical applications.

Method used

The self-attention mechanism is used to learn the weights of the global features and local features of expressions, which improves the robustness of expression recognition, and allows the expression features between different data sets to be aligned in categories through fine-grained adversarial learning.

Benefits of technology

It improves the robustness of expression recognition and the accuracy of expression recognition across data sets, providing stronger migration characteristics.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115273175B_ABST
    Figure CN115273175B_ABST
Patent Text Reader

Abstract

The present invention relates to a facial expression recognition method, apparatus, electronic device, and storage medium. The facial expression recognition method of the present invention includes: obtaining a facial expression image to be recognized; performing face cropping and key point detection on the facial expression image to obtain a face global image of a preset size and face key point coordinates; inputting the face global image and the face key points into a feature extraction model to obtain a facial expression global feature and a facial expression local feature; based on a self-attention mechanism, fusing the facial expression global feature and the facial expression local feature to obtain a facial expression fusion feature; and inputting the facial expression fusion feature into a trained facial expression classifier to obtain a facial expression recognition result. The facial expression recognition method of the present invention learns the features of facial expressions based on a self-attention mechanism, improves the robustness of facial expression recognition, and provides more transferable features for subsequent cross-dataset facial expression recognition.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of facial expression, and in particular to an expression recognition method, device, electronic device and storage medium. Background Art

[0002] Facial expression recognition has always been a hot topic in the field of computer vision. Research on automatic facial expression recognition helps to understand human behavior, and automatic expression recognition is also widely used in fields such as human-computer interaction, online education and safe driving. In the past decade, researchers have collected many different facial expression datasets, including datasets in laboratory environments (such as CK+, JAFFE, MMI, Oulu-CASIA) and natural environments (such as RAF, SFEW2.0, ExpW, FER2013). With these datasets, different deep learning models have been intensively proposed, gradually promoting the performance of expression recognition.

[0003] Although great progress has been made in facial expression recognition technology in recent years, most of the existing facial expression recognition requires that the datasets used for training and testing are the same dataset, that is, the expression images in the training set and the test set are collected under the same conditions. However, in many practical applications, this assumption does not hold, because the test dataset is usually collected online and is more difficult to control than the training data. In different application scenarios, the distribution differences of expression data are common. In this case, when the model trained by limited distribution data is extended to unknown data in reality, the model often achieves unsatisfactory recognition effects.

[0004] In view of the current problems of distribution differences in facial expression datasets and poor performance of expression recognition models in practical applications, researchers use two expression datasets with different distributions to simulate the distribution differences of expression data in real application scenarios. One of the datasets is a labeled source dataset, which is used to simulate the situation with labels during expression recognition training; the other dataset is an unlabeled target dataset, which is used to simulate the situation where data is unknown in actual expression recognition applications. Currently, the methods for cross-dataset expression recognition are mainly divided into the following three types:

[0005] 1. Methods based on transductive learning. The classifier is trained through the expression data and labels of the source dataset to generate pseudo-labels for the target dataset, and the classifier is continuously trained with the pseudo-labels, and the pseudo-labels of the target dataset are further optimized with the classifier. Such cyclic training is performed for joint optimization to achieve cross-dataset effects.

[0006] The disadvantage of these methods is that they rely on the accuracy of the initial expression recognition model in the target dataset. Once the distribution differences between the two datasets are large, the cross-dataset effect of the model will be greatly reduced.

[0007] 2. Domain adaptation methods based on minimizing statistical differences. The domain adaptation methods based on minimizing statistical differences map the feature maps obtained by feature extraction on the training set and the test set into a high-dimensional space, and minimize the statistical differences between the two data sets to achieve cross-data set facial expression recognition.

[0008] The common problems of these methods are that calculating the statistical differences of the data set features after each training will increase the computational cost, increase the running time, and rely on a higher-performance processor.

[0009] 3. Domain adaptation methods based on adversarial learning. These methods incorporate the idea of adversarial learning into domain adaptation, making the trained classifier unable to distinguish the data of the training set and the test set to obtain domain-invariant features, thereby achieving cross-data set facial expression recognition.

[0010] The common problems of these methods are that only the overall distribution differences between facial expression data sets are considered during the cross-data set process, ignoring the decision boundary of the classification task, resulting in the problem of mismatched categories of facial expression data sets after cross-data sets. Summary of the Invention

[0011] Based on this, the object of the present invention is to provide a facial expression recognition method, device, electronic device and storage medium, which learn the features of facial expressions based on the self-attention mechanism, improve the robustness of facial expression recognition, and provide more transferable features for subsequent cross-data set facial expression recognition.

[0012] In the first aspect, the present invention provides a facial expression recognition method, including the following steps:

[0013] Obtain a facial expression image to be recognized;

[0014] Perform face cropping and key point detection on the facial expression image to obtain a face global image of a preset size and face key point coordinates;

[0015] Input the face global image and the face key points into a feature extraction model to obtain a facial expression global feature and a facial expression local feature;

[0016] Based on the self-attention mechanism, fuse the facial expression global feature and the facial expression local feature to obtain a facial expression fusion feature;

[0017] Input the facial expression fusion feature into a trained facial expression classifier to obtain a facial expression recognition result.

[0018] Further, the training steps of the facial expression classifier include:

[0019] Obtain a training data set, where the training data set includes training facial expression images and corresponding facial expression classification labels;

[0020] Perform face cropping and key point detection on the training expression images to obtain a face global image of a preset size and face key point coordinates;

[0021] Input the face global image and the face key points into a feature extraction model to obtain an expression global feature and an expression local feature;

[0022] Based on the self-attention mechanism, fuse the expression global feature and the expression local feature to obtain an expression fusion feature;

[0023] Use the expression fusion feature and the corresponding expression classification label to train the expression classifier to obtain a trained expression classifier.

[0024] Furthermore, the training steps of the expression classifier further include:

[0025] Obtain a target data set without labels;

[0026] Perform face cropping and key point detection on the target data set to obtain a face global image of a preset size and face key point coordinates;

[0027] Input the face global image and the face key points into a feature extraction model to obtain an expression global feature and an expression local feature;

[0028] Based on the self-attention mechanism, fuse the expression global feature and the expression local feature to obtain an expression fusion feature;

[0029] Add a domain discriminator to the trained expression classifier;

[0030] Use the expression fusion features of the training data set and the target data set to perform adversarial training on the domain discriminator and the classifier;

[0031] After the training is completed, save the feature extraction model and the classifier.

[0032] Furthermore, the feature extraction model is an MTCNN model.

[0033] Furthermore, the face key point coordinates include the coordinates of the left eye, right eye, nose, left mouth corner, and right mouth corner;

[0034] The expression local feature is extracted through a LocalNet module, and the input of the LocalNet module is a key region cropped with a size of 0.2n * 0.2n * 3 centered on five key points.

[0035] Furthermore, based on the self-attention mechanism, fusing the expression global feature and the expression local feature to obtain an expression fusion feature includes:

[0036] Obtain 1 global expression feature and 5 local expression features;

[0037] Multiply each 1*128-dimensional feature by three 128*128 transformation matrices W q 、W k 、W v to obtain corresponding 128-dimensional values, denoted as q i 、k i 、v i ;

[0038] Use the following formula to calculate the weights between features and obtain the fused expression feature x i :

[0039]

[0040] where d is the feature dimension, d = 128.

[0041] Furthermore, use the expression fusion features of the training dataset and the target dataset to perform adversarial training on the domain discriminator and the classifier, including:

[0042] Obtain the fusion feature x of the target dataset based on the fusion expression feature obtained by the self-attention mechanism i , and input the fusion feature x i into a two-layer MLP to obtain the soft label of the expression;

[0043] For all source dataset expression images and all target dataset expression images, expand their K-dimensional labels to 2K-dimensional labels, where K is the number of expression categories. The labels of the source dataset use the original label information in the 1st to Kth dimensions and are set to 0 in the (K + 1)th to 2Kth dimensions. The labels of the target dataset are set to 0 in the 1st to Kth dimensions and use the soft labels obtained above in the (K + 1)th to 2Kth dimensions;

[0044] Use the following formula to continue inputting the fusion feature obtained from the source dataset into a two-layer MLP to calculate the expression classification loss Lcls. Lcls is the classification loss of the expression, and the cross-entropy loss is used to minimize the difference between the predicted classification and the true expression classification on the source dataset:

[0045]

[0046] where N is the number of samples in the source dataset, y ik is the label that the i-th expression image is the k-th type of expression; p ik is the probability that the fused expression is the k-th type of expression;

[0047] Using the following formula, the fused features of the source dataset and the target dataset are respectively input into the domain-class discriminator to calculate the domain-class adversarial loss Ladv:

[0048]

[0049] where S represents the data of the source dataset, and T represents the data of the target dataset. a ik , a jk is the class information that the source domain sample i or the target domain sample j belongs to the k-th class, f i , f j is the expression fusion feature, P is the probability that the predicted fusion feature belongs to the k-th class and comes from the d dataset, d = 0 means the sample comes from the source dataset, and d = 1 means the sample comes from the target dataset;

[0050] Using the following formula, calculate the total loss after combining the two losses:

[0051] L = αL cls + βL adv

[0052] where α and β are the loss weights, initialized to 1 and 10, and the adversarial goal is to minimize L;

[0053] When the training reaches the minimum of the loss L, stop the training and save the feature extractor and the classifier for the expression recognition of the target dataset.

[0054] In a second aspect, the present invention also provides an expression recognition device, including:

[0055] An image acquisition module, configured to acquire an expression image to be recognized;

[0056] A key point detection module, configured to perform face cropping and key point detection on the expression image to obtain a face global image of a preset size and face key point coordinates;

[0057] A feature extraction module, configured to input the face global image and the face key points into a feature extraction model to obtain an expression global feature and an expression local feature;

[0058] A feature fusion module, configured to fuse the expression global feature and the expression local feature based on the self-attention mechanism to obtain an expression fusion feature;

[0059] An expression recognition module, configured to input the expression fusion feature into a trained expression classifier to obtain an expression recognition result.

[0060] In a third aspect, the present invention also provides an electronic device, including:

[0061] At least one memory and at least one processor;

[0062] The memory is used to store one or more programs;

[0063] When the one or more programs are executed by the at least one processor, the at least one processor implements the steps of an expression recognition method according to any one of the first aspects of the present invention.

[0064] In a fourth aspect, the present invention further provides a computer-readable storage medium,

[0065] The computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, it implements the steps of an expression recognition method according to any one of the first aspects of the present invention.

[0066] An expression recognition method, device, electronic device and storage medium provided by the present invention use a self-attention mechanism to learn and fuse the weights of the global features and local features of expressions, improve the robustness of expression recognition, and provide more transferable features for subsequent cross-dataset expression recognition; perform fine-grained adversarial learning between the source dataset and the target dataset of expressions, align the expression features between different datasets in terms of categories, and improve the accuracy of cross-dataset expression recognition.

[0067] For better understanding and implementation, the present invention will be described in detail below with reference to the accompanying drawings. BRIEF DESCRIPTION OF THE DRAWINGS

[0068] Figure 1 It is a schematic diagram of the steps of an expression recognition method provided by the present invention;

[0069] Figure 2 It is a schematic flowchart of an expression recognition method provided by the present invention in one embodiment;

[0070] Figure 3 It is a schematic diagram of a neural network model used in one embodiment of the present invention;

[0071] Figure 4 It is a schematic structural diagram of an expression recognition device provided by the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0072] To make the objectives, technical solutions and advantages of the present application clearer, the embodiments of the present application will be further described in detail below with reference to the accompanying drawings.

[0073] It should be clear that the described embodiments are only a part of the embodiments of the present application, rather than all of the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative efforts fall within the scope of protection of the embodiments of the present application.

[0074] The terms used in the embodiments of this application are only for the purpose of describing specific embodiments and are not intended to limit the embodiments of this application. The singular forms "a", "the", and "said" used in the embodiments of this application and the appended claims are also intended to include the plural forms unless the context clearly indicates otherwise. It should also be understood that the term "and / or" as used herein refers to and encompasses any and all possible combinations of one or more of the associated listed items.

[0075] When the following description refers to the accompanying drawings, unless otherwise indicated, the same numbers in different drawings represent the same or similar elements. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with this application. On the contrary, they are merely examples of devices and methods consistent with some aspects of this application as detailed in the appended claims. In the description of this application, it should be understood that the terms "first", "second", "third", etc. are only used to distinguish similar objects and do not have to be used to describe a specific order or sequence, nor can they be understood as indicating or implying relative importance. For those of ordinary skill in the art, the specific meanings of the above terms in this application can be understood according to specific circumstances.

[0076] In addition, in the description of this application, unless otherwise specified, "a plurality of" means two or more. "And / or" describes the association relationship of associated objects and indicates that three relationships can exist. For example, A and / or B can represent three situations: A exists alone, A and B exist simultaneously, and B exists alone. The character " / " generally represents an "or" relationship between the associated objects before and after.

[0077] In response to the problems in the background art, the embodiments of this application provide an expression recognition method, as Figures 1-3 shown, the method includes the following steps:

[0078] S01: Obtain an expression image to be recognized.

[0079] S02: Crop the face and detect key points of the expression image to obtain a face global image of a preset size and face key point coordinates.

[0080] S03: Input the face global image and the face key points into a feature extraction model to obtain an expression global feature and an expression local feature.

[0081] In a preferred embodiment, the network framework is as Figure 3As shown, based on this network framework, a facial expression image of n*n*3 pixels is transformed into the standard input data of the model through an expression input layer, and then input into 4 modules to extract the global features of the expression (Block1-4 in the figure). The 4 Blocks are implemented based on Resnet-18. Each Block includes 2 residual modules. The input data of each residual module is processed in two ways. First, the input data is normalized using the BatchNormalization layer, and then passed through two 3*3 depth convolutional layers, and then processed using the max pooling layer. The data dimensions obtained after the two processes are the same. The two data are added together to obtain the final output data. The width and height of the output data are halved, and the number of channels is doubled (the number of channels in the first module remains unchanged, which is 64 of the original). After passing through 4 modules, the n*n*3 expression image becomes a 7*7*512-dimensional expression feature map, and finally a 128-dimensional expression global feature is obtained through the output layer.

[0082] Meanwhile, the model extracts local expression features based on expression key points. Since the decisive information of expressions is concentrated at the positions of the five facial features, the model locates five key points on the face (left eye, right eye, nose, left mouth corner, right mouth corner). The MTCNN (Multi-Task Convolutional Neural Network) is used for face key point localization. After inputting an expression image through MTCNN, the x and y coordinates of the five key points in the image are obtained. After locating the key points, in the original n*n*3 expression image, five local regions of the expression are cropped with the five key points (x, y) as the center to obtain the key regions of the expression information. The cropping size is 0.2n*0.2n*3. After obtaining the key regions, they are input into the LocalNet module for local feature extraction. After inputting 5 local regions of 0.2n*0.2n*3, 5 128-dimensional expression local features are obtained.

[0083] S04: Based on the self-attention mechanism, fuse the expression global features and the expression local features to obtain expression fusion features.

[0084] Preferably, based on the 1 expression global feature and 5 expression local features obtained in the above step 3, the fusion features include the following sub-steps:

[0085] S041: Obtain 1 expression global feature and 5 expression local features.

[0086] S042: Multiply each 1*128-dimensional feature by three 128*128 transformation matrices W q 、W k 、W v, the corresponding 128 - dimensional values are obtained, denoted as q respectively i , k i , v i .

[0087] S043: Use the following formula to calculate the weights between features and obtain the fused expression feature x i .

[0088]

[0089] where d is the feature dimension, d = 128.

[0090] S05: Input the expression fusion feature into the trained expression classifier to obtain the expression recognition result.

[0091] In a preferred embodiment, the training steps of the expression classifier include the following sub - steps.

[0092] S11: Obtain a training data set, which includes training expression images and corresponding expression classification labels.

[0093] S12: Perform face cropping and key - point detection on the training expression images to obtain a face global image of a preset size and face key - point coordinates.

[0094] S13: Input the face global image and the face key - points into a feature extraction model to obtain an expression global feature and an expression local feature.

[0095] S14: Based on the self - attention mechanism, fuse the expression global feature and the expression local feature to obtain an expression fusion feature.

[0096] S15: Use the expression fusion feature and the corresponding expression classification label to train the expression classifier to obtain a trained expression classifier.

[0097] Among them, in steps S11 - S14, the processing method for face expression images is the same as that in steps S01 - S04, so it will not be elaborated here.

[0098] After obtaining the fusion feature, input the fusion feature x i feature into a two - layer MLP layer ( Figure 3 FC layer) for discrimination of expression categories. In this way, an expression recognition model is trained on the source data set.

[0099] In view of the problem of poor accuracy in cross-dataset facial expression recognition in the prior art, in order to apply the facial expression recognition model trained as described above to the facial expression data of unlabeled unknown faces, an unlabeled target dataset is used to simulate unknown face data, and a domain discriminator is added for the migration of the facial expression recognition model. Through the input of the source facial expression dataset (with labels) and the target dataset (without labels) for fine-grained adversarial training, the model is migrated to the target dataset. The traditional domain discriminator inputs the features of the two datasets and discriminates whether the features come from the source domain or the target domain, that is, it can only discriminate [1,0] and [0,1], which represent that the features come from the source domain or the target domain respectively. The invention splits each channel into K channels, a total of 2K channels, so that the confrontation can occur not only between the source domain and the target domain, but also at a finer-grained class level.

[0100] In a preferred embodiment, the training steps of the facial expression classifier include the following sub-steps:

[0101] S21: Obtain an unlabeled target dataset.

[0102] S22: Perform face cropping and key point detection on the target dataset to obtain a face global image of a preset size and face key point coordinates.

[0103] S23: Input the face global image and the face key points into a feature extraction model to obtain facial expression global features and facial expression local features.

[0104] S24: Based on the self-attention mechanism, fuse the facial expression global features and the facial expression local features to obtain facial expression fusion features.

[0105] S25: Add a domain discriminator to the trained facial expression classifier.

[0106] S26: Use the facial expression fusion features of the training dataset and the target dataset to perform adversarial training on the domain discriminator and the classifier.

[0107] S27: After the training is completed, save the feature extraction model and the classifier

[0108] In a specific embodiment, first, we use the fused facial expression features obtained based on the self-attention mechanism to obtain the fused feature x of the target dataset (without labels). i , according to the fused feature x iThe soft labels of expressions are obtained by inputting them into a two-layer MLP for the first-step training. Subsequently, we combine all the expression images from the source datasets and all the expression images from the target datasets, and expand their K-dimensional labels to 2K-dimensional labels, where K is the number of expression categories. For the source datasets, the labels in the 1st to Kth dimensions use the original label information, and the data in the (K + 1)th to 2Kth dimensions are set to 0. For the target datasets, the data in the 1st to Kth dimensions are set to 0, and the soft labels obtained previously are used in the (K + 1)th to 2Kth dimensions.

[0109] We continue to input the fused features obtained from the source datasets into a two-layer MLP to calculate the expression classification loss Lcls. Lcls is the classification loss of expressions, and the cross-entropy loss is used to minimize the difference between the predicted classification and the true expression classification on the source datasets. The formula is shown in (2):

[0110]

[0111] where N is the number of samples in the source datasets, y ik is the label indicating that the i-th expression image belongs to the k-th expression category. p ik is the probability that the fused expression belongs to the k-th expression category.

[0112] Meanwhile, the N expression images from the source datasets and the N expression images from the target datasets are used to extract fused features through Step 1. The fused features of the source datasets and the target datasets are respectively input into the domain-category discriminator to calculate the domain-category adversarial loss Ladv. At the same time, Ladv is the domain-category adversarial loss. To enable the adversarial process to occur not only between the source domain and the target domain but also at a finer-grained class level. The domain-category adversarial loss is shown in formula (3):

[0113]

[0114] where S represents the data in the source datasets, and T represents the data in the target datasets. a ik , a jk is the category information indicating that the i-th sample in the source domain or the j-th sample in the target domain belongs to the k-th category, f i , f j is the expression fused feature, and P is the probability that the predicted fused feature belongs to the k-th category and comes from the d dataset (d = 0 indicates that the sample comes from the source datasets, and d = 1 indicates that the sample comes from the target datasets).

[0115] The total loss after combining the two losses is shown in formula (4):

[0116] L = αL cls + βL adv (4)

[0117] where α and β are the loss weights, initialized as 1 and 10, and the adversarial objective is to minimize L.

[0118] When the training reaches the minimum loss L, stop the training and save the feature extractor and classifier for facial expression recognition of the target dataset.

[0119] In a preferred embodiment, the model uses the flowchart as Figure 2 shown. First, perform face detection and key point extraction on the facial expression image through MTCNN, then extract global and local facial expression features based on the face image and the positions of facial expression key points. After obtaining the two types of features, fuse the global and local facial expression features based on the self-attention mechanism to obtain the fused feature. Then, use the trained facial expression classifier to perform facial expression recognition. When performing cross-dataset facial expression recognition, extract the fused feature for the facial expression image of the target dataset without labels using the same method, calculate the total loss L by combining the fused features between the two datasets. When L is trained to the minimum, save the facial expression fusion feature extraction module and the facial expression classifier module for facial expression recognition of the target dataset. Thus, cross-dataset facial expression recognition is achieved through the labeled source label dataset and the unlabeled target label dataset.

[0120] The embodiment of the present application also provides a facial expression recognition device, as Figure 4 shown. The facial expression recognition device 400 includes:

[0121] An image acquisition module 401, configured to acquire a facial expression image to be recognized;

[0122] A key point detection module 402, configured to perform face cropping and key point detection on the facial expression image to obtain a face global image of a preset size and face key point coordinates;

[0123] A feature extraction module 403, configured to input the face global image and the face key points into a feature extraction model to obtain a facial expression global feature and a facial expression local feature;

[0124] A feature fusion module 04, configured to fuse the facial expression global feature and the facial expression local feature based on the self-attention mechanism to obtain a facial expression fused feature;

[0125] A facial expression recognition module 405, configured to input the facial expression fused feature into a trained facial expression classifier to obtain a facial expression recognition result.

[0126] Preferably, the training steps of the facial expression classifier include:

[0127] Obtain a training dataset, where the training dataset includes training facial expression images and corresponding facial expression classification labels;

[0128] Perform face cropping and key point detection on the training facial expression images to obtain a face global image of a preset size and face key point coordinates;

[0129] Input the global face image and the face key points into a feature extraction model to obtain global expression features and local expression features;

[0130] Based on the self-attention mechanism, fuse the global expression features and the local expression features to obtain fused expression features;

[0131] Use the fused expression features and the corresponding expression classification labels to train the expression classifier to obtain a trained expression classifier.

[0132] Preferably, the training steps of the expression classifier further include:

[0133] Obtain a target data set without labels;

[0134] Perform face cropping and key point detection on the target data set to obtain a global face image and face key point coordinates of a preset size;

[0135] Input the global face image and the face key points into a feature extraction model to obtain global expression features and local expression features;

[0136] Based on the self-attention mechanism, fuse the global expression features and the local expression features to obtain fused expression features;

[0137] Add a domain discriminator to the trained expression classifier;

[0138] Use the fused expression features of the training data set and the target data set to perform adversarial training on the domain discriminator and the classifier;

[0139] After the training is completed, save the feature extraction model and the classifier.

[0140] Preferably, the feature extraction model is an MTCNN model.

[0141] Preferably, the face key point coordinates include the coordinates of the left eye, right eye, nose, left mouth corner, and right mouth corner;

[0142] The local expression features are extracted through a LocalNet module, and the input of the LocalNet module is a key region cropped with a size of 0.2n * 0.2n * 3 centered on five key points.

[0143] Preferably, based on the self-attention mechanism, fusing the global expression features and the local expression features to obtain fused expression features includes:

[0144] Obtain 1 global expression feature and 5 local expression features;

[0145] Multiply each 1*128-dimensional feature by three 128*128 transformation matrices W obtained through training q 、W k 、W v to obtain corresponding 128-dimensional values, denoted as q i 、k i 、v i ;

[0146] Use the following formula to calculate the weights between features and obtain the fused expression feature x i :

[0147]

[0148] where d is the feature dimension, d = 128.

[0149] Preferably, use the expression fusion features of the training data set and the target data set to perform adversarial training on the domain discriminator and the classifier, including:

[0150] Obtain the fusion feature x of the target data set based on the fusion expression feature obtained by the self-attention mechanism i , and input the fusion feature x i into a two-layer MLP to obtain the soft label of the expression;

[0151] For all source data set expression images and all target data set expression images, expand their K-dimensional labels to 2K-dimensional labels, where K is the number of expression categories. The labels of the source data set use the original label information in the 1st to Kth dimensions and are set to 0 in the (K + 1)th to 2Kth dimensions. The labels of the target data set are set to 0 in the 1st to Kth dimensions and use the soft labels obtained above in the (K + 1)th to 2Kth dimensions;

[0152] Use the following formula to continue inputting the fusion features obtained from the source data set into a two-layer MLP to calculate the expression classification loss Lcls. Lcls is the classification loss of the expression, and the cross-entropy loss is used to minimize the difference between the predicted classification and the true expression classification on the source data set:

[0153]

[0154] where N is the number of samples in the source data set, y ik is the label that the i-th expression image is the k-th type of expression; p ik is the probability that the fused expression is the k-th type of expression;

[0155] Use the following formula to input the fusion features of the source data set and the target data set into the domain-class discriminator respectively to calculate the domain-class adversarial loss Ladv:

[0156]

[0157] Among them, S represents the data of the source dataset, and T represents the data of the target dataset. a ik ,a jk is the class information that the source domain sample i or the target domain sample j belongs to the k-th class, f i ,f j is the expression fusion feature, P is the probability that the predicted fusion feature belongs to the k-th class and comes from the d dataset. When d = 0, the sample comes from the source dataset, and when d = 1, the sample comes from the target dataset;

[0158] Use the following formula to calculate the total loss after combining the two losses:

[0159] L = αL cls + βL adv

[0160] Among them, α and β are the loss weights, initialized to 1 and 10, and the adversarial objective is to minimize L;

[0161] When the training reaches the minimum of the loss L, stop the training and save the feature extractor and classifier for the expression recognition of the target dataset.

[0162] For the device embodiment, since it basically corresponds to the method embodiment, the relevant parts can refer to the partial description of the method embodiment. The device embodiments described above are only illustrative. The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units. Those skilled in the art can clearly understand that for the convenience and simplicity of description, the specific working processes of the above-described system, device, and unit can refer to the corresponding processes in the foregoing method embodiments, and will not be elaborated here.

[0163] This application embodiment also provides an electronic device, including:

[0164] At least one memory and at least one processor;

[0165] The memory is used to store one or more programs;

[0166] When the one or more programs are executed by the at least one processor, the at least one processor implements the steps of an expression recognition method as described above.

[0167] For the device embodiments, since they basically correspond to the method embodiments, the relevant parts can be referred to the descriptions of the method embodiments. The device embodiments described above are merely illustrative. The components described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed to multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the present disclosure. A person of ordinary skill in the art can understand and implement it without creative work.

[0168] The embodiments of the present application also provide a computer-readable storage medium.

[0169] The computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the steps of an expression recognition method as described above are implemented.

[0170] Computer-usable storage media include permanent and non-permanent, removable and non-removable media, and information storage can be implemented by any method or technology. The information can be computer-readable instructions, data structures, program modules, or other data. Examples of computer storage media include, but are not limited to: phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technologies, compact disc read-only memory (CD-ROM), digital versatile disc (DVD) or other optical storage, magnetic cassette tapes, magnetic tape magnetic disk storage or other magnetic storage devices, or any other non-transmission media that can be used to store information accessible by a computing device.

[0171] An expression recognition method, device, electronic device, and storage medium provided by the present invention learn the features of expressions based on the self-attention mechanism. The present invention performs key point positioning on facial expressions and detects key points of the main regions (both eyes, nose, and corners of the mouth) that determine the expression category to obtain the positions of expression key points. Subsequently, an expression feature extraction module is constructed based on a residual convolutional neural network to extract the global features and local features of expressions respectively. After obtaining the two sets of features, the self-attention mechanism is used to perform weighted fusion of the expression features to obtain the fused expression features, and an expression classifier is trained through the fused features on the source dataset.

[0172] Aiming at the problem of class mismatch between different data sets, the present invention designs a fine-grained adversarial learning strategy based on expression fusion features to align the expression features between different data sets at the class level. A fine-grained domain discriminator is adopted, which not only distinguishes domains between data sets, but also distinguishes domains at the class level. In order to make the discriminator not only focus on distinguishing domains, the invention first uses a previously trained expression classifier to generate soft labels of the expressions in the target data set, combines the generated soft labels with the labels of the source data set to generate domain-class labels, and then performs fine-grained adversarial learning by splitting both output channels of the traditional binary-class domain discriminator into K channels (K is the number of expression classes).

[0173] The self-attention mechanism is used to learn and fuse the weights of the global and local features of the expression, improving the robustness of expression recognition and providing more transferable features for subsequent cross-data-set expression recognition; fine-grained adversarial learning is carried out between the source data set and the target data set of the expression, aligning the expression features between different data sets at the class level and improving the accuracy of cross-data-set expression recognition.

[0174] The above embodiments only represent several implementation manners of the present invention. The description is relatively specific and detailed, but it should not be construed as a limitation on the scope of the invention patent. It should be noted that for those of ordinary skill in the art, without departing from the concept of the present invention, several deformations and improvements can still be made, and these all belong to the protection scope of the present invention.

Claims

1. A method for facial expression recognition, characterized in that, It includes the following steps: Obtain the facial expression image to be recognized; Perform face cropping and key point detection on the facial expression image to obtain a face global image of a preset size and face key point coordinates; Input the face global image and the face key points into a feature extraction model to obtain an expression global feature and an expression local feature; wherein, the feature extraction model includes an MTCNN model and a LocalNet module; the MTCNN model is used for face key point positioning; the LocalNet module is used for local feature extraction; Based on the self-attention mechanism, fuse the expression global feature and the expression local feature to obtain an expression fusion feature; Input the expression fusion feature into a trained expression classifier to obtain an expression recognition result; wherein, the training steps of the expression classifier include: Obtain an unlabeled target data set; Perform face cropping and key point detection on the target data set to obtain a face global image of a preset size and face key point coordinates; Input the face global image and the face key points into a feature extraction model to obtain an expression global feature and an expression local feature; Based on the self-attention mechanism, fuse the expression global feature and the expression local feature to obtain an expression fusion feature; Add a domain discriminator to the trained expression classifier; Use the expression fusion features of the training data set and the target data set to perform adversarial training on the domain discriminator and the classifier, including: Obtain the fused feature x of the target dataset based on the fused expression features obtained by the self-attention mechanism i , according to the fused feature x i Input it into a two-layer MLP to obtain the soft label of the expression; For all source data set facial expression images and all target data set facial expression images, expand their K-dimensional labels to 2K-dimensional labels, where K is the number of expression categories. The labels of the source data set use the original label information in the 1st to Kth dimensions and are set to 0 in the (K + 1)th to 2Kth dimensions. The labels of the target data set are set to 0 in the 1st to Kth dimensions and use the soft labels obtained previously in the (K + 1)th to 2Kth dimensions; Use the following formula to continue inputting the fusion features obtained from the source data set into a two-layer MLP to calculate the expression classification loss Lcls. Lcls is the classification loss of the expression, and the cross-entropy loss is used to minimize the difference between the predicted classification and the true expression classification on the source data set: where N is the number of samples in the source dataset, and y ik is the label that the i-th facial expression image belongs to the k-th type of facial expression; p ik is the probability that the fused facial expression belongs to the k-th type of facial expression; Use the following formula to input the fusion features of the source data set and the target data set into the domain-category discriminator respectively to calculate the domain-category adversarial loss Ladv: Among them, S represents the data of the source dataset, and T represents the data of the target dataset; a ik , a jk is the class information that the source domain sample i or the target domain sample j belongs to the k-th class, f i , f j is the expression fusion feature, P is the probability that the predicted fusion feature belongs to the k-th class and comes from the d dataset. d = 0 means the sample comes from the source dataset, and d = 1 means the sample comes from the target dataset; Use the following formula to calculate the total loss after combining the two losses: L = αL cls + βL adv where α and β are loss weights, initialized to 1 and 10, and the adversarial goal is to minimize L; When the training reaches the minimum of the loss L, stop the training and save the feature extractor and the classifier for the expression recognition of the target data set; After the training is completed, save the feature extraction model and the classifier.

2. The facial expression recognition method according to claim 1, characterized in that The training steps of the expression classifier include: Obtain a training data set, where the training data set includes training facial expression images and corresponding expression classification labels; Perform face cropping and key point detection on the training facial expression images to obtain a face global image of a preset size and face key point coordinates; Input the face global image and the face key points into a feature extraction model to obtain an expression global feature and an expression local feature; Based on the self-attention mechanism, fuse the global expression feature and the local expression feature to obtain a fused expression feature; Use the fused expression feature and the corresponding expression classification label to train the expression classifier to obtain a trained expression classifier.

3. An expression recognition method according to claim 2, wherein: The facial key point coordinates include the coordinates of the left eye, right eye, nose, left mouth corner, and right mouth corner; The local expression feature is extracted by the LocalNet module, and the input of the LocalNet module is a key region centered on five key points and cropped to a size of 0.2n * 0.2n * 3.

4. A facial expression recognition method according to claim 3, characterized in that, Based on the self-attention mechanism, fusing the global expression feature and the local expression feature to obtain a fused expression feature includes: Obtain 1 global expression feature and 5 local expression features; Multiply each 1*128-dimensional feature by three 128*128 transformation matrices W obtained through training q 、W k 、W v to obtain the corresponding 128-dimensional values, denoted as q i 、k i 、v i ; Use the following formula to calculate the weights between features and obtain the fused expression feature x i : Where d is the feature dimension, d = 128.

5. An expression recognition device, characterized in that, Including: An image acquisition module for acquiring an expression image to be recognized; A key point detection module for performing face cropping and key point detection on the expression image to obtain a face global image of a preset size and facial key point coordinates; A feature extraction module for inputting the face global image and the facial key points into a feature extraction model to obtain a global expression feature and a local expression feature; wherein, the feature extraction model includes an MTCNN model and a LocalNet module; the MTCNN model is used for facial key point positioning; the LocalNet module is used for local feature extraction; A feature fusion module for fusing the global expression feature and the local expression feature based on the self-attention mechanism to obtain a fused expression feature; An expression recognition module for inputting the fused expression feature into a trained expression classifier to obtain an expression recognition result; wherein, the training steps of the expression classifier include: Obtain an unlabeled target data set; Perform face cropping and key point detection on the target data set to obtain a face global image of a preset size and facial key point coordinates; Input the face global image and the facial key points into a feature extraction model to obtain a global expression feature and a local expression feature; Based on the self-attention mechanism, fuse the global expression feature and the local expression feature to obtain a fused expression feature; Add a domain discriminator to the trained expression classifier; Use the fused expression features of the training data set and the target data set to perform adversarial training on the domain discriminator and the classifier, including: Obtain the fused feature x of the target dataset based on the fused expression features obtained by the self-attention mechanism i , according to the fused feature x i Input it into a two-layer MLP to obtain the soft label of the expression; For all source data set expression images and all target data set expression images, expand their K-dimensional labels to 2K-dimensional labels, where K is the number of expression categories. Among them, the labels of the source data set use the original label information in the 1st to Kth dimensions, and the data in the (K + 1)th to 2Kth dimensions are set to 0. The labels of the target data set are set to 0 in the 1st to Kth dimensions and use the soft labels obtained previously in the (K + 1)th to 2Kth dimensions; Using the following formula, the fused features obtained from the source dataset are continuously input into a two-layer MLP to calculate the expression classification loss \(L_{cls}\). \(L_{cls}\) is the classification loss of the expression, and the cross-entropy loss is used to minimize the difference between the predicted classification and the true expression classification on the source dataset: where N is the number of samples in the source dataset, and y ik is the label that the i-th facial expression image belongs to the k-th type of facial expression; p ik is the probability that the fused facial expression belongs to the k-th type of facial expression; Using the following formula, the fused features of the source dataset and the target dataset are respectively input into the domain-class discriminator to calculate the domain-class adversarial loss \(L_{adv}\): Among them, S represents the data of the source dataset, and T represents the data of the target dataset; a ik , a jk is the class information that the source domain sample i or the target domain sample j belongs to the k-th class, f i , f j is the expression fusion feature, P is the probability that the predicted fusion feature belongs to the k-th class and comes from the d dataset. d = 0 means the sample comes from the source dataset, and d = 1 means the sample comes from the target dataset; Using the following formula, calculate the total loss after combining the two losses: L = αL cls + βL adv where \(\alpha\) and \(\beta\) are loss weights, initialized to 1 and 10, and the adversarial objective is to minimize \(L\); When the training reaches the minimum of the loss \(L\), stop the training and save the feature extractor and the classifier for expression recognition on the target dataset; After the training is completed, save the feature extraction model and the classifier.

6. An electronic device, characterized in that, Comprising: At least one memory and at least one processor; The memory is used to store one or more programs; When the one or more programs are executed by the at least one processor, the at least one processor implements the steps of an expression recognition method as described in any one of claims 1-4.

7. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, it implements the steps of an expression recognition method as described in any one of claims 1-4.