A Facial Expression Recognition Method in Real Environments Based on a Spatial Distribution Loss Function
The proposed method enhances face expression recognition in real-world conditions by using a high-efficiency attention mechanism convolutional neural network with a joint loss function to improve feature extraction and classification accuracy.
Patent Information
- Application Number
- CN202111522270.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-12-13
- Publication Date
- 2025-07-15
- Estimated Expiration
- 2041-12-13
AI Technical Summary
The prior art facial expression recognition model in natural environments is difficult to effectively deal with problems such as multi-faceted head, uneven background lighting and facial occlusion, resulting in poor recognition effect.
A convolutional neural network with efficient attention mechanism is designed, and supervised learning is combined with Softmax, centerloss and the newly proposed spatial distribution loss function (SDloss). Expression features are extracted through the ResNet-18 backbone network and the ECA-Net attention mechanism module to optimize the spatial distribution of features.
It improves the accuracy and robustness of facial expression recognition, and can effectively classify expression images in complex natural environments.
Smart Images

Figure CN114187638B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to a method for facial expression recognition in a real environment based on a spatially distributed loss function, and mainly relates to the field of computer vision. Background Art
[0002] Facial expressions are the most intuitive and effective means for humans to express emotions, and expressions are also the most important non-verbal emotional expression ways of humans. Therefore, in order to achieve the purpose of true human-computer interaction, in the past ten-odd years, facial expression recognition technology has received extensive attention and research from scholars. At present, most of the research related to facial expression recognition mostly focuses on the recognition of basic emotions, such as anger, disgust, surprise, fear, sadness, happiness, and neutral. With the continuous development of face recognition technology, this technology has also been applied to multiple fields, such as teaching evaluation, medical rehabilitation, game entertainment, traffic safety, etc. Although facial expression recognition technology already has great application value, its research still faces huge challenges, such as multi-poses of the head, uneven background illumination, personalized expressions of individual expressions, local occlusion of the face, etc. The current facial expression recognition still only has good recognition effects when used in a laboratory environment with single illumination, forward head pose, and no occlusion of the face. If these models are applied to the expression recognition task in a natural environment, it is difficult to achieve good effects.
[0003] The process of facial expression recognition can be divided into two stages: extraction and representation of expression features; and expression classification. In the first stage, it is divided into two categories according to different feature extraction methods: manual visual features; and learning-based features. Commonly used manual visual features can be further divided into texture-based manual features, geometry-based manual features, and hybrid features.
[0004] In recent years, with the development of computer technology, convolutional neural networks have been widely applied to the algorithm models of facial expression recognition, mainly for extracting expression features.
[0005] Convolutional neural network is a prominent deep learning technology for automatically extracting deep features, and its application effect in facial expression recognition is significantly better than traditional methods. For any visual recognition system with a fixed set of categories, its input space can be mapped into a high-dimensional feature vector with the semantic information of the input picture. The method of extracting spatial features based on a deep convolutional neural network is to obtain the abstract semantics of the input image through the combined features from lower levels to higher levels. Then the pooling layer converts the spatial features into a single deep feature vector. Finally, the Softmax loss function is used to evaluate the probability distribution of all categories. Therefore, the performance of the facial expression recognition algorithm can be improved by establishing an embedding space with better discriminative expression features. However, the application of facial expression recognition in natural environments requires obtaining a large number of annotated images in an unconstrained environment, that is, a wild facial expression dataset. Therefore, facial expression images in the wild environment often show significant intra-class variability and inter-class similarity. This is manifested as a small inter-class distance and a large intra-class distance in the embedding space of the samples. This phenomenon seriously affects the classification effect of facial expression recognition. Summary of the Invention
[0006] In view of the above deficiencies of the prior art, the present invention proposes a method for facial expression recognition in a real environment based on a spatial distribution loss function, which can...
[0007] To achieve the above object, the technical solution of the present invention is: characterized in that
[0008] S1 Preprocess the images of the facial expression dataset;
[0009] S2 Design a convolutional neural network with an efficient attention mechanism;
[0010] S3 Deploy a joint loss function for supervised learning during the learning process of the efficient attention mechanism network. This loss function is composed of Softmax loss, center loss, and SD loss, and its formula is as follows:
[0011] L = L s + λL c + γL SD
[0012] where λ = 3 and γ = 5;
[0013] S4 Divide the facial expression dataset into a training set, a validation set, and a test set; pre-train the convolutional neural network designed above;
[0014] S5 Fine-tune the parameters of the training model using the facial expression dataset to obtain the final facial expression recognition model;
[0015] S6 Use the final facial expression recognition model for facial expression recognition.
[0016] Preferably, in step S2, the specific process of designing the efficient attention mechanism convolutional neural network is as follows:
[0017] (S1 This network uses ResNet-18 as the backbone network and embeds the attention mechanism module into each BasicBlock of ResNet-18;
[0018] (S2 The attention mechanism module uses the feature map generated by the convolutional neural network to generate an attention map; and then uses the attention map to regenerate the feature map that has a significant impact on facial expression recognition in the feature map generated by the convolutional neural network.
[0019] Preferably, the attention mechanism module adopts ECA-Net, and its detailed process is as follows:
[0020] ECA-Net first uses global average pooling on the input feature map to compress and extract the features from the two-dimensional matrix to a single value, and then generates channel weights by performing a fast one-dimensional convolution of size K without reducing the dimension, obtaining the correlation dependencies between each channel. Finally, the generated weights are weighted to the original input feature map by multiplication, and the calibration of the features in the channel space is completed by weighting the features extracted by ECA-Net and the original features;
[0021] ECA-Net performs local interaction through K-nearest neighbors, effectively reducing the computational amount and complexity of interacting across all channels. It generates weights for each feature channel through a one-dimensional convolution of size K to obtain the correlation between feature channels, that is:
[0022] ω = σ(Conv1D k (y))
[0023] In the formula, conv1D represents one-dimensional convolution, and K determines the coverage range of cross-channel local interaction. Since the size of the channel dimension C is proportional to k, the corresponding exponential function relationship is obtained:
[0024] C = φ(k) = 2 γ*k-b
[0025] Therefore, given the channel dimension C in this article, the size of the parameter K is adaptively determined through the following functional relationship:
[0026]
[0027] In the formula, odd is the nearest odd number t; and here γ and b are set to 2 and 1 respectively; the mapping function ψ is that the larger the channel dimension, the larger k is, and the larger the range of cross-channel local interaction is.
[0028] Preferably, in step S3, the specific meaning of the joint loss function deployed in the efficient attention mechanism network is:
[0029]
[0030] L=L s +λL c +γL SD
[0031] Where L s is the Softmax loss, L c is the center loss, L SD is the newly proposed spatial distribution loss.
[0032] Preferably, in step S3, the facial expression datasets used in the present invention are RAD-DB and AffectNet data.
[0033] This method uses an efficient attention network to extract subtle and deep features of facial expressions. The newly proposed spatial distribution loss function is used to shorten the distance between similar categories and push away the distance between different categories in the high-dimensional feature space to achieve a spatial distribution that is more conducive to classification. Then the classifier is used to calculate its probability distribution, and the category with the highest probability is taken as the predicted value of the image. After training, a facial expression recognition model is obtained to achieve effective classification of the expression images to be classified. BRIEF DESCRIPTION OF THE DRAWINGS
[0034] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the drawings required for use in the description of the embodiments will be briefly introduced below. Obviously, the drawings described below are only two of the drawings of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative work.
[0035] Figure 1 It is an overall framework diagram of an embodiment of the present invention;
[0036] Figure 2 This is a framework diagram of ECA_Net according to an embodiment of the present invention. DETAILED DESCRIPTION
[0037] The technical solutions in the present invention will be described clearly and completely below in conjunction with the accompanying drawings. Obviously, the described embodiments are only preferred embodiments of the present invention, not all embodiments. Based on the embodiments in the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.
[0038] Example
[0039] like Figure 1As shown in the figure, the embodiments of the present invention include the following steps:
[0040] S1 Preprocess the images in the facial expression dataset;
[0041] S2 Design a convolutional neural network with an efficient attention mechanism. Embed the attention mechanism module ECA-Net into each BasicBlock in ResNet-18. The attention mechanism module uses the feature map generated by the convolutional neural network to generate an attention map. Then, use the attention map to regenerate the feature map that has a significant impact on facial expression recognition in the feature map generated by the convolutional neural network. Its structural diagram is as Figure 1 shown.
[0042] S3 Deploy a joint loss function for supervised learning during the learning process of the efficient attention mechanism network. This loss function consists of softmax loss, center loss, and SD loss. Its formula is as follows:
[0043] L = L s + λL c + γL D
[0044] S4 Divide the facial expression dataset into a training set, a validation set, and a test set. Pre-train the convolutional neural network designed above.
[0045] S5 Fine-tune the parameters of the training model using the facial expression dataset to obtain the final facial expression recognition model.
[0046] S5 Use the final facial expression recognition model for facial expression recognition.
[0047] In step S2, the specific process of the designed efficient attention mechanism network is as follows:
[0048] (S1 This network uses ResNet-18 as the backbone network and embeds the attention mechanism module into each BasicBlock of ResNet-18.
[0049] (S2 The attention mechanism module uses the feature map generated by the convolutional neural network to generate an attention map. Then, use the attention map to regenerate the feature map that has a significant impact on facial expression recognition in the feature map generated by the convolutional neural network. The attention mechanism module adopts ECA-Net, and its detailed process is as follows:
[0050] ECA-Net first uses global average pooling on the input feature map to compress and extract features from a two-dimensional matrix into a single value. Then, without reducing the dimension, it generates channel weights by performing a fast one-dimensional convolution of size K to obtain the correlation dependencies between channels. Finally, the generated weights are multiplied and weighted onto the original input feature map to complete the calibration of features in the channel space between the features extracted by ECA-Net and the original features.
[0051] ECA-Net performs local interactions through K-nearest neighbors, effectively reducing the computational amount and complexity of interactions across all channels. It generates weights for each feature channel through a one-dimensional convolution of size K to obtain the correlation between feature channels, that is
[0052] ω = σ(Conv1D k (y))
[0053] conv1D in the formula represents one-dimensional convolution, and K determines the coverage range of cross-channel local interactions. Since the size of the channel dimension C is proportional to k, the corresponding exponential function relationship is obtained
[0054] C = φ(k) = 2 γ*k-b
[0055] Therefore, given the channel dimension C in this paper, the size of parameter K is adaptively determined through the following functional relationship
[0056]
[0057] In the formula, odd is the nearest odd number t. And here γ and b are set to 2 and 1 respectively. The mapping function ψ is such that the larger the channel dimension, the larger k is, and the larger the range of cross-channel local interactions.
[0058] In step S3, the specific meaning of the joint loss function deployed in the efficient attention mechanism network:
[0059] Suppose a training batch (mini-batch) is given, which contains m training samples, and Y is the label of the training samples. The feature map x is the output of the convolutional neural network.
[0060] Deep convolutional neural networks usually use Softmax for the training process of supervised multi-classification tasks. Softmax can effectively separate deep features of different classes. Softmax is used after the fully connected layer. The result obtained from the last fully connected layer is converted into probability values.
[0061] z i =W T x i +B
[0062] where \(W = [w_1, w_2, \ldots, w K \in \mathbb{R} d×K , B = [b_1, b_2, \ldots, b K \in \mathbb{R} K×1 , are the class weights and bias parameters of the last fully connected layer.
[0063] This probability distribution \(P(y = j|x i )\) is calculated by all classes through the softmax function.
[0064] Finally, the cross-entropy loss function calculates the difference between the predicted value and the true label \(y i to form the following Softmax loss function \(L s :\)
[0065]
[0066] Researchers in deep feature learning proposed the center loss function (center loss). The center loss function (center loss) is a typical method proposed in deep metric learning. It calculates the similarity between the deep features of samples and those in the same class of their own class, and this similarity is measured using the Euclidean distance. The center loss function minimizes the Euclidean distance between the deep features and their corresponding class centers to supervise the training process. Its purpose is to divide the embedding space into \(K\) clusters to solve the \(K\)-classification problem. Suppose a training batch containing \(M\) samples is given, and the deep feature vector of the \(i\)-th sample is represented as \(x i = [x i1 , x i2 , \ldots, x id T \in \mathbb{R} d , the labels of each class are \(y = \{1, \ldots, K\}\) and its class center \(c yi = [c yi1 , c yi2 , \ldots, c yid T \in \mathbb{R} d . The center loss function can be expressed as the following function
[0067]
[0068] The center loss function only considers the distance between the deep features of a sample and the deep features of samples of the same category, ignoring the distance between the deep features of a sample and the deep features of samples of different categories. Moreover, the wild dataset shows an extremely unbalanced sample distribution. Deploying the center loss function on the training model will cause the phenomenon of small sample overlap, resulting in a decline in the performance of facial expression recognition. Therefore, on this basis, this paper takes into account both the within-class distance and between-class distance of sample deep features as well as the problem of extremely unbalanced sample distribution, and proposes a new spatial distribution loss function.
[0069] In this formula, m is the total number of samples in a sample batch, and x i is the high-dimensional sample feature output by the convolutional neural network.
[0070]
[0071] Here, a i is the high-dimensional feature of the sample after passing through the convolutional neural network, and a pi represents the high-dimensional features of samples with the same label as a i . a qj represents the high-dimensional features of samples with different categories from a i . In the facial expression recognition model, three loss functions are used simultaneously, as follows:
[0072] L = L s + λL c + γL SD
[0073] When Centerloss calculates the distance between the deep features of a sample and its class center, the selection of the class center is random. However, the newly proposed loss function in this paper sets T class centers and takes the average of the deep features of the sample to the T within-class centers as the within-class distance, effectively avoiding the error caused by inappropriate randomly selected class centers.
[0074] In step S3, the facial expression datasets used in the present invention are the RAD-DB and AffectNet data.
[0075] The RAD-DB dataset contains approximately 30,000 facial expression images downloaded from the Internet. This dataset contains two parts: a single-label subset (basic expressions) and a double-label subset (compound expressions). This paper uses the single-label subset with seven basic expression categories. The training set of this subset contains 12,271 images, and the test set contains 3,068 images.
[0076] The AffectNet dataset is the largest publicly available FER dataset in the wild, with 450,000 facial images obtained from the Internet and manually annotated with categorical expressions and dimensional affect (valence and arousal). In our experiments, we used 280,000 training images and 3,500 images validation set consisting of six basic expressions and neutral expressions.
[0077] The RAD-DB dataset is used as the input data of the efficient attention network, and the joint loss function is used to supervise the training of the efficient attention mechanism network.
[0078] L=L s +λL c +γL SD Where λ=3,γ=5
[0079] Use the trained algorithm model to classify the facial expressions to be classified into anger, disgust, surprise, fear, sadness, happiness, and neutrality.
[0080] This method uses an efficient attention network to extract subtle and deep features of facial expressions. The newly proposed spatial distribution loss function is used to shorten the distance between similar categories and push away the distance between different categories in the high-dimensional feature space to achieve a spatial distribution that is more conducive to classification. Then the classifier is used to calculate its probability distribution, and the category with the highest probability is taken as the predicted value of the image. After training, a facial expression recognition model is obtained to achieve effective classification of the expression images to be classified.
[0081] The above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent substitutions, improvements, etc. made within the spirit and principle of the present invention should be included in the protection scope of the present invention.
Claims
1. A face expression recognition method in a real environment based on a spatially distributed loss function, characterized in that S1 Preprocess the images in the face expression dataset; S2 Design a convolutional neural network with an efficient attention mechanism; S3 Deploy a joint loss function for supervised learning during the learning process of the efficient attention mechanism network. This loss function consists of Softmax loss, center loss, and SD loss, and its formula is as follows: L = L s + λL c + γL SD where λ = 3, γ = 5, and L s is the Softmax loss, and L c is the center loss; The specific meaning of deploying the joint loss function in the efficient attention mechanism network: L SD is the newly proposed spatial distribution loss; a i is the high-dimensional feature of the sample after passing through the convolutional neural network, a pi table and a i high-dimensional features of samples with the same label; a qj represents and a i high-dimensional features of samples of different classes; S4 Divide the face expression dataset into a training set, a validation set, and a test set; pre-train the convolutional neural network designed above; S5 Fine-tune the parameters of the training model using the face expression dataset to obtain the final face expression recognition model; S6 Use the final face expression recognition model for face expression recognition.
2. The method for facial expression recognition in a real environment based on a spatially distributed loss function according to claim 1, wherein: In step S2, the specific process of designing the convolutional neural network with an efficient attention mechanism is as follows: (S1 This network uses ResNet-18 as the backbone network and embeds the attention mechanism module into each BasicBlock of ResNet-18; (S2 The attention mechanism module uses the feature map generated by the convolutional neural network to generate an attention map; then uses the attention map to regenerate the feature map that has a significant impact on face expression recognition in the feature map generated by the convolutional neural network.
3. The method for facial expression recognition in a real environment based on a spatially distributed loss function according to claim 2, characterized in that: The attention mechanism module adopts ECA-Net, and its detailed process is as follows: ECA-Net first uses global average pooling on the input feature map to compress and extract the features from a two-dimensional matrix to a single value, and then generates channel weights by performing a fast one-dimensional convolution of size K without reducing the dimension, obtaining the correlation dependencies between channels. Finally, the generated weights are multiplied and weighted to the original input feature map to complete the calibration of the features in the channel space by weighting the features extracted by ECA-Net and the original features; ECA-Net performs local interaction through K-nearest neighbors, effectively reducing the computational amount and complexity of interacting across all channels. A weight is generated for each feature channel through a one-dimensional convolution of size K to obtain the correlation between feature channels, that is: ω=σ(Conv1D k (y)) conv1D in the formula represents one-dimensional convolution, and K determines the coverage range of cross-channel local interaction. Since the size of the channel dimension C is proportional to k, the corresponding exponential function relationship is obtained: C = φ(k) = 2 γ*k-b Given the channel dimension C, the size of parameter K is adaptively determined through the following functional relationship: In the formula, odd is the nearest odd number t; and here γ and b are set to 2 and 1 respectively; the mapping function ψ is such that the larger the channel dimension, the larger k is, and the larger the range of cross-channel local interaction.
4. The face expression recognition method in a real environment based on a spatially distributed loss function according to claim 1, characterized in that: In step S3, the specific meaning of deploying the joint loss function in the efficient attention mechanism network:
5. A method for facial expression recognition in a real environment based on a spatially distributed loss function according to claim 1, characterized in that: In step S3, the face expression datasets used in the present invention are RAD-DB and AffectNet data.