A Facial Expression Recognition Method Based on Label Distribution Learning
By training deep learning models to generate Gaussian function marker distribution, the problem of inaccurate marker distribution in existing expression recognition models is solved, and the accuracy and efficiency of expression recognition are improved.
Patent Information
- Application Number
- CN202211216764.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-09-30
- Publication Date
- 2025-07-29
- Estimated Expiration
- 2042-09-30
AI Technical Summary
Existing expression recognition models use Gaussian functions with fixed standard deviation to generate marker distributions, which cannot truly represent the differences between expressions of different intensity, affecting the recognition effect.
By training deep learning models, the sample classification loss is converted into standard deviation, the label distribution is calculated using the Gaussian function, and used it as the ground-truth, the feature extraction network is optimized to generate a label distribution that is more in line with the facts.
The generated mark distribution not only indicates the degree of expression, but also the intensity of expression, which improves the recognition effect of the facial expression recognition model and saves manpower and time costs.
Smart Images

Figure CN115482575B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of deep learning and facial expression recognition, and particularly relates to a facial expression recognition method based on label distribution learning. Background Art
[0002] Facial expression is one of the most natural, powerful, and common signals for humans to express emotional states and intentions, and is an important means of human communication. Facial expression recognition has received increasing attention due to its importance in some real-world applications, such as human-computer interaction, healthcare, driver fatigue detection, etc. Automatic recognition of facial expressions is a popular research direction in the field of machine learning, with important theoretical research significance and wide practical application value. As early as the twentieth century, Ekman and Friesen defined six basic emotions based on cross-cultural research: Anger, Disgust, Fear, Happiness, Sadness, and Surprise. Contempt was subsequently added as one of the basic emotions. In the past few decades, a considerable number of deep learning methods have been applied to facial expression recognition, and most of these methods use a single or several basic expressions to describe an expression image. In recent years, research has shown that real-world expressions may be ambiguous and mixed with multiple basic expressions.
[0003] The method based on label distribution learning uses multiple labels with different intensities as ground-truth to alleviate the problem of label ambiguity, is very suitable for solving the facial expression recognition problem, and has achieved remarkable results. However, since most existing expression datasets only have one-hot labels rather than label distributions, it is impractical to directly apply label distribution learning. One method is to use a Gaussian function to generate label distributions for samples. Most existing methods fix the standard deviation value in the Gaussian function (such as 0.7, 3, etc.), which will make the label distributions of the same type of expressions the same and cannot truly represent the differences between expressions with different intensities. Therefore, it is particularly important to study effective label distribution generation methods to generate more realistic label distributions for the dataset. Summary of the Invention
[0004] The present invention discloses a facial expression recognition method based on label distribution learning to improve the recognition performance of facial expressions based on deep learning.
[0005] The technical solution adopted by the present invention is as follows:
[0006] A facial expression recognition method based on label distribution learning, the method comprising the following steps:
[0007] Step 1, construct a facial expression image dataset and preprocess the facial expression image dataset: perform face detection and alignment on each image in the image dataset, and then normalize the image size (e.g., 224*224) to match the input of the image classification feature extraction network, obtaining an image sample, and set the corresponding facial expression label for each image sample;
[0008] Step 2, construct an image classification network model: sequentially connect a fully connected layer and a classification layer after the image classification feature extraction network. Among them, the output dimension of the fully connected layer is the same as the number of expression categories, and each neuron represents a class. Its output is the possibility that the input image (expression image) of the image classification network model belongs to each expression category, that is, the expression category probability of the current input image. The classification layer normalizes the expression category probability output by the fully connected layer and makes it conform to the Gaussian distribution;
[0009] Step 3, train the network parameters of the image classification network model based on a certain number of image samples. When the change amount of the classification cross-entropy loss is less than a given threshold, execute Step 4;
[0010] Step 4, calculate the classification cross-entropy loss of each image sample, and convert the classification cross-entropy loss value to calculate the labeled distribution of the corresponding expression image by applying the Gaussian function;
[0011] Step 5, use the labeled distribution of the image sample as the ground-truth label of the image sample, and re-train the network parameters of the image classification network model constructed in Step 2. During training, optimize the image classification feature extraction network with the goal of reducing the classification cross-entropy loss and the KL (relative entropy) divergence loss. That is, during training, the loss of the image classification network model is the weighted sum of the classification cross-entropy and the relative entropy divergence loss. Stop when the change amount of the loss of the image classification network model is less than a given threshold to obtain the trained image classification network model;
[0012] Step 6, normalize the size of the face image to be recognized to match the input of the image classification network model, and then input the size-normalized face image to be recognized into the trained image classification network model to obtain the facial expression recognition result of the face image to be recognized: the expression corresponding to the maximum expression category probability.
[0013] Further, the preprocessing of the facial expression image dataset also includes: using random cropping, random horizontal flipping, and random erasing to avoid overfitting.
[0014] Further, the image classification feature extraction network can select the first layer to the penultimate layer of ResNet18 and perform pre-training on the face recognition dataset (e.g., MS-Celeb-1M).
[0015] Further, the normalized expression category probabilities output by the classification layer are as follows: where p ij represents the probability that the i-th input image after normalization belongs to category j, e represents the natural base, θ k represents the probabilities of various categories output by the fully connected layer, Y represents the number of categories, and θ j represents the probability of category j output by the fully connected layer.
[0016] Further, in step 4, the classification cross-entropy loss value is converted and applied to calculate the label distribution of the corresponding expression image by using a Gaussian function, specifically as follows:
[0017] Convert the classification cross-entropy loss value into a standard deviation: where α represents a preset weight, and loss i represents the classification cross-entropy loss value of the i-th input image;
[0018] Calculate the label distribution by using the Gaussian function:
[0019]
[0020] where represents the label distribution of the input image x i (sample), that is, the degree to which category j describes the input image x i , c j represents category j, y i represents the facial expression label (true label) of the image x i , M represents a normalization factor, and
[0021] Further, in step 4, when calculating the label distribution by using the Gaussian function, the fixed expression category order of Mikels’ wheel can be adopted.
[0022] Further, in step 5, the loss of the image classification network model is:
[0023] L = (1 - λ)L C (x, y) + λL D (x, l)
[0024] where λ represents a preset weight, the cross-entropy loss the KL loss where N represents the number of image samples in one round (epochs) during training, C represents the number of categories, y i represents the true label, x represents the input image, y represents the label representation of x, and l represents the label distribution representation of x calculated in step 4.
[0025] Further, in Steps 3 and 5, the given threshold is set to 0.001.
[0026] The technical solution provided by the present invention at least brings the following beneficial effects:
[0027] (1) Automatically generate a label distribution for the expression dataset based on the Gaussian function, saving labor and time costs.
[0028] (2) Automatically generate a label distribution based on the Gaussian function. The generated label distribution not only represents the degree of expression in various expression description images, but also represents the intensity of the expression, which is more in line with the facts, is conducive to the model learning meaningful features, and improves the effect of the facial expression recognition model. BRIEF DESCRIPTION OF THE DRAWINGS
[0029] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the following drawings are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained according to these drawings.
[0030] Figure 1 is a flowchart of a facial expression recognition method based on label distribution learning provided by an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0031] To make the objectives, technical solutions, and advantages of the present invention clearer, the following will further describe the embodiments of the present invention in detail with reference to the drawings.
[0032] The present invention aims to solve the problem in the existing expression recognition model based on label distribution learning that using a univariate Gaussian function with a fixed standard deviation to generate a label distribution for expression images makes the label distributions of the same type of expressions the same, unable to truly represent the differences between expressions of different intensities, and affecting the recognition effect of the model. For this reason, the present invention proposes a facial expression recognition method based on label distribution learning, which learns the features of expression images by training a deep learning model, considers converting the sample classification loss into a standard deviation, calculates the corresponding label distribution through the Gaussian function, and the obtained label distribution not only represents the degree of various expression description samples, but also more represents the intensity of the expression, which is more in line with the facts. Subsequently, by using the generated label distribution as a kind of ground-truth to train the model, the model can learn more meaningful expression features.
[0033] Such as Figure 1As shown in the figure, the facial expression recognition method based on label distribution learning provided by the embodiment of the present invention includes: 1) Preprocessing the face image, performing face detection and alignment to obtain an expression image; 2) Inputting the expression image and extracting the features of the expression image; 3) Classifying the features and optimizing the feature extraction network with the goal of reducing the feature classification entropy; 4) Using the Gaussian function to generate a label distribution for the expression image and taking it as a kind of ground-truth; 5) Reconstructing the network model, inputting the expression image, and extracting the features of the expression image; 6) Classifying the image and optimizing the feature extraction network with the goal of reducing the cross-entropy loss and KL divergence loss; 7) When the classification loss is less than the stop iteration threshold, outputting the classification result.
[0034] As a possible implementation, the facial expression recognition method based on label distribution learning provided by the embodiment of the present invention includes the following steps:
[0035] Step 1: Construct an experimental dataset, divide the experimental dataset into a training set and a validation set according to 90% training set and 10% validation set. The dataset selected in this embodiment is the CK+ dataset (Extended Cohn-Kanade dataset);
[0036] Step 2: Perform face detection and alignment. When processing the image size of 224*224, use random cropping, random horizontal flipping, and random erasing to avoid overfitting;
[0037] Step 3: Establish a ResNet18 network model for image feature extraction, modify the fully connected layer of the feature extraction network model and a classification layer for calculating the target distribution, and perform pre-training on the face recognition dataset MS-Celeb-1M;
[0038] Step 4: Input all training set samples into the model, output the probability distribution of each sample belonging to each class, according to the formula:
[0039]
[0040] Step 5: Calculate the classification cross-entropy loss and optimize the model parameters according to the backpropagation rule;
[0041] Step 6: Calculate the change rate of the loss of this training and the loss of the previous round of training:
[0042]
[0043] where, loss pre represents the loss of the previous round of training, and loss represents the loss during the current training. If is less than 0.001, the training ends and proceeds to Step 8, otherwise proceeds to Step 5;
[0044] Step 7: Calculate the sample label distribution using the Gaussian function, and convert the sample loss value in Step 5 into a standard deviation. The calculation formula is as follows:
[0045]
[0046]
[0047] Wherein,
[0048]
[0049] Step 8: Reconstruct the model according to Step 3;
[0050] Step 9: Input all the training set samples into the model, and output the probability distribution of each sample belonging to each class;
[0051] Step 10: According to the model loss formula: L = (1 - λ)L C (x, y) + λL D (x, l), calculate the model loss, and optimize the model parameters according to the backpropagation rule;
[0052] Step 11: Calculate the change rate of the loss of this training and the loss of the previous round of training. If is less than 0.001, the training ends and enters Step 12; otherwise, enter Step 9;
[0053] Step 12: Input the validation set into the trained network and output the classification result.
[0054] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements for some of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.
[0055] The above are only some embodiments of the present invention. For those of ordinary skill in the art, without departing from the inventive concept of the present invention, several deformations and improvements can still be made, and these all belong to the protection scope of the present invention.
Claims
1. A facial expression recognition method based on label distribution learning, characterized in that It includes the following steps: Step 1: Construct a facial expression image dataset and preprocess the facial expression image dataset. Perform face detection and alignment on each image in the image dataset, and then normalize the image size to match the input of the image classification feature extraction network, obtaining an image sample, and setting a corresponding facial expression label for each image sample; Step 2: Construct an image classification network model. Connect a fully connected layer and a classification layer in sequence after the image classification feature extraction network. Among them, the output dimension of the fully connected layer is the same as the number of expression categories, and its output is the probability of the expression category of the current input image. The classification layer normalizes the probability of the expression category output by the fully connected layer and makes it conform to the Gaussian distribution; Step 3: Train the network parameters of the image classification network model based on a certain number of image samples. When the change in the classification cross-entropy loss is less than a given threshold, execute Step 4; Step 4: Calculate the classification cross-entropy loss of each image sample, and convert the classification cross-entropy loss value to calculate the labeled distribution of the corresponding expression image by applying the Gaussian function; Step 5: Use the labeled distribution of the image sample as the ground-truth label of the image sample, and retrain the network parameters of the image classification network model constructed in Step 2. During training, the loss of the image classification network model is the weighted sum of the classification cross-entropy and the relative entropy divergence loss. Stop when the change in the loss of the image classification network model is less than a given threshold to obtain the trained image classification network model; Step 6: Normalize the size of the face image to be recognized to match the input of the image classification network model, and then input the size-normalized face image to be recognized into the trained image classification network model to obtain the facial expression recognition result of the face image to be recognized: the expression corresponding to the maximum expression category probability.
2. The method according to claim 1, characterized in that, Preprocessing the facial expression image dataset also includes: using random cropping, random horizontal flipping, and random erasing to avoid overfitting.
3. The method according to claim 1, characterized in that The first layer to the penultimate layer of ResNet18 is selected for the image classification feature extraction network and pre-trained on the face recognition dataset.
4. The method according to claim 1, wherein The normalized expression category probabilities output by the classification layer are as follows: Among them, p ij represents the probability that the i-th input image after normalization belongs to category j, e represents the natural base, and θ k represents the probabilities of all categories output by the fully connected layer, Y represents the number of categories, and θ j represents the probability of category j output by the fully connected layer.
5. The method according to any one of claims 1 to 4, characterized in that, In Step 4, converting the classification cross-entropy loss value to calculate the labeled distribution of the corresponding expression image by applying the Gaussian function is specifically: Convert the categorical cross-entropy loss value into a standard deviation: where α represents a preset weight, and loss i represents the categorical cross-entropy loss value of the i-th input image; Calculating the labeled distribution using the Gaussian function: Among them, represents the label distribution of the input image x i of, c j represents the category j, y i represents the facial expression label of the image x i M represents the normalization factor, and 6. The method according to claim 1, wherein In Step 5, the sum of the weights of the classification cross and the relative entropy divergence loss is 1.
7. The method according to claim 1, wherein In Steps 3 and 5, the given threshold is set to 0.001.