An expression recognition method and system

By introducing a new triple loss function into the expression recognition method, adjusting the distance between negative expression classes, the problem of low accuracy of negative expression recognition is solved, and the accuracy of negative expression recognition is improved.

CN114529969BActive Publication Date: 2025-05-30CHONGQING UNIV OF POSTS & TELECOMM
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202210061916.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-01-19
Publication Date
2025-05-30
Estimated Expiration
2042-01-19

AI Technical Summary

Technical Problem

In the existing expression recognition methods, the accuracy of negative expressions is significantly lower than that of positive expressions. The main reason is that the negative expression samples are relatively small and the characteristics are similar, which makes the neural network unable to learn strongly robust features.

Method used

A new triple loss function is proposed. By controlling the class source of triples, the distance between negative expression classes is adjusted so that it tends to be equal to the positive expression classes, thereby enhancing the discrimination of negative expression characteristics.

Benefits of technology

By adjusting the inter-class distance of negative expressions, the accuracy of recognition of negative expressions is improved, and the difference in accuracy between negative expressions and positive expressions is narrowed.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114529969B_ABST
    Figure CN114529969B_ABST
Patent Text Reader

Abstract

The present invention claims protection for a facial expression recognition method and system, belonging to the technical field of biometric recognition. The method includes the steps of: obtaining a face expression picture sample and a corresponding category label; performing face detection and face alignment according to the face expression sample; establishing a deep neural network model, sending the face expression picture into the deep neural network model to extract features, and obtaining expression features; selecting multiple triplets according to the expression category label, and the three samples included in each triplet come from different categories; calculating a first loss value according to the expression features and the triplets; calculating cross-entropy loss as a second loss value according to the true category label and the obtained expression features; sending the expression features into a classifier for classification, and outputting a classification result. The main purpose of the present invention is to readjust the inter-class distance between negative expressions and improve the problem of poor negative expression recognition effect.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of biometric recognition, and particularly relates to an expression recognition method and system. Background Art

[0002] Facial expression recognition is to analyze the input facial images to determine which type of expression the currently input facial image belongs to. Common expression recognition methods usually classify seven basic expressions, including: happy, surprised, neutral, fearful, disgusted, angry, and sad.

[0003] Most of the existing expression recognition methods use deep learning convolutional neural networks to extract facial expression features and classify them. The obvious defect common to current methods is that the accuracy of negative expressions is significantly lower than that of positive expressions. The so-called positive expressions are expressions with positive emotions, including happy and surprised, and negative expressions are expressions with negative emotions, including fearful, disgusted, angry, and sad.

[0004] The reasons for the poor effect of negative expressions are as follows:

[0005] (1) The number of negative expression samples is relatively small. This problem exists in almost all expression datasets. The number of negative expression samples is less than that of positive expression samples. Especially for large datasets, the difference in the number of positive and negative expression samples is very significant. This reason causes the neural network to be unable to learn strongly robust features from the limited negative expressions, and ultimately leads to the neural network being unable to accurately judge negative expressions during the test process.

[0006] (2) The facial features between negative expressions are similar. The change of expression is composed of the movements of multiple local positions on the face, and the local movements corresponding to negative expressions overlap. For example, both anger and disgust include the action of frowning. Similar facial movements, that is, similar features, cause the neural network to be unable to accurately judge which type of expression the current feature belongs to.

[0007] In reality, when people have negative emotions, they are more likely to act overly and impulsively. In practical applications, assuming the monitoring of the driver's emotions or the monitoring of the patient's state, in these applications, others need to intervene when negative expressions appear in the observed person. Therefore, accurately recognizing negative expressions is particularly important.

[0008] After retrieval, the application publication number is CN107358169A, a facial expression recognition method and a facial expression recognition device, including: constructing and training an emotion recognition model based on a convolutional neural network; inputting a face image to be recognized into the emotion recognition model to output an emotion category of the face image, where the emotion category includes one of a positive emotion, a negative emotion, and a neutral emotion; obtaining an expression recognition model corresponding to the emotion category; and inputting the face image into the expression recognition model to output an expression category of the face image. The present invention recognizes facial expressions in a hierarchical manner, selects different expression recognition models according to different emotions, reduces the content that each recognition model needs to remember, reduces the computational complexity of the entire expression recognition process, and improves the computational efficiency.

[0009] The method disclosed in Patent CN107358169A first makes a large-category discrimination of positive, negative, and neutral emotions, and then matches different recognition models according to different large categories. Since the reason for the poor recognition effect of negative expressions is that the features between expressions are similar, and this patent does not perform targeted operations on negative expressions, but only divides out an independent recognition model. However, in this independent recognition model, the features of negative expressions are still similar, and the recognition effect of negative expressions has not been improved. The present invention designs a new triplet loss function, and can adjust the inter-class distance of negative expressions by controlling the category source of the triplet. This loss function can increase the inter-class distance of negative expressions, thereby enhancing the discriminability of negative expression features.

[0010] The application publication number is CN111353390A, a micro-expression recognition method based on deep learning, which includes the following steps: 1: Cropping video data containing expression actions, video frame division, and extraction; 2: Performing preprocessing such as face alignment, face cropping, and normalization on the extracted expression sequence; 3: Performing data augmentation operations on the obtained data set; 4: Building a neural network model; 5: Dividing all facial expression data into a training set and a test set according to a ratio; 6: Using the test set to test the model, and outputting information such as recognition accuracy, recognition time, and error. When the recognition rate meets the requirements, select the current model.

[0011] The method of Patent CN111353390A is to improve the neural network model structure, but the focus is on the overall expression recognition effect, and does not focus on the current pain points. The present invention focuses on the problem of poor negative expression effect and proposes a loss function specifically applicable to distinguishing similar classes. And the present invention does not need to change the neural network model structure, can directly use the classic network structure, and only needs to record the feature values before the fully connected layer. The method is simple and easy to operate. Summary of the Invention

[0012] The present invention aims to solve the problems of the above prior art, and provides an expression recognition method and system. The technical solution of the present invention is as follows:

[0013] An expression recognition method, comprising the following steps:

[0014] Obtain a face expression picture sample and a corresponding class label; perform face detection and face alignment according to the face expression sample;

[0015] Establish a deep neural network model, input the face expression picture into the deep neural network model to extract features, and obtain expression features;

[0016] Select multiple triplets according to the expression class label, and the three samples included in each triplet come from different classes;

[0017] Calculate a first loss value according to the expression features and the triplets;

[0018] Calculate the cross-entropy loss as a second loss value according to the true class label and the obtained expression features;

[0019] Input the expression features into a classifier for classification, and output a classification result.

[0020] Further, the face expression picture sample and the corresponding class label specifically include:

[0021] Use a camera to record a video of a face showing an expression;

[0022] Observe the time when the face shows an expression in the video, extract the frame corresponding to this time and save it as an image to obtain a face expression picture sample, and label the class of the expression.

[0023] Further, the establishment of the deep neural network model, inputting the face expression picture into the neural network model to extract features, and obtaining expression features specifically include:

[0024] Create an 18-layer residual network model, and initialize the parameters included in the model, as well as parameters such as the neural network learning rate and optimizer;

[0025] Input the expression sample into the deep neural network to obtain the eigenvalue before the last layer of the deep neural network;

[0026] Input the eigenvalue before the last layer into the last fully connected layer of the model to obtain the finally output eigenvalue.

[0027] Further, the selection of multiple triplets according to the expression class label, and the three samples included in each triplet come from different classes specifically include:

[0028] Traverse the entire set of labels. When the three samples being traversed belong to different categories, store the subscripts corresponding to the current three samples in a set to obtain a set of triples storing subscripts.

[0029] To ensure that the distances between the three sample features in each triple are unbalanced, filter out combinations where all three samples belong to positive expressions or negative expressions from the set of triples obtained in the previous step. The features of negative expressions are similar in themselves, and the method for measuring feature similarity is the distance between features. The closer the distance, the more similar, that is, the distances between negative expressions will be closer. The significance of ensuring the unbalanced distances of the three samples in the triple is that, assuming the currently selected triple contains two negative expressions and one positive expression, then the distance between the positive expression and the negative expressions will be relatively far, while the distance between the two negative expressions is relatively close, which is the so-called distance imbalance. In this case, through subsequent control of the loss function, the distance between the two negative expressions can be made to become farther, numerically approaching the distance between the positive expression and the negative expressions, achieving the purpose of readjusting the distance between negative expressions, increasing the distance between negative expressions, and making the inter-class distance between positive and negative expressions tend to be equal to the inter-class distance between negative-negative expressions. However, if the currently selected triple all belongs to negative expressions, then there is no obvious difference in the distances between the samples at the beginning, so it is difficult to make the inter-class distance between negative-negative expressions reach the inter-class distance between positive-negative expressions. The same is true when the triple all belongs to positive expressions. This is the purpose of not selecting triples that all belong to negative expressions or positive expressions.

[0030] Furthermore, calculating the first loss value according to the expression features and the triples specifically includes:

[0031] Calculate a loss value for each triple in the set of triples. Extract the corresponding three features from the expression features obtained before the last layer of the neural network according to the subscripts stored in the triple. The three features can form three Euclidean distances. Each time, randomly select two of the distances, making the ratio of the two Euclidean distances tend to 1. Therefore, take the square of the value obtained by subtracting 1 from the ratio of the two Euclidean distances as the loss value, and finally take the average of the loss values calculated for multiple triples as the final value of the first loss.

[0032] Furthermore, take out the corresponding three 512-dimensional eigenvalue x 1 、x 2 、x 3 from the features obtained before the last layer of the neural network according to the subscripts stored in each triple in the obtained set of triples. Three Euclidean distances can be obtained through the three features. Each time, randomly select two of the Euclidean distances d1, d2. The Euclidean distance calculation formula is as follows:

[0033]

[0034] Using the ratio of d1 and d2 approaching 1 as the loss function, the purpose is to make the distances d1 and d2 between two classes approach equality. A loss value is calculated for each triple, and finally the average value is taken. The formula of the loss function is as follows, where T is the number of triples:

[0035]

[0036] This loss function can only make the two distances among the three samples approach equality within each triple. However, different two distances are randomly selected in different triples. Therefore, overall, the three distances will all approach equality, and the result will be that all the distances between classes approach equality.

[0037] Furthermore, calculating the cross-entropy loss based on the true class label and the facial expression features obtained from the output of the last layer of the neural network as the second loss value specifically includes:

[0038] Assume that the output value of the i-th facial expression sample after the last layer of the neural network is P i =(p i1 , p i2 , …, p iM ), and calculate the cross-entropy loss jointly with its true label Y i =(y i1 , y i2 , …, y iM ). Among them, M represents the number of facial expression categories. In the representation of the true label, assume that this sample belongs to the j-th class. Then, the value of y i in Y ij is 1, and the rest of the values are 0. The calculation formula is as follows, where N is the number of samples:

[0039]

[0040] According to the orders of magnitude of the two loss values, adjust the orders of magnitude of the two loss values to be the same through the coefficient λ and then add them. The calculation formula is as follows:

[0041] Loss = λ * Loss1 + Loss2

[0042] Calculate the predicted class according to the output P i of the network model. The subscript corresponding to the maximum value is the predicted class, and finally output the predicted class.

[0043] An facial expression recognition system adopting the method described in any one of the above, which includes:

[0044] An input module for inputting a facial expression image into the facial expression recognition system;

[0045] An expression acquisition module, which is used to process the input multiple facial expression images, perform face detection and face alignment on the input picture samples, and obtain facial expression picture samples;

[0046] A normalization processing module, which is used to perform normalization processing on the facial expression picture samples so that the sizes of the facial expression pictures are the same, and obtain normalized expression picture samples;

[0047] A format conversion module, which is used to convert the obtained normalized facial expression pictures into the tensor format required by the neural network model;

[0048] A model management module, which is used to create, manage and save the neural network model;

[0049] A feature extraction module, which is used to extract the features corresponding to the expression samples through the neural network model;

[0050] A combination generation module, which is used to generate expression triples that meet the requirements;

[0051] A loss calculation module, which is used to calculate two loss values. The first loss value is calculated according to the triples, and the second loss value is to calculate the cross-entropy loss value according to the true labels. Finally, the two losses are combined, and the network parameters are updated by backpropagating the gradient according to the loss;

[0052] A prediction category module, which is used to predict the classification result according to the features extracted by the feature extraction module and output the prediction result.

[0053] Furthermore, the combination generation module is used to select the triples that meet the requirements according to the sample labels. Each sample in a batch of samples corresponds to a unique subscript. The selection result of the triples is the subscripts of three samples from different categories, and the three samples do not belong to negative expressions at the same time or belong to positive expressions at the same time. The purpose of storing the subscripts is to facilitate subsequent loss calculation and extract the corresponding features according to the subscripts;

[0054] The loss calculation module is used to calculate two loss functions, and add the two values after a certain adjustment. The adjustment is mainly to make the orders of magnitude of the two loss values consistent and ensure that their weights are approximately equal, that is, to ensure that both play a role.

[0055] The advantages and beneficial effects of the present invention are as follows:

[0056] The present invention aims to improve the problem of poor recognition effect of negative expressions in the existing facial expression recognition method, and provides a facial expression recognition method and system. By controlling the distance of negative expressions in the feature space, the feature discriminability is enhanced, and the difference in accuracy between negative expressions and positive expressions is reduced.

[0057] In the traditional expression recognition method based on neural network, the distances between positive expressions and between positive and negative expressions are relatively far, while the distances between negative expressions are relatively small, that is, there is a problem of unbalanced distances between different categories, resulting in a relatively low accuracy rate for negative expressions. The present invention aims to improve this problem.

[0058] The main innovation of the paper Learning Informative and Discriminative Features for FacialExpression Recognition in the Wild is to improve the center loss function. The loss function designed is to minimize the distance between each sample and the center of its class and maximize the distance to the center of other classes. The loss function designed by patent CN113887325A is the distance between each sample and the center of its class. The difference between the present invention and the above two methods is that the loss functions of the above two methods have the same restrictions on all categories, and the result of the adjustment of this loss function is that the distance between classes is still unbalanced. The main advantage of the present invention is that the distance between three samples in a triplet is adjusted each time, and the category source of the three selected samples is controlled. The feature of the present invention is that this method does not perform the same operation on all categories, but focuses on categories that are originally close in distance, that is, between negative and positive expressions, so it can mainly improve the problem of poor recognition effect of negative expressions in a targeted manner. BRIEF DESCRIPTION OF THE DRAWINGS

[0059] Figure 1 The present invention provides a flow chart of a preferred embodiment of an expression recognition method;

[0060] Figure 2 This is the structure diagram of the expression recognition system. DETAILED DESCRIPTION

[0061] The following will describe the technical solutions in the embodiments of the present invention in detail in conjunction with the accompanying drawings in the embodiments of the present invention. The described embodiments are only part of the embodiments of the present invention.

[0062] The technical solution of the present invention to solve the above technical problems is:

[0063] Embodiment 1:

[0064] like Figure 1 As shown, the expression recognition method of the present invention comprises the following steps:

[0065] Step 1: Use a camera to collect clear facial expression videos of multiple people, each of whom has seven basic expressions. Extract frames from the collected video, extract frames with expressions in the video, and obtain valid facial expression image samples;

[0066] Step 2: Perform face detection and facial landmark detection on the facial expression picture samples obtained in Step 1. The specific method is to use the algorithm in the Dlib library, which can initially obtain the face region and facial landmarks, and then perform face alignment according to the similarity transformation matrix method to make the nose of all face images in the middle position of the picture;

[0067] Step 3: Label the categories of the facial expression picture samples obtained in Step 2. Use 0 - 6 to represent seven expressions respectively to obtain a complete expression dataset for subsequent training of the network model;

[0068] Step 4: Create a deep neural network model and input all the expression picture samples into the model. The specific steps are as follows:

[0069] (1) Perform normalization processing on the facial expression picture samples obtained in Step 2, set the size of all samples to 224×224 to obtain normalized expression picture samples;

[0070] (2) Divide the normalized expression picture samples into multiple sets, each set contains different types of expressions, and convert the samples into the tensor format required by the neural network to obtain multiple tensor sets;

[0071] (3) Create a 18 - layer residual network model and load the network model parameters trained on a large - scale face dataset into the current model to obtain an initialized neural network model;

[0072] (4) Input multiple tensor sets into the neural network model in sequence for feature extraction to obtain corresponding expression features. The i - th expression sample obtains a feature value x i =(x i1 ,x i2 ,…,x in ), where n = 512, and each sample obtains a 512 - bit feature;

[0073] (5) Input the expression features obtained in the previous step into the last fully - connected layer of the neural network to obtain classification information P i =(p i1 ,p i2 ,…,p iM ), where M is the number of categories, and each sample corresponds to seven values, respectively representing the possibility that the sample belongs to seven categories;

[0074] Step 5: Select triplets according to the labels corresponding to the current tensor set. The specific operations are as follows:

[0075] (1) Traverse the entire label set. When the three samples traversed all belong to different categories, store the subscripts corresponding to the current three samples in a set to obtain a set of triples storing subscripts;

[0076] (2) To ensure that the distances between the three sample features in each triple are unbalanced, filter out the combinations where all three samples belong to positive expressions or negative expressions from the set of triples obtained in the previous step;

[0077] Step 6: According to the subscripts stored in each triple in the set of triples obtained in Step 5, extract the corresponding three 512-dimensional eigenvalue features from the features obtained in the third step of Step 4. Three Euclidean distances can be obtained from the three features. Each time, randomly select two of the Euclidean distances d1 and d2. The Euclidean distance calculation formula is as follows:

[0078]

[0079] Use the ratio of d1 and d2 approaching 1 as the loss function. Calculate a loss value for each triple, and finally take the average. The loss function formula is as follows, where T is the number of triples:

[0080]

[0081] This loss function only controls two distances among the three samples within each triple, but randomly selects different two distances in different triples. Therefore, overall, it controls the three distances, that is, controls the distances between all categories;

[0082] Step 7: According to the final classification information P obtained in the fourth step of Step 4 i , and the true label Y of the expression sample i =(y i1 ,y i2 ,…,y iM ) jointly calculate the cross-entropy loss. In the representation of the true label, if the sample belongs to the j-th category, then the value of y i in Y ij is 1, and the rest are 0. The calculation formula is as follows, where N is the number of samples:

[0083]

[0084] Step 8: According to the orders of magnitude of the two loss values, adjust the orders of magnitude of the two loss values to be the same through the coefficient λ for addition. The calculation formula is as follows:

[0085] Loss = λ * Loss1 + Loss2

[0086] Step 9: According to the output P of the network model iCalculate the predicted category, where the subscript corresponding to the maximum value is the predicted category, and finally output the predicted category.

[0087] Example 2:

[0088] As Figure 2 shown, an expression recognition system includes:

[0089] An input module for inputting a facial expression image and a corresponding category label into the expression recognition system;

[0090] An expression acquisition module for processing the input multiple facial expression images, performing face detection and face alignment on the input picture samples, and obtaining facial expression picture samples;

[0091] A normalization processing module for normalizing the facial expression picture samples so that the sizes of the facial expression pictures are the same, and obtaining normalized expression picture samples;

[0092] A format conversion module for converting the obtained normalized facial expression pictures into the tensor format required by the neural network model;

[0093] A model management module for creating, managing, and saving a neural network model;

[0094] A feature extraction module for extracting features corresponding to the expression samples through the neural network model;

[0095] A combination generation module for generating expression triples that meet the requirements;

[0096] A loss calculation module for calculating two loss values. The first loss value is calculated according to the triples, and the second loss value calculates the cross-entropy loss value according to the true labels, and finally combines the two losses;

[0097] A predicted category module for predicting the classification result according to the features extracted by the feature extraction module and outputting the prediction result.

[0098] The expression acquisition module is used to obtain the facial expression region from the input picture, and perform face detection and face alignment on the picture samples using the Dlib library according to the facial expression picture samples to obtain normalized facial expression samples; the normalization processing module is used to normalize the facial expression samples, normalize the facial expression picture sample data, and set the sizes of all samples to 224×224;

[0099] The model management module includes a model creation unit, a model parameter loading unit, and a model storage unit. The specific steps are as follows: First, create a 18-layer residual network structure in the model creation unit. Second, in the model parameter loading unit, first obtain the network parameters from the model trained on a large face dataset and load the pre-trained parameters into the created model. Finally, the model storage unit is responsible for saving the model parameters that meet the conditions during multiple predictions and trainings;

[0100] The feature extraction module is used to obtain the features corresponding to the input expression samples. By inputting the samples into the neural network model, the features corresponding to the samples can be obtained, which mainly includes two parts. The first is the output before the last layer of the network, which is used to calculate the first loss value. The second is the output of the last layer of the network, which is used to calculate the cross-entropy loss and the classification result;

[0101] The combination generation module is used to select the triplets that meet the requirements. First, select all the triplets that can be formed by samples of different categories according to the labels of the samples, and delete the combinations where all three belong to positive expressions or all belong to negative expressions from them to obtain a set of triplets, where each triplet stores the subscript corresponding to the selected sample;

[0102] The loss calculation module is used to calculate two loss functions. The first is calculated through the set of triplets obtained by the combination generation module. Calculate a loss for each triplet, and finally take the average of all triplets as the first loss value. The second loss is the cross-entropy loss, which is calculated by the output of the last layer of the network and the true labels. Finally, the two values are added after a certain adjustment, and the adjustment is mainly to make the orders of magnitude of the two loss values consistent;

[0103] The prediction category module is used to predict the final category, use the second part of the features obtained by the feature extraction module for prediction, and output the final prediction result.

[0104] The systems, devices, modules, or units described in the above embodiments can be specifically implemented by computer chips or entities, or by products with certain functions. A typical implementation device is a computer. Specifically, the computer can be, for example, a personal computer, a laptop computer, a cellular phone, a camera phone, a smart phone, a personal digital assistant, a media player, a navigation device, an email device, a game console, a tablet computer, a wearable device, or any combination of these devices.

[0105] It should also be noted that the term "comprising", "including" or any other variation thereof is intended to cover non-exclusive inclusion, so that a process, method, commodity or device comprising a series of elements not only includes those elements, but also includes other elements not expressly listed, or also includes elements inherent to such process, method, commodity or device. Without further limitation, an element defined by the statement "comprising an..." does not exclude the presence of additional identical elements in the process, method, commodity or device comprising said element.

[0106] The above embodiments should be understood as being only for illustrative purposes of the present invention and not for limiting the protection scope of the present invention. After reading the content described in the present invention, those skilled in the art can make various changes or modifications to the present invention, and these equivalent changes and modifications also fall within the scope defined by the claims of the present invention.

Claims

1. An expression recognition method, characterized in that, it includes the following steps: Obtain a face expression picture sample and the corresponding category label; perform face detection and face alignment according to the face expression sample; Establish a deep neural network model, send the face expression picture into the deep neural network model to extract features, and obtain expression features; Select multiple triplets according to the expression category label, and the three samples included in each triplet come from different categories; Calculate a first loss value according to the expression feature and the triplet; Specifically, it includes: Calculate a loss value for each triplet in the triplet set, extract the corresponding three features from the expression features obtained before the last layer of the neural network according to the subscripts stored in the triplet, the three features can form three Euclidean distances, randomly select two of them each time, so that the ratio of the two Euclidean distances tends to 1, so take the square value of the ratio of the two Euclidean distances minus the value 1 as the loss value, and finally take the average of the loss values calculated by multiple triplets as the final value of the first loss; Calculate the cross-entropy loss according to the true category label and the obtained expression feature as the second loss value; Send the expression feature into a classifier for classification and output a classification result.

2. The expression recognition method according to claim 1, characterized in that, the face expression picture sample and the corresponding category label specifically include: Use a camera to record a video of a face showing an expression; Observe the time when the face shows an expression in the video, extract the frame corresponding to this time and save it as an image to obtain a face expression picture sample, and label the category of the expression.

3. The expression recognition method according to claim 2, characterized in that, the establishment of the deep neural network model, sending the face expression picture into the neural network model to extract features, and obtaining expression features specifically include: Create a 18-layer residual network model, and initialize the parameters included in the model, as well as parameters such as the learning rate and optimizer of the neural network; Input the expression sample into the deep neural network to obtain the eigenvalue before the last layer of the deep neural network; Input the eigenvalue before the last layer into the last fully connected layer of the model to obtain the finally output eigenvalue.

4. The expression recognition method according to claim 3, characterized in that, the selection of multiple triplets according to the expression category label, and the three samples included in each triplet come from different categories specifically include: Traverse the entire label set. When the three samples traversed all belong to different categories, store the subscripts corresponding to the current three samples into the set to obtain a triplet set storing subscripts; To ensure the distance imbalance among the three sample features in each triple, combinations where all three samples belong to positive expressions or negative expressions are filtered out from the triple set obtained in the previous step; the features of negative expressions are similar in themselves, and the method of measuring feature similarity is the distance between features. The closer the distance, the more similar, that is, the distance between negative expressions will be closer; the significance of ensuring the distance imbalance among the three samples in the triple is that currently selected triples contain two negative expressions and one positive expression, so the distance between the positive expression and the negative expressions will be relatively far, while the distance between the two negative expressions is relatively close, which is the so-called distance imbalance; in this case, through the control of the subsequent loss function, the distance between the two negative expressions can be made farther, numerically approaching the distance between the positive expression and the negative expressions, achieving the purpose of readjusting the distance between negative expressions, increasing the distance between negative expressions, and making the inter-class distance between positive and negative expressions and the inter-class distance between negative and negative expressions tend to be equal; however, the currently selected triples all belong to negative expressions.

5. An expression recognition method according to claim 4, characterized in that Take out the corresponding three 512-dimensional eigenvalue x from the features obtained before the last layer of the neural network according to the subscripts stored in each triple in the obtained triple set 1 , x 2 , x 3 . Three Euclidean distances can be obtained from the three features. Each time, two of the Euclidean distances d1 and d2 are randomly selected. The Euclidean distance calculation formula is as follows: using the ratio of d1 and d2 approaching 1 as the loss function, the purpose is to make the two inter-class distances d1 and d2 tend to be equal. One loss value is calculated for each triple, and finally the average value is taken. The loss function formula is as follows, where T is the number of triples: This loss function can only make the two distances among the three samples tend to be equal within each triple, but different two distances are randomly selected in different triples, so overall, the three distances will tend to be equal, and the result will be that all inter-class distances tend to be equal.

6. An expression recognition method according to claim 5, characterized in that calculating the cross-entropy loss based on the true class label and the expression features obtained from the output of the last layer of the neural network as the second loss value, specifically including: Suppose the output value of the $i$-th expression sample after the last layer of the neural network is $P$ i =(p i1 , p i2 , …, p iM ), and calculate the cross-entropy loss jointly with its true label $Y$ i =(y i1 , y i2 , …, y iM ). Among them, $M$ represents the number of expression categories. In the representation of the true label, assuming that this sample belongs to the $j$-th category, then the value of $y$ i in $Y$ ij is 1, and the rest of the values are 0. The calculation formula is as follows, where $N$ is the number of samples: According to the order of magnitude of the two loss values, the order of magnitude of the two loss values is adjusted to be the same by the coefficient λ for addition, and the calculation formula is as follows: Loss = λ * Loss1 + Loss2 According to the output P of the network model i Calculate the predicted class, where the subscript corresponding to the maximum value is the predicted class, and finally output the predicted class.

7. An expression recognition system using the method according to any one of claims 1-6, characterized in that comprising: an input module for inputting a face expression image into the expression recognition system; an expression acquisition module for processing the input multiple face expression images, performing face detection and face alignment on the input picture samples, and obtaining face expression picture samples; a normalization processing module for performing normalization processing on the face expression picture samples to make the sizes of the face expression pictures the same, and obtaining normalized expression picture samples; a format conversion module for converting the obtained normalized face expression pictures into the tensor format required by the neural network model; a model management module for creating, managing, and saving the neural network model; a feature extraction module for extracting the features corresponding to the expression samples through the neural network model; a combination generation module for generating expression triples that meet the requirements; The loss calculation module is used to calculate two loss values. The first loss value is calculated according to the above-mentioned triples, and the second loss value is the cross-entropy loss value calculated according to the true labels. Finally, the two losses are combined, and the network parameters are updated by backpropagating the gradient according to the loss. The prediction class module is used to predict the classification result according to the features extracted by the feature extraction module and output the prediction result.

8. The facial expression recognition system according to claim 7, wherein, the combination generation module is used to select eligible triples according to the sample labels. Each sample in a batch of samples corresponds to a unique subscript. The selection result of the triples is the subscripts of three samples from different categories, and the three samples do not belong to negative facial expressions or positive facial expressions at the same time. The purpose of storing the subscripts is to facilitate the subsequent calculation of the loss and extract the corresponding features according to the subscripts. The loss calculation module is used to calculate two loss functions, and add the two values after adjustment. The adjustment is mainly to make the orders of magnitude of the two loss values consistent and ensure that their weights are approximately equal, that is, to ensure that both play a role.

Citation Information

Patent Citations

  • Facial expression recognition method and facial expression recognition device

    CN107358169A

  • Micro-expression recognition method based on deep learning

    CN111353390A

  • Model training method and device, and expression recognition method and device

    CN113887325A