Training method of facial expression recognition model, facial expression recognition method and device
By obtaining clean and blurry training samples from the target training images and iteratively training the collaborative model, the problem of poor robustness of existing facial expression recognition models is solved, and higher recognition accuracy is achieved.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- MASHANG CONSUMER FINANCE CO LTD
- Filing Date
- 2022-10-09
- Publication Date
- 2026-05-12
AI Technical Summary
In existing technologies, facial expression recognition models rely solely on clean training samples, resulting in poor robustness and low accuracy in facial expression recognition.
Clean and conflict training samples are obtained from target training images containing facial expressions. Fuzzy training samples are obtained based on the emotional polarity of facial expressions in the conflict training samples. The collaborative model is iteratively trained, and the target expression recognition model is jointly trained using clean and fuzzy training samples.
This improves the accuracy of the facial expression recognition model in recognizing different categories of facial expressions, and solves the problem of low accuracy caused by using only clean samples.
Smart Images

Figure CN116152881B_ABST
Abstract
Description
Technical Field
[0001] This application belongs to the field of image processing technology, specifically relating to a training method for an expression recognition model, a facial expression recognition method, and an apparatus. Background Technology
[0002] Because facial expression recognition has potential applications in many fields such as healthcare, education, and smart services, it has received increasing attention.
[0003] In related technologies, trained facial expression recognition models are often used for facial expression recognition. However, these models are often based on clean training samples, which means they miss some important samples, resulting in poor robustness and low accuracy in facial expression recognition. Summary of the Invention
[0004] This application provides a training method for an expression recognition model, a facial expression recognition method, and an apparatus, which can solve the problem in related technologies where the accuracy of facial expression recognition results is low due to using only clean samples to obtain an expression recognition model.
[0005] In a first aspect, embodiments of this application provide a method for training an expression recognition model, the method comprising:
[0006] Clean training samples and conflict training samples are obtained from target training images containing facial expressions. The clean training samples are training samples whose losses in both the first recognition model and the second recognition model are less than a first threshold. The conflict training samples are training samples whose losses in one of the first recognition model and the second recognition model are less than the first threshold, and whose losses in the other of the first recognition model and the second recognition model are greater than a second threshold. The second threshold is greater than or equal to the first threshold.
[0007] Based on the emotional polarity of the facial expressions contained in the conflict training samples, fuzzy training samples are obtained from the conflict training samples.
[0008] Based on the clean training samples and the fuzzy training samples, the collaborative model for facial expression recognition is iteratively trained to obtain the target facial expression recognition model. The collaborative model includes the first recognition model and the second recognition model, and the target facial expression recognition model is a pre-trained collaborative model.
[0009] Secondly, embodiments of this application provide a facial expression recognition method, the method comprising:
[0010] Acquire a target face image, wherein the target face image is the image to be used for facial expression recognition;
[0011] The target face image is input into the target expression recognition model, which is a pre-trained expression recognition model.
[0012] The target facial expression recognition model is used to perform facial expression recognition processing on the target facial image, and the facial expression recognition result of the target facial image is output.
[0013] The target facial expression recognition model is obtained using the training method described in the first aspect.
[0014] Thirdly, embodiments of this application provide a training apparatus for an expression recognition model, the apparatus comprising:
[0015] The acquisition module is used to acquire clean training samples and conflict training samples from target training images containing facial expressions. The clean training samples are training samples whose losses in both the first recognition model and the second recognition model are less than a first threshold. The conflict training samples are training samples whose losses in one of the first recognition model and the second recognition model are less than the first threshold, and whose losses in the other of the first recognition model and the second recognition model are greater than a second threshold. The second threshold is greater than or equal to the first threshold.
[0016] The processing module is configured to obtain fuzzy training samples from the conflict training samples based on the emotional polarity of the facial expressions contained in the conflict training samples; and to iteratively train the collaborative model for expression recognition based on the clean training samples and the fuzzy training samples to obtain a target expression recognition model, wherein the collaborative model includes the first recognition model and the second recognition model, and the target expression recognition model is a pre-trained collaborative model.
[0017] Fourthly, embodiments of this application provide a facial expression recognition device, which includes:
[0018] The acquisition module is used to acquire a target face image, wherein the target face image is an image to be used for facial expression recognition;
[0019] The input module is used to input the target face image into the target expression recognition model, wherein the target expression recognition model is a pre-trained expression recognition model;
[0020] The processing module is used to perform expression recognition processing on the target face image through the target expression recognition model, and output the expression recognition result of the target face image;
[0021] The target facial expression recognition model is obtained using the training device described in the third aspect.
[0022] Fifthly, embodiments of this application provide an electronic device including a processor and a memory, the memory storing programs or instructions that, when executed by the processor, implement the steps of the method described in the first or second aspect.
[0023] In a sixth aspect, embodiments of this application provide a readable storage medium on which a program or instructions are stored, which, when executed by a processor, implement the steps of the method described in the first or second aspect.
[0024] In a seventh aspect, embodiments of this application provide a chip, the chip including a processor and a communication interface, the communication interface being coupled to the processor, the processor being used to run programs or instructions to implement the methods described in the first or second aspect.
[0025] Eighthly, embodiments of this application provide a computer program product stored in a storage medium, which is executed by at least one processor to implement the method as described in the first or second aspect.
[0026] In this embodiment, clean training samples and conflict training samples are obtained from target training images containing facial expressions. The clean training samples are those whose losses in both the first and second recognition models are less than a first threshold. The conflict training samples are those whose losses in one of the first and second recognition models are less than the first threshold, but whose losses in the other of the first and second recognition models are greater than a second threshold, where the second threshold is greater than or equal to the first threshold. Based on the emotional polarity of the facial expressions contained in the conflict training samples, fuzzy training samples are obtained from the conflict training samples. Based on the clean training samples and the fuzzy training samples, a collaborative model for expression recognition is iteratively trained to obtain a target expression recognition model. The collaborative model includes the first recognition model and the second recognition model, and the target expression recognition model is a pre-trained collaborative model. In this way, the first and second recognition models can be used together. Conflict training samples are selected based on the loss of the training samples in the first and second recognition models. Based on the emotional polarity of the facial expressions contained in the conflict training samples, fuzzy training samples are obtained from the conflict training samples. The target expression recognition model is then trained using clean and fuzzy training samples. Since both clean and fuzzy training samples are used during the training process, rather than just clean training samples, the target expression recognition model trained in this way can recognize different categories of facial expressions with high accuracy. This solves the problem in related technologies where the accuracy of facial expression recognition is low due to using only clean samples to obtain the expression recognition model. Attached Figure Description
[0027] Figure 1 This is a flowchart illustrating a training method for an expression recognition model provided in an embodiment of this application;
[0028] Figure 2 This is a flowchart of another training method for an expression recognition model provided in an embodiment of this application;
[0029] Figure 3 This is a flowchart of another training method for an expression recognition model provided in an embodiment of this application;
[0030] Figure 4 This is a schematic diagram of a training method for an expression recognition model provided in an embodiment of this application;
[0031] Figure 5 This is a flowchart of a facial expression recognition method provided in an embodiment of this application;
[0032] Figure 6 This is a schematic diagram of the experimental results obtained from experiments based on six datasets;
[0033] Figure 7 This is a schematic diagram of different types of samples in the dataset;
[0034] Figure 8 This is a schematic diagram of the experimental results of synthesizing noise on the dataset;
[0035] Figure 9 This is a schematic diagram of the results of the ablation experiment;
[0036] Figure 10 It is a schematic diagram of the distribution of visual features;
[0037] Figure 11 This is a diagram illustrating how accuracy changes as the number of training rounds increases;
[0038] Figure 12 This is a diagram illustrating a visual example;
[0039] Figure 13 This is a structural block diagram of a training device for an expression recognition model provided in an embodiment of this application;
[0040] Figure 14 This is a structural block diagram of a facial expression recognition device provided in an embodiment of this application;
[0041] Figure 15 This is a structural block diagram of an electronic device provided in an embodiment of this application. Detailed Implementation
[0042] The technical solutions of the embodiments of this application will be clearly described below with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of this application. All other embodiments obtained by those skilled in the art based on the embodiments of this application are within the scope of protection of this application.
[0043] The terms "first," "second," etc., used in the specification and claims of this application are used to distinguish similar objects and not to describe a specific order or sequence. It should be understood that such use of data can be interchanged where appropriate so that embodiments of this application can be implemented in orders other than those illustrated or described herein, and the objects distinguished by "first," "second," etc., are generally of the same class and the number of objects is not limited; for example, a first object can be one or more. Furthermore, in the specification and claims, "and / or" indicates at least one of the connected objects, and the character " / " generally indicates that the preceding and following objects are in an "or" relationship.
[0044] As described in the background section, related technologies typically only use clean training samples for training to obtain facial expression recognition models. Because these models omit other important types of samples, their robustness is poor, resulting in low accuracy in facial expression recognition. To address this issue, this application introduces a first recognition model and a second recognition model. Clean and conflict training samples are obtained by analyzing the losses of the target training image across different recognition models. Furthermore, based on the emotional polarity of the facial expressions contained in the conflict training samples, fuzzy training samples are obtained from the conflict training samples. The collaborative model, comprising the first and second recognition models, is then iteratively trained using both clean and fuzzy training samples to obtain the target facial expression recognition model. Since both clean and fuzzy training samples are used during training, rather than just clean samples, the target facial expression recognition model trained in this way can recognize different categories of facial expressions with high accuracy. This solves the problem of low accuracy in facial expression recognition caused by using only clean samples to obtain the facial expression recognition model in related technologies.
[0045] This application's embodiments can not only determine fuzzy training samples from conflicting training samples, but also further divide the conflicting training samples into first conflicting training samples and second conflicting training samples. From the first conflicting training samples, a first fuzzy training sample for a first recognition model is further determined, and from the second conflicting training samples, a second fuzzy training sample for a second recognition model is further determined. For the first fuzzy training sample, the loss of the first recognition model is less than the loss of the second recognition model; for the second fuzzy training sample, the loss of the second recognition model is less than the loss of the first recognition model. Thus, subsequently, a first loss function and a second loss function for the first recognition model can be determined differentially based on the first and second fuzzy training samples, thereby ensuring better machine learning performance. For example, the first loss function can be obtained based on the cross-entropy loss function of the first fuzzy training sample and the KL loss function of the second fuzzy training sample. The second loss function can be obtained based on the cross-entropy loss function of the second fuzzy training sample and the KL loss function of the first fuzzy training sample. Since cross-entropy can make the model focus on the main emotion, while KL divergence can make the model focus on other emotions besides the main emotion, the first recognition model can be well adjusted based on the first loss function (obtained based on the cross-entropy loss function of the first fuzzy training sample and the KL loss function of the second fuzzy training sample), and the second recognition model can be well adjusted based on the second loss function (obtained based on the cross-entropy loss function of the second fuzzy training sample and the KL loss function of the first fuzzy training sample).
[0046] Furthermore, it's important to understand that during the iterative training of the collaborative model in this application embodiment, the results of the first recognition model can be used to optimize the results of the second recognition model, or vice versa. For example, if a first fuzzy training sample obtained from a first conflict training sample shows a smaller loss in the first recognition model but a larger loss in the second recognition model, this indicates that the first recognition model is relatively more accurate than the second. In this case, the KL loss function can be used to calculate the probability value of the prediction result in the second recognition model fitting the prediction result in the first recognition model, and the parameters of the second recognition model can be updated based on the KL divergence loss value to optimize the second recognition model. This allows the second recognition model to notice non-primary expressions in the target training image. Similarly, for the same fuzzy training sample, if the recognition result of the second recognition model is more accurate, the first recognition model can also be optimized based on the prediction result of the second recognition model. In this way, the first and second recognition models can learn from each other during prediction and mutually optimize and correct each other, which is beneficial for further improving the accuracy of the target expression recognition model.
[0047] Importantly, to address the issue of suboptimal results where the first and second recognition models gradually converge with increasing training iterations, this application introduces an enhancement function and proposes a diverse enhancement scheme. Specifically, in this application, the feature correlation between the first and second recognition models can be determined first, and then a penalty term in the enhancement function can be determined based on this feature correlation. In this way, the model parameters can be adjusted according to the penalty term in the enhancement function, making the features extracted by the first recognition model differentiated from those of the second recognition model, allowing both models to extract as much information as possible from the target training image.
[0048] The methods provided in this application will be described in detail below with reference to the accompanying drawings, through specific embodiments and application scenarios.
[0049] Figure 1 This is a flowchart illustrating a training method for an facial expression recognition model provided in an embodiment of this application. For example... Figure 1 As shown, this application provides a method for training an expression recognition model. The subject of this method can be an electronic device, such as a server or a personal computer. The method for training an expression recognition model provided in this application may include the following steps:
[0050] Step 110: Obtain clean training samples and conflict training samples from the target training images containing facial expressions. The clean training samples are training samples whose losses in both the first recognition model and the second recognition model are less than a first threshold. The conflict training samples are training samples whose losses in one of the first recognition model and the second recognition model are less than the first threshold, and whose losses in the other of the first recognition model and the second recognition model are greater than a second threshold. The second threshold is greater than or equal to the first threshold.
[0051] The target training images can be multiple images randomly selected from a batch in the target training dataset, or all images in the target training dataset. Target training images can be categorized into clean training samples and conflict training samples. Conflict training samples can include at least one of noisy training samples and blurred training samples. Clean training samples can be training samples containing a single or easily identifiable facial expression. Blurred training samples can be difficult samples, i.e., training samples containing multiple facial expressions.
[0052] Each training sample in the target training image can be initially labeled, carrying label information. This label information can include an expression label and / or the emotional polarity of the expression. One expression label can correspond to one expression category. The emotional polarity of an expression can be divided into negative and positive. Specific expression categories can include surprise, happiness, anger, fear, sadness, etc. Surprise and happiness correspond to positive emotional polarities, while anger, fear, and sadness correspond to negative emotional polarities. In single-label classification, where the main facial expressions are determined, the training sample may contain only one label.
[0053] The first and second recognition models can have the same structure. Both models may contain a backbone model, which can be, for example, a VGG16 model, a VGG model, a ResNet model, or an improved version of an existing model. The first and second recognition models can identify the input target training image and output the label and / or the probability corresponding to the label. The loss of each training sample in the target training image in the first and second recognition models can be determined by calculating the cross-entropy loss function. Using the first and second recognition models together can filter out clean and conflicting training samples from the target training image.
[0054] For example, when the target training images contain a large number of training samples, the target training images can be divided into multiple batches (e.g., 4000 target training images divided into 5 batches, each containing 800 training samples). Then, the training samples from one batch are input into a first recognition model and a second recognition model. The loss of each training sample in a batch is calculated based on the outputs of the first and second recognition models. Finally, training samples whose losses in both the first and second recognition models are less than a first threshold are identified as clean training samples. Training samples whose losses in either the first or second recognition model are less than the first threshold, but whose losses in either the second or first recognition model are greater than the second threshold, are identified as conflicting training samples. The second threshold can be greater than or equal to the first threshold, and the specific values of the first and second thresholds can be set according to actual needs.
[0055] Step 120: Based on the emotional polarity of the facial expressions contained in the conflict training samples, obtain fuzzy training samples from the conflict training samples.
[0056] The emotional polarity of facial expressions included in the conflict training samples can be either pre-labeled emotional polarity or determined based on the output of the target recognition model. Since the conflict training samples suffer different losses in the first and second recognition models, the recognition model with the smaller corresponding loss can be identified as the target recognition model. Fuzzy training samples are then obtained based on the pre-labeled emotional polarity or the emotional polarity determined by the output of the target recognition model. Specifically, when the emotional polarity of the facial expressions is pre-labeled, for any training sample in the conflict training samples, it can be determined whether it is a fuzzy training sample in the following two ways: First, based on the output of the target recognition model, obtain the sum of probabilities of each expression label belonging to the pre-labeled emotional polarity in a training sample. If the sum of probabilities is greater than a first preset value, the training sample is determined to be a fuzzy training sample. Second, based on the output of the target recognition model, obtain the sum of probabilities of each expression label not belonging to the pre-labeled emotional polarity in a training sample. If the sum of probabilities is not greater than a second preset value, the training sample is determined to be a fuzzy training sample. The second preset value can be less than the first preset value. Similarly, when the emotional polarity of a facial expression is determined based on the output of the target recognition model, for any training sample in the conflict training samples, it can be determined whether it is a fuzzy training sample in the following two ways: First, based on the output of the target recognition model, take the emotional polarity of the expression with the highest probability in a training sample as the target emotional polarity, and obtain the sum of probabilities of all expression labels belonging to the target emotional polarity in the training sample. If the sum of probabilities is greater than a first preset value, the training sample is determined to be a fuzzy training sample. Second, based on the output of the target recognition model, take the emotional polarity of the expression with the highest probability in a training sample as the target emotional polarity, and obtain the sum of probabilities of all expression labels not belonging to the target emotional polarity in the training sample. If the sum of probabilities is not greater than a second preset value, the training sample is determined to be a fuzzy training sample. The second preset value can be less than the first preset value.
[0057] Step 130: Based on the clean training samples and the fuzzy training samples, iteratively train the collaborative model for expression recognition to obtain the target expression recognition model, wherein the collaborative model includes the first recognition model and the second recognition model, and the target expression recognition model is a pre-trained collaborative model.
[0058] The collaborative model can be equivalent to the first and second recognition models, or it can be a model that adds other functional modules to the first and second recognition models. After determining the clean and fuzzy training samples, the clean and fuzzy training samples can be used as the training set. The training samples in the training set are input into the collaborative model in batches. The parameters of the collaborative model are updated according to the loss of the training samples in the first and second recognition models. The collaborative model is iteratively trained to finally obtain the target expression recognition model.
[0059] The training method for the facial expression recognition model provided in this application embodiment can use a first recognition model and a second recognition model in combination. Conflict training samples are selected based on the loss of the training samples in the first and second recognition models. Based on the emotional polarity of the facial expressions contained in the conflict training samples, fuzzy training samples are obtained from the conflict training samples. The target facial expression recognition model is then trained using clean training samples and fuzzy training samples. Since both clean and fuzzy training samples are used during the training process, rather than just clean training samples, the target facial expression recognition model trained in this way can recognize different categories of facial expressions with a high accuracy rate. This solves the problem in related technologies where training the facial expression recognition model using only clean samples results in low accuracy of facial expression recognition.
[0060] The main reason for the significant loss of fuzzy training samples in facial expression recognition models is the presence of multiple expressions, and the ambiguity between expressions belonging to the same emotional polarity within the samples, making accurate identification difficult. Therefore, first determining the emotional polarity of facial expressions in the samples, and then using the probability of the expression labels belonging to that emotional polarity for judgment, can yield fuzzy training samples from conflict training samples. In one embodiment of this application, step 220, which involves obtaining fuzzy training samples from the conflict training samples based on the emotional polarity of the facial expressions contained in the conflict training samples, may include: for each sample in the conflict training samples, performing the following operations: obtaining the emotional polarity of the facial expressions contained in the sample, where the emotional polarity includes positive or negative; obtaining the sum of probabilities of each expression label belonging to that emotional polarity in the sample; and determining the sample as a fuzzy training sample if the sum of probabilities is greater than a preset value. Thus, by obtaining the sum of probabilities of each expression label belonging to that emotional polarity in the sample and comparing the sum of probabilities with a preset value, it is possible to determine whether there are multiple expressions belonging to the same emotional polarity in the sample, thereby accurately obtaining fuzzy training samples from the conflict training samples.
[0061] The emotional polarity of the facial expressions contained in the sample can be a pre-labeled emotional polarity. Each sample can have one emotional polarity, such as positive or negative. After determining the emotional polarity of the facial expressions contained in the sample, the sum of probabilities of each expression label belonging to the emotional polarity in the sample can be obtained through the output of the first recognition model or the second recognition model. If the sum of probabilities is greater than a preset value, the sample can be determined as a fuzzy training sample; if the sum of probabilities is not greater than the preset value, the sample can be determined as a noisy training sample and will not be used for subsequent training. Taking a preset value of 0.8 as an example, if the emotional polarity of the facial expressions contained in the sample is positive, and the loss of the sample in the first recognition model is less than the first threshold and the loss in the second recognition model is greater than the second threshold, then the output result of the first recognition model can be obtained, such as surprise 0.55, happiness 0.3, anger 0.05, fear 0.06, sadness 0.04. Since the emotional polarity of the expression labels surprise and happiness is positive, the probabilities of the expression labels surprise and happiness can be added together, i.e., 0.55 + 0.3. Because 0.85 is greater than the preset value of 0.8, this sample can be determined as a fuzzy training sample.
[0062] In one embodiment of this application, obtaining the sum of probabilities of each expression tag belonging to the emotional polarity in the sample may include: calculating the sum of probabilities of each expression tag belonging to the emotional polarity in the sample in the following manner: Among them, S i Let y be the sum of the probabilities of each emoji tag; C represents the specific number of emoji tag categories, and y is the sum of the probabilities of each emoji tag. i ∈{1,2,...,C} is a single-class emoji label; p j (x i ) represents sample x i The probability of belonging to category j; pol(j) represents the emotional polarity of category j; when pol(j) = pol(y) i When ) is true, l(pol(j) = pol(y) i =1. Thus, the calculation formula can be used to quickly obtain fuzzy training samples from conflict training samples.
[0063] The conflict training samples in this application embodiment can be further subdivided into first conflict training samples and second conflict training samples. The following discussion focuses on the case where the conflict training samples include both first and second conflict training samples. In one embodiment of this application, the conflict training samples include first and second conflict training samples; the first conflict training sample is a training sample whose loss in the first recognition model is less than the first threshold and whose loss in the second recognition model is greater than the second threshold; the second conflict training sample is a training sample whose loss in the second recognition model is less than the first threshold and whose loss in the first recognition model is greater than the second threshold; the fuzzy training samples include: a first fuzzy training sample for the first recognition model and a second fuzzy training sample for the second recognition model. Step 120, obtaining fuzzy training samples from the conflict training samples based on the emotional polarity of the facial expressions contained in the conflict training samples, may include: obtaining a first fuzzy training sample from the first conflict training samples based on the emotional polarity of the facial expressions contained in the first conflict training samples; and obtaining a second fuzzy training sample from the second conflict training samples based on the emotional polarity of the facial expressions contained in the second conflict training samples. In this way, conflict training samples can be distinguished into first conflict training samples and second conflict training samples based on the loss, and first fuzzy training samples and second fuzzy training samples can be obtained in a targeted manner by combining emotional polarity.
[0064] In this embodiment, the first fuzzy training sample for the first recognition model can be a training sample whose loss in the first recognition model is less than a first threshold and whose loss in the second recognition model is greater than a second threshold. Similarly, the second fuzzy training sample for the second recognition model can be a sample whose loss in the second recognition model is less than the first threshold and whose loss in the first recognition model is greater than the second threshold.
[0065] Since the loss of the first conflict training sample in the first recognition model is relatively small, it indicates that the output of the first recognition model for the first conflict training sample is relatively accurate. Therefore, the emotional polarity of the facial expressions contained in the first conflict training sample can be obtained. Combining this with the output of the first recognition model for the first conflict training sample, it is determined whether each sample in the first conflict training sample contains multiple expressions belonging to the same emotional polarity. Samples containing the same emotional polarity are then obtained from the first conflict training sample as the first fuzzy training sample. The specific acquisition process can be referred to the description above.
[0066] Similarly, since the loss of the second conflict training samples in the second recognition model is relatively small, it indicates that the output of the second recognition model for the second conflict training samples is relatively accurate. Therefore, the emotional polarity of the facial expressions contained in the second conflict training samples can be obtained. Combining this with the output of the second recognition model for the second conflict training samples, it can be determined whether each sample in the second conflict training samples contains multiple expressions belonging to the same emotional polarity. Samples containing the same emotional polarity are then obtained from the second conflict training samples as the second fuzzy training samples. The specific acquisition process can be referred to the description above.
[0067] Figure 2 This is another method for training an expression recognition model provided in the embodiments of this application. For example... Figure 2 As shown, the training method for the facial expression recognition model provided in this application embodiment may include the following steps:
[0068] Step 210: Obtain clean training samples, first conflict training samples, and second conflict training samples from the target training images containing facial expressions. The clean training samples are training samples whose losses in both the first and second recognition models are less than a first threshold. The first conflict training samples are training samples whose losses in the first recognition model are less than the first threshold and whose losses in the second recognition model are greater than a second threshold. The second conflict training samples are training samples whose losses in the second recognition model are less than the first threshold and whose losses in the first recognition model are greater than the second threshold.
[0069] The second threshold is greater than or equal to the first threshold. The target training images can be multiple images randomly selected from a batch of the target training dataset, or all images in the target training dataset. The loss of each training sample in the target training images in the first recognition model and the second recognition model can be calculated using the cross-entropy loss function.
[0070] Step 220: Based on the emotional polarity of the facial expressions contained in the first conflict training samples, obtain a first fuzzy training sample from the first conflict training samples; wherein, the first fuzzy training sample is a fuzzy training sample for the first recognition model.
[0071] In this embodiment of the application, it can be determined one by one whether each sample in the first conflict training sample contains multiple expressions belonging to the same emotional polarity. If a sample contains multiple expressions belonging to the same emotional polarity, it is determined to be a fuzzy training sample, thereby realizing the acquisition of the first fuzzy training sample from the first conflict training sample.
[0072] Step 230: Based on the emotional polarity of the facial expressions contained in the second conflict training samples, obtain a second fuzzy training sample from the second conflict training samples; wherein, the second fuzzy training sample is a fuzzy training sample for the second recognition model.
[0073] In this embodiment of the application, it can be determined one by one whether each sample in the second conflict training sample contains multiple expressions belonging to the same emotional polarity. If a sample contains multiple expressions belonging to the same emotional polarity, it is determined to be a fuzzy training sample, thereby realizing the acquisition of a second fuzzy training sample from the second conflict training sample.
[0074] Step 240: Input the target samples into the first recognition model and the second recognition model respectively to perform facial expression recognition learning. The target samples include the clean training samples, the first fuzzy training samples and the second fuzzy training samples.
[0075] Step 250: Determine the first type of loss value of the target sample using the first loss function for the first recognition model, and determine the second type of loss value of the target sample using the second loss function for the second recognition model;
[0076] Both the first and second loss functions can include at least one of the cross-entropy loss function or the relative entropy loss function (KL divergence loss function). The first type of loss function can be a function used to reflect the expression recognition learning performance of the first recognition model. The second type of loss function can be a function used to reflect the expression recognition learning performance of the second recognition model. The smaller the values of the first and second loss functions, the better the learning performance of the first and second recognition models.
[0077] Step 260: Based on the first type of loss value and the second type of loss value, iteratively train the collaborative model to obtain the target expression recognition model; wherein, the collaborative model includes the first recognition model and the second recognition model, and the target expression recognition model is a pre-trained collaborative model.
[0078] After determining the first and second type of loss values, the parameters of the first recognition model can be adjusted using the first type of loss value, and the parameters of the second recognition model can be adjusted using the second type of loss value. The adjusted first and second recognition models can then be iteratively trained. Furthermore, after determining the first and second type of loss values, other parameters that can improve model performance can be further identified. These parameters can then be combined with multi-dimensional data to adjust the first and second recognition models respectively, and the adjusted collaborative model can be iteratively trained.
[0079] The training method for the facial expression recognition model provided in this application allows the first and second recognition models to learn from clean and fuzzy data by inputting clean and fuzzy training samples into the first and second recognition models. Simultaneously, by iteratively training the collaborative model using loss values, the trained target facial expression recognition model can better fit the clean and fuzzy data, thereby improving the accuracy of facial expression recognition.
[0080] In one embodiment of this application, the first loss function is a joint loss function including a first objective function part and a second objective function part; the first objective function part is represented as the product of a first cross-entropy loss function and a first weighting coefficient corresponding to the first cross-entropy loss function; the second objective function part is represented as the product of a first KL divergence loss function and a second weighting coefficient corresponding to the first KL divergence loss function; wherein, the first objective function part is for the clean training samples and the first fuzzy training samples, and the second objective function part is for the clean training samples and the second fuzzy training samples; the second loss function is a joint loss function including a first specified function part and a second specified function part; the first specified function part is represented as the product of a second cross-entropy loss function and the first weighting coefficient; the second specified function part is represented as the product of a second KL divergence loss function and the second weighting coefficient; wherein, the first specified function part is for the clean training samples and the second fuzzy training samples, and the second specified function part is for the clean training samples and the first fuzzy training samples. Thus, on the one hand, the cross-entropy loss can be calculated to make the expression recognition model focus on the main expressions in the target expression image during training. On the other hand, the KL divergence loss can be calculated to make the expression recognition model focus on the non-main expressions in the target expression image during training. This allows the first and second recognition models to be optimized based on cross-entropy loss and KL divergence loss, which can improve the recognition accuracy of the trained target expression recognition model to a certain extent.
[0081] In this embodiment of the application, the sum of the first weighting coefficient and the second weighting coefficient can be 1.
[0082] The first loss function can be expressed as: L1=(1-λ)·L ce (S1)+λ·L KL (S2).
[0083] The second loss function can be expressed as: L2=(1-λ)·L ce (S2)+λ·L KL (S1).
[0084] Where (1-λ) is the first weighting coefficient, λ is the second weighting coefficient, S1 is the clean training sample and the first fuzzy training sample, and S2 is the clean training sample and the second fuzzy training sample. The first cross-entropy loss function L... ce (S1) may include the sum of the cross-entropy loss function for clean training samples and the cross-entropy loss function for the first fuzzy training samples. The second cross-entropy loss function L... ce (S2) may include the sum of the cross-entropy function for the clean training samples and the cross-entropy loss function for the second fuzzy training samples. The first KL divergence loss function L KL (S2) may include the sum of the KL divergence loss function for the clean samples and the KL divergence loss function for the second fuzzy training samples. The second KL divergence loss function L KL (S1) may include the sum of the KL divergence loss function for clean samples and the KL divergence loss function for the first fuzzy training samples.
[0085] For example, if the target training images are multiple images randomly selected from a batch of the target training dataset, the target training images can be represented by set I. 1 The set I represents the training samples whose loss is less than the first threshold in the first recognition model. 2 This represents the training samples whose loss is less than the first threshold in the second recognition model. A clean training sample can then be represented as I. 1 ∩I 2 .
[0086] The cross-entropy loss function for clean training samples can be expressed as:
[0087]
[0088] The first KL divergence loss function for clean training samples can be expressed as:
[0089]
[0090] The second KL divergence loss function for clean training samples can be expressed as:
[0091]
[0092] Among them, p1(I 1 ∩I 2 p2(I) can represent the predicted value of the first recognition model. 1 ∩I 2 ) can represent the predicted value of the second recognition model. The cross-entropy loss function and the KL divergence loss function for fuzzy training samples have similar expressions.
[0093] Since the training samples of the first and second recognition models are the same, as the number of training iterations increases, the first and second recognition models may gradually reach a consensus (i.e., the output results of the two models will tend to be consistent), affecting the performance of the collaborative model. Therefore, the first and second recognition models can be penalized to ensure their diversity and focus on extracting different features. In one embodiment of this application, the first loss function further includes a third objective function, which is represented as the product of a first enhancement function for diversity enhancement and a third weighting coefficient corresponding to the enhancement function; the second loss function further includes a third specified function, which is represented as the product of a second enhancement function for diversity enhancement and the third weighting coefficient. Thus, by introducing enhancement functions, the first and second recognition models can be optimized based on the first type of loss value, the second type of loss value, and the enhancement value, thereby improving the generalization ability of the first and second recognition models.
[0094] In this embodiment, the sum of the first weighting coefficient, the second weighting coefficient, and the third weighting coefficient can be 1. The enhancement function can be a function obtained based on the correlation between the first recognition model and the second recognition model.
[0095] In one embodiment of this application, when the first loss function includes a third objective function portion and the second loss function includes a third specified function portion,
[0096] The first loss function can be expressed as: L1=(1-λ-γ)·L ce (S1)+λ·L KL (S2)+γ·E1;
[0097] The second loss function can be expressed as: L2=(1-λ-γ)·L ce (S2)+λ·L KL (S1)+γ·E2;
[0098] Where (1-λ-γ) is the first weighting coefficient, λ is the second weighting coefficient, γ is the third weighting coefficient, and (1-λ-γ)+λ+γ=1; S1 represents the clean training sample and the first fuzzy training sample, S2 represents the clean training sample and the second fuzzy training sample, and L ce (S1) represents the first cross-entropy loss function for clean training samples and the first fuzzy training samples, L ce (S2) represents the second cross-entropy loss function for clean training samples and second fuzzy training samples; L KL (S2) represents the first KL loss function for the clean training samples and the second fuzzy training samples, L KL(S1) represents the second KL loss function for clean training samples and the first fuzzy training samples; E1 represents the first augmentation function, and E2 represents the second augmentation function. Thus, optimization can be performed based on cross-entropy and KL divergence loss, enabling the facial expression recognition model to accurately identify primary expressions while also paying attention to secondary expressions. Simultaneously, the augmentation function optimizes the first and second recognition models, improving their generalization ability and preventing them from converging in the later stages of training.
[0099] In one embodiment of this application, the first enhancement function can be represented as:
[0100]
[0101] The second enhancement function can be expressed as:
[0102]
[0103] Where η is a hyperparameter, This represents the average output across all individuals for the feature map of the first recognition model. f' represents the average output of all individuals for the feature map of the second recognition model. k f” represents the output of each individual in the feature map of the first recognition model. k This represents the output of each individual in the feature map of the second recognition model. This increases the diversity between each individual and the overall average output while minimizing the distance between each individual and the overall average output.
[0104] For ease of understanding, the process of obtaining the first enhancement function and the second enhancement function is explained in the embodiments of this application with reference to the following formulas.
[0105] In this embodiment, feature maps of K channels can be connected to the feature extraction modules of the first and second recognition models to obtain a first set of feature maps F1 = {f'1, f'2, f'3, ..., f'} for the first recognition model. k For the second set of feature maps F2={f”1,f”2,f”3,...,f” of the second recognition model, k For each set of feature maps, output the average value and calculate the feature correlation between the first recognition model and the second recognition model.
[0106] Among them, average output
[0107] Will Considering the first recognition model as a separately learned part, the bias-variance from F1 to F2 can be decomposed as:
[0108]
[0109] In formula (2), the first term is the deviation and the second term is the variance.
[0110] Substituting formula (1) into formula (2), we can obtain formula (3):
[0111]
[0112] After reorganization, formula (3) can be simplified to formula (4):
[0113]
[0114] In formula (4), the first term is the weighted error between F1 and F2, and the second term can be used to measure the correlation between the average output and each individual in F1.
[0115] Based on formula (4), the first enhancement function can be obtained:
[0116]
[0117] In formula (5), η is a hyperparameter, representing the trade-off between the two terms. Its value can be, for example, 1, and its specific value is determined based on experimental parameter tuning. The first term aims to improve the output of each individual in F1 and the overall average output. The diversity among them, the second term focuses on minimizing f' for each individual feature map of the first recognition model. k Compared with the total average output The distance.
[0118] Similarly, Treating this as a separately learned part of the second recognition model, and calculating it in the manner described above, we can obtain the second enhancement function:
[0119]
[0120] In formula (6), η is a hyperparameter, representing the trade-off between the two terms. Its value can be, for example, 1, and its specific value is determined based on experimental parameter tuning. The first term aims to improve the output of each individual in F2 and the overall average output. The diversity among them, the second term focuses on minimizing the individual f” of the feature map of the second recognition model. k Compared with the total average output The distance.
[0121] After determining the first enhancement function and the second enhancement function, the first enhancement function can be substituted into the first loss function, and the second enhancement function can be substituted into the first loss function. The first recognition model can be optimized based on the first loss function, and the second recognition model can be optimized based on the second loss function.
[0122] In the case where the first loss function includes the third objective function and the second loss function includes the third specified function, the embodiments of this application further elaborate on the overall process of the training method for the facial expression recognition model in conjunction with the above formula.
[0123] First, we can obtain the training set D and set the learning rate l. r Adaptive noise rate r t Fixed noise rate r, number of wheels T r Let T be the number of iterations in each round, and W be the number of iterations. Where r... t Let T be the noise rate in the t-th iteration. r The number of iterations is predefined, and T represents the total number of iterations. r The value of r is always less than or equal to T. Since the classification abilities of the first and second recognition models are relatively poor in the initial training phase, r can be gradually increased as the first and second recognition models perform better. t .specific
[0124] If t is not greater than T, the following operation can be performed:
[0125] Extract any batch I from the training dataset, and input all images in batch I as target training images into the first and second recognition models to obtain clean training samples and conflict training samples. Then, according to the formula... Obtain fuzzy training samples from the conflict training samples. Calculate the cross-entropy loss and KL divergence loss of the clean training samples, and the cross-entropy loss and KL divergence loss of the fuzzy training samples. Calculate the feature correlation (degree of association) between the first and second recognition models according to formula (2). Then substitute the above data into the first and second loss functions for calculation. Adjust the parameters of the first recognition model according to the calculation result of the first loss function, and simultaneously adjust the parameters of the second recognition model according to the calculation result of the second loss function.
[0126] The foregoing has discussed how to obtain fuzzy training samples from conflict training samples and how to determine the loss. The following sections will further discuss the specific structures and iterative processes of the first and second recognition models mentioned in this application.
[0127] In one embodiment of this application, the first recognition model includes a first feature extraction part and a first normalization part connected in sequence, and the second recognition model includes a second feature extraction part and a second normalization part connected in sequence; the first feature extraction part includes a first convolutional part, a first enhancement part, and a first fully connected layer connected in sequence; the second feature extraction part includes a second convolutional part, a second enhancement part, and a second fully connected layer connected in sequence; the first enhancement part is used to enhance the diversity of the first recognition model, and the second enhancement part is used to enhance the diversity of the second recognition model; the collaborative model is iteratively trained based on the first type of loss value and the second type of loss value to obtain... The target expression recognition model includes: adjusting the parameters of the first target part in the first recognition model based on the first type of loss value and performing iterative training to obtain a trained first target recognition model; the first target part includes at least one of a first convolutional part, a first enhancement part, and a first fully connected layer; adjusting the parameters of the second target part in the second recognition model based on the second type of loss value and performing iterative training to obtain a trained second target recognition model; the second target part includes at least one of a second convolutional part, a second enhancement part, and a second fully connected layer; and using the first target recognition model and the second target recognition model as the target expression recognition model. In this way, the parameters of the first target part in the first recognition model can be specifically adjusted based on the first type of loss value, and the parameters of the second target part in the second recognition model can be specifically adjusted based on the second type of loss value, thereby improving the recognition accuracy of the expression recognition model.
[0128] In the embodiments of this application, the first convolution part and the second convolution part can be used to map the original data in the input target training image to the hidden layer feature space, and the fully connected layer can map the feature information generated by the convolution part into a feature vector.
[0129] Figure 3 This is another method for training an expression recognition model provided in the embodiments of this application. For example... Figure 3 As shown, the training method for the facial expression recognition model provided in this application embodiment may include the following steps:
[0130] Step 310: Obtain clean training samples, first conflict training samples, and second conflict training samples from the target training images containing facial expressions. The clean training samples are training samples whose losses in both the first and second recognition models are less than a first threshold. The first conflict sample is a training sample whose loss in the first recognition model is less than the first threshold and whose loss in the second recognition model is greater than a second threshold. The second conflict sample is a training sample whose loss in the second recognition model is less than the first threshold and whose loss in the first recognition model is greater than the second threshold.
[0131] The second threshold is greater than or equal to the first threshold.
[0132] Step 320: For each sample in the first conflict training sample, perform the following operations: obtain the emotional polarity of the facial expressions contained in the sample, the emotional polarity including positive or negative; obtain the sum of probabilities of each expression label belonging to the emotional polarity in the sample; if the sum of probabilities is greater than a preset value, determine the sample as the first fuzzy training sample; wherein, the first fuzzy training sample is a fuzzy training sample for the first recognition model.
[0133] Step 330: For each sample in the second conflict training sample, perform the following operations: obtain the emotional polarity of the facial expressions contained in the sample, the emotional polarity including positive or negative; obtain the sum of probabilities of each expression label belonging to the emotional polarity in the sample; if the sum of probabilities is greater than a preset value, determine the sample as the second fuzzy training sample; wherein, the second fuzzy training sample is a fuzzy training sample for the second recognition model.
[0134] Step 340: Input the target samples into the first recognition model and the second recognition model respectively to perform facial expression recognition learning. The target samples include the clean training samples, the first fuzzy training samples and the second fuzzy training samples.
[0135] Step 350: Determine the first type of loss value of the target sample using the first loss function for the first recognition model, and determine the second type of loss value of the target sample using the second loss function for the second recognition model;
[0136] Step 360: Based on the first type of loss value and the second type of loss value, iteratively train the collaborative model to obtain the target expression recognition model, wherein the collaborative model includes the first recognition model and the second recognition model, and the target expression recognition model is a pre-trained collaborative model.
[0137] The training method for the facial expression recognition model provided in this application embodiment can use a first recognition model and a second recognition model in combination. Conflict training samples are selected based on the loss of the training samples in the first and second recognition models. Based on the emotional polarity of the facial expressions contained in the conflict training samples, fuzzy training samples are obtained from the conflict training samples. The target facial expression recognition model is then trained using clean training samples and fuzzy training samples. Since both clean and fuzzy training samples are used during the training process, rather than just clean training samples, the target facial expression recognition model trained in this way can recognize different categories of facial expressions with a high accuracy rate. This solves the problem in related technologies where the accuracy of facial expression recognition results is low due to using only clean samples to obtain the facial expression recognition model.
[0138] Figure 4 This is a schematic diagram of a training method for an expression recognition model provided in an embodiment of this application. Figure 4 Network 1 can correspond to the first recognition model mentioned above, and Network 2 can correspond to the second recognition model mentioned above. For example... Figure 4 As shown, after inputting the target training image, Network 1 and Network 2 can perform feature extraction and output corresponding results. Based on the features of the training samples in the target training image within Network 1 and Network 2, the training samples can be divided into: clean training samples, conflict training samples, and noisy training samples. After determining the conflict training samples, fuzzy training samples can be obtained from them based on the emotional polarity of the facial expressions contained within them. The loss values of the fuzzy training samples are then used to optimize Network 1 and Network 2 respectively. Simultaneously, based on the outputs of Network 1 and Network 2, an enhancement function for diversity enhancement can be determined. This enhancement function is used to differentiate the features extracted by Network 1 and Network 2, thereby improving the network's generalization ability.
[0139] Figure 5 This application provides a facial expression recognition method. For example... Figure 5 As shown, the facial expression recognition method provided in this application embodiment may include the following steps:
[0140] Step 510: Obtain the target face image, wherein the target face image is the image to be used for facial expression recognition;
[0141] The target face image can be a single image containing facial expressions, or multiple images containing facial expressions. Acquiring the target face image can include: receiving a target face image input by a user, receiving a target face image sent by an external device, or having a local device determine a face image from a local image set as the target face image. The target face image can be subsequently input into a target expression recognition model for facial expression recognition.
[0142] Step 520: Input the target face image into the target expression recognition model, wherein the target expression recognition model is a pre-trained expression recognition model;
[0143] A pre-trained facial expression recognition model can be a model that has been pre-trained multiple times on a training dataset and has demonstrated good performance in facial expression recognition. The training dataset can include both clean and blurred training samples. After inputting the target facial image into the facial expression recognition model, the model can then perform subsequent operations.
[0144] Step 530: Perform expression recognition processing on the target face image using the target expression recognition model, and output the expression recognition result of the target face image.
[0145] The target facial expression recognition model is obtained using any of the facial expression recognition model training methods described above.
[0146] The expression recognition model's processing of a target face image may include at least one of feature extraction and normalization. The expression recognition result may include the label distribution of the target face image, such as the expression labels and probabilities of each label. For a single target face image, the sum of the probabilities corresponding to each expression label in the expression recognition result can be 1.
[0147] As described above, the collaborative model can include a first recognition model and a second recognition model, with the target expression recognition model being a pre-trained collaborative model. In this scenario, after inputting a target face image into the target expression recognition model, the first and second recognition models can simultaneously process the target face image, outputting a first expression recognition result and a second expression recognition result, respectively. By averaging the probabilities of the same expression labels in the first and second expression recognition results, the expression recognition result of the target face image can be obtained.
[0148] The facial expression recognition method provided in this application uses a pre-trained expression recognition model to process the expression of a target face image, and can quickly output the expression recognition result. Furthermore, because both clean and fuzzy training samples are used during training, rather than just clean training samples, the pre-trained expression recognition model can recognize different categories of facial expressions with high accuracy.
[0149] Furthermore, this application has evaluated the training method of the facial expression recognition model mentioned above through numerous experiments. Multiple experimental results demonstrate the effectiveness of the training method for the facial expression recognition model in the embodiments of this application. The specific experimental process is as follows:
[0150] First, obtain the dataset.
[0151] (I) Obtaining Datasets from Natural Environments. To verify the robustness of the facial expression recognition model proposed in this application, experiments were conducted on five natural datasets. The first natural dataset is the Context-Aware Emotion Recognition (CAER-S) dataset. CAER-S is a context-based facial expression dataset containing 70,000 emotion images, randomly divided into a training set (70%), a validation set (10%), and a test set (20%), including seven categories: surprise, fear, disgust, happiness, sadness, anger, and neutral. The second natural dataset is FERPlus. The training, validation, and test sets of FERPlus contain 28,709, 3,589, and 3,589 images, respectively, each image resized to 48×48 pixels, belonging to one of eight emotion (facial expression) categories: neutral, happiness, surprise, sadness, anger, disgust, fear, and contempt. The third natural dataset is the large-scale Real-world Affective Faces Database (RAF-DB). RAF-DB contains 15,339 facial expression images, with 12,271 in the training set and 3,068 in the test set. Its categories are the same as the CAER-S dataset. The fourth natural dataset is the Static Facial Expressions in the Wild (SFEW) dataset. The SFEW dataset contains 879 training samples and 406 validation samples. Similar to CAER-S, it is manually labeled with seven emotions. Since the SFEW test set has not been released, this application uses the validation set for testing. The fifth natural dataset is the AffectNet dataset. AffectNet is the largest FER dataset to date, containing 450,000 images, with the same labeled categories as FERPlus.
[0152] (II) Obtaining a dataset collected in a laboratory environment. To further demonstrate that the facial expression recognition model proposed in this application has better generalization ability, it is compared with other methods on the Extended Cohn-Kanade Dataset (CK+ dataset). The CK+ dataset was collected in a laboratory and is generally considered a clean Facial Expression Recognition (FER) dataset. The CK+ dataset contains 327 facial images that can be accurately described by one of seven expressions (emotions): surprise, fear, disgust, happiness, sadness, anger, and contempt.
[0153] Second, evaluation settings
[0154] (I) Comparison Method. To verify the effectiveness of the training method for the facial expression recognition model proposed in the embodiments of this application, it is compared with state-of-the-art methods. The comparison method can be divided into two groups.
[0155] The first group comprises noise label learning methods. Some methods use a single network to train robust models. Here, the training method for the facial expression recognition model proposed in this application is compared with three representative single-network methods, such as CurriculumNet, MetaCleaner, and A Compact Embedding for Facial Expression Similarity (SL). Recent methods utilize two networks to determine which samples can be used to train the model. Mutual Learning is used here as the baseline for peer (cooperative) network-based methods. Furthermore, this application also compares the method with DeCoupling, Co-Teaching+, JoCoR, Co-Teaching, and DivideMix.
[0156] The second group consists of methods specifically designed for FER. Early models were trained using high-quality face images collected in the lab, making them sensitive to label noise. This application demonstrates the performance of two representative methods: the Weakly Supervised Local-Global Relation Network (WS-LGRN) and the Deep Subdomain Adaptation Network (DSAN). With the increasing availability of face images collected from natural environments, researchers have begun to focus on the robustness of the models. Here, the training method of the facial expression recognition model proposed in this application embodiment is compared with the following: Preserve the depth locality of Convolutional Neural Networks (DLP-CNN), Inconsistent Pseudo Annotations to Latent Truth (IPA2LT) framework, Pan et al., Training Deep Convolutional Neural Networks with Genetic Algorithm (gACNN), Random Access Array (RAN), Explicit Shape Regression (ESRs), Label Distribution Learning on Auxiliary Label Space Graphs (LDL-ALSG), and Self-Cure Network (SCN).
[0157] (II) Noise Setting. Similar to existing representative methods, the noise rate of a specific dataset was inferred using a validation set in the experiments. The noise rates calculated based on the CAER-S, FERPlus, RAF-DB, SFEW, AffectNet, and CK+ datasets were 0.1, 0.1, 0.1, 0.4, 0.4, and 0.0, respectively. Following standard protocols, this application also artificially introduces label noise into the CAER-S, FERPlus, and RAF-DB benchmarks to create artificially synthesized noise datasets to explore the robustness of the facial expression recognition model proposed in the embodiments of this application. For each type of expression, the labels of 10%, 20%, and 30% of the data can be randomly flipped, and the noise rates of these experiments can be recalculated.
[0158] (III) Structure. The training method for the facial expression recognition model proposed in this application employs two peer networks with the same prediction heads (corresponding to the first and second recognition models described above). The two networks are trained simultaneously. Various network backbone structures exist in FER, such as VGG and ResNet. To enable the two peer networks to train each other, this application selects two networks with the same structure. For a fair comparison with recent methods, this application uses VGG-16 as the backbone structure of the CNN.
[0159] Third, implementation details
[0160] (I) Training. This application uses PyTorch to implement the training method of the proposed facial expression recognition model, and trains it end-to-end on two GTX 1080ti GPUs. The two networks are pre-trained on the ImageNet and VGGFace2 datasets, respectively. The input image size is set to 224×224, employing common random cropping and horizontal flipping strategies. The model is optimized using stochastic gradient descent and undergoes 100 iterations of training. The mini-batch size, momentum, and weight decay are set to 32, 0.9, and 5e-4, respectively. The learning rate is initialized to 0.01 and decays to one-tenth of the previous rate every 30 iterations. For the hyperparameters in the facial expression recognition model proposed in this application embodiment, T is set... r =10, eta=1.0, λ=0.1, γ=0.005, τ=0.8.
[0161] (ii) Inference. During the inference phase, all test data is assumed to be clean. Two networks each predict the probability of the test data. The outputs of the two networks are then averaged to obtain the final probability for each image.
[0162] Fourth, compare with state-of-the-art methods.
[0163] The training method of the facial expression recognition model proposed in this application is compared with state-of-the-art methods on the CAER-S, FERPlus, RAF-DB, SFEW, AffectNet, and CK+ datasets. Some of the comparison methods are robust to noisy data, while others are specifically designed for the FER task. Figure 6 This is a schematic diagram of the experimental results obtained from experiments based on six datasets. Figure 6 N / A indicates that the experimental results have not been published. Figure 6 The results shown have three aspects of observation.
[0164] (I) The training method of the facial expression recognition model proposed in this application is compared with the label noise method. 1) The label noise method using peer networks outperforms the single-network method. For example, compared with SL, DivideMix has higher accuracy on all datasets. Errors in a single network accumulate, leading to selection bias. Using peer networks allows for collaborative discovery of different types of errors, resulting in better performance. 2) An interesting finding is that Co-Teaching outperforms Co-Teaching+. Co-Teaching outperforms Co-Teaching+ on all six datasets. This is because Co-Teaching+ has a limited number of usable samples on real-world noisy datasets. Therefore, Co-Teaching+ has relatively lower accuracy on the FER dataset collected in natural environments compared to Co-Teaching. 3) DivideMix outperforms other label noise methods. DivideMix uses a semi-supervised algorithm, replacing noise labels with pseudo-labels, and further trains the network using these samples. DivideMix achieves higher accuracy on five datasets except AffectNet. Specifically on the CAER-S and RAF-DB datasets, DivideMix outperformed the previous best method, Co-Teaching, by 2.59% and 1.26%, respectively. On the AffectNet dataset, which has a high noise level, the large number of pseudo-labels inevitably introduces the influence of noisy labels, resulting in a 1.22% decrease in accuracy for DivideMix compared to Co-Teaching.
[0165] (II) The training method of the expression recognition model proposed in this application and the FER method were compared. 1) On datasets collected in natural environments, the noise-robust method performs better. On RAF-DB, the accuracy of the representative noise-robust method SCN is improved by 2.77% compared with the noise-sensitive method DSAN. In addition, this application found that on RAF-DB, the accuracy of DCP-CNN and gACNN is lower than that of the noise-sensitive method. These methods do not directly discard occluded images, but train the model to identify the emotions contained in these samples. However, most occluded images cannot actually convey specific emotions, such as... Figure 7 The samples shown are labeled with noise. Therefore, these methods have suboptimal performance. Among them, Figure 7The images in the middle are from the RAF-DB dataset. The left column shows samples with clean labels, and the middle column shows samples with noisy labels. Specifically, occlusion of faces and low-quality shooting may lead to noisy labels. The right column shows samples with blurred expressions. The borders represent the true labels of the samples. 2) The noise-sensitive method designed for FER performs better on clean datasets collected in the lab. WS-LGRN and DSAN both achieve accuracy exceeding 98% on the CK+ dataset, outperforming the noise-robust method. This is because the noise-robust method aims to enable the model to capture important cues for emotion classification in natural environment images. For CK+ images, important cues such as region and reliability differ from those in natural environment images. Therefore, on CK+, the noise-sensitive method outperforms the noise-robust method.
[0166] (III) Comparison of the training method of the facial expression recognition model proposed in this application with state-of-the-art methods. 1) The training method of the facial expression recognition model proposed in this application achieves state-of-the-art performance on six datasets. On CK+, the training method of the facial expression recognition model proposed in this application achieves a classification accuracy of 98.92%, the same as DSAN. On natural environment datasets such as RAF-DB and AffectNet, the training method of the facial expression recognition model proposed in this application achieves accuracies of 89.56% and 61.82%, respectively, which is at least 1% higher than other methods. 2) Compared with the label noise method, the training method of the facial expression recognition model proposed in this application shows a significant improvement on six datasets. Although DivideMix uses a semi-supervised algorithm, replacing noisy labels with pseudo-labels, this strategy can also cause cognitive bias. Specifically, cognitive bias refers to the gradual accumulation of incorrect pseudo-labels during training, which ultimately limits the performance of the model. Unlike DivideMix, the training method of the facial expression recognition model proposed in this application selects fuzzy training samples from conflicting training samples and further uses KL divergence to train the network. Therefore, on AffectNet, the training method of the facial expression recognition model proposed in this application achieves an accuracy 3.81% higher than DivideMix. 3) Compared to the FER method, the training method of the facial expression recognition model proposed in this application considers the ambiguity of emotions when processing datasets collected from natural environments. SCN is a representative FER method that utilizes self-attention to directly measure the uncertainty of facial expressions. The training method of the facial expression recognition model proposed in this application utilizes polarity cues to further select ambiguous samples and employs KL divergence to address the ambiguity of emotions. Therefore, the training method of the facial expression recognition model proposed in this application achieves state-of-the-art performance on six datasets.
[0167] Fifth, comparison of synthetic data.
[0168] Artificially synthesized labeled noise was introduced into the FERPlus, RAF-DB, and AffectNet datasets to explore the performance of the training method for the facial expression recognition model proposed in this application in the face of synthetic noise. The synthetic noise labels can simulate extreme noise conditions with continuously increasing noise rates, and evaluate the robustness of the target facial expression recognition model obtained by the training method of the facial expression recognition model proposed in this application when the noise is high. Figure 8 This is a schematic diagram of the experimental results of synthesizing noise on the dataset. Figure 8 The accuracy is shown on three datasets. The accuracy of all methods decreases as the noise rate increases. This is because a higher noise rate results in more samples being discarded. A reduction in the number of samples leads to a decrease in performance.
[0169] The training method for the facial expression recognition model proposed in this application achieved state-of-the-art performance on all three datasets. As described above, the training method for the facial expression recognition model proposed in this application selects noisy labels based on different perspectives of the collaborative network. Despite a high noise rate, the training method for the facial expression recognition model proposed in this application can still effectively filter out noisy labels. Subsequently, the network selects clean and blurred samples to teach each other, which helps the network learn discriminative features for recognition. Furthermore, the diversity enhancement module maintains collaborative intelligence. Experimental results demonstrate the significance of the diversity enhancement module.
[0170] Sixth, ablation experiment
[0171] The training method for the facial expression recognition model proposed in this application comprises two parts: first, fuzziness-sensitive learning, used to train the network for clean and fuzzy samples; and second, diversity enhancement based on negative correlation, used to prevent the collaborative network from converging. The baseline method involves fine-tuning the backbone network (VGG-16) on the dataset. Ablation analysis of both modules is performed on the FERPlus, RAF-DB, and AffectNet datasets. Figure 9 This is a schematic diagram of the results of the ablation experiment. Figure 9 The baseline model refers to the model that is not fine-tuned using two modules. AS (Ambiguity-Sensitive) represents an ambiguity-sensitive strategy, and DE (diversity enhanced) represents a diversity-enhancing module. For example... Figure 9 As shown, two conclusions can be drawn: First, both the fuzziness-sensitive learning strategy and the diversity enhancement module can improve accuracy. Second, the training method for the facial expression recognition model proposed in this application, which includes both modules, achieves the best performance, indicating that there is no obvious conflict between the two modules, and they can cooperate to improve performance.
[0172] Next, the role of fuzziness-sensitive learning is demonstrated visually. This module aims to select clean and fuzzy samples, and then train the network using cross-entropy and KL divergence. The activation features of the penultimate layer trained on CAER-S are extracted and visualized using t-SNE. Figure 10 This is a schematic diagram of the distribution of visual features. Figure 10 The image sequentially displays the feature distributions of the training methods for SCN, DivideMix, and the facial expression recognition model proposed in this application. Although the accuracy of DivideMix is only about 0.5% higher than that of SCN, the features learned by DivideMix are significantly more discriminative than those of SCN. This means that methods aimed at reducing the influence of noisy data can better expand the distance between categories. It is observed that the training method of the facial expression recognition model proposed in this application simultaneously expands the inter-class distance and reduces the intra-class variance. The KL divergence used in the training method of the facial expression recognition model proposed in this application is designed to focus on non-primary categories, thus it is a label smoothing method designed for FER. Since label smoothing can better cluster instances belonging to the same category within clusters, the training method of the facial expression recognition model proposed in this application is inherently suitable for solving the ambiguity of facial expressions.
[0173] Furthermore, the test accuracy on RAF-DB is visualized as the number of iterations increases. The diversity enhancement module aims to maintain the diversity of the collaborative network during the training phase. Figure 11 This is a diagram illustrating how accuracy changes as the number of training rounds increases. Figure 11 The curves, from top to bottom, correspond to the training method, Co-Teaching, Co-Teaching+, baseline model, and Decoupling of the facial expression recognition model proposed in this application. The baseline model refers to training a single VGG network on RAF-DB. For example... Figure 11 As shown, most other methods achieve maximum accuracy at round 30. However, the accuracy of the training method for the facial expression recognition model proposed in this application continues to increase until it reaches its highest accuracy at round 50.
[0174] Seventh, visualization
[0175] This document presents some visualization examples of the training method for the facial expression recognition model proposed in this application on a natural environment dataset. Figure 12 This is a diagram illustrating a visual example. Figure 12 The middle border represents the true labeled sentiment category, and the histogram represents the probability in the predicted distribution. Figure 12In (a), some correctly predicted images are shown. For images with clear expressions (i.e., the first column), the training method of the expression recognition model proposed in this application can output a prediction distribution with very low entropy. Furthermore, for occluded images with relatively clear expressions (i.e., the second column), the training method of the expression recognition model proposed in this application can identify the correct category. The training method of the expression recognition model proposed in this application is also effective for blurry and low-quality images, as illustrated by the third and fourth images, respectively. Training with blurry samples helps the training method of the expression recognition model proposed in this application extract emotionally discriminative features.
[0176] exist Figure 12 (b) illustrates some failure examples of the training method for the facial expression recognition model proposed in this application. (Refer to...) Figure 12 (b) shows that the severely low-quality images in the first and fourth columns lead to incorrect predictions. If the image is occluded at key points of the expression (i.e., the second column), it is also difficult to identify the correct category. Furthermore, for the third image, which is difficult even for humans to recognize, the training method of the expression recognition model proposed in this application embodiment also fails to predict the correct emotion.
[0177] VIII. Conclusion
[0178] This application discusses the problem of facial expression recognition collected in natural environments. Embodiments of this application propose an emotion ambiguity-sensitive collaborative network to account for the differences between fuzzy samples and noisy labels, and introduce a diversity enhancement module based on negative correlation. Extensive experiments demonstrate the advantages of the proposed module. The training method for the facial expression recognition model provided in this application achieves state-of-the-art performance on six commonly used facial expression datasets.
[0179] Figure 13 This is a structural block diagram of the training device for the facial expression recognition model provided in an embodiment of this application. For example... Figure 13 As shown, the training device 1300 for the facial expression recognition model provided in this application embodiment includes: an acquisition module 1310 and a processing module 1320.
[0180] The acquisition module 1310 is used to acquire clean training samples and conflict training samples from target training images containing facial expressions. The clean training samples are training samples whose losses in both the first recognition model and the second recognition model are less than a first threshold. The conflict training samples are training samples whose losses in one of the first recognition model and the second recognition model are less than the first threshold, and whose losses in the other of the first recognition model and the second recognition model are greater than a second threshold. The second threshold is greater than or equal to the first threshold.
[0181] The processing module 1320 is used to obtain fuzzy training samples from the conflict training samples based on the emotional polarity of the facial expressions contained in the conflict training samples; and to iteratively train a collaborative model for expression recognition based on the clean training samples and the fuzzy training samples to obtain a target expression recognition model, wherein the collaborative model includes the first recognition model and the second recognition model, and the target expression recognition model is a pre-trained collaborative model.
[0182] The training device for the facial expression recognition model provided in this application embodiment can use a first recognition model and a second recognition model in combination. Conflict training samples are selected based on the loss of the training samples in the first and second recognition models. Fuzzy training samples are obtained from the conflict training samples based on the emotional polarity of the facial expressions contained in the conflict training samples. The target facial expression recognition model is then trained using clean training samples and fuzzy training samples. Since both clean and fuzzy training samples are used during the training process, rather than just clean training samples, the target facial expression recognition model trained in this way can recognize different categories of facial expressions with a high accuracy rate. This solves the problem in related technologies where the accuracy of facial expression recognition results is low due to using only clean samples to obtain the facial expression recognition model.
[0183] Optionally, in one embodiment of this application, during the process of obtaining fuzzy training samples from the conflict training samples based on the emotional polarity of facial expressions contained in the conflict training samples, the processing module 1310 is specifically configured to: for each sample in the conflict training samples, perform the following operations: obtain the emotional polarity of the facial expressions contained in the sample, wherein the emotional polarity includes positive or negative; obtain the sum of probabilities of each expression label belonging to the emotional polarity in the sample; and determine that the sample is a fuzzy training sample if the sum of probabilities is greater than a preset value. Thus, by obtaining the sum of probabilities of each expression label belonging to the emotional polarity in the sample and comparing the sum of probabilities with a preset value, it is possible to determine whether there are multiple expressions belonging to the same emotional polarity in the sample, thereby accurately obtaining fuzzy training samples from the conflict training samples.
[0184] Optionally, in one embodiment of this application, during the process of obtaining the sum of probabilities of each expression tag belonging to the emotional polarity in the sample, the processing module 1310 is specifically used to: calculate the sum of probabilities of each expression tag belonging to the emotional polarity in the sample in the following manner: Where Si is the sum of probabilities of each emoji tag; C represents the specific number of emoji tag categories, and y i ∈{1,2,...,C} is a single-class emoji label; p j (xi ) represents sample x i The probability of belonging to category j; pol(j) represents the emotional polarity of category j; when pol(j) = pol(y) i When ) is true, l(pol(j) = pol(y) i =1. Thus, the calculation formula can be used to quickly obtain fuzzy training samples from conflict training samples.
[0185] Optionally, in one embodiment of this application, the conflict training samples include a first conflict training sample and a second conflict training sample; the first conflict training sample is a training sample whose loss in the first recognition model is less than the first threshold and whose loss in the second recognition model is greater than the second threshold; the second conflict training sample is a training sample whose loss in the second recognition model is less than the first threshold and whose loss in the first recognition model is greater than the second threshold; the fuzzy training samples include: a first fuzzy training sample for the first recognition model and a second fuzzy training sample for the second recognition model; in the process of obtaining fuzzy training samples from the conflict training samples based on the emotional polarity of the facial expressions contained in the conflict training samples, the processing module 1310 is specifically used to: obtain a first fuzzy training sample from the first conflict training samples based on the emotional polarity of the facial expressions contained in the first conflict training samples; and obtain a second fuzzy training sample from the second conflict training samples based on the emotional polarity of the facial expressions contained in the second conflict training samples. Thus, the conflict training samples can be distinguished into a first conflict training sample and a second conflict training sample based on the loss, and the first fuzzy training sample and the second fuzzy training sample can be obtained specifically in combination with the emotional polarity.
[0186] Optionally, in one embodiment of this application, during the process of iteratively training the collaborative model for expression recognition based on the clean training samples and the fuzzy training samples to obtain the target expression recognition model, the processing module 1310 is specifically configured to: input the target samples into the first recognition model and the second recognition model respectively for expression recognition learning, wherein the target samples include the clean training samples, the first fuzzy training samples, and the second fuzzy training samples; determine a first type of loss value for the target samples using a first loss function for the first recognition model, and determine a second type of loss value for the target samples using a second loss function for the second recognition model; and iteratively train the collaborative model based on the first type of loss value and the second type of loss value to obtain the target expression recognition model. Thus, by inputting the clean training samples and fuzzy training samples into the first and second recognition models, the first and second recognition models can learn from clean and fuzzy data. Simultaneously, by iteratively training the collaborative model using the loss values, the trained target expression recognition model can better fit the clean and fuzzy data, improving the expression recognition accuracy.
[0187] Optionally, in one embodiment of this application, the first loss function is a joint loss function including a first objective function part and a second objective function part; the first objective function part is represented as the product of a first cross-entropy loss function and a first weighting coefficient corresponding to the first cross-entropy loss function; the second objective function part is represented as the product of a first KL divergence loss function and a second weighting coefficient corresponding to the first KL divergence loss function; wherein, the first objective function part is for the clean training samples and the first fuzzy training samples, and the second objective function part is for the clean training samples and the second fuzzy training samples; the second loss function is a joint loss function including a first specified function part and a second specified function part; the first specified function part is represented as the product of a second cross-entropy loss function and the first weighting coefficient; the second specified function part is represented as the product of a second KL divergence loss function and the second weighting coefficient; wherein, the first specified function part is for the clean training samples and the second fuzzy training samples, and the second specified function part is for the clean training samples and the first fuzzy training samples. Thus, on the one hand, the cross-entropy loss can be calculated to make the expression recognition model focus on the main expressions in the target expression image during training. On the other hand, the KL divergence loss can be calculated to make the expression recognition model focus on the non-main expressions in the target expression image during training. This allows the first and second recognition models to be optimized based on cross-entropy loss and KL divergence loss, which can improve the recognition accuracy of the trained target expression recognition model to a certain extent.
[0188] Optionally, in one embodiment of this application, the first loss function further includes a third objective function portion, which is represented as the product of a first enhancement function for diversity enhancement and a third weighting coefficient corresponding to the enhancement function; the second loss function further includes a third specified function portion, which is represented as the product of a second enhancement function for diversity enhancement and the third weighting coefficient. Thus, by introducing enhancement functions, the first and second recognition models can be optimized based on the first type of loss value, the second type of loss value, and the enhancement value, thereby improving the generalization ability of the first and second recognition models.
[0189] Optionally, in one embodiment of this application, the first loss function is expressed as: L1=(1-λ-γ)·L ce (S1)+λ·L KL (S2)+γ·E1; The second loss function is expressed as: L2=(1-λ-γ)·L ce (S2)+λ·L KL (S1)+γ·E2; where (1-λ-γ) is the first weighting coefficient, λ is the second weighting coefficient, γ is the third weighting coefficient, (1-λ-γ)+λ+γ=1; S1 represents the clean training sample and the first fuzzy training sample, S2 represents the clean training sample and the second fuzzy training sample, L ce (S1) represents the first cross-entropy loss function for clean training samples and the first fuzzy training samples, L ce (S2) represents the second cross-entropy loss function for clean training samples and second fuzzy training samples; L KL (S2) represents the first KL loss function for the clean training samples and the second fuzzy training samples, L KL (S1) represents the second KL loss function for clean training samples and the first fuzzy training samples; E1 represents the first augmentation function, and E2 represents the second augmentation function. Thus, optimization can be performed based on cross-entropy and KL divergence loss, enabling the facial expression recognition model to accurately identify primary expressions while also paying attention to secondary expressions. Simultaneously, the augmentation function optimizes the first and second recognition models, improving their generalization ability and preventing them from converging in the later stages of training.
[0190] Optionally, in one embodiment of this application, the first enhancement function is represented as: The second enhancement function is expressed as: Where η is a hyperparameter, This represents the average output across all individuals for the feature map of the first recognition model. f represents the average output of all individuals for the feature map of the second recognition model.k ' represents the output for each individual in the feature map of the first recognition model, f k "" represents the output of each individual in the feature map of the second recognition model. In this way, the diversity between each individual and the overall average output can be improved, while minimizing the distance between each individual and the overall average output.
[0191] Optionally, in one embodiment of this application, the first recognition model includes a first feature extraction part and a first normalization part connected in sequence, and the second recognition model includes a second feature extraction part and a second normalization part connected in sequence; the first feature extraction part includes a first convolutional part, a first enhancement part, and a first fully connected layer connected in sequence; the second feature extraction part includes a second convolutional part, a second enhancement part, and a second fully connected layer connected in sequence; the first enhancement part is used to enhance the diversity of the first recognition model, and the second enhancement part is used to enhance the diversity of the second recognition model; the collaborative model is iteratively trained based on the first type of loss value and the second type of loss value to obtain the target expression. During the recognition model process, the processing module 1310 is specifically used to: adjust the parameters of the first target part in the first recognition model based on the first type of loss value, and perform iterative training of the model to obtain a trained first target recognition model; the first target part includes at least one of a first convolutional part, a first enhancement part, and a first fully connected layer; adjust the parameters of the second target part in the second recognition model based on the second type of loss value, and perform iterative training of the model to obtain a trained second target recognition model; the second target part includes at least one of a second convolutional part, a second enhancement part, and a second fully connected layer; and use the first target recognition model and the second target recognition model as the target expression recognition model. In this way, the parameters of the first target part in the first recognition model can be specifically adjusted based on the first type of loss value, and the parameters of the second target part in the second recognition model can be specifically adjusted based on the second type of loss value, thereby improving the recognition accuracy of the expression recognition model.
[0192] It should be noted that the training device for the facial expression recognition model provided in this application corresponds to the training method for the facial expression recognition model mentioned above. Related details can be found in the description of the training method for the facial expression recognition model above, and will not be repeated here.
[0193] Figure 14 This is a structural block diagram of a facial expression recognition device provided in an embodiment of this application. Figure 14 As shown, the facial expression recognition device 1400 provided in this application embodiment includes: an acquisition module 1410, an input module 1420, and a processing module 1430;
[0194] The acquisition module 1410 is used to acquire a target face image, wherein the target face image is an image to be used for facial expression recognition;
[0195] The input module 1420 is used to input the target face image into the target expression recognition model, wherein the target expression recognition model is a pre-trained expression recognition model.
[0196] The processing module 1430 is used to perform expression recognition processing on the target face image through the target expression recognition model, and output the expression recognition result of the target face image;
[0197] The target facial expression recognition model 1400 is obtained using the training device 1300 for the facial expression recognition model described above.
[0198] The facial expression recognition device provided in this application uses a pre-trained expression recognition model to process the expression of a target face image, and can quickly output the expression recognition result. Furthermore, because both clean and fuzzy training samples are used during the training process, rather than just clean training samples, the pre-trained expression recognition model can recognize different categories of facial expressions with high accuracy.
[0199] In addition, such as Figure 15As shown in the figure, this application embodiment also provides an electronic device 1500, which can be various types of computers, etc. The electronic device 1500 includes: a processor 1510 and a memory 1520. The memory 1520 stores programs or instructions, and when the programs or instructions are executed by the processor 1510, they implement the steps of any of the methods described above. For example, when the program is executed by the processor 1520, it performs the following process: obtaining clean training samples and conflict training samples from target training images containing facial expressions, wherein the clean training samples are training samples whose losses in both the first recognition model and the second recognition model are less than a first threshold, and the conflict training samples are training samples whose losses in one of the first recognition model and the second recognition model are less than the first threshold, and whose losses in the other of the first recognition model and the second recognition model are greater than a second threshold, wherein the second threshold is greater than or equal to the first threshold; obtaining fuzzy training samples from the conflict training samples based on the emotional polarity of the facial expressions contained in the conflict training samples; and iteratively training a collaborative model for expression recognition based on the clean training samples and the fuzzy training samples to obtain a target expression recognition model, wherein the collaborative model includes the first recognition model and the second recognition model, and the target expression recognition model is a pre-trained collaborative model. In this way, the first and second recognition models can be used together. Conflict training samples are selected based on the loss of the training samples in the first and second recognition models. Based on the emotional polarity of the facial expressions contained in the conflict training samples, fuzzy training samples are obtained from the conflict training samples. The target expression recognition model is then trained using clean and fuzzy training samples. Since both clean and fuzzy training samples are used during the training process, rather than just clean training samples, the target expression recognition model trained in this way can recognize different categories of facial expressions with high accuracy. This solves the problem in related technologies where the accuracy of facial expression recognition is low due to using only clean samples to obtain the expression recognition model.
[0200] This application embodiment also provides a readable storage medium storing a program or instructions that, when executed by the processor 1510, implement the steps of any of the methods described above. For example, when the program is executed by the processor 1510, it performs the following process: obtaining clean training samples and conflict training samples from a target training image containing facial expressions, wherein the clean training samples are training samples whose losses in both the first and second recognition models are less than a first threshold, and the conflict training samples are training samples whose losses in one of the first and second recognition models are less than the first threshold, and whose losses in the other of the first and second recognition models are greater than a second threshold, where the second threshold is greater than or equal to the first threshold; obtaining fuzzy training samples from the conflict training samples based on the emotional polarity of the facial expressions contained in the conflict training samples; and iteratively training a collaborative model for expression recognition based on the clean training samples and the fuzzy training samples to obtain a target expression recognition model, wherein the collaborative model includes the first recognition model and the second recognition model, and the target expression recognition model is a pre-trained collaborative model. In this way, the first and second recognition models can be used together. Conflict training samples are selected based on the loss of the training samples in the first and second recognition models. Based on the emotional polarity of the facial expressions contained in the conflict training samples, fuzzy training samples are obtained from the conflict training samples. The target expression recognition model is then trained using clean and fuzzy training samples. Since both clean and fuzzy training samples are used during the training process, rather than just clean training samples, the target expression recognition model trained in this way can recognize different categories of facial expressions with high accuracy. This solves the problem in related technologies where the accuracy of facial expression recognition is low due to using only clean samples to obtain the expression recognition model.
[0201] Those skilled in the art will understand that embodiments of the present invention can be provided as methods, systems, or computer program products. Therefore, the present invention can take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, the present invention can take the form of a computer program product embodied on one or more computer-usable storage media (including, but not limited to, disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.
[0202] This invention is described with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of the invention. It will be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, special-purpose computer, embedded processor, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, generate instructions for implementing the flowchart illustrations and / or block diagrams. Figure 1 One or more processes and / or boxes Figure 1 A device that provides the functions specified in one or more boxes.
[0203] These computer program instructions may also be stored in a computer-readable storage medium that can direct a computer or other programmable data processing device to function in a particular manner, such that the instructions stored in the computer-readable storage medium produce an article of manufacture including instruction means, which are implemented in a process Figure 1 One or more processes and / or boxes Figure 1 The function specified in one or more boxes.
[0204] These computer program instructions may also be loaded onto a computer or other programmable data processing equipment to cause a series of operational steps to be performed on the computer or other programmable equipment to produce a computer-implemented process, thereby providing instructions that execute on the computer or other programmable equipment for implementing the process. Figure 1 One or more processes and / or boxes Figure 1 The steps of the function specified in one or more boxes.
[0205] In a typical configuration, a computing device includes one or more processors (CPU), input / output interfaces, network interfaces, and memory.
[0206] Memory may include non-persistent storage in computer-readable media, such as random access memory (RAM) and / or non-volatile memory, such as read-only memory (ROM) or flash RAM. Memory is an example of computer-readable media.
[0207] Computer-readable media includes both permanent and non-permanent, removable and non-removable media that can store information using any method or technology. Information can be computer-readable instructions, data structures, modules of programs, or other data. Examples of computer storage media include, but are not limited to, phase-change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technologies, CD-ROM, digital versatile optical disc (DVD) or other optical storage, magnetic tape, magnetic magnetic disk storage or other magnetic storage devices, or any other non-transferable medium that can be used to store information accessible by a computing device. As defined herein, computer-readable media does not include transient computer-readable media, such as modulated data signals and carrier waves.
[0208] It should also be noted that the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitation, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes said element.
[0209] Those skilled in the art will understand that embodiments of this application can be provided as methods, systems, or computer program products. Therefore, this application can take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, this application can take the form of a computer program product embodied on one or more computer-usable storage media (including, but not limited to, disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.
[0210] The above description is merely an embodiment of this application and is not intended to limit this application. Various modifications and variations can be made to this application by those skilled in the art. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principle of this application should be included within the scope of the claims of this application.
Claims
1. A training method for an facial expression recognition model, characterized in that, include: Clean training samples and conflict training samples are obtained from target training images containing facial expressions. The clean training samples are training samples whose losses in both the first recognition model and the second recognition model are less than a first threshold. The conflict training samples are training samples whose losses in one of the first recognition model and the second recognition model are less than the first threshold, and whose losses in the other of the first recognition model and the second recognition model are greater than a second threshold. The second threshold is greater than or equal to the first threshold. Based on the emotional polarity of the facial expressions contained in the conflict training samples, fuzzy training samples are obtained from the conflict training samples; the fuzzy training samples are training samples containing multiple facial expressions. Based on the clean training samples and the fuzzy training samples, the collaborative model for facial expression recognition is iteratively trained to obtain the target facial expression recognition model. The collaborative model includes the first recognition model and the second recognition model, and the target facial expression recognition model is a pre-trained collaborative model.
2. The training method according to claim 1, characterized in that, The step of obtaining fuzzy training samples from the conflict training samples based on the emotional polarity of facial expressions contained in the conflict training samples includes: For each sample in the conflict training samples, perform the following operation: Obtain the emotional polarity of the facial expressions contained in the sample, wherein the emotional polarity includes positive or negative; Obtain the sum of probabilities of each expression label belonging to the emotional polarity in the sample; If the sum of the probabilities is greater than a preset value, the sample is determined to be a fuzzy training sample.
3. The training method according to claim 2, characterized in that, The step of obtaining the sum of probabilities of each expression label belonging to the emotional polarity in the sample includes: The sum of probabilities of each expression label belonging to the emotional polarity in the sample is calculated as follows: Where Si is the sum of the probabilities of each emoji tag; C represents the specific number of emoji tag categories. It is a single-category emoji tag; Indicates sample The probability of belonging to category j; Indicates the emotional polarity of category j; when When true, .
4. The training method according to claim 1, characterized in that, The conflict training samples include a first conflict training sample and a second conflict training sample; the first conflict training sample is a training sample whose loss in the first recognition model is less than the first threshold and whose loss in the second recognition model is greater than the second threshold; the second conflict training sample is a training sample whose loss in the second recognition model is less than the first threshold and whose loss in the first recognition model is greater than the second threshold. The fuzzy training samples include: a first fuzzy training sample for the first recognition model and a second fuzzy training sample for the second recognition model; The step of obtaining fuzzy training samples from the conflict training samples based on the emotional polarity of facial expressions contained in the conflict training samples includes: Based on the emotional polarity of the facial expressions contained in the first conflict training sample, a first fuzzy training sample is obtained from the first conflict training sample. Based on the emotional polarity of the facial expressions contained in the second conflict training samples, a second fuzzy training sample is obtained from the second conflict training samples.
5. The training method according to claim 4, characterized in that, The step of iteratively training the collaborative model for expression recognition based on the clean training samples and the fuzzy training samples to obtain the target expression recognition model includes: The target samples are respectively input into the first recognition model and the second recognition model to perform facial expression recognition learning. The target samples include the clean training samples, the first fuzzy training samples and the second fuzzy training samples. The first type of loss value of the target sample is determined by a first loss function for the first recognition model, and the second type of loss value of the target sample is determined by a second loss function for the second recognition model. Based on the first type of loss value and the second type of loss value, the collaborative model is iteratively trained to obtain the target expression recognition model.
6. The training method according to claim 5, characterized in that, The first loss function is a joint loss function including a first objective function part and a second objective function part; the first objective function part is represented as the product of a first cross-entropy loss function and a first weighting coefficient corresponding to the first cross-entropy loss function; the second objective function part is represented as the product of a first KL divergence loss function and a second weighting coefficient corresponding to the first KL divergence loss function; wherein, the first objective function part is for the clean training samples and the first fuzzy training samples, and the second objective function part is for the clean training samples and the second fuzzy training samples; The second loss function is a joint loss function including a first specified function part and a second specified function part; the first specified function part is represented as the product of the second cross-entropy loss function and the first weighting coefficient; the second specified function part is represented as the product of the second KL divergence loss function and the second weighting coefficient; wherein, the first specified function part is for the clean training samples and the second fuzzy training samples, and the second specified function part is for the clean training samples and the first fuzzy training samples.
7. The training method according to claim 6, characterized in that, The first loss function further includes a third objective function, which is represented as the product of a first enhancement function for diversity enhancement and a third weighting coefficient corresponding to the enhancement function; The second loss function also includes a third specified function portion, which is represented as the product of a second enhancement function for diversity enhancement and the third weighting coefficient.
8. The training method according to claim 7, characterized in that, The first loss function is expressed as: ; The second loss function is expressed as: ; Where (1-λ-γ) is the first weighting coefficient, λ is the second weighting coefficient, γ is the third weighting coefficient, and (1-λ-γ)+λ+γ=1; S1 represents the clean training sample and the first fuzzy training sample, S2 represents the clean training sample and the second fuzzy training sample, and L ce (S1) represents the first cross-entropy loss function for clean training samples and the first fuzzy training samples, L ce (S2) represents the second cross-entropy loss function for clean training samples and second fuzzy training samples; L KL (S2) represents the first KL loss function for the clean training samples and the second fuzzy training samples, L KL (S1) represents the second KL loss function for clean training samples and the first fuzzy training samples; E1 represents the first boosting function and E2 represents the second boosting function.
9. The training method according to claim 8, characterized in that, The first enhancement function is represented as: The second enhancement function is expressed as: in, For hyperparameters, This represents the average output across all individuals for the feature map of the first recognition model. This represents the average output across all individuals for the feature map of the second recognition model. This represents the output for each individual feature map of the first recognition model. This represents the output for each individual in the feature map of the second recognition model.
10. The training method according to claim 7, characterized in that, The first recognition model includes a first feature extraction part and a first normalization part connected in sequence, and the second recognition model includes a second feature extraction part and a second normalization part connected in sequence; the first feature extraction part includes a first convolutional part, a first enhancement part, and a first fully connected layer connected in sequence; the second feature extraction part includes a second convolutional part, a second enhancement part, and a second fully connected layer connected in sequence; the first enhancement part is used to enhance the diversity of the first recognition model, and the second enhancement part is used to enhance the diversity of the second recognition model; The step of iteratively training the collaborative model based on the first type of loss value and the second type of loss value to obtain the target expression recognition model includes: Based on the first type of loss value, the parameters of the first target part in the first recognition model are adjusted and the model is iteratively trained to obtain a trained first target recognition model; the first target part includes at least one of the following: a first convolutional part, a first enhancement part, and a first fully connected layer; Based on the second type of loss value, the parameters of the second target part in the second recognition model are adjusted and the model is iteratively trained to obtain a trained second target recognition model; the second target part includes at least one of the following: a second convolutional part, a second enhancement part, and a second fully connected layer; The first target recognition model and the second target recognition model are used as the target expression recognition model.
11. A method for facial expression recognition, characterized in that, include: Acquire a target face image, wherein the target face image is the image to be used for facial expression recognition; The target face image is input into the target expression recognition model, which is a pre-trained expression recognition model. The target facial expression recognition model is used to perform facial expression recognition processing on the target facial image, and the facial expression recognition result of the target facial image is output. The target facial expression recognition model is obtained using the training method according to any one of claims 1-10.
12. An electronic device, characterized in that, It includes a processor and a memory, the memory storing a program or instructions that, when executed by the processor, implement the steps of the method as described in any one of claims 1-11.
13. A readable storage medium, characterized in that, The readable storage medium stores a program or instructions that, when executed by a processor, implement the steps of the method as described in any one of claims 1-11.