Method and device for optimizing facial expression recognition
The method addresses noise and class imbalance in facial expression recognition by using data augmentation and adaptive binary cross-entropy loss to improve deep learning models' accuracy in face expression recognition.
Patent Information
- Application Number
- CN202510473597.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-16
- Publication Date
- 2025-07-15
AI Technical Summary
The existing facial expression recognition technology has low recognition rate when facing noise and long-tail imbalance problems, which affects the practical applications of emotional disorder diagnosis, intelligent security and abnormal behavior detection.
Through data augmentation, fine-grained feature extraction, semantic consistency correction and adaptive binary cross-entropy loss, the facial expression recognition model is optimized and its anti-noise interference capability is enhanced.
It improves the accuracy of facial expression recognition, solves the problem of difficult expression recognition, reduces the negative impact of noise labels, and improves the performance of deep learning models in practical application scenarios.
Smart Images

Figure CN120318883A_ABST
Abstract
Description
Technical Field
[0001] The present invention discloses a method and device for optimizing facial expression recognition, relating to the technical field of image recognition. Background Art
[0002] Facial expression recognition technology can help a computer effectively interpret human emotions, and thus is widely applied in many fields such as the design of an adaptive human-computer interaction interface, diagnosis of emotional disorders, intelligent security, and detection of abnormal behaviors. However, the inherent ambiguity of facial expressions and the subjectivity of annotators lead to noise in facial expression images. At the same time, there is a serious class imbalance phenomenon in the publicly available facial expression datasets, and the number of positive emotion classes is far more than that of negative emotion classes. And due to the problem of noise, there is a problem of long-tail distribution imbalance in the facial expression dataset, that is, the tail classes in the facial expression dataset will be habitually misrecognized as other classes, reducing the performance of the facial expression recognition method in actual application scenarios. However, few existing methods simultaneously focus on the problems of noise and long-tail imbalance, so the recognition rate on the tail classes of the facial expression dataset is low, which is very unfavorable for the actual applications in aspects such as diagnosis of emotional disorders, intelligent security, and detection of abnormal behaviors. Summary of the Invention
[0003] Aiming at the problems of the prior art, the present invention provides a method and device for optimizing facial expression recognition, enhancing the anti-noise interference ability of the deep learning model and improving the facial expression recognition accuracy of the deep learning model in actual application scenarios.
[0004] The specific solution proposed by the present invention is as follows:
[0005] The present invention provides a method for optimizing facial expression recognition, including:
[0006] Step 1: For the facial image I, perform data augmentation operations to obtain the facial image I' after random erasing and the facial image I'' after random erasing and random flipping,
[0007] Randomly crop the facial image I'', perform a binary difference operation on the cropped image, and restore the image to the image size before cropping to obtain the facial image I t ,
[0008] Create a visual cue in the pixel space for the facial image I, add the visual cue to the facial image I to form a cue image set, and extract the fine-grained texture and edge features of the facial image I by using a Gabor operator and add them to the cue image set.
[0009] Step 2: Use a convolutional neural network CNN to construct a facial expression recognition model, use the facial expression recognition model to extract the global feature F according to the facial image I', and use the facial expression recognition model to extract the global feature F according to the facial image I tExtract the local feature F”, and use the facial expression recognition model to extract the fine-grained feature F’ according to the prompt image set.
[0010] Step 3: Use the facial expression recognition model to calculate and obtain the class activation maps M and M" based on the global feature F and the local feature F”. After performing a random horizontal flip operation on the class activation map M", obtain the class activation map M′, and compare the semantic consistency between M and M′ to compensate for the feature blur caused by pose changes and occlusions.
[0011] Step 4: Use the facial expression recognition model to add an adaptive factor to the binary cross-entropy loss for long-tail correction to avoid the classification bias of the tail classes caused by noise in the long-tail distribution.
[0012] Step 5: Use the facial expression recognition model to send F and F’ into the global average pooling layer to reduce the feature dimension, and then send the dimension-reduced feature f into the fully connected layer to obtain two classification probabilities. After adding the two classification probabilities, use the classification loss to obtain the facial expression classification. The calculation formula is as follows:
[0013]
[0014] is the weight of y in the fully connected layer i y i is the label of the i-th given facial image, f i is the feature of the i-th facial image, C is the number of facial expression categories, and N is the number of samples used to train the model.
[0015] Furthermore, in step 1 of the method for optimizing facial expression recognition, use RandomErasing() in the image transformation tool transforms of the torchvision tool to perform random erasing on the facial image I to obtain the erased facial image I’;
[0016] After using RandomErasing() in the image transformation tool transforms of the torchvision tool to perform random erasing on the facial image I, then use RandomHorizontalFlip() to perform random horizontal flipping to obtain the erased and flipped facial image I”.
[0017] Furthermore, in step 1 of the method for optimizing facial expression recognition, use ν φ to represent the visual prompt, ν φ with a size of where H, W, and C are the height, width, and number of channels of the facial image I respectively. Add the visual prompt ν φ to the facial image I to form the prompt image set
[0018] Then prompt the image set
[0019] Extract the fine-grained texture and edge features of the facial image I using the Gabor operator and add them to the prompt image set in
[0020] Then g(x) is the fine-grained texture and edge features of the facial image I extracted by the Gabor operator.
[0021] Furthermore, in step 3 of the method for optimizing facial expression recognition, the following formula is adopted:
[0022]
[0023] MSE is the mean square error, calculate the loss Use the loss Compare the semantic consistency of M and M'.
[0024] Furthermore, in step 4 of the method for optimizing facial expression recognition, based on the formula of balanced binary cross-entropy loss:
[0025]
[0026] represents the distribution of the y i th category, is the number of the y i th category, N is the total number of samples, introduce an adaptive factor C is the number of facial image categories, and the formula for the new binary cross-entropy loss is obtained:
[0027]
[0028] Among them, is the distribution of the y i th category in the training set of the facial expression recognition model, and the adaptive binary cross-entropy loss formula is sorted out:
[0029]
[0030] Among them, is the one-hot label, is the bias term of the logarithm of y i , is the logarithmic probability of the y i th category.
[0031] The present invention also provides a device for optimizing facial expression recognition, including an enhancement prompt module, a model management module, a semantic consistency module, and an adaptive binary cross-entropy loss module.
[0032] The enhancement prompt module performs data enhancement operations on the facial image I to obtain the randomly erased facial image I' and the randomly erased and randomly flipped facial image I".
[0033] Randomly crop the facial image I", perform binary difference operations on the cropped image, and restore the image to the size before cropping to obtain the facial image I. t ,
[0034] Create visual prompts for the facial image I in the pixel space, add the visual prompts to the facial image I to form a prompt image set, and use the Gabor operator to extract the fine-grained texture and edge features of the facial image I and add them to the prompt image set.
[0035] The model management module uses a convolutional neural network CNN to build a facial expression recognition model. The model management module uses the facial expression recognition model to extract the global feature F according to the facial image I', uses the facial expression recognition model to extract the local feature F" according to the facial image I t Extract the fine-grained feature F' according to the prompt image set using the facial expression recognition model.
[0036] The semantic consistency module uses the facial expression recognition model to calculate the class activation maps M and M" according to the global feature F and the local feature F", performs a random horizontal flip operation on the class activation map M" to obtain the class activation map M′, and compares the semantic consistency of M and M′ to make up for the feature blur caused by pose changes and occlusions.
[0037] The adaptive binary cross-entropy loss module uses the facial expression recognition model to add an adaptive factor to the binary cross-entropy loss for long-tail correction to avoid classification bias of tail classes caused by noise in the case of long-tail distribution.
[0038] The classification module uses the facial expression recognition model to send F and F' into the global average pooling layer to reduce the feature dimension, and then sends the reduced-dimensional feature f into the fully connected layer to obtain two classification probabilities. After adding the two classification probabilities, use the classification loss Obtain the facial expression classification. The calculation formula is as follows:
[0039]
[0040] is the weight of yi in the fully connected layer, yi is the label of the given i-th facial image, f iis the feature of the i-th facial image, C is the number of facial expression categories, and N is the number of samples used to train the model.
[0041] Furthermore, the enhancement prompt module of the device for optimizing facial expression recognition randomly erases the facial image I using RandomErasing() in the image transformation tool transforms of torchvision to obtain the erased facial image I';
[0042] After randomly erasing the facial image I using RandomErasing() in the image transformation tool transforms of torchvision, it is then randomly horizontally flipped using RandomHorizontalFlip() to obtain the erased and flipped facial image I".
[0043] Furthermore, the enhancement prompt module of the device for optimizing facial expression recognition uses ν φ to represent the visual prompt, and the size of ν φ is H, W, and C are the height, width, and number of channels of the facial image I respectively. The visual prompt ν φ is added to the facial image I to form the prompt image set
[0044] Then the prompt image set
[0045] Extracts the fine-grained texture and edge features of the facial image I using the Gabor operator and adds them to the prompt image set in,
[0046] Then g(x) is the fine-grained texture and edge features of the facial image I extracted by the Gabor operator.
[0047] Furthermore, the semantic consistency module of the device for optimizing facial expression recognition uses the following formula:
[0048]
[0049] MSE is the mean square error, and the loss is used to compare the semantic consistency between M and M'.
[0050] Furthermore, the adaptive binary cross-entropy loss module of the device for optimizing facial expression recognition is based on the formula of balanced binary cross-entropy loss:
[0051]
[0052] Represents the distribution of the y-th i category, is the number of the y-th i category, N is the total number of samples, and an adaptive factor is introduced C is the number of facial image categories, and the formula for the new binary cross-entropy loss is obtained:
[0053]
[0054] where is the distribution of the y-th i category in the training set of the facial expression recognition model, and the adaptive binary cross-entropy loss formula is sorted out:
[0055]
[0056] where is the one-hot label, is the bias term of the logarithm of y i of, is the logarithmic probability of the y-th i category.
[0057] The beneficial effects of the present invention are:
[0058] The present invention performs data augmentation on facial images, can obtain multi-scale augmented data, and performs fine-grained feature extraction of the original facial images, enhancing the fine-grained feature learning of the feature extractor shared by the facial expression recognition model.
[0059] And using random erasing, randomly erased and randomly horizontally flipped images to simulate difficult expression recognition situations such as occlusion and pose changes, and using semantic consistency contrast to obtain complete facial images to solve the problem of difficult expression recognition.
[0060] At the same time, an adaptive binary cross-entropy loss is defined to perform long-tail correction to optimize the problem of misidentifying facial expressions in the tail classes by the facial expression recognition model. Integrating the process of optimizing the model of the present invention, reducing the negative impact of noisy labels on facial expression recognition, enhancing the anti-noise interference ability of the deep learning model, and improving the facial expression recognition accuracy of the deep learning model in actual application scenarios. BRIEF DESCRIPTION OF THE DRAWINGS
[0061] Figure 1 is a schematic diagram of the application process framework of the method of the present invention.
[0062] Figure 2 is a schematic diagram of the facial image classification process. DETAILED DESCRIPTION OF THE INVENTION
[0063] The present invention will be further described below in conjunction with the accompanying drawings and specific embodiments, so that those skilled in the art can better understand the present invention and be able to implement it, but the embodiments cited do not limit the present invention.
[0064] Embodiment 1
[0065] The present invention provides a method for optimizing facial expression recognition, including:
[0066] Step 1: For the facial image I, perform data augmentation operations to obtain the randomly erased facial image I' and the randomly erased and randomly flipped facial image I".
[0067] Randomly crop the facial image I", perform a binary difference operation on the cropped image, and restore the image to the size before cropping to obtain the facial image I t ,
[0068] Create a visual cue in the pixel space of the facial image I, add the visual cue to the facial image I to form a set of cue images, and use the Gabor operator to extract the fine-grained texture and edge features of the facial image I and add them to the set of cue images.
[0069] Among them, the facial image I can be read using Pytorch series tools, and the input facial image I can be transformed through the transforms image transformation tool in the torchvision tool. Use RandomErasing() in the image transformation tool transforms of the torchvision tool to randomly erase the facial image I to obtain the erased facial image I'; after randomly erasing the facial image I using RandomErasing() in the image transformation tool transforms of the torchvision tool, then use RandomHorizontalFlip() for random horizontal flipping to obtain the erased and flipped facial image I". It is also possible to perform transformation operations such as ToPILImage(), Resize(224,224), ToTensor(), Normalize() on the facial image according to the requirements of the facial expression recognition model to facilitate subsequent image processing.
[0070] In step 1, ν φ is also used to represent the visual cue, and the size of ν φ is where H, W, and C are the height, width, and number of channels of the facial image I respectively. Add the visual cue ν φ to the facial image I to form a set of cue images
[0071] Then the set of cue images
[0072] Extract the fine-grained texture and edge features of the facial image I using the Gabor operator and add them to the hint image set in
[0073] Then g(x) is the fine-grained texture and edge features of the facial image I extracted by the Gabor operator.
[0074] Step 2: Use the convolutional neural network CNN to construct a facial expression recognition model. Use the facial expression recognition model to extract the global feature F according to the facial image I’, use the facial expression recognition model to extract the local feature F” according to the facial image I t Extract the fine-grained feature F’ according to the hint image set using the facial expression recognition model. The convolutional neural network CNN can be selected according to application requirements. For example, use the convolutional neural network (CNNs) ResNet18 to construct the facial expression recognition model, and use the facial expression recognition model constructed by the ResNet18 deep learning model as a shared feature extractor for feature extraction.
[0075] Step 3: Use the facial expression recognition model to calculate and obtain the class activation maps M and M" according to the global feature F and the local feature F”. After performing a random horizontal flip operation on the class activation map M", obtain the class activation map M′, and compare the semantic consistency between M and M′ to make up for the feature blur caused by pose changes and occlusions.
[0076] The following formula can be used:[[]]
[0077]
[0078] MSE is the mean square error, calculate the loss Use the loss Compare the semantic consistency between M and M′.
[0079] Step 4: Use the facial expression recognition model to add an adaptive factor to the binary cross-entropy loss for long-tail correction to avoid classification bias of the tail classes caused by noise in the long-tail distribution case.
[0080] The following formula can be based on the balanced binary cross-entropy loss:[[]]
[0081]
[0082] represents the distribution of the y i class, is the number of samples in the y i class, N is the total number of samples, introduce an adaptive factor C is the number of facial image classes, and the formula for the new binary cross-entropy loss is obtained:[[]]
[0083]
[0084] Among them, is the distribution of the y-th i category in the facial expression recognition model training set, and the adaptive binary cross-entropy loss formula is obtained by sorting:
[0085]
[0086] Among them, is the one-hot label, is the bias term of the logarithm of y i , is the logarithm probability of the y-th i category.
[0087] Step 5: Use the facial expression recognition model to send F and F' into the global average pooling layer to reduce the feature dimension, and then send the dimension-reduced feature f into the fully connected layer to obtain two classification probabilities. After adding the two classification probabilities, use the classification loss to obtain the facial expression classification, and the calculation formula is as follows:
[0088]
[0089] is the weight of yi in the fully connected layer, yi is the label of the i-th given facial image, and f i is the feature of the i-th facial image, C is the number of facial expression categories, and N is the number of samples used to train the model.
[0090] The facial expression recognition model of the present invention can be applied and deployed. Combined with Figure 2 , the model is used in the test or inference stage. The facial image is input, the feature extractor is used to extract features, the features are input into the global average pooling layer for feature processing, then input into the fully connected layer to obtain the classification probability, and then the classification probability is used to perform the expression classification of the facial image, and the result is obtained to end the recognition process.
[0091] Embodiment 2
[0092] The present invention also provides a device for optimizing facial expression recognition, including an enhancement prompt module, a model management module, a semantic consistency module, and an adaptive binary cross-entropy loss module.
[0093] The enhancement prompt module performs data enhancement operations on the facial image I to obtain the randomly erased facial image I' and the randomly erased and randomly flipped facial image I".
[0094] Randomly crop the facial image I”, perform a binary difference operation on the cropped image, and restore the image to the size before cropping to obtain the facial image I t ,
[0095] Create visual cues for the facial image I in the pixel space, add the visual cues to the facial image I to form a set of cue images, and use the Gabor operator to extract the fine-grained texture and edge features of the facial image I and add them to the set of cue images
[0096] The model management module uses a convolutional neural network CNN to construct a facial expression recognition model. The model management module uses the facial expression recognition model to extract the global feature F from the facial image I’, and uses the facial expression recognition model to extract the local feature F” from the facial image I t Extract the fine-grained feature F’ from the set of cue images using the facial expression recognition model
[0097] The semantic consistency module uses the facial expression recognition model to calculate the class activation maps M and M" based on the global feature F and the local feature F”. After performing a random horizontal flip operation on the class activation map M", the class activation map M′ is obtained, and the semantic consistency between M and M′ is compared to compensate for the feature blur caused by pose changes and occlusions
[0098] The adaptive binary cross-entropy loss module uses the facial expression recognition model to add an adaptive factor to the binary cross-entropy loss for long-tail correction to avoid the classification bias of the tail classes caused by noise in the long-tail distribution
[0099] The classification module uses the facial expression recognition model to send F and F’ into the global average pooling layer to reduce the feature dimension, and then sends the reduced-dimensional feature f into the fully connected layer to obtain two classification probabilities. After adding the two classification probabilities, use the classification loss Obtain the facial expression classification The calculation formula is as follows
[0100]
[0101] is the weight of yi in the fully connected layer, yi is the label of the given i-th facial image, and f i is the feature of the i-th facial image, C is the number of facial expression categories, and N is the number of samples used to train the model
[0102] Regarding the information interaction and execution process among the above-mentioned modules in the device, since they are based on the same concept as the method embodiment of the present invention, the specific content can be referred to the description in the method embodiment of the present invention and will not be elaborated here
[0103] Similarly, the device of the present invention performs data enhancement on facial images, can obtain multi-scale enhanced data, and performs fine-grained feature extraction on the original facial images, enhancing the fine-grained feature learning of the feature extractor shared by the facial expression recognition model.
[0104] And by using random erasing, images with random erasing and random horizontal flipping are used to simulate difficult expression recognition situations such as occlusion and pose changes. By using semantic consistency contrast, complete facial images are obtained to solve the problem of difficult expression recognition.
[0105] At the same time, an adaptive binary cross-entropy loss is defined to perform long-tail correction to optimize the problem of misidentifying facial expressions in the tail classes by the facial expression recognition model. Integrating the process of optimizing the model of the present invention reduces the negative impact of noisy labels on facial expression recognition, enhances the anti-noise interference ability of the deep learning model, and improves the facial expression recognition accuracy of the deep learning model in actual application scenarios.
[0106] It should be noted that not all steps and modules in the above-mentioned processes and device structures are necessary, and some steps or modules can be ignored according to actual needs. The execution order of each step is not fixed and can be adjusted according to needs. The system structure described in the above-mentioned embodiments can be a physical structure or a logical structure. That is, some modules may be implemented by the same physical entity, or some modules may be implemented by multiple physical entities respectively, or some components in multiple independent devices can be jointly implemented.
[0107] The above-mentioned embodiments are only preferred embodiments given to fully illustrate the present invention, and the protection scope of the present invention is not limited thereto. Equivalent substitutions or transformations made by those skilled in the art on the basis of the present invention are all within the protection scope of the present invention. The protection scope of the present invention is subject to the claims.
Claims
1. A method for optimizing facial expression recognition, characterized in that Including: Step 1: For the facial image I, perform data augmentation operations to obtain the randomly erased facial image I' and the randomly erased and randomly flipped facial image I". Randomly crop the facial image I”, perform a binary difference operation on the cropped image, and restore the image to the size before cropping to obtain the facial image I t , Create visual cues in the pixel space of the facial image I, add the visual cues to the facial image I to form a set of cue images, and use the Gabor operator to extract the fine-grained texture and edge features of the facial image I and add them to the set of cue images. Step 2: Use a convolutional neural network (CNN) to build a facial expression recognition model. Use the facial expression recognition model to extract global feature F from facial image I’, and use the facial expression recognition model to extract local feature F” from facial image I t Extract fine-grained feature F’ from the prompt image set using the facial expression recognition model Step 3: Use the facial expression recognition model to calculate the class activation maps M and M" based on the global feature F and the local feature F", perform a random horizontal flip operation on the class activation map M" to obtain the class activation map M′, and compare the semantic consistency between M and M′ to compensate for the feature blur caused by pose changes and occlusions. Step 4: Use the facial expression recognition model to perform long-tail correction on the binary cross-entropy loss by adding an adaptive factor to avoid the classification bias of the tail classes caused by noise in the case of long-tail distribution. Step 5: Use the facial expression recognition model to send F and F’ into the global average pooling layer to reduce the feature dimension, and then send the dimension-reduced feature f into the fully connected layer to obtain two classification probabilities. After adding the two classification probabilities, use the classification loss to obtain the facial expression classification, and the calculation formula is as follows: is the weight of yi in the fully connected layer, where yi is the label of the i-th given facial image, and f i is the feature of the i-th facial image, C is the number of facial expression categories, j represents the j-th facial image category, where j ranges from 0 to C - 1, and N is the number of samples used to train the model.
2. The method for optimizing facial expression recognition according to claim 1, characterized in that In Step 1, use RandomErasing() in the image transformation tool transforms of the torchvision tool to perform random erasing on the facial image I to obtain the erased facial image I'. After using RandomErasing() in the image transformation tool transforms of the torchvision tool to perform random erasing on the facial image I, then use RandomHorizontalFlip() to perform a random horizontal flip to obtain the erased and flipped facial image I".
3. The method for optimizing facial expression recognition according to claim 1, characterized in that In step 1, νφ is used to represent the visual cue, and the size of νφ is H, W, and C are the height, width, and number of channels of the facial image I respectively. The visual cue νφ is added to the facial image I to form a set of cue images Then prompt the image set Extract the fine-grained texture and edge features of the facial image I using the Gabor operator and add them to the prompt image set in Then g(x) extracts the fine-grained texture and edge features of the facial image I by using the Gabor operator.
4. A method for optimizing facial expression recognition according to claim 1, wherein the following formula is adopted in Step 3: MSE is the mean squared error, which calculates the loss Utilize the loss Compare the semantic consistency between M and M'.
5. A method for optimizing facial expression recognition according to claim 1, characterized in that The formula for the balanced binary cross-entropy loss in Step 4: Indicates the distribution of the y i category, is the number of the y i category, N is the total number of samples, and an adaptive factor is introduced C is the number of facial image categories, and the formula for the new binary cross-entropy loss is obtained: Among them, is the distribution of the y i -th category in the facial expression recognition model training set, and the adaptive binary cross-entropy loss formula is obtained by sorting: Among them, is a one-hot label, is the bias term of the logarithm y i , is the logarithmic probability of the y i -th category.
6. An apparatus for optimizing facial expression recognition, characterized in that Including an enhanced cue module, a model management module, a semantic consistency module, and an adaptive binary cross-entropy loss module. The enhanced cue module performs data augmentation operations on the facial image I to obtain the randomly erased facial image I' and the randomly erased and randomly flipped facial image I". Randomly crop the facial image I”, perform a binary difference operation on the cropped image, and restore the image to the size before cropping to obtain the facial image I t , Create visual cues in the pixel space of the facial image I, add the visual cues to the facial image I to form a set of cue images, and use the Gabor operator to extract the fine-grained texture and edge features of the facial image I and add them to the set of cue images. The model management module uses a convolutional neural network (CNN) to build a facial expression recognition model. The model management module uses the facial expression recognition model to extract the global feature F from the facial image I’, and uses the facial expression recognition model to extract the local feature F” from the facial image I t and uses the facial expression recognition model to extract the fine-grained feature F’ from the prompt image set. The semantic consistency module uses the facial expression recognition model to calculate the class activation maps M and M" based on the global feature F and the local feature F", performs a random horizontal flip operation on the class activation map M" to obtain the class activation map M′, and compares the semantic consistency between M and M′ to compensate for the feature blur caused by pose changes and occlusions. The adaptive binary cross-entropy loss module uses the facial expression recognition model to perform long-tail correction on the binary cross-entropy loss by adding an adaptive factor to avoid the classification bias of the tail classes caused by noise in the case of long-tail distribution. The classification module uses a facial expression recognition model to send F and F' into the global average pooling layer to reduce the feature dimension, and then sends the dimension-reduced feature f into the fully connected layer to obtain two classification probabilities. After adding the two classification probabilities, the classification loss is used to obtain the facial expression classification, and the calculation formula is as follows: is the weight of yi in the fully connected layer, where yi is the label of the i-th given facial image, and f i is the feature of the i-th facial image, C is the number of facial expression categories, and N is the number of samples used to train the model.
7. The device for optimizing facial expression recognition according to claim 6, characterized in that The enhanced cue module uses RandomErasing() in the image transformation tool transforms of the torchvision tool to perform random erasing on the facial image I to obtain the erased facial image I'. After randomly erasing the facial image I using RandomErasing() in the image transformation tool transforms of torchvision, then randomly horizontally flipping it using RandomHorizontalFlip() to obtain the erased and flipped facial image I”.
8. The device for optimizing facial expression recognition according to claim 6, characterized in that The enhanced hint module uses νφ to represent visual hints, and the size of νφ is H, W, and C are the height, width, and number of channels of the facial image I respectively. The visual hint νφ is added to the facial image I to form a set of hint images Then prompt the image set Extract the fine-grained texture and edge features of the facial image I using the Gabor operator and add them to the hint image set in Then g(x) extracts the fine-grained texture and edge features of the facial image I by means of the Gabor operator.
9. The device for optimizing facial expression recognition according to claim 6, characterized in that The semantic consistency module adopts the following formula: The MSE is the mean squared error, which calculates the loss Utilize the loss Compare the semantic consistency between M and M'.
10. The device for optimizing facial expression recognition according to claim 6, characterized in that it is adaptively binary The cross-entropy loss module is based on the formula of balanced binary cross-entropy loss: Represents the distribution of the y-th i category, is the number of the y-th i category, N is the total number of samples, and an adaptive factor is introduced C is the number of facial image categories, and the formula for the new binary cross-entropy loss is obtained: Among them, is the distribution of the y i -th category in the training set of the facial expression recognition model, and the adaptive binary cross-entropy loss formula is obtained by sorting: Among them, is a one-hot label, is the bias term of the logarithm y i , is the logarithmic probability of the y i th category.