Shielded facial expression recognition method based on multi-angle feature extraction
Through the multi-angle feature extraction method, the area division strategy of key facial feature points and the self-attention weight selection mechanism are used to solve the occlusion problem in facial expression recognition, and efficient occlusion expression recognition is achieved.
Patent Information
- Application Number
- CN202510203925.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-24
- Publication Date
- 2025-07-01
AI Technical Summary
Facial expression recognition technology faces occlusion problems in actual application, resulting in the loss of key facial expression features and affecting the recognition accuracy.
Using a multi-angle feature extraction method, the area division strategy and self-attention weight selection mechanism of key facial feature points are focused on important facial areas, ignoring the occluded areas, and reducing the impact of the occluded areas on the network.
It effectively improves the recognition ability of occluded expressions, enhances the robustness and accuracy of feature extraction, and improves the recognition accuracy and efficiency.
Smart Images

Figure CN120236308A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of image processing and recognition, and particularly relates to a method for occluded facial expression recognition based on multi-angle feature extraction. Background Art
[0002] Facial expressions are one of the most common and important signals for expressing human emotions and intentions, and play an important role in daily interpersonal communication. Facial expression recognition technology has made great progress and wide applications in the field of computer vision. It can be used to recognize people's emotional states, provide a better user experience for human-computer interaction, and has broad application prospects in various fields such as medical and health, driver state detection, and intelligent classrooms, attracting more and more attention. Facial expression recognition has important applications in the fields of sentiment analysis, human-computer interaction, virtual reality, etc. However, facial expression recognition technology still faces some severe challenges in practical applications. Facial occlusion is one of the main challenges for accurate facial expression recognition because occlusion will lead to the loss of key facial expression features, making the model unable to learn complete facial features, and thus seriously affecting the accuracy of facial expression recognition. In real-world scenarios, the facial area of a person is easily occluded by items such as sunglasses, hats, scarves, and masks, which will result in incomplete or inaccurate acquisition of the information of facial expression images, and further make traditional facial expression recognition methods face more severe challenges in practical applications. This occlusion problem will not only affect the accurate judgment of the recognition system on the emotional state, but also have an adverse impact on the applications in fields such as security monitoring and medical diagnosis.
[0003] Since traditional facial expression recognition systems and methods usually rely on the analysis of facial features, such as the shapes and movements of areas like eyes and mouths, to accurately recognize expressions. However, when the face is occluded, these key features may not be accurately extracted, resulting in a decrease in recognition accuracy. Therefore, there is an urgent need to provide a method for accurately recognizing occluded facial expressions. Summary of the Invention
[0004] Aiming at the problems existing in the above-mentioned prior art, the present invention provides a method for occluded facial expression recognition based on multi-angle feature extraction. This method is based on the regional division strategy of facial key feature points and the self-attention weight selection mechanism, can focus on important facial areas, ignore occluded areas, reduce or eliminate the influence of occluded and irrelevant areas of the face on the network, can effectively solve the OFER problem, and can greatly improve the recognition ability of occluded expressions.
[0005] To achieve the above object, the present invention provides a method for occluded facial expression recognition based on multi-angle feature extraction, including the following steps:
[0006] Step 1: Construct an image data set;
[0007] Select a publicly available dataset containing several complete face images. For each complete face image in the publicly available dataset, use Dlib to detect five-point feature points, and perform a cropping operation on the face image based on the five-point detection results to obtain the top-left, top-right, bottom-left, and bottom-right target region images. Then, store the complete face image, as well as the top-left, top-right, bottom-left, and bottom-right target region images, to finally obtain an image dataset;
[0008] Step Two: Construct an occluded facial expression recognition module;
[0009] S21: Construct a multi-feature extraction module;
[0010] Use the SwinTransformerEncoder network to construct a fine-grained extraction branch module;
[0011] Use IResnet-50 as the backbone convolutional neural network, and combine a multi-scale feature pyramid and Transformer-encoder to construct a multi-scale branch PTIR-50 module;
[0012] Combine the fine-grained extraction branch module and the multi-scale branch PTIR-50 module to construct a multi-feature extraction module;
[0013] S22: Construct a regional detail feature fusion module;
[0014] Based on the channel attention network, and introduce the RB-Loss function to construct a regional detail feature fusion module;
[0015] S23: Construct a consistency feature recognition module;
[0016] Based on the Con-feature loss function to construct a consistency feature recognition module;
[0017] S24: Combine the multi-feature extraction module, the regional detail feature fusion module, and the consistency feature recognition module to construct an occluded facial expression recognition module;
[0018] Step Three: Combine the occluded facial expression recognition module and a classifier to construct an occluded facial expression recognition model;
[0019] Step Four: Use a joint optimization strategy to construct a loss function;
[0020] Step Five: Use the training set to train the occluded facial expression recognition model. During the training process, continuously update the model parameters through the backpropagation algorithm to reduce the value of the loss function, and use the test set to test the trained occluded facial expression recognition model to obtain a trained occluded facial expression recognition model;
[0021] Step Six: Use the occluded facial expression recognition model to recognize occluded expressions;
[0022] S61: Use a face detection device to collect occluded face images in complex scenarios;
[0023] S62: Use the occluded facial expression recognition model to recognize occluded expressions;
[0024] S62-1: In the multi-feature extraction module, use the multi-scale branch PTIR-50 module to extract the global features of the face image and obtain the self-attention weights corresponding to the global features. Use the multi-scale branch PTIR-50 module to extract the regional features of the target area image and obtain the self-attention weights corresponding to the regional features. At the same time, use the fine-grained extraction branch module to extract the fine-grained features of the input image, and then send the global features and the self-attention weights corresponding to the global features to the consistency feature recognition module, and send the fine-grained features, regional features, the self-attention weights corresponding to the regional features, and the fine-grained features to the regional detail feature fusion module;
[0025] S62-2: In the regional detail feature fusion module, use the RB-Loss function to guide the self-attention weights corresponding to the important target area features to be greater than the self-attention weights of the global features. Then, select the important target area features from the target area features according to the size of the self-attention weights. Then, fuse the important target area features and the fine-grained features to obtain a fusion feature that contains both important area information and complete image detail information, and then send the fusion feature to the consistency feature recognition module;
[0026] S62-3: In the consistency feature recognition module, use the Con-feature loss function to make the global features and the fusion features guide each other, forcing the multi-scale branch PTIR-50 module to focus on more discriminative expression features to enhance the accuracy of the global feature extraction by the multi-scale branch PTIR-50 module and obtain more discriminative global features, and then input the global features into the classifier;
[0027] S62-4: The classifier recognizes the important facial areas with rich information and not occluded based on the global features and outputs the prediction results.
[0028] As an optimization, in S62-1 of Step Six, obtain the global feature F of the face image X according to formula (1) X and the self-attention Attn X corresponding to the global feature F X , obtain the fine-grained feature F of the face image X according to formula (2) Swin , obtain the top-left image X according to formula (3) LtThe target region feature F Lt and the target region feature F Lt corresponding self-attention Attn Lt , obtain the bottom-left image X according to formula (4) Lb The target region feature F Lb and the target region feature F Lb corresponding self-attention Attn Lb , obtain the top-right image X according to formula (5) Rt The target region feature F Rt and the target region feature F Rt corresponding self-attention Attn Rt , obtain the bottom-right image X according to formula (6) Rb The target region feature F Rb and the target region feature F Rb corresponding self-attention Attn Rb ; obtain the feature set F according to formula (7), obtain the self-attention set A according to formula (8); obtain the weight μ of each attention in the self-attention set A according to formula (9), and obtain the weight set W according to formula (10);
[0029] [F X , Attn X = P(X; θ p )(1);
[0030] F Swin = S(X; θ S )(2);
[0031] [F Lt , Attn Lt = P(X Lt ; θ p )(3);
[0032] [F Lb , Attn Lb = P(X Lb ; θ p )(4);
[0033] [F Rt , Attn Rt = P(X Rt ; θ p )(5);
[0034] [F Rb , Attn Rb = P(X Rb ; θ p )(6);
[0035] F = [FX , F Lb , F Lt , F Rb , F Rt (7);
[0036] A = [Attn X , Attn Lb , Attn Lt , Attn Rb , Attn Rt (8);
[0037] μ = f(Attn, q)(9);
[0038] W = [μ X , μ Lb , μ Lt , μ Rb , μ Rt (10);
[0039] In the formula, P(·; θ p ) is IResnet-50, S(·; θ s ) is SWIN-E, q represents the parameters of the SE Module, f represents the sigmoid function, μ X is the attention weight of the face image X, μ Lb is the attention weight of the bottom-left image X Lb , μ Lt is the attention weight of the top-left image X Lt , μ Rb is the attention weight of the bottom-right image X Rb , μ Rt is the attention weight of the top-right image X Rt .
[0040] As an optimization, in S62-2 of step six, the maximum attention weight W i ;
[0041] W i = maxW(11).
[0042] As an optimization, in S62-2 of step six, the important target region feature F i and the fine-grained feature F Swin are fused according to formula (12), and the fused feature F H is obtained;
[0043] F H = SE(concat(F i , F Swin ))(12);
[0044] In the formula, SE(·) represents the SE Module.
[0045] As an optimization, in S62-2 of step six, the RB-Loss function is as shown in formula (13);
[0046] L RB = max{0, ω - (μ max - μ0)}(13);
[0047] In the formula, ω is a hyperparameter, μ0 is the attention weight of the global feature, and μ max represents the maximum weight of all features.
[0048] As an optimization, in S62-3 of step six, the Con-feature loss function is as shown in formula (14);
[0049]
[0050] In the formula, N represents the number of images, and L represents the length of the feature sequence.
[0051] As an optimization, in step four, the loss function is constructed according to formula (15);
[0052] L train = L cls + αL sm + βL RB + γL con-f (15);
[0053] In the formula, L cls is the classification loss function, L sm is the label smoothing loss function, L RB is the RB-loss loss function, L con-f is the Con-feature loss loss function; where C represents the number of expression categories in the image data, i represents the index of the image data, j and k represent the indices of the expression categories in the image data, x represents the output of the model, x i,j represents the output value of the i-th sample in the j-th category, y i represents the sample label value of the i-th sample, τ represents the smoothing factor, α is the weight hyperparameter of L sm , β is the weight hyperparameter of L RB , and γ is the weight hyperparameter of L con-f .
[0054] As a preference, in step one, the public datasets RAF-DB, FERPlus, Occlusion-FERPlus, and Occlusion-RAFDB are used to construct the image dataset. At the same time, the public datasets RAF-DB and FERPlus are used as the training set, and the public datasets Occlusion-FERPlus and Occlusion-RAFDB are used as the test set.
[0055] The present invention proposes a method for occluded facial expression recognition based on multi-angle feature extraction. Among them, the occluded facial expression recognition module mainly consists of a multi-feature extraction module, a regional detail feature fusion module, and a consistency feature recognition module. Based on this combination, the multi-feature extraction module can obtain the fine-grained information, global information, and important region information of the image, and then can comprehensively use these feature information to perform accurate and efficient expression recognition operations on the occluded facial image. The multi-feature extraction module consists of a fine-grained extraction branch module and a multi-scale branch PTIR-50 module. It can not only use the fine-grained extraction branch module to extract the fine-grained features of the original image, but also use the multi-scale branch PTIR-50 module to simultaneously extract the global features of the original image and the local features of the target region image, and can obtain the self-attention weights corresponding to the global features and the self-attention weights corresponding to the target region features. Since the fine-grained extraction branch module is composed of a Swin Transformer Encoder network, and the multi-scale branch PTIR-50 module is mainly composed of an IResnet-50 network, and the structures and parameters of each network are different, thus, even facing the same task, the features that different networks can extract are different. Therefore, the multi-feature extraction module can learn more different features. Through the setting of the regional detail feature fusion module, it can capture important facial regions to solve the FER task with occlusion and pose changes. At the same time, the RB-loss can be used to optimize the self-attention weights, improve the self-attention weights of important regions, and can select important target region features according to the self-attention weights of each feature. Further, the important target region features and the fine-grained features can be fused, greatly enriching the feature expression, and obtaining a fusion feature that simultaneously contains important region information and complete image detail information. Through the setting of the consistency feature recognition module, the Con-feature loss function can be used to process the features extracted from different branches of the same image, effectively ensuring the consistency of the global features and the fusion features, realizing the mutual guidance between features to prompt the model to capture more detailed and accurate expression information, enabling the network to learn more consistent features, so as to make full use of the global information, important target region information, and detail information of the image for expression recognition, enhancing the robustness and accuracy of feature extraction, obtaining a more discriminative and better consistency feature representation, which can enhance the sensitivity of the model to the subtle differences of different expressions, and at the same time effectively prevent feature redundancy from occurring during the feature fusion process, thus facilitating the improvement of the recognition accuracy and efficiency. After obtaining the consistent feature representation, inputting it into the classifier can enable the recognition module to focus on important facial regions and ignore the occluded regions, thereby greatly improving the recognition accuracy and efficiency. In addition, during the training process of the recognition model, a joint optimization strategy is adopted, effectively ensuring the performance of the recognition model.
[0056] The present invention utilizes a dual-branch feature extraction network and a loss function to guide the network to focus on more important and detailed facial expression regions, innovatively implementing a regional division strategy based on facial key feature points and a self-attention weight selection mechanism, enabling the model to focus on important facial regions and ignore occluded regions, thereby reducing or eliminating the influence of occluded and irrelevant facial regions on the network, effectively solving the OFER problem, and greatly improving the recognition ability of occluded expressions. Brief Description of the Drawings
[0057] Figure 1 is a flowchart of the present invention;
[0058] Figure 2 is a structural schematic diagram of the multi-scale extraction branch module in the present invention;
[0059] Figure 3 is a schematic diagram of the cropping process of the input image by the image cropping module in the present invention;
[0060] Figure 4 are partial pictures in the JAFFE facial expression recognition dataset;
[0061] Figure 5 is a schematic diagram of different features extracted from different regional blocks;
[0062] Figure 6 is a schematic diagram of the effect of using the Con-feature loss function, where a shows the effect of F H and F X attention, and b shows the effect of using con-feature loss attention;
[0063] Figure 7 is a visualization comparison schematic diagram of the feature maps after con-feature loss and feature fusion;
[0064] Figure 8 is a visualization schematic diagram of the attention of some FER networks on OFER samples;
[0065] Figure 9 is a comparison schematic diagram of the prediction results of the recognition method in the present invention, the prediction results of the baseline PTIR-50, and the sample labels. Detailed Embodiment
[0066] The present invention will be further described below with reference to the accompanying drawings.
[0067] As Figures 1 to 9As shown in the figure, the present invention provides a method for occluded facial expression recognition based on multi-angle feature extraction (MAFE), including the following steps:
[0068] Step 1: Construct an image data set;
[0069] Select a publicly available data set containing several complete face images. For each complete face image in the publicly available data set, use Dlib to detect five-point feature points (left eye feature point, right eye feature point, nose feature point, left mouth corner feature point, right mouth corner feature point), and perform a cropping operation on the face image based on the five-point detection result to obtain target region images of the upper left corner (Lt), upper right corner (Rt), lower left corner (Lb), and lower right corner (Rb). This cropping operation can not only help find the regions in the image that are more important for expression recognition, but also create pictures similar to the situation of facial occlusion, thereby enriching the data set and improving the robustness of the model to OFER samples. Then store the complete face image, as well as the target region images of the upper left corner (Lt), upper right corner (Rt), lower left corner (Lb), and lower right corner (Rb), and finally obtain the image data set;
[0070] Through the image cropping operation, it is convenient to crop the original image into multiple target region images containing multiple feature points. In this way, it is convenient to extract the global features of the original image and the local features of the target region images during the subsequent feature extraction process. This cropping operation can not only help find the regions in the image that are more important for expression recognition, but also create pictures similar to the situation of facial occlusion, thereby effectively achieving the purpose of enriching the data set and being beneficial to improving the robustness of the model to OFER samples.
[0071] Step 2: Construct an occluded facial expression recognition module;
[0072] S21: Construct a multi-feature extraction module;
[0073] Generally speaking, image recognition mainly refers to coarse-grained image recognition, such as cat and dog recognition. However, there are more detailed classification tasks in the field of image recognition, which are called fine-grained image recognition. The main purpose of this recognition problem is to make a more detailed subclass division of images belonging to the same basic category. For example, distinguish the breeds of kittens. The research on fine-grained recognition can be extended to the FER field because differentiating human faces with different expressions under the large category of human faces is essentially a more detailed subclass division of face images. Figure 5Some pictures from the JAFFE facial expression recognition dataset are shown, which can be regarded as happy people, sad people, scared people, etc. In addition, the FER task is different from other recognition tasks. Facial expression features are more hidden in the subtle areas of the human face. In order to extract more fine-grained features, the Swin Transformer Encoder (SWIN-E) network is used to construct a fine-grained extraction branch module; the fine-grained extraction branch module is used to extract fine-grained features from the original image;
[0074] IResnet-50 is used as the backbone convolutional neural network, and a multi-scale branch PTIR-50 module is constructed by combining a multi-scale feature pyramid and a Transformer-encoder; the multi-scale branch PTIR-50 module is used to extract global features and the corresponding self-attention weights of the global features from the original image. At the same time, it is used to extract target region features and the corresponding self-attention weights of the target region features from the target region image. The purpose is to pay attention to the fine-grained information of the image while obtaining the global information of the image, enrich the feature representation of the image, and make the final feature expression more stable and comprehensive; among them, the multi-scale feature pyramid and the Transformer-encoder module in the multi-scale branch PTIR-50 network can capture the long-range dependence relationship in the image while capturing the multi-scale information of the image, making the final feature representation more comprehensive and accurate. In the actual operation process, after the input image is processed by IResnet-50, the multi-scale feature pyramid structure technology is used to generate multi-scale feature representations. Specifically, three levels of extraction features, large, medium, and small, can be constructed, and the three levels of extraction features are input into the Transformer-encoder to capture the feature scales of the features, and the output of the Transformer-encoder is aggregated to form the feature (feature) and self-attention (attention) of the input image respectively.
[0075] A multi-feature extraction module is constructed by combining the fine-grained extraction branch module and the multi-scale branch PTIR-50 module;
[0076] S22: Construct a regional detail feature fusion module;
[0077] Based on the channel attention network and introducing the RB-Loss function, a regional detail feature fusion module is constructed; the regional detail feature fusion module is used to guide the self-attention weight corresponding to the important target region feature to be greater than the self-attention weight corresponding to the original image according to the regional deviation loss function (RB-Loss), and select the important target region feature from the target region features based on the size of the self-attention weight, and fuse the important target region feature and the fine-grained feature to obtain a fusion feature that contains both important target region information and complete image detail information;
[0078] S23: Construct a consistency feature recognition module;
[0079] Construct a consistency feature recognition module based on the Con-feature loss function; the consistency feature recognition module is used to make the global feature and the fusion feature guide each other through the Con-feature loss function to obtain a more consistent and better feature representation;
[0080] Figure 9 The visualization of the feature map after Con-feature loss and feature fusion is shown. The left figure of feature fusion pays more attention to areas such as the facial contour and eyebrows, and the right figure of Con-feature loss helps to focus the attention on the facial features. It can be seen that the visualization of Con-feature loss pays attention to places with more expression features, while simple feature fusion pays attention to some areas irrelevant to expression recognition.
[0081] S24: Combine the multi-feature extraction module, the regional detail feature fusion module and the consistency feature recognition module to construct an occluded facial expression recognition module;
[0082] Step 3: Combine the occluded facial expression recognition module and a classifier to construct an occluded facial expression recognition model;
[0083] Step 4: Adopt a joint optimization strategy to construct a loss function;
[0084] Step 5: Use the training set to train the occluded facial expression recognition model. During the training process, continuously update the model parameters through the backpropagation algorithm to reduce the value of the loss function, and use the test set to test the trained occluded facial expression recognition model to obtain a trained occluded facial expression recognition model;
[0085] Step 6: Use the occluded facial expression recognition model to recognize occluded expressions;
[0086] S61: Use a face detection device to collect occluded face images in a complex scene;
[0087] S62: Use the occluded facial expression recognition model to recognize occluded expressions;
[0088] S62-1: In the multi-feature extraction module, the global features of the face image are extracted using the multi-scale branch PTIR-50 module, and the self-attention weights corresponding to the global features are obtained. The regional features of the target region image are extracted using the multi-scale branch PTIR-50 module, and the self-attention weights corresponding to the target region features are obtained. At the same time, the fine-grained features of the input image are extracted using the fine-grained extraction branch module, and then the global features and the self-attention weights corresponding to the global features are sent to the consistency feature recognition module, and the fine-grained features, regional features, the self-attention weights corresponding to the regional features, and the fine-grained features are sent to the regional detail feature fusion module;
[0089] Figure 5 Shows the schematic diagrams of different features extracted from different regional blocks. It can be seen from this that the focus points of different regions are different, and the features that can be extracted are also different.
[0090] S62-2: In the regional detail feature fusion module, the RB-Loss function is used to guide the self-attention weight corresponding to the important target region feature to be greater than the self-attention weight of the global feature. Then, the important target region features are selected from the target region features according to the magnitude of the self-attention weights. Then, the important target region features and the fine-grained features are fused to obtain the fusion features that simultaneously contain important region information and complete image detail information, and then the fusion features are sent to the consistency feature recognition module;
[0091] S62-3: In the consistency feature recognition module, the Con-feature loss function is used to make the global features and the fusion features guide each other, forcing the multi-scale branch PTIR-50 module to focus on more discriminative expression features to enhance the accuracy of the global feature extraction by the multi-scale branch PTIR-50 module and obtain more discriminative global features, and then the global features are input into the classifier;
[0092] S62-4: The classifier identifies the important facial regions with rich information and unobstructed according to the global features and outputs the prediction results.
[0093] As an optimization, in S62-1 of step six, the global feature F of the face image X is obtained according to formula (1) X and the self-attention Attn X corresponding to the global feature F X , the fine-grained feature F of the face image X is obtained according to formula (2) Swin , the target region feature F of the upper left corner image X Lt is obtained according to formula (3), and the self-attention Attn Lt corresponding to the target region feature F Lt and the target region feature F Lt, according to formula (4), we can get the lower left corner image X Lb The target area features F Lb and target region feature F Lb The corresponding self-attention Attn Lb , according to formula (5), we can get the upper right corner image X Rt The target area features F Rt and target region feature F Rt The corresponding self-attention Attn Rt , according to formula (6), we can get the lower right corner image X Rb The target area features F Rb and target region feature F Rb The corresponding self-attention Attn Rb ; Obtain the feature set F according to formula (7), and obtain the self-attention set A according to formula (8); Use SEModule and sigmoid function to process each attention in the self-attention set A obtained in the Multi-Feature Extraction Module into an attention weight μ. Specifically, obtain the weight μ of each attention in the self-attention set A according to formula (9), and obtain the weight set W according to formula (10);
[0094] [F X ,Attn X ]=P(X;θ p )(1);
[0095] F Swin =S(X;θ S )(2);
[0096] [F Lt ,Attn Lt ]=P(X Lt θ p )(3);
[0097] [F Lb ,Attn Lb ]=P(X Lb θ p )(4);
[0098] [F Rt ,Attn Rt ]=P(X Rt θ p )(5);
[0099] [F Rb ,Attn Rb ]=P(X Rb θ p )(6);
[0100] F = [F X , F Lb , F Lt , F Rb , F Rt (7);
[0101] A = [Attn X , Attn Lb , Attn Lt , Attn Rb , Attn Rt (8);
[0102] μ = f(Attn, q)(9);
[0103] W = [μ X , μ Lb , μ Lt , μ Rb , μ Rt (10);
[0104] Wherein, P(·; θ p ) is IResnet-50, S(·; θ s ) is SWIN-E, q represents the parameters of the SE Module, f represents the sigmoid function, μ X is the attention weight of the face image X, μ Lb is the attention weight of the bottom-left image X Lb , μ Lt is the attention weight of the top-left image X Lt , μ Rb is the attention weight of the bottom-right image X Rb , μ Rt is the attention weight of the top-right image X Rt .
[0105] As a preference, in S62-2 of step six, the maximum attention weight W i is obtained according to formula (11); by taking the maximum value of W, the maximum attention weight W i and the corresponding important region feature F i are obtained;
[0106] W i = maxW(11).
[0107] As a preference, in S62-2 of step six, the important target region feature F i and the fine-grained feature F Swin are fused according to formula (12), and the fused feature F H, the fused feature contains both the fine-grained feature of X and the important region feature of X;
[0108] F H = SE(concat(F i , F Swin ))(12);
[0109] In the formula, SE(·) represents the SE Module.
[0110] For the fused feature F of the obtained input image X H and the image feature F X , in order to enable the network to learn more accurate and consistent features, the Con-feature loss is used to optimize the feature selection process of the model. Finally, the image feature F X is input into the multi-layer perceptron (MLP) to return the predicted emotion label Y.
[0111] As an optimization, in S62-2 of step six, the RB-Loss function is as shown in formula (13); since different facial expressions are mainly defined by different facial regions, using RB-loss can force the attention weight of important regions to be greater than the attention weight of the original facial image, thereby reducing the influence of other unimportant regions on the model and improving the performance of the model.
[0112] L RB = max{0, ω - (μ max - μ0)}(13);
[0113] In the formula, ω is a hyperparameter, μ0 is the attention weight of the global feature, and μ max represents the maximum weight of all features.
[0114] As an optimization, in S62-3 of step six, the Con-feature loss function is as shown in formula (14);
[0115]
[0116] In the formula, N represents the number of images, and L represents the length of the feature sequence.
[0117] As an optimization, in step four, the loss function is constructed according to formula (15); since there are a large number of ambiguous expression images in the FER dataset, these expressions will cause the model to have uncertainty for some samples, thereby affecting the model's judgment of this type of image. To reduce this influence, the label smoothing strategy Loss sm is innovatively introduced. The label smoothing gives these images a certain tolerance probability, which can prevent the model from overly believing in the sample labels.
[0118] L train = L cls + αL sm + βL RB + γL con-f (15);
[0119] In the formula, L cls is the classification loss function, L sm is the label smoothing loss function, L RB is the RB-loss loss function, L con-f is the Con-feature loss loss function; where C represents the number of expression categories in the image data, i represents the index of the image data, j and k represent the indices of the expression categories in the image data, x represents the output of the model, x i,j represents the output value of the i-th sample on the j-th category, y i represents the sample label value of the i-th sample, τ represents the smoothing factor, α is the weight hyperparameter of L sm β is the weight hyperparameter of L RB γ is the weight hyperparameter of L con-f of the weight hyperparameter.
[0120] As a preference, in step one, the public datasets RAF-DB, FERPlus, Occlusion-FERPlus, and Occlusion-RAFDB are used to construct the image dataset. Meanwhile, the public datasets RAF-DB and FERPlus are used as the training set, and the public datasets Occlusion-FERPlus and Occlusion-RAFDB are used as the test set.
[0121] Among them, RAF-DB (Real-world Affective Faces Database) is the first real-world face expression dataset containing basic or composite expressions. The images in this dataset have great variability in terms of age, gender, race, head pose, lighting conditions, occlusion (such as glasses, facial hair, or self-occlusion), post-processing operations (such as various filters and special effects), etc. The images of the six basic expressions (happy, surprised, sad, angry, disgusted, fearful) and neutral expressions in the dataset are used for experiments.
[0122] FERPlus is an extension of the FER2013 dataset for the ICML2013 challenge. FERPlus consists of a large-scale facial expression image collection by Google Search Engine and new labels provided by Microsoft for FER2013. It contains 28,709 training images, 3,589 validation images, and 3,589 test images, with a size of 48×48 pixels. The main difference between FER2013 and FERPlus lies in the annotations. FER2013 was annotated with seven facial expression labels (neutral, happy, surprised, sad, angry, disgusted, fearful) by one annotator, while FERPlus added the contempt label and was annotated by 10 annotators.
[0123] Occlusion-RAFDB is a occluded test set manually annotated based on the occlusion types (non-occluded, wearing a mask, wearing glasses, left / right object, upper face object, and lower face object) on the test set of the original RAF-DB dataset, and images with at least one occlusion type are selected. There are 735 images in total, including corresponding facial expression annotations and occlusion type annotations. Only the facial expression annotations in the test set are used.
[0124] Similarly, Occlusion-FERPlus is a occluded test set manually annotated based on the occlusion types (non-occluded, wearing a mask, wearing glasses, left / right object, upper face object, and lower face object) on the test set of the original FERPlus dataset, and images with at least one occlusion type are selected. There are 605 images in total, including corresponding facial expression annotations and occlusion type annotations. Only the facial expression annotations in the test set are used.
[0125] 4.2 Implementation details
[0126] Verification process:
[0127] 1. Use the recognition method in the present invention and the baseline PTIR-50 to verify the known sample labels. The verification results are as Figure 1 shown. It can be seen from Figure 2 that the prediction results of the present invention are consistent with the true state of the sample labels;
[0128] 2. Complete the experiment using pytorch on two Nvidia Tesla V100 GPUs. During the experiment, the sizes of the input original images and face region images are both adjusted to 224*224, L trainThe parameters in it are set as α = 2, β = 1, γ = 2. Initialize the learning rate to 0.000025 and stop training at the 70th epoch. Compare the performance of MAFE with other methods on the Occlusion-FERPlus and Occlusion-RAFDB datasets. Summary on Occlusion-RAFDB: Table 1 lists the performance of the methods proposed in the FER field in the past five years from 2020 to 2024 on the Occlusion-RAFDB dataset. MAFE performs excellently in terms of accuracy, reaching 89.42%. Summary on Occlusion-FERPlus: Table 1 lists the performance of the methods proposed in the FER field in the past five years from 2020 to 2024 on the Occlusion-FERPlus dataset. MAFE performs excellently in terms of accuracy, reaching 86.94%.
[0129] In addition, MAFE also achieved good performance on the original RAF-DB and FERPlus datasets. Summary on RAFDB: Table 2 lists the performance of the methods proposed in the FER field in the past five years from 2020 to 2024 on the RAFDB dataset. MAFE performs excellently in terms of accuracy, reaching 92.11%, which is 2.51% and 2.57% higher than the previous Latent-OFER and MPA respectively. Summary on FERPlus: Table 2 lists the performance of the methods proposed in the FER field in the past five years from 2020 to 2024 on the FERPlus dataset. MAFE performs excellently in terms of accuracy, reaching 90.15%, which is 0.73% and 1.02% higher than the previous SCAN-CCI and MPA respectively.
[0130] Table 1: Schematic diagram of the performance comparison between MAFE and other methods on the Occlusion-FERPlus and Occlusion-RAFDB datasets
[0131]
[0132] Table 2: Schematic diagram of the performance comparison between MAFE and other methods on the RAF-DB and FERPlus datasets
[0133]
[0134] 3. Ablation experiment;
[0135] To separately explore the role of each part in the MAFE structure, ablation experiments were designed and conducted on the RAF-DB dataset and the Occlusion-RAFDB dataset. The results are shown in Table 3. Observing the table, it is found that compared with using only PTIR-50 or SWIN-E for facial expression recognition, MAFE achieved good performance on both the original RAF-DB dataset and the Occlusion-RAFDB dataset, with accuracies reaching 92.11% and 89.42% respectively.
[0136] Table 3: Ablation experiment data of MAFE
[0137]
[0138]
[0139] 4. To demonstrate the role of con-feature loss, a comparative experiment between con-feature loss and feature fusion was conducted. Specifically, the performance of processing F H and F X using con-feature loss was compared with the performance of fusing F H and F X . The specific results are shown in Table 4. It can be seen from the table that the method of using Con-feature loss to process F H and F X , which enables the model to learn more accurate features, has a great advantage in accuracy compared with the method of fusing F H and F X and then inputting them into MLP. It was found during the training process that the time taken to train one epoch using con-feature loss is also three minutes shorter than that using feature fusion.
[0140] Table 4: Data comparing feature loss and feature fusion
[0141]
[0142] 5. An in-depth study was conducted on the four cut-out regions to explore the impact of important regions on the model performance. The features F X , F Lb , F Lt , F Rb , F Rt , F Swin were separately combined with the fine-grained feature F SwinFeature fusion is performed. The training set of the RAF-DB dataset is used for training, and the test set of the RAF-DB dataset and the Occlusion-RAFDB dataset are used for testing. The experimental results are shown in Table 5. Max in the table represents the selected important region. It can be seen from Table 5 that the strategy of fusing important region information and fine-grained information is superior to other fusion strategies in the table, and the best results are achieved on the RAF-DB and Occlusion-RAFDB datasets, which are 92.11% and 89.42% respectively.
[0143] Table 5: Data on the impact of different regions on model performance
[0144]
[0145] 6. Visual comparison;
[0146] To demonstrate the effectiveness of MAFE for the OFER problem, attention visualization is performed on some face-occluded samples. Figure 8 Attention visualizations of some FER networks on OFER samples are respectively shown. It can be intuitively seen from the figure the superiority of MAFE in the present invention on OFER samples. In Figure 8 , the first row shows some OFER sample pictures. The next two rows respectively show the attention of FDRL and ARM on OFER samples, and then the attention of PTIR-50 is shown. It can be seen that these three networks also invest a lot of attention in the occluded parts of OFER samples, resulting in poor recognition accuracy. In the pictures shown by MAFE in the last row, almost no attention is invested in the occluded parts, indicating that MAFE has certain superiority in dealing with OFER samples.
[0147] The present invention proposes a method for occluded facial expression recognition based on multi-angle feature extraction. Among them, the occluded facial expression recognition module mainly consists of a multi-feature extraction module, a regional detail feature fusion module, and a consistency feature recognition module. Based on this combination, the multi-feature extraction module can obtain the fine-grained information, global information, and important region information of the image, and then can comprehensively utilize these feature information to perform accurate and efficient expression recognition operations on the occluded facial image. The multi-feature extraction module consists of a fine-grained extraction branch module and a multi-scale branch PTIR-50 module. It can not only extract the fine-grained features of the original image using the fine-grained extraction branch module, but also extract the global features of the original image and the local features of the target region image simultaneously using the multi-scale branch PTIR-50 module, and can obtain the self-attention weights corresponding to the global features and the self-attention weights corresponding to the target region features. Since the fine-grained extraction branch module is composed of a Swin Transformer Encoder network, and the multi-scale branch PTIR-50 module is mainly composed of an IResnet-50 network, and the structures and parameters of each network are different. In this way, even for the same task, the features that different networks can extract are different. Therefore, the multi-feature extraction module can learn more different features. Through the setting of the regional detail feature fusion module, it can capture important facial regions to solve the FER task with occlusion and pose changes. At the same time, the RB-loss can be used to optimize the self-attention weights, improve the self-attention weights of important regions, and can select important target region features according to the self-attention weights of each feature. Further, the important target region features and the fine-grained features can be fused, greatly enriching the feature expression, and obtaining fusion features that simultaneously contain important region information and complete image detail information. Through the setting of the consistency feature recognition module, the Con-feature loss function can be used to process the features extracted from different branches of the same image, effectively ensuring the consistency of the global features and the fusion features, realizing the mutual guidance between features to prompt the model to capture more detailed and accurate expression information, enabling the network to learn more consistent features, so as to fully utilize the global information, important target region information, and detail information of the image for expression recognition, enhancing the robustness and accuracy of feature extraction, obtaining more discriminative and better consistency feature representations, which can enhance the sensitivity of the model to the subtle differences of different expressions, and at the same time effectively prevent feature redundancy from occurring during the feature fusion process, thus being beneficial to improving the accuracy and efficiency of recognition. After obtaining the consistent feature representation, inputting it into the classifier can enable the recognition module to focus on important facial regions and ignore the occluded regions, thus greatly improving the accuracy and efficiency of recognition. In addition, during the training process of the recognition model, a joint optimization strategy is adopted, effectively ensuring the performance of the recognition model.
[0148] The present invention utilizes a dual-branch feature extraction network and a loss function to guide the network to focus on more important and detailed facial expression regions, innovatively implementing a regional division strategy based on facial key feature points and a self-attention weight selection mechanism, enabling the model to focus on important facial regions and ignore occluded regions, thereby reducing or eliminating the impact of occluded and irrelevant facial regions on the network, effectively solving the OFER problem and greatly improving the recognition ability of occluded expressions.
Claims
1. A method for recognizing occluded facial expressions based on multi-angle feature extraction, characterized in that: The following steps are involved: Step 1: Build an image dataset; A public dataset containing several complete face images is selected. For each complete face image in the public dataset, five-point feature point detection is performed using Dlib. Based on the five-point detection results, the face image is cropped to obtain the upper left corner, upper right corner, lower left corner and lower right corner target area images. The complete face image and the upper left corner, upper right corner, lower left corner and lower right corner target area images are then stored to finally obtain an image dataset. Step 2: Build an occluded facial expression recognition module; S21: construct a multi-feature extraction module; The SwinTransformerEncoder network is used to construct a fine-grained extraction branch module; IResnet-50 is used as the backbone convolutional neural network, and a multi-scale feature pyramid and Transformer-encoder are combined to build a multi-scale branch PTIR-50 module; A multi-feature extraction module is constructed by combining the fine-grained extraction branch module and the multi-scale branch PTIR-50 module; S22: constructing a regional detail feature fusion module; Based on the channel attention network, the RB-Loss function is introduced to build a regional detail feature fusion module; S23: building a consistent feature recognition module; Construct a consistent feature recognition module based on the Con-feature loss function; S24: Combining the multi-feature extraction module, the regional detail feature fusion module and the consistency feature recognition module to form an occluded facial expression recognition module; Step 3: Combine the occluded facial expression recognition module and the classifier to build an occluded facial expression recognition model; Step 4: Use joint optimization strategy to construct loss function; Step 5: Use the training set to train the occluded facial expression recognition model. During the training process, the model parameters are continuously updated through the back propagation algorithm to reduce the value of the loss function, and the trained occluded facial expression recognition model is tested using the test set to obtain a trained occluded facial expression recognition model; Step 6: Using the occluded facial expression recognition model to recognize the occluded expression; S61: using a face detection device to collect occluded face images in complex scenes; S62: recognizing the occluded facial expression using the occluded facial expression recognition model; S62-1: In the multi-feature extraction module, the global features of the face image are extracted using the multi-scale branch PTIR-50 module, and the self-attention weights corresponding to the global features are obtained. The regional features of the target region image are extracted using the multi-scale branch PTIR-50 module, and the self-attention weights corresponding to the target region features are obtained. At the same time, the fine-grained features of the input image are extracted using the fine-grained extraction branch module, and the global features and the self-attention weights corresponding to the global features are sent to the consistency feature recognition module. The fine-grained features, regional features, the self-attention weights corresponding to the regional features, and the fine-grained features are sent to the regional detail feature fusion module; S62-2: In the regional detail feature fusion module, the RB-Loss function is used to guide the self-attention weight corresponding to the important target region feature to be greater than the self-attention weight of the global feature, and then the important target region feature is selected from the target region features according to the size of the self-attention weight. Then, the important target region feature and the fine-grained feature are fused to obtain a fused feature that contains both the important region information and the complete image detail information, and then the fused feature is sent to the consistency feature recognition module; S62-3: In the consistency feature recognition module, the Con-feature loss function is used to make the global features and the fusion features guide each other, forcing the multi-scale branch PTIR-50 module to focus on more discriminative expression features, so as to enhance the accuracy of the multi-scale branch PTIR-50 module in extracting global features and obtain more discriminative global features, which are then input into the classifier; S62-4: The classifier identifies important facial areas that are rich in information and not blocked based on global features, and outputs a prediction result.
2. The method for recognizing occluded facial expressions based on multi-angle feature extraction according to claim 1, characterized in that: In step 6 S62-1, the global feature F of the face image X is obtained according to formula (1): X and the global feature F X The corresponding self-attention Attn X According to formula (2), we can obtain the fine-grained features F of the face image X Swin , according to formula (3), we can get the upper left corner image X Lt The target area features F Lt and target region feature F Lt The corresponding self-attention Attn Lt , according to formula (4), we can get the lower left corner image X Lb The target area features F Lb and target region feature F Lb The corresponding self-attention Attn Lb , according to formula (5), we can get the upper right corner image X Rt The target area features F Rt and target region feature F Rt The corresponding self-attention Attn Rt , according to formula (6), we can get the lower right corner image X Rb The target area features F Rb and target region feature F Rb The corresponding self-attention Attn Rb ; Obtain the feature set F according to formula (7), obtain the self-attention set A according to formula (8); Obtain the weight μ of each attention in the self-attention set A according to formula (9), and obtain the weight set W according to formula (10); [F X ,Attn X ]=P(X;θ p )(1); F Swin =S(X;θ S )(2); [F Lt ,Attn Lt ]=P(X Lt ;θ p )(3); [F Lb ,Attn Lb ]=P(X Lb ;θ p )(4); [F Rt ,Attn Rt ]=P(X Rt ;θ p )(5); [F Rb ,Attn Rb ]=P(X Rb ;θ p )(6); F=[F X ,F Lb ,F Lt ,F Rb ,F Rt ](7); A=[Attn X ,Attn Lb ,Attn Lt ,Attn Rb ,Attn Rt ](8); μ=f(Attn,q)(9); W=[μ X ,m Lb ,m Lt ,m Rb ,m Rt ](10); In the formula, P(·;θ p ) is IResnet-50, S(·;θ s ) is SWIN-E, q represents the parameters of SE Module, f represents the sigmoid function, μ X is the attention weight of face image X, μ Lb X is the lower left corner of the image Lb The attention weight, μ Lt X is the upper left corner of the image Lt The attention weight, μ Rb X is the lower right corner image Rb The attention weight, μ Rt X is the upper right corner of the image Rt The attention weight.
3. The method for recognizing occluded facial expressions based on multi-angle feature extraction according to claim 1, characterized in that: In step 6 S62-2, the maximum attention weight W is obtained according to formula (11): i ; IN i =maxW(11)。 4. The method for recognizing occluded facial expressions based on multi-angle feature extraction according to claim 1, characterized in that: In step 6 S62-2, the important target area feature F is converted into i and fine-grained features F Swin Fusion is performed and the fusion feature F is obtained H ; F H =SE(concat(F i ,F Swin ))(12); Where, SE(·) represents SE Module.
5. The method for recognizing occluded facial expressions based on multi-angle feature extraction according to claim 1, characterized in that: In step 6 S62-2, the RB-Loss function is as shown in formula (13); L RB =max{0,ω-(μ max -μ0)}(13); Where ω is a hyperparameter, μ0 is the attention weight of the global feature, and μ max Represents the maximum weight of all features.
6. The method for recognizing occluded facial expressions based on multi-angle feature extraction according to claim 1, characterized in that: In step 6 S62-3, the Con-feature loss function is as shown in formula (14); Where N represents the number of images and L represents the length of the feature sequence.
7. The method for recognizing occluded facial expressions based on multi-angle feature extraction according to claim 1, characterized in that: In step 4, the loss function is constructed according to formula (15); L train =L cls +αL sm +βL RB +γL con-f (15); Where, L cls is the classification loss function, L sm is the label smoothing loss function, L RB is the RB-loss loss function, L con-f is the Con-feature loss function; where C represents the number of expression categories in the image data, i represents the index of the image data, j and k represent the index of the expression category in the image data, x represents the output of the model, and x i,j represents the output value of the i-th sample in the j-th category, y i represents the sample label value of the i-th sample, τ represents the smoothing factor, and α is L sm The weight hyperparameter of L RB The weight hyperparameter of L con-f The weight hyperparameters of .
8. The method for recognizing occluded facial expressions based on multi-angle feature extraction according to claim 1, characterized in that: In step one, the public datasets RAF-DB, FERPlus, Occlusion-FERPlus and Occlusion-RAFDB are used to construct image datasets. At the same time, the public datasets RAF-DB and FERPlus are used as training sets, and the public datasets Occlusion-FERPlus and Occlusion-RAFDB are used as test sets.