Facial expression recognition method in shielding environment
By introducing a multi-level automatic encoder and attention mechanism into the ResNet-18 network, the attention mapping of basic expression features is calculated, and the problem of reduced accuracy and robustness of facial expression recognition in the wild environment is solved, achieving higher recognition accuracy and robustness.
Patent Information
- Application Number
- CN202510304827.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-14
- Publication Date
- 2025-07-01
AI Technical Summary
When facial expression recognition is performed in outdoor environments, it is affected by factors such as occlusion, posture changes, and uneven light sources, resulting in a decrease in recognition accuracy and robustness.
The lightweight backbone network ResNet-18 is used to obtain basic expression features, and the displacement, channel and spatial attention feature mapping is calculated through a multi-level automatic encoder and attention mechanism. Combining the compactness loss and softening classification loss, facial expression recognition results are output.
It significantly improves the accuracy and robustness of facial expression recognition when there is occlusion, and can more effectively deal with complex light sources and posture changes in wild environments.
Smart Images

Figure CN120236312A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical fields of artificial intelligence and emotion computing, and particularly relates to a facial expression recognition method in an occluded environment. Background Art
[0002] With the maturity of technology, the application of emotion recognition technology has become increasingly widespread and has great potential and value. With the development and progress of emotion recognition technology, facial expression recognition technology has received increasing attention and plays an important role in emotion recognition systems. As one of the most direct and natural ways for humans to express emotions, facial expressions are also highly distinguishable and easy to collect, so they are widely used in various scenarios.
[0003] Most early facial expression recognition was carried out in a laboratory environment. During the early research on facial expression recognition, the training of most models mainly relied on small controlled expression datasets. In recent years, with the increasing demand for recognizing large-scale real emotional expressions, the ability to perform facial expression recognition in the wild has received increasing attention. However, once facial expression recognition is carried out in the wild, its accuracy will be greatly reduced due to problems such as incomplete feature extraction.
[0004] Facial expression recognition in the wild is interfered by many factors. For example, the posture and movement of the sample during recognition, and the hair of the sample may block some of its facial features; different from the cold light illumination environment in the laboratory, in the wild environment, there are various light sources, the light distribution is uneven, and in some environments, the light sources are mainly warm colors, which also greatly interferes with the completion of the entire experiment. Summary of the Invention
[0005] In order to solve the deficiencies of the existing facial expression recognition methods in the wild environment, the purpose of the present invention is to provide a facial expression recognition method in an occluded environment, which can significantly improve the accuracy and robustness of facial expression recognition when occlusion exists.
[0006] To achieve the above purpose, the present invention adopts the following technical solutions:
[0007] A facial expression recognition method in an occluded environment, comprising the following steps:
[0008] S1: For a given facial emotion recognition sample, use a lightweight backbone network ResNet-18 to obtain the basic expression features of the sample, and at the same time construct the intra-class and inter-class spatial distributions of the basic expression features;
[0009] S2: Calculate the displacement attention feature map, channel attention feature map and spatial attention feature map of the basic expression features, and obtain the attention map of the basic expression features;
[0010] S3: Attach emotional labels to the obtained attention maps;
[0011] S4: Calculate the combined loss and output the facial expression recognition result of the sample based on the combined loss and the attention map.
[0012] In S1, a lightweight backbone network ResNet-18 is used to obtain the basic expression features of the sample, and constructing the intra-class and inter-class spatial distributions of the basic expression features includes the following steps:
[0013] S11: From the input samples, use the lightweight backbone network ResNet-18 to obtain the basic expression features of the sample;
[0014] S12: Establish a multi-level autoencoder to calculate the adaptive weights of the basic expression features;
[0015] S13: Use the compactness loss function Calculate the compactness loss of the basic expression features and reconstruct the intra-class and inter-class spatial distributions of the basic expression features; where: represents the feature vector corresponding to the category y i after average pooling of the basic expression features; ω i represents the adaptive weight of the basic expression features; represents the corresponding class center, y i ∈ {1, 2, ··· n}, represents the category corresponding to the basic expression features; ||·||2 represents L2 regularization; i represents the input sample number; represents the dot product of elements, and M represents the number of training samples in a mini-batch;
[0016] In S2, calculate the displacement attention feature map, channel attention feature map and spatial attention feature map of the basic expression features, and obtain the attention map of the final basic expression features, including the following steps:
[0017] S21: Calculate the displacement attention feature map of the basic expression features;
[0018] S22: For the input basic expression features, process the basic expression features with two 3×3 convolutions, then perform average pooling on the processed expression features, then use a 1×1 convolution to compress the channels of the pooled expression features, and then perform a 1×1-based convolution channel expansion again under the action of the ReLU activation function, and finally process the expression features after compression and ReLU activation function channel expansion with the sigmoid function to obtain the channel attention feature map;
[0019] S23: For the input basic expression features, use a large convolutional kernel of 7×7, then perform convolutional processing using a grouped convolutional kernel with 7×7 convolutional kernels, and finally multiply the expression features after grouped convolutional processing by the basic expression features activated by the sigmoid function to obtain a spatial attention feature map;
[0020] S24: Add the displacement attention feature map, channel attention feature map, and spatial attention feature map obtained in steps S21, S22, and S23 to obtain an attention map of the basic expression features;
[0021] In S21, the calculation of the displacement attention feature map of the basic expression features is as follows:
[0022] S211: Take the basic expression features of the input of a given size, move the basic expression features of the given size in four spatial directions in pixel units, and delete the unused pixels to obtain enhanced features;
[0023] S212: Encode the obtained enhanced features, then feed the encoded enhanced features into a multi-head self-attention learning unit, calculate the query Q, key K, and value V vectors respectively, calculate the attention weights of the enhanced features from the Q, K, and V vectors, and then calculate the attention embedding map from the attention weights and enhanced features, and input the attention embedding map into the subsequent steps;
[0024] S213: From the following formula and the attention embedding map, calculate the displacement attention feature map;
[0025] In the formula: LN represents layer normalization, represents doubling the original number of channels using a linear function, represents projecting the extended number of channels to the original number of channels using a linear function, δ represents the GELU activation function, represents element-wise multiplication, Z represents the attention embedding map, is the displacement attention feature map.
[0026] In S4, the calculation of the combined loss, and the output of the facial expression recognition result of the sample according to the combined loss and the attention map includes the following steps:
[0027] S41: From the formula calculate the softened classification loss;
[0028] In the formula: p i represents the predicted probability after using the softmax function, k represents the smoothing factor, set to 0.1, y iIndicates the category corresponding to the basic expression features, which is 1 when the label is correct and 0 when the label is incorrect. N represents the number of categories; Indicates the softened classification loss;
[0029] S42: Combine the calculated softened classification loss with the lack of compactness to obtain the final joint loss, complete the training and optimization of the network, and output the facial emotion recognition result according to the attention mapping and the final joint loss result.
[0030] Compared with the existing inventions, the present invention has the following advantages:
[0031] 1. The selected ResNet-18 as the feature extraction model balances the lightness and efficiency of feature extraction, and can better solve the problems of gradient disappearance and explosion;
[0032] 2. The displacement attention feature map of the basic expression features obtained in S21 effectively compensates for the limitations of local feature extraction in traditional convolutional neural networks; specifically, the calculation method used in S213 is different from that of traditional feedforward networks, which only process the self-attention mechanism through two layers of linear networks for channel conversion, and can more precisely and effectively control the flow of basic expression feature information;
[0033] 3. In S23, a large convolutional kernel of 7×7 is adopted, and then grouped convolutional kernels of 7×7 convolutional kernels are used for convolution processing, which can enhance the capture of basic expression features in the whole step and improve the robustness and accuracy of facial expression recognition. Description of the Drawings
[0034] Appendix Figure 1 : The specific process of a facial expression recognition method in an occluded environment Detailed Embodiments
[0035] The embodiments of the present application are described below, where the same or similar reference numerals represent the same or similar elements or elements with the same or similar functions throughout. The embodiments described below are exemplary and are intended to explain the present application, but should not be construed as limiting the present application.
[0036] A facial expression recognition method in an occluded environment, as shown in the appendix Figure 1 shown, includes the following steps:
[0037] S1: For a given facial emotion recognition sample, use the lightweight backbone network ResNet-18 to obtain the basic expression features of the sample, and at the same time construct the intra-class and inter-class spatial distributions of the basic expression features;
[0038] S2: Calculate the displacement attention feature map, channel attention feature map, and spatial attention feature map of the basic expression features, and obtain the attention map of the basic expression features;
[0039] S3: Attach emotion labels to the obtained attention map;
[0040] S4: Calculate the joint loss, and output the facial expression recognition result of the sample according to the joint loss and the attention map.
[0041] In S1, a lightweight backbone network ResNet-18 is used to obtain the basic expression features of the sample. At the same time, constructing the intra-class and inter-class spatial distributions of the basic expression features includes the following steps:
[0042] S11: From the input samples, when the i-th sample x is input, obtain the basic expression features through the backbone network S:
[0043]
[0044] where w represents the weight parameters of the network, x i is the i-th input sample, and x' i is the basic expression feature of the i-th input sample.
[0045] S12: Establish a multi-level autoencoder to calculate the adaptive weights of the basic expression features, specifically as follows: Construct a multi-level autoencoder here. This asymmetric encoder first flattens the basic expression features x' and converts them into a low-dimensional space with 128 dimensions, then remaps them again into a sub-high-dimensional feature with 1024 dimensions, and finally remaps them again to the output subspace with 512 dimensions, thereby generating the final spatial feature weights. This not only reduces redundant information but also further enhances the adaptive expression ability of the enhancement function.
[0046] The whole process is as follows:
[0047]
[0048] where h l represents the feature output of the l-th layer. When l = 1, the network input h0 is the flattened representation of x'. b l is the bias set to 0 here, δ represents the ReLU activation function that enhances the nonlinear ability of the network, and τ refers to the tanh activation function of the last layer. ψ is the softmax function, that is, the normalized exponential function, which is used to further map the output weight result between 0 and 1; ω i represents the adaptive weight of the basic expression features.
[0049] S13: Use the compactness loss function Calculate the compactness loss of the basic expression features and reconstruct the intra-class and inter-class spatial distributions of the basic expression features. Wherein: represents the feature vector corresponding to the category y after average pooling of the basic expression features i ; ω i represents the adaptive weight of the basic expression features; represents the corresponding class center, y i ∈{1, 2, ··· n}, represents the category corresponding to the basic expression features; ||·||2 represents L2 regularization; i represents the input sample number; represents the dot product of elements, and M represents the number of training samples in a mini-batch;
[0050] S2 specifically includes the following steps:
[0051] S21: Calculate the displacement attention feature map of the basic expression features, specifically as follows:
[0052] (1) S211: For the input basic expression feature x′ of size C×H×W, select some channels and move them in order along four spatial directions (such as left, right, up, and down) in pixel units. The deleted pixels are no longer used, and the empty pixels are filled with 0, and the remaining channels remain unchanged. Finally, obtain the enhanced feature E with the same shape after the above transformation. The specific process is as follows:
[0053] e[0:H,1:W] 0:aC ←x′[0:H,0:W-1] 0:aC
[0054] e[0:H,0:W-1] aC:2aC ←x′[0:H,1:W] aC:2aC
[0055] e[0:H-1,0:W] 2aC:3aC ←x′[1:H,0:W] 2aC:3aC
[0056] e[1:H,0:W] 3aC:4aC ←x′[0:H-1,0:W] 3aC:4aC
[0057] where a represents the scaling factor of the number of selected channels. Set the scaling factor to 1 / 12. Finally, combine the channels obtained by the above transformation with the remaining unchanged channels, and the output is E. It has the advantages of simple and efficient mechanism and no need to set parameters.
[0058] S212: Encode the obtained enhanced features, and then feed the encoded enhanced features into a multi-head self-attention learning unit to calculate query Q, key K, and value V vectors respectively. Calculate the attention weights of the enhanced features from the Q, K, and V vectors, and then calculate the attention embedding mapping from the attention weights and the enhanced features. Input the attention embedding mapping into the subsequent steps;
[0059] Based on the enhanced feature E obtained from the previous step, first encode the input according to the learnable positional embedding, and then feed the encoded enhanced features into a multi-head self-attention learning unit to calculate query Q, key K, and value V vectors respectively. In addition, based on obtaining the query Q, key K, and value V vectors, perform the following operations to obtain the attention weights, and input the attention weights Att(Q, K, V) into the subsequent steps, which are specifically described as follows:
[0060]
[0061] Where W represents the corresponding parameter matrix, LN represents layer normalization, D = C / H, which represents the dimension of each attention head, C is the original embedding dimension, and H is the number of attention heads.
[0062] From the enhanced feature E and the attention weights, obtain the attention embedding mapping Z, and the specific method is as follows:
[0063] Z = Att(Q, K, V) + E
[0064] S213: From the following formula and the attention embedding mapping, calculate the displacement attention feature map;
[0065] In the formula: LN represents layer normalization, represents doubling the original number of channels using a linear function, represents projecting the expanded number of channels to the original number of channels using a linear function, δ represents the GELU activation function, represents element-wise multiplication, Z represents the attention embedding mapping, is the displacement attention feature map.
[0066] S22: For the input basic expression features, first, use two 3×3 convolutions to process the basic expression features to further enhance the non-linear feature extraction ability and expand the network's field of view. On this basis, perform global average pooling on the processed expression features, then perform 1×1 convolution to compress the channels of the pooled expression features, and perform 1×1-based convolution channel expansion again under the action of the ReLU activation function. Finally, activate the reconstructed channel features with the sigmoid activation function and multiply them with the feature elements before pooling to construct the final channel attention feature map The entire process is shown in detail as follows:
[0067]
[0068] Among them, C is the original number of channels, which is set to 512 here, r and s are scaling factors, with values of 4 and 16 respectively. GAP represents global average pooling, and δ and σ are the ReLU and sigmoid activation functions respectively, indicating element-wise multiplication.
[0069] S23: Apply a large 7×7 convolutional kernel to the basic expression feature x′, with the grouped convolution size of 64 and reduce the channel dimension, then perform convolution processing using a grouped convolution kernel with 7×7 convolutional kernels, enhance the feature extraction through channel reconstruction after non-linear enhancement by the activation function, and finally multiply it with the original x′ activated by the sigmoid activation function to obtain the final spatial attention feature map Its formula is as follows:
[0070]
[0071] Among them, c represents the original number of channels, r is the scaling factor, which is set to 4 here. δ refers to the GELU activation function, and σ is the sigmoid function, indicating element-wise multiplication.
[0072] S24: Add the displacement attention feature map, channel attention feature map, and spatial attention feature map obtained in steps S21, S22, and S23 to obtain the attention map of the basic expression feature, specifically as follows:
[0073]
[0074] S4 specifically includes the following steps:
[0075] S41: Calculate the softened classification loss from the formula
[0076] In the formula: p i represents the predicted probability after using the softmax function, k represents the smoothing factor, which is set to 0.1, y i represents the category corresponding to the basic expression feature, which is 1 when the label is correct and 0 when the label is incorrect, N represents the number of categories, is the calculated softened classification loss.
[0077] S42: Combine the calculated softened classification loss with the lack of compactness to obtain the final combined loss Complete the training and optimization of the network, and output the facial emotion recognition result according to the attention map and the final combined loss result. The specific formula is as follows:
[0078]
[0079] where λ is the compactness loss The control ratio parameter in the combined loss.
[0080] The implementation of a facial expression recognition method in an occluded environment according to the present invention is conducive to completing facial expression recognition in the wild environment with occlusion with high accuracy, can more effectively perform facial expression recognition on a wider range of samples, improve the operation accuracy, and has important application value for broadening the application scenarios and fields of facial expression recognition and expanding the facial expression dataset, etc.
[0081] The above description of the present invention is provided to enable any ordinary person skilled in the art to implement or use the present invention. Various modifications to the present invention are obvious to those of ordinary skill in the art, and the general principles defined herein can also be applied to other variations without departing from the protection scope of the present invention. Therefore, the present invention is not limited to the examples and designs described herein, but is consistent with the broadest scope that conforms to the principles and novel features disclosed herein.
Claims
1. A facial expression recognition method in an occluded environment, characterized in that: The following steps are involved: S1: For a given facial emotion recognition sample, a lightweight backbone network ResNet-18 is used to obtain the basic expression features of the sample, and at the same time, the intra-class and inter-class spatial distribution of the basic expression features is constructed; S2: Calculate the displacement attention feature map, channel attention feature map and spatial attention feature map of the basic expression features, and obtain the attention map of the basic expression features; S3: attach sentiment labels to the obtained attention maps; S4: Calculate the joint loss and output the facial expression recognition result of the sample based on the joint loss and attention mapping.
2. The method for facial expression recognition in an occluded environment according to claim 1, characterized in that: In S1, a lightweight backbone network ResNet-18 is used to obtain the basic expression features of the sample. At the same time, the intra-class and inter-class spatial distribution of the basic expression features is constructed, including the following steps: S11 uses a lightweight backbone network ResNet-18 to obtain the basic expression features of the sample from the input sample; S12: Establish a multi-level autoencoder to calculate the adaptive weights of basic expression features; S13: Using compactness loss function Calculate the compactness loss of basic expression features and reconstruct the intra-class and inter-class spatial distribution of basic expression features; where: Indicates the corresponding category y after the average pooling of basic expression features i The characteristic vector of i Adaptive weights representing basic expression features; represents the corresponding class center, y i ∈{1, 2, ···n}, represents the category corresponding to the basic expression features; ||·||2 represents L2 regularization; i represents the input sample number; represents the dot product of the elements, and M represents the number of training samples in a small batch.
3. The facial expression recognition method in an occluded environment according to claim 1, characterized in that: In S2, calculating the displacement attention feature map, the channel attention feature map, and the spatial attention feature map of the basic expression features, and obtaining the final attention map of the basic expression features includes the following steps: S21: Calculate the displacement attention feature map of basic expression features; S22: For the input basic expression features, two 3×3 convolutions are used to process the basic expression features, and then the processed expression features are averaged and pooled. Then, 1×1 convolution is used to compress the channels of the pooled expression features, and the 1×1 convolution channel is expanded again under the action of the ReLU activation function. Finally, the sigmoid function is used to process the compressed and ReLU activation function channel-expanded expression features to obtain the channel attention feature map; S23: For the input basic expression features, a large 7×7 convolution kernel is used, and then a grouped convolution kernel of 7×7 convolution kernels is used for convolution processing. Finally, the expression features processed by the grouped convolution are multiplied with the basic expression features under the activation of the sigmoid function to obtain the spatial attention feature map; S24: Add the displacement attention feature map, channel attention feature map, and spatial attention feature map obtained in steps S21, S22, and S23 to obtain the attention map of the basic expression features.
4. The facial expression recognition method in an occluded environment according to claim 3, characterized in that: In S21, the calculation obtains the displacement attention feature map of the basic expression feature, which is as follows: S211: taking the input basic expression features of a given size, moving the basic expression features of the given size in pixel units along four spatial directions, and deleting unused pixels to obtain enhanced features; S212: Encode the obtained enhanced features, and then feed the encoded enhanced features into a multi-head self-attention learning unit, respectively calculate the query Q, key K and value V vectors, calculate the attention weight of the enhanced features from the Q, K, V vectors, and then calculate the attention embedding map from the attention weight of the enhanced features and the enhanced features, and input the attention embedding map into the subsequent steps; S213: By the following formula and the attention embedding map, and calculate the displacement attention feature map; Where: LN represents layer normalization, Indicates that the original number of channels is doubled using a linear function. Indicates that the number of expanded channels is projected to the original number of channels using a linear function, δ represents the GELU activation function, represents element-wise multiplication, Z represents the attention embedding map, is the shifted attention feature map.
5. A facial expression recognition method in an obstructed environment according to claim 1, characterized in that: Calculating the joint loss as described in S4 and outputting the facial expression recognition result of the sample according to the joint loss and the attention map includes the following steps: S41: By formula Calculate the softened classification loss; Where: p i represents the predicted probability after using the softmax function, k represents the smoothing factor, which is set to 0.1, and y i Indicates the category corresponding to the basic expression feature. It is 1 when the label is correct and 0 when the label is incorrect. N indicates the number of categories. represents the softened classification loss; S42: Combine the calculated softened classification loss with the compactness loss to obtain the final joint loss, complete the network training and optimization, and output the facial emotion recognition result based on the attention mapping and the final joint loss result.