Shielding expression recognition method based on adaptive region attention enhancement mechanism
By using the method of using adaptive region-enhanced attention mechanism and cosine clustering loss function in facial expression recognition, the problem of low accuracy of expression recognition under occlusion conditions is solved, and higher accuracy of expression recognition and discrimination ability are achieved.
Patent Information
- Application Number
- CN202411758343.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2024-05-13
- Filing Date
- 2024-12-03
- Publication Date
- 2025-05-13
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
In the case where the existing facial expression recognition methods are blocked, the accuracy of facial expression recognition has significantly decreased, making it difficult to effectively locate and identify key facial features.
Using a method based on the adaptive region enhancement attention mechanism, the expression feature map is extracted through the feature extraction network, combined with the space-channel attention mechanism and the cosine cluster loss function, the adaptive region enhancement module calculates the weights of each sub-region of the face, thereby achieving accurate expression recognition.
It improves the accuracy of expression recognition under occlusion conditions, can better locate and identify key facial features, and enhances the model's ability to judge expressions.
Smart Images

Figure SMS_1 
Figure SMS_2 
Figure SMS_3
Abstract
Description
Technical Field
[0001] The invention relates to an occluded expression recognition method based on an adaptive region enhanced attention mechanism, and belongs to the technical field of expression recognition. Background Art
[0002] The present invention relates to facial expression recognition (FER) technology in the field of computer vision. At present, FER technology is widely used in the fields of human-computer interaction, sentiment analysis, etc. However, the accuracy of expression recognition of the existing FER method is significantly reduced when the face is occluded, because the facial expression features are reduced, which increases the similarity between various expressions. In addition, the key to identifying an expression lies in the positioning and identification of key facial features. However, under the condition of occlusion, accurately positioning and identifying these key features becomes a huge challenge. The above problems limit the further development and application of FER technology. Therefore, the problem faced by the present invention is how to improve the accuracy of facial expression recognition when the face is occluded. By applying the attention mechanism technology and loss function in deep learning to the FER technology, the autonomous learning ability of the attention mechanism can be used to perceive the key features of the face, and the improved loss function can be used to make the same type of expression features in the feature space more compact, thereby improving the accuracy of expression recognition under occlusion conditions.
[0003] With the rapid development of deep learning and the release of some large-scale data sets, deep neural networks represented by convolutional neural networks (CNN) have become the mainstream of image recognition technology. At present, image recognition technology based on deep learning has high performance in terms of accuracy and robustness. Applying this technology to the FER task can achieve accurate recognition of expression images under occlusion conditions. Summary of the invention
[0004] In view of the above prior art, the present invention provides a method for occluded expression recognition based on an adaptive region enhanced attention mechanism, which comprises the following steps:
[0005] Step 1: Extract expression feature graph using feature extraction network;
[0006] Step 2: Use the spatial-channel attention mechanism to locate and identify key facial features;
[0007] Step 3: Use the cosine clustering loss function to calculate the distance loss between the expression feature map and the class center;
[0008] Step 4: Use the adaptive region enhancement module to calculate the weights of each facial sub-region;
[0009] Step 5: Use the classification layer to classify expressions and obtain classification results.
[0010] The step 1 specifically includes the following sub-steps:
[0011] Step 1.1: Import necessary libraries and modules in the Python environment, including deep learning framework PyTorch, image processing library, etc.
[0012] Step 1.2: Download or load the pre-trained ResNet-18 model. The ResNet-18 used in this invention is pre-trained on the MS-Celeb-1M face recognition dataset;
[0013] Step 1.3: Modify the model structure, remove the original fully connected layer of ResNet-18, and keep only the convolutional layer. Then input the prepared expression image into the ResNet-18 model and perform forward propagation to obtain the expression feature map.
[0014] The step 2 specifically includes the following sub-steps:
[0015] Step 2.1: Use the spatial attention unit in the spatial-channel attention mechanism module to obtain spatial features. First, compress the input feature map to reduce its dimension. Then, perform 3×3, 5×5, and 7×7 convolution operations on the output feature map. Finally, send the output feature map to the ReLU activation function to obtain spatial features;
[0016] Step 2.2: Use the channel attention unit in the spatial-channel attention mechanism module to obtain spatial-channel features. First, perform average pooling and maximum pooling operations on the spatial feature map. Then, pass the spatial feature map through the fully connected layer and batch normalization layer to obtain the channel weight. At the same time, multiply the channel weight with the spatial feature map of the unit to obtain the final spatial-channel feature.
[0017] The step three specifically includes the following sub-steps:
[0018] Step 3.1: Process the input feature vector and convert the three-dimensional feature vector into a one-dimensional feature vector;
[0019] Step 3.2: After processing the feature vector, calculate the Euclidean distance between the sample feature and the category center. It can be expressed as:
[0020] Among them, x represents the feature vector of the sample, c represents the category center vector, and L e represents the Euclidean distance, ||·||2 represents the L2 norm, n is the dimension of the feature vector, x i and c iare the eigenvalues of the eigenvector and the category center vector at the i-th eigenvalue;
[0021] Step 3.3: Calculate the cosine distance between the sample feature and the category center. It can be expressed as:
[0022] Among them, L c represents the cosine distance, and cos(xc) represents the cosine similarity between vectors x and c;
[0023] Step 3.4: weighted fusion of Euclidean distance and cosine distance. It can be expressed as: L ec =αL e +λL c
[0024] Among them, α and λ are the weights of Euclidean distance and cosine distance respectively, and α+λ=1, L ec is the final comprehensive distance loss.
[0025] The step 4 specifically includes the following sub-steps:
[0026] Step 4.1: Process the input feature vector and convert the three-dimensional feature tensor into a one-dimensional feature vector. And use the scaling operation to convert the high-dimensional vector into a low-dimensional vector to reduce the computational overhead of the model;
[0027] Step 4.2: Use Mahalanobis distance to calculate the similarity of each region. The formula is:
[0028] in, and represents two eigenvectors, S represents the covariance matrix, and after calculating the similarity between each two regions, a correlation coefficient matrix is obtained. Then, the elements of each row of the matrix are added together to obtain the weight of each region, and then all weights are normalized to ensure that the sum of all weights is 1.
[0029] The step five: using the feature map and fully connected layer obtained in step 4.2 to obtain the final expression classification result.
[0030] The advantages and benefits of this method are as follows: This method uses the adaptive region enhancement attention mechanism and the cosine clustering loss function to achieve accurate recognition of facial expressions under occlusion conditions. This method first uses the spatial-channel attention mechanism module to simultaneously focus on multiple facial sub-regions, which can better locate and identify key facial features compared to other attention mechanism models. Therefore, the local features of the face can be fully and effectively utilized to improve the model performance. Then, the use of the adaptive region enhancement module enables the FER model to automatically focus on the non-occluded areas of the face and ignore the occluded areas of the face. This allows the model to better learn facial expression features, thereby improving the accuracy of occluded expression recognition. Finally, the cosine clustering loss function is used to comprehensively consider the Euclidean distance and the cosine distance, and the weight hyperparameter is introduced when calculating the distance between the sample and the class center to balance the contribution of the two distances, so as to improve the model's ability to distinguish between various categories of expression samples, thereby improving the recognition accuracy of the model. BRIEF DESCRIPTION OF THE DRAWINGS Figure 1 It is the overall flow chart; Figure 2 It is a flow chart of the spatial-channel attention mechanism; Figure 3 It is the two-dimensional quadrant graph of the ReLU function and its derivative graph; Figure 4 It is a comparison of feature clustering results between the method of the present invention and the baseline method. DETAILED DESCRIPTION
[0031] The following is based on the attached Figure 1-4 The present invention is further described: The technical solution of the present invention will be described clearly and completely below in conjunction with the accompanying drawings. The described embodiments are part of the embodiments of the present invention, but not all of them. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.
[0032] The present invention proposes a method for occluded expression recognition based on adaptive region enhanced attention mechanism, the specific steps are as follows: Figure 1 shown, including:
[0033] Step 1: Extract expression feature graph using feature extraction network;
[0034] Step 2: Use the spatial-channel attention mechanism to locate and identify key facial features;
[0035] Step 3: Use the cosine clustering loss function to calculate the distance loss between the expression feature map and the class center;
[0036] Step 4: Use the adaptive region enhancement module to calculate the weights of each facial sub-region;
[0037] Step 5: Use the classification layer to classify the expressions and obtain the classification results;
[0038] The core of the present invention includes three parts: the first part is to locate and identify key facial features, the second part is to calculate the distance loss between the expression feature map and the class center, and the third part is to calculate the weights of each facial sub-region.
[0039] In step 1, the feature extraction network uses ResNet-18 to extract the expression feature map. The process includes importing the deep learning framework library and loading the pre-trained ResNet-18 model, preparing the expression data set and preprocessing it, performing face detection and face alignment on the image, and resizing the image to 224×224 to obtain the expression image. This can reduce the complexity of the data and increase the training speed of the model. The network structure is modified to remove the last full connection, and the output of the intermediate convolutional layer is obtained as the expression feature map through forward propagation.
[0040] The second step is to use the spatial-channel attention mechanism module to locate and identify key facial features. This module first divides the expression feature map into several feature sub-maps, and then applies the spatial-channel attention mechanism to each feature sub-map. Therefore, it can focus on multiple areas of the face at the same time, thereby capturing facial expression features more accurately and comprehensively. The specific process of this module is as follows Figure 2 shown.
[0041] The spatial-channel attention mechanism module includes a spatial attention unit and a channel attention unit. In step 2.1, the spatial attention unit first compresses and reduces the dimension of the input feature map through a 1×1 convolution operation. The spatial size of the feature map after dimensionality reduction remains unchanged, and the number of channels becomes 1 / 2 of the input feature map. This process is conducive to the information exchange between the features of each channel and reduces the amount of model calculation. Then, the output feature map is subjected to 3×3, 5×5 and 7×7 convolution operations. The spatial size of the feature map output by the three convolution operations remains unchanged, and the number of channels becomes 1. The feature maps output by the three convolution operations are then added together. Since information of different scales is integrated, it is easier for the model to understand the overall structure and contextual information in the image. Finally, the output feature map is sent to the ReLU activation function. This process can introduce more nonlinear transformations to overcome the linear limitations in the neural network. At the same time, the output feature map is multiplied by the original input feature map to obtain the spatial feature. In step 2.2, the channel attention unit uses the spatial feature obtained in step 2.1 as its original input feature. First, the input features are subjected to average pooling and maximum pooling operations at the same time. This process changes the spatial size of the input feature map to 1×1, and the number of channels remains unchanged. The average and maximum pooling operations enable the model to learn some texture and abstract information of the image. Then, the feature map is passed through a linear layer to obtain the channel weight. At the same time, the channel weight is multiplied with the original input feature map of the unit to obtain the final output feature.
[0042] Step 3 is to use the cosine clustering loss function to calculate the distance loss between the expression feature map obtained in step 1.3 and the class center. The cosine clustering loss function achieves the purpose of expanding the distance between various samples in the feature space by expanding the distance between class centers and combining Euclidean distance and cosine distance. This loss function mainly optimizes the distance between samples of different categories, mainly based on Euclidean distance and cosine distance. Specifically, in step 3.1, the input feature vector (where H, W and C represent the height, width and number of channels of the three-dimensional feature vector respectively) is converted into a one-dimensional feature vector after global flat pooling operation. In step 3.2, after the feature vector is processed, the Euclidean distance between the sample feature and the category center is calculated. Assuming that the sample feature vector is x and the category center vector is c, the Euclidean distance between them is expressed as:
[0043] Among them, L e represents the Euclidean distance, ||·||2 represents the L2 norm, n is the dimension of the feature vector, x i and c iare the feature vector and the category center vector at the i-th eigenvalue respectively. Then in step 3.3, the cosine distance between the sample feature and the category center is calculated. Assuming that the sample feature vector is x and the category center vector is c, the cosine distance between them is expressed as:
[0044] Among them, L c Represents the cosine distance, and cos(xc) represents the cosine similarity of vectors x and c. Compared with the Euclidean distance, the cosine distance can pay more attention to the directionality of the vector rather than the size. Therefore, it is insensitive to the specific values of the data features in this article and is more suitable for measuring the relative relationship between expression features. And it often provides better performance when processing sparse data or high-dimensional data. Then in step 3.4, the Euclidean distance and the cosine distance are weightedly fused. The final comprehensive distance L is obtained. ec , expressed as: L ec =αL e +λL c
[0045] Among them, α and λ are the weights of Euclidean distance and cosine distance respectively, and α+λ=1, L ec is the final comprehensive distance loss.
[0046] Step four is to calculate the weights of each facial sub-region using the adaptive region enhancement module. The spatial-channel attention mechanism module in step two can accurately locate and identify the key facial regions, but in the process, the features of the occluded parts may be encountered, which will cause certain interference to the model. To address this problem, the present invention proposes an adaptive region enhancement module, the core function of which is to quantify the similarity between regions, thereby obtaining a correlation coefficient matrix, and then assigning corresponding weights to each region. Through this adaptive weight allocation method, the model pays more attention to unoccluded or slightly occluded regions, thereby improving the accuracy and robustness of occluded expression recognition. The adaptive region enhancement module uses the Mahalanobis distance to measure the degree of similarity between each two regions. The Mahalanobis distance is a statistical indicator used to measure the distance between sample points. It takes into account the correlation between each feature and adjusts the covariance matrix when calculating the distance. Compared with the Euclidean distance, the Mahalanobis distance is more accurate when considering the correlation between features. Specifically, for two feature vectors and The covariance matrix is S, and the calculation formula of its Mahalanobis distance is:
[0047] According to the above formula, step 4.1 converts the three-dimensional feature tensor into a one-dimensional vector and uses a scaling operation to convert the high-dimensional vector into a low-dimensional vector to reduce the computational overhead of the model. The process is described as follows: Assuming that the dimension of the tensor T is H×W×C, it is expressed as Among them, H represents the length, W represents the width, and C represents the number of channels. After the flattening operation, the tensor becomes a one-dimensional vector with a dimension of 1×1×C, which is expressed as Then the dimension is reduced by scaling operation. Assuming the scaling factor is α, the dimension of the scaled vector is 1×1×C / α, which is expressed as The flattening operation arranges all elements in the three-dimensional tensor into a one-dimensional vector in a certain order, so that all the information in the tensor is contained in the one-dimensional vector. In step 4.2, after calculating the similarity between each two regions, a correlation coefficient matrix is obtained, and then the elements of each row of the matrix are added to obtain the weight of each region, and then all weights are normalized to ensure that the sum of all weights is 1.
[0048] Step 5 is to first perform a Flatten operation on the feature map to convert it into a one-dimensional vector so that it can be input into the fully connected layer for classification. Then, a ReLU activation function is added after the fully connected layer. The ReLU function can be expressed by the following formula:
[0049] like Figure 3 As shown in (a), this function is very simple. If the input y is greater than zero, the output is y. If the input is less than or equal to zero, the output is zero. This simple feature allows the neural network to converge faster when using the ReLU function, and the gradient disappearance problem will not occur during the gradient descent process. Figure 3 As shown in (b), an important feature of the ReLU function is that its derivative is always 1 in the positive part and always 0 in the negative part. This means that in the positive part, the gradient remains unchanged, but in the negative part, the gradient disappears, and this property leads to the "death of neurons". Although this may cause some neurons to stop updating during training, in fact, the sparsity exhibited by the ReLU function helps the generalization of the network and reduces overfitting.
[0051] The fully connected layer includes multiple neurons, each corresponding to an expression category. By updating the parameters of the fully connected layer during the learning process, the model can map the extracted features to the final expression category.
[0052] Figure 4 (a) is the clustering result of the baseline method for various expression features. Figure 4 (b) is the clustering result of various expression features by the method of the present invention.
[0053] In summary, the occluded expression recognition method provided by the present invention aims at the problem of decreased expression recognition accuracy under occlusion conditions, constructs an effective deep learning model, and improves the loss function. First, a feature extraction network is used to extract an expression feature map. Then, a spatial-channel attention mechanism module is used to locate and identify multiple key areas of the face. At the same time, a cosine clustering loss function is used to measure the distance between the expression feature vector and the center vector of the category to which it belongs, so that the expression features of the same category in the feature space are more compact. Finally, an adaptive region enhancement module is used to identify facial occlusion or unimportant areas, and ignore the area. This method can accurately identify facial expression features under occlusion conditions, and can improve the model's ability to discriminate expressions, thereby improving the model's expression recognition accuracy.
Claims
1. A method for occluded expression recognition based on adaptive region enhanced attention mechanism, characterized in that: The following steps are involved: Step 1: Extract expression feature graph using feature extraction network; Step 2: Use the spatial-channel attention mechanism to locate and identify key facial features; Step 3: Use the spatial-channel attention mechanism to locate and identify key facial features; Step 4: Use the adaptive region enhancement module to calculate the weights of each facial sub-region; Step 5: Use the classification layer to classify expressions and obtain classification results.
2. According to the method of claim 1, the method is characterized in that: In step one, The method for extracting the expression feature map is to extract the expression feature map using the feature extraction network ResNet-18. First, import the deep learning framework library and load the pre-trained ResNet-18 model. Then, prepare the expression data set and pre-process it, perform face detection and face alignment on the image, and resize the image to 224×224. Finally, modify the network structure to remove the last full connection, and obtain the output of the intermediate convolution layer as the expression feature map through forward propagation.
3. According to the method for occluded expression recognition based on adaptive region enhanced attention mechanism described in claim 1, it is characterized in that: In step 2, The spatial-channel attention mechanism module includes a spatial attention unit and a channel attention unit. The spatial attention unit first compresses and reduces the dimension of the input feature map through a 1×1 convolution operation. The spatial size of the feature map after dimensionality reduction remains unchanged, and the number of channels becomes 1 / 2 of the input feature map. Then, the output feature map is subjected to 3×3, 5×5 and 7×7 convolution operations. The spatial size of the feature map output by the three convolution operations remains unchanged, and the number of channels becomes 1. The feature maps output by the three convolution operations are then added. Finally, the output feature map is sent to the ReLU activation function, and the output feature map is multiplied by the original input feature map to obtain the spatial feature. The channel attention unit uses the obtained spatial feature as its original input feature. First, the input feature is subjected to average pooling and maximum pooling operations at the same time. This process changes the spatial size of the input feature map to 1×1, and the number of channels remains unchanged. Then, the feature map is passed through a linear layer to obtain the channel weight, and the channel weight is multiplied by the original input feature map of the unit to obtain the final output feature.
4. According to the method for occluded expression recognition based on adaptive region enhanced attention mechanism described in claim 1, it is characterized in that: In step three, The cosine clustering loss function is the input feature vector (where H, W and C represent the height, width and number of channels of the three-dimensional feature vector respectively) is transformed into a one-dimensional feature vector after global flat pooling operation. After the feature vector is processed, the Euclidean distance between the sample feature and the category center is calculated. Assuming that the sample feature vector is x and the category center vector is c, the Euclidean distance between them is expressed as: Among them, L e represents the Euclidean distance, ||·||2 represents the L2 norm, n is the dimension of the feature vector, x i and c i are the feature vector and the category center vector at the i-th eigenvalue respectively. Next, the cosine distance between the sample feature and the category center is calculated. Assuming that the sample feature vector is x and the category center vector is c, the cosine distance between them is expressed as: Among them, L c represents the cosine distance, and cos(xc) represents the cosine similarity between vectors x and c. Then, the Euclidean distance and the cosine distance are weighted and fused to obtain the final comprehensive distance L. ec , expressed as: L ec =αL e +λL c Among them, α and λ are the weights of Euclidean distance and cosine distance respectively, and α+λ=1, L ec is the final comprehensive distance loss.
5. The method for occluded expression recognition based on adaptive region enhanced attention mechanism according to claim 1, characterized in that: In step four, The adaptive region enhancement module calculates the weight of each facial sub-region. The adaptive region enhancement module uses Mahalanobis distance to measure the similarity between each two regions. Specifically, for two feature vectors and The covariance matrix is S, and the calculation formula of its Mahalanobis distance is: According to the above formula, first, the three-dimensional feature tensor is converted into a one-dimensional vector, and a scaling operation is used to convert the high-dimensional vector into a low-dimensional vector. The process is described as follows: Assuming that the dimension of the tensor T is H×W×C, it is expressed as Among them, H represents the length, W represents the width, and C represents the number of channels. After the flattening operation, the tensor becomes a one-dimensional vector with a dimension of 1×1×C, which is expressed as Then, the dimension is reduced by scaling operation. Assuming the scaling factor is α, the dimension of the scaled vector is 1×1×C / α, which is expressed as The flattening operation arranges all the elements in the three-dimensional tensor into a one-dimensional vector in a certain order. Finally, after calculating the similarity between each two regions, a correlation coefficient matrix is obtained. Then, the elements of each row of the matrix are added together to obtain the weight of each region, and then all weights are normalized to ensure that the sum of all weights is 1.
6. The method for occluded expression recognition based on adaptive region enhanced attention mechanism according to claim 1, characterized in that: In step five, The process of obtaining expression classification is to first perform a Flatten operation on the feature map to convert it into a one-dimensional vector so that it can be input into the fully connected layer for classification. Then, a ReLU activation function is added after the fully connected layer. The ReLU function can be expressed by the following formula: Finally, by updating the parameters of the fully connected layer during the learning process, the model is able to map the extracted features to the final expression categories.