A Facial Expression Recognition Method Based on Enhanced Self-Attention Transformer
By combining the IR50 convolutional neural network and enhanced self-attention Transformer, multi-stage feature extraction and screening are carried out, which solves the problem of background noise interference and high calculation volume in expression recognition, and achieves efficient expression recognition effect.
Patent Information
- Application Number
- CN202310805217.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-06-30
- Publication Date
- 2025-07-25
- Estimated Expiration
- 2043-06-30
AI Technical Summary
The existing expression recognition method based on Transformer has background noise interference in an open environment, and the calculation volume is large, and the recognition accuracy and speed need to be improved.
The combination of IR50 convolutional neural network and enhanced self-attention Transformer is adopted to reduce the number of tokens, improve the recognition accuracy and speed up the reasoning through multi-stage feature extraction and feature fusion and screening.
While ensuring the recognition accuracy, the inference speed is significantly improved and the calculation complexity is reduced by about 17.6%.
Smart Images

Figure CN116884063B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of computer vision, and specifically relates to a face expression recognition method based on an enhanced self-attention Transformer. Background Art
[0002] Facial expressions are important carriers for expressing emotions and key factors for enhancing the efficiency of communication. With the development of virtual reality technology, digital platforms such as virtual anchors and virtual online multi-person social platforms have emerged. Precise and fast facial expression recognition technology plays a key role in these digital fields, and can effectively enhance the interaction experience and communication efficiency of products.
[0003] However, in an open environment, facial expression images contain a lot of background noise, that is, a lot of useless information. The current expression recognition methods mainly input the face-aligned images into the network, and there will still be a lot of useless information; in addition, in practical applications, there are high requirements for the accuracy and speed of expression recognition. The existing Transformer-based expression recognition methods have obtained high accuracy after using pre-trained models, but the computational cost is large, and the recognition accuracy needs to be further improved. Summary of the Invention
[0004] In order to solve the above technical problems, the present invention provides a face expression recognition method based on an enhanced self-attention Transformer, which extracts multi-stage features through a convolutional neural network to obtain multi-scale features, and performs feature fusion and screening through an enhanced Transformer. During this process, the number of tokens is reduced. This method can not only obtain high recognition accuracy, but also significantly improve the inference speed.
[0005] In order to achieve the above object, the present invention is realized through the following technical solutions:
[0006] The present invention is a face expression recognition method based on an enhanced self-attention Transformer, and the method includes the following steps:
[0007] Step S1, obtain a face expression training data set and perform preprocessing operations on the data set;
[0008] Step S2, establish a face expression recognition network model composed of an IR50 convolutional neural network and an enhanced self-attention Transformer;
[0009] Step S3, perform preliminary feature extraction by the IR50 convolutional neural network and splice the features in the intermediate stage;
[0010] Step S4: Input the spliced features into the enhanced self-attention Transformer, and perform similar feature fusion and key feature screening in sequence.
[0011] Step S5: Input the result after the Transformer finishes execution into the fully connected layer, and finally use the Softmax classifier to classify and obtain the facial expression classification result.
[0012] In step S1, the preprocessing operations on the training dataset include the following steps:
[0013] S1.1: Adjust the image size in the training dataset to 112×112, and perform data augmentation operations on it. The data augmentation operation methods include: rotating the images in the facial expression training dataset by 6° and horizontal flipping, and the probability of performing the data augmentation is 0.5.
[0014] S1.2: Perform normalization processing on the augmented images.
[0015] S1.3: Perform masking processing on the images in the facial expression training dataset. The probability of performing the masking operation is 0.7, and the masking method is to fill with random values.
[0016] In step S2, establish a facial expression recognition network model composed of the IR50 convolutional neural network and the enhanced self-attention Transformer. The facial expression recognition network model consists of two main parts: the convolutional neural network IR50 and the enhanced Transformer. The input of the Transformer is the multi-stage output of the IR50.
[0017] In step S3, the IR50 convolutional neural network consists of four stages:
[0018] S3.1): The first stage is shallow feature extraction, which includes convolution, normalization, and activation functions.
[0019] S3.2): The second stage, the third stage, and the fourth stage are multiple residual blocks, and the corresponding feature dimensions are 64×56×56, 128×28×28, and 256×14×14 respectively.
[0020] S3.3): Splice the features of the second stage, the third stage, and the fourth stage. Since the feature dimensions are inconsistent, the second stage and the third stage need to go through downsampling and then be spliced with the fourth stage in the channel dimension. Further, expand the channels from 448 to 768 through a 1×1 convolution, and then use the Flatten dimensionality reduction operation and transpose to adjust its feature dimension to 196×768.
[0021] In step S4, the Transformer model consists of two parts, each with 4 blocks. The first part is progressive feature fusion, which fuses tokens with high content similarity; the second part is progressive feature screening, which gradually discards tokens with low relevance to the class token. Both parts can reduce the number of tokens and speed up the inference speed.
[0022] In the first part, a feature fusion module (Token Merging) is introduced. Using K in Q (query), K (key), and V (value) in the self-attention calculation formula as a measure of similarity, the calculation formula for the K value is as follows:
[0023] K = F × W K
[0024] where F is the input feature and W K is a learnable parameter matrix. K contains the information of each token. By calculating the dot product of K, the similarity between tokens can be quantified, and then similar tokens can be fused. The operation steps are as follows:
[0025] S4.1), all tokens of the input module are evenly divided into two sets A and B, where odd indices are in A and even indices are in B;
[0026] S4.2), for each token in set A, search for the most similar token in B, and use the corresponding K value for dot product calculation to obtain the similarity;
[0027] S4.3), retain the r combinations with the highest similarity, and average and fuse the 2r token features in the r groups to obtain r fused tokens. Here, r is a hyperparameter.
[0028] Furthermore, in the second part of feature selection, a feature selection module (Token Pruning) is introduced. In the self-attention mechanism, the self-attention layer realizes the dynamic aggregation of information through the interaction between Q (query), K (key), and V (value). The self-attention calculation formula for the class token is:
[0029]
[0030] where q c is the Q (query) query vector corresponding to the class token, K corresponds to the key vector of each token, V is the value vector corresponding to each token, d is the dimension of K, and the output of the above formula is the value vector V = [v1, v2,..., v n T a linear combination, and score is the interactive calculation of the class token with respect to all the remaining tokens. v i from the i-th token, the value score i (i.e., the i-th in score) determines how much of the information of the i-th token is fused into the output of the class token through the linear combination. Therefore, the value score i represents the importance of the i-th token. Denote it as the importance score of the i-th token, and the corresponding mathematical formula is:
[0031]
[0032] Furthermore, in descending order of score, retain the top N tokens and discard the remaining tokens. The discarded information is interference information such as background information.
[0033] In step S5, after the Transformer execution is completed, extract the class token and input it into the fully connected layer, and finally use the Softmax classifier to classify and obtain the facial expression classification result.
[0034] The beneficial effects of the present invention are:
[0035] The present invention extracts multi-scale and richer information through multi-stage feature splicing of the convolutional neural network.
[0036] The present invention designs an enhanced Transformer structure. First, it fuses the redundant information and similar features extracted by the convolutional neural network, and then through feature screening, removes the image background features and useless information. This method significantly improves the inference speed while ensuring the recognition accuracy. BRIEF DESCRIPTION OF THE DRAWINGS
[0037] Figure 1 is a schematic diagram of the facial expression recognition network model of the present invention.
[0038] Figure 2 is a schematic diagram of the convolutional neural network module in the network of the present invention.
[0039] Figure 3 is a schematic diagram of the enhanced self-attention Transformer structure of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0040] The following will disclose the embodiments of the present invention with diagrams. For the sake of clarity, many practical details will be described together in the following narrative. However, it should be understood that these practical details are not used to limit the present invention. That is to say, in some embodiments of the present invention, these practical details are not necessary.
[0041] The present invention is a facial expression recognition method based on an enhanced self-attention Transformer, and the method includes the following steps:
[0042] Step 1: Obtain a facial expression training data set and perform preprocessing operations on the facial expression training data set.
[0043] Performing preprocessing operations on the training data set includes the following steps:
[0044] Step 1.1: Adjust the image size in the facial expression training data set to 112×112 and perform data augmentation operations on the facial expression training data set;
[0045] Step 1.2: Perform normalization processing on the augmented images;
[0046] Step 1.3: Mask the images in the facial expression training data set, and the probability of performing the masking operation is 0.7. The masking method is to fill random values.
[0047] Step 2: Establish a facial expression recognition network model composed of a combination of an IR50 convolutional neural network and an enhanced self-attention Transformer.
[0048] As Figure 1 shown, the facial expression recognition network model consists of two main parts: a convolutional neural network IR50 and an enhanced Transformer. Multiscale features of the second, third, and fourth stages of the convolutional neural network are obtained, and the features of the second and third stages are downsampled and then concatenated with the features of the third stage. Further, the number of channels is expanded from 448 to 768 through a 1×1 convolution, and then the dimensionality of the features with 768 channels is adjusted to 196×76 using Flatten and transpose. The features extracted by the convolutional neural network are transformed and then input into the Transformer for feature fusion and screening operations.
[0049] Step 3: Perform preliminary feature extraction by the IR50 convolutional neural network and concatenate the features in the intermediate stage.
[0050] As Figure 2 shown, the IR50 convolutional neural network consists of convolution, PReLu, and BatchNorm. In the first stage, a 3×3 convolution is first used for shallow feature extraction. The second, third, and fourth stages are residual blocks, and the multistage feature dimensions are 64×56×56, 128×28×28, and 256×14×14 respectively.
[0051] Downsample the feature maps of the second and third stages to a size of 14×14 using average pooling, and then perform a concatenation operation in the channel dimension to obtain a feature of 448×14×14.
[0052] Use a 1×1 convolution to expand the channel dimension to 768 dimensions, and then perform Flatten and transpose operations to obtain a feature of 768×196 dimensions.
[0053] Step 4: Input the concatenated features into the enhanced self-attention Transformer, and perform similar feature fusion and key feature screening in sequence.
[0054] As Figure 3 shown, in Step 4, the enhanced self-attention Transformer includes two parts, each part having 4 blocks. The first part is progressive feature fusion, which fuses tokens with high content similarity; the second part is progressive feature screening, which gradually discards tokens with low relevance to the class token. Both parts can reduce the number of tokens and speed up the inference speed.
[0055] Introduce a feature fusion module (Token Merging) in the first part, and use K in Q (query), K (key), and V (value) in the self-attention calculation formula as a measure of similarity. The calculation formula for the K value is as follows:
[0056] K = F×W K
[0057] where F is the input feature, and W K is a learnable parameter matrix. K contains the information of each token. The similarity between tokens can be quantified by calculating the dot product of K, and then the similar tokens are fused. The operation steps are as follows:
[0058] S4.1), Divide all tokens input to the module into two sets A and B, where odd indices are in A and even indices are in B;
[0059] S4.2), For each token in set A, search for the most similar token in B, and calculate the dot product using the corresponding K value to obtain the similarity;
[0060] S4.3), Retain the r combinations with the highest similarity, and average-fuse the 2r token features in the r groups to obtain r fused tokens, where r is a hyperparameter.
[0061] In this example, the input feature dimension is 197×768, the r parameter is set to 15. After passing through 4 Token Merging modules, the feature dimension becomes 137×768, reducing 60 tokens. The Class token does not participate in the fusion calculation.
[0062] In the second part, the feature selection part, a feature selection module (Token Pruning) is introduced. In the self-attention mechanism, the self-attention layer realizes the dynamic aggregation of information through the interaction between Q (query), K (key), and V (value). The self-attention calculation formula for the classtoken is as follows:
[0063]
[0064] where q c is the Q (query) query vector corresponding to the class token, K is the key-value vector K (key) corresponding to each token, V is the value vector V (value) corresponding to each token, d is the dimension of K. The output of the above formula is a linear combination of the value vector V = [v1, v2,..., v n T and score is the interaction calculation of the class token with respect to the remaining tokens except the class token. v i comes from the i-th token, and the value score i i.e., the i-th in score determines how much information of the i-th token is fused into the output of the class token through the linear combination. Therefore, the value score i represents the importance of the i-th token. It is expressed as the importance score of the i-th token, and the corresponding mathematical formula is:
[0065]
[0066] Sorting by score from high to low, the top N tokens are retained, and the remaining tokens are discarded. The discarded information is interference information such as background information.
[0067] In this example, the input feature dimension is 137×768, the retention rate is set to 0.9. After passing through 4 modules, the feature dimension becomes 91×768, reducing 46 tokens. The Class token does not participate in the screening calculation.
[0068] Step 5: Extract the class token from the result after the Transformer is executed and input it into the fully connected layer. Finally, use the Softmax classifier to classify and obtain the expression classification result.
[0069] The experiment of this embodiment was carried out on the dataset RAF-DB. This example was compared with IR50+VIT_Small without feature fusion and screening. In terms of accuracy, the average of the experimental results of each model was taken 4 times. It can be seen that after feature fusion and screening, the accuracy of the model of the present invention has been improved by 0.001. The experimental results are shown in Table 1 below.
[0070] Table 1
[0071]
[0072] In Table 1, FLOPs (Floating Point Operations): the number of floating-point operations, which is used to measure the computational complexity of the model and is often used as an indirect measure of the speed of neural network models. It can be seen that the computational complexity of the model of the present invention has been reduced by about 17.6% compared with the IR50+VIT_Small model while ensuring the accuracy.
[0073] The above is only the preferred embodiment of the present invention, and the protection scope of the present invention is not limited to the above embodiment. Any equivalent modification or change made by those of ordinary skill in the art according to the disclosure of the present invention shall be included in the protection scope recorded in the claims.
Claims
1. A face expression recognition method based on an enhanced self-attention Transformer, characterized in that: The described facial expression recognition method includes the following steps: Step 1: Obtain a facial expression training data set and perform preprocessing operations on the facial expression training data set; Step 2: Establish a facial expression recognition network model composed of a combination of an IR50 convolutional neural network and an enhanced self-attention Transformer; Step 3: Perform preliminary feature extraction by the IR50 convolutional neural network and splice the features in the intermediate stage; Step 4: Input the spliced features into the enhanced self-attention Transformer, and perform similar feature fusion and key feature screening in sequence; Step 5: Input the result after the enhanced self-attention Transformer is executed into the fully connected layer, and finally use a classifier to classify to obtain the expression classification result, where: In Step 4, the enhanced self-attention Transformer includes two parts, each part has 4 blocks. The first part is progressive feature fusion, which fuses word tokens with high content similarity. The second part is progressive feature screening, which gradually discards tokens with low relevance to the class token Class token. The two parts of the enhanced self-attention Transformer model reduce the number of tokens and speed up the inference speed.
2. The face expression recognition method based on an enhanced self-attention Transformer according to claim 1, wherein: In Step 1, the preprocessing operations on the facial expression training data set include the following steps: Step 1.1: Adjust the image size in the facial expression training data set to 112×112 and perform data augmentation operations on the facial expression training data set; Step 1.2: Perform normalization processing on the augmented images; Step 1.3: Perform masking processing on the images in the facial expression training data set, and the probability of performing the masking operation is 0.
7. The masking method is to fill random values.
3. The face expression recognition method based on an enhanced self-attention Transformer according to claim 2, characterized in that: In Step 1.1, the data augmentation operation is specifically as follows: The images in the facial expression training data set are rotated by 6° and horizontally flipped, and the probability of performing data augmentation is 0.
5.
4. A method for facial expression recognition based on an enhanced self-attention Transformer according to claim 1, characterized in that: The IR50 convolutional neural network in Step 3 consists of four stages: The first stage is shallow feature extraction, which includes convolution, normalization, and activation functions; The second stage, the third stage, and the fourth stage are multiple residual blocks, and the corresponding feature dimensions are 64×56×56, 128×28×28, and 256×14×14 respectively; Splice the features of the second stage, the third stage, and the fourth stage. Among them, the second stage and the third stage are downsampled and then spliced with the fourth stage in the channel dimension. Specifically: The channel is expanded from 448 to 768 through a 1×1 convolution, and then the feature dimension of channel 768 is adjusted to 196×768 by using Flatten for dimensionality reduction and transposition.
5. The face expression recognition method based on an enhanced self-attention Transformer according to claim 1, characterized in that: In the enhanced self-attention Transformer, a feature fusion module is introduced in the first part, and K in Q (query), K (key), and V (value) in the self-attention calculation formula is used as a measure of similarity. The K value calculation formula is as follows: K = F × W K Among them, F is the input feature, and W K is a learnable parameter matrix. K contains the information of each token. By calculating the dot product of K, the similarity between tokens is quantified, and then the similar tokens are fused.
6. The face expression recognition method based on an enhanced self-attention Transformer according to claim 5, wherein: The specific steps for fusing similar tokens include the following: Step 4.1: Divide all tokens input into the enhanced self-attention Transformer into two sets A and B, where tokens with odd indices belong to A and tokens with even indices belong to B; Step 4.2: For each token in set A, search for the most similar token in B, and calculate the dot product using the corresponding K value to obtain the similarity; Step 4.3: Retain the r combinations with the highest similarity, and average and fuse the 2r token features in the r groups respectively to obtain r fused tokens, where r is a hyperparameter.
7. A method for facial expression recognition based on an enhanced self-attention Transformer according to claim 1, characterized in that: In the step-by-step feature screening section, a feature selection module is introduced. In the self-attention mechanism, the self-attention layer realizes the dynamic aggregation of information through the interaction between Q (query), K (key), and V (value). The self-attention calculation formula for the class token is: Among them, q c is the Q (query) query vector corresponding to the class token, K is the key vector corresponding to each token, V is the value vector corresponding to each token, d is the dimension of K, and the output of the above formula is the value vector V = [v1, v2,..., v n T is a linear combination of, score is the interaction calculation of the class token with respect to the remaining tokens except the class token, v i comes from the i-th token, and the value score i determines how much information of the i-th token is fused into the output of the class token through the linear combination. The value score i represents the importance of the i-th token, which is expressed as the importance score of the i-th token. The corresponding mathematical formula is: Retain the top N tokens according to the score from high to low, and discard the remaining tokens. The discarded information is interference information.
Citation Information
Patent Citations
Micro-expression recognition method based on frequency domain characteristics
CN116030521A
KR20230088616A
Cited By
Video facial expression recognition method, system and device based on face key point optimization region features, processor and storage medium
CN117877081A