Face Expression Recognition Method Based on Spatial Attention and Deep Neural Network

The method addresses model complexity and class imbalance in face expression recognition by using an inverse residual layer-based network with spatial attention and Focal loss, achieving improved accuracy and classification performance.

CN115171180BActive Publication Date: 2025-07-15SOUTHEAST UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202210598202.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-05-30
Publication Date
2025-07-15
Estimated Expiration
2042-05-30

AI Technical Summary

Technical Problem

The existing facial expression recognition methods based on deep learning have high model complexity and are difficult to meet real-time requirements. The accuracy of certain categories in uneven expression data sets is low, resulting in limited network performance.

Method used

The network structure and spatial attention mechanism based on the inverse residual layer are adopted, combined with the Focal loss loss function, to improve the learning effect of difficult-to-classify categories.

Benefits of technology

Improve the accuracy of facial expression recognition, especially in categories with small samples, and improve the overall and average classification accuracy.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115171180B_ABST
    Figure CN115171180B_ABST
Patent Text Reader

Abstract

The present invention discloses a face expression recognition method based on spatial attention and deep neural network, which can be used to realize real-time detection of human face expressions. The invention mainly includes: obtaining a face expression data set, dividing it into a training set, a validation set and a test set according to a certain proportion, extracting the face part in the image, and performing data augmentation operations such as rotation, horizontal flipping, cropping, color jittering, etc.; training an expression classification network, using a convolutional neural network based on inverse residual layers to extract image features, then distinguishing the importance of each region of the image through a spatial attention mechanism, and classifying the face image based on the obtained features; evaluating the recognition effect of the network on the face expression classification data set. Compared with the current main face expression classification algorithms, the present invention obtains a higher average classification accuracy, and the algorithm has high real-time performance, which is a high-quality face expression recognition algorithm.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to a face expression recognition method based on spatial attention and deep neural network, and is applicable to the technical field of face expression recognition in computer vision. Background Art

[0002] As one of the popular research topics in the field of computer vision, face expression analysis aims to determine the expression state of a face based on the input face image. Face expression analysis has a very wide range of application scenarios, which has a significant effect on improving the quality of human life and is of great research value. Specific application fields include but are not limited to: in the vehicle driving scenario, analyzing whether the driver of a motor vehicle is fatigued or driving under the influence of alcohol based on the facial expression of the driver, and giving an early warning; in the social public area scenario, performing real-time expression analysis on multiple targets to timely identify potential dangers; in the daily life scenario, analyzing the facial expressions of people around to help disabled people understand the emotional states of others and facilitate their communication.

[0003] With the development of technology, people have increasingly tried to use machine vision and image processing technologies to achieve automated face expression recognition. In recent years, deep learning algorithms have been widely applied to various fields such as natural language processing, data mining, and image processing due to their powerful learning ability and adaptability. Some methods based on convolutional neural networks have also been introduced into the field of face expression analysis, including VGG-16, Inception-V3, ResNet-50, etc. Thanks to the local connection and weight sharing mechanisms of convolutional neural networks, the parameters and computational amount of face image processing are greatly reduced, and the accuracy of expression recognition is also significantly improved compared with traditional methods. Although the face expression analysis methods based on deep learning have achieved good results, there are still some problems in the current methods, including: blindly stacking convolutional layers will lead to a higher model complexity and cannot meet the real-time requirements in actual use scenarios; in the existing public expression datasets, the expressions of each category are often unevenly distributed, and some expression classifications are relatively difficult, and the low accuracy on these expression categories results in a lower average accuracy index of the network, which limits the overall performance of the network.

[0004] Therefore, it is of great significance to carry out research on algorithms based on spatial attention features. Summary of the Invention

[0005] The present invention proposes a facial expression recognition method based on spatial attention and deep neural network. The network model in the present invention uses a network structure based on reverse residual layers for image feature extraction, and adopts a spatial attention mechanism to pay higher attention to the important regions affecting facial expression classification. At the same time, Focal loss is used as the loss function, which improves the learning effect of the network for difficult-to-classify categories. The network model in the present invention has good recognition performance after training, and has a higher recognition rate compared with existing algorithms such as IACNN (Identity-aware Convolutional Neural Network), indicating that the present invention effectively improves the accuracy of facial expression recognition.

[0006] To achieve the above object, the present invention provides the following technical solutions:

[0007] A facial expression recognition method based on spatial attention and deep neural network, the method comprising the following steps:

[0008] Step S1: According to the publicly available facial expression dataset, construct a training set, a validation set and a test set; extract the face part in the image through preprocessing, and perform data augmentation operations such as rotation, horizontal flipping, cropping, and color jittering;

[0009] Step S2: Construct a facial expression classification and recognition network based on reverse residual layers and spatial attention mechanism; use the constructed training set to supervise and train the network model until the network converges to the optimal performance; during the training process, use the validation set to reflect the training process;

[0010] Step S3: Test the converged network model on the constructed test set, and evaluate the network performance according to the overall classification accuracy, average classification accuracy and classification confusion matrix.

[0011] Preferably, each piece of data in the facial expression dataset in step S1 is given in the form of a data pair, including an input RGB image to be classified and the labeled expression category as the true value.

[0012] Preferably, the preprocessing in step S1 includes: first, extract the face contour in the input RGB face image to be classified; then, segment the face part from the original image, and change the pixel size of the segmented image to 224×224 by means of bilinear interpolation; finally, perform data augmentation processing, and the augmentation means include: random rotation, random cropping, horizontal flipping, randomly modifying the saturation and hue of the RGB image, randomly adding Gaussian noise, and finally normalizing the RGB image.

[0013] Preferably, the classification and recognition network includes:

[0014] a) Image feature extraction module: Its input is the preprocessed face image. This module is based on the inverted residual layer. When processing image features, it first increases the input dimension through the dilation layer, then fully extracts feature information using depthwise separable convolution, and finally maps the extracted features from the high-dimensional space back to the low-dimensional space through the projection layer, thereby reducing the subsequent computational amount. The input size of this module is 3×224×224, and it contains a total of 5 inverted residual layers. The output size of the last inverted residual layer is 640×7×7;

[0015] b) Spatial attention mechanism: Its input is the output features of the image feature extraction module, with a size of 96×14×14. This part utilizes the correlation in the spatial position of the image to adaptively judge the importance of each region in the image and assigns different weights to it, thereby obtaining more valuable information in the spatial dimension and extracting more effective spatial features. The final output size is 1×640;

[0016] c) Focal loss function: The network model uses Focal loss as the loss function during training. By adding a dynamic adjustment factor and a sample balance factor to the difficult-to-classify categories, the classification accuracy of the network model for categories with a small number of samples is improved, thereby increasing the average classification accuracy of the network model.

[0017] Preferably, the Focal loss is specifically: The Focal loss in the multi-classification task scenario can be expressed as

[0018]

[0019] In the above formula, p(i) is the true sample class distribution, q(i) is the predicted sample class distribution, and α k is the sample balance factor for class k, and the larger α k is set, the greater the weight of this class in the final loss; γ k is the dynamic adjustment factor for class k, and the larger γ k is set, the greater the weight of this class in the final loss.

[0020] The beneficial effects of the present invention are:

[0021] The network model proposed by the present invention uses a feature extraction module based on the inverted residual layer and a spatial attention mechanism for facial expression recognition, strengthens the representation of facial expression features, improves the representation ability, and further improves the detection effect. The gain is specifically reflected in that in terms of the overall classification accuracy and the average classification accuracy, the present invention can obtain a higher accuracy rate compared with the existing methods. Brief Description of the Drawings

[0022] Figure 1 It is a flowchart of the facial expression recognition method based on spatial attention and deep neural network in Embodiment 1.

[0023] Figure 2 It is a flowchart of the training and prediction of the expression classification network in Embodiment 1.

[0024] Figure 3 It is a structural diagram of the expression classification network model in Embodiment 1.

[0025] Figure 4 It is a structural diagram of the inverted residual layer in Embodiment 1.

[0026] Figure 5 It is an effect diagram of facial expression recognition in Embodiment 1.

[0027] Figure 6 It is the recognition accuracy value in Embodiment 1. Detailed implementation manners

[0028] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention. Apparently, the described embodiments are some, but not all, of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.

[0029] Embodiment 1

[0030] Refer to Figures 1 - 6 , this embodiment provides a facial expression recognition method based on spatial attention and deep neural network. This method trains an expression classification model through a deep learning method, and this model can output the predicted expression categories according to the input facial RGB image.

[0031] Specifically, refer to Figure 1 , this method specifically includes:

[0032] Step S1: Construct a training set, a validation set, and a test set according to the publicly available facial expression dataset; extract the facial part in the image through preprocessing, and perform data augmentation operations such as rotation, horizontal flipping, cropping, and color jittering;

[0033] More specifically, Step S1 includes: obtaining the dataset for facial expression recognition. Each piece of data in the dataset is given in the form of a data pair, including an RGB image to be classified as the input data and the labeled expression category as the true value. Among them, the expression category is represented by a number from 0 to 6 (0 - angry, 1 - contemptuous, 2 - disgusting, 3 - afraid, 4 - happy, 5 - sad, 6 - surprised).

[0034] Before each piece of data in the dataset is input into the model, preprocessing is required. First, extract the facial contour from the input RGB face image to be classified; then, segment the face part from the original image, and change the pixel size of the segmented image to 224×224 by means of bilinear interpolation; finally, perform data augmentation processing, and the augmentation means include: random rotation, random cropping, horizontal flipping, randomly modifying the saturation and hue of the RGB image, randomly adding Gaussian noise, and finally standardize the RGB image.

[0035] Step S2: Train the facial expression classification network, extract image features using a convolutional neural network based on inverse residual layers and spatial attention, and output the classification result;

[0036] More specifically, in this embodiment, step S2 includes: splitting the dataset into a training set containing 854 images and a test set containing 127 images, training the facial expression classification network based on the Pytorch framework using the training set, fixing the network parameters of the model when the network parameters in the network model reach the convergence standard, obtaining the trained facial expression classification network, and finally using the trained facial expression network to predict the test set to obtain the facial expression classification result.

[0037] It should be noted that the facial expression classification method provided in this embodiment is not limited to the Pytorch framework, as long as it can train the dataset X, and after iterating several times (the order of magnitude of times) during the training process, the loss function converges and finally can accurately classify the facial expression according to the RGB image.

[0038] It should be noted that in this example, the facial expression classification network described in step S2 includes the following main components:

[0039] a) Image feature extraction module: Its input is the preprocessed face image. This module is based on the inverse residual layer architecture. When processing image features, first increase the input dimension through the dilation layer, then fully extract feature information using depthwise separable convolution, and finally map the extracted features from the high-dimensional space back to the low-dimensional space through the projection layer, thereby reducing the subsequent calculation amount. The input size of this module is 3×224×224, and it contains a total of 5 inverse residual layers. The output size of the last inverse residual layer is 640×7×7;

[0040] b) Spatial attention mechanism: Its input is the output feature of the image feature extraction module, with a size of 96×14×14. This part uses the correlation in the spatial position of the image to adaptively judge the importance of each region in the image and assign different weights to it, so as to obtain more valuable information in the spatial dimension, thereby extracting more effective spatial features. The final output size is 1×640;

[0041] c) Focal loss loss function: The network model uses Focal loss as the loss function during training. By adding a dynamic adjustment factor and a sample balance factor to the difficult-to-classify categories, the classification accuracy of the network model for categories with fewer samples is improved, thereby increasing the average classification accuracy of the network model.

[0042] Step S3: Calculate the average classification accuracy according to the classification results and the true labels.

[0043] More specifically, in this embodiment, using the prediction results described in step S2 and the true values described in step S1, the overall classification accuracy and the average classification accuracy are calculated based on the number of correctly classified samples and the total number of samples in the test set.

[0044] It should be noted that the definition of the overall classification accuracy is the ratio of the number of correctly classified samples to the total number of samples in the test set, and the definition of the average classification accuracy is the average of the classification accuracies of each category; the indicators for measuring the facial expression recognition effect are not limited to the overall classification accuracy and the average classification accuracy. Any indicator that can show the difference between the obtained facial expression classification results and the true facial expression labels can be used as a measurement indicator.

[0045] As Figure 3 shown, this model is the overall structure diagram of the facial expression classification network, mainly including three modules: an image feature extraction module, a spatial attention mechanism, and a Focal loss loss function. Since the Focal loss loss function is only used during the training process, it is not shown in the network structure diagram.

[0046] The input of the image feature extraction module is the preprocessed face image. The basic structure of the network is an inverse residual layer. There are 5 inverse residual layers in this module, which are shown as IRes Layer 1 to IRes Layer 5 in the figure. After obtaining the image, this module downsamples the image and increases the channel dimension through layer1-layer5, and then performs high-dimensional feature extraction. As the number of layers increases, the size of the image features will gradually decrease, while the feature dimension will gradually increase, and finally an output feature of 640×7×7 is obtained.

[0047] The input of the spatial attention mechanism is the output features of the image feature extraction module. After obtaining the image feature information, the features are first unfolded in the width and height dimensions to become 640×49, and then Softmax is used to activate them. This is to classify each segment of the features according to their importance, that is, to determine the importance of each region of the image. Then, the features are transformed into 1×640 through global average pooling as the output features of the network. Finally, the probability of each expression (a total of 7 kinds) is judged through a fully connected layer, and a classification judgment is made on the expression.

[0048] As Figure 4 shown, taking IRes Layer 1 as an example, the structure of the inverse residual layer is shown. Each inverse residual layer is composed of 2 inverse residual blocks. The first inverse residual block downsamples the size of the input features, while the second inverse residual block keeps the size of the input features unchanged. When processing image features, the inverse residual block first increases the input dimension through a dilation layer, then uses depthwise separable convolution to fully extract the feature information, and finally projects the extracted features from the high-dimensional space back to the low-dimensional space through a projection layer, thereby reducing the subsequent calculation amount.

[0049] As Figure 5 shown, it is a partial effect diagram of face expression recognition using the expression classification network. Obviously, from a qualitative perspective, compared with the comparison algorithm, the classification results output by the method proposed in the present invention are better and more accurate. As Figure 6 shown, from a quantitative perspective, the classification accuracy of the method proposed in the present invention is higher than that of other algorithms, so the performance is better.

[0050] Where the present invention is not described in detail are all well-known technologies to those skilled in the art.

[0051] The preferred specific embodiments of the present invention have been described in detail above. It should be understood that those of ordinary skill in the art can make many modifications and variations based on the concept of the present invention without creative work. Therefore, all technical solutions that can be obtained by those skilled in the art in the technical field of the present invention based on the concept of the present invention through logical analysis, reasoning, or limited experiments on the basis of the prior art should be within the protection scope determined by the claims.

Claims

1. A face expression recognition method based on spatial attention and deep neural network, characterized in that, The method includes the following steps: Step S1: Construct a training set, a validation set, and a test set based on the publicly available facial expression dataset; extract the facial part in the image through preprocessing, and perform data augmentation operations such as rotation, horizontal flipping, cropping, and color jittering; Step S2: Construct a facial expression classification and recognition network based on the inverted residual layer and the spatial attention mechanism; use the constructed training set to supervise the training of the network model until the network converges to the optimal performance; during the training process, use the validation set to reflect the training process; Step S3: Test the converged network model on the constructed test set, and evaluate the network performance according to the overall classification accuracy, the average classification accuracy, and the classification confusion matrix; The preprocessing in step S1 includes: first, extract the facial contour in the input RGB facial image to be classified; then, segment the facial part from the original image, and perform bilinear interpolation on the pixels of the segmented image; finally, perform data augmentation processing, and the augmentation means include: random rotation, random cropping, horizontal flipping, randomly modifying the saturation and hue of the RGB image, randomly adding Gaussian noise, and finally normalizing the RGB image; The classification and recognition network includes: a) Image feature extraction module: Its input is the preprocessed facial image. This module includes five inverted residual layers, and each inverted residual layer is composed of two inverted residual blocks. The first inverted residual block downsamples the size of the input feature, and the second inverted residual block keeps the size of the input feature unchanged. When processing image features, the inverted residual block first increases the input dimension through the dilation layer, then uses the depthwise separable convolution to fully extract the feature information, and finally projects the extracted feature from the high-dimensional space back to the low-dimensional space through the projection layer, so as to reduce the subsequent calculation amount; b) Spatial attention mechanism: Its input is the output feature of the image feature extraction module. After obtaining the image feature information, first expand the feature in the width and height dimensions, then use softmax to activate it, and finally use global average pooling to process the output of softmax; c) Focal loss function: The network model uses Focal loss as the loss function during training. By adding a dynamic adjustment factor and a sample balance factor to the difficult-to-classify categories, the classification accuracy of the network model in the categories with fewer samples is improved, thereby improving the average classification accuracy of the network model.

2. The face expression recognition method based on spatial attention and deep neural network according to claim 1, wherein Each piece of data in the facial expression dataset in step S1 is given in the form of a data pair, including an input RGB image to be classified and the labeled expression category as the ground truth.

3. The face expression recognition method based on spatial attention and deep neural network according to claim 1, wherein The Focal loss is specifically: The Focal loss in the multi-classification task scenario is expressed as In the above formula, p(i) is the true sample class distribution, q(i) is the predicted sample class distribution, and α k is the sample balance factor for class k, and the larger α k is set, the greater the weight of this class in the final loss; γ k is the dynamic adjustment factor for class k, and the larger γ k is set, the greater the weight of this class in the final loss.