A lightweight macro-expression recognition method based on an effective attention mechanism

Through the DSC-LNN model combined with a small-scale convolution kernel, deep separable convolution and mixed attention mechanism, the problem of heavy computing and storage burden in macro-expression recognition is solved, and the accuracy and generalization ability of expression recognition are improved.

CN117058734BActive Publication Date: 2025-07-11ARTIFICIAL INTELLIGENCE RES INST OF HEFEI COMPREHENSIVE NAT SCI CENT (ANHUI ARTIFICIAL INTELLIGENCE LAB)
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202310876233.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-07-17
Publication Date
2025-07-11
Estimated Expiration
2043-07-17

AI Technical Summary

Technical Problem

The existing macro-expression recognition methods rely on manual design features, poor generalization capabilities, heavy computing and storage burdens of deep learning models, and traditional methods fail to effectively utilize the key channels and regions of the image. The SoftMax loss function clusters features in the expression recognition task.

Method used

The DSC-LNN model based on deep separable convolution and lightweight network is adopted, combining small-scale convolution kernels, deep separable convolution, mixed attention mechanism, island loss and SoftMax loss function, and extract key information through an effective attention mechanism and enhance discriminant ability.

Benefits of technology

It realizes lightweight high-performance expression recognition, which reduces computing and storage requirements, and improves the accuracy and generalization capabilities of expression recognition.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN117058734B_ABST
    Figure CN117058734B_ABST
Patent Text Reader

Abstract

The present invention relates to the technical field of facial expression recognition, and discloses a lightweight macro-expression recognition method based on an effective attention mechanism. The construction method of the DSC-LNN network model includes: building an image feature extraction block composed of three layers of small-scale convolutional kernels, and two convolutional kernel feature extraction blocks based on depthwise separable convolutions; constructing a channel attention module; constructing a spatial attention module; serially connecting the constructed channel attention module and spatial attention module in sequence to form an effective attention mechanism module, and embedding it into the convolutional kernel feature extraction block. Input the feature vector output by the Dropout layer into the fully connected layer and the classifier for feature fusion and expression classification to obtain the final classification result. The parameters of the depthwise separable convolution are significantly fewer and the computational requirements are lower. In addition, since the dataset for facial expression recognition is usually small and there is a risk of overfitting, a lightweight network is more suitable for this application scenario.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of facial expression recognition, and particularly relates to a lightweight macro-expression recognition method based on an effective attention mechanism. Background Art

[0002] Macro-expression recognition is an important branch of facial emotion analysis. Macro-expressions include six basic human expressions: anger, disgust, fear, happiness, sadness, and surprise. Previous facial expression recognition methods usually relied on manually designed features, which were only applicable to specific application fields and could not fully represent image information. The recognition accuracy of these methods was also easily affected by scene changes, and their generalization ability was poor.

[0003] In recent years, the emergence of deep learning has provided new ideas for the problem of macro-expression recognition. Classic networks such as VGGNet, Alexnet, and ResNet have shown good performance in various tasks. At the same time, the emergence of technologies such as data augmentation, transfer learning, and model fusion has further improved the performance of deep learning models. However, with the continuous increase in the number of network layers and the number of iterations, deep learning models face significant challenges in storage and calculation on terminals. Moreover, traditional deep learning assigns equal importance to all channels of an image in the facial expression recognition task, thus ignoring the key channels that have a more significant impact on recognition. To solve this problem, some studies have proposed applying the attention mechanism to existing models.

[0004] The hybrid attention mechanism includes two parts: the spatial attention mechanism and the channel attention mechanism. The main role of the spatial attention mechanism is to weight different regions of the input image, so that the network pays more attention to key regions during training. The channel attention mechanism enhances the performance of the neural network by selectively amplifying important channels in the feature map and suppressing less relevant channels. As a combination of the spatial attention mechanism and the channel attention mechanism, the hybrid attention mechanism can simultaneously extract key information and important channels in the input image, and its principle is similar to the "integrated specificity" in the human visual system, that is, the comprehensive importance of different colors and shapes during the perception process.

[0005] In classification tasks, different from face recognition and face verification tasks where one category represents only one person, there are multiple subjects in one category in face expression recognition. Images belonging to the same expression category also have certain differences in appearance, gender, skin color, and age, while the differences between different expression categories of the same subject are relatively small. Therefore, previous models using the SoftMax loss function as a classifier are not sensitive to the clustering features in the face expression recognition task. Summary of the Invention

[0006] To solve the above technical problems, the present invention provides a lightweight macro-expression recognition method based on an effective attention mechanism.

[0007] To solve the above technical problems, the present invention adopts the following technical solutions:

[0008] A lightweight macro-expression recognition method based on an effective attention mechanism, the DSC-LNN network model adopted includes an image feature extraction block FEBlock, a convolutional kernel feature extraction block DSC-FEBlock, a Dropout layer, a fully connected layer, and a classifier; the construction method of the DSC-LNN network model includes the following steps:

[0009] Step 1, build an image feature extraction block FEBlock composed of three layers of small-scale convolutional kernels, and two convolutional kernel feature extraction blocks DSC-FEBlock based on depthwise separable convolutions; the image feature extraction block FEBlock includes a convolutional layer, a batch normalization layer, an activation function layer, and a max pooling layer, and the activation function layer uses a Swish function with smooth characteristics; the convolutional kernel feature extraction block DSC-FEBlock has the same structure as the image feature extraction block FEBlock except for the following: DSC-FEBlock replaces the convolutional layer in FEBlock with a depthwise separable convolutional layer; the number of channels of the three layers of FEBlock and the two layers of DSC-FEBlock are 32, 64, 128, 256, and 256 in sequence;

[0010] Step 2, construct a channel attention module; the channel attention module first performs global average pooling and max pooling on the feature map input by the previous layer to generate two different feature maps, then processes the two different feature maps separately through a shared one-dimensional convolution, then uses element-wise summation to merge the output feature vectors, and finally generates the final channel attention feature map M through a Sigmoid activation operation c ;

[0011] Step 3, construct a spatial attention module; the spatial attention module performs max pooling and average pooling on the channel attention feature map M c along the channel dimension, splices the obtained feature vectors along the channel dimension to obtain a feature map, then uses convolutional operations to extract the spatial correlation between the feature maps, and generates a spatial attention map through a Sigmoid activation function > s ;

[0012] Step 4, serially connect the constructed channel attention module and spatial attention module in sequence to form an effective attention mechanism module EAM, and embed the effective attention mechanism module EAM into the convolutional kernel feature extraction block DSC-FEBlock, and the embedding position is after the activation function and before the max pooling layer of DSC-FEBlock;

[0013] Step 5: Input the feature vectors output by the Dropout layer into the fully connected layer and the classifier for feature fusion and expression classification to obtain the final classification result; the loss function of the classifier consists of the island loss and the SoftMax loss function.

[0014] Further, in Step 2, the channel attention module uses global average pooling AvgPool(F) and max pooling MaxPool(F) on the feature map F input by the previous layer to generate two different feature maps and ; then, process and respectively through a shared one-dimensional convolution, then use element-wise summation to merge the output feature vectors, and finally generate the final channel attention feature map M through the Sigmoid activation operation σ c , and the above process is expressed by the formula as follows:

[0015]

[0016] where C1D k represents a 1D convolution with a convolution kernel size of k.

[0017] Further, the loss function of the classifier in Step 5 consists of the island loss and the SoftMax loss;

[0018] The island loss is based on the center loss L C :

[0019]

[0020] In the formula, N represents the size of the batch data, x i represents the feature vector of sample i, y i represents the true label of sample i, represents the center vector of the class to which sample i belongs, and ‖·‖ represents the Euclidean distance;

[0021] The island loss L IL is:

[0022]

[0023] where, c k and c j respectively represent the kth and jth centers of the L2 norms ‖c k ‖2 and ||c j ||2, (·) represents the dot product, and λ1 is a hyperparameter used for balancing;

[0024] Then the loss function L of the classifier is:

[0025] L = L S + λL IL ;

[0026] where L S is the SoftMax loss, and λ is a parameter used to balance the SoftMax loss and the island loss.

[0027] Compared with the prior art, the beneficial technical effects of the present invention are:

[0028] The present invention designs a lightweight neural network model DSC-LNN based on the depthwise separable convolution technology. Compared with the traditional convolutional layer, the depthwise separable convolution has significantly fewer parameters and lower computational requirements. In addition, since the dataset for facial expression recognition is usually small and there is a risk of overfitting, the lightweight network is more suitable for this application scenario.

[0029] In the selection of the activation function, the present invention uses the Swish function with smooth characteristics to replace the traditional ReLU activation function. The Swish function also has the advantages of simplicity and easy calculation. Different from the ReLU function, Swish not only does not disappear in the negative value interval, but also smoothly transitions to non-zero values, effectively alleviating the problems of gradient vanishing and gradient explosion.

[0030] The channel attention module of the traditional hybrid attention mechanism adopts a shared MLP structure composed of two fully connected layers. The dimensionality reduction operation of the first fully connected layer weakens the direct correspondence between the channels and their weights while reducing the complexity of the model. The technical solution provided by the present invention uses a one-dimensional convolution with a receptive field of k to replace the MLP of the channel attention module, achieving the purpose of reducing model parameters while enhancing the network's ability to focus on effective information.

[0031] Aiming at the problem of large intra-class differences and small inter-class differences in facial expression recognition, the present invention introduces the joint use of the island loss and the Softmax loss function. The island loss is based on the center loss. The center loss only considers the intra-class distance, while the island loss not only considers the intra-class distance but also the distance between different classes. It further enhances the discriminative ability of the learned deep features by increasing the inter-class differences. BRIEF DESCRIPTION OF THE DRAWINGS

[0032] Figure 1 is a schematic structural diagram of the DSC-LNN network model of the present invention;

[0033] Figure 2Schematic diagrams of the image feature extraction block FEBlock and the convolutional kernel feature extraction block DSC-FEBlock of the present invention; among them, (a) is the schematic diagram of the image feature extraction block FEBlock, and (b) is the schematic diagram of the convolutional kernel feature extraction block DSC-FEBlock;

[0034] Figure 3 Schematic diagram of the DSC-FEBlock of the present invention incorporating the effective attention mechanism module EAM. Specific implementation manners

[0035] The following will make a detailed description of a preferred implementation manner of the present invention with reference to the accompanying drawings.

[0036] The DSC-LNN network model of the present invention combines the small-scale convolutional kernel and the depthwise separable convolution technology, and proposes a lightweight convolutional kernel feature extraction block DSC-FEBlock, which effectively reduces the number of parameters and the amount of computation, while reducing the risk of overfitting. Aiming at the problem that traditional deep learning methods in macro-expression recognition pay the same attention to different regions and different feature channels, the present invention proposes an effective attention module EAM based on the hybrid attention mechanism. This module uses a one-dimensional convolution with a receptive field of k as the bridge for local cross-channel interaction. Without dimensionality reduction, only a very small number of parameters need to be added to achieve high performance. Embedding EAM into the DSC-LNN network model realizes the dual benefits of network lightweight and high performance. In addition, the present invention also considers the problem that the inter-class difference is small while the intra-class difference is large in the facial expression recognition task, and proposes to jointly use the island loss and the Softmax loss function to train the proposed network, thereby effectively improving the recognition rate of facial expressions.

[0037] The DSC-LNN network model adopted by the present invention is as Figure 1 shown, and is composed of an image feature extraction block FEBlock, a convolutional kernel feature extraction block DSC-FEBlock, a Dropout layer, a fully connected layer, and a classifier. The specific construction method of the DSC-LNN network model includes the following steps:

[0038] Step 1, construct an image feature extraction block FEBlock composed of small-scale convolutional kernels and a small-scale convolutional kernel feature extraction block DSC-FEBlock based on depthwise separable convolution. As Figure 2As shown in (a), the FEBlock consists of a convolutional layer (Conv), a batch normalization layer (Batch Normalization), an activation function layer (Swish), and a max pooling layer (Max pooling). The size of its convolutional kernel is 3x3, the size of the pooling kernel is 2x2, and the Swish function with smoothing characteristics is used in the activation function layer. The DSC-FEBlock replaces the traditional convolutional layer in the FEBlock with a depthwise separable convolutional layer (DSC). After multiple experimental verifications, the number of channels in the three-layer FEBlock and the two-layer DSC-FEBlock structures is set to 32, 64, 128, 256, 256, as shown in Figure 1 shown.

[0039] Among them, the small-scale convolutional kernel is a relative concept, compared with the 5x5 or larger-scale convolutional kernels used in the past. The small-sized convolutional kernel can better capture the local features in the input image. In this embodiment, the size of the small-scale convolutional kernel is greater than or equal to 3x3 and less than 5x5; among them, the 3x3 convolutional kernel is the smallest convolutional kernel that can capture the information of the eight neighborhoods of pixels.

[0040] Step 2: Construct a channel attention module. The channel attention module first performs global average pooling and max pooling on the feature map input by the previous layer, then processes them respectively through a shared one-dimensional convolution, then uses element-wise summation to merge the output feature vectors, and finally generates the final channel attention feature map M through a Sigmoid activation operation. c .

[0041] Step 3: Construct a spatial attention module. The spatial attention module performs max pooling and average pooling on the feature map M generated in Step 2 along the channel dimension, concatenates the obtained feature vectors along the channel dimension to obtain a feature map, then uses convolution operations to extract the spatial correlation between the feature maps, and generates a spatial attention map M through a Sigmoid activation function. c s .

[0042] Step 4: Serialize and connect the channel attention module and the spatial attention module constructed in Step 2 and Step 3 to form an effective attention mechanism module EAM, and embed it into the convolutional kernel feature extraction block DSC-FEBlock built in Step 1. As shown in Figure 3 shown, the specific embedding position is after the activation function and before the max pooling layer of the convolutional kernel feature extraction block DSC-FEBlock.

[0043] ​Step 5: Input the feature vectors obtained in the previous layer into a fully connected layer and a classifier composed of island loss and SoftMax loss function for feature fusion and expression classification to obtain the final classification result.

[0044] In Step 2, the channel attention module uses global average pooling AvgPool(F) and max pooling MaxPool(F) on the feature map F input in the previous layer to generate two different feature maps and . Then, through a shared one-dimensional convolution, and are processed respectively to learn the weights of each channel. Finally, element-wise summation is used to merge the output feature vectors, and after the Sigmoid activation operation σ, the final channel attention feature map M is generated c . The above process is shown in the following formula:

[0045]

[0046] where C1D k represents a 1D convolution with a kernel size of k.

[0047] In Step 5, the loss function of the classifier in the prediction stage consists of two parts: island loss and SoftMax loss. The island loss is based on the center loss. The center loss maintains a class center vector for each class, and optimizes the model by minimizing the distance between the feature vector of each sample and the class center vector of its corresponding class. The center loss L C is:

[0048]

[0049] In the formula, N represents the size of the batch data, x i represents the feature vector of sample i, y i represents the true label of sample i, represents the center vector of the class to which sample i belongs, and ‖·‖ represents the Euclidean distance. By minimizing the center loss, samples of the same class will be pulled towards their corresponding centers, thereby reducing the overall within-class difference.

[0050] The island loss, meanwhile, compresses each cluster and separates the cluster centers as isolated "islands". The island loss L IL is:

[0051]

[0052] where, c k and c j represent the L2 norm ‖c k ‖2 and ||cj ||2, (·) represents the dot product. Specifically, the first term aims to learn the center of each class in the feature space, which serves as a representative point for all samples belonging to that class. The center of each class is learned by minimizing the distance between the feature representation of the image and its corresponding class center. This encourages the feature representations of samples belonging to the same class to be close in the feature space. The second term aims to increase the distance between the class centers in the feature space so that they are well separated in the feature space. This helps to improve the discriminative ability of the feature representation, making it easier to distinguish different samples in the dataset. The balance between these two terms is controlled by the hyperparameter λ1, which is used to balance the two terms. By minimizing the island loss, samples of the same class are clustered and samples of different classes are separated.

[0053] Then the loss function L of the classifier is:

[0054] L=L S +λL IL ;

[0055] Where L S is the SoftMax loss, and λ is used to balance the SoftMax loss and the island loss. In this embodiment, λ=0.01 and λ1=10 are set based on experience.

[0056] It is obvious to those skilled in the art that the present invention is not limited to the details of the exemplary embodiments described above, and that the present invention can be implemented in other specific forms without departing from the spirit or essential features of the present invention. Therefore, from any point of view, the embodiments should be regarded as exemplary and non-limiting, and the scope of the present invention is defined by the appended claims rather than the above description, and it is intended that all changes falling within the meaning and scope of the equivalent elements of the claims are included in the present invention, and any reference numerals in the claims should not be regarded as limiting the claims involved.

[0057] In addition, it should be understood that although this specification is described in terms of implementation methods, not every implementation method contains only one independent technical solution. This narrative method of the specification is only for the sake of clarity. Those skilled in the art should regard the specification as a whole. The technical solutions in each embodiment may also be appropriately combined to form other implementation methods that can be understood by those skilled in the art.

Claims

1. A lightweight macro-expression recognition method based on an effective attention mechanism. The adopted DSC-LNN network model includes an image feature extraction block FEBlock, a convolutional kernel feature extraction block DSC-FEBlock, a Dropout layer, a fully connected layer, and a classifier. The construction method of the DSC-LNN network model includes the following steps: Step 1: Build a three-layer image feature extraction block FEBlock composed of small-scale convolutional kernels and a two-layer convolutional kernel feature extraction block DSC-FEBlock based on depthwise separable convolution. The image feature extraction block FEBlock includes a convolutional layer, a batch normalization layer, an activation function layer, and a max pooling layer. The activation function layer uses the Swish function with smooth characteristics. The convolutional kernel feature extraction block DSC-FEBlock has the same structure as the image feature extraction block FEBlock except for the following: The DSC-FEBlock replaces the convolutional layer in the FEBlock with a depthwise separable convolutional layer. The number of channels of the three-layer FEBlock and the two-layer DSC-FEBlock are 32, 64, 128, 256, and 256 in sequence. Step 2, construct a channel attention module; the channel attention module first performs global average pooling and max pooling on the feature map input from the previous layer to generate two different feature maps, then processes the two different feature maps respectively through a shared one-dimensional convolution, then uses element-wise summation to merge the output feature vectors, and finally through a Sigmoid activation operation to generate the final channel attention feature map. c ; Step 3: Construct a spatial attention module; the spatial attention module performs max-pooling and average-pooling on the channel attention feature map > c along the channel dimension, concatenates the obtained feature vectors along the channel dimension to obtain a feature map, then uses convolutional operations to extract the spatial correlation between the feature maps, and generates a spatial attention map c through the Sigmoid activation function s ; Step 4: Serially connect the constructed channel attention module and spatial attention module in sequence to form an effective attention mechanism module EAM, and embed the effective attention mechanism module EAM into the convolutional kernel feature extraction block DSC-FEBlock. The embedding position is after the activation function and before the max pooling layer of the DSC-FEBlock. Step 5: Input the feature vector output by the Dropout layer into the fully connected layer and the classifier for feature fusion and expression classification to obtain the final classification result. The loss function of the classifier consists of the island loss and the SoftMax loss function.

2. The lightweight macro-expression recognition method based on an effective attention mechanism according to claim 1, characterized in that In step two, the channel attention module uses global average pooling AvgPool(F) and max pooling MaxPool(F) on the feature map F input from the previous layer to generate two different feature maps and Then, they are processed respectively through a shared one-dimensional convolution on and The output feature vectors are then merged using element-wise summation, and finally, through the Sigmoid activation operation σ, the final channel attention feature map M is generated c , which is expressed by the formula as follows: Among them, C1D k represents a 1D convolution with a convolution kernel size of k.

3. The lightweight macro-expression recognition method based on an effective attention mechanism according to claim 1, characterized in that, The loss function of the classifier in Step 5 consists of the island loss and the SoftMax loss. The island loss is based on the central loss L C as follows: where N represents the size of the batch data, and x i represents the feature vector of sample i, and y i represents the true label of sample i, represents the center vector of the class to which sample i belongs, and ‖·‖ represents the Euclidean distance; Island loss L IL is as follows: Among them, c k and c j respectively represent the k-th and j-th centers of the L2 norm ‖c k ‖2 and ||c j ||2, (·) represents the dot product, and λ1 is a hyperparameter for balancing; Then the loss function L of the classifier is: L = L S + λL IL ; where L S is the SoftMax loss, and λ is a parameter used to balance the SoftMax loss and the island loss.

Citation Information

Patent Citations

  • Three-dimensional face expression recognition method based on SSF-IL-CNN

    CN110188621A

  • Lightweight network facial expression recognition method fusing equilibrium loss

    CN113128369A