A lightweight facial expression recognition method based on linear self-attention

By introducing a linear self-attention mechanism in facial expression recognition, combined with local space and channel attention modules, the problem of difficulty in learning local regional correlation and global characteristics of expressions in the prior art is solved, and efficient and accurate expression recognition is achieved.

CN117912083BActive Publication Date: 2025-05-09HU BEI SHENG SAN WA SHI PIN YOU XIAN GONG SI
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202410105719.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-01-25
Publication Date
2025-05-09
Estimated Expiration
2044-01-25

AI Technical Summary

Technical Problem

The prior art is difficult to effectively learn the correlation between different local areas of expression in facial expression recognition, and the CNN-based model cannot understand facial expression images globally, resulting in poor performance when processing inter-class similarity and intra-class differences.

Method used

A lightweight facial expression recognition method with linear self-attention is proposed. The long-distance dependence of expression area patches is captured through the global linear self-attention module, combined with the local spatial attention module and the channel attention module, feature fusion and weighting are performed, and expression classification is performed under the supervision of identification loss.

Benefits of technology

Effectively learning the global features of expressions reduces the computational complexity of self-attention, improves the efficiency and accuracy of expression recognition, and enhances the ability to handle inter-class similarities and intra-class differences.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN117912083B_ABST
    Figure CN117912083B_ABST
Patent Text Reader

Abstract

The present invention claims a linear self-attention lightweight facial expression recognition method (LSViT), which aims to design a lightweight facial expression recognition neural network and belongs to the field of pattern recognition. The method comprises the following steps: first, in view of the problem that the visual transformer (ViT) has many parameters, a lightweight network model LSViT based on CNN and ViT is designed, and by adopting multi-stage local-global feature parallel processing operations, the local and global features of the expression are effectively integrated, which can greatly reduce the complexity of multi-head self-attention while improving the accuracy of expression recognition. Secondly, a linear multi-head self-attention module is designed to fully capture the correlation characteristics between the various regions of the expression. In addition, a local spatial attention module is designed to allow the network to focus more on key information when processing facial expressions with interference factors. Finally, a discrimination loss is designed to expand the inter-class distance and minimize the intra-class distance, further improving the accuracy of expression recognition.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of computer vision, and in particular relates to a lightweight facial expression recognition method based on a self-attention mechanism. Background Art

[0002] Facial expression recognition (FER) has attracted increasing attention in the field of computer vision due to its wide applications in human-computer interaction, public safety, driver fatigue monitoring, and other fields. However, it is still challenging to perform FER in the wild using deep neural networks due to the intra-class variations and inter-class similarities in facial expression images. The inter-class similarity indicates that different kinds of expressions may have similar muscle movements, while the intra-class variation indicates that there may be different ways of expressing the same expression.

[0003] It is worth noting that the application of convolutional neural networks (CNNs) has made great progress in FER. However, each convolution filter in CNN operates only on a small area. This spatial locality makes it difficult for the model to learn the structural dependencies between different facial units in most neural layers. Therefore, CNN-based FER models can only capture local facial features but cannot understand facial expression images globally. In previous studies, most CNN-based models obtain the features of images by simply stacking many convolutional layers. As the image passes through more and more convolutional layers, the loss of image features occurs, and the model cannot obtain complete global features in most cases. In addition, the convolution filters in CNN rely heavily on spatial locality and cannot learn the global features of facial expressions at the beginning of the model.

[0004] In general, convolutional filters in the backbone layers play a key role in learning low-level local features such as textures and edges. These features are crucial for performing fine-grained recognition tasks, such as facial expression recognition. Expression recognition network models must learn a global understanding of facial images by paying more attention to important local features at the initial stage of the network. Research works such as LSTM or Visual Transformer (ViT) are dedicated to solving this problem. In particular, ViT utilizes a multi-head self-attention mechanism to generate an attention map with long-range inductive biases learned from different facial patches. Therefore, the ViT-based mechanism may be suitable for compensating for the shortcomings of convolutional filters in learning long-range inductive biases. However, due to the quadratic computational complexity of the multi-head self-attention mechanism, ViT cannot achieve satisfactory results on resource-limited mobile devices. At the same time, since ViT has a strong ability to learn long-range inductive biases, it needs to learn spatial inductive biases from large-scale datasets (e.g., ImageNet). Unfortunately, such large datasets are usually not available for expression recognition tasks. Therefore, how to ensure the lightweight of ViT and retain its strong modeling ability in learning long-range inductive biases to overcome the weaknesses of CNN-based FER models without relying on large-scale expression datasets remains a challenging problem to be solved.

[0005] In addition, deep metric learning (DML) is one of the widely used methods to handle significant intra-class variations and inter-class similarities by improving the discriminative power of learned embedded features. Specifically, DML methods achieve intra-class compactness and inter-class separation by maximizing the similarity between deep features in the embedding space and their corresponding class prototypes. In recent years, researchers have proposed many deep DML methods to address the above problems. ArcFace is one of the loss functions used to learn generalized feature embeddings. It learns a well-structured intra-class feature distribution by pulling simple samples to the center of the class. This enhances face recognition in the wild and prevents the model from overfitting in noisy, low-quality scenes.

[0006] Therefore, how to alleviate the challenges of intra-class and inter-class differences and large parameters of visual transformers has become an urgent problem that has not been fully studied. In order to solve the above problems, the present invention proposes a lightweight facial expression recognition method based on linear self-attention.

[0007] After searching, CN115240261A, a method and device for facial expression recognition based on a hybrid attention mechanism. It includes: obtaining three-dimensional feature parameters and two-dimensional feature vectors of each connection layer channel through a three-dimensional feature image; using an attention module to encode the three-dimensional feature parameters and the two-dimensional feature vectors of the channels of each connection layer, obtaining the channel weight vectors of each connection layer, and mapping them to the original dimension through dot multiplication to obtain a channel attention feature image; using a convolutional layer to perform dimensionality reduction and feature extraction operations on the three-dimensional feature image to obtain spatial feature information, and expanding the spatial feature information to the same size as the three-dimensional feature image to obtain a spatial attention feature image; fusing the channel attention feature image with the spatial attention feature image, and outputting a fused feature image. The present invention combines channel attention with spatial attention, effectively avoiding the increase in the number of network layers, and improving the efficiency and accuracy of recognizing expressions.

[0008] However, this patent has deficiencies in the extraction of global features of expressions, and cannot well learn the correlation between different local areas of expressions, and fails to properly solve the problems of inter-class similarity and intra-class differences in expression recognition. Therefore, this patent takes advantage of the self-attention in global feature learning and designs an efficient linear self-attention, which can effectively learn the global features of expressions while reducing the computational complexity of self-attention. At the same time, a discrimination loss function is further designed to expand the distance between classes and reduce the distance within classes, thereby optimizing the accuracy of expression recognition. Summary of the invention

[0009] The present invention aims to solve the above problems of the prior art. A linear self-attention lightweight facial expression recognition method is proposed. The technical solution of the present invention is as follows:

[0010] A linear self-attention lightweight facial expression recognition method, comprising the following steps:

[0011] Step 1: Input the facial expression image into the expression recognition network, and feed the initial features after 3×3 convolution into the global linear self-attention module and the local spatial attention module respectively;

[0012] Step 2: The global linear self-attention module obtains global expression features by capturing the long-range dependencies of patches in different expression regions;

[0013] Step 3: The local spatial attention module performs spatial attention weighting through the MobileNext hourglass block and pooling operation. The hourglass block performs identity mapping and spatial transformation. The features learned by the hourglass block are respectively subjected to maximum pooling and average pooling operations to obtain the attention scores of important expression areas. Finally, the output attention feature map is obtained through feature weighting.

[0014] Step 4: The global and local expression features learned by the multi-stage feature extraction operation are fused, and finally the channel features are weighted by the channel attention module and input into the classifier for expression classification.

[0015] Furthermore, the step 1 specifically includes the following steps:

[0016] A1. Detect facial key points of the facial expression image through the face detection and alignment network MTCNN, align the facial expression image, and crop it into an input image I of 224×224 size;

[0017] A2. Input image I into two 3×3 convolutions to extract the primary features of the image, represented by x. Feed x to the global linear self-attention module and the local spatial attention module respectively. C, H, W represent the number of image channels, height, and width, respectively.

[0018] Furthermore, the step 2 specifically includes:

[0019] B1. Flatten the input low-level feature x into a feature sequence and generate three different feature token matrices Q, K, and V through linear mapping;

[0020] B2. For the token matrix Q, use the weights The linear layer of Q maps each token L in Q to a scalar; this linear projection is an inner product operation that calculates the distance between token L and x, resulting in a k-dimensional vector; the Softmax function is then applied to this k-dimensional vector to produce a context score

[0021] B3. The context score c s Weighting the input token matrix K produces a contextual attention matrix Z that encodes global information;

[0022] B4. Finally, in order to obtain the final global features, the token matrix V after deep convolution and the context attention matrix Z are Hadamard-multiplied to obtain the final global expression encoding features. The linear self-attention calculation formula is as follows:

[0023]

[0024] Among them, λ represents the scaling factor, ⊙ represents the Hadamard product, and DWConv represents the depthwise convolution operation.

[0025] Furthermore, the step 3 specifically includes:

[0026] C1. Set It is the input tensor after the initial convolution. X is passed through the hourglass module for new spatial feature interaction. After performing deep convolution on X, the channel dimension is reduced and expanded, and then it is further convolved to encode richer spatial feature information to obtain a new spatial feature map. The formula for the hourglass block is as follows:

[0027] G=φ e (DW(φ r (DW(X))))+X (2)

[0028] Among them, φ e and φ r They represent two point-by-point convolutions for channel expansion and channel reduction, respectively, and DW represents depthwise convolution;

[0029] C2. Aggregate the channel information of the feature map by using two pooling operations to generate two 2D feature maps: and To represent the average pooling features and maximum pooling features in the channel; then, the Sigmod activation function is used to generate the spatial attention map The spatial attention formula is calculated as:

[0030]

[0031] Among them, σ represents the sigmoid activation function.

[0032] Furthermore, the step 4 specifically includes the following steps:

[0033] D1. For the global and local expression features M(G) and N(G) learned through multi-stage feature extraction operations, the features are fused through the Concat operation. The formula is:

[0034] F=Concat([M(G);N(G)]) (4)D2. After the fused feature F is subjected to channel attention weighted interactive channel information by ECANet, it is input into the classifier for expression classification; the effective channel attention network ECANet is a neural network that can generate channel attention through fast one-dimensional convolution, and its convolution kernel size is adaptively determined by the nonlinear mapping of the channel dimension.

[0035] Furthermore, the calculation formula for the convolution kernel size of D2 is as follows:

[0036]

[0037] where |t| odd represents the odd number closest to t, C represents the number of input feature channels, γ represents the feature weight, and b represents the bias.

[0038] Furthermore, the step 4 also includes: the entire network is supervised by the identification loss Di-Loss to expand the distance between classes and reduce the distance within classes, specifically including the following steps:

[0039] E1. The softmax loss function formula for classification loss function is as follows:

[0040]

[0041] in represents the characteristics of the i-th sample belonging to the j-th class, represents the weight matrix of the i-th sample, represents the weight matrix of the jth class, b yi , b j represents bias, and N represents the number of classifications;

[0042] E2, Additive Angular Margin Loss ArcFace loss sets the bias of Softmax loss to 0, where θ j is the angle between the weight matrix and the feature; then, the weights ∥∥W are transformed by L2 j ∥∥Normalize, and convert ∥x i ∥Normalized to the scaling factor s; At the same time, ArcFace adds an additional corner margin penalty between the intra-class and inter-class distances to simultaneously enhance the intra-class compactness and inter-class differences; The ArcFace loss function formula is as follows:

[0043]

[0044] E3. The identification loss function Di-Loss is designed, which adds a constraint term based on the L2 norm to expand the distance between expressions; the Di Loss formula is expressed as follows:

[0045]

[0046] Among them, m is the total number of expression samples in the batch, N is the total number of expression categories, and c is j is the center of the jth expression sample, p j is the proportion of the j-th expression sample in batch m, and μ is the balance parameter.

[0047] An electronic device comprises a memory, a processor and a computer program stored in the memory and executable on the processor, wherein when the processor executes the program, a lightweight facial expression recognition method based on linear self-attention as described in any one of the items is implemented.

[0048] A non-transitory computer-readable storage medium having a computer program stored thereon, wherein the computer program, when executed by a processor, implements a lightweight facial expression recognition method based on linear self-attention as described in any one of the items.

[0049] A computer program product comprises a computer program, wherein when the computer program is executed by a processor, the linear self-attention lightweight facial expression recognition method as described in any one of the items is implemented.

[0050] The advantages and beneficial effects of the present invention are as follows:

[0051] The method of the present invention first proposes a new facial expression recognition method, called a linear self-attention visual transformer. The network consists of three parts: a linear multi-head self-attention module, a local spatial attention module, and a channel attention module. The overall network is supervised by the discriminant loss Di-Loss. First, the linear multi-head self-attention module is used to fully capture the correlation between the various regions of the expression and learn the global characteristics of the expression. Secondly, the local spatial attention module allows the network to focus more on key information when processing facial expressions with interference factors, and uses channel attention to further increase channel information interaction. Finally, under the supervision of the discriminant loss, the inter-class distance is further expanded and the intra-class distance is minimized to improve the accuracy of expression recognition. The main advantages and beneficial effects are as follows:

[0052] 1. This paper designs a lightweight linear multi-head self-attention algorithm, which can effectively learn the correlation between different local patches. In addition, compared with other ViT-based expression recognition methods, the linear multi-head self-attention algorithm has lower resource consumption and computational complexity, enhances the robustness of features, and improves the performance of facial expression recognition.

[0053] 2. The local spatial attention module designed in the present invention increases the weight of the key areas of expression, thereby improving the network's ability to process detail features and enhancing the robustness of features to interference from expression changes.

[0054] 3. The present invention designs a discrimination loss function that effectively expands the inter-class distance of expressions and reduces the intra-class distance. The auxiliary network enhances the importance of fusion features and further improves the recognition accuracy. BRIEF DESCRIPTION OF THE DRAWINGS

[0055] Figure 1 It is a schematic diagram of the overall network model structure of a preferred embodiment provided by the present invention. DETAILED DESCRIPTION

[0056] The following will describe the technical solutions in the embodiments of the present invention in detail in conjunction with the accompanying drawings in the embodiments of the present invention. The described embodiments are only part of the embodiments of the present invention.

[0057] The technical solution of the present invention to solve the above technical problems is:

[0058] As attached Figure 1 As shown, a lightweight facial expression recognition method based on linear self-attention includes the following steps:

[0059] 1. As attached Figure 1 As shown in the figure, a facial expression image is input into the expression recognition network, and the initial features after 3×3 convolution are fed into the global linear self-attention module and the local spatial attention module respectively, which specifically includes the following steps:

[0060] A1. Detect facial key points of the facial expression image through the face detection and alignment network MTCNN, align the facial expression image, and crop it into an input image I of 224×224 size;

[0061] A2. Input image I into two 3×3 convolutions to extract the primary features of the image, represented by x. Where C, H, and W represent the number of image channels, height, and width, respectively. Feed x to the global linear self-attention module and the local spatial attention module respectively;

[0062] 2. As attached Figure 1 As shown in the figure, the global linear self-attention module mainly obtains global expression features by capturing the long-distance dependencies of patches in different expression areas. At the same time, it reduces the quadratic computational complexity of the classic multi-head self-attention mechanism to linear, effectively achieving a lightweight effect while ensuring the accuracy of expression recognition. Specifically, it includes the following steps:

[0063] B1. Flatten the input low-level feature x into a feature sequence, and generate three different feature token matrices Q, K, and V through linear mapping.

[0064] B2. For the token matrix Q, use the weights The linear layer of Q maps each token L in Q to a scalar. This linear projection is an inner product operation that calculates the distance between token L and x, resulting in a k-dimensional vector. The Softmax function is then applied to this k-dimensional vector to produce a context score

[0065] B3, context score c s is used to calculate the attention matrix Z. Specifically, the context score c s The input token matrix K is weighted to produce a contextual attention matrix Z that encodes global information.

[0066] B4. Finally, in order to obtain the final global features, the token matrix V after deep convolution and the context attention matrix Z are Hadamard-multiplied to obtain the final global expression encoding features. Since the designed linear self-attention operations are all based on element-by-element multiplication or addition, compared to the traditional multi-head self-attention which requires expensive batch-level matrix multiplication and other operations, we reduce the overall computational complexity from quadratic to linear. The linear self-attention calculation formula is as follows:

[0067]

[0068] Among them, λ represents the scaling factor, ⊙ represents the Hadamard product, and DWConv represents the depthwise convolution operation.

[0069] 3. As attached Figure 1 As shown in the figure, the local spatial attention module mainly performs spatial attention weighting through the MobileNext hourglass block and pooling operation. The hourglass block can perform identity mapping and spatial transformation at higher dimensions of the network, which can effectively eliminate feature information loss and gradient confusion and has a lighter effect. The features learned by the hourglass block are further used through maximum pooling and average pooling operations to obtain the attention scores of important expression areas, and finally the output attention feature map is obtained through feature weighting, which specifically includes the following steps:

[0070] C1. Set is the input tensor after the initial convolution, and x is passed through the hourglass module for new spatial feature interaction. Specifically, x is subjected to deep convolution to reduce and expand the channel dimension, and then subjected to deep convolution again to encode richer spatial feature information to obtain a new spatial feature map The formula for the hourglass block can be written as follows:

[0071] G=φ e (DW(φ r (DW(X))))+X (2)

[0072] Among them, φ e and φ r They represent two point-wise convolutions for channel expansion and channel reduction, respectively, and DW represents depth-wise convolution.

[0073] C2. Aggregate the channel information of the feature map by using two pooling operations to generate two 2D feature maps: and To represent the average pooling features and the maximum pooling features in the channel. Then, the Sigmod activation function is used to generate our spatial attention map In short, the spatial attention formula is calculated as:

[0074]

[0075] Among them, σ represents the sigmoid activation function.

[0076] 4. As attached Figure 1 As shown in FIG. 1 , the global and local expression features learned by the multi-stage feature extraction operation are fused, and finally the channel features are further weighted by the channel attention module and then input into the classifier for expression classification, which specifically includes the following steps:

[0077] D1. For the global and local expression features M(G) and N(G) learned through multi-stage feature extraction operations, the features are fused through the Concat operation. The formula is:

[0078] F=Concat([M(G);N(G)]) (4)

[0079] D2. After the fused feature F is subjected to channel attention weighted interactive channel information by ECANet, it is input into the classifier for expression classification. The effective channel attention network ECANet is a neural network that can generate channel attention through fast one-dimensional convolution. The size of its convolution kernel can be adaptively determined by the nonlinear mapping of the channel dimension. Specifically, the adaptive kernel size calculation formula is as follows:

[0080]

[0081] where |t| odd represents the odd number closest to t, C represents the number of input feature channels, γ represents the feature weight, and b represents the bias.

[0082] 5. As attached Figure 1 As shown in the figure, the entire network is supervised by the identification loss Di-Loss, which effectively expands the distance between classes and reduces the distance within classes. Specifically, it includes the following steps:

[0083] E1. Real-world facial expression recognition applications require a large number of expression images acquired in unconstrained environments. Therefore, for the wild expression recognition task, its images show significant intra-class variations and inter-class similarities, and feature discrimination is a key supervision step. However, the classic classification loss function, the softmax function, does not explicitly optimize feature embedding to improve the similarity of intra-class samples and the diversity of inter-class samples, which leads to the performance gap of deep facial expression recognition. ArcFace loss was recently proposed to significantly expand the inter-class distance and reduce the intra-class distance, achieving competitive results. Specifically, the formula of the most widely used classification loss function, the softmax loss function, is as follows:

[0084]

[0085] in represents the characteristics of the i-th sample belonging to the j-th class, represents the weight matrix of the i-th sample, represents the weight matrix of the jth class, b yi , b j represents bias, and N represents the number of classifications;

[0086] E2, Additive Angular Margin Loss ArcFace loss sets the bias of Softmax loss to 0, where θ j is the angle between the weight matrix and the feature. Then, the weights ∥∥W are transformed by L2 j ∥∥Normalize, and convert ∥x i ∥Normalized to the scaling factor s. At the same time, ArcFace adds an additional angular margin penalty between the intra-class and inter-class distances to simultaneously enhance intra-class compactness and inter-class differences. The ArcFace loss function formula is as follows:

[0087]

[0088] E3. Considering that the inter-class distance in ArcFace loss is not particularly discriminative when the cluster center is close to the origin, we designed a discriminative loss function Di-Loss, which adds a constraint term based on the L2 norm to expand the distance between expressions. Specifically, the Di Loss formula can be expressed as follows:

[0089]

[0090] Among them, m is the total number of expression samples in the batch, N is the total number of expression categories, and c is j is the center of the jth expression sample, p j is the proportion of the j-th expression sample in batch m, and μ is the balance parameter.

[0091] The systems, devices, modules or units described in the above embodiments may be implemented by computer chips or entities, or by products with certain functions. A typical implementation device is a computer. Specifically, the computer may be, for example, a personal computer, a laptop computer, a cellular phone, a camera phone, a smart phone, a personal digital assistant, a media player, a navigation device, an email device, a game console, a tablet computer, a wearable device, or a combination of any of these devices.

[0092] Computer readable media include permanent and non-permanent, removable and non-removable media that can be implemented by any method or technology to store information. Information can be computer readable instructions, data structures, program modules or other data. Examples of computer storage media include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technology, compact disk read-only memory (CD-ROM), digital versatile disk (DVD) or other optical storage, magnetic cassettes, magnetic tape magnetic disk storage or other magnetic storage devices or any other non-transmission media that can be used to store information that can be accessed by a computing device. As defined herein, computer readable media does not include temporary computer readable media (transitory media), such as modulated data signals and carrier waves.

[0093] It should also be noted that the terms "include", "comprises" or any other variations thereof are intended to cover non-exclusive inclusion, so that a process, method, commodity or device including a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such process, method, commodity or device. In the absence of more restrictions, the elements defined by the sentence "comprises a ..." do not exclude the existence of other identical elements in the process, method, commodity or device including the elements.

[0094] The above embodiments should be understood to be only used to illustrate the present invention and not to limit the protection scope of the present invention. After reading the contents of the present invention, technicians can make various changes or modifications to the present invention, and these equivalent changes and modifications also fall within the scope defined by the claims of the present invention.

Claims

1. A linear self-attention lightweight facial expression recognition method, characterized in that: The following steps are involved: Step 1: Input the facial expression image into the expression recognition network, and feed the initial features after 3×3 convolution into the global linear self-attention module and the local spatial attention module respectively; Step 2: The global linear self-attention module obtains global expression features by capturing the long-range dependencies of patches in different expression regions; Step 3: The local spatial attention module performs spatial attention weighting through the MobileNext hourglass block and pooling operation. The hourglass block performs identity mapping and spatial transformation. The features learned by the hourglass block are respectively subjected to maximum pooling and average pooling operations to obtain the attention scores of important expression areas. Finally, the output attention feature map is obtained through feature weighting. Step 4: The global and local expression features learned by the multi-stage feature extraction operation are fused, and finally the channel features are weighted by the channel attention module and input into the classifier for expression classification; The step 2 specifically includes: B1. Flatten the input low-level feature x into a feature sequence and generate three different feature token matrices Q, K, and V through linear mapping; B2. For the token matrix Q, use the weights The linear layer of Q maps each token L in Q to a scalar; this linear projection is an inner product operation that calculates the distance between token L and x, resulting in a k-dimensional vector; the Softmax function is then applied to this k-dimensional vector to produce a context score B3. The context score c s Weighting the input token matrix K produces a contextual attention matrix Z that encodes global information; B4. Finally, in order to obtain the final global features, the token matrix V after deep convolution and the context attention matrix Z are Hadamard-multiplied to obtain the final global expression encoding features. The linear self-attention calculation formula is as follows: Among them, λ represents the scaling factor, ⊙ represents the Hadamard product, and DWConv represents the deep convolution operation; The calculation formula of the convolution kernel size of D2 is as follows: where |t| odd represents the odd number closest to t, C represents the number of input feature channels, γ represents the feature weight, and b represents the bias; The step 4 also includes: the entire network is supervised by the identification loss Di-Loss to expand the distance between classes and reduce the distance within classes, which specifically includes the following steps: E1. The softmax loss function formula for classification loss function is as follows: in represents the characteristics of the i-th sample belonging to the j-th class, represents the weight matrix of the i-th sample, represents the weight matrix of the jth class, b j represents bias, N represents the number of classifications; E2, Additive Angular Margin Loss ArcFace sets the bias of Softmax loss to 0, where θ j is the angle between the weight matrix and the feature; then, the weights ∥∥W are transformed by L2 j ∥∥Normalize, and convert ∥x i ∥Normalized to the scaling factor s; At the same time, ArcFace adds an additional corner margin penalty between the intra-class and inter-class distances to simultaneously enhance the intra-class compactness and inter-class differences; The ArcFace loss function formula is as follows: E3. The identification loss function Di-Loss is designed, which adds a constraint based on the L2 norm to expand the distance between expressions; the DiLoss formula is as follows: Among them, m is the total number of expression samples in each batch, N is the total number of expression categories, and c is j is the center of the jth expression sample, p j is the proportion of the j-th expression sample in batch m, and μ is the balance parameter.

2. A linear self-attention lightweight facial expression recognition method according to claim 1, characterized in that: The step 1 specifically comprises the following steps: A1. Detect facial key points of the facial expression image through the face detection and alignment network MTCNN, align the facial expression image, and crop it into an input image I of 224×224 size; A2. Input image I into two 3×3 convolutions to extract the primary features of the image, represented by x. Feed x to the global linear self-attention module and the local spatial attention module respectively, where C, H, and W represent the number of image channels, height, and width, respectively.

3. An electronic device, characterized in that: The invention comprises a memory, a processor and a computer program stored in the memory and executable on the processor, wherein when the processor executes the program, the linear self-attention lightweight facial expression recognition method as claimed in any one of claims 1 to 2 is implemented.

4. A non-transitory computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the lightweight facial expression recognition method based on linear self-attention as described in any one of claims 1 to 2 is implemented.

5. A computer program product, comprising a computer program, characterized in that When the computer program is executed by a processor, the lightweight facial expression recognition method based on linear self-attention as described in any one of claims 1 to 2 is implemented.

Citation Information

Patent Citations

  • Face key point detection method based on attention guidance lightweight network

    CN115966004A

  • Feature-enhanced lightweight network FGNet facial expression recognition method

    CN116311414A