A Facial Expression Recognition Method Based on Multi-Granularity Perception and Label Distribution Learning

Through the multi-grained perception and label distribution learning methods, the inconsistency and ambiguity of expression image labels in the natural environment are solved. Through global perceptual attention and gradual training, different granularity features are fused, and the accuracy and robustness of facial expression recognition are improved.

CN118898864BActive Publication Date: 2025-07-11CHENGDU SHUSHENG LANGLANG TECHNOLOGY CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202410928567.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-07-11
Publication Date
2025-07-11
Estimated Expiration
2044-07-11

AI Technical Summary

Technical Problem

The existing facial expression recognition technology has problems with inconsistent and ambiguity of expression image labels in natural environments, resulting in limited learning ability of the model and difficulty in accurately classifying complex emotions. The existing methods ignore the complementarity and global context information of different granularity features.

Method used

The multi-grained sensing and label distribution learning method is used to segment the images into different granularity levels, and features are extracted through the global perception attention module, and progressive training and label distribution learning are introduced, which integrates the characteristics of different granularity and global context information to construct multi-grained local features and label distribution for supervision.

Benefits of technology

It improves the accuracy and robustness of expression recognition, can better handle complex emotions and uncertainties, enhances the model's ability to recognize subtle expression differences, and reduces the impact of fuzzy labels.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN118898864B_ABST
    Figure CN118898864B_ABST
Patent Text Reader

Abstract

The present invention claims protection for a face expression recognition method based on multi-granularity perception and label distribution learning (GPLDL), aiming to design a face expression recognition model for solving the uncertainty problem of expressions in actual scenarios, belonging to the field of pattern recognition. The method includes the following steps: First, design a multi-granularity hierarchical perception fusion module, which effectively combines low-level detailed features and high-level semantic information, enhancing the model's ability to distinguish subtle expression differences. Second, through progressive training, learn multi-granularity features from different granularity levels of expression images, and maintain the structural integrity of facial features as much as possible. In addition, we also design a global perception attention module to capture the global context information of the face. Finally, design a label distribution learning module to construct a more comprehensive emotion distribution, effectively alleviating the negative impact of ambiguous expressions on model learning. And introduce label distribution loss to further improve the accuracy of expression recognition.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of computer vision, and particularly relates to a face expression recognition method based on multi-granularity perception and label distribution learning. Background Technique

[0002] Facial expressions play a crucial role in interpersonal communication. By controlling facial expressions, people can deeply enrich their interactions with each other. Currently, facial expression recognition technology has significant application value in the field of emotional human-computer interaction such as medical robots, safe driving, criminal psychological analysis, and distance education, and has become an important research direction in the field of computer vision.

[0003] In recent years, with the development of deep learning, methods based on convolutional neural networks (CNNs) have performed excellently in facial expression recognition (FER) and have occupied a dominant position in FER research. The excellent performance of these deep networks mainly benefits from the support of large-scale FER datasets. Usually, a single facial image in these datasets is annotated as one of several basic expressions to train the FER model. However, due to the different definitions of the same expression image category by people with different backgrounds, combined with the uncertainty brought by the ambiguity of wild facial images themselves, it often leads to inconsistent and inaccurate expression image labels, which seriously affects the learning ability of the FER model.

[0004] In addition, emotions in the real world are not just single emotions in a controlled environment, but exist in the form of a combination of basic emotions and complex emotions. Therefore, there are often some similarities between different expressions in a natural environment. People's emotional perceptions of the same image may be different, especially for images with high uncertainty, and it is difficult to accurately classify them. Moreover, there are quite a lot of ambiguous expression images in large-scale expression datasets, which makes it difficult for the model to learn the emotional characteristics of specific categories from them, thus affecting the performance of the deep network. Therefore, compared with a single label, the emotional distribution can better represent the multiple emotions contained in a single image because it takes into account the combination of multiple emotions, thus reducing the impact of ambiguity. However, most current expression datasets only provide a single label for expression images, and it is quite difficult to manually annotate the label distribution on a large-scale dataset. This results in insufficient supervision during training, making the performance of many FER models reach a bottleneck.

[0005] On the other hand, there is a common problem of large inter-class similarity in facial expression images, which also brings uncertainty to facial expression images. Due to the subtle differences between some categories, it is easy to cause confusion in the prediction of different categories of facial expressions, hindering the network from learning robust emotional features. Due to the visual similarity of these facial expression images, different categories can often only be distinguished by fine-grained local differences. Existing work enhances the representation ability of features through attention-based methods, extracts local features with key discriminability to distinguish similar facial expressions. However, these methods only focus on the local features of facial expressions and ignore the importance of the global context information of facial expressions. Therefore, some methods propose to fuse the local features and global features of facial expressions to further enhance the representation ability of facial expression features. In addition, most existing methods for extracting local features divide the image into specific four parts and extract local features from fixed and single-granularity-level parts. But most methods focus on using deep features containing high-level semantic information for recognition, lacking effective integration of information at different granularities and levels, thus hindering further improvement. Considering the uniqueness of the FER task, it is found that the differences in subtle parts such as complex textures can also provide key discriminative information for FER. Moreover, features at different levels contain information at different granularities. By using the fusion learning of low-level detailed information and high-level abstract semantic information, the network can better locate the multi-granularity discriminative regions of facial expressions and enhance the representation ability of the model. Therefore, making full use of the implicit complementary advantages of different-level features and different granularity levels can effectively improve the performance of FER. To solve the above problems, the present invention proposes a face expression recognition method based on multi-granularity perception and label distribution learning.

[0006] Application Publication No. CN117746484A, a micro-expression recognition method and system with enhanced multi-granularity cognition of facial features, belongs to the field of affective computing technology. The method includes data acquisition, multi-dimensional preprocessing, multi-granularity global feature extraction, multi-granularity local feature extraction, multi-granularity cognition enhancement, and feature classification. A multi-dimensional preprocessing method is used for data preprocessing. Multi-granularity global / local feature extraction methods are used to obtain coarse / fine-grained feature vectors respectively, and transfer learning is introduced when extracting multiple local features. In the multi-granularity cognition enhancement process, a channel attention module is used to extract the global feature vector of the face, and then it is weighted to the local feature vector, aiming to use the coarse-grained features of the facial global to provide guidance for the fine-grained features of the local features; the obtained final feature vector is used as an input component of the classification model for model training to obtain the classification model. The system is implemented based on the above method and can monitor and feedback the changes in students' classroom emotions in real time.

[0007] This patent obtains features of different granularities through the extraction methods of global features and local features. Most methods for extracting local features only divide an image into four regions at one block size level, extract the local features of the four regions, and then splice their feature vectors. This results in learning discriminative features only from one granularity level when extracting local features. However, the discriminative regions of wild expressions usually have different sizes. Therefore, the local features learned from only one granularity level are weakly robust. From the perspective of different granularities, we divide the input image into three different sizes of blocks, and adopt a progressive training method. The fine-grained features extracted by the shallow network are used as the guidance for the coarse-grained features extracted by the deep network, and the expression features of different granularities are fused to obtain multi-granularity local discriminative features to address the problem of misrecognition of expressions caused by small inter-class differences in expression images. Moreover, there are often problems of expression uncertainty in micro-expressions and macro-expressions, and this is even more obvious in micro-expressions. Existing datasets often use single labels as annotations for images, while expressions in real-world scenarios are ambiguous and usually appear in the form of a combination of multiple expressions. Therefore, using only single labels as supervision is insufficient. Thus, we propose a multi-granularity perception label distribution construction module. By utilizing the characteristics of the constructed label distribution with different emotional distributions, we fully supervise the learning of the network, enabling the model to learn more robust emotional features. Summary of the Invention

[0008] The present invention aims to solve the above problems of the prior art. A face expression recognition method based on multi-granularity perception and label distribution learning is proposed. The technical solution of the present invention is as follows:

[0009] A face expression recognition method based on multi-granularity perception and label distribution learning, which includes the following steps:

[0010] Step 1: Before inputting each image into the network for feature extraction, divide it into images of three different granularity levels, and randomly shuffle the image blocks;

[0011] Step 2: Input the three different granularity expression images after data augmentation into the last three layers of the basic network. Follow the principle that the coarse-grained input is fed into the deep network to extract coarse-grained expression features, the fine-grained input is fed into the shallow network to extract fine-grained expression features, and at the same time, the coarse-grained expression features are input into the global perception attention module GPAM to extract global expression features;

[0012] Step 3: Adopt a progressive training method for the multi-granularity features extracted at different granularity stages. First, input the fine-grained images into the lower-layer network for training through the classification loss, and then gradually input the larger granularity images into the next stage of the basic network until the deepest layer network; the whole process is trained three times, and finally, the different granularity features obtained by progressive training are spliced to obtain multi-granularity local features;

[0013] Step 4: Fuse the fused multi-granularity local expression features and the global perception expression features, and finally obtain the final prediction distribution through a fully connected layer and normalization processing; calculate the granularity similarity between the different granularity features obtained by the previous progressive training and the multi-granularity fusion features respectively, and construct a label distribution based on this; finally, use the constructed label distribution as a supervision signal, design a distribution loss, and train the overall network; integrate all the output predictions of the network as the final prediction for facial expression classification.

[0014] Further, before inputting each image into the network for feature extraction in Step 1, the image is segmented into 3 different granularity levels of images, and the image patches are randomly shuffled. The specific steps are as follows:

[0015] A1: For the facial expression image, detect the facial key points through the face detection and alignment network MTCNN. The face detection and alignment network MTCNN is a multi-task convolutional neural network composed of three sub-networks (P-Net, R-Net, O-Net), and each sub-network undertakes different tasks. P-Net is responsible for generating candidate windows, R-Net is responsible for filtering out incorrect candidate windows, and O-Net is responsible for facial key point localization and fine adjustment of the face bounding box. These three sub-networks cooperate together to achieve efficient and accurate face detection. It is usually used to preprocess the images in the facial dataset, that is, face detection and alignment operations), and align the facial expression image, and crop it into an input image I with a size of 224×224.

[0016] A2: Input the image into a jigsaw generator. The granularity size of each image is n = 1, 2, 4 respectively, that is, the image is segmented into different granularity images with patch sizes of 1×1, 2×2, and 4×4 respectively. Shuffle each patch and reconstruct the image to obtain 3 enhanced facial expression images with a size of 224×224 and different granularities. .

[0017] Further, Step 2 specifically includes:

[0018] B1: Obtain 3 enhanced images with different granularity levels, and use the images as the inputs for the last three stages of the backbone network respectively; represent the backbone network feature extractor as F, which contains L stages; the intermediate feature map output by the feature extractor F is represented as , where , respectively represent the height, width, and number of channels of the output feature map in the l-th stage;

[0019] B2: Input into the convolutional block to obtain the feature representation:

[0020]

[0021] Among them include a 1x1 and a 3x3 convolutional layer; then, through pooling operation on the above features, a feature vector is obtained and denoted as ;

[0022] B3. Finally, a global perception attention module is introduced to enhance the final multi-granularity fusion features; a dual-branch attention structure is used, namely the channel attention branch and the spatial attention branch; the feature vector output by the feature extractor F after StageL is denoted as ; First, is respectively calculated on two parallel branches to obtain the channel attention map denoted as , and the spatial attention map denoted as ; Among them, the channel attention consists of global average pooling and two fully connected layers; the spatial attention consists of 2 1×1 depthwise separable convolutions and 2 3×3 dilated convolutions. The calculation process of the entire global perception attention is as follows:

[0023]

[0024]

[0025]

[0026] Among them is the sigmoid function; finally, the global feature map can be calculated as:

[0027]

[0028] Global average pooling is used to add up all the pixel values of the feature map and calculate the average to obtain a value, that is, this value is used to represent the corresponding feature map.

[0029] Furthermore, step 3 specifically includes the following steps:

[0030] C1. Input the expression feature vectors of different granularities into a classifier composed of two fully connected layers with BatchNorm and ELU activation functions respectively to generate prediction probabilities ;

[0031] C2. For the training of the output of each stage, the cross-entropy between the true label y and the predicted probability distribution is used as the classification loss, and the calculation formula is as follows:

[0032]

[0033] Among them, N represents the number of training images, and C represents the number of categories. represents the label that the i-th image belongs to the k-th category. is the probability that the l-th stage prediction belongs to the k-th category. Here, the probability refers to how much the probability of belonging to the k-th category is; the final total cross-entropy loss is calculated as:

[0034]

[0035] C3. To further utilize feature fusion, the intermediate layer feature vectors of different stages obtained through progressive training are connected to obtain a multi-granularity fusion vector, and the formula is as follows:

[0036] .

[0037] Furthermore, step 4 specifically includes the following steps:

[0038] D1. After obtaining the enhanced global feature and the multi-granularity fusion vector the final multi-granularity fusion feature is calculated as:

[0039]

[0040] D2. After inputting the multi-granularity fusion feature into a fully connected layer and a normalization layer softmax, the final prediction is ;

[0041] D3. After obtaining the final output probability distribution, it is supervised by the original label in a similar way, and the classification loss is defined as follows:

[0042]

[0043] Among them, is the probability that the final stage prediction of the i-th sample belongs to the k-th category.

[0044] Furthermore, step 4 constructs a label distribution with rich information for each expression image as supervision, specifically including the following steps:

[0045] E1. Calculate the cosine similarity between the multi-granularity fusion feature and each as where , and the specific formula is as follows:

[0046]

[0047] where is and Dot product, where a is the index;

[0048] E2. Then, for each corresponding cosine similarity After that, perform normalization as follows:

[0049]

[0050] E3. According to the average cosine similarity and the probability distribution at each stage , the final label distribution is constructed by the following formula :

[0051]

[0052] where . logit probability The soft probability distribution obtained by passing through a softmax layer.

[0053] Furthermore, in step 4, the final predicted distribution of the model and the constructed distribution are supervised and trained through a designed distribution loss, which specifically includes the following steps:

[0054] F1. After obtaining the constructed label distribution, perform label distribution training on the finally output multi-granularity fusion features, aiming to minimize the difference between the constructed label distribution and the predicted probability distribution. The distribution loss is obtained by the following formula:

[0055]

[0056] where i is the index of the samples in a mini-batch.

[0057] Furthermore, perform an addition operation on the losses to obtain the final total loss , specifically including:

[0058] G1. Through the joint action of each module, the total loss function of the entire framework can be expressed as:

[0059]

[0060] where and are weighted ramp functions that change with the number of epochs, and the calculation formulas are as follows:

[0061]

[0062] where Let $\alpha$ denote the balance parameter, $\beta$ denote the current epoch index, and $E$ denote the threshold of epochs.

[0063] An electronic device includes a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the program, it implements the facial expression recognition method of multi-granularity perception and label distribution learning as described in any one of the preceding claims.

[0064] A non-transitory computer-readable storage medium stores a computer program. When the computer program is executed by a processor, it implements the facial expression recognition method of multi-granularity perception and label distribution learning as described in any one of the preceding claims.

[0065] The advantages and beneficial effects of the present invention are as follows:

[0066] The present invention mainly aims at the challenges brought by the label inconsistency of the expression dataset and the inherent ambiguity of expressions in the real scenario. By using multi-granularity hierarchical feature fusion and label distribution learning, a facial expression recognition method of multi-granularity perception and label distribution learning is designed. The advantage of the present invention is that through a hierarchical fusion strategy, low-level detailed features and high-level semantic information are integrated together, enhancing the model's ability to recognize and distinguish subtle expression differences. The proposed global perception attention module further captures comprehensive context information, improving the robustness of feature representation. In addition, a dynamic label distribution learning module is introduced to generate a detailed emotion distribution for each sample, effectively reducing the impact of ambiguous and noisy labels and improving the accuracy of expression recognition in the real scenario.

[0067] 1. Most existing facial expression methods rely on extracting information from high-level features or enhancing features extracted from low-level features using an attention mechanism, ignoring the complementarity of features at different granularity levels. Neurons in the deeper layers of the network contain rich semantic information, but as the depth increases, it inevitably leads to the loss of low-level information (such as texture, edge connection, etc.). From another perspective, existing FER methods extract local features from predefined image patches. A common practice is to evenly divide an image into 4 image patches as the input of the network, and then extract local features. However, only one granularity level of local features can be obtained from fixed image patches, and smaller granularity facial features are ignored. Different from this, the present invention designs a multi-granularity hierarchical feature fusion learning strategy, using the fusion learning of low-level detailed information and high-level abstract semantic information, enabling the network to better locate the multi-granularity discriminant regions of expression images. Then, we use the global perception attention module to efficiently extract the global context information in the expression image, further enhancing the multi-granularity fusion features.

[0068] 2. The expressions in most real scenarios are uncertain and appear in the form of a combination of multiple emotions. Therefore, misrecognition often occurs. The present invention designs a multi-granularity perception-based label distribution construction module and introduces label distribution learning to mitigate the negative impact of blurred expressions in the wild environment on model learning. Since features of different granularities contain different information, fine-grained features contain a lot of texture information and can effectively suppress the influence of external factors such as illumination and occlusion. The goal of the present invention is to simultaneously utilize different granularity information from the low-level and high-level of the network to reconstruct the target label distribution with high fidelity, enabling the network to learn knowledge from both fine-grained and coarse-grained features, and effectively avoiding the negative impact of expression uncertainty on label distribution construction and discriminative feature extraction. Finally, the constructed label distribution is used as the supervision of the network output, enabling the network to extract richer emotional features and make more accurate recognition of uncertain expressions in real scenarios. BRIEF DESCRIPTION OF THE DRAWINGS

[0069] Figure 1 It is a schematic diagram of the overall network model structure of the preferred embodiment provided by the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0070] Next, the technical solutions in the embodiments of the present invention will be clearly and detailedly described with reference to the accompanying drawings in the embodiments of the present invention. The described embodiments are only a part of the embodiments of the present invention.

[0071] The technical solution for the present invention to solve the above technical problems is:

[0072] As shown in the attached Figure 1 drawings, a face expression recognition method based on multi-granularity perception and label distribution learning includes the following steps:

[0073] 1. As shown in the attached Figure 1 drawings, before inputting each image into the network for feature extraction, it is divided into 3 types of images with different granularity levels, and the image patches are randomly shuffled. Specifically, it includes the following steps:

[0074] A1. For a face expression image, detect the face key points through the face detection and alignment network MTCNN, align the face expression image, and crop it into an input image with a size of 224×224.

[0075] A2. Input the image into a jigsaw generator, where the granularity size of each image is n = 1, 2, 4 respectively, that is, the image is divided into different granularity images with block sizes of 1×1, 2×2, and 4×4 respectively. Shuffle each block and reconstruct the image to obtain 3 enhanced expression images with a size of 224×224 and different granularities.

[0076] 2. As shown in the attached Figure 1As shown in the figure, the three differently-granularity facial expression images after data augmentation are input into the last three layers of the basic network. Coarse-granularity input is fed into the deep network to extract coarse-granularity facial expression features, and fine-granularity input is fed into the shallow network to extract fine-granularity facial expression features. At the same time, the coarse-granularity facial expression features are input into the Global Perception Attention Module (GPAM) to extract global facial expression features. The specific steps are as follows:

[0077] B1. Obtain the enhanced images at 3 granularity levels and use the images as the inputs for the last three stages of the backbone network respectively. Denote the backbone network feature extractor as F, which contains L stages. The intermediate feature maps output by the feature extractor F are denoted as where , represent the height, width, and number of channels of the output feature map at stage l respectively;

[0078] B2. Input into the convolutional block to obtain the feature representation:

[0079]

[0080] where contains a 1x1 and a 3x3 convolutional layer. Then, we perform a pooling operation on the above features to obtain the feature vector denoted as ;

[0081] B3. Finally, a Global Perception Attention Module is introduced to highlight the global features of the facial expressions and enhance the final multi-granularity fusion features. A dual-branch attention structure is used, namely the channel attention branch and the spatial attention branch. As described above, the feature vector output by the feature extractor F after StageL is denoted as . First, is calculated separately in two parallel branches to obtain the channel attention map denoted as , and the spatial attention map denoted as . Among them, the channel attention consists of global average pooling and two fully connected layers. The spatial attention consists of 2 1×1 depthwise separable convolutions and 2 3×3 dilated convolutions. The entire calculation process of the Global Perception Attention is as follows:

[0082]

[0083]

[0084]

[0085] where is the sigmoid function. Finally, the global feature map can be calculated as:

[0086]

[0087] 3. As attached Figure 1 As shown in the figure, a progressive training method is used for the multi-granularity features extracted at different granularity stages. First, the fine-grained image is input into the low-level network for training through classification loss, and then the larger granularity image is gradually input into the next stage of the basic network until the deepest network. The whole process is trained three times, and finally the different granularity features obtained by progressive training are spliced ​​to obtain multi-granularity local features, which specifically includes the following steps:

[0088] C1. Input the expression feature vectors of different granularities into a classifier consisting of two fully connected layers with BatchNorm and ELU activation functions to generate prediction probabilities. .

[0089] C2. For the training of the output of each stage, we use the cross entropy between the true label y and the predicted probability distribution as the classification loss, calculated as follows:

[0090]

[0091] Where N is the number of training images, C is the number of categories, Indicates that the i-th image belongs to the label of the k-th class, is the probability of belonging to the kth class predicted in the lth stage. It should be clear that due to progressive training, all parameters used in the current prediction will be optimized, which can help each stage in the model work together. The final total cross entropy loss can be calculated as:

[0092]

[0093] C3. In order to further utilize feature fusion, the intermediate layer feature vectors of different stages obtained by progressive training are connected To obtain the multi-granularity fusion vector, the formula is as follows:

[0094]

[0095] 4. As attached Figure 1 As shown in the figure, the fused multi-granularity local expression features are fused with the global perception expression features, and finally the final prediction distribution is obtained through a fully connected layer and normalization. The different granularity features obtained by the previous progressive training are respectively calculated for the granularity similarity with the multi-granularity fusion features to construct the label distribution. Finally, the constructed label distribution is used as a supervision signal, a distribution loss is designed, and the overall network is trained. Specifically, the following steps are included:

[0096] D1. Enhanced global features and the multi-granularity fusion vector After that, the final multi-granularity fusion feature can be calculated as follows:

[0097]

[0098] where is the total number of training samples, is the class center.

[0099] D2. After inputting the multi-granularity fusion feature into a fully connected layer and a normalization layer (softmax), the final prediction is ;

[0100] D3. After obtaining the probability distribution of the final output, similarly supervised by the original label, the classification loss is defined as follows:

[0101]

[0102] where is the probability that the final stage prediction of the i-th sample belongs to the k-th class;

[0103] 5. As shown in the appendix Figure 1 The network constructs a label distribution with rich information for each facial expression image as supervision, which specifically includes the following steps:

[0104] E1. Calculate the cosine similarity between the multi-granularity fusion feature and each as where The specific formula is as follows:

[0105]

[0106] where is and 's dot product, and a is the index;

[0107] E2. Then, after obtaining the cosine similarity corresponding to each , perform normalization as follows:

[0108]

[0109] E3. According to the average cosine similarity and the probability distribution of each stage , the final label distribution is constructed by the following formula:

[0110]

[0111] Among them 。

[0112] 6. As shown in the appendix Figure 1 The final predicted distribution of the model and the constructed distribution are supervised and trained through the designed distribution loss, and the total loss of the model is designed 。Specifically, it includes the following steps:

[0113] F1. After obtaining the constructed label distribution, perform label distribution training on the multi-granularity fusion features of the final output. The purpose is to minimize the difference between the constructed label distribution and the predicted probability distribution, and the L2 norm is one of the effective measurements. Therefore, the distribution loss can be obtained by the following formula:

[0114]

[0115] where i is the index of the samples in a mini-batch.

[0116] F2. Through the joint action of each module, the total loss function of the entire framework can be expressed as:

[0117]

[0118] Among them and are weighted ramp functions that change with the number of epochs, and the calculation formula is as follows:

[0119]

[0120] Among them represents the balance parameter, β represents the current epoch index, and E is the threshold of the epoch.

[0121] The systems, devices, modules, or units illustrated in the above embodiments can be specifically implemented by computer chips or entities, or by products with certain functions.

[0122] A computer-readable medium includes both permanent and non-permanent, removable and non-removable media and can implement information storage by any method or technology. The information can be computer-readable instructions, data structures, program modules, or other data. Examples of computer storage media include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technologies, compact disc read-only memory (CD-ROM), digital versatile disc (DVD) or other optical storage, magnetic cassettes, magnetic tape magnetic disk storage or other magnetic storage devices, or any other non-transitory medium that can be used to store information that can be accessed by a computing device. As defined herein, a computer-readable medium does not include transitory computer-readable media such as modulated data signals and carrier waves.

[0123] It should also be noted that the term "comprising", "including" or any other variant thereof is intended to cover non-exclusive inclusion, such that a process, method, article or apparatus comprising a series of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such process, method, article or apparatus. Without further limitation, an element defined by the statement "comprising an..." does not exclude the presence of additional identical elements in the process, method, article or apparatus comprising the element.

[0124] The above embodiments should be understood as being only for illustrative purposes of the present invention and not for limiting the scope of protection of the present invention. After reading the content described in the present invention, those skilled in the art can make various changes or modifications to the present invention, and these equivalent changes and modifications also fall within the scope defined by the claims of the present invention.

Claims

1. A face expression recognition method based on multi-granularity perception and label distribution learning, characterized in that It includes the following steps: Step 1: Before inputting each image into the network for feature extraction, divide it into images at three different granularity levels and randomly shuffle the image patches; Step 2: Input the three differently granularity expression images after data augmentation into the last three layers of the basic network. Follow the principle that the coarsely granularity image is input into the deep network to extract coarsely granularity expression features, the finely granularity image is input into the shallow network to extract finely granularity expression features, and at the same time, input the coarsely granularity expression features into the Global Perception Attention Module (GPAM) to extract global expression features; Step 3: Adopt a progressive training method for the multi-granularity features extracted at different granularity stages. First, input the finely granularity image into the lower layer network for training through classification loss, and then gradually input larger granularity images into the next stage of the basic network until the deepest layer network; the whole process is trained three times. Finally, splice the multi-granularity features obtained by progressive training to get multi-granularity local features; Step 4: Fuse the fused multi-granularity local expression features with the global perception expression features, and finally obtain the final prediction distribution through a fully connected layer and normalization processing; calculate the granularity similarity between the different granularity features obtained by the previous progressive training and the multi-granularity fusion features respectively to construct a label distribution; finally, use the constructed label distribution as a supervision signal to design a distribution loss to train the overall network; Integrate all the output predictions of the network as the final prediction for facial expression classification; The specific content of Step 2 includes: B1. Obtain three enhanced images at different granularity levels, and use the images as the inputs for the last three stages of the backbone network respectively; represent the backbone network feature extractor as F, which contains L stages; represent the intermediate feature maps output by the feature extractor F as where l ∈ [1, L], H l , W l , C l represent the height, width and number of channels of the output feature maps at the l-th stage respectively; B2. Input F l into convolutional block B l to obtain a feature representation: Among them, B(·) includes a 1x1 and a 3x3 convolutional layer; then, the above features are subjected to a pooling operation to obtain a feature vector represented as v l ; B3. Finally, a global perception attention module is introduced to enhance the final multi-granularity fusion features. A dual-branch attention structure is used, namely the channel attention branch and the spatial attention branch. The feature vector output by the feature extractor F after StageL is denoted as V L ; First, V L is respectively calculated in two parallel branches to obtain the channel attention map denoted as the spatial attention map denoted as where the channel attention consists of global average pooling and two fully-connected layers; the spatial attention consists of two 1×1 depthwise separable convolutions and two 3×3 dilated convolutions. The entire calculation process of the global perception attention is as follows: M(V l ) = σ(M c (V l ) + M s (V l )) (4) where σ is the sigmoid function; finally, the global feature map can be calculated as: GAP represents global average pooling, and its function is to add up all the pixel values of the feature map and take the average to get a value, that is, use this value to represent the corresponding feature map.

2. The facial expression recognition method based on multi-granularity perception and label distribution learning according to claim 1, characterized in that, Before inputting each image into the network for feature extraction in Step 1, divide it into images at three different granularity levels and randomly shuffle the image patches. The specific steps are as follows: A1: For the facial expression image, detect the facial key points through the face detection and alignment network MTCNN. The face detection and alignment network MTCNN is a multi-task convolutional neural network composed of three sub-networks, namely P-Net, R-Net, and O-Net. Each sub-network undertakes different tasks; P-Net is responsible for generating candidate windows, R-Net is responsible for filtering out incorrect candidate windows, and O-Net is responsible for facial key point localization and fine adjustment of the face box; these three sub-networks cooperate together to achieve efficient and accurate face detection, and use it to preprocess the images in the facial dataset, that is, perform face detection and alignment operations, and align and crop the facial expression image into an input image I with a size of 224×224; A2: Input the image I into a jigsaw puzzle generator, where the granularity size of each image is n = 1, 2, 4 respectively, that is, divide the image into differently granularity images with patch sizes of 1×1, 2×2, and 4×4 respectively, shuffle each patch, and reconstruct the image to obtain three differently granularity expression images I1, I2, and I3 with a size of 224×224 after enhancement.

3. The face expression recognition method with multi-granularity perception and label distribution learning according to claim 2, characterized in that, The specific content of Step 3 includes the following steps: C1. Input the expression feature vectors with different granularities into a classifier composed of two fully connected layers with BatchNorm and ELU activation functions respectively to generate the prediction probability y L-2 , y L-1 , y L ; C2. For the training of the output of each stage, the cross-entropy between the true label y and the predicted probability distribution is used as the classification loss, and the calculation formula is as follows: where N represents the number of training images, C represents the number of classes, and y i,k represents the label that the i-th image belongs to the k-th class, is the probability of belonging to the k-th class predicted in the l-th stage, where the probability here refers to how much the probability of belonging to the k-th class is; the final total cross-entropy loss is calculated as: C3. To further utilize feature fusion, the intermediate layer feature vectors v at different stages obtained through progressive training are concatenated l to obtain a multi-granularity fusion vector, and the formula is as follows: v concat = concat[v L-2 , v L-1 , v L (8).

4. A facial expression recognition method based on multi-granularity perception and label distribution learning according to claim 3, characterized in that The specific steps of step 4 are as follows: D1. After obtaining the enhanced global feature and the multi-granularity fusion vector v concat the final multi-granularity fusion feature is calculated as follows: D2. After inputting the multi-granularity fusion features into a fully connected layer and a normalization layer softmax, the final prediction is obtained as y concat ; D3. After obtaining the finally output probability distribution, similarly, it is supervised by the original label, and the classification loss is defined as follows: where, is the probability that the final-stage prediction of the i-th sample belongs to the k-th class.

5. A method for facial expression recognition based on multi-granularity perception and label distribution learning according to claim 4, characterized in that Step 4 constructs a label distribution with rich information for each facial expression image as supervision, which specifically includes the following steps: E1. Calculate the multi-granularity fusion feature V concat The cosine similarity with each {v L-2 , v L-1 , v L} is s a , where a ∈ [L - 2, L], and the specific formula is as follows: where <V concat , v a > is the dot product of V concat and v a , and a is the index; E2. Then, obtain each v a corresponding cosine similarity s a After that, perform normalization as follows: E3. According to the average cosine similarity s a and the probability distribution {y L-2 , y L-1 , y L} at each stage, the final label distribution is constructed by the following formula where \(i\in[L - 2,L]\), \(p\) l represents the soft probability distribution obtained by converting the logit probabilities \(\{y\) L-2 ,y L-1 ,y L \} through a softmax layer.

6. The face expression recognition method with multi-granularity perception and label distribution learning according to claim 5, characterized in that The specific steps of step 4 for supervising and training the final prediction distribution of the model and the constructed distribution through the designed distribution loss are as follows: F1. After obtaining the constructed label distribution, perform label distribution training on the multi-granularity fusion features of the final output, aiming to minimize the difference between the constructed label distribution and the predicted probability distribution. The distribution loss is obtained by the following formula: where i is the index of the samples in a mini-batch.

7. A method for facial expression recognition based on multi-granularity perception and label distribution learning according to claim 6, characterized in that Add the losses to obtain the final total loss L total , specifically including: G1. Through the combined action of each module, the total loss function of the entire framework can be expressed as: L total = (λ1L soft + λ2L cl ) + αL ce (15) where λ1 and λ2 are weighted ramp functions that change with the number of epochs, and the calculation formula is as follows: where α represents the balance parameter, β represents the current epoch index, and E is the threshold of the epoch.

8. An electronic device, characterized in that, It includes a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the program, it implements the multi-granularity perception and label distribution learning-based facial expression recognition method according to any one of claims 1 to 7.

9. A non-transitory computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, it implements the multi-granularity perception and label distribution learning-based facial expression recognition method according to any one of claims 1 to 7.

Citation Information

Patent Citations

  • Facial feature multi-granularity cognitive enhancement micro-expression recognition method and system

    CN117746484A

  • Image generation using one or more neural networks

    US20220012858A1