Facial expression recognition method and system based on noise robust graph model

By using a noise-rosive graph model in the facial expression recognition system, the problems of difficult samples and noise labels are solved, and the effect of improving recognition accuracy and model robustness is achieved.

CN120047982APending Publication Date: 2025-05-27HUAZHONG NORMAL UNIV
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510104449.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-23
Publication Date
2025-05-27

AI Technical Summary

Technical Problem

Existing facial expression recognition methods have problems of insufficient learning and overfitting when dealing with difficult samples and noisy labels, resulting in low recognition accuracy.

Method used

Using a method based on the noise-robust graph model, a system that can handle difficult samples and noise labels is constructed through feature extraction and pre-identification network model, an adjacency matrix initialization model, an adjacency matrix regularization model and a classification network model. The system is supervised and trained through the cross entropy loss function to improve the robustness and accuracy of the model.

Benefits of technology

It effectively improves the accuracy of facial expression recognition, enhances the model's learning ability on difficult samples, and suppresses the influence of noise labels, improving the robustness of the model.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120047982A_ABST
    Figure CN120047982A_ABST
Patent Text Reader

Abstract

The invention discloses a facial expression recognition method and system based on a noise robust graph model, and the method comprises the steps: inputting a to-be-recognized facial image into the noise robust graph model, and carrying out the recognition of a facial expression, and the noise robust graph model comprises a feature extraction and pre-recognition network model, which is used for extracting the depth feature of a sample image and an expression pre-recognition tag; the adjacency matrix initialization model is used for constructing an initialized adjacency matrix; the adjacency matrix regularization model is used for regularizing the initialized adjacency matrix and outputting a regularized adjacency matrix; the classification network model is used for outputting an expression identification tag; and calculating a cross entropy loss function according to the expression pre-identification tag, the initialized adjacency matrix, the regularized adjacency matrix and the expression identification tag, and performing supervision training on the noise robust graph model by using the cross entropy loss function. According to the method, the feature similarity and learning difficulty difference between samples are concerned at the same time, and the expression recognition problem in a natural scene containing difficult samples and noise labels can be solved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of computer vision, and more specifically, relates to a facial expression recognition method and system based on a noise robust graph model. Background Art

[0002] Facial expressions, as the outward manifestation of inner emotions, play a vital role in human communication and emotional communication. Facial expression recognition (FER) enables computers to interpret these subtle emotional cues and is therefore regarded as the cornerstone of human-computer interaction (HCI) for natural emotion perception. However, building an accurate and reliable FER system still faces many challenges. On the one hand, facial expressions in natural environments are often affected by head rotation and partial occlusion, resulting in a large number of difficult samples with high intra-class variance. On the other hand, due to the limitations of data collection and annotation methods, some facial expression samples will inevitably be mislabeled, causing the model to inadvertently learn unreliable features. Therefore, enhancing the model's ability to learn from difficult samples and suppressing the influence of noisy labels remain key issues to be addressed in the field of FER.

[0003] Graph convolutional networks (GCNs) provide a promising approach for representation learning that helps address the above challenges. Specifically, GCNs view samples as nodes in a graph and construct an adjacency matrix based on the labels and / or features of the samples. The adjacency matrix describes the topological relationships between nodes by assigning weights to its edges, thereby enhancing the learning of a given node with the assistance of neighboring nodes. This unique ability enables samples with large intra-class variance to be indirectly connected through intermediate nodes, which facilitates the model to learn from difficult samples. In addition, since the construction of the adjacency matrix usually relies on the similarity measure of features, GCNs show a certain degree of robustness to mislabeled samples. The principle of this robustness is that even if some samples are mislabeled, the weights of the wrong connections are often small, thereby hindering the propagation of unreliable features. Therefore, since GCN is not only good at learning from difficult samples but also robust to noisy labels to a certain extent, it shows good performance in the FER task in a natural environment.

[0004] Although GCN itself has the inherent ability to handle difficult samples and noisy labels to a certain extent, solving these two problems simultaneously in the FER task is still a key issue to be solved. In fact, existing work ignores the fact that there is a delicate balance between learning difficult samples and resisting noisy labels. When the model overemphasizes suppressing learning from noisy labels, it tends to misjudge difficult samples as mislabeled samples, thereby inhibiting the learning process on these samples. On the contrary, when the model overemphasizes learning from difficult samples, it may regard mislabeled samples as difficult samples, resulting in overfitting to noisy samples. Summary of the Invention

[0005] In view of the above defects or improvement requirements of the prior art, the present invention provides a facial expression recognition method and system based on a noise-robust graph model, aiming to solve the problems of insufficient learning of difficult samples and overfitting to noise samples in existing facial expression recognition methods, resulting in low accuracy.

[0006] To achieve the above object, according to one aspect of the present invention, there is provided a facial expression recognition method based on a noise-robust graph model, which inputs a facial image to be recognized into the noise-robust graph model to recognize facial expressions. The noise-robust graph model includes a feature extraction and pre-recognition network model, an adjacency matrix initialization model, an adjacency matrix regularization model, and a classification network model;

[0007] The feature extraction and pre-recognition network model is used to extract the depth features of the sample image and the expression pre-recognition label from the input sample image;

[0008] The adjacency matrix initialization model is used to reduce the dimension of the depth features of the input sample image and construct an initial adjacency matrix using the similarity between the reduced-depth features;

[0009] The adjacency matrix regularization model is used to calculate the L2 norm distance of the depth features of the sample image and constrain the initial adjacency matrix based on the L2 norm distance of the depth features of the sample images with the same pre-recognition label, and output a regularized adjacency matrix;

[0010] The classification network model is used to output an expression recognition label according to the depth features of the input sample image and the regularized adjacency matrix;

[0011] Calculate the cross-entropy loss function according to the expression pre-recognition label, the initial adjacency matrix, the regularized adjacency matrix, and the expression recognition label, and use the cross-entropy loss function to supervise and train the noise-robust graph model.

[0012] Further, the feature extraction and pre-recognition network model is a convolutional neural network, and the classification network model is a graph convolutional neural network.

[0013] Further, first optimize and train the feature extraction and pre-recognition network model separately, and then supervise and train the noise-robust graph model.

[0014] Further, use a deep neural network model pre-trained on a large-scale face dataset as the feature extraction and pre-recognition network model, and use sample images with expression labels to optimize and train the feature extraction and pre-recognition network model separately.

[0015] Further, the dimensionality reduction of the depth features of the input sample image includes the steps of:

[0016] Incorporate non-linearity using the cosine kernel function, and the calculation formula is:

[0017]

[0018] where k ij represents the cosine similarity between the i-th sample and the j-th sample, f i is the depth feature of the i-th sample, f j is the depth feature of the j-th sample, k(.) is the cosine kernel function, T represents the transpose operation of a matrix or vector, ‖.‖ represents the vector norm, ‖f i ‖ represents the norm of f i , ||f j || represents the norm of f j ;

[0019] Adopt the non-linear principal component analysis method for feature dimensionality reduction, and the calculation formula is:

[0020] K=(k ij )

[0021]

[0022] where 1 N is a vector of all 1s with dimension N×1, N is the total number of sample images in a batch, I is an identity matrix with dimension N×N, K is the kernel matrix composed of all k ij as elements, is the kernel matrix after centering operation, Λ is a diagonal matrix constructed by arranging the eigenvalues in descending order, V is the eigenvector matrix of the matrix , V m,ij represents the weight of the sample f j to f i in the m-th principal component of V, V m,ji represents the weight of the sample f i to f j in the m-th principal component of V, m is the dimensionality of the reduced features, f i * is the reduced feature of f i , f j * is the reduced feature of f j ;

[0023] Further, the calculation formula for constructing the initial adjacency matrix using the similarity between the reduced depth features is:

[0024]

[0025] Among them, is the feature similarity after dimensionality reduction between the i-th sample and the j-th sample, and S * is all The initialized adjacency matrix formed by the elements.

[0026] Furthermore, the step of constraining the initialized adjacency matrix based on the L2 norm distance of the deep features of the sample images with the same pre-identified labels and outputting the regularized adjacency matrix includes:

[0027] Construct connections using the expression pre-identification labels of the sample images, and the calculation formula is:

[0028] A IN = S * ⊙ A LCM

[0029]

[0030] A LCM ={A LCM (i, j)}

[0031] Among them, A IN is the initialized adjacency matrix of the samples with the same pre-identified labels retained in S * , S * is the initialized adjacency matrix, A LCM is the label consistency matrix representing whether the pre-identified labels are the same, A LCM (i, j) is the element of A LCM , indicating the consistency of the corresponding pre-identified labels of sample i and sample j, and ⊙ represents the Hadamard product, represents the expression pre-identification label of the i-th sample, represents the expression pre-identification label of the j-th sample;

[0032] Construct a regularized adjacency matrix using the L2 norm distance of the deep features of the sample images, and the calculation formula is:

[0033]

[0034] Among them, A REG is the regularized adjacency matrix, is the L2 norm vector corresponding to the deep features of N sample images, diag is the operation of matrixifying and assigning values on the diagonal, I is the all-1 matrix with dimensions N×N, and M represents the L2 norm distance matrix of the sample images.

[0035] Furthermore, the calculation formula for the cross-entropy loss function calculated according to the expression pre-identification label, the initialized adjacency matrix, the regularized adjacency matrix, and the expression recognition label is:

[0036]

[0037] Among them, is the cross-entropy loss function of the noise-robust graph model, is the pre-recognition loss, is the graph network classification loss, is the regularization loss, N is the total number of sample images in a batch, C represents the total number of sample expression categories, w c represents the category weight of the c-th category of the sample, y i,c represents the annotation of the i-th sample in the c-th category, p i,c represents the probability of the i-th sample in the c-th category output by the feature extraction and pre-recognition network model, p' i,c represents the probability of the i-th sample in the c-th category output by the classification network model, A IN is the initial adjacency matrix of samples with the same pre-recognition label, A REG is the regularized adjacency matrix, and F represents the Frobenius norm.

[0038] According to another aspect of the present invention, a facial expression recognition system based on a noise-robust graph model is provided, including a training module and a prediction module;

[0039] The noise-robust graph model includes a feature extraction and pre-recognition network model, an adjacency matrix initialization model, an adjacency matrix regularization model, and a classification network model;

[0040] The feature extraction and pre-recognition network model is used to extract the depth features and expression pre-recognition labels of the sample images from the input sample images;

[0041] The adjacency matrix initialization model is used to reduce the dimension of the depth features of the input sample images and construct an initial adjacency matrix using the similarity between the reduced-depth features;

[0042] The adjacency matrix regularization model is used to calculate the L2 norm distance of the sample images and constrain the initial adjacency matrix based on the L2 norm distance of the depth features of the sample images with the same pre-recognition label, and output a regularized adjacency matrix;

[0043] The classification network model is used to output an expression recognition label according to the depth features and the regularized adjacency matrix of the input sample images;

[0044] The training module is used to calculate the cross-entropy loss function according to the expression pre-recognition label, the initial adjacency matrix, the regularized adjacency matrix, and the expression recognition label, and use the cross-entropy loss function to perform supervised training on the noise-robust graph model;

[0045] The prediction module is used to input the facial image to be recognized into the trained noise-robust graph model to recognize facial expressions.

[0046] Furthermore, the feature extraction and pre-recognition network model is a convolutional neural network, and the classification network model is a graph convolutional neural network.

[0047] Generally speaking, compared with the prior art, the above technical solution conceived by the present invention can solve the problems of insufficient learning of difficult samples and overfitting to noise samples in the existing facial expression recognition methods, resulting in low accuracy, thereby improving the accuracy of facial expression recognition. Specifically, the adjacency matrix represents the differences in features and learning difficulties between difficult samples and noise-label samples, and the initialization of the adjacency matrix adopts the non-linear principal component analysis method to explore the relationship between samples on the low-dimensional manifold, so as to better cope with the challenge of large intra-class differences in natural scene expressions. For the problem of low accuracy caused by overfitting to noise samples, the adjacency matrix regularization designs a label consistency mask and an L2-norm distance regularization method. On the one hand, the label consistency mask can further reduce the intra-class distance in the feature space through the pre-recognition labels; on the other hand, the L2-norm distance regularization method can constrain the model to focus on the learning of difficult samples and ignore the influence of noise labels through the learning difficulty difference, so that the robustness and accuracy of the model are further improved. Description of the Drawings

[0048] Figure 1 is a schematic diagram of the working principle of the facial expression recognition method provided by an embodiment of the present invention;

[0049] Figure 2 is a schematic diagram of the L2-norm distance of facial expressions provided by an embodiment of the present invention;

[0050] Figure 3 is a schematic diagram of the network structure of the noise-robust graph model provided by an embodiment of the present invention. Detailed Embodiments

[0051] In order to make the objectives, technical solutions and advantages of the present invention clearer, the present invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention. In addition, the technical features involved in the various embodiments of the present invention described below can be combined with each other as long as they do not conflict with each other.

[0052] In the description of the embodiments of the present application, the terms "including" and "having" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, apparatus, product, or device that includes a series of steps or modules does not necessarily have to be limited to those steps or modules clearly listed, but may include other steps or modules that are not clearly listed or are inherent to these processes, methods, products, or devices.

[0053] In the embodiments of the present invention, the naming or numbering of steps does not mean that the steps in the method flow must be executed in the chronological / logical order indicated by the naming or numbering. The named or numbered process steps can be changed in the execution order according to the technical purpose to be achieved, as long as the same or similar technical effects can be achieved.

[0054] Referring to "embodiments" herein means that the specific features, structures, or characteristics described in connection with the embodiments may be included in at least one embodiment of the present invention. The phrase appearing in various positions in the specification does not necessarily refer to the same embodiment, nor is it an independent or alternative embodiment mutually exclusive with other embodiments. Those skilled in the art explicitly and implicitly understand that the embodiments described herein can be combined with other embodiments.

[0055] The present invention provides a facial expression recognition method and system based on a noise-robust graph model, which will be described separately below.

[0056] The embodiments of the present invention provide a facial expression recognition method based on a noise-robust graph model. The facial image to be recognized is input into the noise-robust graph model to recognize the facial expression. The noise-robust graph model includes a feature extraction and pre-recognition network model, an adjacency matrix initialization model, an adjacency matrix regularization model, and a classification network model;

[0057] The feature extraction and pre-recognition network model is used to extract the depth features and expression pre-recognition labels of the sample image from the input sample image;

[0058] The adjacency matrix initialization model is used to reduce the dimension of the depth features of the input sample image and construct an initial adjacency matrix using the similarity between the reduced-depth features;

[0059] The adjacency matrix regularization model is used to calculate the L2 norm distance of the depth features of the sample image and constrain the initial adjacency matrix based on the L2 norm distance of the depth features of the sample images with the same pre-recognition label, and output a regularized adjacency matrix;

[0060] The classification network model is used to output an expression recognition label according to the depth features of the input sample image and the regularized adjacency matrix;

[0061] Calculate the cross - entropy loss function based on the expression pre - recognition label, initialize the adjacency matrix, regularize the adjacency matrix, and the expression recognition label, and use the cross - entropy loss function to perform supervised training on the noise - robust graph model.

[0062] Further, before extracting the depth features of the sample image and the expression pre - recognition label from the input sample image, pre - process the dataset samples.

[0063] The embodiments of the present invention adopt a series of refined operations for data pre - processing, including using face detection technology for face localization, cropping the located face region to remove background information, and then adjusting the size of the cropped face image to meet the model input requirements. To increase the diversity and complexity of the data, operations such as flipping at different angles and random erasing can also be performed to simulate expressions in different poses and occlusion situations. These pre - processing steps not only provide high - quality input data for the subsequent model but also enhance the generalization ability and robustness of the model by improving the diversity and complexity of the data.

[0064] Further, the feature extraction and pre - recognition network model is a convolutional neural network (CNN).

[0065] During training, the feature extraction and pre - recognition network model can be optimized and trained separately first, and then the noise - robust graph model can be subjected to supervised training.

[0066] The working principle of the feature extraction and pre - recognition network model is specifically described below.

[0067] Further, use the deep neural network model pre - trained on a large - scale face dataset as the feature extraction and pre - recognition network model, and use the sample images with expression labels to perform separate optimization training on the feature extraction and pre - recognition network model.

[0068] In other words, import the deep neural network model pre - trained on a large - scale face dataset, and use the samples with expression labels to fine - tune this pre - trained face model, and finally use the fine - tuned model to extract depth features.

[0069] Specifically, the pre - trained model reads the network parameters in the pre - trained face model Ms - Celeb - 1M through the ResNet18 network model, directly inputs all global face sample images into the ResNet18 network, and outputs a 512 - dimensional vector before the last fully - connected layer in the ResNet18 network; then, use the labeled samples to fine - tune the pre - trained face model Ms - Celeb - 1M to obtain a deep model, and finally use the obtained deep model to extract the depth embedding features of all global face sample images. This feature is the depth feature vector of the samples in this embodiment.

[0070] The method for obtaining the pre-recognition label of the expression is as follows: When fine-tuning the pre-trained face model Ms-Celeb-1M based on the labeled samples, the predicted label of the model for the samples is obtained, and this label is the pre-recognition label of the samples in this embodiment. The supervised fine-tuning is supervised using the following loss:

[0071]

[0072] where N represents the total number of samples in the batch, C represents the total number of sample categories, w c represents the weight of the c-th category of the sample, y i,c represents the annotation of the i-th sample in the c-th category, p i,c represents the probability of the i-th sample in the c-th category output by the feature extraction and pre-recognition network model. The fine-tuned feature extraction and pre-recognition network model can be used to extract deep features on the one hand and provide pre-recognition labels of expressions for subsequent steps on the other hand.

[0073] The working principle of the adjacency matrix initialization model is specifically described below.

[0074] The adjacency matrix initialization model is used to reduce the dimension of the deep features of the input sample image and construct an initial adjacency matrix using the similarity between the reduced deep features.

[0075] To obtain a reliable topological relationship between nodes, the KPCA method can be used for dimensionality reduction:

[0076]

[0077]

[0078] where 1 N is a vector of all 1s with dimension N×1, N is the total number of sample images in the batch, I is a matrix of all 1s with dimension N×N, K is a kernel matrix composed of all k ij as elements, is the kernel matrix after the centering operation, Λ is a diagonal matrix constructed by arranging the eigenvalues in descending order, V is the eigenvector matrix of the matrix , V m,ij represents the weight of the sample f j in the m-th principal component of V with respect to f i , V m,ji represents the weight of the sample f i in the m-th principal component of V with respect to f j , m is the dimension of the reduced features, m is the dimension of the reduced features, f i * is f i the reduced features of fj * is f j The features after dimensionality reduction:

[0079]

[0080] where f i is the depth feature of the i-th sample, f j is the depth feature of the j-th sample, K(i, j) is the cosine kernel function, T represents the transpose operation of a matrix or vector, ‖.‖ represents the vector norm, and ‖f i ‖ represents the norm of f i and ||f j || represents the norm of f j ;

[0081] The calculation formula for constructing the initial adjacency matrix using the similarity between the depth features after dimensionality reduction is:

[0082]

[0083] where is the similarity of the features after dimensionality reduction between the i-th sample and the j-th sample, and S * is the initial adjacency matrix.

[0084] The working principle of the adjacency matrix regularization model is described in detail below.

[0085] Adjacency matrix regularization is used to constrain the adjacency matrix using the label consistency masking strategy (LCM) and the learning difficulty difference matrix, and output a regularized adjacency matrix to achieve learning of difficult samples while being robust to noisy labels. The label consistency masking strategy connection is used to maintain the connection between simple samples and difficult samples, and the connection strategy based on the L2 norm distance is used to suppress the learning of noisy labels.

[0086] The label consistency masking strategy is as follows:

[0087] Construct connections using the pre-identified sample labels:

[0088] A IN = S * ⊙ A LCM

[0089]

[0090] A LCM = {A LCM (i, j)}

[0091] where A IN is the initial adjacency matrix of the samples with the same pre-identified labels retained in S * and ALCM As the label consistency matrix representing whether the pre-identified labels are the same, A LCM (i, j) is an element of A LCM , representing the consistency of the pre-identified labels corresponding to sample i and sample j. ⊙ represents the Hadamard product, represents the pre-identified expression label of the i-th sample, S * is the initialized adjacency matrix. Since different data sets require different threshold settings, too high a threshold will lead to insufficient connections between difficult samples, while too low a threshold may introduce unreliable relationships. Therefore, after calculating the similarity matrix, the embodiments of the present invention do not use a threshold to filter out connections with low similarity, but propose this strategy to maintain the connections between simple samples and difficult samples.

[0092] The connection strategy based on the L2 norm distance is as follows:

[0093] According to the theory of deep learning, the model tends to first learn simple samples and then difficult samples. There are significant differences in the norms of simple samples and difficult samples during learning, such as Figure 2 . At the same time, regardless of whether the sample is simple or difficult, as the model gradually learns, its feature norm will increase. However, for mislabeled samples, the correct classification of the model will be penalized in the initial stage, and overfitting may occur in the later stage. Therefore, the feature norms of these mislabeled samples may first decrease, but may increase during overfitting. Based on this, the embodiments of the present invention can separate mislabeled samples from the training set by monitoring the change difference of the feature L2 norm. For this purpose, the embodiments of the present invention propose a regularization term based on the L2 norm distance to optimize the adjacency matrix:

[0094]

[0095] where F represents the Frobenius norm. For an m×n matrix A, its calculation formula is:

[0096]

[0097] A REG The calculation formula of is:

[0098]

[0099] where, is the L2 norm vector corresponding to the deep features of N sample images. diag is the operation of matrixifying and assigning values on the diagonal. I is an all-1 matrix with dimensions N×N, and M ij represents the L2 norm distance matrix, that is:

[0100] Mij = |||f i ||-||f j |||

[0101] A REG The first term of A is used to adjust the learning importance of the node itself, where ‖G‖ is the L2 norm of G and is used for the normalization of diagonal elements. When mislabeled samples are fitted, the L2 norm of their features tends to decrease, resulting in a reduction in the learning weight of the corresponding node, thereby suppressing the learning of that node.

[0102] A REG The second term of A also introduces a label consistency mask because it can be used to suppress the influence of noisy labels in the initial stage of training. When the model is underfitting and mislabeled samples are not fully learned in the initial stage of training, the pre-identified labels output by the convolutional neural network (CNN) may reveal the true class of the samples, which can hinder the formation of incorrect connections caused by noisy labels. However, this inhibitory effect is limited because as the model continues to learn, the CNN will overfit these mislabeled samples, resulting in the formation of unreliable connections. Therefore, it is necessary to optimize the adjacency matrix to continuously suppress noisy labels.

[0103] M is used to adjust the aggregation weight between nodes. As the norm distance between two nodes increases, the weight of the connecting edge between them decreases, and vice versa. Using I ensures that all elements in the entire second term are less than 1. This design suppresses the influence of mislabeled samples on other samples while promoting the learning of difficult samples. For mislabeled samples, due to the continuous suppression of learning under the action of the first term, their L2 norm remains small, and the norm distance from other samples is large. Therefore, even if there is a connection between mislabeled samples and other samples, information propagation will be suppressed. For difficult samples, once their pseudo-labels point to the true class, the LCM strategy will establish a connection between them and easy samples. As difficult samples are gradually learned, their L2 norm increases, and the L2 norm distance from easy samples decreases, which enables easy samples to play a greater role in the learning process of difficult samples.

[0104] As Figure 3 shown, the adjacency matrix regularization model is used to impose constraints on the initialized adjacency matrix output by the initialization of the adjacency matrix, which will be used in the next step of model construction.

[0105] The working principle of the classification network model is specifically described below.

[0106] The classification network model is used to output an expression recognition label based on the depth features of the input sample image and the regularized adjacency matrix.

[0107] Furthermore, the classification network model is a graph convolutional neural network. A graph is represented as Among them represents the set of nodes. In the node classification method, each node represents a facial expression sample; ε represents the set of edges connecting the nodes in the graph, indicating some associations between the nodes. Based on the sample associations, the convolutional operation on the graph model can be expressed as:

[0108]

[0109] Among them, is the adjacency matrix with self-connections added, represents the identity matrix; is 's degree matrix, H (l) represents the node feature matrix of the l-th layer; H (l+1) represents the updated node feature matrix of the l-th layer; W (l) represents the trainable weight of the l-th layer; σ(·) is the activation function. The output of the last graph convolutional layer is supervised by the training loss to optimize the GCN weight W.

[0110] The working principle of the supervised training is described in detail below.

[0111] As Figure 3 shown, the initial adjacency matrix is adopted:

[0112]

[0113] The first-step regularization is carried out using the label-consistent mask:

[0114] A IN = S * ⊙ A LCM

[0115] The second-step regularization is carried out using the L2 norm constraint:

[0116]

[0117] In summary, the total loss of the model is as follows:

[0118]

[0119] Specifically, is the cross-entropy loss function of the noise-robust graph model, is the pre-recognition loss, is the graph network classification loss, is the regularization loss, N is the total number of sample images in the batch, C represents the total number of sample expression categories, w c represents the category weight of the c-th category of the sample, y i,c represents the annotation of the c-th category of the i-th sample, p i,cIt represents the probability of the i-th sample of the c-th category output by the feature extraction and pre-recognition network model, p' i,c It represents the probability of the i-th sample of the c-th category output by the classification network model, and F represents the Frobenius norm.

[0120] Preferably, in the training stage, the embodiment of the present invention first optimizes the convolutional neural network alone for one cycle. In each subsequent cycle, the convolutional neural network and the graph convolutional neural network are jointly optimized. The pre-recognition labels adopted by the label consistency masking strategy come from the prediction results of the convolutional neural network model for all samples in the previous cycle.

[0121] In the inference stage, the test samples are input into the model in small batches. The trained convolutional neural network model provides deep embeddings and pre-recognition labels, and then a graph is constructed. Then, the graph data is input into the trained convolutional neural network to predict the category of each sample.

[0122] In one embodiment, the RAF-DB (Real-world Affective Faces Database) expression library is adopted; this expression library contains 29,672 facial images collected from the Internet and is labeled by 315 crowdsourcing workers; it contains a total of 7 expressions, namely: anger, disgust, fear, happiness, sadness, surprise, and neutral;

[0123] The present invention selects all facial expression pictures and conducts experiments according to the training set and test set divided by the expression library; when using the classical GCN framework, the obtained expression recognition accuracy is 88.33%; when the initial adjacency matrix obtained by the KPCA method is input into the GCN framework, the obtained expression recognition accuracy is 88.69%. After further adding the regularization strategy based on pre-recognition labels and the L2 norm, the best expression recognition accuracy is 90.10%.

[0124] In another embodiment, the FERPlus expression library is adopted; this expression library is an extension of the original FER dataset, in which the facial expression images are relabeled as one of 8 emotion types: neutral, happy, surprised, sad, angry, disgusted, fearful, and contemptuous;

[0125] The present invention selects all facial expression pictures in the dataset for training; when using the classical GCN framework, the obtained expression recognition accuracy is 86.42%; when the initial adjacency matrix obtained by the KPCA method is input into the GCN framework, the obtained expression recognition accuracy is 86.62%. After further adding the regularization strategy based on pre-recognition labels and the L2 norm, the best expression recognition accuracy is 88.94%.

[0126] An embodiment of the present invention further provides a facial expression recognition system based on a noise-robust graph model, including a training module and a prediction module;

[0127] The noise-robust graph model includes a feature extraction and pre-recognition network model, an adjacency matrix initialization model, an adjacency matrix regularization model, and a classification network model;

[0128] The feature extraction and pre-recognition network model is used to extract the depth features and expression pre-recognition labels of the sample image from the input sample image;

[0129] The adjacency matrix initialization model is used to reduce the dimension of the depth features of the input sample image, and construct an initial adjacency matrix using the similarity between the reduced-depth features;

[0130] The adjacency matrix regularization model is used to calculate the L2 norm distance of the sample image, and constrain the initial adjacency matrix based on the L2 norm distance of the depth features of the sample images with the same pre-recognition label, and output a regularized adjacency matrix;

[0131] The classification network model is used to output an expression recognition label according to the depth features of the input sample image and the regularized adjacency matrix;

[0132] The training module is used to calculate the cross-entropy loss function according to the expression pre-recognition label, the initial adjacency matrix, the regularized adjacency matrix, and the expression recognition label, and perform supervised training on the noise-robust graph model using the cross-entropy loss function;

[0133] The prediction module is used to input the facial image to be recognized into the trained noise-robust graph model to recognize the facial expression.

[0134] The working principle and technical effect of the facial expression recognition system based on the noise-robust graph model are the same as those of the above-mentioned facial expression recognition method based on the noise-robust graph model, and will not be elaborated here.

[0135] Those skilled in the art can easily understand that the above are only preferred embodiments of the present invention, and are not intended to limit the present invention. Any modifications, equivalent replacements, and improvements made within the spirit and principle of the present invention shall be included in the protection scope of the present invention.

Claims

1. A facial expression recognition method based on noise robust graphical model, characterized in that: Inputting a facial image to be recognized into the noise robust graph model to recognize facial expressions, wherein the noise robust graph model includes a feature extraction and pre-recognition network model, an adjacency matrix initialization model, an adjacency matrix regularization model, and a classification network model; The feature extraction and pre-recognition network model is used to extract the deep features of the sample image and the expression pre-recognition label from the input sample image; The adjacency matrix initialization model is used to reduce the dimension of the deep features of the input sample image, and construct an initialization adjacency matrix using the similarity between the deep features after the dimension reduction; The adjacency matrix regularization model is used to calculate the L2 norm distance of the depth features of the sample image, and constrain the initialized adjacency matrix based on the L2 norm distance of the depth features of the sample images with the same pre-recognition label, and output a regularized adjacency matrix; The classification network model is used to output expression recognition labels according to the deep features of the input sample image and the regularized adjacency matrix; A cross entropy loss function is calculated based on the expression pre-recognition label, the initialized adjacency matrix, the regularized adjacency matrix and the expression recognition label, and the noise robust graph model is supervised and trained using the cross entropy loss function.

2. The facial expression recognition method based on the noise robust graphical model as claimed in claim 1, characterized in that The feature extraction and pre-identification network model is a convolutional neural network, and the classification network model is a graph convolutional neural network.

3. The facial expression recognition method based on the noise robust graphical model as claimed in claim 1, characterized in that The feature extraction and pre-identification network model is first optimized and trained separately, and then the noise robust graph model is supervised and trained.

4. The facial expression recognition method based on the noise robust graphical model as claimed in claim 3, characterized in that, A deep neural network model pre-trained with a large-scale face dataset is used as the feature extraction and pre-recognition network model, and the feature extraction and pre-recognition network model is individually optimized and trained using sample images with expression labels.

5. The facial expression recognition method based on the noise robust graphical model as claimed in claim 1, characterized in that, The step of reducing the dimension of the deep features of the input sample image comprises the following steps: Using the cosine kernel function to incorporate nonlinearity, the calculation formula is: Among them, k ij represents the cosine similarity between the i-th sample and the j-th sample, f i is the depth feature of the i-th sample, f j is the deep feature of the jth sample, k(.) is the cosine kernel function, T represents the transpose operation of the matrix or vector, ‖.‖ represents the vector modulus, ‖f i ‖ represents f i The modulus length, ||f j || represents f j The module length; The nonlinear principal component analysis method is used to reduce the feature dimension, and the calculation formula is: K=(k ij ) Among them, 1 N is a vector of all 1s with a dimension of N×1, N is the total number of sample images in a batch, I is a matrix of all 1s with a dimension of N×N, and K is the matrix of all k ij As the core matrix composed of elements, is the kernel matrix after the centralization operation, Λ is given by The diagonal matrix constructed by descending eigenvalues, V is a matrix The eigenvector matrix, V m,ij Indicates that in the mth principal component of V, sample f j f i The weight, V m,ji Indicates that in the mth principal component of V, sample f i f j The weight of m is the feature dimension after dimensionality reduction, and f i * f i The features after dimensionality reduction, f j Features after dimensionality reduction.

6. The facial expression recognition method based on the noise robust graphical model as claimed in claim 5, characterized in that, The calculation formula for constructing the initialization adjacency matrix using the similarity between the deep features after dimensionality reduction is: in, is the feature similarity between the i-th sample and the j-th sample after dimensionality reduction, S * For all Initialized adjacency matrix consisting of elements.

7. The facial expression recognition method based on the noise robust graphical model as claimed in claim 1, characterized in that: The method of constraining the initialization adjacency matrix based on the L2 norm distance of the deep features of the sample images with the same pre-identified labels and outputting the regularized adjacency matrix comprises the following steps: The connection is constructed using the expression pre-recognition label of the sample image. The calculation formula is: A IN =S * ⊙A LCM A LCM ={A LCM (i,j)} Among them, A IN For S * The adjacency matrix S is initialized with the samples with the same pre-identified labels retained in * To initialize the adjacency matrix, A LCM A is the label consistency matrix representing whether the pre-identified labels are the same. LCM (i,j) is A LcM The elements of represent the consistency of the pre-identified labels corresponding to sample i and sample j, ⊙ represents the Hadamard product, represents the expression pre-recognition label of the i-th sample, represents the expression pre-recognition label of the jth sample; The L2 norm distance of the deep features of the sample image is used to construct a regularized adjacency matrix. The calculation formula is: Among them, A REG is the regularized adjacency matrix, is the L2 norm vector corresponding to the deep features of N sample images, diag is the operation of matrixization and diagonal assignment, I is an all-1 matrix with dimension N×N, and M represents the L2 norm distance matrix of the sample image.

8. The facial expression recognition method based on the noise robust graphical model as claimed in claim 1, characterized in that: The calculation formula for calculating the cross entropy loss function based on the expression pre-recognition label, the initialized adjacency matrix, the regularized adjacency matrix and the expression recognition label is: in, is the cross entropy loss function of the noise robust graphical model, To pre-identify losses, is the graph network classification loss, is the regularization loss, N is the total number of sample images in a batch, C represents the total number of sample expression categories, and w c Represents the category weight of the cth category of the sample, y i,c represents the label of the cth category of the ith sample, p i,c represents the probability of the cth category of the ith sample output by the feature extraction and pre-identification network model, p′ i,c A represents the probability of the cth category of the i-th sample output by the classification network model, IN Initialize the adjacency matrix for samples with the same pre-identified label, A REG is the regularized adjacency matrix, and F represents the Frobenius norm.

9. A facial expression recognition system based on noise robust graphical model, characterized in that: Includes training module and prediction module; The noise robust graph model includes a feature extraction and pre-identification network model, an adjacency matrix initialization model, an adjacency matrix regularization model and a classification network model; The feature extraction and pre-recognition network model is used to extract the deep features of the sample image and the expression pre-recognition label from the input sample image; The adjacency matrix initialization model is used to reduce the dimension of the deep features of the input sample image, and construct an initialization adjacency matrix using the similarity between the deep features after the dimension reduction; The adjacency matrix regularization model is used to calculate the L2 norm distance of the sample image, and constrain the initialized adjacency matrix based on the L2 norm distance of the deep features of the sample images with the same pre-recognition label, and output the regularized adjacency matrix; The classification network model is used to output expression recognition labels according to the deep features of the input sample image and the regularized adjacency matrix; The training module is used to calculate the cross entropy loss function according to the expression pre-recognition label, the initialized adjacency matrix, the regularized adjacency matrix and the expression recognition label, and use the cross entropy loss function to supervise the noise robust graph model; The prediction module is used to input the facial image to be recognized into the trained noise robust graph model to recognize facial expressions.

10. The facial expression recognition system based on noise robust graphical model as claimed in claim 9, characterized in that: The feature extraction and pre-identification network model is a convolutional neural network, and the classification network model is a graph convolutional neural network.

Citation Information

Cited By

  • Confusion expression recognition method, device and equipment based on emotion matrix and high-aggregation sub-graph network and storage medium

    CN121074967A