Personalized facial expression recognition method and device based on style adaptive learning

Through the method of adaptive style learning, a style information coding model and expression mapping model are constructed, which solves the problem of weak generalization ability of expression recognition model caused by individual differences, and achieves efficient personalized facial expression recognition.

CN120236307APending Publication Date: 2025-07-01BEIJING NORMAL UNIVERSITY
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202510182199.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-19
Publication Date
2025-07-01

AI Technical Summary

Technical Problem

The prior art is difficult to effectively solve the problem of weak generalization ability of expression recognition models caused by individual differences, especially the poor recognition performance of models trained on limited sample databases on new individual expression data.

Method used

Adaptive style learning is adopted to construct style information encoding model, discriminator model and expression mapping model, and optimize model parameters using alternating iterative training strategies to realize facial style adaptive expression recognition.

Benefits of technology

Effectively synthesize expression samples that are consistent with the face styles to be tested in the target domain, which improves the recognition performance of the expression recognition model for personalized facial expressions under different face styles in the target domain and enhances the generalization ability of the model.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120236307A_ABST
    Figure CN120236307A_ABST
Patent Text Reader

Abstract

The invention provides a personalized facial expression recognition method and device based on style adaptive learning. The personalized facial expression recognition method based on style adaptive learning comprises the following steps: step 1, acquiring two facial expression databases which do not contain the same subject, and respectively taking the two facial expression databases as a training data set and a test data set; 2, constructing a face style information coding model S, performing implicit learning on the face style information, and coding the face style information into a vector with the length of n; 3, constructing a discriminator model D containing true and false discrimination and expression classification dual-task learning; and the like. The method has the advantages that the expression samples consistent with the to-be-detected face styles in the target domain can be effectively synthesized, the expression information of the source domain samples is effectively migrated to the target domain face samples, face style migration of similar expressions is achieved, and the number of training set samples is further increased on the premise that the samples do not need to be manually labeled.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of computer vision, and in particular to a personalized facial expression recognition method and device based on style adaptive learning. Background Art

[0002] Facial expressions, as the external manifestations of human emotions, carry rich emotional information and can effectively reflect the joys, sorrows, anger, and pleasures of the human heart. Expression recognition shows broad application prospects in multiple fields such as artificial intelligence, psychology, and medicine. However, expression recognition in real scenarios faces many challenges, and one of the important challenges is the interference of the individual differences of the subjects on expression recognition. The personalized expressions of the subjects are diverse, making the generalization ability of the expression recognition model trained on a limited sample database weak for new individual expression data. To overcome this challenge faced by automatic expression recognition, it is urgent to model the personalized difference factors such as the appearance and age of the human face, so that the model can effectively achieve robust recognition of personalized expressions.

[0003] The joint learning method of facial expression synthesis and recognition based on the generative adversarial network is an effective technical means to eliminate expression identity differences or achieve personalized expression recognition. Currently, the generative adversarial network is mainly used for eliminating non-frontal head postures of facial expressions and generating expression samples. There is no report on the expression recognition method that uses the generative adversarial network to achieve face style adaptive learning.

[0004] In the publicly disclosed patent application documents, for example, Chinese Patent Application No. CN202110349518.6 discloses a facial expression recognition method, device, equipment, and storage medium. Among them, the method includes: obtaining an original picture; cropping a face image from the original picture through a first neural network and extracting single-person expression features; performing multi-scale extraction on the single-person expression features through a second neural network to obtain attention feature maps of the single-person expression features at different scales; merging the attention feature maps of the single-person expression features at different scales to obtain a fused feature map; using a Softmax classifier to classify and recognize the fused feature map and generating a classification result. The facial expression recognition method of the present invention uses two neural networks, one of which is used to extract facial expression features, and the other is used to recognize the facial expression features at multiple scales, improving the accuracy of facial expression recognition.

[0005] For another example, Chinese Patent Application No. CN202110279160.4 discloses a facial expression recognition method based on frequency domain features and product neural network. The steps of this method are as follows: preprocess the facial expressions in the facial expression dataset; extract the global frequency domain features and local frequency domain features of the face from the preprocessed facial expressions; divide the facial expression dataset into a training set and a test set; construct and initialize a product neural network, and use the divided training set and test set to train and test the product neural network; evaluate the product neural network, collect facial expression test samples and input them into the trained product neural network to obtain the final expression classification. A novel end-to-end product neural network is designed, which combines the global features and local features of the face and provides an efficient facial expression recognition method.

[0006] For another example, Chinese Patent Application No. CN202010320414.8 discloses a facial expression recognition method, system, storage medium, computer program, and terminal. A pre-trained image generation model is combined with a given depth map and RGB image. The trained image generation model can convert the input depth map into an RGB image according to the RGB image style used in training; generate the eyebrows, eyes, and mouth of the expression in the RGB image, and train a convolutional neural network considering the eyebrows, eyes, and mouth. The convolutional neural network realizes expression recognition. By strengthening the feature information of the human eyes, eyebrows, and mouth, the recognition accuracy is higher; the effect of the image generation model is better. Through the image generation model, not only the important information about the expression is retained, but also the RGB image form for expression recognition is unified; the accuracy of expression recognition is also higher.

[0007] None of the above patent applications can solve the diversity caused by individual differences, such as modeling according to personalized difference factors such as the appearance and age of the face, so as to effectively achieve robust recognition of personalized expressions. Summary of the Invention

[0008] In order to solve the problem that the samples in the labeled expression database are limited and there are individual differences between the source domain and target domain expression samples, the present invention proposes a personalized facial expression recognition method and device based on style adaptive learning, so as to effectively recognize personalized expressions and improve the generalization ability of the expression recognition model.

[0009] To achieve this technical purpose, the present invention provides a personalized facial expression recognition method and device based on style adaptive learning.

[0010] The personalized facial expression recognition method based on style adaptive learning includes the following steps:

[0011] Step 1: Obtain two facial expression databases that do not contain the same subjects, and use them as the training dataset (source domain) and the test dataset (target domain) respectively;

[0012] Step 2: Construct a facial style information encoding model S to implicitly learn the facial style information and encode it into a vector of length n;

[0013] Step 3: Construct a discriminator model D that includes dual-task learning of real / fake discrimination and expression classification;

[0014] Step 4: Construct a facial expression mapping model G that includes dual-task learning of synthesis and recognition;

[0015] Step 5: Adopt an alternating iterative training strategy to optimize the parameters of the expression mapping model G and the discriminator model D;

[0016] Step 6: Based on the expression recognition branch in the trained expression mapping model, realize the prediction of the facial expression category of the sample to be tested.

[0017] Further, in Step 1, when obtaining two facial expression databases that do not contain the same subjects and using them as the training dataset (source domain) and the test dataset (target domain) respectively, specifically:

[0018] The two expression databases need to have the same category labels; perform face preprocessing on the two facial expression databases: use the MTCNN face detection framework to accurately locate the position of the face in the image, effectively crop the face area through the face detection box, and scale the cropped face image to a size of 224X224 (unit: pixel).

[0019] Further, in Step 3, when constructing a discriminator model D that includes dual-task learning of real / fake discrimination and expression classification, specifically:

[0020] The discriminator model D discriminates whether the input sample is a real sample (from the target domain), and at the same time trains the network parameters based on the source domain facial expression data to perform expression recognition task learning, so as to realize the determination of the expression category of the input sample;

[0021] The discriminator model D includes a data input layer, 1 convolutional layer (convolution kernel 3X3), V residual encoding units, 1 convolutional layer (convolution kernel 4X4), 1 convolutional layer (convolution kernel 1X1, output layer), 2 LeakyReLU activation layers, 1 fully connected output layer + ReLU activation layer, 1 fully connected output layer. After the shared backbone network layer (including convolutional layer, V residual encoding units), the discriminator model D includes 2 output ends: expression category discrimination, picture real / fake discrimination. The loss function guiding the parameter optimization of the discriminator model D is as follows:

[0022] The loss function for facial expression classification in the source domain is as follows in Equation (1):

[0023]

[0024] where: y sfe represents the expression label of the source domain face, represents the predicted expression label of the discriminator D for the source domain face;

[0025] Adversarial loss function:

[0026]

[0027] where: represents the target domain samples, represents the source domain samples, v s represents the face style information encoding vector of the target domain samples, F D represents the non-linear mapping of the discriminator, F G represents the non-linear mapping in the expression decoding process.

[0028] Furthermore, the construction of the face expression mapping model G for dual-task learning of synthesis and recognition described in step 4 specifically includes:

[0029] Step 4.1 Based on step 2, the style information encoding vector and the expression information encoding result are used to guide the expression synthesis with face style adaptation; the synthesized face expressions are merged with the source domain dataset to form multi-source data for training the expression recognition branch of the expression encoding model, which includes two sub-task learnings: the expression synthesis task and the expression category classification task. The expression synthesis model includes 3 sub-modules: the front-end shared expression information encoding module, the expression information discriminator module, and the expression information decoding module;

[0030] Step 4.2 Front-end shared expression information encoding module: This module includes a data input layer, 1 convolutional layer (convolution kernel 3X3), and K residual encoding units. The dual-task face expression recognition model has a shared sub-module in the front-end network, that is, the expression encoding module. The output features of this module are respectively used as the input of the expression recognition sub-task module and as the input of the deconvolution layer to achieve face expression synthesis;

[0031] Step 4.3 Expression information discriminator module: This module includes 2 residual encoding units, 2 LeakyReLU activation layers, 1 convolutional layer (convolution kernel 4X4), and 1 fully connected output layer;

[0032] Step 4.4 Expression information decoding module: This module includes Q residual decoding units, 1 InstanceNorm2d normalization layer, 1 LeakyReLU activation layer, and 1 convolutional output layer;

[0033] Step 4.5 To achieve the joint learning of expression synthesis and expression discrimination, it is necessary to design a corresponding loss function for the network to constrain the optimization of parameters. The specific loss constraint function is as follows:

[0034] Step 4.5.1 Facial style encoding learning: The style information encoding network includes a data input layer, a 1-layer convolutional layer (convolution kernel 3x3), M-layer residual encoding units, a 1-layer LeakyReLU activation function layer, a 1-layer convolutional layer (convolution kernel 2x2)+LeakyReLU activation layer, and a fully connected output layer. The output of the facial style information encoding network is expressed as v = F style (X t ), where the input is X t (target domain sample). The learning process of this style information encoding module is regarded as a non-linear mapping F t (·) for X style . The learned style information encoding vector is further used to guide the style-adaptive expression synthesis. The process is as follows: X g = F G (X o , v), where: X o is the output of the expression information encoding module, and X g represents the synthesized facial expression. It is the decoded output result obtained by inputting the expression encoding information of the source domain facial expression sample and the target domain facial style encoding information v into N deconvolution network layers. F G (·) represents the non-linear mapping in the decoding process. To ensure that the network can effectively synthesize style-adaptive facial expressions based on the style encoding information, the following constraint function needs to be designed for the learning process: L re = |v s - v g |1, where: v g = F style (F G (F e (X s ), v s ))), v s represents the style information encoding vector obtained according to the target domain sample style, v g represents the style encoding information of the near-target domain facial sample synthesized after referring to the target domain facial sample style. The smaller the difference between v s and v g , the closer the style of the synthesized expression is to the facial style of the real target domain sample. F e (·) represents the facial expression information encoding process;

[0035] Step 4.5.2 Diversity Constraint of Synthetic Samples: During the process of facial expression synthesis, the differences between synthetic samples need to be considered. The following loss constraint function needs to be designed for the learning process, as shown in Equation (3) below:

[0036] L diff = -F G (F e (X s ), v s1 ) - F G (F e (X s ), v s2 )1......(3),

[0037] where: v s1 and v s2 are the style information encoding vectors of two different faces, X s is the source domain sample, F e (·) represents the encoding of facial expression information, F G (·) represents the non-linear mapping in the decoding process, L diff = -|·1 indicates that the difference between synthetic samples under two different face styles needs to be maximized during optimization, and |·1 represents the sum of the absolute values of vector components;

[0038] Step 4.5.3 Loss Function for Facial Expression Quality Evaluation:

[0039] Loss Function for Synthetic Facial Expression Quality Evaluation Based on Discriminator D: To effectively generate facial expressions of specific categories, it is necessary to determine the facial expression categories of synthetic samples. The quality evaluation loss function is as shown in Equation (4) below:

[0040]

[0041] where: y sfe ij represents the label of the i-th sample of the j-th facial expression of the source domain face, represents the predicted label of the discriminator D for the synthetic expression of the i-th sample of the j-th facial expression, represents the summation of the cross-entropy losses for each facial expression category, C is the number of facial expression categories, and n is the total number of facial expression samples;

[0042] Step 4.5.4 Combine the synthetic facial expressions with the facial expression data set of the source domain face, and thus form a new training set for training the facial expression recognition branch network of the facial expression encoding model, promoting the improvement of the facial expression recognition model for personalized facial expressions of different identities in the target domain:

[0043]

[0044] where: ysfe ij Denotes the label of the \(i\)-th sample of the \(j\)-th type of expression in the source domain. Denotes the predicted class label of the facial expression of the \(i\)-th sample of the \(j\)-th type of expression in the source domain by the expression encoding recognition sub-network. Denotes the predicted class label of the synthesized expression of the \(i\)-th sample of the \(j\)-th type of expression by the expression encoding recognition sub-network. Denotes the summation of the cross-entropy losses for each type of expression, where \(C\) is the number of expression classes and \(n\) is the total number of expression samples.

[0045] Furthermore, in step 5, the training strategy of alternating iteration is adopted to optimize the parameters of the face expression mapping model \(G\) and the discriminator model \(D\) for the dual tasks of synthesis and recognition:

[0046] The loss function guiding the expression synthesis model \(G\) in step 5.1 is shown in the following formula (6):

[0047]

[0048] Where: Denotes the adversarial loss function, whose purpose is to prompt the network to synthesize source domain face samples into realistic target domain face pictures, so that the discriminator cannot identify the authenticity. \(\lambda\) α Is the coefficient of the sample diversity loss constraint term, \(L\) diff Synthetic sample diversity loss constraint, \(L\) re Face style reconstruction constraint loss function, \(L\) feq Expression quality assessment loss function, \(L\) SFE Denotes the multi-source expression classification loss function;

[0049] The loss function guiding the expression discriminator model \(D\) in step 5.2 is shown in the following formula (7):

[0050]

[0051] Where: Adversarial loss function \(L\) fed Source domain face expression classification loss function, the parameters of the face expression mapping model \(G\) for the dual tasks of synthesis and recognition are defined as \(\theta\) G , the parameters of the expression-class - picture authenticity dual-task discriminator \(\theta\) D , the parameters of the face style encoding network \(\theta\) s , and the optimal parameters of the network module are sought through alternating iterative training Where in the iteration process, the parameter \(\theta\) G is fixed, and the parameter \(\theta\) D is updated; the parameter \(\theta\) D is fixed, and the parameters \(\theta\) G and the parameter \(\theta\) s are updated:

[0052] First, fix the network parameters of the face expression mapping model G for dual-task learning of synthesis and recognition, and optimize the network parameters of the discriminator model D; then, fix the network parameters of the discriminator model D, and then optimize the network parameters of the two networks of the face expression mapping model G for dual-task learning of synthesis and recognition and the style information encoding model S. The two are alternately iterated to complete the optimization of the network parameters of all models.

[0053] Furthermore, based on the expression recognition branch in the trained expression mapping model in step 6, predict the face expression category of the sample to be tested:

[0054] The loss function of the face expression mapping model G for guiding dual-task learning of synthesis and recognition in step 6.1 is shown in the following formula (8):

[0055]

[0056] Where: represents the adversarial loss function, which enables the network to synthesize source domain face samples into realistic target domain face pictures, so that the discriminator cannot identify the authenticity. λ α is the coefficient of the sample diversity loss constraint term, L diff is the sample diversity loss constraint for synthesis, L re is the face style reconstruction constraint loss function, L feq is the expression quality evaluation loss function, L SFE represents the multi-source expression classification loss function;

[0057] The loss function of the discriminator model D in step 6.2 is shown in the following formula (9):

[0058]

[0059] Among them, the parameters of the face expression mapping model G for guiding dual-task learning of synthesis and recognition are defined as θ G , the parameters of the expression category-picture authenticity dual-task discriminator θ D , the parameters of the face style encoding network θ s , and the optimal parameters of the network module are sought through alternating iterative training Among them, during the iteration process, fix the parameter θ G , and update the parameter θ D ; fix the parameter θ D , and update the parameter θ G and the parameter θ s ;

[0060] Step 6.3 Based on the expression recognition branch in the trained expression mapping model, predict the facial expression category of the sample to be tested: input the preprocessed test sample into the expression mapping model, and the prediction result of the expression category can be obtained from the output of the expression recognition branch of the model.

[0061] The present invention further provides an apparatus based on the personalized facial expression recognition method based on style adaptive learning, including a processor and a computer program stored in a memory and capable of running on the processor, and the processor executes the computer program to implement the personalized facial expression recognition method based on style adaptive learning.

[0062] Compared with the existing cross-domain expression recognition methods in the technical field, the superior technical effect of the personalized facial expression recognition method based on style adaptive learning of the present invention is that it can effectively synthesize expression samples consistent with the style of the target domain to-be-tested human face, effectively transfer the expression information of the source domain samples to the target domain human face samples, realize the human face style transfer of the same type of expressions, further expand the number of training set samples without manual annotation of samples, and improve the recognition performance of the expression recognition model for personalized facial expressions under different human face styles in the target domain. BRIEF DESCRIPTION OF THE DRAWINGS

[0063] Figure 1 is a schematic flow chart of the personalized facial expression recognition method based on style adaptive learning of the present invention;

[0064] Figure 2 is a schematic diagram of the style adaptive learning-based facial expression synthesis of the personalized facial expression recognition method of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0065] In order to more clearly understand the above objects, features and advantages of the present invention, the present application will be further described in detail below with reference to the drawings and specific embodiments. It should be noted that, without conflict, the embodiments of the present application and the features in the embodiments can be combined with each other.

[0066] Embodiment

[0067] As Figure 1 shown, the personalized facial expression recognition method based on style adaptive learning includes the following steps:

[0068] Step 1, obtain two facial expression databases that do not contain the same subjects, and use them as the training data set (source domain) and the test data set (target domain) respectively;

[0069] Step 2, construct a face style information encoding model S, implicitly learn the face style information, and encode the face style information into a vector with a length of n;

[0070] Step 3, construct a discriminator model D that includes dual-task learning of true / false discrimination and expression classification;

[0071] Step 4, construct a face expression mapping model G that includes dual-task learning of synthesis and recognition;

[0072] Step 5, adopt an alternating iterative training strategy to optimize the parameters of the expression mapping model G and the discriminator model D;

[0073] Step 6, based on the expression recognition branch in the trained expression mapping model, realize the prediction of the face expression category of the sample to be tested.

[0074] As a specific step embodiment of the present invention: in step 1, the two face expression databases that do not contain the same subjects are obtained and used as the training data set (source domain) and the test data set (target domain) respectively. Specifically:

[0075] The two expression databases need to have the same category labels; perform face preprocessing on the two face expression databases: use the MTCNN face detection framework to accurately locate the position of the face in the image, effectively crop the face area through the face detection box, and scale the cropped face image to a size of 224X224 (unit: pixel).

[0076] As a specific step embodiment of the present invention: in step 3, the discriminator model D that includes dual-task learning of true / false discrimination and expression classification is constructed. Specifically:

[0077] The discriminator model D discriminates whether the input sample is a real sample (from the target domain), and at the same time trains the network parameters based on the source domain face expression data to perform expression recognition task learning, so as to realize the determination of the expression category of the input sample;

[0078] The discriminator model D includes a data input layer, 1 convolutional layer (convolution kernel 3X3), V residual encoding units, 1 convolutional layer (convolution kernel 4X4), 1 convolutional layer (convolution kernel 1X1, output layer), 2 LeakyReLU activation layers, 1 fully connected output layer + ReLU activation layer, 1 fully connected output layer. The discriminator model D includes 2 output ends after the shared backbone network layer (including convolutional layer, V residual encoding units): expression category discrimination, picture true / false discrimination. The loss function guiding the parameter optimization of the discriminator model D is as follows:

[0079] The source domain face expression classification loss function is as follows in formula (1):

[0080]

[0081] where: y sfe represents the expression label of the source domain face, Denote the expression prediction label of the discriminator D for the source domain face;

[0082] Adversarial loss function:

[0083]

[0084] Where: Denote the target domain samples, Denote the source domain samples, v s Denote the face style information encoding vector of the target domain samples, F D Denote the non-linear mapping of the discriminator, F G Denote the non-linear mapping in the expression decoding process.

[0085] As a specific step embodiment of the present invention: The face expression mapping model G for constructing dual-task learning of synthesis and recognition in step 4 includes the following specific steps:

[0086] Step 4.1 On the basis of step 2, use the style information encoding vector and the expression information encoding result to guide the expression synthesis of face style adaptation; Merge the synthesized face expressions with the source domain dataset to form multi-source data for training the expression recognition branch of the expression encoding model, which includes two sub-task learnings: expression synthesis task and expression category classification task. The expression synthesis model includes 3 sub-modules: a front-end shared expression information encoding module, an expression information discriminator module, and an expression information decoding module;

[0087] Step 4.2 Front-end shared expression information encoding module: This module includes a data input layer, 1 convolutional layer (convolution kernel 3X3), and K residual encoding units. The dual-task face expression recognition model is a shared sub-module in the front-end network, that is, the expression encoding module. The output features of this module are respectively used as the input of the expression recognition sub-task module and as the input of the transposed convolutional layer to achieve face expression synthesis;

[0088] Step 4.3 Expression information discriminator module: This module includes 2 residual encoding units, 2 LeakyReLU activation layers, 1 convolutional layer (convolution kernel 4X4), and 1 fully connected output layer;

[0089] Step 4.4 Expression information decoding module: This module includes Q residual decoding units, 1 InstanceNorm2d normalization layer, 1 LeakyReLU activation layer, and 1 convolutional output layer;

[0090] Step 4.5 In order to realize the joint learning of expression synthesis and expression discrimination, it is necessary to design corresponding loss functions for the network to constrain the optimization of parameters. The specific loss constraint functions are as follows:

[0091] Step 4.5.1 Facial Style Encoding Learning: The style information encoding network includes a data input layer, 1 convolutional layer (convolution kernel 3x3), M residual encoding units, 1 LeakyReLU activation function layer, 1 convolutional layer (convolution kernel 2x2) + LeakyReLU activation layer, and a fully connected output layer. The output of the facial style information encoding network is expressed as v = F style (X t ), where the input is X t (target domain sample). The learning process of this style information encoding module is regarded as a non-linear mapping F t (·) for X style . The learned style information encoding vector is further used to guide the style-adaptive expression synthesis. The process is as follows: X g = F G (X o , v), where: X o is the output of the expression information encoding module, and X g represents the synthesized facial expression. It is the decoded output result obtained by inputting the expression encoding information of the source domain facial expression sample and the target domain facial style encoding information v into N deconvolution network layers. F G (·) represents the non-linear mapping in the decoding process; To ensure that the network can effectively synthesize style-adaptive facial expressions based on the style encoding information, the following constraint function needs to be designed for the learning process: L re = |v s - v g |1, where: v g = F style (F G (F e (X s ), v s ))), v s represents the style information encoding vector obtained according to the target domain sample style, v g refers to the style encoding information of the near-target domain face sample synthesized after referring to the target domain face sample style. The smaller the difference between v s and v g , the closer the style of the synthesized expression is to the face style of the real target domain sample. F e (·) represents the facial expression information encoding process;

[0092] Step 4.5.2 Diversity Constraint of Synthesized Samples: In the process of facial expression synthesis, the differences between synthesized samples need to be considered. The following loss constraint function needs to be designed for the learning process:

[0093] L diff = -|F G (F e (Xs ),v s1 )-F G (F e (X s ),v s2 )1......(3),

[0094] Where: v s1 and v s2 Style information encoding vectors of two different faces;

[0095] Step 4.5.3 Expression quality assessment loss function:

[0096] Synthetic expression quality assessment loss function based on discriminator D: In order to effectively generate a specific category of expressions, it is necessary to determine the expression category of the synthetic sample. The quality assessment loss function is as follows:

[0097]

[0098] Where: y sfe ij represents the label of the i-th sample of the j-th expression of the source domain face, represents the predicted label of the discriminator D for the i-th sample of the j-th expression, represents the sum of the cross entropy losses for each type of expression, C is the number of expression categories, and n is the total number of expression samples;

[0099] Step 4.5.4 The synthesized facial expressions and the source domain facial expression data set are combined to form a new training set for training the expression recognition branch network of the expression encoding model, so as to promote the expression recognition model to recognize personalized facial expressions of different identities in the target domain:

[0100]

[0101] in: y sfe ij represents the label of the i-th sample of the j-th category expression in the source domain, represents the category prediction label of the facial expression of the i-th sample of the j-th category expression in the source domain by the expression encoding and recognition subnetwork, represents the category prediction label of the expression encoding and recognition sub-network for the i-th sample of the j-th expression, represents the sum of the cross entropy losses for each type of expression, C is the number of expression categories, and n is the total number of expression samples.

[0102] As a specific step embodiment of the present invention: the step 5 adopts the alternating iterative training strategy to optimize the parameters of the expression mapping model G and the discriminator model D, including the following specific steps:

[0103] Step 5.1 guides the loss function of the expression synthesis model G as follows:

[0104]

[0105] Where: denotes the adversarial loss function, whose purpose is to prompt the network to synthesize source domain face samples into realistic target domain face pictures, so that the discriminator cannot identify the authenticity, and λ α is the coefficient of the sample diversity loss constraint term, and L diff is the synthetic sample diversity loss constraint, and L re is the face style reconstruction constraint loss function, and L feq is the expression quality evaluation loss function, and L SFE denotes the multi-source expression classification loss function;

[0106] Step 5.2 guides the loss function of the expression discriminator model D as shown in formula (7) below:

[0107]

[0108] Where: the adversarial loss function L fed is the source domain face expression classification loss function.

[0109] The parameters of the face expression mapping model G for dual-task learning of synthesis and recognition are defined as θ G , the parameters of the expression category - picture authenticity dual-task discriminator θ D , the parameters of the face style encoding network θ s , and the optimal parameters of the network module are sought through alternating iterative training Among them, during the iteration process, the parameters θ G are fixed, and the parameters θ D are updated; the parameters θ D are fixed, and the parameters θ G and the parameters θ s are updated:

[0110] First, fix the network parameters of the face expression mapping model G for dual-task learning of synthesis and recognition, and optimize the network parameters of the discriminator model D; then, fix the network parameters of the discriminator model D, and then optimize the network parameters of the face expression mapping model G for dual-task learning of synthesis and recognition and the style information encoding model S. The two are alternately iterated to complete the optimization of the network parameters of all models.

[0111] As a specific step implementation example of the present invention: Step 6 realizes the prediction of the face expression category of the to-be-tested sample based on the expression recognition branch in the trained expression mapping model, including the following steps:

[0112] Step 6.1 Guide the loss function of the face expression mapping model G for synthesis and recognition dual-task learning as shown in the following formula (8):

[0113]

[0114] Wherein: represents the adversarial loss function, enabling the network to synthesize source domain face samples into realistic target domain face images, making it impossible for the discriminator to distinguish between genuine and fake, λ α is the coefficient of the sample diversity loss constraint term, L diff is the sample diversity loss constraint for synthesized samples, L re is the face style reconstruction constraint loss function, L feq is the expression quality evaluation loss function, L SFE represents the multi-source expression classification loss function;

[0115] Step 6.2 Guide the loss function of the discriminator model D as shown in the following formula (9):

[0116]

[0117] The parameters of the face expression mapping model G for synthesis and recognition dual-task learning are defined as θ G , the expression category - picture authenticity dual-task discriminator parameters θ D , the face style encoding network parameters θ s , and the optimal parameters of the network module are sought through alternating iterative training Wherein during the iteration process, the parameter θ G is fixed, and the parameter θ D is updated; the parameter θ D is fixed, and the parameters θ G and the parameter θ s are updated;

[0118] Step 6.3 Based on the expression recognition branch in the trained expression mapping model, realize the prediction of the face expression category of the test sample to be tested: Input the preprocessed test sample into the expression mapping model, and the prediction result of the expression category can be obtained from the output of the expression recognition branch of the model.

[0119] The present invention further provides an apparatus for the personalized face expression recognition method based on style adaptive learning, including a processor and a computer program stored on a memory and capable of running on the processor, and the processor executes the computer program to implement the personalized face expression recognition method based on style adaptive learning.

[0120] Application example

[0121] Combined with Figure 2, in order to verify the effectiveness of the personalized facial expression recognition method based on style adaptive learning of the present invention, the cross-database recognition results based on two adult expression databases RAF-DB(A) and ExpW(A) across three child expression databases RAF-DB(B), CFEW(B), and ExpW(B) are shown in Table 1 and Table 2 respectively:

[0122] Table 1 Ablation experiment results of adult expression database across three child facial expression databases (%)

[0123]

[0124] Table 2 Ablation experiment results of adult expression database across three child facial expression databases (%)

[0125]

[0126] Among them, the comparison methods include Deep Adaptation Network (DAN), Domain Adversarial Network (DANN), Deep Correlation Alignment Network (DeepCoral), Dynamic Adversarial Adaptation Network (DAAN), Batch Nuclear Norm Maximization (BNM), and Deep Sub-domain Adaptive Model (DSAN).

[0127] From Figure 2 the experimental results, it can be clearly seen that, among them, the first row represents the adult expression pictures of RAF-DB(A), the first column represents the child expression pictures of RAF-DB(B), and the middle part is an example of the child facial expression pictures synthesized with the child face style as a reference. The expression synthesis experimental results show that the personalized facial expression recognition method based on style adaptive learning of the present invention can effectively and adaptively synthesize the expression pictures of the target domain face style and keep the source domain face expression information unchanged.

[0128] It can also be clearly shown from the above application examples (Table 1 and Table 2) that the personalized facial expression recognition method based on style adaptive learning of the present invention has better recognition effect and application effectiveness than the existing transfer learning methods.

[0129] The present invention is not limited by the above embodiments. What is described in the above embodiments and the specification is only to illustrate the principle of the present invention. Without departing from the spirit and scope of the present invention, the present invention will have various changes and improvements, and these changes and improvements all fall within the scope of the present invention claimed. The scope of protection claimed by the present invention is defined by the appended claims.

Claims

1. A personalized facial expression recognition method based on style adaptive learning, characterized in that: include: Step 1, obtain two facial expression databases that do not contain the same subjects, as a training data set and a test data set respectively; Step 2: construct a facial style information encoding model S, implicitly learn the facial style information, and encode the facial style information into a vector of length n; Step 3, construct a discriminator model D that includes dual-task learning of true and false discrimination and expression classification; Step 4, construct a facial expression mapping model G for synthesis and recognition dual-task learning; Step 5, using an alternating iterative training strategy to optimize the parameters of the expression mapping model G and the discriminator model D; Step 6: Based on the expression recognition branch in the trained expression mapping model, the facial expression category of the sample to be tested is predicted.

2. According to the personalized facial expression recognition method based on style adaptive learning as claimed in claim 1, in step 1, two facial expression databases that do not contain the same subjects are obtained, respectively as a training data set and a test data set, specifically: The two expression databases must have the same category labels; the two facial expression databases are preprocessed: the MTCNN face detection framework is used to accurately locate the position of the face in the image, the face area is effectively cropped through the face detection frame, and the cropped face image is scaled to a pixel size of 224X224.

3. According to the personalized facial expression recognition method based on style adaptive learning as claimed in claim 1, the step 3 of constructing a discriminator model D including dual-task learning of true and false discrimination and expression classification is specifically: The discriminator model D determines whether the input sample is a real sample or not. It also trains the network parameters based on the source domain facial expression data to learn the expression recognition task, thereby determining the expression category of the input sample. The discriminator model D includes a data input layer, a convolution layer with a convolution kernel of 3X3, V residual coding units, a convolution layer with a convolution kernel of 4X4, a convolution layer with a convolution kernel of 1X1, an output layer, two LeakyReLU activation layers, a fully connected output layer + ReLU activation layer, and a fully connected output layer. The discriminator model D includes two output terminals after the shared backbone network layer: expression category discrimination and image authenticity discrimination. The loss function guiding the optimization of the discriminator model D parameters is as follows: The source domain facial expression classification loss function is as follows (1): in: y sfe represents the expression label of the source domain face, represents the expression prediction label of the source domain face by the discriminator D; Adversarial loss function: in: represents the target domain sample, represents the source domain sample, v s The face style information encoding vector of the target domain sample, F D represents the nonlinear mapping of the discriminator, F G Represents the nonlinear mapping in the expression decoding process.

4. According to the personalized facial expression recognition method based on style adaptive learning as claimed in claim 1, the facial expression mapping model G for synthesis and recognition dual-task learning is constructed in step 4, specifically comprising: Step 4.1 Based on step 2, the style information encoding vector and the expression information encoding result are used to guide the expression synthesis of facial style adaptation; the synthesized facial expression is merged with the source domain data set to form multi-source data for training the expression recognition branch of the expression encoding model, which includes two sub-task learning: expression synthesis task and expression category classification task. The expression synthesis model includes three sub-modules: the front-end shared expression information encoding module, expression information discrimination module, and expression information decoding module; Step 4.2: The front-end shared expression information encoding module: The expression information encoding module includes a data input layer, a convolution layer - a convolution kernel 3X3, and K layers of residual coding units. The dual-task facial expression recognition model is a shared submodule in the front-end network, that is, the expression encoding module. The output features of the expression encoding module are used as the input of the expression recognition subtask module and as the input of the deconvolution layer to achieve facial expression synthesis. Step 4.3 Expression information discrimination module: The expression information discrimination module includes 2 layers of residual coding units, 2 layers of LeakyReLU activation layers, 1 convolution layer with a convolution kernel of 4X4, and 1 fully connected output layer; Step 4.4 Expression information decoding module: The expression information decoding module includes Q layers of residual decoding units, 1 layer of InstanceNorm2d normalization layer, 1 layer of LeakyReLU activation layer, and 1 layer of convolution output layer; Step 4.5 In order to realize the joint learning of expression synthesis and expression discrimination, it is necessary to design a corresponding loss function for the network to constrain the optimization of parameters. The specific loss constraint function is as follows: Step 4.5.1 Face style coding learning: The style information coding network includes a data input layer, a convolution layer with a convolution kernel of 3x3, M layers of residual coding units, a LeakyReLU activation function layer, a convolution layer with a convolution kernel of 2x2+LeakyReLU activation layer, and a fully connected output layer. The output of the face style information coding network is represented by v=F style (X t ), where the input is X t , the learning process of the style information encoding module is regarded as t Perform nonlinear mapping F style (·), the learned style information encoding vector is further used to guide style-adaptive expression synthesis, the process is as follows: X g =F G (X o ,v), Where: X o is the output of the expression information encoding module, X g represents the synthesized facial expression, which is the decoded output result obtained by inputting the expression encoding information of the source domain facial expression sample and the target domain facial style encoding information v into N deconvolutional network layers, F G (·) represents the nonlinear mapping in the decoding process; in order to ensure that the network can effectively synthesize style-adaptive facial expressions based on style encoding information, it is necessary to design the following constraint function for the learning process: L re =|v s -v g |1, where: v g =F style (F G (F e (X s ),v s )),v s represents the style information encoding vector obtained according to the target domain sample style, v g The style encoding information of the near-target domain face sample synthesized after referring to the style of the target domain face sample, v s and v g The smaller the difference is, the closer the style of the synthesized expression is to the face style of the real target domain sample. e (·) represents the encoding process of facial expression information; Step 4.5.2 Synthetic sample diversity constraint: In the process of facial expression synthesis, it is necessary to consider the differences between the synthetic samples, and it is necessary to design the following loss constraint function for the learning process, as shown in the following formula (3): L diff =-|F G (F e (X s ),v s1 )-F G (F e (X s ),v s2 )|1......(3), Where: v s1 and v s2 The style information encoding vector of two different faces, X s Source domain samples, F e (·) represents the encoding of facial expression information, F G (·) represents the nonlinear mapping in the decoding process, L diff =-|·|1 indicates that the difference between the synthetic samples under two different face styles needs to be maximized in the optimization process, and |·|1 indicates the sum of the absolute values ​​of the vector components; Step 4.5.3 Expression quality assessment loss function: Synthetic expression quality assessment loss function based on the discriminator D: In order to effectively generate a specific category of expression, it is necessary to determine the expression category of the synthetic sample. The quality assessment loss function is shown in the following formula (4): Where: y sfe ij represents the label of the i-th sample of the j-th expression of the source domain face, represents the predicted label of the discriminator D for the i-th sample of the j-th expression, represents the sum of the cross entropy losses for each type of expression, C is the number of expression categories, and n is the total number of expression samples; Step 4.5.4 The synthesized facial expressions and the source domain facial expression data set are combined to form a new training set for training the expression recognition branch network of the expression encoding model, so as to promote the expression recognition model to recognize personalized facial expressions of different identities in the target domain: in: y sfe ij represents the label of the i-th sample of the j-th category expression in the source domain, represents the category prediction label of the facial expression of the i-th sample of the j-th category expression in the source domain by the expression encoding and recognition subnetwork, represents the category prediction label of the expression encoding and recognition sub-network for the i-th sample of the j-th expression, represents the sum of the cross entropy losses for each type of expression, C is the number of expression categories, and n is the total number of expression samples.

5. According to the personalized facial expression recognition method based on style adaptive learning as claimed in claim 1, in step 5, the alternating iterative training strategy is used to optimize the parameters of the facial expression mapping model G and the discriminator model D of the synthesis and recognition dual-task learning: Step 5.1 The loss function of the expression synthesis model G is shown in the following formula (6): in: represents the adversarial loss function, which aims to force the network to synthesize source domain face samples into realistic target domain face images, so that the discriminator cannot identify the authenticity, λ α is the sample diversity loss constraint coefficient, L diff Synthetic sample diversity loss constraint, L re Face style reconstruction constraint loss function, L feq Expression quality assessment loss function, L SFE represents the multi-source expression classification loss function; Step 5.2 guides the loss function of the expression discriminator model D, as shown in equation (7): in: Adversarial Loss Function L fed The loss function for source domain facial expression classification, the parameters of the facial expression mapping model G for dual-task learning of synthesis and recognition are defined as θ G , expression category-image authenticity dual-task discriminator parameter θ D , face style encoding network parameters θ s , through alternating iterative training to find the optimal parameters of the network module In the iterative process, the fixed parameter θ G , for the parameter θ D Update; fix parameter θ D , for the parameter θ G and parameter θ s To update: First, the network parameters of the facial expression mapping model G for dual-task synthesis and recognition learning are fixed, and the network parameters of the discriminator model D are optimized. Then, the network parameters of the discriminator model D are fixed, and the parameters of the facial expression mapping model G for dual-task synthesis and recognition learning and the style information encoding model S are optimized. The two are alternately iterated to complete the network parameter optimization of all models.

6. According to the personalized facial expression recognition method based on style adaptive learning as claimed in claim 1, the expression recognition branch in the trained expression mapping model in step 6 is used to predict the facial expression category of the sample to be tested: Step 6.1 The loss function of the facial expression mapping model G that guides the synthesis and recognition dual-task learning is shown in the following formula (8): in: represents the adversarial loss function, so that the network can synthesize the source domain face samples into realistic target domain face images, so that the discriminator cannot identify the authenticity, λ α is the sample diversity loss constraint coefficient, L diff Synthetic sample diversity loss constraint, L re Face style reconstruction constraint loss function, L feq Expression quality assessment loss function, L SFE represents the multi-source expression classification loss function; Step 6.2 The loss function of the expression discriminator model D is shown in the following formula (9): Among them, the parameters of the facial expression mapping model G that guides the dual-task learning of synthesis and recognition are defined as θ G , expression category-image authenticity dual-task discriminator parameter θ D , face style encoding network parameters θ s , through alternating iterative training to find the optimal parameters of the network module In the iterative process, the fixed parameter θ G , for the parameter θ D Update; fix parameter θ D , for the parameter θ G and parameter θ s Make updates; Step 6.3 predicts the facial expression category of the test sample based on the expression recognition branch in the trained expression mapping model: input the preprocessed test sample into the expression mapping model, and obtain the prediction result of the expression category from the output of the expression recognition branch of the model.

7. A device based on the personalized facial expression recognition method based on style adaptive learning as described in claim 1, comprising a processor and a computer program stored in a memory and capable of running on the processor, wherein the processor executes the computer program to implement the personalized facial expression recognition method based on style adaptive learning.

Citation Information

Patent Citations

  • Facial expression recognition methods, systems, storage media, computer programs, and terminals

    CN111582067B

  • Facial expression recognition method based on frequency domain features and product neural network

    CN113011314A

  • Facial expression recognition method and device, equipment and storage medium

    CN113095185A