Expression editing method based on facial components, electronic device and readable storage medium
Through deep learning networks combining semantic segmentation, AU prior knowledge, attention mechanism and gradient inversion technology to extract and fuse facial expression-related and irrelevant features, and combined with diffusion model for image reconstruction, the problem of facial structure distortion during expression editing in the existing technology is solved, and high-quality and natural expression editing effects are achieved.
Patent Information
- Application Number
- CN202411650340.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-11-19
- Publication Date
- 2025-05-13
- Estimated Expiration
- 2044-11-19
AI Technical Summary
Existing facial expression editing methods are difficult to effectively separate expressions from non-expression features, resulting in the problem of facial structure distortion or loss of features during expression editing.
Deep learning network is used to extract expression-related and irrelevant features in face images, and feature extraction and fusion is performed through semantic segmentation, facial action unit (AU) prior knowledge, attention mechanism and gradient inversion technology, and image reconstruction is performed by combining diffusion model.
It realizes the precise separation of expressions and non-expression characteristics, maintains the stability of facial structure, and ensures the natural, smooth and high-quality expression editing.
Smart Images

Figure CN119152082B_ABST
Abstract
Description
Technical Field
[0001] The invention belongs to the field of computer vision, and in particular relates to an expression editing method based on facial components, an electronic device and a readable storage medium. Background Art
[0002] Facial expression editing technology uses advanced deep learning models to decouple facial expressions from other facial features, and is widely used in virtual reality, animation, social media, and emotional computing. This technology is not only used for expression design of virtual characters and augmented reality interaction, but also plays a role in emotion monitoring, psychotherapy, facial recognition and other scenarios. Although existing methods still face challenges in separating expressions from non-expression features, the accuracy and naturalness of facial expression editing will continue to improve through the development of technologies such as cross-modal learning and attention mechanisms, and its application prospects are broad.
[0003] Existing facial expression editing methods mainly rely on generative models trained with large-scale data, but these methods still face great challenges in decoupling expressions and maintaining the consistency of facial structure. In particular, most current methods find it difficult to effectively separate features that are not related to expressions, resulting in facial structure distortion or feature loss when editing expressions. Therefore, how to effectively decouple expressions from other facial features while maintaining the stability of the facial structure has become a technical problem that needs to be solved in this field. Summary of the invention
[0004] In order to make up for the shortcomings of the prior art, the purpose of the present invention is to provide an expression editing method based on facial components, an electronic device and a readable storage medium. For any input image containing a human face, a deep learning network is used to extract features related to facial expressions and features unrelated to facial expressions in the original image, and prior knowledge such as AU is introduced in the feature extraction process, thereby constructing a fine-grained facial expression editing result.
[0005] The technical problem solved by the present invention can be achieved through the following specific technical solutions:
[0006] On the one hand, a facial expression editing method based on facial components is provided, comprising the following steps:
[0007] Step 1: Use the SemanticStyleGAN model to perform semantic segmentation on the input face image to obtain local facial components of the face, and use facial action units as prior knowledge related to expression to extract local facial features through a facial local feature extraction network;
[0008] Step 2: While executing step 1, use the face key point detector MobileFaceNet to detect the face key points of the input face image to obtain the global facial features;
[0009] Step 3: Use the attention mechanism to learn the relationship between global features and local features, and perform feature fusion to obtain features related to expression;
[0010] Step 4: While executing step 3, using gradient inversion technology to extract features unrelated to expression from the input face image;
[0011] Step 5, combining the features related to expression obtained in step 3 with the features unrelated to expression obtained in step 4, and reconstructing the image through a diffusion model;
[0012] Step 6: Perform model inference based on the input face image and the desired edited expression to generate a face image after expression editing.
[0013] Furthermore, the facial local feature extraction network in step 1 includes a facial component feature extractor, a text feature extractor and a cross-modal attention cross module, wherein the facial component feature extractor is the visual encoder ViT in CLIP, which is used to extract features from each facial component and encode it into a feature vector; the text feature extractor is the text encoder Transformer in CLIP, the input text consists of expression labels and AU prior knowledge, and the text encoder extracts features from the input text and encodes it into a text feature vector; the cross-modal attention cross module is used to interactively fuse the facial component feature vector with the text feature vector, and output a local expression feature vector, which can be expressed as follows:
[0014] ,
[0015] in, The local expression feature vector is obtained by cross-modal fusion of visual features and text features, representing the expression information related to the local components of the face; Represents the image feature vector corresponding to the local components of the face. In the visual encoder ViT of the CLIP model, the input image is converted into a feature vector after encoding ; Represents the input text feature vector. The text encoder Transformer encodes the input text and generates a feature vector that represents the text information related to the expression and AU prior knowledge; It is a function used to normalize the similarity calculation to ensure that the sum of the local feature weights is 1; represents the dimension scaling factor of the feature vector, It is the dimension of the feature vector. The square root of the dimension is usually used for scaling in the attention mechanism to prevent the gradient from disappearing or exploding due to excessive values in the Softmax function.
[0016] Furthermore, in step 3, the relationship between global features and local features is learned using the attention mechanism, the input is local features and global features, and the output is features related to expression. The formula can be expressed as:
[0017] ,
[0018] in, represents the feature vector related to expression obtained by feature fusion; is the global feature vector, which represents the overall facial features extracted from the input face image; is the local feature vector, which represents the features extracted from each local component of the face; is the number of local features, indicating the number of local components of the face segmentation; symbol Represents the dot product of global features and local features, measuring the similarity between them; Represents the transpose of the local eigenvector, used to compare with the global eigenvector Perform a dot product; is the scaling factor of the feature vector dimension, which is used to normalize the similarity scores in the attention mechanism.
[0019] Furthermore, in step 4, the network for extracting expression-independent features includes an image feature extractor ViT, the input is a face image, and the output is a feature that is independent of expression. The adversarial classification is achieved by adding the gradient reversal technology to the classifier. The classifier structure includes a gradient reversal layer, a linear layer, a Relu layer, and a Softmax layer. The loss function can be expressed as follows:
[0020] ,
[0021] in, Represents the probability of the kth category predicted by the classifier. When indicating the type of expression and Consistent, among which The feature information related to expression is trained through adversarial classification, so that the expression-independent feature extractor can extract expression-independent features.
[0022] Furthermore, in the image reconstruction process in step 5, the expression-related features obtained in step 3 and the expression-irrelevant features obtained in step 4 are combined, the input of the image reconstruction module is the original image, the expression-related features and the expression-irrelevant features, and the output is the reconstructed image. The image reconstruction module is a diffusion model, and the denoising process starts from a highly noisy latent image, and gradually removes the noise through multiple denoising steps.
[0023] Furthermore, in each stage of the denoising process in step 5, a cross-attention mechanism is used to combine the latent code of the reference image and the current noisy latent image, thereby retaining the structure and feature information of the face; the feature combination step fuses the expression-related features and the irrelevant features through linear projection to ensure that the image maintains its original facial features when the expression is adjusted; and after multiple rounds of diffusion cycles, a noise-free latent image is generated, and a high-fidelity facial image is reconstructed through a decoder.
[0024] On the other hand, an electronic device is provided, including one or more processors and a memory, wherein one or more programs are stored in the memory, and the one or more programs include instructions for executing the above-mentioned facial component-based expression editing method.
[0025] A readable storage medium is also provided, comprising one or more programs for execution by one or more processors of an electronic device, wherein the one or more programs include instructions for executing the above-mentioned facial component-based expression editing method.
[0026] Compared with the prior art, the present invention has the following advantages:
[0027] (1) The present invention uses SemanticStyleGAN to perform semantic segmentation on facial images, combines the cross-modal attention mechanism of the CLIP model with the global feature extraction of MobileFaceNet, and achieves accurate separation of expression features and non-expression features; this method can maintain the integrity of local facial components and avoid facial structure distortion during expression editing, ensuring the natural coordination between expression and overall facial features.
[0028] (2) The present invention innovatively introduces CLIP, attention mechanism and gradient reversal technology, and dynamically balances the relationship between global features and local features through multi-level feature extraction and fusion. In particular, the application of gradient reversal technology effectively isolates facial features irrelevant to expression, ensuring that the overall facial appearance is not affected when editing expressions, thereby ensuring the natural and smooth editing of expressions.
[0029] (3) The present invention adopts a diffusion model for image reconstruction, which greatly improves the image quality after expression editing. The diffusion model gradually denoises and reconstructs the image, and the generated effect is more natural and delicate, which significantly improves the visual effect of facial expression editing. In addition, combined with the efficient global feature extraction method of MobileFaceNet, the present invention greatly improves the processing efficiency while ensuring accuracy, and is suitable for a variety of practical application scenarios. BRIEF DESCRIPTION OF THE DRAWINGS
[0030] Figure 1 A schematic diagram of the flow chart of the expression editing method of the present invention;
[0031] Figure 2 A schematic diagram of a facial expression-related feature extraction structure of the present invention;
[0032] Figure 3 Schematic diagram of facial components extracted for the present invention;
[0033] Figure 4 A schematic diagram of facial key points extracted by the present invention;
[0034] Figure 5 A schematic diagram of a facial expression-independent feature extraction structure of the present invention;
[0035] Figure 6 It is a schematic diagram of the facial image reconstruction structure of the present invention. DETAILED DESCRIPTION
[0036] In order to make the purpose, technical solution and advantages of the present invention clearer, the present invention is further described in detail below in conjunction with the accompanying drawings and specific embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention. Example 1
[0037] The present embodiment provides an expression editing method based on facial components, which includes a facial global feature and local feature extraction module, an expression-irrelevant feature extraction module and a diffusion model to generate a facial expression editing model. The facial image and its corresponding expression label and expected expression label are used as input, and expression-related features are obtained using local feature extraction and global feature extraction modules. At the same time, expression-irrelevant features are obtained using gradient inversion technology. Then, combined with the diffusion model, based on the fusion results of expression-related and irrelevant features, the method gradually removes noise and reconstructs a high-fidelity edited image.
[0038] like Figure 1 and Figure 2 As shown, the facial expression editing method based on facial components includes the following steps:
[0039] Step 1: Use the SemanticStyleGAN model to perform semantic segmentation on the input face image to obtain the local facial components of the face, and use the facial action unit (AU) as the prior knowledge related to the expression to extract the local facial features through the facial local feature extraction network.
[0040] The input of the SemanticStyleGAN model is a face image, and the output is the extracted local facial components. Figure 3As shown in the figure, considering that not all facial components are related to expressions, 9 facial components are selected for semantic alignment with AU prior knowledge. These 9 facial components are left and right eyebrows, left and right eyes, upper and lower lips, mouth, nose, and face shape. The text semantics are enriched using facial action units and text descriptions of expressions, such as "Happiness with cheek raiser, lipcorner puller or lip part". The pre-trained visual encoder ViT in CLIP is used to extract the features of facial components and encode them into feature vectors. The pre-trained text encoder Transformer in CLIP is used to extract features from the input text and encode it into a text feature vector. The cross-modal attention cross module is used to interactively fuse the facial component feature vector and the text feature vector to output a local expression feature vector.
[0041] Among them, the facial local feature extraction network includes a facial component feature extractor, a text feature extractor and a cross-modal attention cross module. The facial component feature extractor is the visual encoder ViT in CLIP, which is used to extract features from each facial component and encode it into a feature vector; the text feature extractor is the text encoder Transformer in CLIP, the input text consists of expression labels and AU prior knowledge, and the text encoder extracts features from the input text and encodes it into a text feature vector; the cross-modal attention cross module is used to interactively fuse the facial component feature vector with the text feature vector, and output the local expression feature vector, which can be expressed as:
[0042] ,
[0043] in, The local expression feature vector is obtained by cross-modal fusion of visual features and text features, representing the expression information related to the local components of the face; Represents the image feature vector corresponding to the local components of the face. In the visual encoder ViT of the CLIP model, the input image is converted into a feature vector after encoding ; Represents the input text feature vector. The text encoder Transformer encodes the input text and generates a feature vector that represents the text information related to the expression and AU prior knowledge; It is a function used to normalize the similarity calculation to ensure that the sum of the local feature weights is 1; represents the dimension scaling factor of the feature vector, It is the dimension of the feature vector. The square root of the dimension is usually used for scaling in the attention mechanism to prevent the gradient from disappearing or exploding due to excessive values in the Softmax function.
[0044] It can be seen that the SemanticStyleGAN model achieves more sophisticated image generation and editing by independently controlling the local semantic regions of the image. It processes the structure and texture of different regions separately, provides the ability to decouple local features, makes expression editing more accurate, and can achieve independent control of details while maintaining overall consistency. Facial action units are standardized expression coding systems that describe facial muscle movements. In the present invention, they are used as prior knowledge for the generation and adjustment of expression features; combined with the AU system, expression editing is made more realistic and natural, ensuring that the generated expressions conform to the laws of facial muscle movement.
[0045] Step 2: While executing step 1, use the face key point detector MobileFaceNet to detect the face key points of the input face image to obtain the global facial features.
[0046] This embodiment uses the model provided by the paper "Mobilefacenets: Efficient CNNs for accurate real-time face verification on mobile devices" to detect key points of the face and obtain global facial features, such as Figure 4 shown.
[0047] Step 3: Use the attention mechanism to learn the relationship between global features and local features, and perform feature fusion to obtain features related to expression.
[0048] In this embodiment, the attention mechanism is used to learn the relationship between global features and local features. The input is local features and global features, and the output is features related to expression. The formula can be expressed as:
[0049] ,
[0050] in, represents the feature vector related to expression obtained by feature fusion; is the global feature vector, which represents the overall facial features extracted from the input face image; is the local feature vector, which represents the features extracted from each local component of the face; is the number of local features, indicating the number of local components of the face segmentation; symbol Represents the dot product of global features and local features, measuring the similarity between them; Represents the transpose of the local eigenvector, used to compare with the global eigenvector Perform a dot product; is the scaling factor of the feature vector dimension, which is used to normalize the similarity scores in the attention mechanism.
[0051] Step 4: While executing step 3, use the gradient inversion technique to extract features unrelated to expression from the input face image.
[0052] like Figure 5 As shown, the network for extracting expression-independent features includes an image feature extractor ViT, the input is a face image, and the output is features that are independent of expression. The adversarial classification is achieved by adding the gradient reversal technology to the classifier. The classifier structure includes a gradient reversal layer, a linear layer, a Relu layer, and a Softmax layer. The loss function can be expressed as follows:
[0053] ,
[0054] in, Represents the probability of the kth category predicted by the classifier. When indicating the type of expression and Consistent, among which The feature information related to expression is trained through adversarial classification, so that the expression-independent feature extractor can extract expression-independent features.
[0055] Step 5: Combine the expression-related features obtained in step 3 with the expression-independent features obtained in step 4, and reconstruct the image through a diffusion model.
[0056] like Figure 6 As shown, in the image reconstruction process of this embodiment, the expression-related features obtained in step 3 and the expression-irrelevant features obtained in step 4 are combined, the input of the image reconstruction module is the original image, the expression-related features and the expression-irrelevant features, and the output is the reconstructed image. The image reconstruction module is a diffusion model, and the denoising process starts from a highly noisy latent image, and gradually removes the noise through multiple denoising steps. In each stage of the denoising process, a cross-attention mechanism is used to combine the latent coding of the reference image and the current noisy latent image, thereby retaining the structure and feature information of the face; the feature combination step fuses the expression-related features and the irrelevant features through linear projection to ensure that the image maintains its original facial features when the expression is adjusted; and after multiple rounds of diffusion cycles, a noise-free latent image is generated, and a high-fidelity facial image is reconstructed through the decoder.
[0057] Step 6: Perform model inference based on the input facial image and the desired edited expression (i.e., perform model inference according to steps 1 to 5 using the model weights obtained after the model training converges according to steps 1 to 5) to generate a facial image after expression editing. Example 2
[0058] This embodiment provides an electronic device, including one or more processors and a memory, wherein the memory stores one or more programs, and the one or more programs include instructions for executing the facial component-based expression editing method of Example 1. Example 3
[0059] This embodiment provides a readable storage medium, including one or more programs for execution by one or more processors of an electronic device, and the one or more programs include instructions for executing the facial component-based expression editing method of Example 1.
[0060] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or replace some or all of the technical features therein with equivalents. However, these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of the present invention.
Claims
1. The facial expression editing method based on facial components is characterized by: The following steps are involved: Step 1: Use the SemanticStyleGAN model to perform semantic segmentation on the input face image to obtain the local facial components of the face, and use the facial action unit as the prior knowledge related to the expression to extract the local facial features through the facial local feature extraction network; the facial local feature extraction network includes a facial component feature extractor, a text feature extractor and a cross-modal attention cross module, the facial component feature extractor is the visual encoder ViT in CLIP, which is used to extract features from each facial component and encode it into a feature vector; the text feature extractor is the text encoder Transformer in CLIP, the input text is composed of expression labels and AU prior knowledge, the text encoder extracts features from the input text and encodes it into a text feature vector; the cross-modal attention cross module is used to interactively fuse the facial component feature vector with the text feature vector, and output the local expression feature vector, which can be expressed as: Among them, l i represents the local expression feature vector, which is obtained by cross-modal fusion of visual features and text features, and represents the expression information related to the local components of the face; i Represents the image feature vector corresponding to the local components of the face. In the visual encoder ViT of the CLIP model, the input image is converted into a feature vector v after encoding. i ; t represents the input text feature vector. The text encoder Transformer encodes the input text and generates a feature vector that represents the text information related to the expression and AU prior knowledge. Softmax is a function used to normalize the similarity calculation to ensure that the sum of the local feature weights is 1. represents the dimension scaling factor of the feature vector, d k is the dimension of the feature vector. The square root of the dimension is usually used for scaling in the attention mechanism to prevent the gradient from vanishing or exploding due to excessive values in the Softmax function. Step 2: While executing step 1, use the face key point detector MobileFaceNet to detect the face key points of the input face image to obtain the global facial features; Step 3: Use the attention mechanism to learn the relationship between global features and local features, and perform feature fusion to obtain features related to expression. Use the attention mechanism to learn the relationship between global features and local features. The input is local features and global features, and the output is features related to expression. The formula can be expressed as: Among them, s represents the feature vector related to expression obtained by feature fusion; g is the global feature vector, which represents the overall facial features extracted from the input face image; l i is a local feature vector, which represents the features extracted from each local component of the face; K is the number of local features, which represents the number of local components into which the face is segmented; the symbol · represents the dot product of the global feature and the local feature, which measures the similarity between them; represents the transpose of the local eigenvector, which is used to perform dot product with the global eigenvector g; is the scaling factor of the feature vector dimension, which is used to normalize the similarity scores in the attention mechanism; Step 4: While executing step 3, using gradient inversion technology to extract features unrelated to expression from the input face image; Step 5, combining the features related to expression obtained in step 3 with the features unrelated to expression obtained in step 4, and reconstructing the image through a diffusion model; Step 6: Perform model inference based on the input face image and the desired edited expression to generate a face image after expression editing.
2. The facial expression editing method based on facial components according to claim 1, characterized in that: In step 4, the network for extracting expression-independent features includes an image feature extractor ViT, the input is a face image, and the output is a feature that is independent of expression. The adversarial classification is achieved by adding the gradient reversal technology to the classifier. The classifier structure includes a gradient reversal layer, a linear layer, a Relu layer, and a Softmax layer. The loss function can be expressed as follows: Among them, p k It represents the probability of the kth category predicted by the classifier. When u==k, it means that the expression category u is consistent with k, where u comes from the feature information related to the expression. Through adversarial classification training, the expression-independent feature extractor can extract expression-independent features.
3. The facial expression editing method based on facial components according to claim 1, characterized in that: In the image reconstruction process in step 5, the expression-related features obtained in step 3 and the expression-irrelevant features obtained in step 4 are combined, the input of the image reconstruction module is the original image, the expression-related features and the expression-irrelevant features, and the output is the reconstructed image. The image reconstruction module is a diffusion model, and the denoising process starts from a highly noisy potential image, and gradually removes the noise through multiple denoising steps.
4. The facial expression editing method based on facial components according to claim 3, characterized in that: In each stage of the denoising process in step 5, a cross-attention mechanism is used to combine the latent code of the reference image and the current noisy latent image, thereby retaining the structure and feature information of the face; the feature combination step fuses the expression-related features and the irrelevant features through linear projection to ensure that the image maintains its original facial features when the expression is adjusted; and after multiple rounds of diffusion cycles, a noise-free latent image is generated, and a high-fidelity facial image is reconstructed through a decoder.
5. An electronic device comprising one or more processors and a memory, wherein the memory stores one or more programs, and the one or more programs include instructions for executing the facial component-based expression editing method as described in any one of claims 1-4.
6. A readable storage medium comprising one or more programs for execution by one or more processors of an electronic device, wherein the one or more programs include instructions for executing the facial component-based expression editing method as described in any one of claims 1-4.
Citation Information
Patent Citations
Real world image super-resolution method based on stable diffusion
CN118918009A
Expression editing method and device and storage medium
CN118968576A