Facial expression transfer method and system based on diffusion model and facial key points
By introducing the Swin Transformer module and ControlNet architecture into the diffusion model, combining the control loss function and the identity maintenance loss function, the fine control and identity information retention problems of facial expression migration in the existing technology are solved, and high-quality and flexible facial expression migration effect is achieved.
Patent Information
- Application Number
- CN202510279546.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-11
- Publication Date
- 2025-05-16
- Estimated Expiration
- 2045-03-11
AI Technical Summary
The prior art method of diffusion model that is difficult to finely control facial expressions is unable to retain the identity information of the target face during the generation process, and lacks flexibility and cannot apply the target face with different expressions.
The text encoder of the original stable diffusion model is replaced by the Swin Transformer module, designed as an identity feature encoder, combined with the ControlNet architecture, an expression migration model is built, and the model is trained by controlling the loss function and identity-keeping loss function that depends on the time step.
It realizes the migration of facial expressions from the source face image to the target face while retaining the identity information of the target face, generating high-quality and flexible facial expression migration results, which are suitable for target faces with different expressions.
Smart Images

Figure CN119784576B_ABST
Abstract
Description
Technical Field
[0001] The present application belongs to the field of computer vision and image processing, and specifically relates to a facial expression migration method and system based on a diffusion model and facial key points. Background Art
[0002] Facial expression is one of the most common emotional behavior signals. Currently, the widely accepted classification of facial expressions includes 7 categories, namely anger, disgust, fear, happiness, neutrality, sadness and surprise, which provides an important basis for specific application fields such as facial expression transfer. Given a source face image and a target face image, facial expression transfer refers to transferring the expression in the source face image to the target face while retaining the identity information of the target face.
[0003] Traditional facial expression transfer methods are based on building face models, such as 3DMM models and Blendshape models, and then face deformation, interpolation, synthesis and other operations are performed on this basis to achieve facial expression transfer. However, the disadvantage of traditional methods is that the computational overhead is large and the parameters need to be manually adjusted. With the emergence and rapid development of deep learning, scholars have begun to try to use generative models based on deep neural networks for image generation, such as variational autoencoders (VAE), but due to their weak ability to fit data distribution, the quality of generated images is low. After the Generative Adversarial Networks (GAN) was proposed, GAN was widely used in the field of image generation due to its ability to fit complex data distribution brought by its adversarial game learning mechanism. Scholars have designed many methods for facial expression transfer based on GAN models, such as StyleGAN and StarGAN. However, the disadvantage of the GAN model is that its training process is unstable and prone to mode collapse, resulting in a single generated sample. In recent years, the introduction of the diffusion model has greatly promoted the development of the field of image generation. The diffusion model generates images by learning to simulate the reverse diffusion process from noise to real data. It has attracted widespread attention from researchers because of its high quality and rich details. ControlNet is a deep learning architecture designed to control the generation effect of the diffusion model. Its core idea is to introduce additional conditional control mechanisms based on existing large-scale text graph models (such as Stable Diffusion), so as to achieve the ability to control the generation results of the model according to specific input conditions. The original ControlNet proposed control conditions in the form of edge graphs, posture graphs, depth graphs, segmentation graphs, etc. to control the generation effect of the model.
[0004] However, the existing technologies have the following deficiencies: First, there is a lack of diffusion model methods that can finely control facial expressions for facial expression transfer; second, it is impossible to retain the identity information in the target face image during the generation process to achieve the effect of transferring the facial expression of the source face to the target face required by the facial expression transfer task; third, most of the existing expression transfer technologies use neutral expression images as a benchmark for transfer, that is, assuming that the target face has a neutral expression, and the transfer result is generated by converting the neutral expression image. Such technologies lack flexibility and cannot cope with diverse transfer scenarios with arbitrary expressions as input. Therefore, the existing technologies cannot transfer the facial expression of the source face image to the target face while retaining the identity information of the target face, and cannot be applied to target faces with different expressions, resulting in difficulty in generating high-quality and flexible facial expression transfer results. Summary of the invention
[0005] The purpose of the embodiments of the present application is to provide a facial expression migration method and system based on a diffusion model and facial key points, which can migrate the facial expression of a source face image to a target face while retaining the identity information of the target face, and is applicable to target faces with different expressions to generate high-quality and flexible facial expression migration results.
[0006] In order to solve the above technical problems, this application is implemented as follows:
[0007] In a first aspect, an embodiment of the present application provides a facial expression migration method based on a diffusion model and facial key points, the method comprising:
[0008] Extract facial key points of facial expression images in public facial expression datasets, and construct training sample pairs including source face key points, target face images, and expression transfer results;
[0009] The text encoder of the original stable diffusion model is replaced by the Swin Transformer module, which is designed as an identity feature encoder to obtain an improved stable diffusion model.
[0010] According to the pre-trained neural network block of the improved stable diffusion model, a replica network block is copied, a zero convolution layer is added to the replica network block, and the replica network block is designed as an expression feature encoder to obtain a ControlNet architecture;
[0011] Building an expression migration model based on the improved stable diffusion model and the ControlNet architecture;
[0012] Based on the training sample pairs, training the expression transfer model using a control loss function and a time-step-dependent identity preservation loss function;
[0013] The trained expression transfer model is used to generate expression transfer results that retain identity features.
[0014] As an optional implementation manner of the first aspect of the present application, the steps of extracting facial key points of facial expression images in a public facial expression dataset, and constructing a training sample pair including source face key points, target face images, and expression transfer results include: collecting facial expression images under 7 basic different expressions of each subject in the public facial expression dataset, and the 7 basic different expressions are neutral, anger, disgust, fear, happiness, sadness, and surprise; using the facial expression image under a specific expression as the target face image in the training sample pair; using the facial expression images under other expressions other than the specific expression as the expression transfer results in the training sample pair; and extracting the facial key points of the facial expression image under the specific expression through the MTCNN model as the source face key points in the training sample pair.
[0015] As an optional implementation of the first aspect of the present application, a Swin Transformer module is used to replace the text encoder of the original stable diffusion model, which is designed as an identity feature encoder to obtain an improved stable diffusion model, including the following steps: the original stable diffusion model includes 12 encoder modules, 1 intermediate block, 12 decoder modules and a text encoder, each module includes multiple residual network layers and a Vision Transformer structure, the decoder module is connected to the encoder module through a jump connection, and the text encoder is used to encode text prompts into an embedding vector as a control condition for the output image of the original stable diffusion model; the Swin Transformer module is used to replace the text encoder as the backbone network of the identity feature encoder, and a window self-attention mechanism is introduced into the identity feature encoder to obtain an improved stable diffusion model.
[0016] As an optional implementation of the first aspect of the present application, the step of introducing a window self-attention mechanism in the identity feature encoder includes: designing a W-MSA operation layer and a SW-MSA operation layer in the identity feature encoder; the W-MSA operation layer is used to limit the calculation of self-attention within a local window, and the SW-MSA operation layer is used to perform cross-window attention calculations based on a moving window.
[0017] As an optional implementation of the first aspect of the present application, according to the pre-trained neural network block of the improved stable diffusion model, a replica network block is copied, a zero convolution layer is added to the replica network block, and it is designed as an expression feature encoder to obtain the step of obtaining the ControlNet architecture: the zero convolution layer is a 1x1 convolution layer whose weights and bias parameters are initialized to zero, and is used to connect the pre-trained neural network block and the replica network block.
[0018] As an optional implementation of the first aspect of the present application, in the step of training the expression transfer model based on the training sample pairs using a control loss function and an identity preservation loss function that depends on a time step, the control loss function is used to calculate the difference between the expression features generated by the expression transfer model and the target expression features, and the calculation formula of the control loss function is: ,in, represents the original real image, represents the time step, express The noise image obtained at the moment, represents the identity features extracted from the target face image, represents the expression features extracted from the facial key points of the source face image, Represents a noisy image The noise added in Represents the constructed expression transfer model, represents the network parameters of the model, Indicates the calculation of the square of the L2 norm, Express The maximum likelihood estimate of , Represents the calculated control loss function value.
[0019] As an optional implementation of the first aspect of the present application, in the step of training the expression transfer model based on the training sample pair using a control loss function and an identity preservation loss function that depends on the time step, the identity preservation loss function is used to calculate the difference between the identity features of the output image at different time steps in the expression sequence generated by the expression transfer model and the identity features of the original input image, and the calculation formula of the identity preservation loss function is: ,in, represents the current time step, The expression transfer model in the present invention is based on Noise image at time Predict the generated real image, represents the target face image in the training sample pair, represents the expression transfer result image in the training sample pair, represents the features extracted by the pre-trained face recognition model, Representation characteristics With features The square of the L2 distance between Represents the weight coefficient that depends on the time step Represents the computed time-step-dependent identity preserving loss value.
[0020] In a second aspect, an embodiment of the present application provides a facial expression transfer system based on a diffusion model and facial key points, the system comprising:
[0021] A data acquisition module is used to extract facial key points of facial expression images in a public facial expression dataset and construct training sample pairs including source facial key points, target facial images, and expression transfer results;
[0022] The identity feature extraction module is used to replace the text encoder of the original stable diffusion model with the Swin Transformer module, which is designed as an identity feature encoder to obtain an improved stable diffusion model;
[0023] An expression feature extraction module is used to copy the pre-trained neural network block of the improved stable diffusion model to obtain a replica network block, add a zero convolution layer to the replica network block, and design it as an expression feature encoder to obtain a ControlNet architecture;
[0024] A migration model building module, used to build an expression migration model according to the improved stable diffusion model and the ControlNet architecture;
[0025] A transfer model training module, for training the expression transfer model based on the training sample pairs using a control loss function and a time-step-dependent identity preservation loss function;
[0026] The migration result generation module is used to generate expression migration results that retain identity features through the trained expression migration model.
[0027] In a third aspect, an embodiment of the present application provides an electronic device, comprising a processor, a memory, and a program or instruction stored in the memory and executable on the processor, wherein the program or instruction, when executed by the processor, implements the steps of the method described in the first aspect.
[0028] In a fourth aspect, an embodiment of the present application provides a readable storage medium, on which a program or instruction is stored, and when the program or instruction is executed by a processor, the steps of the method described in the first aspect are implemented.
[0029] Compared with the prior art, the present invention proposes a facial expression transfer method based on a diffusion model and facial key points, which improves the text encoder of the original stable diffusion model into an original identity feature encoder, so as to more efficiently extract features containing identity information from the target face image as the identity control conditions of the generated model; then use the ControlNet architecture to build an original model for facial expression transfer based on the improved stable diffusion model, and the identity feature encoder extracts the identity features in the target face image as the identity control conditions of the improved stable diffusion model; the facial key points in the source face image are extracted by the MTCNN model as the expression control conditions of the improved stable diffusion model, and a The original sample pair construction technology for facial expression transfer constructs a training sample pair consisting of three parts: facial key points of the source face image, the target face image, and the expression transfer result, which is used to train the improved stable diffusion model; in the model training process, an original control loss function and an original identity preservation loss function dependent on the time step are used to train the improved stable diffusion model. The control loss function enables the model to learn to generate images under the guidance of two control conditions: identity features extracted from the target face image and expression features extracted from the facial key points of the source face image. The identity preservation loss function dependent on the time step ensures the identity consistency between the model-generated image and the target face image and the expression transfer effect of the generated image. The present invention can transfer the facial expression of the source face image to the target face, while retaining the identity information of the target face, generating high-quality facial expression transfer results, and is applicable to any of the seven basic different expressions of the target face, with flexibility and practical value. BRIEF DESCRIPTION OF THE DRAWINGS
[0030] Figure 1 is a flow chart of a facial expression transfer method based on a diffusion model and facial key points provided by a first embodiment of the present invention;
[0031] Figure 2 Schematic diagram of a sample pair construction technique for facial expression transfer in a first embodiment of the present invention;
[0032] Figure 3 It is a structural diagram of an identity feature encoder in the first embodiment of the present invention;
[0033] Figure 4 It is a structural diagram of the Swin Transformer module in the first embodiment of the present invention;
[0034] Figure 5 It is a schematic diagram of the principle of using the ControlNet architecture to construct a network block in the first embodiment of the present invention;
[0035] Figure 6A schematic diagram of the structure of the expression transfer model constructed in the first embodiment of the present invention;
[0036] Figure 7 A schematic diagram of the calculation principle of the time-step-dependent identity preservation loss function in the first embodiment of the present invention;
[0037] Figure 8 This is a schematic diagram of the method of step S05 of the first embodiment of the present invention;
[0038] Fig. 9 This is a schematic diagram of the method of step S06 of the first embodiment of the present invention;
[0039] Fig.10 It is a structural schematic diagram of a facial expression transfer system based on a diffusion model and facial key points provided by the second embodiment of the present invention. DETAILED DESCRIPTION
[0040] The following will be combined with the drawings in the embodiments of the present application to clearly and completely describe the technical solutions in the embodiments of the present application. Obviously, the described embodiments are part of the embodiments of the present application, not all of the embodiments. Based on the embodiments in the present application, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of this application.
[0041] The terms "first", "second", etc. in the specification and claims of the present application are used to distinguish similar objects, and are not used to describe a specific order or sequence. It should be understood that the data used in this way can be interchangeable under appropriate circumstances, so that the embodiments of the present application can be implemented in an order other than those illustrated or described here. In addition, the "and / or" in the specification and claims represents at least one of the connected objects, and the character " / " generally represents that the objects associated with each other are in an "or" relationship.
[0042] In order to illustrate the technical solution described in this application, a specific embodiment is provided below for illustration.
[0043] Example 1
[0044] See also Figure 1 , which is a flowchart of a facial expression migration method based on a diffusion model and facial key points proposed in the first embodiment of the present application. The proposed method includes S01 to S06.
[0045] Step S01: extracting facial key points of facial expression images in a public facial expression dataset, and constructing training sample pairs including source face key points, target face images, and expression transfer results.
[0046] In step S01 of the present invention, a training data set is constructed. First, facial expression images in three public facial expression data sets are used, including public Oulu data sets, public RaFD data sets and public TFED data sets, and 106 facial key points are extracted through the MTCNN model. These facial key points correspond to the key position information of eyebrows, eyes, nose, mouth and facial contour in the facial features, and can accurately describe the facial expression features in a facial expression image. The MTCNN model combines the ideas of multi-task learning and multi-stage detection, and works by cascading the three sub-networks of P-Net network, R-Net network and O-Net network, wherein the P-Net network performs preliminary facial candidate frames and facial key point position predictions on the input image, and the R-Net network further screens and optimizes the facial candidate frames and key point positions predicted by the P-Net network, and the O-Net network is further corrected and refined on the basis of the prediction results of the P-Net network and the R-Net network, and the final result is output.
[0047] After extracting facial key points, an original sample pair construction technology for facial expression transfer is used to construct a training sample pair consisting of source face key points, target face image, and expression transfer result. The three parts of the training sample pair are shown in formula (1):
[0048]
[0049] In formula (1), Represents a pair of constructed training samples; represents the facial key points extracted from the source face image; represents the target face image; Indicates the expression transfer result.
[0050] Specifically, for the facial expression images of each subject under 7 basic different expressions in the public facial expression dataset, these 7 basic different expressions are neutral ( neutral ),anger( anger ),disgust( disgust ),fear( fear ),hapiness( happy) ,sad( sadness ),surprise( surprise), the facial key points of the facial expression images under these 7 expressions are extracted by the MTCNN model. For each facial expression image, it is used as the target face image in the sample pair; for facial images of other expressions, it is used as the expression migration result in the sample pair, and its corresponding facial key points are used as the facial key points of the source face image in the sample pair. Taking the facial expression image of a subject's angry expression as an example, the process of constructing 6 pairs of sample pairs is shown in formula (2):
[0051]
[0052]
[0053]
[0054]
[0055]
[0056]
[0057] In formula (2), represents a set of training sample pairs constructed from facial expression images of a subject with angry expression, A facial expression image representing the subject's angry expression, which is the target face image in the training sample pair; , , , , , The facial expression images represent the subject’s neutral, disgust, fear, happiness, sadness, and surprise expressions, which are the expression transfer results in the training sample pairs. , , , , , They represent the facial key points extracted from the facial expression images of the subject with six expressions, namely neutral, disgust, fear, happiness, sadness, and surprise. They are the source face key points in the training sample pairs. For the facial expression images of a subject with seven basic different expressions, six pairs of sample pairs can be constructed for each facial expression image, that is, a total of 42 pairs of training sample pairs can be constructed from the facial expression images of a subject. The constructed sample pair set is shown in formula (3):
[0058]
[0059] In formula (3), represents a set of sample pairs constructed through the facial expression images of the subject; , , , , , , They represent sample pair sets constructed based on the facial expression images of the subject's neutral, angry, disgusted, fearful, happy, sad, and surprised expressions. The construction process of each set is shown in formula (2). Each set contains 6 sample pairs. The schematic diagram of the technique for constructing training sample pairs for facial expression transfer in step S01 of the present invention is shown in Figure 2 .
[0060] The original training sample pair construction technology for facial expression transfer in the present invention can make full use of facial expression images in public facial expression data sets to construct as many training sample pairs as possible. Compared with most existing expression transfer technologies that require neutral expression images as target facial images for transfer, the training sample pairs constructed in the present invention can train the model to use facial images with any expression as target facial images, and perform expression transfer based on the facial key points corresponding to the source facial images, so that the model can fully learn the complex mapping relationship between the source facial image, the target facial image and the expression transfer result, effectively improving the generalization and flexibility of the model in the facial expression transfer task, thereby achieving higher facial expression image generation quality.
[0061] Step S02: The text encoder of the original stable diffusion model is replaced by a Swin Transformer module, which is designed as an identity feature encoder to obtain an improved stable diffusion model.
[0062] In step S02 of the present invention, an original Expr-Diffusion model for facial expression migration is constructed. First, the present invention improves the text encoder of the original stable diffusion model into an original identity feature encoder for extracting identity features from the target face image as identity control conditions in the facial expression image generation process.
[0063] The original stable diffusion model is a pre-trained diffusion model based on the U-Net network architecture. Its network consists of 12 encoder blocks, 1 middle block, and 12 decoder blocks. Each module contains multiple residual network layers and visual converters. The encoder module extracts the features of the input image and gradually downsamples the input image to obtain low-dimensional intermediate features; the middle block further processes and transforms the intermediate features; the decoder module gradually upsamples the feature representation until it is restored to a high-resolution image. The decoder and encoder are connected through jump connections. In addition, the original stable diffusion model encodes text prompts into embedded vectors through the CLIP text encoder as a control condition for the model to generate images.
[0064] The present invention improves the text encoder in the original stable diffusion model into an original identity feature encoder for extracting identity features from target face images. The identity feature encoder uses the Swin Transformer module (a deep learning model based on the Transformer architecture) as the backbone network, which uses a window self-attention mechanism to divide the input image into multiple smaller windows and limit the self-attention calculation to the local window, thereby reducing the computational complexity from the original Reduced to , significantly reducing the amount of calculation, where Indicates the size of the entire image. Represents the size of the local window. The identity feature encoder includes 4 stages, mainly composed of a block partition layer (Patch Partition), a linear embedding layer (Linear Embedding), a block merging layer (Patch Merging), and a Swin Transformer module connected in series. The block partition layer (Patch Partition) is responsible for dividing the input target face image into image blocks of fixed size, the linear embedding layer (Linear Embedding) maps the image blocks to the feature space, and the block merging layer (Patch Merging) aggregates the local area of the feature map to reduce the spatial resolution of the feature map and increase the number of channels of the feature map. The structural diagram of the identity feature encoder in step S02 of the present invention is shown in Figure 3 .
[0065] The Swin Transformer module consists of two core components: Window based Multi-headSelf-Attention (W-MSA) and Shifted Window based Multi-headSelf-Attention (SW-MSA). The working principle of the Swin Transformer module is shown in formula (4):
[0066]
[0067]
[0068]
[0069]
[0070] In formula (4), Represents input features; Representation layer normalization layer (Layer Norm); represents a multi-layer perceptron; and They represent window multi-head self-attention operation and moving window multi-head self-attention operation respectively; express Features of the output of the operation layer; express Features of the output of the operation layer; and Both said The structural diagram of the SwinTransformer module in the identity feature encoder in step S02 of the present invention is shown in Figure 4 .
[0071] The present invention adopts the Swin Transformer module to construct the identity feature encoder, wherein the W-MSA operation layer limits the calculation of self-attention to the local window to significantly reduce the computational overhead; the SW-MSA operation layer performs cross-window attention calculation based on the moving window, which enhances the information transmission and interaction between different windows. The two cooperate with each other, so that the identity feature encoder in the present invention has a larger global receptive field and obtains richer image semantic information, so that it can better extract features containing identity information from the target face image as the identity control condition of the generation model, and improve the retention of the target face identity information during the expression migration process.
[0072] Step S03: According to the pre-trained neural network block of the improved stable diffusion model, a replica network block is copied, a zero convolution layer is added to the replica network block, and it is designed as an expression feature encoder to obtain a ControlNet architecture.
[0073] In step S03 of the present invention, the present invention constructs an original expression transfer model (Expr-Diffusion model) based on the improved stable diffusion model through the ControlNet architecture. and Represents two control conditions of the Expr-Diffusion model in the present invention, where Indicates identity characteristics, Indicates facial features, and The extraction method is shown in formula (5) and formula (6):
[0074]
[0075]
[0076] In formula (5), represents the target face image; An identity feature encoder constructed by the present invention is represented; represents the identity feature extracted from the target face image, which serves as the control condition for identity information in model generation. In formula (6), Represents the facial key points extracted from the source face image; The expression feature encoder constructed by the present invention is composed of four 4x4 convolutional layers; Represents the expression features extracted from the facial key points of the source face image, which serve as the control conditions for expression in model generation.
[0077] A pre-trained neural network block of the improved stable diffusion model is represented as , its feature processing process is shown in formula (7):
[0078]
[0079] In formula (7), Represents input features; Represents the parameters of the neural network block; Represents the output features. The parameters of the original pre-trained network block Lock it, that is, do not train it. At the same time, copy a trainable replica network block from the network block, and express its parameters as In the present invention, the replica network block receives the identity features extracted from the target face , expression features extracted from facial key points of source face images These two control conditions are used as input. Through these two control conditions, the output of the replica network block can be adjusted, while the output of the original pre-trained network block remains unchanged.
[0080] The original pre-trained network block and the trainable replica network block are connected through a zero convolution layer (ZeroConvolution). The zero convolution layer is a 1x1 convolution layer with weights and bias parameters initialized to zero. The ControlNet architecture in the present invention adds two zero convolution layers to each replica network block. The feature processing process after the pre-trained network block and the replica network block are connected through the zero convolution layer is shown in formula (8):
[0081]
[0082] In formula (8), Represents input features; Represents a pre-trained neural network block; Represents the parameters of the neural network block; Represents a trainable replica network block; Represents the parameters of the replica network block; and Represents two zero-convolutional layers in the ControlNet architecture; and Represent the parameters of the two zero convolutional layers respectively; The schematic diagram of the principle of using the ControlNet architecture to construct the network block in step S03 of the present invention is shown in Figure 5 ,in, Figure 5 (a) indicates that before applying the ControlNet architecture, Figure 5 (b) in the figure shows the application of ControlNet architecture.
[0083] Step S04: construct an expression transfer model based on the improved stable diffusion model and ControlNet architecture.
[0084] In step S04 of the present invention, the ControlNet architecture is applied to the improved stable diffusion model, all parameter weights in the improved stable model are locked, and for the 12 encoder modules and 1 intermediate block, the trainable replica network blocks are copied, and the output features of these replica network blocks are input into the zero convolution layer, and then the features obtained by the corresponding zero convolution layer processing are fused with the output of the intermediate block and decoder module corresponding to the stable diffusion model to obtain the final output result of the entire Expr-Diffusion model. The structural schematic diagram of the facial expression transfer model constructed in step S04 of the present invention is shown in Figure 6 At the beginning of training, since the weights and bias parameters of the zero convolution layer are all zero, the output of the entire network module is consistent with the output of the original pre-trained network block, that is, This design ensures that the network output of the Expr-Diffusion model in the present invention will not be affected by noise due to the control conditions of the trainable replicas and inputs in the early stage of training, and it can still maintain the image generation capability of the original pre-trained model, while also ensuring the stability of the network gradient in the early stage of training. As the training progresses, the parameters of the zero convolution layer will be gradually adjusted so that the identity features used as input and facial features It can control and guide the face images generated by the Expr-Diffusion model.
[0085] In summary, the Expr-Diffusion model in the present invention can flexibly use the facial key points extracted from the source face image and the target face image as control conditions while maintaining the original image generation capability of the stable diffusion model, and guide the model to generate a face image with the facial expression of the source face and retain the identity information of the target face, thereby achieving a high-quality expression migration effect. In addition, since the parameters of the original pre-trained network block are locked, its parameters do not participate in the gradient calculation and weight update, which improves the training speed of the Expr-Diffusion model in the present invention and reduces the computational consumption.
[0086] Step S05: Based on the training sample pairs, the expression transfer model is trained using a control loss function and a time-step-dependent identity preservation loss function.
[0087] In step S05 of the present invention, the constructed Expr-Diffusion model is trained. The core principle of the Expr-Diffusion model is divided into two stages: the forward process and the reverse process. , the forward process gradually adds noise to the image at a certain time step, and obtains The noise image at time is , until The essence of training the Expr-Diffusion model in the present invention is: given a set of inputs, including a noisy image , time step , identity features extracted from the target face image , expression features extracted from facial key points of source face images , so that the model learns to predict the noisy image based on this set of inputs The noise added The present invention refers to the model training loss function corresponding to this part of the training objectives as the control loss function. The calculation process of the control loss function is shown in formula (9):
[0088]
[0089] In formula (9), represents the original real image; represents the time step; express The noise image obtained at the moment; represents the identity features extracted from the target face image; represents the expression features extracted from the facial key points of the source face image; Represents a noisy image The noise added in represents the Expr-Diffusion model constructed by the present invention, Represents the network parameters of the model; Indicates calculation of the square of the L2 norm; Express The maximum likelihood estimate of ; represents the calculated control loss function value. The control loss function enables the model to learn to generate images under the guidance of two control conditions: identity features extracted from the target face image and expression features extracted from the facial key points of the source face image.
[0090] In addition to the control loss function, the present invention also adds an original identity preservation loss function that depends on the time step in the training process to ensure the identity consistency between the generated image and the target face image and the expression transfer effect of the generated image. In the prior art, the identity preservation loss is usually determined by calculating the gap between the features extracted from the target face image and the generated image by the pre-trained face recognition model. However, such a design method will lead to a decrease in the expression transfer effect of the generated image. The reason is that the features extracted by the pre-trained face recognition model usually include other facial features in addition to identity features. When the gap between the features extracted from the target face image and the generated image by the model is minimized as much as possible, the facial feature information in the generated image in addition to the identity information will also be close to the target face image, thereby resulting in a decrease in the expression transfer effect of the generated image. In response to this problem, combined with the situation that the target face image and the expression transfer result image in the sample pair used for training in the present invention belong to the same subject, the present invention designs an original identity preservation loss function that depends on the time step. The calculation process of the loss function is shown in formula (10):
[0091]
[0092] In formula (10), Indicates the current time step; Indicates that the Expr-Diffusion model in the present invention is based on Noise image at time Predict the generated real image; Represents the target face image in the training sample pair; represents the expression transfer result image in the training sample pair; Represents the features extracted by the pre-trained face recognition model; Representation characteristics With features The square of the L2 distance between them; represents the weight coefficient that depends on the time step, and its calculation method is ,in represents the current time step, represents the total time step; The calculated time-step-dependent identity preservation loss function value is shown in FIG. Figure 7 .
[0093] The time-step-dependent identity preservation loss function designed by the present invention is essentially calculated by interpolating between the squares of two L2 distances, which are the L2 distance between the features extracted by the pre-trained face recognition model from the generated image and the target face image, and the L2 distance between the features extracted by the pre-trained face recognition model from the generated image and the expression transfer result image. When , the Expr-Diffusion model predicts the denoised image from the pure noise image. , Focuses on shrinking pre-trained face recognition models from generated images With the target face image This enables the Expr-Diffusion model to learn to fully utilize the target face image The identity feature information in is used to generate, thereby ensuring the generated image and the target face image The identity consistency between hour, ,at this time Focuses on shrinking pre-trained face recognition models from generated images and expression transfer result image This makes the Expr-Diffusion model generate results The facial features of the image are close to the expression transfer result image , thereby ensuring the expression migration effect of the image generated by the Expr-Diffusion model. At the same time, due to the target face image in the training sample pair And expression transfer result image From the same subject, exist The generated image will not be destroyed. and the target face image In summary, the time-step-dependent identity-preserving loss function designed by the present invention can adjust the direction of learning prediction of the Expr-Diffusion model according to the time step during the training process, thereby ensuring the identity consistency between the model-generated image and the target face image and the expression migration effect of the generated image.
[0094] Based on the above two loss functions, the total loss function used in training the Expr-Diffusion model of the present invention is shown in formula (11):
[0095]
[0096] In formula (11), Represents the calculated control loss function value; represents the calculated time-step-dependent identity preservation loss function value; represents the weight coefficient; represents the calculated total loss function value. The method schematic diagram of step S05 of the present invention is shown in Figure 8 .
[0097] Step S06: Generate an expression transfer result that retains identity features through the trained expression transfer model.
[0098] In step S06 of the present invention, facial expression migration is performed using the trained Expr-Diffusion model. The target face image and the facial key points extracted from the source face image are input into the model, and the model generates a face image with the facial expression of the source face and retains the identity information of the target face, thereby realizing facial expression migration. The generation process is shown in formula (12):
[0099]
[0100] In formula (12), represents the target face image; represents the source face image; Represents the facial key points extracted by the MTCNN model in the first part of the present invention; represents the Expr-Diffusion model trained in the third part of the present invention; The method schematic diagram of step S06 of the present invention is shown in Fig. 9 .
[0101] In this embodiment, the present invention uses facial expression images from the public Oulu dataset, the public RaFD dataset, and the public TFED dataset to construct sample pairs. Four fifths of the sample pairs are used to train the Expr-Diffusion model, and the remaining one fifth of the sample pairs are used to test and evaluate the expression migration effect of the model. The optimizer used for model training is the Stochastic Gradient Descent (SGD) optimizer. batch_size The number of training rounds is 16, and a total of 200 rounds are trained on a server equipped with an NVIDIA RTX3090 graphics card. Among the sample pairs used for testing, 6 pairs of sample pairs are randomly selected to evaluate the expression transfer effect of the trained Expr-Diffusion model in the present invention.
[0102] The expression transfer results generated by the present invention all contain the expression in the source face image, while retaining the identity information of the target face, and the generated image quality is good. This shows that the facial expression transfer technology based on the diffusion model and facial key points proposed by the present invention can generate a high-quality facial image with the facial expression of the source face and retaining the identity information of the target face according to the input target face image and the facial key points extracted from the source face image, thereby realizing facial expression transfer.
[0103] In addition, for target face images with different kinds of basic expressions, the present invention can transfer the expressions in the source face image to the target face. This also shows that the training sample pairs constructed in the present invention can train the model to use facial images with expressions other than neutral expressions as target face images, and perform expression transfer based on the facial key points corresponding to the source face image, making the model more flexible in the facial expression transfer task.
[0104] Example 2
[0105] See also Fig.10 , shown is a structural schematic diagram of a facial expression migration system based on a diffusion model and facial key points proposed in the second embodiment of the present application, the system comprising:
[0106] The data acquisition module 100 is used to extract facial key points of facial expression images in a public facial expression data set, and construct training sample pairs including source facial key points, target facial images, and expression transfer results;
[0107] The identity feature extraction module 200 is used to replace the text encoder of the original stable diffusion model with a Swin Transformer module, and is designed as an identity feature encoder to obtain an improved stable diffusion model;
[0108] The expression feature extraction module 300 is used to copy the pre-trained neural network block of the improved stable diffusion model to obtain a replica network block, add a zero convolution layer to the replica network block, and design it as an expression feature encoder to obtain a ControlNet architecture;
[0109] A migration model building module 400, for building an expression migration model according to the improved stable diffusion model and the ControlNet architecture;
[0110] A transfer model training module 500, for training the expression transfer model based on the training sample pairs using a control loss function and a time-step-dependent identity preservation loss function;
[0111] The transfer result generation module 600 is used to generate an expression transfer result that retains identity features through the trained expression transfer model.
[0112] A facial expression migration system based on a diffusion model and facial key points in an embodiment of the present application may be a device, or a component, an integrated circuit, or a chip in a terminal. The device may be a mobile electronic device or a non-mobile electronic device. Exemplarily, the mobile electronic device may be a mobile phone, a tablet computer, a laptop computer, a PDA, an in-vehicle electronic device, a wearable device, an ultra-mobile personal computer (UMPC), a netbook, or a personal digital assistant (PDA), etc., and the non-mobile electronic device may be a server, a network attached storage (NAS), a personal computer (PC), etc., which is not specifically limited in the embodiment of the present application.
[0113] In the embodiment of the present application, a facial expression transfer system based on a diffusion model and facial key points can be a device having an operating system. The operating system can be an Android operating system, an iOS operating system, or other possible operating systems, which are not specifically limited in the embodiment of the present application.
[0114] The facial expression transfer system based on the diffusion model and facial key points provided in the embodiment of the present application can achieve Figure 1 In the method embodiment, each process of implementing the facial expression transfer method based on a diffusion model and facial key points will not be described here to avoid repetition.
[0115] Optionally, an embodiment of the present application also provides an electronic device, including a processor, a memory, and a program or instruction stored in the memory and executable on the processor. When the program or instruction is executed by the processor, each process of the above-mentioned embodiment of a facial expression migration method based on a diffusion model and facial key points is implemented, and the same technical effect can be achieved. To avoid repetition, it will not be repeated here.
[0116] An embodiment of the present application also provides a readable storage medium, on which a program or instruction is stored. When the program or instruction is executed by a processor, each process of the above-mentioned facial expression migration method embodiment based on a diffusion model and facial key points is implemented, and the same technical effect can be achieved. To avoid repetition, it will not be repeated here.
[0117] The processor is a processor in the electronic device described in the above embodiment. The readable storage medium includes a computer readable storage medium, such as a computer read-only memory (ROM), a random access memory (RAM), a magnetic disk or an optical disk.
[0118] It should be noted that, in this article, the terms "comprise", "include" or any other variant thereof are intended to cover non-exclusive inclusion, so that the process, method, article or device including a series of elements includes not only those elements, but also includes other elements not explicitly listed, or also includes elements inherent to such process, method, article or device. In the absence of further restrictions, the elements defined by the sentence "comprise one..." do not exclude the presence of other identical elements in the process, method, article or device including the element. In addition, it should be pointed out that the scope of the methods and devices in the embodiments of the present application is not limited to performing functions in the order shown or discussed, and may also include performing functions in a substantially simultaneous manner or in reverse order according to the functions involved, for example, the described method may be performed in an order different from that described, and various steps may also be added, omitted, or combined. In addition, the features described with reference to certain examples may be combined in other examples.
[0119] Through the description of the above implementation methods, those skilled in the art can clearly understand that the above-mentioned embodiment methods can be implemented by means of software plus a necessary general hardware platform, and of course by hardware, but in many cases the former is a better implementation method. Based on such an understanding, the technical solution of the present application, or the part that contributes to the prior art, can be embodied in the form of a software product, which is stored in a storage medium (such as ROM / RAM, a magnetic disk, or an optical disk), and includes a number of instructions for a terminal (which can be a mobile phone, a computer, a server, an air conditioner, or a network device, etc.) to execute the methods described in each embodiment of the present application.
[0120] The embodiments of the present application are described above in conjunction with the accompanying drawings, but the present application is not limited to the above-mentioned specific implementation methods. The above-mentioned specific implementation methods are merely illustrative and not restrictive. Under the guidance of the present application, ordinary technicians in this field can also make many forms without departing from the purpose of the present application and the scope of protection of the claims, all of which are within the protection of the present application.
Claims
1. A facial expression transfer method based on diffusion model and facial key points, characterized in that: The method comprises: Extract facial key points of facial expression images in public facial expression datasets, and construct training sample pairs including source face key points, target face images, and expression transfer results; The Swin Transformer module is used to replace the text encoder of the original stable diffusion model and is designed as an identity feature encoder to obtain an improved stable diffusion model, which includes: The original stable diffusion model includes 12 encoder modules, 1 intermediate block, 12 decoder modules and a text encoder, each module includes multiple residual network layers and a Vision Transformer structure, the decoder module is connected to the encoder module through a jump connection, and the text encoder is used to encode the text prompt into an embedding vector as a control condition for the output image of the original stable diffusion model; the Swin Transformer module is used to replace the text encoder as the backbone network of the identity feature encoder, and a window self-attention mechanism is introduced into the identity feature encoder to obtain an improved stable diffusion model; Wherein, a W-MSA operation layer and a SW-MSA operation layer are designed in the identity feature encoder; the W-MSA operation layer is used to limit the calculation of self-attention within a local window, and the SW-MSA operation layer is used to perform cross-window attention calculation based on a moving window; According to the pre-trained neural network block of the improved stable diffusion model, a replica network block is copied, a zero convolution layer is added to the replica network block, and the replica network block is designed as an expression feature encoder to obtain a ControlNet architecture; Building an expression migration model based on the improved stable diffusion model and the ControlNet architecture; Based on the training sample pairs, training the expression transfer model using a control loss function and a time-step-dependent identity preservation loss function; The trained expression transfer model is used to generate expression transfer results that retain identity features.
2. A facial expression transfer method based on a diffusion model and facial key points according to claim 1, characterized in that: The steps of extracting facial key points of facial expression images in a public facial expression dataset and constructing training sample pairs including source face key points, target face images, and expression transfer results include: Collect facial expression images of each subject under 7 basic different expressions in the public facial expression dataset, where the 7 basic different expressions are neutral, anger, disgust, fear, happiness, sadness, and surprise; The facial expression image under the specific expression is used as the target face image in the training sample pair; Using facial expression images under other expressions other than the specific expression as expression transfer results in the training sample pairs; The facial key points of the facial expression image under the specific expression are extracted through the MTCNN model as the source face key points in the training sample pair.
3. A facial expression transfer method based on a diffusion model and facial key points according to claim 1, characterized in that: According to the pre-trained neural network block of the improved stable diffusion model, a replica network block is copied, a zero convolution layer is added to the replica network block, and the expression feature encoder is designed to obtain the ControlNet architecture: The zero convolution layer is a 1x1 convolution layer whose weight and bias parameters are initialized to zero, and is used to connect the pre-trained neural network block and the replica network block.
4. A facial expression transfer method based on a diffusion model and facial key points according to claim 1, characterized in that: In the step of training the expression transfer model based on the training sample pairs using a control loss function and a time-step-dependent identity preservation loss function, the control loss function is used to calculate the difference between the expression features generated by the expression transfer model and the target expression features, and the calculation formula of the control loss function is: , in, represents the original real image, represents the time step, express The noise image obtained at the moment, represents the identity features extracted from the target face image, represents the expression features extracted from the facial key points of the source face image, Represents a noisy image The noise added in Represents the constructed expression transfer model, represents the network parameters of the model, Indicates the calculation of the square of the L2 norm, Express The maximum likelihood estimate of , Represents the calculated control loss function value.
5. A facial expression transfer method based on a diffusion model and facial key points according to claim 1, characterized in that: In the step of training the expression transfer model based on the training sample pairs using a control loss function and a time-step-dependent identity preservation loss function, the identity preservation loss function is used to calculate the difference between the identity features of the output image and the identity features of the original input image at different time steps in the expression sequence generated by the expression transfer model, and the calculation formula of the identity preservation loss function is: , in, represents the current time step, Indicates that the Expr-Diffusion model in the present invention is based on Noise image at time Predict the generated real image, represents the target face image in the training sample pair, represents the expression transfer result image in the training sample pair, represents the features extracted by the pre-trained face recognition model, Representation characteristics With features The square of the L2 distance between represents the weight coefficient that depends on the time step, Represents the computed time-step-dependent identity preserving loss value.
6. A facial expression transfer system based on diffusion model and facial key points, characterized in that: The system comprises: A data acquisition module is used to extract facial key points of facial expression images in a public facial expression dataset and construct training sample pairs including source facial key points, target facial images, and expression transfer results; The identity feature extraction module is used to replace the text encoder of the original stable diffusion model with a Swin Transformer module, and is designed as an identity feature encoder to obtain an improved stable diffusion model. The text encoder of the original stable diffusion model is replaced with a Swin Transformer module, and is designed as an identity feature encoder to obtain an improved stable diffusion model, specifically comprising: the original stable diffusion model includes 12 encoder modules, 1 intermediate block, 12 decoder modules and a text encoder, each module includes multiple residual network layers and a Vision Transformer structure, the decoder module is connected to the encoder module through a jump connection, and the text encoder is used to encode text prompts into an embedding vector as a control condition for the output image of the original stable diffusion model; the text encoder is replaced with a Swin Transformer module as the backbone network of the identity feature encoder, and a window self-attention mechanism is introduced into the identity feature encoder to obtain an improved stable diffusion model; wherein a W-MSA operation layer and a SW-MSA operation layer are designed in the identity feature encoder; the W-MSA operation layer is used to limit the calculation of self-attention within a local window, and the SW-MSA operation layer is used to perform cross-window attention calculation based on a moving window; An expression feature extraction module is used to copy the pre-trained neural network block of the improved stable diffusion model to obtain a replica network block, add a zero convolution layer to the replica network block, and design it as an expression feature encoder to obtain a ControlNet architecture; A migration model building module, used to build an expression migration model according to the improved stable diffusion model and the ControlNet architecture; A transfer model training module, for training the expression transfer model based on the training sample pairs using a control loss function and a time-step-dependent identity preservation loss function; The migration result generation module is used to generate expression migration results that retain identity features through the trained expression migration model.
7. An electronic device, characterized in that: The method comprises a processor, a memory and a program or instruction stored in the memory and executable on the processor. When the program or instruction is executed by the processor, the steps of a facial expression migration method based on a diffusion model and facial key points as described in any one of claims 1 to 5 are implemented.
8. A readable storage medium, characterized in that: The readable storage medium stores a program or instruction, and when the program or instruction is executed by the processor, the steps of the facial expression migration method based on a diffusion model and facial key points as described in any one of claims 1-5 are implemented.
Citation Information
Patent Citations
Natural image compressed sensing method based on potential diffusion model
CN116524048A
KR20250016987A