Identity attribute controllable face generation method based on double-flow diffusion model

By combining the dual-stream diffusion model and the U-Net structure, the problem of controlling fine-grained attributes in existing technologies is solved, achieving high-quality face generation with controllable identity attributes, and improving the realism of generated images and the generalization ability of the model.

CN120913264APending Publication Date: 2025-11-07INST OF COMPUTING TECH CHINESE ACAD OF SCI +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202511203691.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-08-27
Publication Date
2025-11-07

AI Technical Summary

Technical Problem

Existing face generation methods based on diffusion models struggle to effectively control fine-grained attributes such as hairstyle, expression, and age, and lack explicit supervision for textual attribute prompts, resulting in inaccurate generation results.

Method used

A dual-stream diffusion model is adopted, including a denoising branch diffusion model and a decoupled branch diffusion model. The ArcFace and CLIP algorithms are used for face identification and attribute encoding. Feature interaction is performed through the DCA layer. A U-Net model is constructed for image generation, and the model parameters are optimized through gradient descent.

Benefits of technology

It achieves precise control over fine-grained attributes, improves the realism and detail of generated images, enhances the model's generalization ability and training efficiency, and reduces training resource costs.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120913264A_ABST
    Figure CN120913264A_ABST
Patent Text Reader

Abstract

The invention discloses an identity attribute controllable face generation method based on a double-flow diffusion model, and relates to the technical field of face generation. According to the method, the double-flow diffusion model can be respectively processed and optimized aiming at different task targets by utilizing a double-branch structural design, interference and conflicts generated by different task requirements in a single model can be avoided, and each branch can more effectively learn and optimize feature representation related to own tasks; therefore, the control capability and the generation quality of the model on complex tasks are improved; and through the DCA layer, interference between the attribute information and the identity information is further reduced, and the generation quality and the condition control capability are improved.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of face generation, and particularly relates to an identity attribute controllable face generation method based on a double-flow diffusion model. BACKGROUND

[0002] Under the background of personalized generation, as one of the core multi-modal application scenarios, the demand for identity customization of faces is increasingly prominent. Subject-driven methods have emerged to control the generated image to be highly consistent with the target identity (ID) by using a small amount of reference images or encoders. At present, there are various methods that can effectively fuse reference face features in diffusion models to improve the identity fidelity of the generated results, but it is difficult to effectively control fine-grained attributes such as hairstyle, expression, and age.

[0003] Although the existing face customization generation method based on the diffusion model has good performance in maintaining the identity consistency of the reference face, it still has obvious deficiencies in accurately controlling fine-grained attributes such as hairstyle, expression, and age. The fundamental technical reason is that: on the one hand, the traditional face identity embedding (such as the face features extracted by CLIP or ArcFace) fails to effectively decouple the identity information and attribute information, resulting in excessive reliance on the original attribute distribution of the reference image during model generation; on the other hand, the lack of explicit supervision for text attribute prompts makes it difficult for the model to accurately extract and reproduce the attribute content expected by the user. SUMMARY

[0004] The purpose of the present application is to provide an identity attribute controllable face generation method based on a double-flow diffusion model to improve the above technical problems.

[0005] In order to achieve the above application purpose, the embodiments of the present application provide the following technical solutions:

[0006] An identity attribute controllable face generation method based on a double-flow diffusion model, comprising:

[0007] obtaining a face reference image of a face to be generated and face original attributes and face target attributes of the face; performing image recognition on the face reference image to obtain a face identity;

[0008] encoding the face identity, the face original attributes, and the face target attributes by using a text encoder to obtain a face identity embedding vector, a face original attribute embedding vector, and a face target attribute embedding vector;

[0009] construct a double-flow diffusion model; the double-flow diffusion model comprises a denoising branch diffusion model and a decoupling branch diffusion model; the decoupling branch diffusion model comprises a decoupling branch diffusion sub-model, a VAE decoder and an output layer connected in series; the output layer comprises an attribute classifier and a face recognition layer in parallel; the denoising branch diffusion model and the decoupling branch diffusion sub-model both adopt a U-Net model; the face recognition layer adopts an arcface recognition model;

[0010] input the face identity embedding vector, the face target attribute embedding vector and the second Gaussian noise into the double-flow diffusion model for face generation to obtain the identity attribute controllable face image.

[0011] Further, the U-Net model comprises an encoder, a bottleneck layer and a decoder; the encoder comprises Conv1 layer, Down-Block1 module, Down-Block2 module, Down-Block3 module and Down-Block4 module Mid-Block connected in series; the Down-Block1 module, the Down-Block2 module and the Down-Block3 module have the same structure and all comprise RB-ST1 layer, RB-ST2 layer and down-sampling layer connected in series; the RB-ST1 layer and the RB-ST2 layer both comprise ResBlock5 layer and SpatialTransformer layer; the Down-Block4 module comprises ResBlock1 layer and ResBlock2 layer connected in series;

[0012] the bottleneck layer is a Mid-Block module, which comprises ResBlock3 layer, SpatialTransformer layer and ResBlock4 layer connected in series;

[0013] the decoder comprises Up-Block1 module, Up-Block2 module, Up-Block3 module, Up-Block4 module and Conv2 layer connected in series; the Up-Block1 module comprises three CRB layers and an up-sampling layer connected in series; the CRB layer comprises Concat layer and ResBlock6 layer connected in series; the Up-Block2 module, the Up-Block3 module and the Up-Block4 module have the same structure and all comprise three CRB-ST layers and an up-sampling layer connected in series; the CRB-ST layer comprises Concat layer, ResBlock7 layer and SpatialTransformer layer connected in series;

[0014] the SpatialTransformer layer comprises SA layer and DCA layer connected in series; the DCA layer comprises CA layer in parallel.

[0015] Further, the image recognition of the face reference image adopts an ArcFace algorithm; and the text encoder adopts a CLIP text encoder.

[0016] Further, the training process of the double-flow diffusion model is as follows:

[0017] obtain a face training identity, a face reference training image, a face training target attribute, a face identity embedding training vector, a face original attribute embedding training vector, a face target attribute embedding training vector, and second Gaussian noise;

[0018] map the face reference training image by using a VAE encoder and add first Gaussian noise to obtain a training noisy latent variable;

[0019] input the training noisy latent variable, the face identity embedding training vector, and the face original attribute embedding training vector into a denoising branch diffusion model to obtain predicted noise;

[0020] input the second Gaussian noise, the face identity embedding training vector, and the face target attribute embedding training vector into a decoupling branch diffusion sub-model and repeat ten times to obtain a tenth training identity attribute controllable face initial image;

[0021] input the tenth training identity attribute controllable face initial image into a VAE decoder to obtain a training identity attribute controllable face initial image;

[0022] input the training identity attribute controllable face initial image into an attribute classifier and an arcface recognition model respectively for processing to obtain face predicted attributes and a predicted identity;

[0023] based on the predicted noise, the first Gaussian noise, the predicted identity, the face training identity, the face predicted attributes, and the face training target attribute, calculate a diffusion total loss; the diffusion total loss includes a diffusion loss, an identity (ID) loss, and an attribute loss;

[0024] based on the diffusion total loss, adjust parameters of the denoising branch diffusion model and the decoupling branch diffusion model by using a gradient descent method and an AdamW optimizer.

[0025] Further, the process of obtaining the predicted noise is as follows:

[0026] input the training noisy latent variable into a Conv1 layer to obtain training noisy latent variable convolutional features;

[0027] input the training noisy latent variable convolutional features, the face identity embedding training vector, and the face original attribute embedding training vector into a Down-Block1 module to obtain first image-text denoising key encoding features;

[0028] The first text and image denoising key coding feature module is sequentially input into a Down-Block2 module and a Down-Block3 module to obtain a second text and image denoising key coding feature, a second text and image denoising key initial coding feature, a second text and image denoising key reinforced coding feature, a third text and image denoising key coding feature, a third text and image denoising key initial coding feature, and a third text and image denoising key reinforced coding feature;

[0029] The third text and image denoising key coding feature module is input into a Down-Block4 module to obtain a fourth text and image denoising key coding feature and a fourth text and image denoising key initial coding feature.

[0030] The fourth text and image denoising key coding feature is input into a Mid-Block module to obtain a fifth text and image denoising key coding feature.

[0031] The fifth text and image denoising key coding feature, the fourth text and image denoising key coding feature, the fourth text and image denoising key initial coding feature, and the third text and image denoising key coding feature are input into an Up-Block1 module to obtain a sixth text and image denoising key decoding feature.

[0032] The sixth text and image denoising key decoding feature, the third text and image denoising key initial coding feature, the third text and image denoising key reinforced coding feature, and the second text and image denoising key coding feature are input into an Up-Block2 module to obtain a seventh text and image denoising key decoding feature.

[0033] The seventh text and image denoising key decoding feature, the second text and image denoising key initial coding feature, the second text and image denoising key reinforced coding feature, and the first text and image denoising key coding feature are input into an Up-Block3 module to obtain an eighth text and image denoising key decoding feature.

[0034] The eighth text and image denoising key decoding feature, the first text and image denoising key initial coding feature, the first text and image denoising key reinforced coding feature, and the training noisy latent variable convolution feature are input into an Up-Block4 module to obtain a ninth text and image denoising key decoding feature.

[0035] The ninth text and image denoising key decoding feature is input into a Conv2 layer to obtain a predicted noise.

[0036] Further, the Down-Block1 module comprises:

[0037] The training noisy latent variable convolution feature is input into a ResBlock5 layer to obtain a training noisy latent variable convolution time feature.

[0038] The training noise-added latent variable convolution time feature, the face identity embedding training vector, and the face original attribute embedding training vector are input into the SpatialTransformer layer to obtain first graph-text denoising key initial encoding features;

[0039] The first graph-text denoising key initial encoding features, the face identity embedding training vector, and the face original attribute embedding training vector are sequentially input into the ResBlock5 layer and the SpatialTransformer layer of the RB-ST2 layer to obtain first graph-text denoising key reinforced encoding features;

[0040] The first graph-text denoising key reinforced encoding features are input into a down-sampling layer to obtain first graph-text denoising key encoding features.

[0041] Further, the formula corresponding to the SpatialTransformer is:

[0042] K e =W k e;

[0043] V e =W v e;

[0044]

[0045] wherein, respectively represent the key vector and the value vector corresponding to the face identity embedding training vector, represents the first graph-text denoising key initial encoding features, and attn(·) represents an attention mechanism, respectively represent the key vector and the value vector corresponding to the face original attribute embedding training vector, K e , V e respectively represent the key vector and the value vector, W v , W k respectively represent the weight matrix corresponding to the K value and the V value, softmax(·) represents a maximum value function, and (·) T represents a transposed matrix, d represents the dimension corresponding to K e , and e represents an input feature, i.e., the face identity embedding training vector or the face original attribute embedding training vector.

[0046] Further, the calculation of the diffusion total loss comprises:

[0047] Based on the formula:

[0048]

[0049] the diffusion loss

[0050] Based on the formula:

[0051]

[0052] Get identity loss

[0053] Based on the formula:

[0054]

[0055] Get attribute loss

[0056] Based on the formula:

[0057]

[0058] Get diffusion total loss Wherein, represents expectation, epsilon represents the first Gaussian noise, epsilon theta (·) represents the U-Net model, z t represents the training noise latent variable of the t-th time step, ||·|| represents the norm, phi (·) represents the ArcFace model, log (·) represents the logarithmic function with natural constant as the base, sum (·) represents the summation function, represents the training identity attribute controllable face initial image, b i represents the i-th face training target attribute, represents the i-th face target attribute classifier, n represents the total number of attributes, G (·) represents the decoupling branch diffusion sub-module, represents the face reference training image, lambda att , lambda id respectively represent the attribute weight and the identity weight.

[0059] The beneficial effects of the present application are:

[0060] 1. The double-branch structure design of the present application enables the double-flow diffusion model to process and optimize different task targets respectively, which can avoid the interference and conflict caused by different task requirements in a single model, and each branch can more effectively learn and optimize the feature representation related to its own task, thereby improving the control ability and generation quality of the model for complex tasks.

[0061] 2. The double-branch structure of the present application gives the model flexibility in different task scenarios; when facing diversified inputs (such as different types of images, text descriptions, etc.), the two branches can process the input according to their respective task characteristics, and then integrate the processing results, so that the model can adapt to more different types of tasks and data, improving the generalization ability of the model.

[0062] 3. The DCA layer replaces the traditional CA layer, enabling more detailed feature interaction in different feature dimensions and spatial positions, while focusing on the mutual relationship between different modalities or different levels of features. Compared with the CA layer, it can mine more rich context information, help the model to capture the long-distance dependence between features more accurately, better combine the global structural features and local detailed features of the image, and generate images with more details and realism.

[0063] 4. Since the double branch decomposes the complex task, the optimization target of each branch is more clear and single, which can reduce the conflict and interference during parameter update in the training process, making the model more easily converge to a better local optimal solution. At the same time, the efficient feature interaction mechanism of DCA helps to propagate gradient information faster, speed up the training process of the model, reduce the time and resource cost required for training, and improve the training efficiency of the model. BRIEF DESCRIPTION OF DRAWINGS

[0064] In order to more clearly illustrate the technical solutions of the embodiments of the present application, the following will briefly introduce the drawings needed to be used in the embodiments. It should be understood that the following drawings only show some embodiments of the present application, and therefore should not be regarded as a limitation on the scope. For those skilled in the art, other related drawings can also be obtained without creative labor on the basis of these drawings.

[0065] Figure 1 The method flowchart in embodiment 1 of the present application;

[0066] Figure 2 The double-flow diffusion model structure diagram in embodiment 1 of the present application;

[0067] Figure 3 The U-Net model structure diagram in embodiment 1 of the present application;

[0068] Figure 4 The U-Net model detail diagram in embodiment 1 of the present application;

[0069] Figure 5 The training process diagram of the double-flow diffusion model in embodiment 1 of the present application;

[0070] Figure 6 The method flowchart in embodiment 2 of the present application;

[0071] Figure 7 The method flowchart in embodiment 3 of the present application. DETAILED DESCRIPTION

[0072] The technical solutions in the embodiments of the present application will be described clearly and completely below with reference to the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. The components of the embodiments of the present application described and shown in the accompanying drawings can be arranged and designed in various different configurations. Therefore, the following detailed description of the embodiments of the present application provided in the accompanying drawings is not intended to limit the scope of the claimed present application, but only represents selected embodiments of the present application. Based on the embodiments of the present application, all other embodiments obtained by those skilled in the art without creative work fall within the scope of the present application.

[0073] Embodiment 1

[0074] Please refer to Figure 1 The embodiment provides a face generation method based on a double-flow diffusion model. Figure 1 The execution subject of the method can be a software and / or a hardware device. The execution subject of the present application can include but is not limited to at least one of the following: a user equipment, a network equipment, and the like. The user equipment can include but is not limited to a computer, a smart phone, a personal digital assistant (PDA), and the above-mentioned electronic devices, and the like. The network equipment can include but is not limited to a single network server, a server group composed of multiple network servers, or a cloud composed of a large number of computers or network servers based on cloud computing. The cloud computing is a kind of distributed computing, which is a super virtual computer composed of a loose coupled computer group. The embodiment does not make any limitation in this regard.

[0075] A face generation method based on a double-flow diffusion model includes:

[0076] S1, obtaining a face reference image of a face to be generated, and face original attributes and face target attributes; performing image recognition on the face reference image to obtain a face identity; the image recognition adopts an ArcFace algorithm.

[0077] S2, encoding the face identity, the face original attributes and the face target attributes by using a text encoder to obtain a face identity embedding vector e id , a face original attribute embedding vector e b and a face target attribute embedding vector e a ; the text encoder adopts a CLIP text encoder to convert an input text description into a text embedding vector. This module is responsible for encoding the attribute description and the identity description respectively, and providing condition information for the diffusion model.

[0078] S3, constructing a double-flow diffusion model; for example, Figure 2As shown, the double-flow diffusion model includes a denoising branch diffusion model and a decoupling branch diffusion model; the decoupling branch diffusion model includes a decoupling branch diffusion sub-model, a VAE decoder and an output layer connected in series; the output layer includes an attribute classifier and an arcface recognition model in parallel; the denoising branch diffusion model and the decoupling branch diffusion sub-model both adopt a U-Net model; the face recognition layer adopts an arcface recognition model;

[0079] As shown in Figure 3 and Figure 4 the U-Net model includes an encoder and a decoder; the encoder includes Conv1 layer, Down-Block1 module, Down-Block2 module, Down-Block3 module and Down-Block4 module connected in series; the Down-Block1 module, Down-Block2 module and Down-Block3 module have the same structure and all include RB-ST1 layer, RB-ST2 layer and down-sampling layer connected in series; the RB-ST1 layer and the RB-ST2 layer both include ResBlock5 layer and SpatialTransformer layer; the Down-Block4 module includes ResBlock1 layer and ResBlock2 layer connected in series;

[0080] the Mid-Block module includes ResBlock3 layer, SpatialTransformer layer and ResBlock4 layer connected in series;

[0081] the decoder includes Up-Block1 module, Up-Block2 module, Up-Block3 module, Up-Block4 module and Conv2 layer connected in series; the Up-Block1 module includes three CRB layers and an up-sampling layer connected in series; the CRB layer includes Concat layer and ResBlock6 layer connected in series; the Up-Block2 module, Up-Block3 module and Up-Block4 module have the same structure and all include three CRB-ST layers and an up-sampling layer connected in series; the CRB-ST layer includes Concat layer, ResBlock7 layer and SpatialTransformer layer connected in series;

[0082] the SpatialTransformer layer includes SA layer (self-attention layer) and DCA layer connected in series; the DCA layer includes CA layer (cross-attention layer) in parallel.

[0083] Among them, in the ResBlock1 layer, the ResBlock5 layer, the ResBlock6 layer and the ResBlock7 layer, the time step is introduced, the time step information is coded into a high-dimensional vector, so that the U-Net model can utilize the time dependence in the time series data, and can represent the progress of the generation process (such as the gradual conversion from noise to clear image). Each SpatialTransformer layer has a TextEmbedding, which maps the input discrete text data (such as words or sentences) into a high-dimensional continuous vector space, which can capture semantic information and guide the generation process of the U-Net model, that is, generate images according to text descriptions.

[0084] The training process of the double-flow diffusion model is as follows:

[0085] S3-1, obtaining a face training identity, a face reference training image, a face training target attribute, and a face identity embedding training vector, a face original attribute embedding training vector, a face target attribute embedding training vector and a second Gaussian noise;

[0086] S3-2, mapping the face reference training image by using the VAE encoder and adding the first Gaussian noise to obtain a training noisy latent variable; wherein the encoder in S3-2 is a VAE encoder.

[0087] The face reference training image is mapped by the VAE encoder, and the high-dimensional face reference training image is mapped to a low-dimensional latent space to obtain a corresponding latent variable mapping.

[0088] The latent variable mapping obtained by the VAE encoder is "clean" (the image information has been compressed and conforms to the latent space distribution), but the denoising branch of the diffusion model needs to process "noisy latent variables", so the first Gaussian noise needs to be injected on the basis of the latent variable mapping to simulate the initial state of the diffusion process. Gaussian noise is introduced to the latent variable mapping to obtain a noisy latent variable.

[0089] S3-3, inputting the training noisy latent variable, the face identity embedding training vector and the face original attribute embedding training vector into the denoising branch diffusion model to obtain a predicted noise;

[0090] Since the denoising branch diffusion model and the decoupling branch diffusion sub-model both use the U-Net model, the training processes of the two are exactly the same, so one branch diffusion model is selected for description.

[0091] As shown in Figure 5 S3-3 includes:

[0092] S3-3-1, input the training noisy latent variable into the Conv1 layer to obtain a training noisy latent variable convolution feature, and the output dimension is 320x64x64;

[0093] In the decoupling branch diffusion sub-model, the input of the Conv1 layer is the second Gaussian noise.

[0094] S3-3-2, input the training noisy latent variable convolution feature, the face identity embedding training vector, and the face original attribute embedding training vector into the Down-Block1 module to obtain a first graph-text denoising key encoding feature (320x32x32), a first graph-text denoising key initial encoding feature (320x64x64), and a first graph-text denoising key reinforced encoding feature (320x64x64).

[0095] The specific process of S3-3-2 is as follows:

[0096] P1, input the training noisy latent variable convolution feature, the face identity embedding training vector, and the face original attribute embedding training vector into the RB-ST1 layer to obtain the first graph-text denoising key initial encoding feature.

[0097] The corresponding process in the RB-ST1 layer is as follows:

[0098] P1-1, input the training noisy latent variable convolution feature into the ResBlock5 layer and embed the time step to obtain a training noisy latent variable convolution time feature (320x64x64).

[0099] P1-2, input the training noisy latent variable convolution time feature, the face identity embedding training vector, and the face original attribute embedding training vector into the SpatialTransformer layer to obtain the first graph-text denoising key initial encoding feature.

[0100] In the SpatialTransformer layer, the training noisy latent variable convolution time feature is input into the SA layer, the key feature is extracted by using the self-attention mechanism, the weight score of the key feature is improved, the training noisy latent variable key feature is obtained, and the training noisy latent variable key feature is used as the Q value (query vector) of two CA layers; the face identity embedding training vector is used as the K value (key vector) and the V value (value vector), and is input into a CA layer together with the Q value to obtain the corresponding person identity encoding feature; the face original attribute embedding training vector is used as the K value and the V value, and is input into another CA layer together with the Q value to obtain the corresponding person attribute encoding feature; the person identity encoding feature and the person attribute encoding feature are merged to obtain the first graph-text denoising key initial encoding feature.

[0101] Therefore, the formula corresponding to the CA layer in the SpatialTransformer layer is as follows:

[0102] K e = W k e;

[0103] V e = W v e;

[0104]

[0105] wherein, respectively represent the key vector and the value vector corresponding to the face identity embedding training vector, represents the first image-text denoising key initial encoding feature, attn(·) represents an attention mechanism, respectively represent the key vector and the value vector corresponding to the face original attribute embedding training vector, K e , V e respectively represent the key vector and the value vector, W v , W k respectively represent the weight matrix corresponding to the K value and the V value, softmax(·) represents a maximum value function, (·) T represents a transpose matrix, d represents the dimension corresponding to K e , and e represents an input feature, i.e., a face identity embedding training vector or a face original attribute embedding training vector.

[0106] In the embodiment, the processing procedures of each SpatialTransformer layer are the same. The application further reduces the interference between attribute information and identity information through the DCA layer, and improves the generation quality and conditional control ability.

[0107] P2, input the first image-text denoising key initial encoding feature, the face identity embedding training vector, and the face original attribute embedding training vector into the RB-ST2 layer to obtain a first image-text denoising key reinforced encoding feature; the RB-ST2 layer and the RB-ST1 layer have the same structure, and therefore the processing procedure of the RB-ST2 layer is not described in detail.

[0108] P3, input the first image-text denoising key reinforced encoding feature into a down-sampling layer to obtain a first image-text denoising key encoding feature.

[0109] S3-3-3, the first image-text denoising key encoding feature module is input into the RB-ST3 layer according to the following formula: Figure 3The data flow in the third text and image denoising key coding feature module is sequentially input to the Down-Block2 module and the Down-Block3 module, and second text and image denoising key coding features (640x16x16), second text and image denoising key initial coding features (640x32x32), second text and image denoising key reinforced coding features (640x32x32), third text and image denoising key coding features (1280x8x8), third text and image denoising key initial coding features (1280x16x16), and third text and image denoising key reinforced coding features (1280x16x16) are obtained, respectively.

[0110] The processing procedures of the Down-Block2 module and the Down-Block3 module are the same as those of the Down-Block1 module, and therefore are not described in detail. The second text and image denoising key coding features output by the Down-Block2 module are input data of the Down-Block3 module.

[0111] S3-3-4, the third text and image denoising key coding feature module is input to the Down-Block4 module to obtain fourth text and image denoising key coding features (1280x8x8) and fourth text and image denoising key initial coding features (1280x8x8).

[0112] Specifically, the third text and image denoising key coding feature module is input to the ResBlock1 layer, and a time step is introduced to obtain the fourth text and image denoising key initial coding features. The fourth text and image denoising key initial coding features are input to the ResBlock2 layer to obtain the fourth text and image denoising key coding features.

[0113] S3-3-5, the fourth text and image denoising key coding features are input to the Mid-Block module to obtain fifth text and image denoising key coding features (1280x8x8).

[0114] The fourth text and image denoising key coding features are input to the ResBlock3 layer to obtain text and image denoising convolution coding features (1280x8x8). The text and image denoising convolution coding features are input to the SpatialTransformer layer to obtain text and image denoising convolution key coding features (1280x8x8). The text and image denoising convolution key coding features are input to the ResBlock4 layer to obtain the fifth text and image denoising key coding features.

[0115] S3-3-6, the fifth text and image denoising key coding features, the fourth text and image denoising key coding features, the fourth text and image denoising key initial coding features, and the third text and image denoising key coding features are input to the Up-Block1 module to obtain sixth text and image denoising key decoding features (1280x16x16).

[0116] Specifically, the fourth graph denoising key encoding feature and the fifth graph denoising key encoding feature are input into a first CRB layer in the Up-Block1 module to obtain a first graph denoising key time decoding feature (1280x8x8); the first graph denoising key time decoding feature and the fourth graph denoising key initial encoding feature are input into a second CRB layer to obtain a second graph denoising key time decoding feature (1280x8x8); the second graph denoising key time decoding feature and the third graph denoising key encoding feature are input into a third CRB layer to obtain a third graph denoising key time decoding feature (1280x8x8); and the third graph denoising key time decoding feature is input into an up-sampling layer to obtain a sixth graph denoising key decoding feature (1280x16x16).

[0117] In each CRB layer, taking the first CRB layer as an example, the fourth graph denoising key encoding feature and the fifth graph denoising key encoding feature are input into a Concat layer to obtain a spliced graph denoising key encoding feature (2560x8x8); the spliced graph denoising key encoding feature is input into a ResBlock6 layer, and a time step is introduced to obtain a first graph denoising key time decoding feature.

[0118] S3-3-7, input the sixth graph denoising key decoding feature, the third graph denoising key initial encoding feature, the third graph denoising key reinforced encoding feature and the second graph denoising key encoding feature into the Up-Block2 module to obtain a seventh graph denoising key decoding feature (1280x32x32);

[0119] The S3-3-7 includes:

[0120] T1, input the sixth graph denoising key decoding feature and the third graph denoising key initial encoding feature into a first CRB-ST layer to obtain a graph denoising key convolution initial feature (1280x16x16);

[0121] In each CRB-ST layer, taking the first CRB-ST layer as an example, T1 includes:

[0122] Input the sixth graph denoising key decoding feature and the third graph denoising key initial encoding feature into a Concat layer to obtain a spliced graph denoising key decoding feature (2560x16x16);

[0123] Input the spliced graph denoising key decoding feature into a ResBlock7 layer, and introduce a time step to obtain a graph denoising key-time decoding feature (1280x16x16);

[0124] The text and image denoising key-time decoding feature, the face identity embedding training vector and the face original attribute embedding training vector are input into a SpatialTransformer layer to obtain a text and image denoising key convolution initial feature.

[0125] T2, the text and image denoising key convolution initial feature and the third text and image denoising key reinforced coding feature are input into a second CRB-ST layer to obtain a text and image denoising key convolution intermediate feature (1280x16x16);

[0126] T3, the text and image denoising key convolution intermediate feature and the second text and image denoising key coding feature are input into a third CRB-ST layer to obtain a text and image denoising key convolution final feature (1280x16x16);

[0127] T4, the text and image denoising key convolution final feature is input into an up-sampling layer to obtain a seventh text and image denoising key decoding feature.

[0128] S3-3-8, the seventh text and image denoising key decoding feature, the second text and image denoising key initial coding feature, the second text and image denoising key reinforced coding feature and the first text and image denoising key coding feature are input into an Up-Block3 module to obtain an eighth text and image denoising key decoding feature (640x64x64).

[0129] S3-3-9, the eighth text and image denoising key decoding feature, the first text and image denoising key initial coding feature, the first text and image denoising key reinforced coding feature and the training noisy latent variable convolution feature are input into an Up-Block4 module to obtain a ninth text and image denoising key decoding feature (320x64x64);

[0130] S3-3-10, the ninth text and image denoising key decoding feature is input into a Conv2 layer to obtain a predicted noise (4x64x64).

[0131] In Figure 5 the Gaussian noise in the formula is a second Gaussian noise of the embodiment.

[0132] S3-4, the second Gaussian noise, the face identity embedding training vector and the face target attribute embedding training vector are input into a decoupling branch diffusion sub-model to obtain a first training identity attribute controllable face initial image;

[0133] S3-5, the step S3-4 is repeated until the number of repetitions reaches 10 to obtain a tenth training identity attribute controllable face initial image;

[0134] S3-6, the tenth training identity attribute controllable face initial image is input into a VAE decoder to obtain a training identity attribute controllable face initial image;

[0135] S3-7, input the training identity attribute controllable face initial image into the attribute classifier and the arcface recognition model respectively for processing to obtain face predicted attributes and predicted identities;

[0136] The S3-7 includes:

[0137] The training identity attribute controllable face initial image is input into the attribute classifier to predict the face attributes of the training identity attribute controllable face initial image.

[0138] The training identity attribute controllable face initial image is input into the arcface recognition model for face recognition to obtain the corresponding predicted face identity. The arcface recognition model is a classical face recognition algorithm, and the full name is "Additive Angular Margin Loss for Deep Face Recognition". The core innovation is to optimize the intra-class aggregation and inter-class separation of feature vectors through angular distance, which significantly improves the precision and robustness of face recognition.

[0139] S3-8, based on the predicted noise, the first Gaussian noise, the predicted identity, the face training identity, the face predicted attribute, and the face target attribute, a diffusion total loss is calculated; the diffusion total loss includes a diffusion loss, an identity loss, and an attribute loss; the diffusion loss adopts a mean square error loss function;

[0140] The S3-8 includes:

[0141] S3-8-1, based on the predicted noise and the first Gaussian noise, a diffusion loss is calculated The corresponding formula is:

[0142]

[0143] Wherein, represents expectation, and ε represents the first Gaussian noise, and ε θ (·) represents a U-Net model, and the output is a predicted noise, z t represents the training noise latent variable at the t-th time step, and ||·|| represents a norm.

[0144] S3-8-2, based on the training identity attribute controllable face image and the face reference training image, an identity loss is calculated The corresponding formula is:

[0145]

[0146] Wherein, φ(·) represents an ArcFace model.

[0147] S3-8-3, based on the predicted attribute of the training identity attribute controllable face image and the face target attribute, an attribute loss is calculated The corresponding formula is:

[0148]

[0149] wherein, log(·) represents a logarithmic function with a natural constant as a base number, ∑(·) represents a summation function, represents a training initial image of an identity attribute controllable face, b i represents the i-th face target attribute, represents a face predicted attribute of the i-th face target attribute (output of the attribute classifier), n represents a total number of attributes, and G(·) represents a decoupling branch diffusion sub-model, represents a face training image.

[0150] S3-8-4, based on the diffusion loss, the identity loss and the attribute loss, a total diffusion loss is calculated by weighted summation The corresponding formula is:

[0151]

[0152] wherein, λ att , λ id respectively represent an attribute weight and an identity weight. In the embodiment, λ att , λ id are both 0.01.

[0153] S3-9, based on the total diffusion loss, the parameters of the denoising branch diffusion model and the decoupling branch diffusion model are adjusted by the gradient descent method and the AdamW optimizer, and the training of the double-flow diffusion model is completed.

[0154] In the parameter adjustment process, the denoising branch diffusion model and the decoupling branch diffusion model are adjusted at the same time each time the gradient is descended, and the two diffusion models share parameters, and the parameter update is synchronized to take effect on the denoising branch and the decoupling branch in the same parameter memory space during training.

[0155] S4, the face identity embedding vector, the face target attribute embedding vector and the second Gaussian noise are input to the diffusion model for face generation, and an identity attribute controllable face image is obtained.

[0156] Embodiment 2:

[0157] Three different reference images f1, f2 and f3 are selected, and corresponding face target attributes are formulated, that is, the skin of the person in f1 is whitened, the person in f2 is made younger, and the person in f3 is made male. An identity attribute controllable face generation method based on a double-flow diffusion model is adopted as a method (this method), PhotoMaker and InstantID are respectively used to process the three different reference images, and the results are as followsFigure 6 As shown, PhotoMaker and InstantID are difficult to maintain identity consistency and accurately control facial attributes at the same time, while the present application can well maintain identity consistency and accurately control facial attributes.

[0158] Embodiment 3:

[0159] Two different reference images f4 and f5 are selected, and the corresponding face target attributes are formulated, that is, the person in f5 is added with a beard, and the person in f6 is changed to a male. The present method and the diffusion model without decoupling branch are used to process the two different reference images, and the results are as shown in Figure 7 As shown, the present method is closer to the input in face basic feature restoration, can better fit the attribute requirements such as beard and male, and retains the original person's expression and feature logic; while the image generated by the diffusion model without decoupling branch has obvious face feature deformation, poor attribute fitting degree, and cannot accurately restore the input set features and attributes, losing the ability of accurate control of facial attributes. It can be seen that the present method performs better in image feature restoration and attribute adaptation.

[0160] It can be seen that the present application is significantly better than the existing method in face attribute accuracy, and at the same time, achieves a performance comparable to the current advanced technology level in identity consistency. In addition, the technical solution proposed by the present application can simultaneously realize multi-attribute joint control and zero-shot attribute control, effectively improving the practicability of the model in the generalization scene under the premise of maintaining the clarity and authenticity of the generated image.

[0161] The above only describes the preferred embodiments of the present application and is not used to limit the present application. For those skilled in the art, the present application can have various changes and variations. Any modification, equivalent replacement, improvement, etc. made within the spirit and principles of the present application shall be included in the protection scope of the present application.

[0162] The above only describes the preferred embodiments of the present application and is not used to limit the present application. For those skilled in the art, the present application can have various changes and variations. Any modification, equivalent replacement, improvement, etc. made within the spirit and principles of the present application shall be included in the protection scope of the present application. The protection scope of the present application shall be subject to the protection scope of the claims.

Claims

1. A method for identity-controllable face generation based on a double-flow diffusion model, characterized in that, The method comprises the following steps: obtaining a face reference image of a face to be generated and original attributes and target attributes of the face; performing image recognition on the face reference image to obtain a face identity; encoding the face identity, the original attributes and the target attributes of the face by using a text encoder to obtain a face identity embedding vector, a face original attribute embedding vector and a face target attribute embedding vector; constructing a double-flow diffusion model; the double-flow diffusion model comprises a denoising branch diffusion model and a decoupling branch diffusion model; the decoupling branch diffusion model comprises a decoupling branch diffusion sub-model, a VAE decoder and an output layer connected in series; the output layer comprises an attribute classifier and a face recognition layer connected in parallel; the denoising branch diffusion model and the decoupling branch diffusion sub-model both adopt a U-Net model; the face recognition layer adopts an arcface recognition model; inputting the face identity embedding vector, the face target attribute embedding vector and second Gaussian noise into the double-flow diffusion model to generate a face, and obtaining an identity attribute controllable face image.

2. The identity attribute controllable face generation method based on a double-flow diffusion model according to claim 1, characterized in that, The U-Net model comprises an encoder, a bottleneck layer and a decoder; the encoder comprises a Conv1 layer, a Down-Block1 module, a Down-Block2 module, a Down-Block3 module and a Down-Block4 module connected in series; the Down-Block1 module, the Down-Block2 module and the Down-Block3 module have the same structure and all comprise an RB-ST1 layer, an RB-ST2 layer and a down-sampling layer connected in series; the RB-ST1 layer and the RB-ST2 layer both comprise a ResBlock5 layer and a SpatialTransformer layer; the Down-Block4 module comprises a ResBlock1 layer and a ResBlock2 layer connected in series; the bottleneck layer is a Mid-Block module comprising a ResBlock3 layer, a SpatialTransformer layer and a ResBlock4 layer connected in series; the decoder comprises an Up-Block1 module, an Up-Block2 module, an Up-Block3 module, an Up-Block4 module and a Conv2 layer connected in series; the Up-Block1 module comprises three CRB layers and an up-sampling layer connected in series; the CRB layer comprises a Concat layer and a ResBlock6 layer connected in series; the Up-Block2 module, the Up-Block3 module and the Up-Block4 module have the same structure and all comprise three CRB-ST layers and an up-sampling layer connected in series; the CRB-ST layer comprises a Concat layer, a ResBlock7 layer and a SpatialTransformer layer connected in series; the SpatialTransformer layer comprises a SA layer and a DCA layer connected in series; the DCA layer comprises a CA layer.

3. The identity attribute controllable face generation method based on a double-flow diffusion model according to claim 1, characterized in that, The image recognition of the face reference image adopts an ArcFace algorithm; and the text encoder adopts a CLIP text encoder.

4. The identity attribute controllable face generation method based on a double-flow diffusion model according to claim 2, characterized in that, The training process of the double-flow diffusion model is as follows: obtaining a face training identity, a face reference training image, a face training target attribute, a face identity embedding training vector, a face original attribute embedding training vector, a face target attribute embedding training vector, and second Gaussian noise; mapping the face reference training image by using a VAE encoder and adding first Gaussian noise to obtain a training noisy latent variable; inputting the training noisy latent variable, the face identity embedding training vector, and the face original attribute embedding training vector into a denoising branch diffusion model to obtain predicted noise; inputting the second Gaussian noise, the face identity embedding training vector, and the face target attribute embedding training vector into a decoupling branch diffusion sub-model and repeating ten times to obtain a tenth training identity attribute controllable face initial image; inputting the tenth training identity attribute controllable face initial image into a VAE decoder to obtain a training identity attribute controllable face image; inputting the training identity attribute controllable face initial image into an attribute classifier and an arcface recognition model for processing to obtain face predicted attributes and a predicted identity; based on the predicted noise, the first Gaussian noise, the predicted identity, the face training identity, the face predicted attributes, and the face training target attribute, a diffusion total loss is calculated; the diffusion total loss includes a diffusion loss, an identity loss, and an attribute loss; based on the diffusion total loss, the parameters of the denoising branch diffusion model and the decoupling branch diffusion model are adjusted by a gradient descent method and an AdamW optimizer.

5. The identity-controllable face generation method based on a double-flow diffusion model according to claim 4, characterized in that, The process of obtaining the predicted noise is as follows: inputting the training noisy latent variable into a Conv1 layer to obtain training noisy latent variable convolution features; inputting the training noisy latent variable convolution features, the face identity embedding training vector, and the face original attribute embedding training vector into a Down-Block1 module to obtain first image-text denoising key encoding features; sequentially inputting the first image-text denoising key encoding feature module into a Down-Block2 module and a Down-Block3 module to respectively obtain second image-text denoising key encoding features, second image-text denoising key initial encoding features, second image-text denoising key reinforced encoding features, third image-text denoising key encoding features, third image-text denoising key initial encoding features, and third image-text denoising key reinforced encoding features; inputting the third image-text denoising key encoding feature module into a Down-Block4 module to obtain fourth image-text denoising key encoding features and fourth image-text denoising key initial encoding features; inputting the fourth image-text denoising key encoding features into a Mid-Block module to obtain fifth image-text denoising key encoding features; inputting the fifth image-text denoising key encoding features, the fourth image-text denoising key encoding features, the fourth image-text denoising key initial encoding features, and the third image-text denoising key encoding features into an Up-Block1 module to obtain sixth image-text denoising key decoding features; The sixth graph-text denoising key decoding feature, the third graph-text denoising key initial encoding feature, the third graph-text denoising key reinforced encoding feature and the second graph-text denoising key encoding feature are input into an Up-Block2 module to obtain a seventh graph-text denoising key decoding feature; The seventh graph-text denoising key decoding feature, the second graph-text denoising key initial encoding feature, the second graph-text denoising key reinforced encoding feature and the first graph-text denoising key encoding feature are input into an Up-Block3 module to obtain an eighth graph-text denoising key decoding feature; The eighth graph-text denoising key decoding feature, the first graph-text denoising key initial encoding feature, the first graph-text denoising key reinforced encoding feature and the training noisy latent variable convolution feature are input into an Up-Block4 module to obtain a ninth graph-text denoising key decoding feature; The ninth graph-text denoising key decoding feature is input into a Conv2 layer to obtain a predicted noise.

6. The identity attribute controllable face generation method based on a double-flow diffusion model according to claim 5, characterized in that, The Down-Block1 module comprises: The training noisy latent variable convolution feature is input into a ResBlock5 layer to obtain a training noisy latent variable convolution time feature; The training noisy latent variable convolution time feature, the face identity embedding training vector and the face original attribute embedding training vector are input into a SpatialTransformer layer to obtain the first graph-text denoising key initial encoding feature; The first graph-text denoising key initial encoding feature, the face identity embedding training vector and the face original attribute embedding training vector are sequentially input into a ResBlock5 layer and a SpatialTransformer layer of an RB-ST2 layer to obtain the first graph-text denoising key reinforced encoding feature; The first graph-text denoising key reinforced encoding feature is input into a down-sampling layer to obtain the first graph-text denoising key encoding feature.

7. The identity-controllable face generation method based on a double-flow diffusion model according to claim 6, characterized in that, The formula corresponding to the SpatialTransformer is: K e = W k e; V e = W v e; wherein, respectively represent the key vector and the value vector corresponding to the face identity embedding training vector, respectively represent the key vector and the value vector corresponding to the face identity embedding training vector, respectively represent the key vector and the value vector corresponding to the face identity embedding training vector, e , V e respectively represent the key vector and the value vector, v , W k respectively represent the weight matrix corresponding to the K value and the V value, and softmax(·) represents the maximum value function, T represents the transpose matrix, d represents the corresponding dimension of K e , and e represents the input feature, i.e., the face identity embedding training vector or the face original attribute embedding training vector.

8. The identity attribute controllable face generation method based on a double-flow diffusion model according to claim 4, characterized in that, The calculation of the diffusion total loss comprises: Based on the formula: obtaining a diffusion loss Based on the formula: obtaining an identity loss Based on the formula: get attribute loss Based on the formula: Based on the formula: obtaining a diffusion total loss wherein, denotes expectation, ε denotes a first Gaussian noise, ε θ denotes a U-Net model, and the output of the U-Net model is a predicted noise, z t denotes a training noisy latent variable at a t-th time step, ||·|| denotes a norm, φ(·) denotes an ArcFace model, log(·) denotes a logarithm function with a natural constant as a base number, and ∑(·) denotes a summation function, denotes a training identity-controllable face initial image, b i denotes an i-th face training target attribute, denotes an attribute classifier of an i-th face target attribute, n denotes a total number of attributes, and G(·) denotes a decoupling branch diffusion sub-module, denotes a face reference training image, λ att , and λ id respectively denote an attribute weight and an identity weight.