Face attribute editing and style control method and system based on potential guidance and reference guidance diffusion model
By constructing a style coding generation unit, a modulation coding unit and a denoising unit of a potential guided and reference guided diffusion model, the limitations of face attribute editing and style manipulation in the prior art are solved, and a higher precision and stable image editing effect is achieved.
Patent Information
- Application Number
- CN202510284831.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-11
- Publication Date
- 2025-07-04
AI Technical Summary
The existing facial attribute editing technology has limitations in style manipulation, including the training instability of the GAN model, limited semantic direction expression ability and dependence on training samples, resulting in insufficient editing accuracy and quality.
Based on the potential guidance and reference guidance diffusion model, a style coding generation unit, a style modulation coding unit and a diffusion denoising unit are constructed. Through the forward reverse consistency training strategy, the style coding and semantic coding of face images are generated and modulated, and the noise is gradually denoised to improve editing accuracy and stability.
Improve the accuracy of facial attribute editing and style manipulation, optimize image editing quality, and alleviate the instability of the training process and the dependence on paired training samples.
Smart Images

Figure CN120259488A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of face image synthesis, attribute editing, and style manipulation. More specifically, it relates to a method and system for face attribute editing and style manipulation based on a latent-guided and reference-guided diffusion model. Background Art
[0002] Face attribute editing technology is a technology that uses computer vision and deep learning methods to modify or enhance certain features of a face image. Its goal is to generate personalized images that meet the user's needs by editing specific attributes in the face image, such as age, gender, expression, hairstyle, skin color, etc. Face attribute editing technology has a wide range of applications in fields such as virtual avatar generation, professional photo editing, and data augmentation in face recognition tasks. These application scenarios have a strong demand for precise editing, that is, being able to keep other face features unchanged while modifying the target attribute. Achieving this editing accuracy faces huge challenges due to the complexity of face features and the high correlation between face attributes and face features, such as the connection between age and glasses.
[0003] To achieve precise editing, many studies using advanced conditional generative adversarial networks (cGANs) have emerged in recent years, mainly focusing on two categories of directions: operating in the latent space of pre-trained GANs and learning image-to-image translation. The first category of directions focuses on the attribute editing task. This method first obtains the latent representation of a given face through GAN inversion, then operates on this latent representation along the semantic direction, and inputs it into the pre-trained GAN to generate the editing result. The second category of directions mainly uses an encoder-decoder framework, taking the attribute label as the conditional input to guide the image editing process. However, it is worth noting that although the above methods have achieved certain success in attribute editing, they have limitations in style manipulation. These limitations mainly stem from the discontinuity of the GAN latent space, the mode collapse problem, the ambiguity of style definition, the sensitivity to hyperparameters, and the deficiencies of the encoder-decoder framework in aspects such as style representation, generated image quality, flexibility, and dependence on training data. To solve this problem, some advanced GAN-based methods achieve style manipulation by introducing Gaussian noise or reference images to manipulate the attribute appearance, such as distinguishing between myopia glasses and sunglasses. Although GAN-based methods have made significant progress in attribute editing and style manipulation, they are still restricted by the inherent problems of GANs. These problems include unstable training, non-convergence, and mode collapse, restricting their application in high-quality and accurate attribute editing and style manipulation.
[0004] The prior art discloses a face stylization method, system, device and medium. In this solution, first, the real face model StyleGAN is trained using the stylized image to obtain a stylized face StyleGAN model. Then, at least one face attribute direction vector is obtained, and the initial vector corresponding to the real face image is superimposed with the face attribute direction vector to obtain a target vector. The target vector is input into the trained stylized face StyleGAN model to obtain a target stylized face image. This solution uses the trained StyleGAN model to edit the attributes of the target face, which improves the accuracy of face attribute editing to a certain extent. However, the StyleGAN model used in this solution has problems such as poor definition ability of styles, unstable training process, and strong dependence on high-quality training samples.
[0005] Recently, diffusion models have shown advantages over GANs in image synthesis quality and have achieved remarkable success in various face editing tasks. Inspired by this progress, more and more research on face attribute editing has started to turn its attention to diffusion models. By adopting a simplified version of the variational lower bound, these methods optimize the conditional diffusion model to learn a semantically rich latent space. Then, the editing effect is achieved by operating on the face representation along the semantic direction. Different from the GAN methods that rely on unstable adversarial training, these diffusion-based methods have achieved impressive attribute editing effects through a more stable and direct self-reconstruction training scheme. However, there is still a key limitation: the expressive ability of the semantic direction is limited, resulting in these methods still being unable to effectively perform style manipulation. Summary of the Invention
[0006] To solve the problems of unstable training process, limited expressive ability of semantic direction, and limitations in style manipulation in the current face attribute editing solutions, the present invention proposes a face attribute editing and style manipulation method and system based on a latent-guided and reference-guided diffusion model, which improves the accuracy of attribute editing and style manipulation and optimizes the quality of image editing; at the same time, improves the stability of the training process and alleviates the dependence on paired training samples.
[0007] To achieve the above technical effects, the technical solution of the present invention is as follows:
[0008] A face attribute editing and style manipulation method based on a latent-guided and reference-guided diffusion model, the method comprising the following steps:
[0009] S1: Construct a latent guidance and reference guidance diffusion model, including a style encoding generation unit, a style modulation encoding unit, and a diffusion denoising unit; the style encoding generation unit generates a style encoding of a face image based on a latent guidance or reference guidance strategy, the style modulation encoding unit modulates the input face image using the style encoding to generate a semantic encoding, and the diffusion denoising unit gradually denoises the input face image based on the semantic encoding;
[0010] S2: Obtain a reference face image and an input face image, and train the style encoding generation unit and the style modulation encoding unit using a forward-backward consistency training strategy to stabilize the denoising process of the diffusion denoising unit for the input face image, and obtain a trained latent guidance and reference guidance diffusion model;
[0011] S3: Use the trained latent guidance and reference guidance diffusion model for attribute editing and style manipulation of face images.
[0012] In this technical solution, first, a latent guidance and reference guidance diffusion model including a style encoding generation unit, a style modulation encoding unit, and a diffusion denoising unit is constructed. Among them, the style encoding generation unit generates a style encoding of a face image based on a latent guidance or reference guidance strategy, enhancing the expression ability of the style; the style modulation encoding unit modulates the style encoding and the input face image to generate a semantic encoding, effectively improving the accuracy of attribute editing and style manipulation and optimizing the quality of image editing; the diffusion denoising unit gradually denoises the input face image based on the semantic encoding; using a forward-backward consistency training strategy to train the style encoding generation unit and the style modulation encoding unit improves the stability of the training process and alleviates the dependence on paired training samples.
[0013] Preferably, the style encoding generation unit described in step S1 includes a mapper and an extractor; wherein, the mapper is used to implement random style editing to generate a random style encoding, and the extractor performs set style editing according to the reference face image to generate a set style encoding.
[0014] Preferably, the mapper is a multi-layer perceptron, and the multi-layer perceptron is composed of five LinearReLU layers and one linear layer. The specific process of generating the style encoding is as follows:
[0015] Sample a Gaussian noise z from a standard normal distribution as the input of the mapper;
[0016] The mapper generates a style code based on the specific attribute j' corresponding to the given abstract attribute category label i, where the specific attribute j' is an attribute different from the specific attribute j: the first two LinearReLU layers are indexed by the abstract attribute category label i, and the subsequent three LinearReLU layers and one Linear layer are indexed by the specific attribute j' to generate the style code s i,j' , which is a random style code, and the expression is:
[0017] s i,j′ = M i,j′ (z).
[0018] Preferably, the extractor includes an input layer, an intermediate layer, and an output layer. The process of generating the style code is as follows:
[0019] Input the reference face image y i,j' , where i represents the abstract attribute category label and j' represents the specific attribute corresponding to the abstract attribute category label i;
[0020] Use the input layer and the intermediate layer to extract the face features of the reference face image;
[0021] In the output layer, use i for indexing to encode the face features of the reference face image into the style code s i,j’ and output it. The style code is a set style code, and the expression is:
[0022] s i,j′ = E i (y i,j′ ).
[0023] According to the above technical means, the extractor and the mapper generate the style code of the face image based on the latent guidance or reference guidance strategy, enhancing the expression ability of the style.
[0024] Preferably, the style modulation coding unit described in step S2 includes an input module, a first style modulation module, an intermediate layer, a second style modulation module, and an output layer;
[0025] The input module receives the input face image x i,j , where i represents the abstract attribute category label, and converts the pixel data of the input face image x i,j into high-dimensional abstract features and outputs them to the style modulation module;
[0026] Both the first style modulation module and the second style modulation module are used to modulate the high-dimensional abstract features of the input face image based on the learnable variables and the style code to obtain the high-dimensional abstract features of the input face image with the modulated local features;
[0027] The middle layer is used to receive the high-dimensional abstract features of the input face image modulated by the first style modulation module, further refine the high-dimensional abstract features, and output them to the second style modulation module; the second style modulation module receives the further refined high-dimensional abstract features, and further modulates the refined high-dimensional abstract features based on the learnable variables and style encoding to obtain the clearest high-dimensional abstract features of the modulated input face image, and outputs them to the output layer;
[0028] The output layer is used to receive the clearest high-dimensional abstract features of the input face image modulated by the second style modulation module, refine them to obtain semantic encoding, and output the semantic encoding to the diffusion denoising unit.
[0029] Preferably, either the first style modulation module or the second style modulation module includes two parallel linear modules, a learnable variable module, several sequentially connected AdaIN modules, and a cross-attention cross module;
[0030] Both of the two parallel linear modules are used to receive the style encoding, perform linear transformation on the style encoding to obtain a style vector and a style embedding, and input them into the AdaIN module and the cross-attention cross module respectively;
[0031] The learnable variable module is connected to the first AdaIN module, and the learnable variable module introduces an additional learnable vector as the conditional input of the first AdaIN module;
[0032] The AdaIN module also receives the high-dimensional abstract features of the input face image from the input module and the style vector, and uses the style vector and the learnable variable to modulate the high-dimensional abstract features of the input face image to obtain the modulated high-dimensional abstract features of the input face image and output them to the cross-attention cross module;
[0033] The cross-attention cross module is used to receive the modulated high-dimensional abstract features of the input face image and the style embedding, and perform a cross-attention operation between the high-dimensional abstract features and the style embedding, and output the operated features to the output layer.
[0034] According to the above technical means, the style modulation encoding unit modulates the style encoding and the input face image to generate semantic encoding, effectively improving the accuracy of attribute editing and style manipulation, and optimizing the quality of image editing.
[0035] Preferably, the style encoding generation unit and the style modulation encoding unit are trained using a forward-backward consistency training strategy, and the process is as follows:
[0036] Select two different specific attributes j and j' under a certain abstract attribute category label i of the face;
[0037] Select a certain original face image x in the training seti,j The original face image has property j under a certain abstract property category i of the face, and a semantic mask m of the face region showing property j is obtained i,j ;
[0038] Select a face image r in the training set i,j’ The face image has a specific property j' under a certain abstract property category label i of the face, and a semantic mask m of the face region showing the specific property j' is obtained i,j’ ; j' represents a property different from the specific property j, and represents the modified property of the specific property j; the semantic direction of the modified property j' of the specific property j is d s , the semantic direction d s The expression of is:
[0039]
[0040] where D i,j′ and D i,j respectively represent the face image sets of property j' and property j, |D i,j' | and |D i,j | represent the cardinality of the face image set, φ represents the input block; φ(y i,j′ ) represents the face image feature obtained after processing the face image with property j' through the input block, and φ(x i,j ) represents the face image feature obtained after processing the face image with property j through the input block;
[0041] Using the semantic mask m i,j Transfer the face region showing property j' in r i,j’ to the face region showing property j in x i,j to obtain the swapped face image x i,j' ;
[0042] Define the pre-vector d for editing the face property j to j' m The expression of is:
[0043] d m = φ(x i,j′ ) - φ(x i,j )
[0044] where φ(x i,j ) represents the face image feature output after processing the original face image through the input block, and φ(x i,j' ) represents the face image feature output after processing the swapped face image through the input block; the movement along this direction mainly affects the j property region;
[0045] Using the semantic direction d s Adjust the pre-vector dm The length is used to obtain the vector d for editing the face attribute from j to j'. t The expression is as follows:
[0046]
[0047] Extract the face image x i,j The face image feature f in it i,j , and move the original face image feature f according to the vector d t to obtain the adjusted face image feature f i,j ; i,j'
[0048] Use the style modulation encoding unit to restore the adjusted face image feature f i,j' to the original face image feature f i,j , and the process is as follows:
[0049] Introduce the perceptual loss for image reconstruction and style transfer, and define the perceptual loss function L perc The expression is as follows:
[0050]
[0051] where SM(f i,j′ , s i,j ) represents the effect characterization after restoring the feature map f i,j by the style modulation module based on the style encoding s i,j' ;
[0052] Modulate and integrate the restored face image feature f i,j to obtain the semantic encoding c i,j and output it;
[0053] Construct the classification loss function L cls The expression is as follows:
[0054]
[0055] where, is the n-dimensional attribute value vector of x i,j , represents the transpose operation, represents the prediction of n attribute values;
[0056] Based on the perceptual loss function L perc and the classification loss function L cls construct the complete objective function L full as follows:
[0057]
[0058] When the complete target loss function reaches the minimum value, the training process converges, and the trained style encoding generation unit and style modulation encoding unit are obtained.
[0059] According to the above technical means, the trained style encoding generation unit and style encoding modulation unit are obtained; the stability of the training process is improved, and the dependence on paired training samples is alleviated.
[0060] Preferably, the use of the diffusion denoising unit to gradually denoise the input face image based on the semantic encoding includes the inverse process of deterministic denoising, specifically:
[0061] Define the time step sequence t = 0, 1,..., T - 1, where T is the maximum time step sequence;
[0062] Define the noise scheduling parameter α t , and the expression is:
[0063]
[0064] where β s is a hyperparameter that controls the noise level;
[0065] Construct the inverse process expression of deterministic denoising, and the expression is:
[0066]
[0067] where x t,(i,j) is the face image after the t-th iteration, ∈ θ (x t,(i,j) , t, c i,j ) is the noise predicted by the diffusion denoising unit, α t+1 is the noise scheduling parameter in this iteration, and x t+1,(i,j) is the face image obtained after the t-th iteration;
[0068] Let t = 0, and input the face image x 0,(i,j) = x i,j into the inverse process expression of deterministic denoising to obtain the face image at the next time step t = 1;
[0069] Loop and execute the inverse process of deterministic denoising in the time step sequence of t = 0, 1,..., T - 1, and calculate using the inverse process expression of deterministic denoising until t takes T - 1 to obtain the final x T,(i,j) .
[0070] Preferably, the use of the diffusion denoising unit to gradually denoise the input face image based on the semantic encoding further includes the deterministic denoising process, specifically:
[0071] Given the face image x to be editedT,(i,j) and the new semantic encoding c output from the style modulation encoding unit i,j′ , the expression is:
[0072] c i,j′ = F i,j′ (x i,j , s i,j′ ),
[0073] where s i,j′ = E i (y i,j′ ) or s i,j′ = M i,j′ (z),
[0074] Let t = T,..., 1;
[0075] Construct a deterministic denoising expression, the expression is:
[0076]
[0077] where x t,(i,j’) is the face image after the (T - t)-th iteration, ∈ θ (x t,(i,j′) , t, c i,j′ ) is the noise predicted by the diffusion denoising unit, α t-1 is the noise scheduling parameter in this iteration, x t-1,(i,j’) is the denoised face image obtained in this iteration;
[0078] Let t = T, input the face image x T,(i,j) into the deterministic denoising expression to obtain the face image at the next time step t = T - 1;
[0079] Loop through the time step sequence of t = T,..., 1 to perform the deterministic denoising process, and use the deterministic denoising expression for calculation until t takes 1 to obtain the final x i,j' .
[0080] According to the above technical means, the diffusion denoising unit is used to gradually denoise the input face image based on the semantic encoding, improving the quality of face image editing.
[0081] This application also proposes a face attribute editing and style manipulation system based on a latent-guided and reference-guided diffusion model. The system includes:
[0082] The latent guidance and reference guidance diffusion model construction module is used to construct a style encoding generation unit, a style modulation encoding unit, and a diffusion denoising unit; the style encoding generation unit generates the style encoding of the face image based on the latent guidance or reference guidance strategy, the style modulation encoding unit modulates the style encoding and the input face image to generate a semantic encoding, and the diffusion denoising unit gradually denoises the input face image based on the semantic encoding;
[0083] The forward and backward consistency training module is used to train the style encoding generation unit and the style modulation encoding unit to stabilize the denoising process of the diffusion denoising unit for the input face image, and obtain the trained latent guidance and reference guidance diffusion model;
[0084] The attribute editing module is used to perform attribute editing and style manipulation of the face image by using the trained latent guidance and reference guidance diffusion model.
[0085] Compared with the prior art, the beneficial effects of the technical solution of the present invention are:
[0086] The present invention proposes a method and system for face attribute editing and style manipulation based on a latent guidance and reference guidance diffusion model. First, a latent guidance and reference guidance diffusion model including a style encoding generation unit, a style modulation encoding unit, and a diffusion denoising unit is constructed. Among them, the style encoding generation unit generates the style encoding of the face image based on the latent guidance or reference guidance strategy, which enhances the expression ability of the style; the style modulation encoding unit modulates the style encoding and the input face image to generate a semantic encoding, effectively improving the accuracy of attribute editing and style manipulation, and optimizing the quality of image editing; the diffusion denoising unit gradually denoises the input face image based on the semantic encoding; the forward and backward consistency training strategy is adopted to train the style encoding generation unit and the style modulation encoding unit, improving the stability of the training process and alleviating the dependence on paired training samples. Description of the Drawings
[0087] Figure 1 It represents a schematic flow chart of the method for face attribute editing and style manipulation based on the latent guidance and reference guidance diffusion model proposed in Embodiment 1 of the present invention;
[0088] Figure 2 It represents a process diagram of generating style encoding by using a mapper proposed in Embodiment 2 of the present invention;
[0089] Figure 3 It represents a process diagram of generating style encoding by using an extractor proposed in Embodiment 2 of the present invention;
[0090] Figure 4 It represents a schematic diagram of the overall structure of the style modulation encoding unit proposed in Embodiment 2 of the present invention;
[0091] Figure 5 Schematic structural diagram of the style modulation module proposed in Embodiment 2 of the present invention;
[0092] Figure 6 Schematic flow diagram of transferring the face image in the area of the worn glasses proposed in Embodiment 2 of the present invention;
[0093] Figure 7 Schematic diagram showing the restoration of the face image feature f after removing glasses to the original state of the face image feature f i,j' by using the style modulation coding unit proposed in Embodiment 2 of the present invention; i,j Flow chart;
[0094] Figure 8 Schematic flow diagram of adding and removing noise by using the diffusion denoising unit proposed in Embodiment 2 of the present invention;
[0095] Figure 9 Schematic algorithm diagram of face attribute editing and style manipulation based on the latent-guided and reference-guided diffusion model;
[0096] Figure 10 Schematic structural diagram of a face attribute editing and style manipulation system based on the latent-guided and reference-guided diffusion model proposed in Embodiment 3 of the present invention. Detailed implementation manners
[0097] The accompanying drawings are only for illustrative purposes and should not be construed as limitations on this patent;
[0098] For better illustration of this embodiment, some parts of the accompanying drawings are omitted, enlarged or reduced, and do not represent the actual size;
[0099] For those skilled in the art, it is understandable that some well-known content descriptions in the accompanying drawings may be omitted.
[0100] The technical solutions of the present invention will be further described below with reference to the accompanying drawings and embodiments.
[0101] The description of the positional relationship in the accompanying drawings is only for illustrative purposes and should not be construed as limitations on this patent;
[0102] Embodiment 1
[0103] This embodiment proposes a face attribute editing and style manipulation method based on the latent-guided and reference-guided diffusion model. The schematic flow diagram of this method is shown in Figure 1 , and includes the following steps:
[0104] S1: Construct a latent-guided and reference-guided diffusion model, including a style encoding generation unit, a style modulation encoding unit, and a diffusion denoising unit. The style encoding generation unit generates a style encoding of a face image based on a latent-guided or reference-guided strategy. The style modulation encoding unit modulates the input face image using the style encoding to generate a semantic encoding. The diffusion denoising unit gradually denoises the input face image based on the semantic encoding.
[0105] S2: Obtain a reference face image and an input face image, and train the style encoding generation unit and the style modulation encoding unit using a forward-backward consistency training strategy to stabilize the denoising process of the diffusion denoising unit for the input face image, thereby obtaining a trained latent-guided and reference-guided diffusion model.
[0106] S3: Use the trained latent-guided and reference-guided diffusion model for attribute editing and style manipulation of face images.
[0107] In this embodiment, first, a latent-guided and reference-guided diffusion model including a style encoding generation unit, a style modulation encoding unit, and a diffusion denoising unit is constructed. Among them, the style encoding generation unit generates a style encoding of a face image based on a latent-guided or reference-guided strategy, enhancing the expression ability of the style. The style modulation encoding unit modulates the style encoding and the input face image to generate a semantic encoding, effectively improving the accuracy of attribute editing and style manipulation and optimizing the quality of image editing. The diffusion denoising unit gradually denoises the input face image based on the semantic encoding. Using a forward-backward consistency training strategy to train the style encoding generation unit and the style modulation encoding unit improves the stability of the training process and alleviates the dependence on paired training samples.
[0108] Embodiment 2
[0109] In this embodiment, some letter parameters are explained in advance. Among them, i represents the category label of the abstract attribute of the image. For example, i = 1 represents "bangs", i = 2 represents "glasses", etc. j represents the specific attribute under a certain abstract attribute category. For example, when i = 1 represents "bangs", j = 1 represents "with bangs", j = 2 represents "without bangs"; x i,j or y i,j represents an image with attribute j of label i category, s i,j represents the feature of label i with attribute j, c i,j represents the semantic encoding of label i with attribute j.
[0110] The style encoding generation unit described in step S1 includes a mapper and an extractor. Among them, the mapper is used to implement random style editing to generate a random style encoding, and the extractor performs set style editing based on the reference face image to generate a set style encoding.
[0111] In this embodiment, the mapper is a multi-layer perceptron. For the process diagram of generating style encoding using the mapper, see Figure 2 , the multi-layer perceptron consists of five LinearReLU layers and one linear layer. The specific process of generating style encoding is as follows:
[0112] Sample a Gaussian noise z from the standard normal distribution as the input to the mapper;
[0113] The mapper generates a style encoding based on the specific attribute j' corresponding to the given abstract attribute category label i, where the specific attribute j' is an attribute different from the specific attribute j: The first two LinearReLU layers are indexed by the abstract attribute category label i, and the subsequent three LinearReLU layers and one Linear layer are indexed by the specific attribute j' to generate the style encoding s i,j' , which is a random style encoding, and the expression is:
[0114] s i,j′ = M i,j′ (z).
[0115] In this embodiment, the extractor includes an input layer, an intermediate layer, and an output layer. For the process diagram of generating style encoding using the extractor, see Figure 3 , and the process of generating style encoding is as follows:
[0116] Input the reference face image y i,j' , where i represents the abstract attribute category label and j' represents the specific attribute corresponding to the abstract attribute category label i;
[0117] Use the input layer and the intermediate layer to extract the face features of the reference face image;
[0118] In the output layer, use i for indexing to encode the face features of the reference face image into the style encoding s i,j’ and output it. The style encoding is a set style encoding, and the expression is:
[0119] s i,j′ = E i (y i,j′ ).
[0120] In this embodiment, the style modulation encoding unit described in step S2 includes an input module, a first style modulation module, an intermediate layer, a second style modulation module, and an output layer; for the overall structure of the style modulation encoding unit, see Figure 4 ;
[0121] The input module receives the input face image x i,j , where i represents the abstract attribute category label, and input the face image x i,jThe pixel data is converted into high-dimensional abstract features and output to the style modulation module;
[0122] Both the first style modulation module and the second style modulation module are used to modulate the high-dimensional abstract features of the input face image based on learnable variables and style encodings, and obtain the high-dimensional abstract features of the input face image with modulated local features;
[0123] The intermediate layer is used to receive the high-dimensional abstract features of the input face image modulated by the first style modulation module, further refine the high-dimensional abstract features, and output them to the second style modulation module; the second style modulation module receives the further refined high-dimensional abstract features, and further modulates the refined high-dimensional abstract features based on learnable variables and style encodings to obtain the clearest high-dimensional abstract features of the input face image after modulation, and outputs them to the output layer;
[0124] The output layer is used to receive the clearest high-dimensional abstract features of the input face image modulated by the second style modulation module, refine them to obtain semantic encodings, and output the semantic encodings to the diffusion denoising unit.
[0125] In this embodiment, any one of the first style modulation module and the second style modulation module includes two parallel linear modules, a learnable variable module, several sequentially connected AdaIN modules and a cross-attention cross module; for the structural schematic diagram of any style modulation module, see Figure 5 ;
[0126] Both of the two parallel linear modules are used to receive style encodings, perform linear transformations on the style encodings to obtain style vectors and style embeddings, and respectively input them into the AdaIN module and the cross-attention cross module;
[0127] The learnable variable module, which is connected to the first AdaIN module, introduces an additional learnable vector as the conditional input of the first AdaIN module;
[0128] The AdaIN module also receives the high-dimensional abstract features of the input face image from the input module and the style vector, and uses the style vector and learnable variables to modulate the high-dimensional abstract features of the input face image, obtains the modulated high-dimensional abstract features of the input face image and outputs them to the cross-attention cross module;
[0129] The cross-attention cross module is used to receive the modulated high-dimensional abstract features of the input face image and the style embedding, and perform a cross-attention operation between the high-dimensional abstract features and the style embedding, and output the operated features to the output layer.
[0130] In this embodiment, the style encoding generation unit and the style modulation encoding unit are trained using a forward-backward consistency training strategy. The more specific process is as follows:
[0131] Select the specific attributes j and j' of whether to wear glasses under the abstract attribute category i of the selected human face; where j represents wearing glasses and j' represents not wearing glasses;
[0132] Select a face image x of a person wearing glasses in the training set i,j , and obtain the semantic mask m of the eye region of the face of the person wearing glasses i,j ;
[0133] Select a face image r of a person not wearing glasses in the training set i,j’ ;
[0134] The semantic direction d s is expressed as:
[0135]
[0136] where D i,j′ and D i,j respectively represent the sets of face images with attributes j' and j, |D i,j' | and |D i,j | represent the cardinality of the set of face images, and φ represents the input block; φ(y i,j′ ) represents the face image feature obtained after processing the face image with attribute j' through the input block, and φ(x i,j ) represents the face image feature obtained after processing the face image with attribute j through the input block;
[0137] Use the semantic mask m i,j to transfer the face region showing the j' attribute in r i,j’ to the face region showing the j attribute in x i,j to obtain the swapped face image x i,j' ; See the schematic flow diagram of transferring the face image of the glasses-wearing area in Figure 6
[0138] Define the pre-vector d m for editing the face attribute from j to j' as:
[0139] d m = φ(x i,j′ ) - φ(x i,j )
[0140] where φ(x i,j ) represents the face image feature output after processing the face image of the person wearing glasses through the input block, and φ(x i,j′ ) represents the face image feature output after processing the face image of the person not wearing glasses through the input block; The movement along this direction mainly affects the j-attribute area;
[0141] Utilize the semantic direction d s Adjust the pre-vector d m in length to obtain the vector d for editing the face attribute from j to j’ t The expression is as follows:
[0142]
[0143] Extract the face image x of the person wearing glasses i,j and the face image feature f therein i,j , and move the face image feature f of the person wearing glasses according to the direction of the vector d t to obtain the face image feature f after removing the glasses i,j ; i,j'
[0144] Restore the face image feature f after removing the glasses to the face image feature f in the original state using the style modulation encoding unit i,j' , and the restoration flowchart is shown in i,j , and the process is as follows: Figure 7
[0145] Introduce the perceptual loss for image reconstruction and style transfer, and define the perceptual loss function L perc The expression is as follows:
[0146]
[0147] where SM(f i,j′ , s i,j ) represents the effect characterization after restoration of the feature map f by the style modulation module based on the style encoding s i,j ; i,j'
[0148] Modulate and integrate the restored face image feature f i,j to obtain the semantic encoding c i,j and output it;
[0149] Construct the classification loss function L cls The expression is as follows:
[0150]
[0151] where is the n-dimensional attribute value vector of x i,j , represents the transpose operation, represents the prediction of n attribute values;
[0152] Based on the perceptual loss function L perc and the classification loss function L cls Construct the complete objective function L full as follows:
[0153]
[0154] When the complete target loss function reaches the minimum value, the training process converges, and the trained style encoding generation unit and style modulation encoding unit are obtained.
[0155] In this embodiment, the diffusion denoising unit gradually denoises the input face image based on the semantic encoding, including the inverse process of deterministic denoising. For the structural diagram of the diffusion denoising unit, see Figure 8 , and the process includes:
[0156] First, execute the inverse process of deterministic denoising. Given the face image x to be edited 0,(i,j) and the corresponding semantic encoding c output from the style modulation encoding unit i,j ;
[0157] Define the time step sequence t = 0, 1,..., T - 1, where T is the maximum time step sequence;
[0158] Define the noise scheduling parameter α t , and the expression is:
[0159]
[0160] where β s is a hyperparameter that controls the noise level;
[0161] Construct the inverse process expression of deterministic denoising, and input the image x to be edited 0,(i,j) into the inverse process expression of deterministic denoising for iteration. The expression is:
[0162]
[0163] where x t,(i,j) is the face image after the t-th iteration, ∈ θ (x t,(i,j) , t, c i,j ) is the noise predicted by the diffusion denoising unit, α t+1 is the noise scheduling parameter in this iteration, and x t+1,(i,j) is the face image obtained after the t-th iteration;
[0164] Let t = 0, and input the face image x 0,(i,j) = x i,j into the inverse process expression of deterministic denoising to obtain the face image at the next time step t = 1;
[0165] Loop through the time step sequence of \(t = 0, 1, \ldots, T - 1\) and execute the inverse process of deterministic denoising. Calculate using the expression for the inverse process of deterministic denoising until \(t\) takes \(T - 1\) to obtain the final \(x\). T,(i,j) 。
[0166] The inverse process part of deterministic denoising can be summarized as follows: Given the image \(x\) to be edited i,j and its corresponding semantic encoding \(c\) i,j , we first execute the inverse process of the deterministic generation process to obtain the initial latent noise:
[0167] \(x_0\) T,(i,j) \(= c_{\text{DDIM}}\) enc \((\epsilon \sim\) θ ;\(x\) i,j , \(c\) i,j ) T,(i,j) \(= c_{\text{DDIM}}\) enc \((\epsilon \sim\) θ ;\(x\) i,j , \(c\) i,j ),
[0168] where \(x_0\) T,(i,j) is intended to encode only the information not included in \(c\) i,j , i.e., the random details.
[0169] In this embodiment, the step-by-step denoising of the input face image using the diffusion denoising unit based on the semantic encoding further includes a deterministic denoising process, specifically:
[0170] Given the face image \(x\) to be edited T,(i,j) and the new semantic encoding \(c\) output from the style modulation encoding unit i,j′ , the expression is
[0171] \(c\) i,j′ \(= F\) i,j′ \((x\) i,j , \(s\) i,j′ ),
[0172] where \(s\) i,j′ \(= E\) i \((y\) i,j′ ) or \(s\) i,j′ \(= M\) i,j′ \((z)\),
[0173] Let \(t = T, \ldots, 1\);
[0174] Construct the deterministic denoising expression:
[0175]
[0176] where \(x_t\) t,(i,j’)is the face image after the (T - t)-th iteration, ∈ θ (x t,(i,j′) , t, c i,j′ ) is the noise predicted by the diffusion denoising unit, and α t-1 is the noise scheduling parameter in this iteration, and x t-1,(i,j) is the denoised face image obtained in this iteration;
[0177] Let t = T, and input the face image x T,(i,j) into the deterministic denoising expression to obtain the face image at the next time step t = T - 1;
[0178] Loop through the time step sequence of t = T,..., 1 and perform the deterministic denoising process, calculating using the deterministic denoising expression until t takes 1 to obtain the final x i,j' .
[0179] The denoising part can be summarized as decoding x i,j' during the generation process with the new semantic encoding c T,(i,j) as the condition to obtain the edited image:
[0180] x i,j′ = cDDIM dec (∈ θ ; x T,(i,j) , c i,j′ ).
[0181] More specifically, for the schematic diagram of the algorithm for face attribute editing and style manipulation based on the latent guidance and reference guidance diffusion model, see Figure 9
[0182] This generation process makes x i,j′ conform to the real data distribution, thus ensuring the quality of the edited image.
[0183] Example 3
[0184] As Figure 10 shown, this example proposes a face attribute editing and style manipulation system based on the latent guidance and reference guidance diffusion model, and the system includes:
[0185] The latent guidance and reference guidance diffusion model construction module is used to construct a style encoding generation unit, a style modulation encoding unit, and a diffusion denoising unit; the style encoding generation unit generates the style encoding of the face image based on the latent guidance or reference guidance strategy, the style modulation encoding unit modulates the style encoding and the input face image to generate the semantic encoding, and the diffusion denoising unit gradually denoises the input face image based on the semantic encoding;
[0186] The forward and backward consistency training module is used to train the style encoding generation unit and the style modulation encoding unit to stabilize the denoising process of the input face image by the diffusion denoising unit, and obtain a trained latent guidance and reference guidance diffusion model;
[0187] The attribute editing module is used to perform attribute editing and style manipulation on face images by using the trained latent guidance and reference guidance diffusion model.
[0188] Obviously, the above embodiments of the present invention are only examples for clearly explaining the present invention, rather than limitations on the implementation manners of the present invention. For those of ordinary skill in the art, other different forms of changes or modifications can be made based on the above description. It is not necessary and impossible to enumerate all implementation manners here. Any modifications, equivalent replacements, and improvements made within the spirit and principle of the present invention shall be included in the protection scope of the claims of the present invention.
Claims
1. A method for face attribute editing and style manipulation based on a latent-guided and reference-guided diffusion model, characterized in that It includes the following steps: S1: Construct a latent guidance and reference guidance diffusion model, including a style encoding generation unit, a style modulation encoding unit, and a diffusion denoising unit; S2: Obtain a reference face image and an input face image, and train the style encoding generation unit and the style modulation encoding unit using a forward-backward consistency training strategy to stabilize the denoising process of the diffusion denoising unit for the input face image, and obtain a trained latent guidance and reference guidance diffusion model; S3: Use the trained latent guidance and reference guidance diffusion model for attribute editing and style manipulation of face images. Among them, the style encoding generation unit generates a style encoding of the face image based on a latent guidance or reference guidance strategy, the style modulation encoding unit modulates the input face image using the style encoding to generate a semantic encoding, and the diffusion denoising unit gradually denoises the input face image based on the semantic encoding.
2. The method for face attribute editing and style manipulation based on the latent guidance and reference guidance diffusion model according to claim 1, characterized in that The style encoding generation unit described in step S1 includes a mapper and an extractor; wherein, the mapper is used to implement random style editing and generate a random style encoding, and the extractor performs set style editing according to the reference face image to generate a set style encoding.
3. The method for face attribute editing and style manipulation based on the potential-guided and reference-guided diffusion model according to claim 2, wherein The mapper is a multi-layer perceptron, and the multi-layer perceptron is composed of five LinearReLU layers and one linear layer. The specific process of generating the style encoding is as follows: Sample a Gaussian noise z from a standard normal distribution as the input of the mapper; The mapper generates a style code based on the specific attribute j' corresponding to the given abstract attribute category label i, where the specific attribute j' is an attribute different from the specific attribute j: the first two LinearReLU layers are indexed by the abstract attribute category label i, and the subsequent three LinearReLU layers and one Linear layer are indexed by the specific attribute j' to generate the style code s i,j ', which is a random style code, and the expression is: s i,j′ = M i,j′ (z).
4. The method for face attribute editing and style manipulation based on the latent-guided and reference-guided diffusion model according to claim 2, wherein The extractor includes an input layer, a middle layer, and an output layer. The process of generating the style encoding is as follows: Input reference face image y j,j ', where i represents the abstract attribute category label, and j' represents the specific attribute corresponding to the abstract attribute category label i; Use the input layer and the middle layer to extract the face features of the reference face image; Index with \(i\) in the output layer, and encode the face features of the reference face image into a style code \(s\) i,j’ and output it. The style code is a set style code, and the expression is: s i,j′ = E i (y i,j′ )。 5. The method for face attribute editing and style manipulation based on the latent-guided and reference-guided diffusion model according to claim 1, characterized in that The style modulation encoding unit described in step S2 includes an input module, a first style modulation module, a middle layer, a second style modulation module, and an output layer; The input module receives the input face image x i,j , where i represents the abstract attribute category label. The input face image x i,j 's pixel data is converted into high-dimensional abstract features and output to the style modulation module; Both the first style modulation module and the second style modulation module are used to modulate the high-dimensional abstract features of the input face image based on learnable variables and style encoding to obtain the high-dimensional abstract features of the input face image with modulated local features; The middle layer is used to receive the high-dimensional abstract features of the input face image modulated by the first style modulation module, further refine the high-dimensional abstract features, and output them to the second style modulation module; The second style modulation module receives the further refined high-dimensional abstract features, further modulates the refined high-dimensional abstract features based on learnable variables and style encoding to obtain the clearest modulated high-dimensional abstract features of the input face image, and outputs them to the output layer; The output layer is used to receive the clearest high-dimensional abstract features of the input face image modulated by the second style modulation module, refine them to obtain a semantic encoding, and output the semantic encoding to the diffusion denoising unit.
6. The method for face attribute editing and style manipulation based on the latent-guided and reference-guided diffusion model according to claim 5, wherein Any one of the first style modulation module and the second style modulation module includes two parallel linear modules, a learnable variable module, several sequentially connected AdaIN modules, and a cross-attention cross module; Both of the two parallel linear modules are used to receive the style encoding, perform a linear transformation on the style encoding to obtain a style vector and a style embedding, and input them into the AdaIN module and the cross-attention cross module respectively; The learnable variable module is connected to the first AdaIN module, and the learnable variable module introduces additional learnable vectors as conditional inputs to the first AdaIN module; The AdaIN module also receives the high-dimensional abstract features of the input face image and the style vector from the input module, modulates the high-dimensional abstract features of the input face image using the style vector and the learnable variables, and outputs the modulated high-dimensional abstract features of the input face image to the cross-attention cross module; The cross-attention cross module is used to receive the high-dimensional abstract features of the modulated input face image and the style embedding, perform a cross-attention operation between the high-dimensional abstract features and the style embedding, and output the processed features to the output layer.
7. The method for face attribute editing and style manipulation based on the latent-guided and reference-guided diffusion model according to claim 6, wherein The style encoding generation unit and the style modulation encoding unit are trained using a forward-backward consistency training strategy. The process is as follows: Select two different specific attributes j and j' under a certain abstract attribute category label i of the face; Select a certain original face image \(x\) in the training set i,j , this original face image has the \(j\) attribute under a certain abstract attribute category \(i\) of the face, and obtain the semantic mask \(m\) of the face region showing the \(j\) attribute i,j ; Select a certain face image r in the training set i,j’ , the face image has a specific attribute j' under a certain abstract attribute category label i of the face, and obtain the semantic mask m of the face area showing the specific attribute j' i,j’ ; j' represents an attribute different from the specific attribute j, representing the modified attribute of the specific attribute j; the semantic direction of the modified attribute j' of the specific attribute j is d s , the semantic direction d s The expression of is: Among them, D i,j′ and D i,j respectively represent the sets of face images of attribute j' and attribute j. |D i,j' | and |D i,j | represent the cardinality of the set of face images. φ represents the input block; φ(y i,j′ ) represents the face image feature obtained after processing the face image with attribute j' through the input block, and φ(x i,j ) represents the face image feature obtained after processing the face image with attribute j through the input block; Using the semantic mask m i,j Transfer the face region showing the j' attribute in r i,j’ to the face region showing the j attribute in x i,j to obtain the swapped face image x i,j' ; Define the pre-vector d for editing the face attribute from j to j' m The expression of which is: d m = φ(x i,j′ ) - φ(x i,j ) Among them, φ(x i,j ) represents the face image feature output after the original face image is processed by the input block, and φ(x i,j′ ) represents the face image feature output after the swapped face image is processed by the input block; the movement along this direction mainly affects the j-attribute area; Using the semantic direction d s Adjust the pre-vector d m in length to obtain the vector d for editing the face attribute from j to j' t The expression is as follows: Extract the face image x i,j The face image feature f in i,j , according to the vector d t Move the original face image feature f in the direction of i,j , to obtain the adjusted face image feature f i,j' ; Restore the adjusted facial image feature f i,j' to the original facial image feature f i,j , and the process is as follows: Introduce the perceptual loss for image reconstruction and style transfer, and define the perceptual loss function \(L\). perc The expression of which is as follows: Among them, SM(f i,j′ , s i,j ) represents the effect characterization after the feature map f i,j is restored by the style modulation module based on the style code s i,j' ; The restored facial image feature f i,j is modulated and integrated to obtain the semantic encoding c i,j and output; Construct the classification loss function L cls The expression is as follows: Among them, is the n-dimensional attribute value vector of x i,j , and represents the transpose operation, represents the prediction of n attribute values; Based on the perceptual loss function L perc and the classification loss function L cls Construct the complete objective function L full as follows: When the complete objective loss function reaches the minimum value, the training process converges, and the trained style encoding generation unit and style modulation encoding unit are obtained.
8. The method for face attribute editing and style manipulation based on the latent-guided and reference-guided diffusion model according to claim 6, characterized in that The above-mentioned use of the diffusion denoising unit to gradually denoise the input face image based on the semantic encoding includes the inverse process of deterministic denoising, specifically: Define a time step sequence t = 0, 1,..., T - 1, where T is the maximum time step sequence; Define the noise scheduling parameter α t , and the expression is: Among them, β s is a hyperparameter for controlling the noise level; Construct the inverse process expression of deterministic denoising, and the expression is: Among them, x t,(i,j) is the face image after the t-th iteration, ∈ θ (x t,(i,j) , t, c i,j ) is the noise predicted by the diffusion denoising unit, α t+1 is the noise scheduling parameter in this iteration, x t+1,(i,j) is the face image obtained after the t-th iteration; Let \(t = 0\), and input the face image \(x\) 0,(i,j) \(= x\) i,j into the inverse process expression of deterministic denoising to obtain the face image at the next time step \(t = 1\); Loop through the time step sequence of t = 0, 1, ..., T - 1 and perform the inverse process of deterministic denoising. Calculate using the expression of the inverse process of deterministic denoising until t takes T - 1 to obtain the final x T,(i,j) 。 9. The method for face attribute editing and style manipulation based on the latent-guided and reference-guided diffusion model according to claim 8, wherein The above-mentioned use of the diffusion denoising unit to gradually denoise the input face image based on the semantic encoding also includes the deterministic denoising process, specifically: Given a face image x to be edited T,(x,j) and a new semantic code c output from the style modulation encoding unit i,j′ , the expression is: c i,j′ = F i,j′ (x i,j , s i,j′ ) where s i,j′ = E i (y i,j′ ) or s i,j′ = M i,j′ (z) Let t = T,..., 1; Construct the deterministic denoising expression, and the expression is: Among them, x t,(i,j’) is the face image after the (T - t)-th iteration, ∈ θ (x t,(i,j′) , t, c i,j′ ) is the noise predicted by the diffusion denoising unit, α t-1 is the noise scheduling parameter in this iteration, x t-1,(i,j’) is the denoised face image obtained in this iteration; Let \(t = T\), and input the face image \(x\) T,(i,j) into the deterministic denoising expression to obtain the face image at the next time step \(t = T - 1\); Loop and execute the deterministic denoising process in the time step sequence of t = T,..., 1, calculate using the deterministic denoising expression until t takes 1 to obtain the final x i,j' .
10. A face attribute editing and style manipulation system based on a latent-guided and reference-guided diffusion model, characterized in that, The system is used to implement the method described in any one of claims 1 to 9, including: A latent-guided and reference-guided diffusion model construction module, which is used to construct a style encoding generation unit, a style modulation encoding unit, and a diffusion denoising unit; the style encoding generation unit generates the style encoding of the face image based on the latent-guided or reference-guided strategy, the style modulation encoding unit modulates the style encoding and the input face image to generate the semantic encoding, and the diffusion denoising unit gradually denoises the input face image based on the semantic encoding; A forward-backward consistency training module, which is used to train the style encoding generation unit and the style modulation encoding unit to stabilize the denoising process of the diffusion denoising unit for the input face image, and obtain a trained latent-guided and reference-guided diffusion model; An attribute editing module, which is used to perform attribute editing and style manipulation of the face image using the trained latent-guided and reference-guided diffusion model.
Citation Information
Cited By
Diffusion-based face structure adjustment image cartooning processing method
CN122657275A