A virtual hairstyle transformation method based on diffusion models
Through the two-stage generation method and diffusion model combined with the cross attention mechanism and the LoRA model, the details processing problem of complex textured hairstyles in virtual hairstyle changing technology is solved, and the hairstyle generation with high detail fidelity and stability is achieved.
Patent Information
- Application Number
- CN202411308060.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-09-19
- Publication Date
- 2025-07-04
- Estimated Expiration
- 2044-09-19
AI Technical Summary
Existing virtual hairstyle changes technology is difficult to deal with complex textured hairstyles, and it is easy to ignore details or create artifacts.
A two-stage generation method is adopted. The first stage is to generate rough hairstyle pictures through the baldness generator and the hairstyle reference network, and the second stage is to refine the hairstyle redrawing and diffusion model, combining the cross attention mechanism and the LoRA model to ensure the transmission and integration of hairstyle details.
The generated hairstyle images have higher detail fidelity, can handle complex textures and details, avoid blur or loss of details, and improve generation stability and training efficiency.
Smart Images

Figure CN119130783B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of virtual hair styling, and in particular to a virtual hair styling method based on a diffusion model. Background Art
[0002] With the development of technology, virtual reality technology has gradually matured, bringing more convenient virtual reality applications. Virtual try-on has always been a popular virtual reality application among the public. Virtual hair styling technology is a more prominent and challenging application in this field. Users can choose different hairstyles, colors, and textures according to their preferences for personalized matching and experimentation. The application of this technology greatly enriches the user's hair styling experience and also avoids the disappointment and economic losses caused by unsatisfactory haircut effects. The key to this task is to perfectly transfer the shape, color, and texture of the target hairstyle to the user's head while ensuring that the user's identity information and background remain unchanged.
[0003] In recent years, most virtual hair styling technologies are mainly based on GAN methods. However, the resulting problem is that it cannot handle hairstyles with complex textures, easily ignores details, or generates artifacts. For the above reasons, the present invention proposes a virtual hair styling technology based on a diffusion model. Summary of the Invention
[0004] In view of the above existing problems, the present invention is proposed.
[0005] Therefore, the present invention provides a virtual hair styling method based on a diffusion model to solve the problems of being unable to handle hairstyles with complex textures, easily ignoring details, or generating artifacts.
[0006] To solve the above technical problems, the present invention provides the following technical solutions:
[0007] On the one hand, the present invention provides a virtual hair styling system based on a diffusion model, which includes a first-stage image processing module that processes the input source image and hairstyle reference image, and performs face detection, alignment, and image cropping operations;
[0008] A baldness generator for generating a bald image of the person in the source image, that is, an image without hair;
[0009] A hairstyle reference network that extracts hairstyle details from the hairstyle reference image. This network extracts hairstyle information through a latent space encoding and cross-attention mechanism and injects the hairstyle information details into the hairstyle generation diffusion model;
[0010] A hairstyle generation diffusion model that generates a rough hair styling image based on the bald image generated by the baldness generator and the hairstyle details provided by the hairstyle reference network, capturing the overall contour of the hairstyle;
[0011] The two-stage image processing module processes the rough hairstyle-changing pictures, including face detection, alignment, and cropping operations. In addition, it generates a binary mask for the hairstyle;
[0012] The hairstyle redrawing diffusion model refines and redraws the rough hairstyle-changing pictures generated in the first stage. Combining the hairstyle features trained by the hairstyle LoRA, it redraws the hairstyle area to make the hairstyle blend harmoniously with other parts and outputs fine hairstyle pictures.
[0013] On the other hand, the present invention provides a virtual hairstyle-changing method based on a diffusion model, including,
[0014] Step S1, select pictures of the same person with different hairstyles during training, such as the SketchHairSalon dataset;
[0015] Step S2, randomly select one picture as the source picture and one picture as the hairstyle reference picture from the pictures of the same person with different hairstyles, and then process them into the required format through the image processing module;
[0016] Step S3, input the source picture I src into the baldness generator to obtain the bald picture I bald ;
[0017] Step S4, input the hairstyle reference picture I ref into the hairstyle reference network to extract hairstyle-related details;
[0018] Step S5, input the bald picture I bald into the hairstyle generation diffusion model, and the extracted hairstyle details are input into the hairstyle generation diffusion model through the cross-attention mechanism to finally obtain the rough hairstyle-changing picture;
[0019] Step S6, prepare 510 pictures of different angles (excluding the back) for each hairstyle as the hairstyle LoRA training set, input the pictures in the hairstyle training set into the image processing module to obtain hairstyle pictures and hairstyle description words;
[0020] Step S7, input the hairstyle pictures and hairstyle description words obtained in Step S6 into the diffusion model for LoRA training of the targeted hairstyle;
[0021] Step S8, generate fine hairstyle-changing pictures through the VAE decoder;
[0022] Steps S1 - S5 are the first-stage hairstyle generation steps, and Steps S6 - S8 are the second-stage hairstyle generation steps;
[0023] Since the hairstyle pictures generated in the first stage are relatively rough, a two-stage virtual hairstyle-changing picture generation method is adopted;
[0024] In the first stage, a trained hairstyle generation diffusion model is used to generate rough hairstyle-changing pictures, and in the second stage, a hairstyle redrawing diffusion model is used to redraw the hairstyle area of the rough hairstyle-changing pictures to obtain the final and more refined hairstyle-changing pictures;
[0025] The roughness and refinement here are relative.
[0026] Furthermore, in the first-stage image processing module:
[0027] Input the source image I src , and the hairstyle reference image I ref into the image processing module, detect the face, and then perform face alignment, cropping and resizing to the required size;
[0028] The baldness generator in the first stage includes a VAE encoder ε, a baldness generation diffusion model, and a baldness ControlNet τ b and a VAE decoder;
[0029] The baldness generation diffusion model is a latent space diffusion model StableDiffusion, which uses the pre-trained weights of SDv1.5. StableDiffusion is a classic latent space diffusion model (LDM, Latent Diffusion Model). LDM uses a variational autoencoder VAE. Given an input image, it first passes through the encoder to map the image to the latent space, so that the diffusion is carried out in the latent space. In the forward process of diffusion, Gaussian noise will be gradually added to the latent variable in T iterations. The reverse diffusion process is to gradually denoise, and finally pass through the decoder to restore to the RGB space to obtain the final image;
[0030] The ControlNet branch creates copies of the trainable diffusion model UNet encoding block and intermediate block, and adds additional zero convolutional layers. The encoding block and intermediate block of ControlNet are added to the decoding block and intermediate block of the diffusion model. The diffusion model here is the baldness generation diffusion model, and the output of each copy block is added to the skip connection of the original diffusion model UNet;
[0031] Randomly sample 4-channel latent space Gaussian noise and input it into the baldness generation diffusion model;
[0032] Input the source image into the VAE encoder to obtain the latent space encoding z src = ε(I src ), and the latent space encoding z srcInput into the Bald ControlNet, which is a trainable copy of the decoding block and the middle block of the bald generation diffusion model. After passing through the zero convolution layer, the encoding block and the middle block of the Bald ControlNet are added to the decoding block and the middle block of the bald generation diffusion model through residual connections, and the reference information of the source image is input into the bald generation diffusion model;
[0033] Finally, the output of the bald generation diffusion model passes through the VAE decoder to output the bald image I of the person in the source image in the RGB space b ;
[0034] During the training of the bald generator, only the Bald ControlNet is trained, and the weights of the bald generation diffusion model, the VAE encoder, and the VAE decoder are fixed. The pre-trained weights of SDv1.5 are used. The Bald ControlNet participates in the training, and the initial weights also use the pre-trained weights of SDv1.5. By minimizing the loss function and using the Adam optimizer for optimization iteration, the loss function for training the bald generator is: Among them, represents the loss function, represents the expected value, that is, given the latent space encoding of the source image I src and the noise ∈ sampled from the standard Gaussian distribution as well as the expectation at the time step t, ∈ b represents the noise estimator of the bald generation diffusion model, which is responsible for predicting the input noise distribution at each time step t, τ b (ε(I src )) represents the bald feature of the input source image generated by the Bald ControlNet τ b After passing through the VAE encoder ε to perform latent space encoding on the source image I src and processed at the time step t, ∈ represents Gaussian noise, represents the L2 norm.
[0035] Furthermore, in the hairstyle reference network in the first stage,
[0036] The hairstyle reference network τ h is trained based on the latent space diffusion model StableDiffusion:
[0037] Input the hairstyle reference image I ref into the pre-trained VAE encoder ε to obtain the latent space encoding z = ε(I ref );
[0038] Input the latent space encoding z into the hairstyle reference network τ h to extract the hairstyle detail feature c h, which is input into the hairstyle generation diffusion model through the cross-attention mechanism. The calculation method of the cross-attention mechanism is as follows:
[0039] Among them, Z” represents the output result of the cross-attention mechanism, Q = ZW q , K' = c h W' k , V' = c h W' v are the query, key, and value matrices of the image features respectively, and W q , W' k and W' v are trainable linear mapping layers. The query matrix uses the same query matrix as the self-attention mechanism. d represents the feature dimension, and T represents the transpose operation.
[0040] Furthermore, in the hairstyle generation diffusion model in the first stage,
[0041] The hairstyle generation diffusion model ∈ b is trained based on the pre-trained StableDiffusion model:
[0042] The bald image I generated by the baldness generator b is input into the pre-trained VAE encoder to obtain the latent space encoding z b = ε(I ref );
[0043] Randomly generate latent space Gaussian noise and add it to the obtained latent space encoding z b , which is input into the hairstyle generation diffusion model. After passing through the feature layer, the feature Z is obtained, and attention calculation is performed through the self-attention mechanism:
[0044] Among them, Q = ZW q , K = ZW k , V = ZW v are the query, key, and value matrices of the image features respectively, and W k and W v are trainable linear mapping layers. The query matrix Q uses the same query matrix as the hairstyle cross-attention mechanism of the hairstyle reference network. After passing through the self-attention mechanism, and then through the cross-attention mechanism Z', the obtained attentions are added: Z new = Z' + Z”, to obtain the new attention Z new , and transfer the hair accurately to the area in the source image;
[0045] After passing through the subsequent feature layer, the finally output latent space feature layer is input into the pre-trained VAE decoder to obtain a rough hairstyle-changed image in the RGB space.
[0046] Further, during the two-stage hairstyle generation training process:
[0047] For each hairstyle, a LoRA small model is trained: W = W0 + ΔW = W0 + BA, where W represents the fine-tuned weight matrix, ΔW represents the weight update amount, and the pre-trained weight is W0 ∈ R d×k , and another low-rank matrix B ∈ R d×r , A ∈ R r×k , and the rank r << min(d, k), where d and k represent the feature dimension and output dimension respectively, reducing the training burden and enabling the model to quickly capture the unique features of each hairstyle, thus significantly improving the quality of the generated images while maintaining the training efficiency.
[0048] Further, in the two-stage image processing module:
[0049] Use a large model to label the input hairstyle, that is, obtain the description words of the hairstyle;
[0050] Adopt the open-source image segmentation model SAM to segment the hair and obtain the hairstyle image I h .
[0051] Further, during the two-stage LoRA training process, the diffusion model ∈ θ Adopt the pre-trained SDv1.5 as the pre-trained weight:
[0052] Input the hairstyle description words obtained from the two-stage image processing module into the pre-trained CLIP text encoder to obtain the text encoding c t , and input it into the diffusion model through the text cross-attention mechanism;
[0053] Input the hairstyle image obtained from the two-stage image processing module into the VAE encoder to obtain the latent space encoding z h = ε(I h ), add randomly generated latent space Gaussian noise Input it into the diffusion model, and apply the weight of the hairstyle LoRA to the diffusion model;
[0054] The output of the last feature layer of the diffusion model, after passing through the VAE decoder, obtains the hairstyle image in the RGB space;
[0055] During the training process, by minimizing the loss function, use the Adam optimizer for optimization iteration. The weights of the diffusion model are fixed, and only the hairstyle LoRA is trained. The loss function for training the hairstyle LoRA is:
[0056] Among them, is the loss function, Denotes the expected value, ∈ θ+Δθ Denotes the diffusion model integrated with the hairstyle LoRA, which estimates the noise after adjusting the weights θ + Δθ. Δθ represents the weight increment for fine-tuning during the training of the LoRA model, z t Denotes the encoding result of the hairstyle image in the latent space, c t Denotes the encoding result of the text description words, and ∈ represents Gaussian noise.
[0057] Furthermore, in the second-stage image processing module:
[0058] In the inference stage, different from the training stage, the diffusion model uses a pre-trained StableDiffusion redrawing model;
[0059] The goal of the redrawing diffusion model is to redraw the image required by the user in the specified area of a picture, and ensure that the redrawn area can be harmoniously integrated with the original image. The redrawing diffusion model is used for virtual hairstyle change, aiming to strip the hairstyle from the source picture and redraw the target hairstyle trained by the hairstyle LoRA on the model, so that other parts remain in the original state and the generated hairstyle is more realistic and natural;
[0060] In the inference stage, the image processing module first performs face detection and face alignment on the input rough hairstyle-changed picture, and crops and adjusts it to the required size;
[0061] Then use the open-source image segmentation method SAM to segment the hairstyle from the source picture and the rough hairstyle-changed picture, and superimpose them to obtain the hairstyle binary mask I' m , and then expand the hairstyle binary mask to obtain the expanded hairstyle binary mask I m ;
[0062] Superimpose the hairstyle binary mask I m with the rough hairstyle-changed picture to obtain the picture I with the hairstyle blocked bg ;
[0063] Obtain the description words for this hairstyle that meet the standards during the LoRA training process.
[0064] Furthermore, in the second-stage hairstyle redrawing diffusion model:
[0065] Input the hairstyle description words into the CLIP text encoder to obtain the text encoding c t ;
[0066] Input the model picture I of the blocked hairstyle bg into the VAE encoder ε to obtain the 4-channel redrawing background latent encoding z bg = ε(I bg ), and generate random Gaussian noise in the latent space The 4-channel redrawn background latent code z bg the 4-channel latent space noise ∈, and the 1-channel hairstyle binary mask map I obtained by the two-stage image processing module m are concatenated along the channels to obtain a 9-channel input;
[0067] The 9-channel input is input into the hairstyle redrawing diffusion model integrated with hairstyle LoRA. After cyclic denoising, the feature layer output by the hairstyle redrawing diffusion model is finally passed through the VAE decoder to obtain a refined hairstyle-changing picture.
[0068] The beneficial effects of the present invention are as follows:
[0069] In the present invention, through a two-stage generation method, a relatively rough hairstyle picture is generated in the first stage, and in the second stage, it is refined through a redrawing mechanism, so that the generated hairstyle picture has higher detail fidelity. In the second stage, a hairstyle redrawing diffusion model is adopted, so that the hairstyle can be better presented in terms of complex textures and details, avoiding the blurring or detail loss that easily occur in the generation of hairstyle pictures by traditional methods.
[0070] In the present invention, aiming at the problem that traditional GAN-based methods often have difficulty in processing hairstyles with complex textures and are prone to generating artifacts or ignoring details, a diffusion model-based solution is combined with a cross-attention mechanism to better capture hairstyle details. By injecting hairstyle reference information during the generation process, the transmission of hairstyle details is ensured, and more complex hairstyle styles can be processed while maintaining high stability.
[0071] In the present invention, the LoRA model is introduced in the second stage for hairstyle fine-tuning, which reduces the number of training parameters, improves the training efficiency. The lightweight fine-tuning method can quickly adapt to the generation requirements of different hairstyles and does not require completely retraining the entire model, saving computing resources and time, and is suitable for the needs of personalized hairstyle generation. BRIEF DESCRIPTION OF THE DRAWINGS
[0072] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the following drawings are only some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.
[0073] Figure 1 It is a schematic flowchart of the virtual hairstyle-changing method of the present invention;
[0074] Figure 2 It is a schematic flowchart of one stage of the virtual hairstyle-changing method of the present invention;
[0075] Figure 3Schematic diagram of the two - stage process of the virtual hairstyle change method of the present invention;
[0076] Figure 4 Schematic diagram of the processing flow of the image processing module of the virtual hairstyle change method of the present invention;
[0077] Figure 5 Schematic diagram of the processing flow of the hairstyle redrawing diffusion model in the two - stage of the virtual hairstyle change method of the present invention. Detailed implementation manners
[0078] To make the above - mentioned objects, features and advantages of the present invention more obvious and understandable, the following will make a detailed description of the specific implementation manners of the present invention in conjunction with the accompanying drawings of the specification.
[0079] In the following description, many specific details are set forth in order to fully understand the present invention. However, the present invention can also be implemented in other ways different from those described herein. Those skilled in the art can make similar promotions without departing from the connotation of the present invention. Therefore, the present invention is not limited by the specific embodiments disclosed below.
[0080] Secondly, the so - called "one embodiment" or "embodiment" herein refers to a specific feature, structure or characteristic that can be included in at least one implementation manner of the present invention. The "in one embodiment" that appears in different places in this specification does not all refer to the same embodiment, nor is it an individual or selectively mutually exclusive embodiment with other embodiments.
[0081] Embodiment 1, referring to Figure 1 , this embodiment provides a virtual hairstyle change system based on a diffusion model, including,
[0082] A first - stage image processing module that processes the input source image and hairstyle reference image, and performs face detection, alignment and image cropping operations;
[0083] A baldness generator for generating a bald image of the person in the source image, that is, an image without hair;
[0084] A hairstyle reference network that extracts hairstyle details from the hairstyle reference image. This network extracts hairstyle information through a latent space encoding and cross - attention mechanism, and injects the hairstyle information details into the hairstyle generation diffusion model;
[0085] A hairstyle generation diffusion model that generates a rough hairstyle - changed image based on the bald image generated by the baldness generator and the hairstyle details provided by the hairstyle reference network, and captures the overall contour of the hairstyle;
[0086] A second - stage image processing module that processes the rough hairstyle - changed image, including face detection, alignment and cropping operations, and also generates a hairstyle binary mask;
[0087] The hairstyle redrawing diffusion model refines and redraws the rough hairstyle-changed pictures generated in the first stage. Combining the hairstyle features trained by the hairstyle LoRA, it redraws the hairstyle area to make the hairstyle blend harmoniously with other parts and outputs fine hairstyle pictures.
[0088] Example 2, refer to Figures 1 - 5 , this example provides a virtual hairstyle-changing method based on the diffusion model, including the following steps,
[0089] Step S1, select pictures of the same person with different hairstyles during training, such as the SketchHairSalon dataset;
[0090] Step S2, randomly select one picture as the source picture and one picture as the hairstyle reference picture from the pictures of the same person with different hairstyles, and then process them into the required format through the image processing module;
[0091] Step S3, input the source picture I src into the baldness generator to obtain the bald picture I bald ;
[0092] Step S4, input the hairstyle reference picture I ref into the hairstyle reference network to extract hairstyle-related details;
[0093] Step S5, input the bald picture I bald into the hairstyle generation diffusion model, and the extracted hairstyle details are input into the hairstyle generation diffusion model through the cross-attention mechanism, and finally obtain the rough hairstyle-changed picture;
[0094] Step S6, prepare 510 pictures of different angles (without the back) for each hairstyle as the hairstyle LoRA hairstyle training set, input the pictures in the hairstyle training set into the image processing module to obtain the hairstyle pictures and hairstyle description words;
[0095] Step S7, input the hairstyle pictures and hairstyle description words obtained in Step S6 into the diffusion model for LoRA training of the targeted hairstyle;
[0096] Step S8, generate fine hairstyle-changed pictures through the VAE decoder;
[0097] Steps S1 - S5 are the first-stage hairstyle generation steps, and Steps S6 - S8 are the second-stage hairstyle generation steps;
[0098] Since the hairstyle pictures generated in the first stage are relatively rough, a two-stage virtual hairstyle-changing picture generation method is adopted;
[0099] In the first stage, the trained hairstyle generation diffusion model is used to generate rough hairstyle-changing pictures. In the second stage, the hairstyle redrawing diffusion model is used to redraw the hairstyle area of the rough hairstyle-changing pictures to obtain the final and more refined hairstyle-changing pictures.
[0100] The roughness and refinement here are relative.
[0101] In the image processing module of the first stage:
[0102] Input the source image I src , and the hairstyle reference image I ref into the image processing module, detect the face, and then perform face alignment, crop and resize it to the required size.
[0103] The baldness generator in the first stage includes a VAE encoder ε, a baldness generation diffusion model, a baldness ControlNet τ b and a VAE decoder.
[0104] The baldness generation diffusion model is the latent space diffusion model StableDiffusion, which uses the pre-trained weights of SDv1.5. StableDiffusion is a classic latent space diffusion model (LDM, Latent Diffusion Model). LDM uses a variational autoencoder VAE. Given an input image, it is first mapped to the latent space by the encoder, so that the diffusion occurs in the latent space. In the forward process of diffusion, Gaussian noise is gradually added to the latent variable in T iterations. The reverse diffusion process is to gradually denoise, and finally it is restored to the RGB space through the decoder to obtain the final image.
[0105] The ControlNet branch creates copies of the trainable diffusion model UNet encoding blocks and intermediate blocks, and adds additional zero convolutional layers. The encoding blocks and intermediate blocks of ControlNet are added to the decoding blocks and intermediate blocks of the diffusion model. The diffusion model here is the baldness generation diffusion model, and the output of each copy block is added to the skip connections of the original diffusion model UNet.
[0106] Randomly sample 4-channel latent space Gaussian noise and input it into the baldness generation diffusion model.
[0107] Input the source image into the VAE encoder to obtain the latent space encoding z src = ε(I src ), and the latent space encoding z srcInput into the Bald ControlNet, which is a trainable copy of the decoding block and the middle block of the bald generation diffusion model. After passing through the zero convolution layer, the encoding block and the middle block of the Bald ControlNet are added to the decoding block and the middle block of the bald generation diffusion model through residual connections, and the reference information of the source image is input into the bald generation diffusion model;
[0108] Finally, the output of the bald generation diffusion model passes through the VAE decoder to output the bald image I of the person in the source image in the RGB space b ;
[0109] During the training of the bald generator, only the Bald ControlNet is trained, and the weights of the bald generation diffusion model, the VAE encoder, and the VAE decoder are fixed. The pre-trained weights of SDv1.5 are used. The Bald ControlNet participates in the training, and the initial weights also use the pre-trained weights of SDv1.5. By minimizing the loss function and using the Adam optimizer for optimization iteration, the loss function for training the bald generator is: Among them, represents the loss function, represents the expected value, that is, given the latent space encoding of the source image I src the noise sampled from the standard Gaussian distribution ∈, and the expectation at the time step t, ∈ b represents the noise estimator of the bald generation diffusion model, which is responsible for predicting the input noise distribution at each time step t, τ b (ε(I src )) represents the bald feature of the input source image generated by the Bald ControlNet τ b After passing through the VAE encoder ε to perform latent space encoding on the source image I src and processed at the time step t, ∈ represents Gaussian noise, represents the L2 norm;
[0110] In the hairstyle reference network in the first stage,
[0111] The hairstyle reference network τ h is trained based on the latent space diffusion model StableDiffusion:
[0112] Input the hairstyle reference image I ref into the pre-trained VAE encoder ε to obtain the latent space encoding z = ε(I ref );
[0113] Input the latent space encoding z into the hairstyle reference network τ h to extract the hairstyle detail feature c h, it is input into the hairstyle generation diffusion model through the cross-attention mechanism, and the calculation method of the cross-attention mechanism is as follows:
[0114] Among them, Z” represents the output result of the cross-attention mechanism, Q = ZW q , K' = c h W' k , V' = c h W' v are the query, key, and value matrices of the image features respectively, and W q , W' k and W' v are trainable linear mapping layers. The query matrix uses the same query matrix as the self-attention mechanism. d represents the feature dimension, and T represents the transpose operation;
[0115] In the hairstyle generation diffusion model in the first stage,
[0116] the hairstyle generation diffusion model ∈ b is trained based on the pre-trained StableDiffusion model:
[0117] The bald picture I generated by the bald generator b is input into the pre-trained VAE encoder to obtain the latent space encoding z b = ε(I ref );
[0118] Randomly generate latent space Gaussian noise and add it to the obtained latent space encoding z b , then input it into the hairstyle generation diffusion model. After passing through the feature layer, the feature Z is obtained, and attention calculation is performed through the self-attention mechanism:
[0119] Among them, Q = ZW q , K = ZW k , V = ZW v are the query, key, and value matrices of the image features respectively, and W k and W v are trainable linear mapping layers. The query matrix Q uses the same query matrix as the hairstyle cross-attention mechanism of the hairstyle reference network. After passing through the self-attention mechanism, and then through the cross-attention mechanism Z', the obtained attentions are added: Z new = Z' + Z”, to obtain the new attention Z new , and transfer the hair accurately to the area in the source image;
[0120] After passing through the subsequent feature layers, the finally output latent space feature layer is input into the pre-trained VAE decoder to obtain a rough hairstyle-changed picture in the RGB space;
[0121] During the two-stage hairstyle generation training process:
[0122] For each hairstyle, a LoRA small model is trained: W = W0 + ΔW = W0 + BA, where W represents the fine-tuned weight matrix, ΔW represents the weight update amount, and the pre-trained weight is W0 ∈ R d×k , and another low-rank matrix B ∈ R d×r , A ∈ R r×k , and the rank r << min(d, k), where d and k represent the feature dimension and output dimension respectively, reducing the training burden and enabling the model to quickly capture the unique features of each hairstyle, thus significantly improving the quality of the generated images while maintaining the training efficiency;
[0123] In the two-stage image processing module:
[0124] Use a large model to label the input hairstyle, that is, obtain the description words of the hairstyle;
[0125] Adopt the open-source image segmentation model SAM to segment the hair and obtain the hairstyle image I h ;
[0126] During the two-stage LoRA training process, the diffusion model ∈ θ Adopt the pre-trained SDv1.5 as the pre-trained weight:
[0127] Input the hairstyle description words obtained from the two-stage image processing module into the pre-trained CLIP text encoder to obtain the text encoding c t , and input it into the diffusion model through the text cross-attention mechanism;
[0128] Input the hairstyle image obtained from the two-stage image processing module into the VAE encoder to obtain the latent space encoding z h = ε(I h ), plus randomly generated latent space Gaussian noise Input it into the diffusion model and apply the weight of the hairstyle LoRA to the diffusion model;
[0129] The output of the last feature layer of the diffusion model, after passing through the VAE decoder, obtains the hairstyle image in the RGB space;
[0130] During the training process, by minimizing the loss function and using the Adam optimizer for optimization iteration, the weights of the diffusion model are fixed, and only the hairstyle LoRA is trained. The loss function for training the hairstyle LoRA is:
[0131] Where, is the loss function, represents the expected value, ∈ θ+ΔθThe diffusion model representing hairstyle LoRA integration estimates the noise after weight adjustment θ+Δθ, where Δθ represents the weight increment fine-tuned during LoRA model training, and z t represents the encoding result of the hairstyle picture in the latent space, c t represents the encoding result of the text description words, and ∈ represents Gaussian noise;
[0132] In the two-stage image processing module:
[0133] In the inference stage, different from the training stage, the diffusion model uses a pre-trained StableDiffusion redrawing model;
[0134] The goal of the redrawing diffusion model is to redraw the image required by the user in the user-specified area of a picture, and ensure that the redrawn area can be harmoniously integrated with the original image. The redrawing diffusion model is used for virtual hairstyle change, aiming to strip the hairstyle of the source picture and redraw the target hairstyle trained by hairstyle LoRA on the model, so that other parts remain in the original state and the generated hairstyle is more realistic and natural;
[0135] In the inference stage, the image processing module first performs face detection and face alignment on the input rough hairstyle-changed picture, and crops and adjusts it to the required size;
[0136] Then use the open-source image segmentation method SAM to segment the hairstyle from the source picture and the rough hairstyle-changed picture, and overlay to obtain the hairstyle binary mask I' m , and then dilate the hairstyle binary mask to obtain the dilated hairstyle binary mask I m ;
[0137] Overlay the hairstyle binary mask I m with the rough hairstyle-changed picture to obtain the picture I covering the hairstyle bg ;
[0138] Obtain the description words for this hairstyle that meet the standards during LoRA training;
[0139] In the two-stage hairstyle redrawing diffusion model:
[0140] Input the hairstyle description words into the CLIP text encoder to obtain the text encoding c t ;
[0141] Input the model picture I covering the hairstyle bg into the VAE encoder ε to obtain the 4-channel redrawing background latent encoding z bg = ε(I bg ), and generate random Gaussian noise in the latent space Input the 4-channel redrawing background latent encoding z bg, the 4-channel latent space noise ∈ and the 1-channel hair-style binary mask image I obtained by the two-stage image processing module m , concatenated along the channels to obtain a 9-channel input;
[0142] The 9-channel input is input into the hair-style redrawing diffusion model integrated with hair-style LoRA. After cyclic denoising, the feature layer output by the hair-style redrawing diffusion model is finally passed through the VAE decoder to obtain a refined hair-style change picture.
[0143] It should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and not to limit them. Although the present invention has been described in detail with reference to the preferred embodiments, those of ordinary skill in the art should understand that the technical solutions of the present invention can be modified or equivalently replaced without departing from the spirit and scope of the technical solutions of the present invention, and they should all be covered within the scope of the claims of the present invention.
Claims
1. A virtual hair styling method based on a diffusion model, characterized in that, Performing virtual hairstyle change using a virtual hairstyle change system based on a diffusion model, the system includes: A first-stage image processing module that processes the input source image and hairstyle reference image, and performs face detection, alignment, and image cropping operations; A baldness generator for generating a bald image of the person in the source image, that is, an image without hair; A hairstyle reference network that extracts hairstyle details from the hairstyle reference image, extracts hairstyle information through a latent space encoding and cross-attention mechanism, and injects the hairstyle information details into the hairstyle generation diffusion model; A hairstyle generation diffusion model that generates a rough hairstyle change image based on the bald image generated by the baldness generator and the hairstyle details provided by the hairstyle reference network; A second-stage image processing module that processes the rough hairstyle change image, including face detection, alignment, and cropping operations, and also generates a hairstyle binary mask; A hairstyle redrawing diffusion model that refines and redraws the rough hairstyle change image generated in the first stage, and combines the hairstyle features trained by hairstyle LoRA to redraw the hairstyle area; The virtual hairstyle change method based on the diffusion model includes the following steps Step S1, selecting pictures of the same person with different hairstyles during training; Step S2, randomly selecting one picture as the source image and one picture as the hairstyle reference image from the pictures of the same person with different hairstyles, and then processing them into the required format through the image processing module; Step S3, input the source image I src into the baldness generator to obtain the bald image I bald ; Step S4, input the hairstyle reference picture I ref into the hairstyle reference network to extract hairstyle-related details; Step S5, input the bald picture I bald into the hairstyle generation diffusion model, and the extracted hairstyle details are input into the hairstyle generation diffusion model through the cross-attention mechanism, and finally a rough picture with a changed hairstyle is obtained; Step S6, preparing pictures at different angles for each hairstyle as the hairstyle LoRA hairstyle training set, inputting the pictures in the hairstyle training set into the image processing module to obtain hairstyle pictures and hairstyle description words; Step S7, inputting the hairstyle pictures and hairstyle description words obtained in Step S6 into the diffusion model for LoRA training of the targeted hairstyle; Step S8, generating a fine hairstyle change image through the VAE decoder; Steps S1 - S5 are the first-stage hairstyle generation steps, and Steps S6 - S8 are the second-stage hairstyle generation steps; During the second-stage hairstyle generation training process: A LoRA small model is trained for each hairstyle: W = W0 + ΔW = W0 + BA, where W represents the fine-tuned weight matrix, ΔW represents the weight update amount, and the pre-trained weight is W0 ∈ R d×k , and another low-rank matrix B ∈ R d×r , A ∈ R r×k , and the rank r << min(d, k), where d and k represent the feature dimension and the output dimension respectively; In the second-stage image processing module: Using a large model to label the input hairstyle, that is, obtaining a description word for the hairstyle; Use the open-source image segmentation model SAM to segment the hair and obtain the hairstyle image I h ; During the two-stage LoRA training process, the diffusion model ∈ θ uses the pre-trained SDv1.5 as the pre-trained weights: Input the hairstyle descriptor obtained by the two-stage image processing module into the pre-trained CLIP text encoder to obtain the text encoding c t , and input it into the diffusion model through the text cross-attention mechanism; Input the hairstyle image obtained by the two-stage image processing module into the VAE encoder to obtain the latent space encoding z h = ε(I h ), add randomly generated latent space Gaussian noise Input it into the diffusion model, and apply the weights of the hairstyle LoRA to the diffusion model; The output of the last feature layer of the diffusion model, after passing through the VAE decoder, obtains a hairstyle picture in the RGB space; During the training process, by minimizing the loss function, using the Adam optimizer for optimization iteration, the weights of the diffusion model are fixed, and only the hairstyle LoRA is trained. The loss function for training the hairstyle LoRA is: Among them, is the loss function, represents the expected value, ∈ θ+Δθ represents the diffusion model integrated with the hairstyle LoRA, which estimates the noise after adjusting the weights θ + Δθ. Δθ represents the weight increment fine-tuned during the training of the LoRA model, z t represents the encoding result of the hairstyle image in the latent space, c t represents the encoding result of the text description words, ∈ represents Gaussian noise, and t represents the time step; In the second-stage image processing module: In the inference stage, the image processing module first performs face detection and face alignment on the input rough hairstyle change image, and crops and adjusts it to the required size; Then, use the open-source image segmentation method SAM to segment the hairstyle from the source image and the rough hairstyle-changed image, and superimpose them to obtain the hairstyle binary mask I'. m Then, dilate the hairstyle binary mask to obtain the dilated hairstyle binary mask I m ; Overlay the hairstyle binary mask I m with the rough hairstyle-changing image to obtain the image I that occludes the hairstyle bg ; Obtaining the description word for this hairstyle that meets the standards during the LoRA training process; In the second-stage hairstyle redrawing diffusion model: Input the hairstyle descriptor into the CLIP text encoder to obtain the text encoding c t ; Input the model picture I that obscures the hairstyle bg into the VAE encoder ε to obtain the 4-channel redrawn background latent code z bg = ε(I bg ), generating a random Gaussian noise in the latent space Concatenate the 4-channel redrawn background latent code z bg , the 4-channel latent space noise ∈, and the 1-channel hairstyle binary mask map I obtained by the second-stage image processing module m along the channels to obtain a 9-channel input; Inputting the 9-channel input into the hairstyle redrawing diffusion model integrated with hairstyle LoRA, through cyclic denoising, the output feature layer of the hairstyle redrawing diffusion model, and finally passing through the VAE decoder to obtain a fine hairstyle change image.
2. The virtual hairstyle transformation method based on a diffusion model according to claim 1, characterized in that, In the first-stage image processing module: Input the source image I src , the hairstyle reference image I ref into the image processing module, detect the face, then perform face alignment, and crop and resize it to the required size; The baldness generator in the first stage includes a VAE encoder ε, a baldness generation diffusion model, and a baldness ControlNet τ b and a VAE decoder; The baldness generation diffusion model is the latent space diffusion model StableDiffusion, which uses the pre-trained weights of SDv1.
5. The latent space diffusion model LDM uses a variational autoencoder (VAE). Given an input image, it first passes through an encoder to map the image to the latent space, enabling the diffusion to occur in the latent space. During the forward process of diffusion, Gaussian noise is gradually added to the latent variables in T iterations. The reverse diffusion process is to gradually denoise, and finally, it passes through a decoder to be restored to the RGB space to obtain the final image; The ControlNet branch creates copies of the trainable diffusion model UNet encoding blocks and intermediate blocks and adds additional zero convolutional layers. The encoding blocks and intermediate blocks of ControlNet are added to the decoding blocks and intermediate blocks of the diffusion model. Here, the diffusion model is the baldness generation diffusion model, and the output of each copy block is added to the skip connections of the original diffusion model UNet; Randomly sample Gaussian noise in the latent space of 4 channels Input it into the baldness generation diffusion model; Input the source image into the VAE encoder to obtain the latent space encoding z src = ε(I src ), input the latent space encoding z src into the Baldness ControlNet, which is a trainable copy of the decoding block and the intermediate block of the baldness generation diffusion model. After passing through zero convolutional layers, add the encoding block and the intermediate block of the Baldness ControlNet to the decoding block and the intermediate block of the baldness generation diffusion model through residual connections, and input the reference information of the source image into the baldness generation diffusion model; The output of the final baldness generation diffusion model, after passing through the VAE decoder, outputs the baldness image I of the person in the source image in the RGB space b ; During the process of training the baldness generator, only the baldness ControlNet is trained, and the weights of the baldness generation diffusion model, VAE encoder, and VAE decoder are fixed. The pre-trained weights of SDv1.5 are adopted. The baldness ControlNet participates in the training, and the initial weights also use the pre-trained weights of SDv1.
5. By minimizing the loss function and using the Adam optimizer for optimization iteration, the loss function for training the baldness generator is as follows: Among them, denotes the loss function, denotes the expected value, that is, given the latent space encoding of the source image I src and the noise ∈ sampled from the standard Gaussian distribution and the expectation at the time step t, ∈ b denotes the noise estimator of the baldness generation diffusion model, which is responsible for predicting the noise distribution of the input at each time step t, τ b (ε(I src )) represents the baldness feature of the input source image generated by the baldness ControlNet τ b After the latent space encoding of the source image I is performed by the VAE encoder ε and processed at the time step t, ∈ represents Gaussian noise, src and denotes the L2 norm.
3. The virtual hairstyle transformation method based on a diffusion model according to claim 2, wherein In the hairstyle reference network in the first stage, Hair style reference network τ h is trained based on the latent space diffusion model StableDiffusion: Input the hairstyle reference picture I ref into the pre-trained VAE encoder ε to obtain the latent space encoding z = ε(I ref ); Input the latent space encoding z into the hairstyle reference network τ h to extract the hairstyle detail feature c h and input it into the hairstyle generation diffusion model through the cross-attention mechanism. The calculation method of the cross-attention mechanism is as follows: Among them, Z” represents the output result of the cross-attention mechanism, and Q = ZW q , K' = c h W' k , V' = c h W' v are the query, key, and value matrices of the image features respectively, Z represents the features obtained through the feature layer, and W q , W' k and W' v are trainable linear mapping layers. The query matrix uses the same query matrix as the self-attention mechanism. d represents the feature dimension, and T represents the transpose operation.
4. The virtual hairstyle transformation method based on a diffusion model according to claim 3, wherein In the hairstyle generation diffusion model in the first stage, Hair style generation diffusion model ∈ b is trained based on the pre-trained StableDiffusion model: The bald image I generated by the baldness generator b is input into the pre-trained VAE encoder to obtain the latent space encoding z b ; Randomly generate Gaussian noise in the latent space Add the obtained latent space encoding z b , input it into the hairstyle generation diffusion model, obtain feature Z through the feature layer, and perform attention calculation through the self-attention mechanism: where Q = ZW q , K = ZW k , V = ZW v are the query, key, and value matrices of the image features respectively, and W k and W v are trainable linear mapping layers. The query matrix Q uses the same query matrix as the hair style cross-attention mechanism of the hair style reference network. After passing through the self-attention mechanism and then through the cross-attention mechanism Z'', the obtained attentions are added: Z new = Z' + Z'', to obtain the new attention Z new ; After passing through subsequent feature layers, the finally output latent space feature layer is input into the pre-trained VAE decoder to obtain a rough hairstyle-changed image in the RGB space.
Citation Information
Patent Citations
Image style migration method and device based on stable diffusion
CN116630464A