A virtual hair replacement method and system for diffusion model combined with Transformer architecture

By combining the diffusion model of Transformer architecture, the baldness generator and hairstyle generation model are used to generate virtual hairstyle pictures, which solves the problem of difficulty in generating accurate virtual hairstyles in the existing technology, and achieves a more detailed and natural hairstyle generation effect.

CN119251345BActive Publication Date: 2025-05-06MEIZHONG (TIANJIN) TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411764606.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-12-04
Publication Date
2025-05-06
Estimated Expiration
2044-12-04

AI Technical Summary

Technical Problem

The existing virtual hairstyle changing technology is difficult to generate accurate virtual hairstyles, which are prone to artifacts, and the details of hairstyles are not restored well.

Method used

Using a diffusion model combining Transformer architecture, virtual hairstyle changes pictures are generated through baldness generator and hairstyle generation model. The bald generator includes VAE encoder, bald generator model, bald controlNet and VAE decoder. The hairstyle generation model also adopts multiple series-connected diffusion models based on Transformer architecture, and introduces a hairstyle reference network to inject hairstyle information through the hairstyle cross attention module.

Benefits of technology

A more accurate virtual hairstyle generation is achieved, avoiding the influence of the user's original picture hairstyle, and the generated hairstyle is more detailed and natural, reducing the appearance of artifacts.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119251345B_ABST
    Figure CN119251345B_ABST
Patent Text Reader

Abstract

The present invention discloses a virtual hairstyle changing method and system combining a diffusion model with a Transformer architecture, and relates to the technical field of virtual hairstyle changing. The technical points of the present invention include: obtaining a source image with hair; processing the source image using a baldness generator to generate a baldness image; generating a hairstyle changing image using a hairstyle generation model according to a hairstyle reference image and a baldness image; wherein, the baldness image is first generated and then the hairstyle changing image is generated, thereby avoiding the influence of the user's original image hairstyle during the generation process; the baldness generator and the hairstyle generation model both adopt a diffusion model based on the Transformer architecture, and at the same time introduce a hairstyle reference network, and inject hairstyle information through a hairstyle cross-attention module, so that the generated hairstyle is more refined. The present invention provides users with an efficient and convenient virtual hairstyle changing solution, and also brings an innovative service model to the hairdressing industry.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of virtual hair-changing technology, and in particular to a diffusion model virtual hair-changing method and system combined with a Transformer architecture. Background Art

[0002] The right hairstyle can well reflect a person's style, and it also plays a very important role in the overall outfit. As the pursuit of beauty continues to deepen, people pay more and more attention to their hairstyle choices. If you can preview the effect of a new hairstyle before trying it, it will greatly reduce unsatisfactory haircut experiences.

[0003] Traditional hairstyle changing technology is usually done with the help of photo editing tools. It not only requires finding a picture that matches the angle of the new hairstyle and the person's photo, but also takes a certain amount of time to edit the picture to make it look real and natural. With the development of artificial intelligence technology, virtual hairstyle changing technology has emerged in the public eye. The focus of virtual hairstyle changing technology is to naturally transform the target hairstyle to the user's photo while maintaining the details of the hairstyle and the recognizability of the user's face. In recent years, most methods are based on generative adversarial networks (GANs), which are not easy to restore the details of the hairstyle and are prone to artifacts. Summary of the invention

[0004] Therefore, the present invention provides a virtual hairstyle changing method and system of a diffusion model combined with a Transformer architecture, so as to solve the problem that the existing methods are not accurate enough in generating virtual hairstyles.

[0005] According to one aspect of the present invention, a virtual transformation method of a diffusion model combined with a Transformer architecture is proposed, the method comprising:

[0006] Get the source image with hair;

[0007] Use the baldness generator to process the source image to generate a baldness image;

[0008] Based on the hairstyle reference pictures and bald pictures, the hairstyle generation model is used to generate the hairstyle change pictures.

[0009] Furthermore, after obtaining the source image with hair, image processing is performed on the source image to obtain a source image that meets the size requirement.

[0010] Furthermore, the baldness generator includes a VAE encoder, a baldness generation model, a baldness ControlNet, and a VAE decoder; wherein the baldness generation model and the baldness ControlNet both include multiple serially connected diffusion models based on the Transformer architecture, and the baldness ControlNet is a trainable copy of the baldness generation model.

[0011] Further, the processing of the source image by using the baldness generator to generate the baldness image includes:

[0012] Input the source image into the VAE encoder to obtain the latent space encoding;

[0013] The latent space code is input into the baldness ControlNet, and after block processing and linear layer processing, it is input into multiple serially connected diffusion models based on the Transformer architecture for processing to obtain source image reference information; the source image reference information is input into the baldness generation model;

[0014] Randomly generate Gaussian noise in latent space, and input the noise into the baldness generation model, after block processing and linear layer processing, obtain a feature map, and input the feature map and the source image reference information output by the baldness ControlNet into multiple serially connected diffusion models based on the Transformer architecture for processing, and the obtained output is processed by a multi-layer perceptron and then deblocked;

[0015] The result of the deblocking process is input into the VAE decoder to obtain the bald picture corresponding to the source picture.

[0016] Furthermore, the hairstyle generation model includes a plurality of diffusion models based on the Transformer architecture connected in series; and generating the hairstyle change picture using the hairstyle generation model according to the hairstyle reference picture and the bald picture includes:

[0017] Input the hairstyle reference image and the bald image generated by the baldness generator into the pre-trained VAE encoder respectively to obtain the corresponding latent space encoding;

[0018] Inputting the latent space encoding corresponding to the hairstyle reference image into the hairstyle reference network for processing to obtain hairstyle detail features; and inputting the hairstyle detail features into the hairstyle generation model;

[0019] Randomly generate Gaussian noise in latent space, and input the noise and the latent space code corresponding to the bald picture into the hairstyle generation model, obtain a feature map after block processing and linear layer processing, input the feature map and the hairstyle detail features output by the hairstyle reference network into multiple serially connected diffusion models based on the Transformer architecture for processing, and the obtained output is processed by a multi-layer perceptron and then deblocked;

[0020] The result of the deblocking process is input into the VAE decoder to obtain the transformed image corresponding to the source image.

[0021] Furthermore, the step of inputting the latent space code corresponding to the hairstyle reference image into the hairstyle reference network for processing to obtain the hairstyle detail features includes: inputting the latent space code corresponding to the hairstyle reference image into multiple serially connected diffusion models based on the Transformer architecture after block processing and linear layer processing to obtain the hairstyle detail features.

[0022] Furthermore, the baldness generator and the hairstyle generation model are both pre-trained models, the baldness generator is trained separately, and the hairstyle reference network participates in the training process of the hairstyle generation model; wherein the loss function in the training process of the baldness generator is as follows:

[0023]

[0024] in, represents Gaussian noise; represents the VAE encoder; Represents the diffusion model based on the Transformer architecture in the bald generator; It means bald ControlNet; Indicates the source image; represents the latent space encoding; t represents the time step; express Distribution of expectations;

[0025] The loss function during the hairstyle generation model training process is as follows:

[0026]

[0027] in, Represents the diffusion model based on the Transformer architecture in the hairstyle generation model; represents the hairstyle reference network; They represent the hairstyle reference pictures and bald pictures respectively; represents the latent space encoding; t represents the time step; express Distribute expectations.

[0028] Furthermore, the diffusion model based on the Transformer architecture is divided into an encoding block and a decoding block; wherein the encoding block is used to compress the input image to obtain features of different levels of the image, and the encoding block includes a self-attention module, a cross-attention module, and a forward propagation network; the decoding block is used to restore the image size, and the decoding block includes a self-attention module, a cross-attention module, a forward propagation network, and a jump module.

[0029] Furthermore, the baldness ControlNet adds an additional zero convolution layer at the connection with each decoding block of the baldness generative model; and adds the output of each baldness ControlNet to the skip connection of the decoding block of the baldness generative model.

[0030] According to another aspect of the present invention, a diffusion model virtual hair changing system combined with a Transformer architecture is proposed, the system comprising: a source image acquisition module configured to acquire a source image with hair;

[0031] A bald picture generation module, configured to process a source picture using a bald generator to generate a bald picture;

[0032] The hairstyle change picture generation module is configured to generate a hairstyle change picture using a hairstyle generation model based on a hairstyle reference picture and a bald picture.

[0033] The present invention has the following technical effects:

[0034] The present invention proposes a virtual hairstyle changing method and system combining a diffusion model with a Transformer architecture. First, a source image with hair is obtained; then the source image is processed by a baldness generator to generate a baldness image; finally, a hairstyle changing image is generated by a hairstyle generation model according to a hairstyle reference image and a baldness image. Among them, a baldness image is generated first, and then a hairstyle changing image is generated to avoid the influence of the user's original image hairstyle during the generation process; the baldness generator and the hairstyle generation model both adopt a diffusion model based on the Transformer architecture, and at the same time introduce a hairstyle reference network, and inject hairstyle information through a hairstyle cross-attention module to make the generated hairstyle more refined. The present invention can provide users with an efficient and convenient virtual hairstyle changing solution, and also bring an innovative service model to the hairdressing industry; the present invention can bring a new hairdressing experience to more people and promote the development of personalized hairdressing services. BRIEF DESCRIPTION OF THE DRAWINGS

[0035] In order to more clearly illustrate the specific implementation methods of the present invention or the technical solutions in the prior art, the drawings required for use in the specific implementation methods or the description of the prior art will be briefly introduced below. Obviously, the drawings described below are some implementation methods of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative work.

[0036] Figure 1 It is a flowchart of a virtual hair transformation method of a diffusion model combined with a Transformer architecture provided by an embodiment of the present invention.

[0037] Figure 2It is a complete flow chart of a virtual hair transformation method of a diffusion model combined with a Transformer architecture provided by an embodiment of the present invention.

[0038] Figure 3 It is a flowchart of a baldness generator generating a baldness picture in an embodiment of the present invention.

[0039] Figure 4 It is a flowchart of the hairstyle generation model generating a hairstyle change picture in an embodiment of the present invention.

[0040] Figure 5 It is a flowchart of the training process of the baldness generator and the hairstyle generation model in the embodiment of the present invention.

[0041] Figure 6 It is a structural diagram of a diffusion model virtual transformation system combined with a Transformer architecture provided by an embodiment of the present invention. DETAILED DESCRIPTION

[0042] In order to make the purpose, technical solution and advantages of the present invention clearer, the technical solution of the present invention will be described clearly and completely below. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work belong to the scope of protection of the present invention.

[0043] The embodiment of the present invention proposes a virtual transformation method of a diffusion model combined with a Transformer architecture, such as Figure 1 As shown, the method includes:

[0044] S1. Get the source image with hair;

[0045] S2, using the baldness generator to process the source image to generate a baldness image;

[0046] S3. Generate a hairstyle change picture using a hairstyle generation model according to the hairstyle reference picture and the bald picture.

[0047] Figure 2 A complete flow chart of a virtual hair transformation method of a diffusion model combined with a Transformer architecture provided by an embodiment of the present invention is shown.

[0048] The method starts from S1. In S1, a source image with hair is obtained.

[0049] According to an embodiment of the present invention, after acquiring a source image of a person with hair, the source image may be processed to obtain a source image that meets the size requirements. For example, the source image is a person image uploaded by a user. , first detect the face, then align the face, and then crop and resize it to the required size.

[0050] Then, S2 is executed, in which the source image is processed by the baldness generator to generate a baldness image. The baldness generator includes a VAE encoder, a baldness generation model, a baldness ControlNet, and a VAE decoder; the baldness generation model and the baldness ControlNet both include multiple serially connected diffusion models based on the Transformer architecture, and the baldness ControlNet is a trainable copy of the baldness generation model.

[0051] According to an embodiment of the present invention, in order to avoid the influence of the user's own hairstyle, the baldness generator will first generate a picture of the user without hair, and then input the generated baldness picture into the hairstyle generation model to generate the final hairstyle picture. The baldness generator includes a VAE encoder - , Baldness Generative Model - , Bald ControlNet- and a VAE decoder. Its function is to generate a bald picture of the person in the source picture, that is, a picture without hair. In this way, in the subsequent hairstyle generation process, the influence of the hairstyle of the user's original picture can be avoided.

[0052] Usually, diffusion models use the latent space diffusion model - Stable Diffusion, which is the classic latent space diffusion model (LDM, Latent Diffusion Model). LDM uses variational autoencoder VAE. Given an input image, the image is first mapped to the latent space through the encoder, so that diffusion is carried out in the latent space. In the forward process of diffusion, Gaussian noise will be gradually added to the latent variable in T iterations. The reverse diffusion process is to gradually remove noise, and finally restore it to RGB space through the decoder to obtain the final image.

[0053] The backbone structure of the diffusion model in the embodiment of the present invention adopts Transformer, and this architecture is called DiT (Diffusion Transformer, a diffusion model based on Transformer architecture). Compared with the traditional diffusion model structure, the Transformer structure has better effects in maintaining spatial features and generating stability. At the same time, the Transformer structure has strong scalability. The more high-quality training data, the more information can be learned; the larger the architecture, the better the generation effect. Diffusion in the latent space at the same time can save a lot of calculations and obtain the same good generation effect.

[0054] Baldness Generative Model - There are 32 DiT modules in it, of which there are two types of DiT modules, namely encoding block and decoding block. Both the encoding block and the decoding block contain self-attention module, cross-attention module and forward propagation network. The decoding block additionally contains a skip module to add the information of the encoding block to the decoding block. The encoding block will gradually compress the size of the image and obtain the features of different levels of the image; the decoding block will gradually restore the size of the image.

[0055] The ControlNet branch of DiT has the same effect as the ControlNet branch of the traditional diffusion model, creating a trainable copy of some DiT modules (i.e., the bald ControlNet- ), and add an additional zero convolution layer at the connection with each decoding block of the bald generative model. When the bald ControlNet starts training, the weights and biases of the zero convolution layer are initialized to zero, so the feedforward process is the same as the process without the ControlNet branch. Before any optimization, there will be no impact on the deep neural features; the output of each replica block is added to the jump connection of the bald generative model, so that the feature map output by the encoding block of the bald ControlNet is added to the feature map output by the decoding block of the DiT module in the bald generative model.

[0056] In this embodiment, preferably, Figure 3 As shown, the source image is processed by the baldness generator, and the process of generating a baldness image includes:

[0057] S21. Input the source image into the VAE encoder to obtain the latent space encoding ; Encode the latent space After block processing, it is processed through a linear layer and input into the baldness ControlNet, which is a trainable copy of the baldness generation model. After a zero convolution layer, the baldness ControlNet is connected to the baldness generation model through a residual connection, and the reference information of the source image is input into the baldness generation model.

[0058] S22, randomly sampled 4-channel latent space Gaussian noise , input into the baldness generation model, and after block division, the input Blocking ( are the width and height of the input image respectively) blocks, It is usually set to 2, and then processed by a linear layer to obtain a feature map; then it passes through multiple DiT modules (that is, multiple diffusion models based on the Transformer architecture connected in series), and the output is processed by a multi-layer perceptron (MLP) and then de-blocked, that is, the blocks are rearranged into their original spatial layout.

[0059] S23, then pass through the VAE decoder to output the bald picture of the person in the source picture in RGB space .

[0060] In this embodiment, the baldness generator and the hairstyle generation model are both pre-trained models, wherein only the baldness ControlNet is trained during the training of the baldness generator. The weights of the baldness generation model, VAE encoder, and VAE decoder are fixed, and the pre-trained weights of the mixed-element DiT are used. The baldness ControlNet participates in the training, and the initial weights also use the pre-trained weights of the mixed-element DiT. By minimizing the loss function, the Adam optimizer is used for optimization iteration. The loss function for training the baldness generator is as follows:

[0061]

[0062] in, represents Gaussian noise; represents the VAE encoder; Represents the diffusion model based on the Transformer architecture in the bald generator; It means bald ControlNet; Indicates the source image; represents the latent space encoding; t represents the time step; express Distribution of expectations;

[0063] Then, S3 is executed, in which a hairstyle change picture is generated using a hairstyle generation model according to the hairstyle reference picture and the bald picture.

[0064] According to an embodiment of the present invention, Figure 4 As shown, the process of generating a hairstyle change picture is as follows:

[0065] S31. Use the hairstyle reference picture And bald pictures generated by baldness generator Input the pre-trained VAE encoder respectively to obtain the corresponding latent space encoding and Here, the hairstyle reference picture is a pre-processed picture with a specific hairstyle. For example, a user can select a hairstyle picture of his / her own preference from multiple hairstyle reference pictures in the hairstyle library.

[0066] S32, encode the latent space corresponding to the hairstyle reference image Input hairstyle reference network Processing is performed to obtain detailed features of the hairstyle ; and the hairstyle details feature Input into the hairstyle generation model;

[0067] S33, Randomly generate Gaussian noise in latent space , and encode the latent space corresponding to the noise and bald picture The hair style generation model is input together, and after block processing and linear layer processing, a feature map is obtained. The feature map and the hair style detail features output by the hair style reference network are combined. The data is then fed into multiple serially connected diffusion models based on the Transformer architecture (i.e., DiT modules) for denoising. The output is processed by a multi-layer perceptron and then deblocked to restore it to the input size.

[0068] S34. Input the result after the deblocking process into the pre-trained VAE decoder to obtain a hair-changing picture corresponding to the RGB space source picture.

[0069] In this embodiment, the hairstyle reference network The structure of is the same as the hairstyle generation model, which is also trained based on the pre-trained weights of Hunyuan DiT. Figure 4 As shown, the specific process of S32 includes:

[0070] S321. Encoding the latent space It is divided into blocks and then processed by the linear layer;

[0071] S322, after being processed by the linear layer, is input into multiple serially connected diffusion models based on the Transformer architecture to extract the detailed features of the hairstyle , and is input into the hairstyle generation model through its cross-attention module.

[0072] In this embodiment, the hairstyle generation model It is trained based on the pre-trained weights of the mixed DiT. Its function is to generate pictures of users with changed hairstyles. During the training of the bald generator, both the hairstyle reference network and the hairstyle generation model participate in the training. The weights of the VAE encoder and the VAE decoder are fixed, and the hairstyle reference network and the hairstyle generation model use the pre-trained weights of the mixed DiT. By minimizing the loss function, the Adam optimizer is used for optimization iteration. The loss function for the hairstyle generation model training is as follows:

[0073]

[0074] in, Represents the diffusion model based on the Transformer architecture in the hairstyle generation model; represents the hairstyle reference network; They represent the hairstyle reference pictures and bald pictures respectively; represents the latent space encoding; t represents the time step; express Distribute expectations.

[0075] In the process of pre-training the baldness generator and the hairstyle generation model, the baldness generator generates bald pictures from pictures with hair, and the hairstyle generation model generates pictures with changed hairstyles from bald pictures. Therefore, the training data set selects bald pictures and pictures with hair of the same person; in the inference stage, the hairstyle reference picture and the source picture may not be pictures of the same person. It should be noted that during the training process, the source picture , Hairstyle Reference Pictures And bald pictures Both require image processing to crop and resize to the required size. Figure 5 A flowchart of the training process of the baldness generator and the hairstyle generation model is shown.

[0076] The present invention uses the Transformer architecture to provide users with an efficient, economical and highly realistic hairstyle selection service. Through this service, users can select their favorite styles from a variety of hairstyles and use the virtual try-on function to directly preview the effects of these hairstyles on their own heads. This service not only allows users to preview and create new virtual images, but also helps avoid disappointment caused by hairstyles that are not as expected. Through precise simulation technology, it is ensured that what users see is what they will get, and the hairstyle effects are lifelike, bringing users an unprecedented personalized experience.

[0077] Another embodiment of the present invention provides a virtual transformation system of a diffusion model combined with a Transformer architecture, such as Figure 6 As shown, the system includes:

[0078] A source picture acquisition module 610, configured to acquire a source picture with hair;

[0079] A bald picture generation module 620, which is configured to process the source picture using a bald generator to generate a bald picture;

[0080] The hairstyle change picture generation module 630 is configured to generate a hairstyle change picture using a hairstyle generation model according to the hairstyle reference picture and the bald picture.

[0081] In this embodiment, optionally, the system further includes an image processing module 640, which is configured to perform image processing on the source image after acquiring the source image with hair, so as to obtain a source image that meets the size requirements.

[0082] In this embodiment, optionally, the baldness generator in the baldness picture generation module 620 includes a VAE encoder, a baldness generation model, a baldness ControlNet, and a VAE decoder; wherein the baldness generation model and the baldness ControlNet both include multiple serially connected diffusion models based on the Transformer architecture, and the baldness ControlNet is a trainable copy of the baldness generation model. The method uses the baldness generator to process the source image to generate the baldness image, including: inputting the source image into the VAE encoder to obtain a latent space code; inputting the latent space code into the baldness ControlNet, and after block processing and linear layer processing, inputting it into multiple diffusion models based on the Transformer architecture in series for processing; inputting the output source image reference information into the baldness generation model; randomly generating latent space Gaussian noise, and inputting the noise into the baldness generation model, after block processing and linear layer processing, obtaining a feature map, and inputting the feature map and the source image reference information output by the baldness ControlNet into multiple diffusion models based on the Transformer architecture in series, and the obtained output is processed by a multi-layer perceptron and then deblocked; inputting the result after the deblocking processing into the VAE decoder to obtain the baldness image corresponding to the source image.

[0083] In this embodiment, optionally, the hairstyle generation model in the hairstyle change picture generation module 630 includes multiple diffusion models based on the Transformer architecture connected in series; the generation of the hairstyle change picture using the hairstyle generation model according to the hairstyle reference picture and the bald picture includes: inputting the hairstyle reference picture and the bald picture generated by the bald generator into the pre-trained VAE encoder respectively to obtain the corresponding latent space code; inputting the latent space code corresponding to the hairstyle reference picture into the hairstyle reference network for processing to obtain the hairstyle detail feature; and inputting the hairstyle detail feature into the hairstyle generation model; randomly generating latent space Gaussian noise, and inputting the noise and the latent space code corresponding to the bald picture into the hairstyle generation model together, after block processing and linear layer processing, a feature map is obtained, and the feature map and the hairstyle detail feature output by the hairstyle reference network are input into multiple diffusion models based on the Transformer architecture connected in series, and the obtained output is processed by a multi-layer perceptron and then de-blocked; the result after de-blocking processing is input into the VAE decoder to obtain the hairstyle change picture corresponding to the source picture.

[0084] It should be noted that the terms used in the present invention are only for describing specific embodiments, rather than limiting the scope of the present application. As shown in the present specification, unless the context clearly indicates an exception, the words "one", "a", "a kind of" and / or "the" do not specifically refer to the singular, but may also include the plural. The terms "include", "comprise" or any other variant thereof are intended to cover non-exclusive inclusion, so that the process, method or device including a series of elements includes not only those elements, but also includes other elements not explicitly listed, or also includes elements inherent to such process, method or device. In the absence of more restrictions, the elements defined by the sentence "include one..." do not exclude the presence of other identical elements in the process, method or device including the elements.

[0085] The above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit the same. Although the present invention has been described in detail with reference to the above embodiments, those skilled in the art should understand that the technical solutions described in the above embodiments may still be modified, or some or all of the technical features may be replaced by equivalents. However, these modifications or replacements do not deviate the essence of the corresponding technical solutions from the technical solutions of the embodiments of the present invention.

Claims

1. A virtual transformation method of a diffusion model combined with a Transformer architecture, characterized in that: include: Get the source image with hair; Performing image processing on the source image to obtain a source image that meets size requirements; The source image is processed by a baldness generator to generate a baldness image; the baldness generator includes a VAE encoder, a baldness generation model, a baldness ControlNet, and a VAE decoder; wherein the baldness generation model and the baldness ControlNet both include a plurality of diffusion models based on a Transformer architecture connected in series, and the baldness ControlNet is a trainable copy of the baldness generation model; the source image is processed by the baldness generator to generate a baldness image, including: Input the source image into the VAE encoder to obtain the latent space code; input the latent space code into the baldness ControlNet, and after block processing and linear layer processing, input it into multiple diffusion models based on the Transformer architecture in series for processing to obtain the source image reference information; input the source image reference information into the baldness generation model; randomly generate latent space Gaussian noise, and input the noise into the baldness generation model, after block processing and linear layer processing, obtain the feature map, and input the feature map and the source image reference information output by the baldness ControlNet into multiple diffusion models based on the Transformer architecture in series for processing, and the obtained output is processed by the multi-layer perceptron and then deblocked; input the result after deblocking processing into the VAE decoder to obtain the baldness image corresponding to the source image; Based on the hairstyle reference pictures and bald pictures, the hairstyle generation model is used to generate the hairstyle change pictures.

2. According to claim 1, a diffusion model virtual transformation method combined with Transformer architecture is characterized in that: The hairstyle generation model includes multiple diffusion models based on the Transformer architecture connected in series; The step of generating a hairstyle change picture using a hairstyle generation model according to the hairstyle reference picture and the bald picture comprises: Input the hairstyle reference image and the bald image generated by the baldness generator into the pre-trained VAE encoder respectively to obtain the corresponding latent space encoding; Inputting the latent space encoding corresponding to the hairstyle reference image into the hairstyle reference network for processing to obtain hairstyle detail features; and inputting the hairstyle detail features into the hairstyle generation model; Randomly generate Gaussian noise in latent space, and input the noise and the latent space code corresponding to the bald picture into the hairstyle generation model, obtain a feature map after block processing and linear layer processing, input the feature map and the hairstyle detail features output by the hairstyle reference network into multiple serially connected diffusion models based on the Transformer architecture for processing, and the obtained output is processed by a multi-layer perceptron and then deblocked; The result of the deblocking process is input into the VAE decoder to obtain the transformed image corresponding to the source image.

3. According to claim 2, a diffusion model virtual transformation method combined with a Transformer architecture is characterized in that: The step of inputting the latent space code corresponding to the hairstyle reference image into the hairstyle reference network for processing to obtain the hairstyle detail features comprises: inputting the latent space code corresponding to the hairstyle reference image into a plurality of serially connected diffusion models based on the Transformer architecture after block processing and linear layer processing to obtain the hairstyle detail features.

4. According to claim 3, a diffusion model virtual transformation method combined with a Transformer architecture is characterized in that: The baldness generator and the hairstyle generation model are both pre-trained models. The baldness generator is trained separately, and the hairstyle reference network participates in the training process of the hairstyle generation model. The loss function in the training process of the baldness generator is as follows: ; in, represents Gaussian noise; represents the VAE encoder; Represents the diffusion model based on the Transformer architecture in the bald generator; It means bald ControlNet; Indicates the source image; represents the latent space encoding; t represents the time step; express Distribution of expectations; The loss function during the hairstyle generation model training process is as follows: ; in, Represents the diffusion model based on the Transformer architecture in the hairstyle generation model; represents the hairstyle reference network; They represent the hairstyle reference pictures and bald pictures respectively; express Distribute expectations.

5. A virtual transformation method of a diffusion model combined with a Transformer architecture according to any one of claims 1 to 4, characterized in that: The diffusion model based on the Transformer architecture is divided into an encoding block and a decoding block; wherein the encoding block is used to compress the input image to obtain features of different levels of the image, and the encoding block includes a self-attention module, a cross-attention module, and a forward propagation network; the decoding block is used to restore the image size, and the decoding block includes a self-attention module, a cross-attention module, a forward propagation network, and a jump module.

6. According to claim 5, a diffusion model virtual transformation method combined with a Transformer architecture is characterized in that: The bald ControlNet adds an additional zero convolution layer at the connection with each decoding block of the bald generative model; And add the output of each bald ControlNet to the skip connection of the decoding block of the bald generative model.

7. A diffusion model virtual hair replacement system combined with Transformer architecture, characterized in that: include: A source picture acquisition module, configured to acquire a source picture with hair; perform image processing on the source picture to acquire a source picture that meets the size requirements; A bald picture generation module is configured to use a bald generator to process a source picture to generate a bald picture; the bald generator includes a VAE encoder, a bald generation model, a bald ControlNet, and a VAE decoder; wherein the bald generation model and the bald ControlNet both include multiple diffusion models based on a Transformer architecture connected in series, and the bald ControlNet is a trainable copy of the bald generation model; the use of the bald generator to process the source picture to generate a bald picture includes: inputting the source picture into the VAE encoder to obtain a latent space code; inputting the latent space code into the bald ControlNet, after block processing, After the linear layer processing, the images are input into a plurality of diffusion models based on the Transformer architecture in series for processing to obtain the source image reference information; the source image reference information is input into the baldness generation model; Gaussian noise in the latent space is randomly generated, and the noise is input into the baldness generation model, and after the block processing and the linear layer processing, a feature map is obtained, and the feature map and the source image reference information output by the baldness ControlNet are input into a plurality of diffusion models based on the Transformer architecture in series for processing, and the obtained output is processed by the multi-layer perceptron and then deblocked; the result after the deblocking processing is input into the VAE decoder to obtain the baldness image corresponding to the source image; The hairstyle change picture generation module is configured to generate a hairstyle change picture using a hairstyle generation model based on a hairstyle reference picture and a bald picture.

Citation Information

Patent Citations

  • Method for generating continuous pictures by long text based on diffusion model

    CN117521672A

  • Virtual fitting method of deep learning 2D picture

    CN118505835A