Virtual fitting method, model training method, electronic device, and storage medium

By combining human posture and clothing analysis models, and using a diffusion model to generate high-quality virtual try-on effects, the challenges of clothing deformation and dynamic processing in existing technologies are solved, and the natural integration of clothing and human is achieved.

CN119130581BActive Publication Date: 2026-02-24HANGZHOU PIXEL INTERACTIVE TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411155373.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-08-21
Publication Date
2026-02-24
Estimated Expiration
2044-08-21

AI Technical Summary

Technical Problem

Existing virtual try-on technology fails to effectively consider the actual body shape of the try-on user, resulting in clothing deformation that affects the realism of the image, especially when dealing with complex textures and human body dynamics.

Method used

By acquiring clothing images and human posture information, and using human posture analysis models, clothing analysis models, and text analysis models, combined with a diffusion model, high-quality virtual try-on effects are generated to ensure a natural blend between clothing and human.

Benefits of technology

The generated virtual try-on images can realistically show how clothing looks on a person, maintaining the authenticity and naturalness of the overall image, and solving the challenges of clothing deformation and dynamic processing in existing technologies.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119130581B_ABST
    Figure CN119130581B_ABST
Patent Text Reader

Abstract

Embodiments of the present application provide a virtual fitting method, a model training method, an electronic device and a storage medium, obtain an image of clothes to be fitted, description information of the clothes to be fitted and a first image; a first character wears first clothes in the first image; separate the first clothes in the first image to obtain a binary mask image of the first clothes and a second image in which the first clothes are blocked; input the first image into a character posture analysis model to obtain posture feature information of the first character; input the image of the clothes to be fitted into a clothes analysis model to obtain clothes feature information of the clothes to be fitted; input the description information into a text analysis model to obtain semantic feature information; input the posture feature information as a conditional input side of a pre-trained diffusion model, input the second image, the binary mask image, the clothes feature information and the semantic feature information as a main input side, and obtain a target image in which the first character wears the clothes to be fitted, thereby improving the virtual fitting effect.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] Embodiments of the present application relate to the technical field of visual processing, and in particular to a virtual try-on method, a model training method, an electronic device and a storage medium. BACKGROUND

[0002] In the clothing industry of e-commerce, the visual presentation of goods is crucial to sales. Although traditional flat clothing images are sufficient, images shown by models usually attract more attention from customers. However, it is time-consuming and laborious to make a large number of model display images. To solve this problem, virtual try-on technology emerges as the times require, which not only improves the shopping experience of consumers, but also effectively reduces the operating costs of clothing retailers.

[0003] In related technologies, the try-on method is mainly to develop standard styles, and then paste the customer's avatar to give consumers the try-on and try-wear effect of clothing. However, the actual body state of the tryer is not considered, and due to the soft nature of clothing, it will deform to a certain extent when it fits the human body, which makes the virtual try-on technology face challenges in processing clothing or human body dynamics with complex textures, thereby affecting the realism of the final image.

[0004] Therefore, how to generate a high-quality virtual try-on effect becomes a problem to be solved. SUMMARY

[0005] Therefore, the embodiments of the present application provide a virtual try-on method, a model training method, an electronic device and a storage medium to at least partially solve the above problems.

[0006] According to a first aspect of the embodiments of the present application, a virtual try-on method is provided, comprising:

[0007] obtaining a to-be-tried clothing image, description information of the to-be-tried clothing and a first image; wherein a first character in the first image wears a first clothing; separating the first clothing in the first image to obtain a binary mask image of the first clothing and a second image in which the first clothing is occluded; inputting the first image into a character pose analysis model to obtain pose feature information of the first character; inputting the to-be-tried clothing image into a clothing analysis model to obtain clothing feature information of the to-be-tried clothing; inputting the description information into a text analysis model to obtain semantic feature information; inputting the pose feature information as a conditional input side of a pre-trained diffusion model, and inputting the second image, the binary mask image, the clothing feature information and the semantic feature information as a main input side of the pre-trained diffusion model, to obtain a target image in which the first character wears the to-be-tried clothing.

[0008] According to a second aspect of the embodiments of the present application, a model training method is provided, comprising:

[0009] An article sample image and a first image are obtained, wherein the first image shows a first person wearing an article sample; the article sample in the first image is occluded to obtain a binary mask image of the article sample and a second image in which the article sample is occluded; the first image is input into a person pose analysis model to obtain pose feature information of the first person; the article sample image is input into an article analysis model to obtain article feature information; description information of the article sample obtained is input into a text analysis model to obtain semantic feature information; the pose feature information is taken as a condition input side of a diffusion model to be trained, and the second image, the binary mask image, the article feature information, and the semantic feature information are taken as main input sides of the diffusion model to be trained to determine a noise loss and a pixel-level loss; and the diffusion model to be trained is trained through the noise loss and the pixel-level loss to obtain a trained diffusion model.

[0010] According to a third aspect of the embodiments of the present application, an electronic device is provided, comprising a processor, a memory, a communication interface, and a communication bus, the processor, the memory, and the communication interface complete communication with each other through the communication bus; the memory is used to store at least one executable instruction, and the executable instruction causes the processor to perform operations corresponding to the method according to the first aspect or the second aspect.

[0011] According to a fourth aspect of the embodiments of the present application, a computer storage medium is provided, and the computer storage medium stores a computer program, and the program is executed by a processor to implement the method according to the first aspect or the second aspect.

[0012] According to the scheme provided in the embodiment of the present application, the image of clothes to be tried on, the description information of the clothes to be tried on and the first image are obtained; wherein the first character wears the first clothes in the first image; the first clothes in the first image are separated to obtain the binary mask image of the first clothes and the second image after the first clothes are shielded; the first image is input into the character posture analysis model to obtain the posture feature information of the first character; the image of clothes to be tried on is input into the clothes analysis model to obtain the clothes feature information of the clothes to be tried on; the description information is input into the text analysis model to obtain the semantic feature information; the posture feature information is taken as the conditional input side of the pre-trained diffusion model, and the second image, the binary mask image, the clothes feature information and the semantic feature information are taken as the main input side of the pre-trained diffusion model to obtain the target image of the first character wearing the clothes to be tried on. In this process, the image of the clothes to be tried on can be redrawn in the specific area where the clothes are separated in the second image, combined with the description information of the clothes and the binary mask image, so that the area of the newly drawn clothes to be tried on image can be harmoniously fused with other parts of the original second image to realize the unity and naturalness in vision. And the posture feature information is taken as the conditional constraint input to the conditional input side of the diffusion model, so that the effect of the clothes to be tried on can be visualized on the character in the finally generated target image, while the authenticity and naturalness of the whole image are maintained. BRIEF DESCRIPTION OF DRAWINGS

[0013] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the following will briefly introduce the drawings needed to be used in the embodiments or prior art description. Obviously, the drawings in the following description are only some embodiments of the present application, and other drawings can also be obtained by those skilled in the art according to these drawings.

[0014] Figure 1 The total architecture diagram of a virtual try-on method provided by the embodiment of the present application is shown in FIG. 1.

[0015] Figure 2 The flowchart of a virtual try-on method provided by the embodiment of the present application is shown in FIG. 2.

[0016] Figure 3 The flowchart of a model training method provided by the embodiment of the present application is shown in FIG. 3.

[0017] Figure 4 The total flowchart of a virtual try-on method provided by the embodiment of the present application is shown in FIG. 4.

[0018] Figure 5 The structure diagram of an electronic device provided by the embodiment of the present application is shown in FIG. 5. DETAILED DESCRIPTION

[0019] In order for those skilled in the art to better understand the technical solutions in the embodiments of the present application, the technical solutions in the embodiments of the present application will be clearly and completely described in the following with reference to the drawings in the embodiments of the present application. Obviously, the described embodiments are only some of the embodiments of the present application, not all. All other embodiments obtained by those of ordinary skill in the art based on the embodiments in the present application shall fall within the scope of protection of the present application.

[0020] To adapt to the virtual try-on method provided by the embodiments of the present application, the embodiments of the present application provide a general architecture diagram of a virtual try-on method, that is, the virtual try-on method of the embodiments of the present application can be applied to the virtual try-on system provided by the embodiments of the present application. As shown in Figure 1 The virtual try-on system includes an image acquisition module, an image processing module and a target image generation module. The image acquisition module is used to acquire a clothing image and a first image, and a first person (for example, a model) wears a first clothing in the first image. In the training stage, the acquired clothing image is the same as the first clothing image worn by the first person. In the inference stage, the acquired clothing image and the first clothing image worn by the first person can be different. The image processing module is used to process the clothing image and the first image. Specifically, by using an open source semantic segmentation technology, such as a Segment Anything Model (SAM), the first clothing worn by the first person can be separated in the first image, and then a binary mask image of the clothing and a second image after the clothing is hidden, that is, a background image of the redrawn image, are obtained. A model for extracting human poses is used to extract human poses from the first image. Save the description information of the clothing image. For clothes without description information, a large model can be used to generate description information corresponding to the clothing image. The target image generation module specifically includes: a clothing image prompt adapter is used to extract features from the clothing image to obtain clothing feature information. A text prompt module is used to extract features from the description information to obtain semantic feature information. A pose prompt model (which can be a ControlNet model) is used to extract features from the pose image to obtain pose feature information. The obtained pose feature information, semantic feature information, clothing feature information, second image after the clothing is hidden, and binary mask image of the clothing are input into a diffusion model to obtain a target image.

[0021] The virtual try-on method provided by the embodiments of the present application will be introduced below. Figure 1 And Figure 2 The virtual try-on method provided by the embodiments of the present application will be introduced below. Figure 2 As shown in Figure 2 is a flowchart of the virtual try-on method provided by the embodiments of the present application Figure 1 The virtual try-on method provided by the embodiments of the present application can be executed by an electronic device, which can be a computer, a server, etc.

[0022] As Figure 2 shown, the virtual try-on method comprises:

[0023] S101, obtaining an image of clothes to be tried on, description information of the clothes to be tried on, and a first image; wherein a first person is wearing a first clothes in the first image.

[0024] In embodiments of the present application, the image of the clothes to be tried on can be clothes that need to be shown or sold to users, and the description information is information describing the appearance, details or wearing scene of the clothes; the first image is an image of the first person wearing the first clothes, wherein the first person can be a real model for clothes display, or a simulated model for clothes display, or a digital model for clothes display, or a customer for clothes purchase, etc.; the first clothes can be different from the clothes to be tried on.

[0025] Illustratively, the image of the clothes to be tried on is obtained, and the description information of the clothes to be tried on can be the style (design style and details, such as round collar, V-neck, straight tube, slim, etc.), color, size (size range from small to large, such as S, M, L, XL, etc.), material (fabric composition, such as cotton, silk, polyester fiber, etc.) and category of the clothes.

[0026] S102, separating the first clothes in the first image to obtain a binary mask image of the first clothes and a second image in which the first clothes are occluded.

[0027] In embodiments of the present application, the binary mask image is used to accurately separate the outline of the clothes from the background, so that the person can "try on" different clothes in a digital environment without actually wearing them. The second image is a background image of the redrawn clothes image, i.e. the drawing of the clothes to be tried on can be performed in the second image.

[0028] The first clothes can be accurately separated from the image of the first person wearing the clothes by using an open-source semantic segmentation technology, such as SAM, and then the binary mask image of the clothes and the second image in which the first clothes are occluded are obtained.

[0029] S103, inputting the first image into a person pose analysis model to obtain pose feature information of the first person.

[0030] After the first person's original first clothes are stripped off, the process of redrawing the clothes to be tried on on the image of the first person needs to ensure that the face, hairstyle and limb movements of the first person remain unchanged. In this way, the final generated target image can visualize the effect of the clothes to be tried on on the first person, while maintaining the authenticity and naturalness of the overall image, so the acquisition of the person's pose information is involved.

[0031] The posture feature information relates to the position, direction and movement of each part of the body, and how these factors combine to express the individual's emotions, intentions and personality. It includes, but is not limited to, head posture feature information, torso feature information, facial expression feature information, and micro-expression and micro-movement feature information.

[0032] In the embodiments of the present application, the character posture analysis model can include a pose prompt ControlNet model, input the first image into the encoder ε of a variational autoencoder (Variational Autoencoder), to obtain the latent space encoding c of the pose p = ε (I pose ), input the latent space encoding into the pose prompt ControlNet, i.e. the encoding block and the middle block copy of the diffusion model, pass through a zero convolution layer, add the encoding block and the middle block of the pose prompt ControlNet to the decoding block and the middle block of the diffusion model through a residual connection, and input the pose feature information into the diffusion model. Wherein, the ControlNet model branch creates a trainable copy of the diffusion model U-Net encoding block and middle block, and adds an additional zero convolution layer. The encoding block (Downblock) and the middle block of ControlNet are added to the decoding block (Upblock) and the middle block of the diffusion model.

[0033] S104, input the image of the clothes to be tried into the clothes analysis model, to obtain the clothes feature information of the clothes to be tried.

[0034] In the embodiments of the present application, the feature information of the clothes is multifaceted, not only related to practicality and comfort, but also reflecting culture and fashion trends. It includes, but is not limited to, color feature information, material feature information, pattern and print feature information, and accessory feature information.

[0035] The clothes analysis model can include a contrastive language-image pre-training encoder (Contrastive Language-Image Pre-training, CLIP), a linear layer, and a layer normalization layer. The clothes analysis model processes the image of the clothes to be tried, and obtains the feature information of the clothes to be tried.

[0036] S105, input the description information into the text analysis model, to obtain the semantic feature information.

[0037] In the embodiments of the present application, the description information is information describing the appearance, details or wearing scene of the clothes. When generating the target image, the description information of the clothes is added, and the semantic information can provide rich guiding information for the generation of the target image, so that a more realistic and delicate clothes display effect is generated.

[0038] The text analysis model can be a CLIP text encoder. The description information is input into the text analysis model to obtain semantic feature information, which provides a text reference for the generation of the target image.

[0039] In S106, the pose feature information is input as a condition input side of the pre-trained diffusion model, and the second image, the binary mask image, the clothing feature information, and the semantic feature information are input as main input sides of the pre-trained diffusion model to obtain the target image in which the first person wears the to-be-try-on clothing.

[0040] The pre-trained diffusion model allows the model to generate data according to additional input information (conditions), which can be a category label, a text description, another image, or any form of guiding signal, with the purpose of guiding the generation process to make the generated result meet specific conditions or attributes.

[0041] In the embodiments of the present application, the second image, the binary mask image, the clothing feature information, and the semantic feature information are input as main input sides of the pre-trained diffusion model, and the pose feature information is input as a condition input side of the pre-trained diffusion model to guide the pre-trained diffusion model as an additional input condition to obtain the target image in which the first person wears the to-be-try-on clothing.

[0042] It can be understood that in the embodiments of the present application, the to-be-try-on clothing image, the description information of the to-be-try-on clothing, and the first image are obtained; wherein the first person wears the first clothing in the first image; the first clothing in the first image is separated to obtain a binary mask image of the first clothing and a second image in which the first clothing is occluded; the first image is input into a person pose analysis model to obtain pose feature information of the first person; the to-be-try-on clothing image is input into a clothing analysis model to obtain clothing feature information of the to-be-try-on clothing; the description information is input into a text analysis model to obtain semantic feature information; and the pose feature information is input as a condition input side of the pre-trained diffusion model, and the second image, the binary mask image, the clothing feature information, and the semantic feature information are input as main input sides of the pre-trained diffusion model to obtain the target image in which the first person wears the to-be-try-on clothing. In this process, the to-be-try-on clothing image can be redrawn in the specific region where the clothing is separated in the second image, combined with the description information and the binary mask image of the clothing, so that the redrawn to-be-try-on clothing image region can be harmoniously fused with other parts of the original second image to achieve visual unity and naturalness. And the pose feature information is input as a condition constraint into the condition input side of the diffusion model, so that the final generated target image can visualize the effect of the to-be-try-on clothing worn on the person, while maintaining the authenticity and naturalness of the overall image.

[0043] In some embodiments of this application, S106 uses pose feature information as the conditional input of the pre-trained diffusion model, and uses the second image, binary mask image, clothing feature information and semantic feature information as the main input of the pre-trained diffusion model. The target image of the first person wearing the clothing to be tried on can be obtained through S1061 to S1066, which will be explained through the following steps.

[0044] S1061. Input the second image into the pre-trained encoder to obtain the first latent space feature vector.

[0045] The pre-trained encoder can be a variational autoencoder (VAE). An encoder takes input data (such as images, audio, or text) and transforms it into a compact, low-dimensional feature representation. This process typically involves multiple layers of neural networks, each layer progressively abstracting the input data and extracting higher-level features. Unlike traditional autoencoders, VAE encoders do more than just reconstruct the input data. They also attempt to learn the distribution of a latent space, where points correspond to samples in the original dataset. This means the encoder not only compresses the data but also estimates the probability distribution of each data point in the latent space.

[0046] In some embodiments of this application, the second image is input into a pre-trained encoder, which extracts key features of the second image through a series of neural network layers, such as convolutional layers, fully connected layers or other types of layers. The encoder of the VAE outputs the parameters of the latent variables, that is, the first latent space feature vector is obtained.

[0047] S1062. The first latent space feature vector, latent space noise and binary mask image are concatenated according to a preset channel to obtain the concatenated latent space feature vector.

[0048] Channel concatenation allows feature vectors from different levels to be combined, thus utilizing both high-level abstract features and low-level detailed features simultaneously. By increasing the number of channels, the network can learn more feature representations, which helps improve the model's expressive power and discriminative ability. Concatenating latent space noise can enhance the model's generalization ability, robustness, and data augmentation.

[0049] In some embodiments of this application, the second image is input into a pre-trained encoder to obtain a 4-channel first latent space feature vector, and random Gaussian noise in the latent space is randomly generated. The first latent space feature vector of the 4 channels, the generated latent space noise, and the binary mask image of the 1 channel are concatenated along the channels to finally obtain the latent space feature vector of the 9 channels, that is, the concatenated latent space feature vector.

[0050] S1063. Based on the concatenated latent space feature vector, feature extraction is performed to obtain the second latent space feature vector.

[0051] S1064. Based on the second latent space feature vector, semantic feature information, and text cross-attention layer, determine the third latent space feature vector.

[0052] In some embodiments of this application, key features are extracted from the concatenated latent space feature vector to obtain a second latent space feature vector. The second latent space feature vector represents image features, and the text cross-attention layer can focus on key information in semantic information. The second latent space feature vector and semantic feature information are input into the text cross-attention layer to establish a connection between the two different features, enhance the understanding between the two features, and obtain a third latent space feature vector that integrates image feature information and semantic feature information.

[0053] S1065. Based on the second latent space feature vector, clothing feature information, and image cross-attention layer, determine the fourth latent space feature vector.

[0054] In some embodiments of this application, key features are extracted from the concatenated latent space feature vector to obtain a second latent space feature vector. The second latent space feature vector represents image features. The image cross-attention layer can focus on key information in the image information. The second latent space feature vector and clothing feature information are input into the image cross-attention layer to establish a connection between the two different features, enhance the understanding between the two features, and obtain a fourth latent space feature vector that fuses image feature information and clothing feature information.

[0055] S1066. Determine the target person image based on the third latent space feature vector, the fourth latent space feature vector, and pose feature information.

[0056] In some embodiments of this application, the pre-trained diffusion model peels off the first garment worn by the first person in the first image and then redraws the garment to be tried on. This process requires ensuring that other parts of the model, such as the face, hairstyle, and background, remain unchanged. In this way, the final generated image can visualize the effect of the garment being tried on the person while maintaining the realism and naturalness of the overall image. Therefore, a more realistic and natural target image is determined by combining the third and fourth hidden space vectors with pose feature information.

[0057] Understandably, in some embodiments of this application, the second person image is input into a pre-trained encoder to obtain a first latent space feature vector, enabling the extraction of higher-level features. The first latent space feature vector, latent space noise, and binary mask image are concatenated according to preset channels to obtain a concatenated latent space feature vector, which improves the model's expressive power, discriminative power, and generalization ability. Feature extraction is performed based on the concatenated latent space feature vector to obtain a second latent space feature vector. Based on the second latent space feature vector, semantic feature information, and a text cross-attention layer, a third latent space feature vector is determined. Based on the second latent space feature vector, clothing feature information, and an image cross-attention layer, a fourth latent space feature vector is determined. Through text cross-attention mechanisms and image cross-attention mechanisms, the model can establish a connection between two different features, enhancing the understanding between the two features. The target person image is determined based on the third latent space feature vector, the fourth latent space feature vector, and pose feature information. Introducing pose feature information can determine a more realistic and natural target person image.

[0058] In some embodiments of this application, S1064 determines the third latent space feature vector based on the second latent space feature vector, semantic feature information, and text cross-attention layer. This can be achieved through S1064A and S1064B, which will be explained in detail through the following steps.

[0059] S1064A: Input semantic feature information and the second latent space feature vector into the text cross-attention layer to determine the first query matrix.

[0060] S1064B: Based on the first query matrix, a weighted summation is performed to obtain the third latent space feature vector.

[0061] For example, semantic feature information and the second latent space feature vector are input into the text cross-attention layer, and the first query matrix and the third latent space feature vector are calculated using the following formula (1):

[0062]

[0063] Where Q = ZW q K = c t W k V = c t W v Let W be the query, key, and value matrix of the text features, respectively, and W be the key-value matrix. q W k W v Z is a trainable linear mapping layer, Z is the second latent space feature vector, Q is the first query matrix, and Z′ is the third latent space feature vector.

[0064] It is understood that in some embodiments of this application, semantic feature information and second latent space feature vector are input into the text cross-attention layer to determine the first query matrix, and a weighted sum is performed based on the first query matrix to obtain the third latent space feature vector. This enables the model to establish a connection between the two different features, semantic feature information and second latent space feature vector, thereby enhancing the understanding between the two features.

[0065] In some embodiments of this application, S1065 determines the fourth latent space feature vector based on the second latent space feature vector, clothing feature information and image cross-attention layer. This can be achieved through S1065A and S1065B, which will be explained in detail through the following steps.

[0066] S1065A: Input the clothing feature information and the second latent space feature vector into the image cross-attention layer to determine the second query matrix.

[0067] S1065B: Based on the second query matrix, a weighted summation is performed to obtain the fourth latent space feature vector.

[0068] For example, the clothing feature information and the second latent space feature vector are input into the image cross-attention layer, and the first query matrix and the third latent space feature vector are calculated using the following formula (2):

[0069]

[0070] Where Q = ZW q K′=c i W′ k V′=c i W′ v These represent the query, key, and value matrix for image features, respectively. W′ k and W′ v For a trainable linear mapping layer, Z is the second latent space feature vector, Q is the second query matrix, and Z″ is the fourth latent space feature vector.

[0071] It is understood that in some embodiments of this application, clothing feature information and the second latent space feature vector are input into the image cross-attention layer to determine the second query matrix. Based on the second query matrix, a weighted sum is performed to obtain the fourth latent space feature vector, so that the model can establish a connection between the clothing feature information and the second latent space feature vector, which are two different features, and enhance the understanding between the two features.

[0072] In some embodiments of this application, determining the target person image based on the third latent space feature vector, the fourth latent space feature vector, and the pose feature information in S1066 can be achieved through S1066A to S1066C, which will be explained through the following steps.

[0073] S1066A. Perform a summation operation on the third latent space feature vector and the fourth latent space feature vector to determine the fifth latent space feature vector.

[0074] In some embodiments of this application, the third latent space feature vector and the fourth latent space feature vector are summed to create a comprehensive representation, resulting in the fifth latent space vector.

[0075] S1066B: Based on a pre-trained diffusion model, the fifth latent space feature vector and pose feature information are processed to obtain the sixth latent space vector.

[0076] S1066C: Decode the sixth latent space feature vector to obtain the image of the target person.

[0077] In some embodiments of this application, after obtaining the fifth latent space vector, the fifth latent space vector and pose feature information are processed by traditional modules to obtain the synthesized latent space vector, namely the sixth latent space vector. Finally, the sixth latent space vector is decoded to obtain the target person image.

[0078] It is understood that, in some embodiments of the present invention, the third and fourth latent space feature vectors are summed to determine the fifth latent space feature vector. Based on a pre-trained diffusion model, the fifth latent space feature vector and pose feature information are processed to obtain the sixth latent space vector. The sixth latent space feature vector is then decoded to obtain the target image. In this process, the fifth latent space feature vector has already fused image and text feature information and incorporated pose feature information. The pre-trained diffusion model processes the pose feature information and the fifth latent space feature vector, and finally, the sixth latent space feature vector is decoded to obtain the target image. This not only allows for the redrawing of the clothing image in the specific area where clothing is separated in the second image, combined with the clothing description information and a binary mask image, but also results in a more realistic and natural target image.

[0079] The following is combined with Figure 1 and Figure 3 This application introduces the model training method provided in its embodiments. For example... Figure 3 As shown, Figure 3 This is a schematic diagram of a model training process provided in an embodiment of this application.

[0080] like Figure 3 As shown, the model training methods include:

[0081] S201. Obtain a clothing sample image and a first image; wherein, in the first image, a first person is wearing a clothing sample.

[0082] In the embodiments of this application, during the model training phase, clothing sample images are collected and used to train the model. The clothing samples can be clothing intended for display to users or for sale, or clothing used only for model training and not displayed to users. The first person can be a live model displaying the clothing, a simulated model displaying the clothing, a digital model displaying the clothing, or an image retained by a customer who has purchased clothing, etc. The clothing sample worn by the first person is the same as the acquired clothing sample; that is, during the model training phase, a pair of clothing sample images and an image of the first person wearing that clothing sample are required as input. The clothing sample images during the model training phase are equivalent to the clothing images to be tried on during the model inference phase.

[0083] S202. Occlude the clothing sample in the first image to obtain a binary mask image of the clothing sample and a second image after the clothing sample is occluded.

[0084] In embodiments of this application, a binary mask image is used to accurately separate the outline of clothing from the background, allowing a person to "try on" different garments in a digital environment without actually wearing them. The second image is a background image for redrawing the clothing image; that is, the clothing to be tried on can be drawn in the second image.

[0085] By using open-source semantic segmentation techniques, such as SAM, the clothing sample worn by the first person can be extracted from the first image to obtain a binary mask image of the clothing sample and a second image after the clothing sample is occluded.

[0086] S203. Input the first image into the human posture analysis model to obtain the posture feature information of the first human.

[0087] S204. Input the image of the clothing sample into the clothing analysis model to obtain clothing feature information.

[0088] S205. Input the descriptive information of the obtained clothing samples into the text analysis model to obtain semantic feature information.

[0089] In the embodiments of this application, the implementation methods for obtaining posture feature information in S203 and S103 are similar, the implementation methods for obtaining clothing feature information in S204 and S104 are similar, and the implementation methods for obtaining semantic feature information in S205 and S105 are similar, which will not be described in detail here.

[0090] During the training phase, for clothing samples without descriptive information, a large model can be used to process the clothing samples and obtain descriptive information.

[0091] S206. Use pose feature information as the conditional input of the diffusion model to be trained, and use the second image, binary mask image, clothing feature information and semantic feature information as the main input of the diffusion model to be trained to determine noise loss and pixel-level loss; and train the diffusion model to be trained through noise loss and pixel-level loss to obtain the trained diffusion model.

[0092] A loss function is a function used to quantify the difference between a model's prediction and the actual target value. The goal of training a model is to minimize this loss function by adjusting the model's parameters, thereby making the model's predictions as close as possible to the true output.

[0093] In the embodiments of this application, pose feature information is used as the conditional input of the diffusion model to be trained, and the second image, binary mask image, clothing feature information, and semantic feature information are used as the main input of the diffusion model to be trained, resulting in predicted noise and predicted image. Noise loss and pixel-level loss are determined based on the predicted noise and predicted image. Finally, the diffusion model to be trained is trained using the noise loss and pixel-level loss to obtain the trained diffusion model.

[0094] It is understood that, in the embodiments of this application, based on the obtained posture feature information, clothing feature information and semantic feature information of the person obtained from the acquired clothing sample image and the first image respectively, the posture feature information is then used as a conditional constraint input to the conditional input side of the diffusion model. The second image, binary mask image, clothing feature information and semantic feature information are used as the main input side of the diffusion model to be trained to determine the noise loss and pixel-level loss. Using multiple loss values ​​for training can improve the performance and generalization ability of the model.

[0095] In some embodiments of this application, S206 can be implemented by S2061 to S2063, which will be explained in detail through the following steps.

[0096] S2061. Add noise to the binary mask image and the binary image to obtain the image after noise addition.

[0097] In some embodiments of this application, to enhance the model's generalization ability, improve robustness, and perform data augmentation, noise is added to the binary mask image and the second image to obtain a noise-added image. During noise addition, random latent space noise is generated. The second image is processed using a pre-trained encoder to obtain latent space feature vectors. The latent space feature vectors, the binary mask image, and the randomly generated noise can be concatenated to obtain the noise-added image.

[0098] S2062. Based on the image after noise addition, clothing feature information and semantic feature information, the predicted image and predicted noise are obtained as the main input side of the diffusion model to be trained.

[0099] S2063. Determine the noise loss based on the predicted noise and the real noise; and determine the pixel-level loss based on the predicted image and the real image.

[0100] In some embodiments of this application, the predicted image and predicted noise are obtained based on the image after noise addition, clothing feature information, and semantic feature information as the main input to the diffusion model to be trained. Noise loss is determined based on the error between the predicted noise and the actual noise, and pixel-level loss is determined based on the error between the predicted image and the real image. When calculating the noise loss and pixel-level loss, loss functions such as mean squared error, cross-entropy loss, or absolute error loss can be used.

[0101] For example, the noise loss is shown in formula (3):

[0102]

[0103] When the pose cueing model is ControlNet, in the above formula (3), where τ θ ,∈ θ These are the pose cues ControlNet and the diffusion model, respectively. p For latent space encoding of pose, I bg For the second image, c t For semantic features, c i For image features, For noise prediction, t is the time step for denoising, and z is the time step for denoising. t This is the latent space encoding at time t. In the diffusion model, the text cross-attention module is fixed and does not participate in training; the remaining parts are trainable. The pose cueing ControlNet model weights are fixed and do not participate in training. If needed, the rest of the entire model can be fixed, and only the pose cueing ControlNet model part can be trained.

[0104] The pixel-level loss is calculated as shown in formula (4) below:

[0105]

[0106] In the above formula (4), For pixel-level loss, I gt For real images, I gen To predict the image.

[0107] It is understood that, in the embodiments of this application, the predicted image and predicted noise are obtained based on the image after noise addition, clothing feature information and semantic feature information as the main input side of the diffusion model to be trained. The noise loss is determined by minimizing the difference between the predicted noise and the actual added noise, and the pixel-level loss is determined by the predicted image and the actual image. The noise loss and pixel-level loss can accurately reflect the difference between the model prediction and the real target, and can guide the model to develop in the direction of the optimal solution, thus ensuring the effectiveness, efficiency and final practicality of the model training.

[0108] In some embodiments of this application, S2063 can be implemented by S2063A to S2063B, which will be explained in detail through the following steps.

[0109] S2063A: Summing the noise loss and pixel-level loss yields the total loss.

[0110] S2063B: Train the diffusion model to be trained using the total loss to obtain the trained diffusion model.

[0111] For example, the formula for calculating the total loss is as follows (5):

[0112]

[0113] In the above formula (5), Let λ be the total loss and λ be the hyperparameter.

[0114] After calculating the total loss, the parameters of the diffusion model are adjusted using the total loss until the training conditions are met, resulting in the trained diffusion model.

[0115] It is understood that in some embodiments of this application, the noise loss and pixel-level loss are summed to obtain the total loss. The diffusion model to be trained is then trained using the total loss to obtain the trained diffusion model. By summing the losses, different loss functions are defined for each type of data, so that the model can effectively integrate multiple information sources, thereby improving the model's performance and generalization ability.

[0116] In the embodiments of this application, the above steps are combined, such as Figure 4 As shown, Figure 4 This is a general flowchart of a virtual try-on method provided in an embodiment of this application. Figure 4In the process, the image acquisition module is used to acquire the image of the clothing to be tried on and the first image. In the first image, the first person (e.g., a model) is wearing the first clothing (during the training phase, the first clothing worn by the first person is the same as the clothing to be tried on; during the inference phase, the first clothing worn by the first person is different from the clothing to be tried on). The clothing image and the first image are input into the image processing module to obtain the clothing image, the binary mask image of the clothing, the second image of the clothing being occluded, and the pose image of the first person. The image and text prompts employ a decoupled cross-attention mechanism. The clothing image is input into an image encoder for feature extraction. The extracted features are then passed through a linear layer and input into the image cross-attention layer of the diffusion model (which includes text and image cross-attention layers). A second image, after the clothing is occluded, is processed by a VAE encoder to obtain a latent space feature vector. This latent space feature vector, the binary mask of the clothing, and noise are concatenated and input into the diffusion model. The clothing description information, i.e., the text prompt words, is input into a text encoder for feature extraction. The extracted features are then input into the text cross-attention layer of the diffusion model. Finally, the pose image is input into a VAE encoder for feature extraction. The extracted human pose features are then fed into the diffusion model through residual connections. The diffusion model processes the input data to obtain the target image.

[0117] Reference Figure 5 This document illustrates a schematic diagram of an electronic device according to an embodiment of this application. The specific embodiments of this application do not limit the specific implementation of the electronic device.

[0118] like Figure 5 As shown, the electronic device may include: a processor 502, a communications interface 504, a memory 506, and a communications bus 508.

[0119] in:

[0120] The processor 502, communication interface 504, and memory 506 communicate with each other via communication bus 508.

[0121] Communication interface 504 is used to communicate with other electronic devices or servers.

[0122] The processor 502 is used to execute program 510, specifically the relevant steps in the above method embodiments.

[0123] Specifically, program 510 may include program code that includes computer operation instructions.

[0124] The processor 502 may be a central processing unit (CPU), an application-specific integrated circuit (ASIC), or one or more integrated circuits configured to implement the embodiments of this application. The smart device includes one or more processors, which may be processors of the same type, such as one or more CPUs; or processors of different types, such as one or more CPUs and one or more ASICs.

[0125] Memory 506 is used to store program 510. Memory 506 may include high-speed RAM memory, and may also include non-volatile memory, such as at least one disk storage device.

[0126] Specifically, program 510 can be used to cause processor 502 to perform the operations corresponding to the methods described in the above method embodiments.

[0127] The specific implementation of each step in program 510 can be found in the corresponding descriptions of the steps and units in the above method embodiments, and will not be repeated here. Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the specific working process of the devices and modules described above can be referred to the corresponding process descriptions in the foregoing method embodiments, and will not be repeated here.

[0128] It should be noted that, depending on the implementation needs, the various components / steps described in the embodiments of this application can be broken down into more components / steps, or two or more components / steps or parts of the operation of components / steps can be combined into new components / steps to achieve the purpose of the embodiments of this application.

[0129] The methods described in the embodiments of this application can be implemented in hardware, firmware, or as software or computer code that can be stored in a recording medium (such as a CD-ROM, RAM, floppy disk, hard disk, or magneto-optical disk), or as computer code downloaded over a network that is originally stored in a remote recording medium or a non-transitory machine-readable medium and will be stored in a local recording medium. Thus, the methods described herein can be processed by software stored on a recording medium using a general-purpose computer, a dedicated processor, or programmable or dedicated hardware (such as an ASIC or FPGA). It is understood that the computer, processor, microprocessor controller, or programmable hardware includes storage components (e.g., RAM, ROM, flash memory, etc.) capable of storing or receiving software or computer code that, when accessed and executed by the computer, processor, or hardware, implements the methods described herein. Furthermore, when a general-purpose computer accesses code used to implement the methods shown herein, the execution of the code transforms the general-purpose computer into a dedicated computer for executing the methods shown herein.

[0130] Those skilled in the art will recognize that the units and method steps of the various examples described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of the embodiments of this application.

[0131] The above embodiments are only used to illustrate the embodiments of this application, and are not intended to limit the embodiments of this application. Those skilled in the art can make various changes and modifications without departing from the spirit and scope of the embodiments of this application. Therefore, all equivalent technical solutions also fall within the scope of the embodiments of this application, and the patent protection scope of the embodiments of this application should be defined by the claims.

Claims

1. A virtual try-on method, characterized in that, include: Acquire an image of the clothing to be tried on, a description of the clothing to be tried on, and a first image; wherein, in the first image, a first person is wearing the first clothing; The first garment in the first image is separated to obtain a binary mask image of the first garment and a second image after the first garment is occluded; The first image is input into the human posture analysis model to obtain the posture feature information of the first human. The image of the garment to be tried on is input into the garment analysis model to obtain the garment feature information of the garment to be tried on; The descriptive information is input into the text analysis model to obtain semantic feature information; Using the pose feature information as the conditional input of a pre-trained diffusion model, and using the second image, the binary mask image, the clothing feature information, and the semantic feature information as the main input of the pre-trained diffusion model, a target image of the first person wearing the clothing to be tried on is obtained, including: The second image is input into a pre-trained encoder to obtain the first latent space feature vector; The first latent space feature vector, latent space noise, and the binary mask image are concatenated according to a preset channel to obtain the concatenated latent space feature vector. Based on the concatenated latent space feature vector, feature extraction is performed to obtain the second latent space feature vector; Based on the second latent space feature vector, the semantic feature information, and the text cross-attention layer, the third latent space feature vector is determined; Based on the second latent space feature vector, the clothing feature information, and the image cross-attention layer, the fourth latent space feature vector is determined; The target image is determined based on the third latent space feature vector, the fourth latent space feature vector, and the pose feature information.

2. The method according to claim 1, characterized in that, The step of determining the third latent space feature vector based on the second latent space feature vector, the semantic feature information, and the text cross-attention layer includes: The semantic feature information and the second latent space feature vector are input into the text cross-attention layer to determine the first query matrix; The third latent space feature vector is obtained by performing a weighted summation based on the first query matrix.

3. The method according to claim 1, characterized in that, The step of determining the fourth latent space feature vector based on the second latent space feature vector, the clothing feature information, and the image cross-attention layer includes: The clothing feature information and the second latent space feature vector are input into the image cross-attention layer to determine the second query matrix; The fourth latent space feature vector is obtained by performing a weighted summation based on the second query matrix.

4. The method according to claim 1, characterized in that, Determining the target image based on the third latent space feature vector, the fourth latent space feature vector, and the pose feature information includes: The fifth latent space feature vector is determined by performing a summation operation on the third latent space feature vector and the fourth latent space feature vector. The fifth latent space feature vector and the pose feature information are processed based on the pre-trained diffusion model to obtain the sixth latent space vector. The target image is obtained by decoding the sixth hidden space vector.

5. A model training method, characterized in that, include: Acquire a clothing sample image and a first image; wherein, in the first image, a first person is wearing the clothing sample; The clothing sample in the first image is masked to obtain a binary mask image of the clothing sample and a second image after the clothing sample is masked. The first image is input into the human posture analysis model to obtain the posture feature information of the first human. The clothing sample image is input into the clothing analysis model to obtain clothing feature information; The descriptive information of the acquired clothing samples is input into the text analysis model to obtain semantic feature information; The pose feature information is used as the conditional input of the diffusion model to be trained, and the second image, the binary mask image, the clothing feature information, and the semantic feature information are used as the main input of the diffusion model to be trained to determine noise loss and pixel-level loss; the diffusion model to be trained is then trained using the noise loss and the pixel-level loss to obtain the trained diffusion model, including: Noise is added to the binary mask image and the second image to obtain a noise-added image; The predicted image and predicted noise are obtained by using the image with added noise, the clothing feature information, and the semantic feature information as the main input to the diffusion model to be trained. The noise loss is determined based on the predicted noise and the actual noise; and the pixel-level loss is determined based on the predicted image and the actual image. The total loss is obtained by summing the noise loss and the pixel-level loss. The diffusion model to be trained is trained using the total loss to obtain the trained diffusion model.

6. An electronic device, comprising: The processor, memory, communication interface, and communication bus are provided, wherein the processor, memory, and communication interface communicate with each other via the communication bus. The memory is used to store at least one executable instruction that causes the processor to perform an operation corresponding to the method as described in any one of claims 1-4 or 5.

7. A computer storage medium having a computer program stored thereon, which, when executed by a processor, implements the method as claimed in any one of claims 1-5 or any one of claims 5.

Citation Information

Patent Citations

  • Virtual fitting method and system for reversely generating portrait fitting effect according to image

    CN117974950A

  • Garment fitting method and system based on semantic enhancement and diffusion model, and storage medium

    CN118350903A