Image generation method and apparatus, and electronic device and storage medium
By obtaining posture images, prompt texts and reference images as control conditions and using a preset diffusion model for multiple rounds of denoising, realistic try-on images are generated. This overcomes the limitations of single-modal control in existing technologies, realizes multi-modal try-on image generation, and improves generation flexibility and stability.
Patent Information
- Application Number
- PCT/CN2025/072858
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-03-29
- Filing Date
- 2025-01-16
- Publication Date
- 2025-10-02
AI Technical Summary
Existing virtual fitting technologies only support clothing images as control conditions for generating fitting images, and cannot achieve multimodal image generation.
By obtaining posture images, prompt texts and reference images as control conditions and using a preset diffusion model to perform multiple rounds of denoising, realistic try-on images are generated, thus realizing multi-modal controlled try-on image generation.
Improved flexibility and stability of fitting image generation, ensuring garments are correctly mapped in the image and retaining fine texture details.
Smart Images

Figure CN2025072858_02102025_PF_FP_ABST
Abstract
Description
Image generation method, device, electronic device and storage medium
[0001] This application claims priority to Chinese Patent Application No. 202410384982.2 filed on March 29, 2024. The contents of the above-mentioned Chinese patent application disclosure are hereby incorporated by reference in their entirety as a part of this application. Technical Field
[0002] Embodiments of the present disclosure relate to an image generation method, apparatus, electronic device, and storage medium. Background Art
[0003] Virtual fitting technology aims to generate a virtual image of a person wearing the clothing in the image based on input clothing images and person images. For example, virtual fitting technology only supports single-modal data (garment images) as a control condition for generating fitting images, and cannot achieve image generation based on multi-modal control. Summary of the Invention
[0004] The embodiments of the present disclosure provide an image generation method, device, electronic device, and storage medium, which can realize the generation of try-on images based on multimodal control.
[0005] At least one embodiment of the present disclosure provides an image generation method, including:
[0006] Acquiring control conditions; wherein the control conditions include a posture image, a prompt text, and a reference image; wherein the prompt text is used to prompt a way to try on the clothing in the reference image;
[0007] Generate a try-on image by using a preset diffusion model, according to a preset noise image and the control conditions;
[0008] The subject being tried on in the try-on image is in a posture corresponding to the posture image, and the subject being tried on wears the garment in the try-on manner.
[0009] At least one embodiment of the present disclosure further provides an image generating device, including:
[0010] An acquisition module, configured to acquire control conditions, wherein the control conditions include a posture image, a prompt text, and a reference image; wherein the prompt text is used to indicate a way to try on the garment in the reference image;
[0011] A generation module, configured to generate a try-on image using a preset diffusion model, a preset noise image and the control conditions;
[0012] The subject being tried on in the try-on image is in a posture corresponding to the posture image, and the subject being tried on wears the garment in the try-on manner.
[0013] At least one embodiment of the present disclosure further provides an electronic device, the electronic device comprising:
[0014] one or more processors;
[0015] a storage device for storing one or more programs,
[0016] When the one or more programs are executed by the one or more processors, the one or more processors implement the image generation method as described in any one of the embodiments of the present disclosure.
[0017] At least one embodiment of the present disclosure further provides a storage medium comprising computer-executable instructions, which, when executed by a computer processor, are used to execute the image generation method as described in any one of the embodiments of the present disclosure. BRIEF DESCRIPTION OF THE DRAWINGS
[0018] The above and other features, advantages, and aspects of the various embodiments of the present disclosure will become more apparent with reference to the following detailed description in conjunction with the accompanying drawings. Throughout the drawings, the same or similar reference numerals represent the same or similar elements. It should be understood that the drawings are schematic and that the originals and elements are not necessarily drawn to scale.
[0019] FIG1 is a schematic flow chart of an image generation method provided by at least one embodiment of the present disclosure;
[0020] FIG2 is a schematic diagram of input and output of a preset diffusion model in an image generation method provided by at least one embodiment of the present disclosure;
[0021] FIG3 is a schematic block diagram of data flow of an image generation method provided by at least one embodiment of the present disclosure;
[0022] FIG4 is a schematic diagram of multi-modal, multi-reference attention fusion in an image generation method provided by at least one embodiment of the present disclosure;
[0023] FIG5 is a schematic diagram of a process for constructing a preset diffusion model in an image generation method provided by at least one embodiment of the present disclosure;
[0024] FIG6 is a schematic block diagram of data flow for constructing a sample reference image in an image generation method provided by at least one embodiment of the present disclosure;
[0025] FIG7 is a schematic structural diagram of an image generating device provided by at least one embodiment of the present disclosure; and
[0026] FIG8 is a schematic structural diagram of an electronic device provided by at least one embodiment of the present disclosure. DETAILED DESCRIPTION
[0027] The following describes embodiments of the present disclosure in more detail with reference to the accompanying drawings. Although certain embodiments of the present disclosure are shown in the accompanying drawings, it should be understood that the present disclosure can be implemented in various forms and should not be construed as limited to the embodiments described herein. Rather, these embodiments are provided to provide a more thorough and complete understanding of the present disclosure. It should be understood that the drawings and embodiments of the present disclosure are for illustrative purposes only and are not intended to limit the scope of protection of the present disclosure.
[0028] It should be understood that the various steps described in the method embodiments of the present disclosure may be performed in different orders and / or in parallel. In addition, the method embodiments may include additional steps and / or omit the steps shown. The scope of the present disclosure is not limited in this respect.
[0029] As used herein, the term "including" and its variations are open-ended, i.e., "including but not limited to." The term "based on" means "based, at least in part, on." The term "one embodiment" means "at least one embodiment," the term "another embodiment" means "at least one additional embodiment," and the term "some embodiments" means "at least some embodiments." Other terms are defined in the following description.
[0030] It should be noted that the concepts of "first" and "second" mentioned in this disclosure are only used to distinguish different devices, modules or units, and are not used to limit the order or interdependence of the functions performed by these devices, modules or units.
[0031] It should be noted that the modifications of "one" and "plurality" mentioned in the present disclosure are illustrative rather than restrictive, and those skilled in the art should understand that unless otherwise clearly indicated in the context, they should be understood as "one or more".
[0032] The names of the messages or information exchanged between multiple devices in the embodiments of the present disclosure are only used for illustrative purposes and are not used to limit the scope of these messages or information.
[0033] It is understandable that the data involved in the technical solutions of the embodiments of the present disclosure (including but not limited to the data itself, the acquisition or use of the data) must comply with the requirements of relevant laws, regulations and relevant provisions.
[0034] Figure 1 is a flow chart of an image generation method provided by at least one embodiment of the present disclosure. This embodiment of the present disclosure is applicable to the generation of try-on images based on multimodal control of text and images. The method can be executed by an image generation device, which can be implemented in software and / or hardware and can be configured in an electronic device, such as a computer.
[0035] As shown in FIG1 , the image generation method provided in this embodiment may include:
[0036] S110, obtaining control conditions; wherein the control conditions include a posture image, a prompt text, and a reference image; wherein the prompt text is used to prompt a way to try on the clothing in the reference image.
[0037] For example, a posture image can be considered as an image composed of skeleton key points and connecting lines, used to present posture information.
[0038] For example, the prompt text may include text such as clothing category and try-on method. For example, the prompt text template may be: a person wearing <clothing category 1><try-on method 1>[label 1], <clothing category 2><try-on method 2>[label 2], ..., and <clothing category N><try-on method N>[label N].
[0039] For example, clothing categories include tops, pants, hats, shoes, and bags. The fitting method can be considered the way the clothing is fitted to the person being tried on. For example, for tops, fitting methods might include zipping up or zipping up, tucking into pants, or draping over pants. Furthermore, for clothing categories with a single fitting method (e.g., pants), the fitting method field in the prompt text may be empty.
[0040] For example, the reference image can be understood as an image that includes clothing to be tried on. The reference image may include only a single piece of clothing to be tried on, or may include other clothing items in addition to the clothing to be tried on. For example, the reference image may be labeled with clothing categories to indicate the clothing items to be tried on.
[0041] It can be considered that the prompt text and the reference image have a corresponding relationship, that is, the clothing that needs to be tried on in the reference image is expected to be tried on in the same way as the corresponding clothing category in the prompt text.
[0042] In the embodiment of the present disclosure, the posture image, prompt text, and reference image can be obtained through a user interface or other means, and the obtained posture image, prompt text, and reference image can be used as control conditions.
[0043] In some optional implementations, the process of receiving the reference image may include at least one of the following: receiving reference images corresponding to clothing categories in the order of clothing categories in the prompt text; receiving the reference image and the clothing category corresponding to the reference image; and sequentially generating corresponding reference images according to the clothing categories in the prompt text.
[0044] For example, a prompt text may be received first, and the clothing category order in the prompt text may be parsed sequentially. The user may then be prompted to input a corresponding reference image based on the parsed order. This allows the received reference image to correspond to the prompt text.
[0045] For example, when receiving the reference image, the clothing category corresponding to the reference image may also be received, so that the received reference image and the prompt text have a corresponding relationship.
[0046] Alternatively, the system can first receive a prompt text and parse the description of the clothing category contained in the prompt text, such as "red sweater." Then, based on an existing text-image generation model, it generates a corresponding reference image. This also allows the received reference image to correspond to the prompt text.
[0047] In these optional implementations, a variety of flexible ways of receiving reference images are provided. In addition, other image receiving methods that can make the received reference image and the prompt text have a corresponding relationship can also be applied here.
[0048] S120 , generating a try-on image using a preset diffusion model, according to a preset noise image and control conditions.
[0049] In order to achieve the task of synthesizing a realistic try-on image based on the posture image, prompt text and reference image, in the embodiment of the present disclosure, this task is modeled as a conditional diffusion model (i.e., a preset diffusion model). The model uses the posture image, prompt text and reference image as control conditions, and performs multiple rounds of step-by-step denoising on the preset noise image to estimate the try-on image.
[0050] Exemplarily, a preset diffusion model can be constructed based on sample try-on images and corresponding sample control conditions. Sample try-on images can be understood as images of the subjects actually wearing clothing. Sample control conditions can include sample posture images, sample reference images, and sample prompt texts. For example, the sample control conditions correspond to the sample try-on images and can include at least one of the following: based on an existing posture recognition algorithm, the posture parameters in the sample try-on images can be identified, and corresponding sample posture images can be generated; data enhancement can be performed based on the sample try-on images to generate sample reference images containing various categories of clothing in the sample try-on images; based on an existing image-text generation model, descriptive text of the clothing in each sample reference image can be generated, and sample prompt text can be generated based on the descriptive text.
[0051] By constructing sample try-on images and corresponding sample control conditions, and using the diffusion model to reconstruct the sample try-on images based on the sample control conditions, the preset diffusion model is able to synthesize realistic try-on images based on the pose image, prompt text, and reference image. Accordingly, the preset diffusion model can perform multiple rounds of progressive denoising on the preset noisy image based on the control conditions to estimate the try-on image.
[0052] In the disclosed embodiment, the subject being tried on in the try-on image assumes a posture corresponding to the posture image, and the subject is wearing the clothing according to the try-on method.
[0053] Exemplarily, FIG2 is a schematic diagram of the input and output of a preset diffusion model in an image generation method provided by at least one embodiment of the present disclosure. Referring to FIG2 , the input of the preset diffusion model may include a reference image and a prompt text. For example, the reference image may include three images, which may be a top image, a trousers image, and a bag image. Moreover, each reference image may be labeled with a corresponding clothing category. The prompt text may be, for example, "a person wearing a top with an unzipped zipper, trousers, and carrying a bag in his right hand." Accordingly, in the try-on image output by the preset diffusion model, the person trying on the clothes may be wearing the top in the top image, the trousers in the trousers image, and the bag in the bag image; and the zipper of the top is unzipped, and the bag is carried in the right hand.
[0054] In the disclosed embodiments, by setting prompt text as a control condition for generating try-on images, different clothing styles can be customized, increasing the flexibility of try-on image generation. By combining prompt text with pose images and reference images to generate try-on images, the try-on effect can be inferred based on instructions from both image and text modalities, enabling try-on image generation based on multimodal control.
[0055] In some optional implementations, the control condition may further include a protection area image; accordingly, the non-try-on area of the try-on object presents the protection area image.
[0056] In these optional implementations, the protected area image can be understood as a portion of the image that can be retained within the desired try-on image. For example, the protected area image may include an image of the head area, or an image of an area where the clothing is not desired to be changed. Accordingly, when generating the try-on image, this portion of the image may not be generated. Instead, the protected area image may be mapped onto the try-on image through deformation or other methods, with the non-try-on area of the subject appearing as the protected area image. This further enhances the flexibility of try-on image generation.
[0057] Since the subject in the try-on image needs to assume the posture in the posture image while also retaining the protected area image, inconsistencies between the protected area image and the posture image will affect the generation of the try-on image to a certain extent. Therefore, the process of receiving the posture image and the protected area image may, for example, include: receiving the original clothing image of the subject, and determining the posture image and the protected area image based on the original clothing image. Specifically, based on an existing posture recognition algorithm, the posture parameters in the original clothing image can be identified and the corresponding posture image generated; and the protected area image can be removed from the original clothing image. This ensures consistency between the protected area image and the posture image, thereby improving the generation of the try-on image.
[0058] The technical solution of the embodiment of the present disclosure obtains control conditions; wherein the control conditions include a posture image, a prompt text and a reference image; wherein the prompt text is used to prompt a way of trying on the clothing in the reference image; through a preset diffusion model, a try-on image is generated according to a preset noise image and the control conditions; wherein the try-on object in the try-on image assumes a posture corresponding to the posture image, and the try-on object wears the clothing according to the try-on method.
[0059] In the process of generating a try-on image by denoising a preset noise image based on a diffusion model, by setting a pose image, prompt text, and a reference image as control conditions for generating the try-on image, the subject in the try-on image can assume the pose in the pose image and can try on the garment in the reference image according to the try-on method indicated in the prompt text. This allows the try-on effect to be inferred based on instructions from both image and text modalities, enabling try-on image generation based on multimodal control and improving try-on flexibility.
[0060] The various optional schemes in the image generation method provided in the embodiment of the present disclosure and the above-mentioned embodiment can be combined. The image generation method provided in this embodiment provides a detailed description of the different ways of attention fusion based on prompt text and reference image. By fusing the text features of the prompt text with the image features of the reference image, and performing multimodal attention fusion of the fused multimodal features with the output features of the diffusion model network layer, it not only helps to generate images of different try-on methods, but also enhances the stability of the try-on and ensures the correct mapping of the clothing. By performing attention fusion of the clothing features in the reference image with the output features of the diffusion model network layer, the try-on image can retain the detailed texture details in the clothing, thereby improving the try-on image generation effect.
[0061] For example, Figure 3 is a schematic block diagram of a data flow for an image generation method provided by at least one embodiment of the present disclosure. Referring to Figure 3 , the image generation method provided by this embodiment generates a try-on image using a preset diffusion model, a preset noise image, and control conditions, and may include:
[0062] Feature extraction is performed on the prompt text to obtain text features; feature extraction is performed on the reference image to obtain image features; the text features and image features are fused to obtain multimodal features; the multimodal features are fused with the output features of each network layer in the preset diffusion model through first attention, so that the preset diffusion model outputs a try-on image based on the first attention fusion result.
[0063] As shown in Figure 3, the prompt text may include, for example, "a man wearing a short-sleeved round-neck T-shirt [label 1], gray trousers [label 2], and low-top sneakers [label 3]"; the reference image may correspond to, for example, a T-shirt image G1, a trousers image G2, and a shoe image G3. For example, the prompt text may be subjected to feature extraction using an existing text encoder to obtain text features; and the reference image may be subjected to feature extraction using an existing image encoder to obtain image features. For example, the text features may be fused with the image features of multiple reference images to obtain a multimodal feature. The multimodal feature may be subjected to first attention fusion with the output features of each network layer in a preset diffusion model (such as the denoising U-net model in Figure 3).
[0064] Furthermore, the process of denoising a preset noisy image to generate a try-on image requires multiple rounds of denoising. This means that after each round of denoising, the noisy image is re-input into the preset diffusion model shown in Figure 3. This means that the output features of the same network layer differ during different rounds of denoising. However, the multimodal features fused with the output features through the first attention step are considered unchanged.
[0065] By fusing the text features of the prompt text with the image features of the reference image, and performing multimodal attention fusion on the fused multimodal features with the output features of the diffusion model network layer, it not only helps to generate images of different try-on methods, but also enhances the stability of try-on and ensures the correct mapping of clothing.
[0066] In some optional implementations, generating a try-on image by using a preset diffusion model, according to a preset noise image and control conditions may further include:
[0067] The reference image is extracted through a clothing encoder with the same structure as the preset diffusion model to obtain clothing features. The clothing features are fused with the output features of each network layer in the preset diffusion model through a second attention fusion, so that the preset diffusion model outputs a try-on image based on the second attention fusion result.
[0068] Referring again to Figure 3, the features of the garments to be tried on can be extracted from G1, G2, and G3 using the same clothing encoder structure. Because the reference image corresponds to the clothing categories in the prompt text, in some implementations, the clothing feature extraction process can include: using the clothing encoder to perform cross-attention calculations on the reference image and the clothing categories in the prompt text to obtain clothing features.
[0069] As shown in Figure 3, the clothing encoder can be used to perform cross-attention calculations on G1 and the corresponding clothing category (i.e., T-shirts), G2 and the corresponding clothing category (i.e., pants), and G3 and the corresponding clothing category (i.e., shoes) to obtain clothing features for each clothing item to be tried on. Furthermore, the multimodal features can be fused with the output features of each network layer in the preset diffusion model through a second attention fusion.
[0070] Referring to the process of first attention fusion, the output features of the same network layer in different rounds of denoising can be considered different; while the clothing features that are fused with the output features for the second attention can be considered unchanged.
[0071] The traditional process of generating try-on images typically relies on a priori segmentation models, which limits the generation instructions for the try-on images. However, in the disclosed embodiments, this does not rely on a segmentation model. Instead, the cross-attention layer of the serving encoder extracts clothing features corresponding to the clothing category in the reference image. These clothing features are injected into a preset diffusion model and fused with the output features of each network layer through a second attention process, so that the try-on images retain the detailed texture details of the clothing.
[0072] Furthermore, traditional try-on image generation processes only accept input of a single garment image. However, in the disclosed embodiment, as shown in Figure 3, the reference images include reference images corresponding to at least two clothing categories. This allows the try-on image to present a combined try-on effect of at least two clothing categories, thus enabling flexible, multi-modal, multi-reference control of try-on image generation.
[0073] For example, Figure 4 is a schematic diagram of multimodal, multi-reference attention fusion in an image generation method provided by at least one embodiment of the present disclosure. Referring to Figure 4 , in the image generation method provided by this embodiment, when the reference images include reference images corresponding to at least two clothing categories, the process of determining multimodal features may include:
[0074] According to the clothing category labels in the prompt text, the text features are aligned with the corresponding image features; the aligned text features and image features are fused to obtain a first value vector and a first key vector; accordingly, the multimodal features are fused with the output features of each network layer in the preset diffusion model through a first attention fusion, including: determining a first query vector based on the output features, and performing a first attention fusion based on the first query vector, the first key vector, and the first value vector.
[0075] As mentioned above, the prompt text template example can include clothing category labels, such as [Label 1], [Label 2], and [Label N]. For example, the clothing category labels can be used as placeholders. After extracting text features, the corresponding text features of each clothing category can be located based on the clothing category labels. Accordingly, the clothing category labels can be used to align the text features of each clothing category with the corresponding image features.
[0076] As shown in Figure 4, the aligned text features and each image feature can be fused through the multimodal embedding module to obtain the first value vector V1 and the first key vector K1; the first query vector Q1 can be determined based on the output features; and the first attention fusion can be performed based on the first query vector Q1, the first key vector K1 and the first value vector V1 according to the multimodal attention module.
[0077] Referring again to FIG. 4 , in the image generation method provided in this embodiment, when the reference image includes reference images corresponding to at least two clothing categories, performing a second attention fusion on the clothing features and the output features of each network layer in the preset diffusion model may include:
[0078] According to each clothing feature and the output feature, a second value vector and a second key vector are determined; according to the output feature, a second query vector is determined, and a second attention fusion is performed according to the second query vector, the second key vector and the second value vector.
[0079] As shown in Figure 4, the second value vector V2 and the second key vector K2 can be determined according to the clothing features and the output features through the multi-reference embedding module; the second query vector Q2 can be determined according to the output features; and the second attention fusion can be performed according to the second query vector Q2, the second key vector K2 and the second value vector V2 according to the multi-reference attention module.
[0080] For example, based on the existing attention module mechanism, the first query vector Q1, the first key vector K1 and the first value vector V1, as well as the second query vector Q2, the second key vector K2 and the second value vector V2 can be determined. For example, in FIG4 , the first attention fusion is first performed, and the result of the first attention fusion can be used as the output feature of the multi-reference attention module. In addition, the second attention fusion can also be performed first, and the result of the second attention fusion can be used as the output feature of the multimodal attention module. It can be considered that corresponding adjustments can be made according to the specific diffusion model structure, and no specific limitation is made here. In this way, multi-reference, multi-modal try-on image generation can be achieved.
[0081] The technical solution of the embodiment of the present disclosure describes in detail the different ways of attention fusion based on prompt text and reference image. By fusing the text features of the prompt text with the image features of the reference image, and performing multimodal attention fusion of the fused multimodal features with the output features of the diffusion model network layer, it not only helps to generate images of different try-on methods, but also enhances the stability of the try-on and ensures the correct mapping of the clothing. By performing attention fusion of the clothing features in the reference image with the output features of the diffusion model network layer, the try-on image can retain the detailed texture details of the clothing, thereby improving the try-on image generation effect.
[0082] In addition, the image generation method provided in the embodiment of the present disclosure and the image generation method provided in the above embodiment belong to the same disclosed concept. Technical details not fully described in this embodiment can be referred to the above embodiment, and the same technical features have the same beneficial effects in this embodiment and the above embodiment.
[0083] The various optional schemes in the image generation method provided in the embodiment of the present disclosure and the above-mentioned embodiment can be combined. The image generation method provided in this embodiment provides a detailed description of the construction process of the preset diffusion model and the construction process of the sample reference image. By constructing a reconstruction loss based on the sample try-on image prediction value and the sample try-on image, and constructing a noise prediction loss based on the predicted noise in each round of denoising and the noise added in the corresponding round, the diffusion model can be enabled to have the ability to obtain a try-on image by denoising the preset noise image, thereby realizing the construction of the diffusion model. By performing data enhancement on the basis of the sample try-on image, the sample parameter image can be automatically generated, which helps to improve the efficiency and effect of model construction.
[0084] FIG5 is a schematic diagram of a process for constructing a preset diffusion model in an image generation method provided by at least one embodiment of the present disclosure. As shown in FIG5 , in the image generation method provided by this embodiment, the process of constructing the preset diffusion model may include:
[0085] S510: Add a preset number of rounds of noise to the sample fitting image to obtain a sample noise image.
[0086] In this embodiment, based on existing noise adding methods, a preset number of rounds of noise can be added to the sample fitting image, wherein the preset number of rounds can be pre-set according to the application scenario.
[0087] S520: Generate a sample try-on image prediction value based on the sample noise image and the sample control condition through a preset diffusion model.
[0088] In this embodiment, the sample control conditions may include a sample pose image, a sample reference image, and sample prompt text. As shown above, the sample control conditions and the sample try-on images have a corresponding relationship, indicating that the sample pose image, the sample reference image, and the sample prompt text may correspond to the sample try-on image. Furthermore, the process for generating the predicted value for the sample try-on image can be referenced to the process for generating the try-on image above and will not be further elaborated here.
[0089] For example, FIG6 is a schematic block diagram of the data flow for constructing a sample reference image in an image generation method provided by at least one embodiment of the present disclosure. Referring to FIG6 , in some optional implementations, the process of generating a sample reference image may include:
[0090] Based on the sample try-on images, descriptive texts are generated for the corresponding clothing of each clothing category; based on each descriptive text, detection frames of the corresponding clothing are determined; clothing is segmented based on each detection frame to obtain each clothing; based on the sample try-on images, each clothing is enhanced to obtain an enhanced image; and images of each clothing are selected from the enhanced images to be retained as sample reference images.
[0091] Exemplarily, an existing prompt-based generation model can be used to generate clothing of the main clothing categories in the try-on image. As shown in Figure 6, the prompt words may include, for example, "Please extract items related to clothing, such as short-sleeved round-neck T-shirts." For example, an existing text generation model can be used to generate descriptive text for describing the clothing and wearing methods in the sample try-on image. As shown in Figure 6, the descriptive text includes descriptive text for T-shirts, pants, and shoes; exemplarily, the descriptive text for T-shirts includes, "This man is wearing a short-sleeved round-neck T-shirt." For example, based on the existing detection model, a bounding box corresponding to each clothing category in the image can be generated according to the descriptive text of each clothing category. For example, based on an existing segmentation model, an accurate mask of the clothing corresponding to each clothing category can be generated.
[0092] For example, a cue-based generative model can be used to randomly generate images of people wearing clothing from various clothing categories. Images with high similarity can be removed to obtain an enhanced dataset. From the randomly enhanced dataset, images of tops, pants, and shoes from the sample try-on images can be filtered out and retained as sample reference images. This establishes a correspondence between the person being tried on and the clothing description. As shown in Figure 6, it is understood that a correspondence must be established between the sample try-on images and the sample reference images in two different poses; otherwise, the try-on task degenerates into a copy-and-paste task.
[0093] In these optional implementations, an enhanced dataset is provided for the preset diffusion model, allowing the model to be trained without any explicit clothing segmentation, which can improve the efficiency and effectiveness of model building.
[0094] S530: Determine the reconstruction loss according to the sample try-on image and the predicted value of the sample try-on image.
[0095] In this embodiment, the reconstruction loss between the sample fitting image and the predicted value of the sample fitting image may be determined based on an existing image loss determination method.
[0096] S540: Determine the noise prediction loss according to the predicted noise in each round of denoising using the preset diffusion model and the noise added in the corresponding round.
[0097] In this embodiment, the noise prediction loss may include determining the loss of the predicted noise image and the noise image added in the corresponding round based on the existing image loss determination method; it may also include the loss of the predicted noise image conforming to a preset distribution (such as a Gaussian distribution, etc.).
[0098] S550: Adjust parameters of the preset diffusion model according to the reconstruction loss and the noise prediction loss.
[0099] In this embodiment, back propagation may be performed based on the reconstruction loss and the noise prediction loss to adjust the parameters in the preset diffusion model and complete the construction of the preset diffusion model.
[0100] The technical solution of the embodiment of the present disclosure describes in detail the construction process of the preset diffusion model and the construction process of the sample reference image. By constructing the reconstruction loss based on the sample try-on image prediction value and the sample try-on image, and constructing the noise prediction loss based on the predicted noise in each round of denoising and the noise added in the corresponding round, the diffusion model can be enabled to obtain the try-on image by denoising the preset noise image, thereby realizing the construction of the diffusion model. By performing data enhancement on the basis of the sample try-on image, the sample parameter image can be automatically generated, which helps to improve the model construction efficiency and construction effect. In addition, the image generation method provided by the embodiment of the present disclosure and the image generation method provided by the above embodiment belong to the same public concept. The technical details not fully described in this embodiment can be referred to the above embodiment, and the same technical features have the same beneficial effects in this embodiment and the above embodiment.
[0101] Figure 7 is a schematic diagram of the structure of an image generation device provided by at least one embodiment of the present disclosure. The image generation device provided by this embodiment is applicable to the generation of try-on images based on multimodal control of text and images.
[0102] As shown in FIG7 , the image generation device provided by the embodiment of the present disclosure may include:
[0103] The acquisition module 710 is used to acquire control conditions; wherein the control conditions include a posture image, a prompt text, and a reference image; wherein the prompt text is used to indicate how to try on the clothing in the reference image;
[0104] A generation module 720 is used to generate a try-on image using a preset diffusion model, a preset noise image and control conditions;
[0105] The subject in the try-on image is in a posture corresponding to the posture image, and wears the clothing according to the try-on method.
[0106] In some optional implementations, the generation module can be used to:
[0107] Extract features from the prompt text to obtain text features;
[0108] Perform feature extraction on the reference image to obtain image features;
[0109] Fuse text features and image features to obtain multimodal features;
[0110] The multimodal features are fused with the output features of each network layer in the preset diffusion model by first attention, so that the preset diffusion model outputs a try-on image based on the first attention fusion result.
[0111] In some optional implementations, the generation module can also be used to:
[0112] The clothing encoder with the same structure as the preset diffusion model is used to extract features from the reference image to obtain clothing features.
[0113] The clothing features are fused with the output features of each network layer in the preset diffusion model through a second attention fusion, so that the preset diffusion model outputs a try-on image based on the second attention fusion result.
[0114] In some optional implementations, the generation module may extract clothing features based on the following process:
[0115] Through the clothing encoder, cross-attention calculation is performed on the clothing categories in the reference image and the prompt text to obtain clothing features.
[0116] In some optional implementations, when the reference images include reference images corresponding to at least two clothing categories, the generation module may determine the multimodal features based on the following process:
[0117] According to the clothing category labels in the prompt text, align the text features with the corresponding image features;
[0118] The aligned text features and each image feature are fused to obtain a first value vector and a first key vector;
[0119] Accordingly, the multimodal features are fused with the output features of each network layer in the preset diffusion model using the first attention method, including:
[0120] A first query vector is determined according to the output features, and a first attention fusion is performed according to the first query vector, the first key vector, and the first value vector.
[0121] In some optional implementations, when the reference images include reference images corresponding to at least two clothing categories, the generation module may be configured to:
[0122] Determining a second value vector and a second key vector according to each clothing feature and the output feature;
[0123] A second query vector is determined according to the output features, and a second attention fusion is performed according to the second query vector, the second key vector, and the second value vector.
[0124] In some optional implementations, the image generating apparatus may further include:
[0125] A receiving module is configured to receive a reference image based on at least one of the following processes:
[0126] Receive reference images corresponding to the clothing categories in the order of the clothing categories in the prompt text;
[0127] Receiving a reference image and a clothing category corresponding to the reference image;
[0128] Generate corresponding reference images sequentially according to the clothing categories in the prompt text.
[0129] In some optional implementations, the control condition further includes a protection area image; accordingly, the non-try-on area of the try-on object presents the protection area image.
[0130] In some optional implementations, the receiving module may also be configured to receive the posture image and the protection area image based on the following process:
[0131] An original clothing image of a person being tried on is received, and a posture image and a protection area image are determined according to the original clothing image.
[0132] In some optional implementations, the image generating apparatus may further include:
[0133] A model building module that builds a pre-defined diffusion model based on the following process:
[0134] Add a preset number of rounds of noise to the sample fitting image to obtain a sample noise image;
[0135] Generate the predicted value of the sample fitting image based on the sample noise image and sample control conditions through the preset diffusion model;
[0136] Determine the reconstruction loss based on the sample try-on image and the predicted value of the sample try-on image;
[0137] Determine the noise prediction loss based on the predicted noise in each round of denoising using the preset diffusion model and the noise added in the corresponding round;
[0138] The parameters of the preset diffusion model are adjusted according to the reconstruction loss and noise prediction loss.
[0139] In some optional implementations, the sample control condition includes a sample reference image; and the image generating device may further include:
[0140] The sample construction module is used to construct a sample reference image based on the following process:
[0141] Generate description text for each clothing category based on the sample try-on images;
[0142] According to each description text, determine the detection frame of the corresponding clothing;
[0143] Perform clothing segmentation according to each detection frame to obtain each garment;
[0144] Based on the sample try-on images, each garment is enhanced to obtain an enhanced image;
[0145] Images of each garment are selected from the enhanced images and retained as sample reference images.
[0146] The image generation device provided in the embodiments of the present disclosure can execute the image generation method provided in any embodiment of the present disclosure, and has the corresponding functional modules and beneficial effects of the execution method.
[0147] It is worth noting that the various units and modules included in the above-mentioned device are only divided according to functional logic, but are not limited to the above-mentioned division, as long as the corresponding functions can be achieved; in addition, the specific names of the functional units are only for the convenience of distinguishing each other, and are not used to limit the protection scope of the embodiments of the present disclosure.
[0148] Reference is now made to FIG8 , which illustrates a schematic diagram of the structure of an electronic device (e.g., a terminal device or server in FIG8 ) 800 suitable for implementing at least one embodiment of the present disclosure. The terminal device in the embodiments of the present disclosure may include, but is not limited to, mobile terminals such as mobile phones, laptop computers, digital broadcast receivers, PDAs (personal digital assistants), PADs (tablet computers), PMPs (portable multimedia players), in-vehicle terminals (e.g., in-vehicle navigation terminals), and fixed terminals such as digital TVs and desktop computers. The electronic device illustrated in FIG8 is merely an example and should not limit the functionality and scope of use of the embodiments of the present disclosure.
[0149] As shown in Figure 8, the electronic device 800 may include a processing device (e.g., a central processing unit, a graphics processing unit, etc.) 801, which can perform various appropriate actions and processes according to a program stored in a read-only memory (ROM) 802 or a program loaded from a storage device 808 into a random access memory (RAM) 803. Various programs and data required for the operation of the electronic device 800 are also stored in the RAM 803. The processing device 801, the ROM 802, and the RAM 803 are connected to each other via a bus 804. An input / output (I / O) interface 805 is also connected to the bus 804.
[0150] Typically, the following devices may be connected to the I / O interface 805: an input device 806 including, for example, a touch screen, a touchpad, a keyboard, a mouse, a camera, a microphone, an accelerometer, a gyroscope, etc.; an output device 807 including, for example, a liquid crystal display (LCD), a speaker, a vibrator, etc.; a storage device 808 including, for example, a magnetic tape, a hard disk, etc.; and a communication device 809. The communication device 809 may allow the electronic device 800 to communicate with other devices wirelessly or by wire to exchange data. Although FIG8 shows the electronic device 800 with various devices, it should be understood that it is not required to implement or have all of the devices shown. More or fewer devices may be implemented or have alternatively.
[0151] In particular, according to an embodiment of the present disclosure, the process described above with reference to the flowchart can be implemented as a computer software program. For example, an embodiment of the present disclosure includes a computer program product, which includes a computer program carried on a non-transitory computer-readable medium, and the computer program includes program code for executing the method shown in the flowchart. In such an embodiment, the computer program can be downloaded and installed from the network through the communication device 809, or installed from the storage device 808, or installed from the ROM 802. When the computer program is executed by the processing device 801, the above-mentioned functions defined in the image generation method of the embodiment of the present disclosure are performed.
[0152] The electronic device provided by the embodiment of the present disclosure and the image generation method provided by the above embodiment belong to the same disclosed concept. For technical details not fully described in this embodiment, please refer to the above embodiment, and this embodiment has the same beneficial effects as the above embodiment.
[0153] An embodiment of the present disclosure provides a computer storage medium having a computer program stored thereon. When the program is executed by a processor, the image generation method provided by the above embodiment is implemented.
[0154] It should be noted that the computer-readable medium mentioned above in the present disclosure may be a computer-readable signal medium or a computer-readable storage medium or any combination of the two. The computer-readable storage medium may be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, device or component, or any combination of the above. More specific examples of computer-readable storage media may include, but are not limited to: an electrical connection with one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM) or flash memory (FLASH), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In the present disclosure, a computer-readable storage medium may be any tangible medium containing or storing a program that can be used by or in conjunction with an instruction execution system, device or device. In the present disclosure, a computer-readable signal medium may include a data signal propagated in baseband or as part of a carrier wave, which carries a computer-readable program code. Such propagated data signals may take a variety of forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination thereof. A computer-readable signal medium may also be any computer-readable medium other than a computer-readable storage medium that can send, propagate, or transmit a program for use by or in conjunction with an instruction execution system, apparatus, or device. The program code contained on the computer-readable medium may be transmitted using any suitable medium, including but not limited to wires, optical cables, RF (radio frequency), etc., or any suitable combination thereof.
[0155] In some embodiments, the client and server can communicate using any currently known or later developed network protocol, such as HTTP (Hypertext Transfer Protocol), and can be interconnected with any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include a local area network ("LAN"), a wide area network ("WAN"), an internet (e.g., the Internet), and a peer-to-peer network (e.g., an ad hoc peer-to-peer network), as well as any currently known or later developed network.
[0156] The computer-readable medium may be included in the electronic device, or may exist independently without being incorporated into the electronic device.
[0157] The computer-readable medium carries one or more programs. When the one or more programs are executed by the electronic device, the electronic device:
[0158] Acquire control conditions; wherein the control conditions include a posture image, a prompt text, and a reference image; wherein the prompt text is used to indicate a way of trying on the clothing in the reference image; generate a try-on image by using a preset diffusion model, according to a preset noise image and the control conditions; wherein the try-on object in the try-on image assumes a posture corresponding to the posture image, and the try-on object wears the clothing according to the try-on method.
[0159] Computer program code for performing the operations of the present disclosure may be written in one or more programming languages, or a combination thereof, including, but not limited to, object-oriented programming languages such as Java, Smalltalk, C++, and conventional procedural programming languages such as "C" or similar programming languages. The program code may be executed entirely on the user's computer, partially on the user's computer, as a stand-alone software package, partially on the user's computer and partially on a remote computer, or entirely on the remote computer or server. In cases involving a remote computer, the remote computer may be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or may be connected to an external computer (e.g., through the Internet using an Internet service provider).
[0160] The flowcharts and block diagrams in the accompanying drawings illustrate the possible implementation architecture, functions and operations of the systems, methods and computer program products according to various embodiments of the present disclosure. In this regard, each box in the flowchart or block diagram can represent a module, program segment, or a part of code, and the module, program segment, or a part of code contains one or more executable instructions for realizing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the box can also occur in a different order than that marked in the accompanying drawings. For example, two boxes represented in succession can actually be executed substantially in parallel, and they can sometimes be executed in the opposite order, depending on the functions involved. It should also be noted that each box in the block diagram and / or flowchart, and the combination of the boxes in the block diagram and / or flowchart, can be implemented with a dedicated hardware-based system that performs the specified function or operation, or can be implemented with a combination of dedicated hardware and computer instructions.
[0161] The units involved in the embodiments described in this disclosure may be implemented in software or hardware, wherein the names of the units and modules do not, in certain circumstances, limit the units and modules themselves.
[0162] The functions described above herein may be performed, at least in part, by one or more hardware logic components. For example, and without limitation, exemplary types of hardware logic components that may be used include: Field Programmable Gate Array (FPGA), Application Specific Integrated Circuit (ASIC), Application Specific Standard Parts (ASSP), System on Chip (SOC), Complex Programmable Logic Device (CPLD), and the like.
[0163] In the context of the present disclosure, a machine-readable medium can be a tangible medium that can contain or store a program for use by or in conjunction with an instruction execution system, device or equipment. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can include, but is not limited to, an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, device or equipment, or any suitable combination of the foregoing. A more specific example of a machine-readable storage medium can include an electrical connection based on one or more lines, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.
[0164] According to one or more embodiments of the present disclosure, there is provided an image generation method, the method comprising:
[0165] Acquire control conditions; wherein the control conditions include a posture image, a prompt text, and a reference image; wherein the prompt text is used to prompt a way to try on the clothing in the reference image;
[0166] Generate a try-on image using a preset diffusion model, based on a preset noise image and control conditions;
[0167] The subject in the try-on image is in a posture corresponding to the posture image, and wears the clothing according to the try-on method.
[0168] According to one or more embodiments of the present disclosure, there is provided an image generation method, further comprising:
[0169] In some optional implementations, generating a try-on image by using a preset diffusion model, according to a preset noise image and control conditions, includes:
[0170] Extract features from the prompt text to obtain text features;
[0171] Perform feature extraction on the reference image to obtain image features;
[0172] Fuse text features and image features to obtain multimodal features;
[0173] The multimodal features are fused with the output features of each network layer in the preset diffusion model by first attention, so that the preset diffusion model outputs a try-on image based on the first attention fusion result.
[0174] According to one or more embodiments of the present disclosure, there is provided an image generation method, further comprising:
[0175] In some optional implementations, generating a try-on image using a preset diffusion model, according to a preset noise image and control conditions, further includes:
[0176] The clothing encoder with the same structure as the preset diffusion model is used to extract features from the reference image to obtain clothing features.
[0177] The clothing features are fused with the output features of each network layer in the preset diffusion model through a second attention fusion, so that the preset diffusion model outputs a try-on image based on the second attention fusion result.
[0178] According to one or more embodiments of the present disclosure, there is provided an image generation method, further comprising:
[0179] In some optional implementations, the clothing feature extraction process includes:
[0180] Through the clothing encoder, cross-attention calculation is performed on the clothing categories in the reference image and the prompt text to obtain clothing features.
[0181] According to one or more embodiments of the present disclosure, there is provided an image generation method, further comprising:
[0182] In some optional implementations, when the reference images include reference images corresponding to at least two clothing categories, the process of determining the multimodal features includes:
[0183] According to the clothing category labels in the prompt text, align the text features with the corresponding image features;
[0184] The aligned text features and each image feature are fused to obtain a first value vector and a first key vector;
[0185] Accordingly, the multimodal features are fused with the output features of each network layer in the preset diffusion model using the first attention method, including:
[0186] A first query vector is determined according to the output features, and a first attention fusion is performed according to the first query vector, the first key vector, and the first value vector.
[0187] According to one or more embodiments of the present disclosure, there is provided an image generation method, further comprising:
[0188] In some optional implementations, when the reference image includes reference images corresponding to at least two clothing categories, performing a second attention fusion on the clothing features and the output features of each network layer in the preset diffusion model includes:
[0189] Determining a second value vector and a second key vector according to each clothing feature and the output feature;
[0190] A second query vector is determined according to the output features, and a second attention fusion is performed according to the second query vector, the second key vector, and the second value vector.
[0191] According to one or more embodiments of the present disclosure, there is provided an image generation method, further comprising:
[0192] In some optional implementations, the process of receiving the reference image includes at least one of the following:
[0193] Receive reference images corresponding to the clothing categories in the order of the clothing categories in the prompt text;
[0194] Receiving a reference image and a clothing category corresponding to the reference image;
[0195] Generate corresponding reference images sequentially according to the clothing categories in the prompt text.
[0196] According to one or more embodiments of the present disclosure, there is provided an image generation method, further comprising:
[0197] In some optional implementations, the control condition further includes a protection area image; accordingly, the non-try-on area of the try-on object presents the protection area image.
[0198] According to one or more embodiments of the present disclosure, there is provided an image generation method, further comprising:
[0199] In some optional implementations, the process of receiving the posture image and the protection area image includes:
[0200] An original clothing image of a person being tried on is received, and a posture image and a protection area image are determined according to the original clothing image.
[0201] According to one or more embodiments of the present disclosure, there is provided an image generation method, further comprising:
[0202] In some optional implementations, the process of constructing the preset diffusion model includes:
[0203] Add a preset number of rounds of noise to the sample fitting image to obtain a sample noise image;
[0204] Generate the predicted value of the sample fitting image based on the sample noise image and sample control conditions through the preset diffusion model;
[0205] Determine the reconstruction loss based on the sample try-on image and the predicted value of the sample try-on image;
[0206] Determine the noise prediction loss based on the predicted noise in each round of denoising using the preset diffusion model and the noise added in the corresponding round;
[0207] The parameters of the preset diffusion model are adjusted according to the reconstruction loss and noise prediction loss.
[0208] According to one or more embodiments of the present disclosure, there is provided an image generation method, further comprising:
[0209] In some optional implementations, the sample control condition includes a sample reference image; and a process of generating the sample reference image includes:
[0210] Generate description text for each clothing category based on the sample try-on images;
[0211] According to each description text, determine the detection frame of the corresponding clothing;
[0212] Perform clothing segmentation according to each detection frame to obtain each garment;
[0213] Based on the sample try-on images, each garment is enhanced to obtain an enhanced image;
[0214] Images of each garment are selected from the enhanced images and retained as sample reference images.
[0215] According to one or more embodiments of the present disclosure, there is provided an image generating apparatus, the apparatus comprising:
[0216] An acquisition module is used to acquire control conditions; wherein the control conditions include a posture image, a prompt text, and a reference image; wherein the prompt text is used to prompt a way to try on the clothing in the reference image;
[0217] A generation module, configured to generate a try-on image using a preset diffusion model, a preset noise image, and control conditions;
[0218] The subject in the try-on image is in a posture corresponding to the posture image, and wears the clothing according to the try-on method.
[0219] The above description is merely a preferred embodiment of the present disclosure and an illustration of the technical principles employed. Those skilled in the art should understand that the scope of disclosure involved in the present disclosure is not limited to the technical solutions formed by the specific combination of the above-mentioned technical features, but also includes other technical solutions formed by any combination of the above-mentioned technical features or their equivalents without departing from the above-mentioned disclosed concepts. For example, a technical solution formed by replacing the above-mentioned features with (but not limited to) technical features with similar functions disclosed in this disclosure.
[0220] In addition, although each operation is described in a specific order, this should not be understood as requiring these operations to be performed in the specific order shown or in a sequential order. Under certain circumstances, multitasking and parallel processing may be advantageous. Similarly, although some specific implementation details have been included in the above discussion, these should not be interpreted as limiting the scope of the present disclosure. Some features described in the context of a separate embodiment can also be implemented in a single embodiment in combination. On the contrary, the various features described in the context of a single embodiment can also be implemented in multiple embodiments individually or in any suitable sub-combination mode.
[0221] Although the subject matter has been described in language specific to structural features and / or methodological logical acts, it should be understood that the subject matter defined in the appended claims is not necessarily limited to the specific features or acts described above. Rather, the specific features and acts described above are merely example forms of implementing the claims.
Claims
1. A method for generating an image, comprising: Acquire control conditions; wherein the control conditions include a posture image, a prompt text, and a reference image; the prompt text is used to prompt a way to try on the clothing in the reference image; and Generate a try-on image based on a preset noise image and the control conditions using a preset diffusion model; The subject being tried on in the try-on image is in a posture corresponding to the posture image, and the subject being tried on wears the garment in the try-on manner.
2. The method according to claim 1, wherein The step of generating a try-on image by using a preset diffusion model and according to a preset noise image and the control conditions includes: Performing feature extraction on the prompt text to obtain text features; Performing feature extraction on the reference image to obtain image features; fusing the text features and the image features to obtain multimodal features; and The multimodal features are fused with the output features of each network layer in the preset diffusion model using a first attention fusion, so that the preset diffusion model outputs the try-on image based on the result of the first attention fusion.
3. The method according to claim 1 or 2, wherein: The method of generating a try-on image according to a preset noise image and the control conditions by using a preset diffusion model further includes: Extracting features from the reference image using a clothing encoder with the same structure as the preset diffusion model to obtain clothing features; A second attention fusion is performed on the clothing features and the output features of each network layer in the preset diffusion model, so that the preset diffusion model outputs the try-on image based on the result of the second attention fusion.
4. The method according to claim 3, wherein: The clothing feature extraction process includes: The clothing encoder performs cross-attention calculation on the reference image and the clothing categories in the prompt text to obtain the clothing features.
5. The method according to claim 2, wherein: In the case where the reference images include reference images corresponding to at least two clothing categories, the process of determining the multimodal features includes: According to each clothing category label in the prompt text, aligning the text features with each image feature corresponding to each clothing category label; Fusing the aligned text features and the image features to obtain a first value vector and a first key vector; The first attention fusion of the multimodal features and the output features of each network layer in the preset diffusion model includes: A first query vector is determined according to the output feature, and a first attention fusion is performed according to the first query vector, the first key vector, and the first value vector.
6. The method according to claim 3, wherein: In a case where the reference image includes reference images corresponding to at least two clothing categories, performing a second attention fusion on the clothing features and the output features of each network layer in the preset diffusion model includes: Determining a second value vector and a second key vector according to the clothing features corresponding to each clothing category and the output features; Determine a second query vector based on the output features; and A second attention fusion is performed based on the second query vector, the second key vector, and the second value vector.
7. The method according to any one of claims 1 to 6, wherein: The process of obtaining the reference image includes at least one of the following: receiving reference images corresponding to the clothing categories according to the order of the clothing categories in the prompt text; receiving the reference image and the clothing category corresponding to the reference image; According to the clothing categories in the prompt text, corresponding reference images are generated in sequence.
8. The method according to any one of claims 1 to 7, wherein: The control condition further includes a protection area image; the protection area image is presented in the non-try-on area of the try-on object.
9. The method according to claim 8, wherein The process of acquiring the posture image and the protection area image includes: An original clothing image of the subject being tried on is received, and the posture image and the protection area image are determined according to the original clothing image.
10. The method according to any one of claims 1 to 9, wherein: The process of constructing the preset diffusion model includes: Add a preset number of rounds of noise to the sample fitting image to obtain a sample noise image; Generate a sample try-on image prediction value based on the sample noise image and sample control conditions using the preset diffusion model; Determining a reconstruction loss according to the sample try-on image and the predicted value of the sample try-on image; Determining a noise prediction loss based on the predicted noise in each round of denoising by the preset diffusion model and the noise added in the corresponding round; and Parameters of the preset diffusion model are adjusted according to the reconstruction loss and the noise prediction loss.
11. The method according to claim 10, wherein: The sample control condition includes a sample reference image; and the generation process of the sample reference image includes: Generating description text of clothing corresponding to each clothing category based on the sample try-on image; Determining a detection frame of the corresponding clothing according to each of the description texts; Perform clothing segmentation according to each detection frame to obtain each clothing; Based on the sample try-on images, each of the garments is enhanced to obtain an enhanced image; Images of the clothing are respectively retained from the enhanced images as sample reference images.
12. An image generating device, comprising: An acquisition module is configured to acquire control conditions, wherein the control conditions include a posture image, a prompt text, and a reference image; the prompt text is used to indicate a way to try on the clothing in the reference image; a generating module configured to generate a try-on image according to a preset noise image and the control conditions using a preset diffusion model; The subject being tried on in the try-on image is in a posture corresponding to the posture image, and the subject being tried on wears the garment in the try-on manner.
13. An electronic device comprising: one or more processors; a storage device for storing one or more programs, When the one or more programs are executed by the one or more processors, the one or more processors implement the image generating method according to any one of claims 1 to 11.
14. A storage medium comprising computer executable instructions, wherein the computer executable instructions are used to perform the image generation method according to any one of claims 1 to 11 when executed by a computer processor.
Citation Information
Patent Citations
Virtual fitting image generation method based on key point clustering drive matching
CN113822175A
Virtual fitting method based on implicit diffusion model
CN117011420A
Virtual fitting system
JP2010070889A
Terminal try-on simulation system and operating and applying method thereof
US20080163344A1
System and Method for Continuous Virtual Fitting Using Virtual Fitting Catalogs
US20210334887A1
Cited By
AI model chart generation method, system and device and medium
CN121685752A
Garment image generation method based on basic element retrieval and replacement
CN122049119A