Model fitting map generation method, device, equipment, storage medium and product

By generating a diffusion model based on preset model fitting images, determining information and performing feature encoding based on model image generation requests, the problem of low efficiency in existing model fitting image generation is solved, and efficient and flexible model fitting image generation is achieved.

CN122244390APending Publication Date: 2026-06-19VIPSHOP (GUANGZHOU) SOFTWARE CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
VIPSHOP (GUANGZHOU) SOFTWARE CO LTD
Filing Date
2026-03-17
Publication Date
2026-06-19

AI Technical Summary

Technical Problem

Existing methods for generating model fitting photos are inefficient, relying on real human photography, which leads to high costs, low efficiency, and poor flexibility.

Method used

A diffusion model is generated by pre-setting model fitting images. The model image generation information is determined based on the model image generation request, feature encoding is performed, and model fitting images are generated using a pre-trained variational autoencoder and diffusion model.

Benefits of technology

It improves the efficiency of generating model fitting photos, reduces costs, and enhances the flexibility and efficiency of the generation process.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122244390A_ABST
    Figure CN122244390A_ABST
Patent Text Reader

Abstract

This application discloses a method, apparatus, device, storage medium, and product for generating model fitting images, relating to the field of data processing technology. The method for generating model fitting images includes: responding to a model image generation request, determining model image generation information based on the request, the model image generation information including at least one of model image generation requirements, clothing flat pattern, and model body posture image; performing feature encoding on the model image generation information to obtain an encoding result; and generating a model fitting image based on the encoding result and a preset model fitting image generation diffusion model to obtain a model fitting image. Since this application generates model fitting images based on the provided model image generation information using a preset model fitting image generation diffusion model, compared to existing methods that rely on manually photographing real models for fitting images, the above method of this application can improve the generation efficiency of model fitting images.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of data processing technology, and in particular to methods, apparatus, devices, storage media, and products for generating model fitting images. Background Technology

[0002] With the rapid development of e-commerce, the visual quality of product displays has become a key factor influencing consumer purchasing decisions. Model images, as a core display format for non-standardized product categories such as clothing, shoes, bags, and cosmetics, use real models in realistic wearing / usage scenarios to intuitively convey the product's fit, texture, styling effects, and user experience, significantly enhancing product appeal and conversion rates. Currently, online shopping model images primarily rely on real-life photography, a process that includes hiring professional models, makeup and styling, venue / scene setup, multi-device shooting, and post-production retouching. While this method can present a realistic visual effect, its high cost, low efficiency, and lack of flexibility are becoming increasingly prominent pain points. Therefore, improving the update efficiency of model images has become a key issue for the platform to address. Summary of the Invention

[0003] The main objective of this application is to provide a method, apparatus, device, storage medium, and product for generating model fitting images, aiming to solve the technical problem of low efficiency in the existing model fitting image generation.

[0004] To achieve the above objectives, this application proposes a method for generating model fitting images, the method comprising: In response to a model image generation request, model image generation information is determined based on the model image generation request. The model image generation information includes at least one of the following: model image generation requirements, clothing flat pattern, and model body posture diagram. The generated model image information is feature-encoded to obtain the encoding result; Based on the encoding results and the preset model fitting image, a diffusion model is generated to produce the model fitting image.

[0005] Optionally, the step of performing feature encoding on the model image generation information to obtain the encoding result includes: The model image generation requirement is feature-encoded using a text encoder to obtain the generation requirement encoding result; And / or, by using an image encoder to perform feature encoding on the garment plan view and / or the model's body posture view, a conditional encoding result is obtained; The encoding result is determined based on the requirement encoding result and / or the condition encoding result.

[0006] Optionally, the preset model fitting image generation diffusion model includes a normalization layer, a multilayer perceptron, an attention mechanism, and a projection module; The step of generating model fitting images based on the encoding results and preset model fitting images to generate a diffusion model, thereby obtaining model fitting images, includes: The encoding result is normalized by the normalization layer to obtain the normalized result; The normalization result is processed by the multilayer perceptron and the attention mechanism respectively to obtain the processing result; The processing result is mapped based on the projection module to obtain a model fitting image.

[0007] Optionally, the step of normalizing the encoding result through the normalization layer to obtain a normalized result includes: Determine the dimensional information of the encoded result; The encoding results are concatenated based on the dimensional information to obtain a concatenated result; The splicing result is normalized by the normalization layer to obtain the normalized result.

[0008] Optionally, before generating the model fitting image based on the encoding result and the preset model fitting image, the process further includes: Obtain sample model images, perform multi-labeling on the sample model images, and obtain target sample data; The ground truth map in the target sample data is compressed by a preset variational autoencoder to obtain a latent space representation; The initial diffusion model is trained based on the latent space representation and the target sample data to obtain a preset model fitting image generation diffusion model.

[0009] Optionally, the step of determining model image generation information based on the model image generation request in response to the model image generation request includes: In response to a model image generation request, a conditional image is determined based on the model image generation request, the conditional image including a garment plan view and / or a model body pose view; The model image generation requirements are determined based on the conditional image.

[0010] Furthermore, to achieve the above objectives, this application also proposes a model fitting image generation device, which includes: A response module is used to respond to a model image generation request and determine model image generation information based on the model image generation request. The model image generation information includes at least one of a model image generation requirement, a clothing flat pattern, and a model body posture diagram. The encoding module is used to perform feature encoding on the model image generation information to obtain the encoding result; The generation module is used to generate model fitting images by generating a diffusion model based on the encoding results and preset model fitting images, thereby obtaining model fitting images.

[0011] In addition, to achieve the above objectives, this application also proposes a model fitting image generation device, the device comprising: a memory, a processor, and a computer program stored in the memory and executable on the processor, the computer program being configured to implement the steps of the model fitting image generation method described above.

[0012] In addition, to achieve the above objectives, this application also proposes a storage medium, which is a computer-readable storage medium, on which a computer program is stored, and when the computer program is executed by a processor, it implements the steps of the model fitting image generation method described above.

[0013] In addition, to achieve the above objectives, this application also provides a computer program product, which includes a computer program that, when executed by a processor, implements the steps of the model fitting image generation method described above.

[0014] This application responds to a model image generation request by determining model image generation information based on the request. This information includes at least one of a model image generation requirement, a garment flat pattern, and a model body pose diagram. The model image generation information is then feature-encoded to obtain an encoding result. Based on the encoding result and a preset model fitting image generation diffusion model, a model fitting image is generated, resulting in a model fitting image. Because this application generates model fitting images based on the provided model image generation information using a preset model fitting image generation diffusion model, compared to existing methods that rely on manually photographing real models for fitting images, this application's method improves the efficiency of model fitting image generation. Attached Figure Description

[0015] The accompanying drawings, which are incorporated in and form part of this specification, illustrate embodiments consistent with this application and, together with the description, serve to explain the principles of this application.

[0016] To more clearly illustrate the technical solutions in the embodiments of this application or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, for those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0017] Figure 1 This is a flowchart illustrating an embodiment of the method for generating model fitting images in this application. Figure 2 This is a schematic diagram of a sample model image provided in Embodiment 1 of the method for generating model fitting images in this application; Figure 3 This is a schematic diagram of a sample model image generated using the data generation method provided in Embodiment 1 of the model fitting image generation method of this application; Figure 4 A schematic diagram of the model training process provided in Embodiment 1 of the model fitting image generation method of this application; Figure 5 This is a flowchart illustrating Embodiment 2 of the method for generating model fitting images in this application. Figure 6 This is a schematic diagram of the model structure provided in Embodiment 2 of the method for generating model fitting images in this application; Figure 7 This is a schematic diagram of the model generation result provided in Embodiment 2 of the model fitting image generation method of this application; Figure 8 This is a schematic diagram of the model generation result provided in Embodiment 2 of the model fitting image generation method of this application; Figure 9 This is a schematic diagram of the module structure of the model fitting image generation device according to an embodiment of this application; Figure 10 This is a schematic diagram of the device structure of the hardware operating environment involved in the model fitting image generation method in this application embodiment.

[0018] The purpose, features, and advantages of this application will be further explained in conjunction with the embodiments and with reference to the accompanying drawings. Detailed Implementation

[0019] It should be understood that the specific embodiments described herein are merely illustrative of the technical solutions of this application and are not intended to limit this application.

[0020] To better understand the technical solution of this application, a detailed description will be provided below in conjunction with the accompanying drawings and specific implementation methods.

[0021] The main solution of this application embodiment is as follows: In response to a model image generation request, model image generation information is determined based on the model image generation request. The model image generation information includes at least one of a model image generation requirement, a clothing flat pattern, and a model body posture diagram; the model image generation information is feature-encoded to obtain an encoding result; and a model fitting image is generated based on the encoding result and a preset model fitting image generation diffusion model to obtain a model fitting image. Since this application generates model fitting images based on the provided model image generation information using a preset model fitting image generation diffusion model, compared to existing methods that rely on real human-taken model fitting images, the above method of this application can improve the generation efficiency of model fitting images.

[0022] It should be noted that the executing entity in this embodiment can be a computing service device with data processing, network communication, and program execution functions, such as a tablet computer, personal computer, or mobile phone, or an electronic device or a model fitting image generation device capable of performing the above functions. The following description uses a model fitting image generation device as an example to illustrate this embodiment and the subsequent embodiments.

[0023] Based on this, embodiments of this application provide a method for generating model fitting images, referring to... Figure 1 , Figure 1 This is a flowchart illustrating the method for generating model fitting images in this application, as provided in Embodiment 1.

[0024] In this embodiment, the method for generating model fitting images includes the following steps: Step S10: In response to the model image generation request, determine model image generation information based on the model image generation request. The model image generation information includes at least one of the following: model image generation requirements, clothing flat pattern, and model body posture diagram. It should be noted that the model image generation request can be user-triggered, a command to generate a model fitting image based on provided model image generation information. The model image generation requirement can be in scenarios such as e-commerce, fashion, virtual fitting, and advertising marketing, based on business objectives, requiring the generation of high-quality, realistic virtual model images that meet specific conditions (such as pose, clothing, background, style, openness, tightness / looseness, garment length, shoe length, etc.). The model image generation information can include at least one of the following: model image generation requirement, clothing flat image, and model body pose image. Typically, the model image generation information includes the model image generation requirement, clothing flat image, and model body pose image, indicating that the corresponding clothing in the clothing flat image is generated as a model fitting image on the model corresponding to the model body pose image based on the current requirement and the provided model body pose image. If only the model image generation requirement is included, it indicates that a model fitting image is directly generated based on the user-described model image generation requirement. This embodiment also includes the ability to generate realistic model images based on a plain product lay image (without a model) (i.e., generating model fitting images when only a clothing flat image is provided; in addition, model image generation requirements may also be included). It can also replace the model on the original product model image to improve the conversion effect of the product. In this case, the model image generation information may include the model fitting image provided by the user, as well as multiple additional model images to replace the model in the original model fitting image.

[0025] Step S20: Perform feature encoding on the model image generation information to obtain the encoding result; It should be noted that the feature encoding of the model image generation information to obtain the encoding result can be obtained by using an encoder to perform feature encoding on the model image generation information.

[0026] Step S30: Generate model fitting images by generating a diffusion model based on the encoding results and preset model fitting images, and obtain model fitting images.

[0027] It should be noted that the preset model fitting image generation diffusion model can be a pre-trained diffusion model used to generate model fitting images based on the feature encoding corresponding to the provided model image generation information. For example, the FLUX model. The FLUX model is a family of text-to-image models launched by Black Forest Labs. It is based on the Diffusion Transformer (DiT) / Rectified Flow route, with a family of approximately 12 billion parameters, emphasizing high fidelity, strong instruction compliance, and efficient generation.

[0028] Before step S30, the method further includes: obtaining a sample model image, performing multi-labeling on the sample model image, and obtaining target sample data; The ground truth map in the target sample data is compressed by a preset variational autoencoder to obtain a latent space representation; The initial diffusion model is trained based on the latent space representation and the target sample data to obtain a preset model fitting image generation diffusion model.

[0029] It should be noted that the sample model images may include collected model body posture images (used only to demonstrate the model's posture), flat lay images of clothing (including tops, bottoms, shoes, and skirts, etc.), and / or model fitting images (images showing the model after trying on clothes). See reference. Figure 2 , Figure 2 This is a schematic diagram of a sample model image provided in Embodiment 1 of the method for generating model fitting images in this application. Figure 2 Models A and B in the image can be the model's fitting photos. Figure 2 The clothing image in the image is a flat lay image of the clothing. Sample model images can also be obtained through data generation (i.e., generating clothing image models). See reference... Figure 3 , Figure 3This is a schematic diagram of a sample model image generated using the data generation method provided in Embodiment 1 of the model fitting image generation method of this application. Multi-labeling of the sample model image can include labeling it as tuck, open, tight / loose, garment / skirt length, connected (top and bottom), and shoe prompt. In the clothing industry, "tuck" refers to tucking the hem of a garment into the lower garment (pants / skirt), commonly seen in shirts, T-shirts, sweaters, and other tops. If the model's top hem is tucked into the lower garment in an image, it is labeled "tuck"; otherwise, it is not labeled or is labeled "no-tuck." "Connected (top and bottom)" refers to a garment where the top and bottom are one piece, i.e., a jumpsuit / one-piece dress (such as a jumpsuit, romper, one-piece swimsuit, etc.). If the model in the image is wearing a one-piece garment (top and bottom as one piece), it is labeled "connected (top and bottom); if it is a separate top and pants / skirt," it is not labeled. The "shoe prompt" refers to text descriptions or tags related to shoes, which may include: shoe type (high heels, sneakers, boots, sandals, slippers, etc.), shoe color / material (red leather boots, white sneakers), and the relationship between shoes and clothing (such as "pair with ankle boots" or "nude high heels make legs look longer").

[0030] It should be noted that the preset variational autoencoder can be a pre-trained VAE (Variational Autoencoder), which is an encoder-decoder model pre-trained on large-scale data, used to compress images into low-dimensional latent representations and then reconstruct them. In this embodiment, the conditional images (clothing flat images, model body pose images) are directly encoded into the latent space by reusing the pre-trained VAE of the diffusion model, avoiding architectural complexity. This method eliminates the need for auxiliary modules and reduces the total parameters. These design choices enable easy expansion to support multiple control types with low overhead. (Refer to...) Figure 4 , Figure 4 A schematic diagram of the model training process provided in Embodiment 1 of the model fitting image generation method of this application; Figure 4 The generation result is the model fitting image generated by the pre-defined model fitting image generation diffusion model, which is based on the latent space representation and the target sample data, trained on the initial diffusion model. Figure 4 (The generated results) Figure 4 It also includes a ground truth (GT) image, which is a real model trying on clothes that corresponds to the input flat lay image of the clothes.

[0031] In this embodiment, in response to a model image generation request, model image generation information is determined based on the request. This information includes at least one of the following: model image generation requirements, clothing flat pattern, and model body pose diagram. The model image generation information is then feature-encoded to obtain an encoding result. Based on the encoding result and a preset model fitting image generation diffusion model, a model fitting image is generated, resulting in a model fitting image. Since this embodiment generates model fitting images based on the provided model image generation information using a preset model fitting image generation diffusion model, compared to existing methods that rely on manually photographing real models for fitting images, this embodiment improves the efficiency of model fitting image generation.

[0032] Based on the first embodiment of this application, in the second embodiment of this application, the content that is the same as or similar to that in the first embodiment described above can be referred to the above description, and will not be repeated hereafter. Based on this, please refer to... Figure 5 , Figure 5 This is a flowchart illustrating the second embodiment of the method for generating model fitting images in this application. Step S20 further includes the following steps: Step S201: Encode the model image generation requirements using a text encoder to obtain the generation requirement encoding result; It should be noted that the feature encoding of the model image generation requirement by the text encoder to obtain the generation requirement encoding result can be the word segmentation / feature encoding of the model image generation requirement to obtain the text tokens corresponding to the model image generation requirement. "X tokens" generally refer to the representation after word segmentation / feature encoding (Tokenization) of the image or text, which is convenient for network processing.

[0033] Step S202: and / or, perform feature encoding on the garment plan view and / or the model body posture view using an image encoder to obtain a conditional encoding result; It should be noted that if the model image generation information includes both a garment plan view and a model body pose image, then feature encoding needs to be performed on both the garment plan view and the model body pose image to obtain conditional encoding results corresponding to the garment plan view and the model body pose image, respectively. If the model image generation information includes either a garment plan view or a model body pose image, then feature encoding needs to be performed on either the included garment plan view or the included model body pose image to obtain conditional encoding results corresponding to either the garment plan view or the model body pose image.

[0034] Step S203: Determine the encoding result based on the requirement encoding result and / or the condition encoding result.

[0035] It should be noted that the encoding result refers to all the encoding results calculated in the above steps.

[0036] Furthermore, the preset model fitting image generation diffusion model includes a normalization layer, a multilayer perceptron, an attention mechanism, and a projection module; The step of generating model fitting images based on the encoding results and preset model fitting images to generate a diffusion model, thereby obtaining model fitting images, includes: The encoding result is normalized by the normalization layer to obtain the normalized result; The normalization result is processed by the multilayer perceptron and the attention mechanism respectively to obtain the processing result; The processing result is mapped based on the projection module to obtain a model fitting image.

[0037] It should be noted that this can be referred to Figure 6 , Figure 6 This is a schematic diagram of the model structure provided in Embodiment 2 of the method for generating model fitting images according to this application; the process for generating model fitting images in this embodiment includes: 1. Input data: The model receives the text (i.e., the model image generation requirement), noisy image data (Xnoised, during the model training phase), and conditional information (Candid, clothing flat pattern and / or model body pose image) as input.

[0038] 2. Encoding: The text and the noisy image are encoded by a text encoder and an image encoder, respectively, to obtain the corresponding feature representations, i.e., the encoding results.

[0039] 3. Fusion Processing: The encoded result is fed into DiT Blocks (which can be Transformer-based diffusion model blocks). After processing by multiple DiT Blocks, it is further processed by incorporating information such as noise. The DiT Blocks module receives the encoded result and other inputs, and trains and optimizes the model through operations such as normalization, multilayer perceptron (MLP), attention mechanism (Attn), and projection (proj), finally outputting a processed result that combines text, image, and conditional information.

[0040] in, Figure 6The Text Encoder is responsible for converting input text information into feature vectors that the model can process. The Image Encoder converts input images into feature vectors and extracts relevant image information. DiT Blocks are Transformer blocks based on the Diffusion Model, used for further processing and fusion of encoded text and image features. These blocks may be applied iteratively multiple times to gradually optimize the feature representation. Noise is noise introduced during model training. Cond_text is textual conditions subdivided from the conditional information (Cond), used to more precisely guide model processing, such as specifying the style, theme, and other related textual descriptions of the generated image. Cond_cloth is clothing conditions subdivided from the conditional information (Cond), which may be used for tasks such as image generation and modification related to clothing, such as specifying the type and color of clothing. Norm is a normalization operation used to standardize input data, giving it specific statistical properties, which helps improve the training stability and performance of the model. MLP is a Multi-Layer Perceptron, a feedforward neural network used for nonlinear transformation and feature extraction of input data. Attn: Attention Mechanism, which enables the model to automatically focus on important parts of the input data. In multimodal processing, it can effectively fuse information from different modalities, enhancing the model's ability to capture key information. proj: Projection operation, typically used to map feature vectors to a specific dimension or space to meet the requirements of subsequent processing or output. In this embodiment, tokens representing different control signals (such as Cond_text, Cond_cloth) and tokens representing noisy images are concatenated into a long sequence. Then, multimodal processing components (MLP, Attn, etc.) in modules such as DiT Blocks are used for joint modeling and processing, thereby achieving effective utilization of multimodal input and task completion.

[0041] The step of normalizing the encoding result through the normalization layer to obtain the normalized result includes: determining the dimensional information of the encoding result; The encoding results are concatenated based on the dimensional information to obtain a concatenated result; The splicing result is normalized by the normalization layer to obtain the normalized result.

[0042] It should be noted that this embodiment concatenates tokens from various control signals (model image generation information) with tokens representing noisy images to form a single long sequence. Assuming the token dimension of each control signal is (B, seq_len, hidden_dim), then concatenation is performed along this dimension. B is the batch size, which can be the number of samples / data processed simultaneously during a single training / inference iteration. seq_len is the sequence length, which can be the number of tokens in a single sample. hidden_dim is the feature dimension of each token. See [reference needed]. Figure 7 , Figure 7 This is a schematic diagram of the model generation result provided in Embodiment 2 of the model fitting image generation method of this application. Figure 7 The document provides reference images of dresses and shoes, and sets the model image generation requirement to regenerate the model. The result is as follows. Figure 7 As shown. Figure 8 This is a schematic diagram of the model generation result provided in Embodiment 2 of the model fitting image generation method of this application; Figure 8 The document provides images of models in various poses, including tops and matching bottoms, resulting in the following: Figure 8 As shown.

[0043] This embodiment uses a text encoder to encode the features of the model image generation requirement, obtaining a generation requirement encoding result; and / or uses an image encoder to encode the features of the clothing plan view and / or the model's body pose view, obtaining a conditional encoding result; the encoding result is determined based on the requirement encoding result and / or the conditional encoding result. This embodiment concatenates tokens from various control signals with tokens representing noisy images to form a single long sequence. (Assuming the token dimension of each control signal is (B, seq_len, hidden_dim), then concatenation is performed along the seq_len dimension.) This unified sequence is then processed by multimodal attention (MM-Attention), which enables joint modeling of multimodal inputs.

[0044] It should be noted that the above examples are only for understanding this application and do not constitute a limitation on the model fitting image generation method of this application. Any simple transformations based on this technical concept are within the protection scope of this application.

[0045] This application also provides a device for generating model fitting images; please refer to... Figure 9 The model fitting image generation device includes: Response module 10 is used to respond to a model image generation request and determine model image generation information based on the model image generation request. The model image generation information includes at least one of model image generation requirements, clothing flat pattern, and model body posture diagram. Encoding module 20 is used to perform feature encoding on the model image generation information to obtain an encoding result; The generation module 30 is used to generate model fitting images by generating a diffusion model based on the encoding result and the preset model fitting image, thereby obtaining the model fitting image.

[0046] In this embodiment, in response to a model image generation request, model image generation information is determined based on the request. This information includes at least one of the following: model image generation requirements, clothing flat pattern, and model body pose diagram. The model image generation information is then feature-encoded to obtain an encoding result. Based on the encoding result and a preset model fitting image generation diffusion model, a model fitting image is generated, resulting in a model fitting image. Since this embodiment generates model fitting images based on the provided model image generation information using a preset model fitting image generation diffusion model, compared to existing methods that rely on manually photographing real models for fitting images, this embodiment improves the efficiency of model fitting image generation.

[0047] The model fitting image generation device provided in this application, employing the model fitting image generation method in the above embodiments, can solve the technical problem of low efficiency in existing model fitting image generation. Compared with the prior art, the beneficial effects of the model fitting image generation device provided in this application are the same as those of the model fitting image generation method provided in the above embodiments, and other technical features in the model fitting image generation device are the same as those disclosed in the methods of the above embodiments, and will not be repeated here.

[0048] This application provides a model fitting image generation device, which includes: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to perform the model fitting image generation method in the above embodiment 1.

[0049] The following is for reference. Figure 10The diagram illustrates a structural schematic of a model fitting image generation device suitable for implementing embodiments of this application. The model fitting image generation device in embodiments of this application may include, but is not limited to, mobile terminals such as mobile phones, laptops, digital broadcast receivers, PDAs (Personal Digital Assistants), PADs (Portable Application Description), PMPs (Portable Media Players), in-vehicle terminals (e.g., in-vehicle navigation terminals), and fixed terminals such as digital TVs and desktop computers. Figure 10 The illustrated model fitting image generation device is merely an example and should not impose any limitations on the functionality and scope of use of the embodiments of this application.

[0050] like Figure 10 As shown, the model fitting image generation device may include a processing unit 1001 (e.g., a central processing unit, a graphics processing unit, etc.), which can perform various appropriate actions and processes according to a program stored in a read-only memory (ROM) 1002 or a program loaded from a storage device 1003 into a random access memory (RAM) 1004. The RAM 1004 also stores various programs and data required for the operation of the model fitting image generation device. The processing unit 1001, ROM 1002, and RAM 1004 are interconnected via a bus 1005. An input / output (I / O) interface 1006 is also connected to the bus. Typically, the following systems can be connected to the I / O interface 1006: input devices 1007 including, for example, a touchscreen, touchpad, keyboard, mouse, image sensor, microphone, accelerometer, gyroscope, etc.; output devices 1008 including, for example, a liquid crystal display (LCD), speaker, vibrator, etc.; storage devices 1003 including, for example, magnetic tape, hard disk, etc.; and communication devices 1009. The communication device 1009 allows the model fitting image generation device to communicate wirelessly or wiredly with other devices to exchange data. Although the figure shows model fitting image generation devices with various systems, it should be understood that implementation or possession of all the systems shown is not required. More or fewer systems may be implemented alternatively.

[0051] Specifically, according to the embodiments disclosed in this application, the processes described above with reference to the flowcharts can be implemented as computer software programs. For example, embodiments disclosed in this application include a computer program product comprising a computer program carried on a computer-readable medium, the computer program containing program code for performing the methods shown in the flowcharts. In such embodiments, the computer program can be downloaded and installed from a network via a communication device, or installed from storage device 1003, or installed from ROM 1002. When the computer program is executed by processing device 1001, it performs the functions defined in the methods of the embodiments disclosed in this application.

[0052] The model fitting image generation device provided in this application, employing the model fitting image generation method in the above embodiments, can solve the technical problem of low efficiency in existing model fitting image generation. Compared with the prior art, the beneficial effects of the model fitting image generation device provided in this application are the same as those of the model fitting image generation method provided in the above embodiments, and other technical features in this model fitting image generation device are the same as those disclosed in the previous embodiment method, and will not be repeated here.

[0053] It should be understood that the various parts disclosed in this application can be implemented using hardware, software, firmware, or a combination thereof. In the description of the above embodiments, specific features, structures, materials, or characteristics can be combined in any suitable manner in one or more embodiments or examples.

[0054] The above description is merely a specific embodiment of this application, but the scope of protection of this application is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the scope of the technology disclosed in this application should be included within the scope of protection of this application. Therefore, the scope of protection of this application should be determined by the scope of the claims.

[0055] This application provides a computer-readable storage medium having computer-readable program instructions (i.e., a computer program) stored thereon, the computer-readable program instructions being used to execute the model fitting image generation method in the above embodiments.

[0056] The computer-readable storage medium provided in this application may be, for example, a USB flash drive, but is not limited to, electrical, magnetic, optical, electromagnetic, infrared, or semiconductor systems, devices, or any combination thereof. More specific examples of computer-readable storage media may include, but are not limited to: electrical connections having one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination thereof. In this embodiment, the computer-readable storage medium may be any tangible medium containing or storing a program that can be used by or in conjunction with an instruction execution system, system, or device. The program code contained on the computer-readable storage medium may be transmitted using any suitable medium, including but not limited to: wires, optical cables, RF (Radio Frequency), etc., or any suitable combination thereof.

[0057] The aforementioned computer-readable storage medium may be included in the model fitting image generation device; or it may exist independently and not assembled into the model fitting image generation device.

[0058] Computer program code for performing the operations of this application can be written in one or more programming languages ​​or a combination thereof. These programming languages ​​include object-oriented programming languages—such as Python, Java, Smalltalk, and C++—and conventional procedural programming languages—such as the "C" language or similar programming languages. The program code can be executed entirely on the user's computer, partially on the user's computer, as a standalone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving remote computers, the remote computer can be connected to the user's computer via any type of network—including a Local Area Network (LAN) or a Wide Area Network (WAN)—or can be connected to an external computer (e.g., via the Internet using an Internet service provider).

[0059] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of this application. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of code containing one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions indicated in the blocks may occur in a different order than those indicated in the drawings. For example, two consecutively indicated blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, can be implemented using a dedicated hardware-based system that performs the specified function or operation, or using a combination of dedicated hardware and computer instructions.

[0060] The modules described in the embodiments of this application can be implemented in software or hardware. The names of the modules do not necessarily limit the functionality of the unit itself.

[0061] The readable storage medium provided in this application is a computer-readable storage medium that stores computer-readable program instructions (i.e., a computer program) for executing the above-described model fitting image generation method, which can solve the technical problem of low efficiency in existing model fitting image generation. Compared with the prior art, the beneficial effects of the computer-readable storage medium provided in this application are the same as the beneficial effects of the model fitting image generation method provided in the above embodiments, and will not be repeated here.

[0062] This application also provides a computer program product, including a computer program that, when executed by a processor, implements the steps of the model fitting image generation method described above.

[0063] The computer program product provided in this application can solve the technical problem of low efficiency in generating model fitting images in existing technologies. Compared with the prior art, the beneficial effects of the computer program product provided in this application are the same as those of the model fitting image generation method provided in the above embodiments, and will not be repeated here.

[0064] The above description is only a part of the embodiments of this application and does not limit the scope of protection of this application. All equivalent structural transformations made under the technical concept of this application and using the content of this application specification and drawings, or direct / indirect applications in other related technical fields, are included in the scope of protection of this application.

Claims

1. A method for generating model fitting images, characterized in that, The method for generating model fitting photos includes the following steps: In response to a model image generation request, model image generation information is determined based on the model image generation request. The model image generation information includes at least one of the following: model image generation requirements, clothing flat pattern, and model body posture diagram. The generated model image information is feature-encoded to obtain the encoding result; Based on the encoding results and the preset model fitting image, a diffusion model is generated to produce the model fitting image.

2. The method for generating model fitting images as described in claim 1, characterized in that, The step of performing feature encoding on the generated model image information to obtain the encoding result includes: The model image generation requirement is feature-encoded using a text encoder to obtain the generation requirement encoding result; And / or, by using an image encoder to perform feature encoding on the garment plan view and / or the model's body posture view, a conditional encoding result is obtained; The encoding result is determined based on the requirement encoding result and / or the condition encoding result.

3. The method for generating model fitting images as described in claim 2, characterized in that, The preset model fitting image generation diffusion model includes a normalization layer, a multilayer perceptron, an attention mechanism, and a projection module. The step of generating model fitting images based on the encoding results and preset model fitting images to generate a diffusion model, thereby obtaining model fitting images, includes: The encoding result is normalized by the normalization layer to obtain the normalized result; The normalization result is processed by the multilayer perceptron and the attention mechanism respectively to obtain the processing result; The processing result is mapped based on the projection module to obtain a model fitting image.

4. The method for generating model fitting images as described in claim 3, characterized in that, The step of normalizing the encoding result through the normalization layer to obtain the normalized result includes: Determine the dimensional information of the encoded result; The encoding results are concatenated based on the dimensional information to obtain a concatenated result; The splicing result is normalized by the normalization layer to obtain the normalized result.

5. The method for generating model fitting images as described in any one of claims 1-4, characterized in that, Before generating the model fitting image by generating a diffusion model based on the encoding result and the preset model fitting image, the process further includes: Obtain sample model images, perform multi-labeling on the sample model images, and obtain target sample data; The ground truth map in the target sample data is compressed by a preset variational autoencoder to obtain a latent space representation; The initial diffusion model is trained based on the latent space representation and the target sample data to obtain a preset model fitting image generation diffusion model.

6. The method for generating model fitting images as described in any one of claims 1-4, characterized in that, The step of responding to a model image generation request and determining model image generation information based on the model image generation request includes: In response to a model image generation request, a conditional image is determined based on the model image generation request, the conditional image including a garment plan view and / or a model body pose view; The model image generation requirements are determined based on the conditional image.

7. A device for generating model fitting images, characterized in that, The model fitting image generation device includes: A response module is used to respond to a model image generation request and determine model image generation information based on the model image generation request. The model image generation information includes at least one of a model image generation requirement, a clothing flat pattern, and a model body posture diagram. The encoding module is used to perform feature encoding on the model image generation information to obtain the encoding result; The generation module is used to generate model fitting images by generating a diffusion model based on the encoding results and preset model fitting images, thereby obtaining model fitting images.

8. A device for generating model fitting images, characterized in that, The device includes: a memory, a processor, and a computer program stored in the memory and executable on the processor, the computer program being configured to implement the steps of the model fitting image generation method as described in any one of claims 1 to 6.

9. A storage medium, characterized in that, The storage medium is a computer-readable storage medium, and a computer program is stored on the storage medium. When the computer program is executed by a processor, it implements the steps of the model fitting image generation method as described in any one of claims 1 to 6.

10. A computer program product, characterized in that, The computer program product includes a computer program that, when executed by a processor, implements the steps of the model fitting image generation method as described in any one of claims 1 to 6.