Method and device for virtual fitting of objects
By pre-training and fitting training on the diffusion model, the problem of virtual fitting model insufficient understanding of clothing categories is solved, and a better dress-up effect and user experience is achieved.
Patent Information
- Application Number
- CN202410494388.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-04-23
- Publication Date
- 2025-08-29
- Estimated Expiration
- 2044-04-23
AI Technical Summary
The existing virtual fitting models have weak understanding of clothing categories, resulting in failure in dressing and poor dressing effect, and there are problems such as loss of clothes texture and unnatural.
By obtaining training images and clothing images, determining mask segmentation images and clothing categories, pre-training and fitting training of the diffusion model, improving the model's clothing generation ability and category understanding ability, and using training clothing to replace the clothing in the target image.
Reduce the number of failures in the dressing, improve the dressing effect, and improve user satisfaction.
Smart Images

Figure CN118365940B_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the field of image processing technology, and in particular to a method and device for virtual fitting of an object. Background Art
[0002] Virtual fitting involves replacing a person's clothing in an image with another piece of clothing. It's widely used in apparel e-commerce and can significantly improve the user shopping experience. However, current virtual fitting models have a weak understanding of clothing categories and are unable to comprehend the complex and diverse range of clothing categories, leading to failed and poorly rendered outfit changes. This can cause problems such as texture loss and unnatural fit on the model. Summary of the Invention
[0003] In view of this, the embodiments of the present disclosure provide a method, apparatus, electronic device, and computer-readable storage medium for virtual fitting of an object to solve the problems of failure and poor effect of virtual fitting in the prior art.
[0004] In a first aspect of an embodiment of the present disclosure, a method for virtual fitting of an object is provided, comprising: obtaining a first training image, determining a first mask segmentation map that obscures clothing of an object in the first training image and the category of clothing of the object in the first training image; pre-training a diffusion model using the first mask segmentation map, the first training image and the category of clothing of the object in the first training image to improve the clothing generation capability and clothing category understanding capability of the diffusion model; obtaining a second training image and a training clothing image, determining a second mask segmentation map that obscures clothing of the object in the second training image and the category of training clothing in the training clothing image; training the diffusion model for fitting using the second mask segmentation map, the training clothing image and the category of training clothing in the training clothing image, so that the diffusion model replaces clothing of the object in the second training image with training clothing; using the diffusion model after fitting training as a virtual fitting model, and using the virtual fitting model to provide a virtual fitting service for the target object.
[0005] According to a second aspect of an embodiment of the present disclosure, there is provided a device for virtual fitting of an object, comprising: a first acquisition module configured to acquire a first training image, determine a first mask segmentation map that obscures clothing of an object in the first training image, and the category of clothing of the object in the first training image; a pre-training module configured to pre-train a diffusion model using the first mask segmentation map, the first training image, and the category of clothing of the object in the first training image, so as to improve the clothing generation capability and clothing category understanding capability of the diffusion model; a second acquisition module configured to acquire a second training image and a training clothing image, determine a second mask segmentation map that obscures clothing of the object in the second training image, and the category of training clothing in the training clothing image; a fitting training module configured to perform fitting training on the diffusion model using the second mask segmentation map, the training clothing image, and the category of training clothing in the training clothing image, so that the diffusion model replaces clothing of the object in the second training image with training clothing; and a virtual fitting module configured to use the diffusion model after fitting training as a virtual fitting model, and use the virtual fitting model to provide virtual fitting services for the target object.
[0006] According to a third aspect of an embodiment of the present disclosure, an electronic device is provided, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor implements the steps of the above method when executing the computer program.
[0007] According to a fourth aspect of the embodiments of the present disclosure, a computer-readable storage medium is provided, which stores a computer program. When the computer program is executed by a processor, the steps of the above method are implemented.
[0008] Compared with the prior art, the disclosed embodiment has the following advantages: obtaining a first training image, determining a first mask segmentation map of clothing covering an object in the first training image and the category of clothing of the object in the first training image; using the first mask segmentation map, the first training image and the category of clothing of the object in the first training image to pre-train a diffusion model to improve the clothing generation ability and clothing category understanding ability of the diffusion model; obtaining a second training image and a training clothing image, determining a second mask segmentation map of clothing covering an object in the second training image and the category of training clothing in the training clothing image; using the second mask segmentation map, the training clothing image and the category of training clothing in the training clothing image to perform fitting training on the diffusion model, so that the diffusion model replaces the clothing of the object in the second training image with the training clothing; using the diffusion model after fitting training as a virtual fitting model, and using the virtual fitting model to provide a virtual fitting service for the target object. The above technical means can solve the problems of virtual fitting failure and poor fitting effect in the prior art, thereby reducing the number of fitting failures, improving fitting effect and improving user satisfaction. BRIEF DESCRIPTION OF THE DRAWINGS
[0009] In order to more clearly illustrate the technical solutions in the embodiments of the present disclosure, the following briefly introduces the drawings required for use in the embodiments or descriptions of the prior art. Obviously, the drawings in the following description are only some embodiments of the present disclosure. For ordinary technicians in this field, other drawings can be obtained based on these drawings without any creative work.
[0010] Figure 1 1 is a flow chart of a method for virtual fitting of an object provided by an embodiment of the present disclosure;
[0011] Figure 2 1 is a flow chart of a fitting training method provided by an embodiment of the present disclosure;
[0012] Figure 3 1 is a schematic diagram of a device structure for a subject to virtually try on clothes, provided in an embodiment of the present disclosure;
[0013] Figure 4 This is a schematic diagram of the structure of an electronic device provided by an embodiment of the present disclosure. DETAILED DESCRIPTION
[0014] In the following description, specific details such as particular system structures and techniques are provided for purposes of illustration, not limitation, to facilitate a thorough understanding of the embodiments of the present disclosure. However, it will be apparent to those skilled in the art that other embodiments of the present disclosure may be implemented without these specific details. In other cases, detailed descriptions of well-known systems, devices, circuits, and methods are omitted to avoid obscuring the present disclosure with unnecessary detail.
[0015] A method and apparatus for virtual fitting of an object according to an embodiment of the present disclosure will be described in detail below with reference to the accompanying drawings.
[0016] Figure 1 The present invention provides a flowchart of a method for virtually fitting an object with clothes according to an embodiment of the present invention. Figure 1 The method for virtually fitting a subject may be executed by a computer or a server, or by software on the computer or the server.
[0017] like Figure 1 As shown, the method for the object to virtually try on clothes includes:
[0018] S101, obtaining a first training image, determining a first mask segmentation map of clothing covering an object in the first training image and a category of the clothing of the object in the first training image;
[0019] S102, pre-training a diffusion model using the first mask segmentation image, the first training image, and the category of clothing of the object in the first training image, so as to improve the diffusion model's clothing generation capability and its ability to understand clothing categories;
[0020] S103, acquiring a second training image and a training clothing image, and determining a second mask segmentation map of clothing that obscures the object in the second training image and a category of the training clothing in the training clothing image;
[0021] S104, performing fitting training on the diffusion model using the second mask segmentation image, the training clothing image, and the category of the training clothing in the training clothing image, so that the diffusion model replaces the clothing of the object in the second training image with the training clothing;
[0022] S105 , using the diffusion model trained for fitting as a virtual fitting model, and providing a virtual fitting service for the target object using the virtual fitting model.
[0023] The first and second training images are images of people wearing clothing. The training clothing images are images of training clothing, with the subject being a person. Clothing categories include: long-sleeved T-shirts, short-sleeved T-shirts, suspenders, jackets, shorts, trousers, short skirts, long skirts, suits, coats, and more.
[0024] Pre-training the diffusion model improves its clothing generation capabilities and its ability to understand clothing categories. This addresses the existing virtual fitting model's weak understanding of clothing categories, which can lead to failed outfit changes and inability to grasp complex and diverse clothing categories. Fitting training on the diffusion model involves training the model to replace the clothing in the second training image with the training clothing. This addresses issues with existing virtual fitting models, such as texture loss and unnatural clothing fit on the model, and improves the quality of outfit changes.
[0025] According to the embodiment of the present application, a technical solution is provided, which obtains a first training image, determines a first mask segmentation map of clothing covering an object in the first training image and the category of clothing of the object in the first training image; uses the first mask segmentation map, the first training image and the category of clothing of the object in the first training image to pre-train a diffusion model to improve the diffusion model's clothing generation ability and understanding of clothing categories; obtains a second training image and a training clothing image, determines a second mask segmentation map of clothing covering an object in the second training image and the category of training clothing in the training clothing image; uses the second mask segmentation map, the training clothing image and the category of training clothing in the training clothing image to train the diffusion model for fitting, so that the diffusion model replaces the clothing of the object in the second training image with the training clothing; uses the diffusion model after fitting training as a virtual fitting model, and uses the virtual fitting model to provide a virtual fitting service for the target object. The above technical means can solve the problems of virtual fitting failure and poor fitting effect in the existing technology, thereby reducing the number of fitting failures, improving fitting effect, and improving user satisfaction.
[0026] In an optional embodiment, a target image of a target object and a target clothing image of a target clothing are obtained; a mask segmentation map of the clothing of the target object in the target image and the category of the target clothing are determined; and based on the mask segmentation map, the target clothing image and the category of the target clothing, a virtual fitting model is used to replace the clothing of the target object in the target image with the target clothing.
[0027] Furthermore, determining a first mask segmentation map of clothing that obscures the object in the first training image and the category of clothing for the object in the first training image includes: masking the clothing for the object in the first training image by an image mask segmentation method to obtain a first mask segmentation map; and determining the category of clothing for the object in the first training image by using a clothing classification model, wherein the clothing classification model has been trained to determine the category of clothing for the object in the image and extract category features of clothing for the object in the image.
[0028] Image mask segmentation is an image processing technique that uses masks to divide an image into distinct regions for separate processing. Clothing classification models can employ a transformer architecture. Training the model to determine the category of clothing in an image and extracting the categorical features of clothing can be done using existing training methods.
[0029] Furthermore, the diffusion model is pre-trained using the first mask segmentation map, the first training image, and the category of clothing of the object in the first training image to improve the diffusion model's clothing generation ability and its ability to understand clothing categories, including: using the prompt words composed of the category of clothing of the object in the first training image and the first mask segmentation map as inputs of the diffusion model, and using the first training image as output of the diffusion model, and pre-training the diffusion model to improve the diffusion model's clothing generation ability and its ability to understand clothing categories.
[0030] This can be understood as follows: the diffusion model generates a predicted image based on a cue word consisting of the clothing category of the object in the first training image and the first mask segmentation map. The cross-entropy loss function is used to calculate the loss between the predicted image and the first training image. The diffusion model parameters are then optimized based on this loss, completing pre-training. Using the first training image as the output of the diffusion model effectively sets it as the theoretically optimal output of the diffusion model, essentially using it as the label for the diffusion model's output.
[0031] Furthermore, determining a second mask segmentation map of clothing that obscures the object in the second training image and the category of the training clothing in the training clothing image includes: masking the clothing of the object in the second training image by an image mask segmentation method to obtain a second mask segmentation map; and determining the category of the clothing of the object in the second training image by using a clothing classification model, wherein the clothing classification model has been trained to determine the category of the clothing of the object in the image and extract category features of the clothing of the object in the image.
[0032] Determining the second mask segmentation map of clothing that obscures the object in the second training image and the category of training clothing in the training clothing image is similar to determining the first mask segmentation map of clothing that obscures the object in the first training image and the category of clothing for the object in the first training image, and will not be repeated here.
[0033] Furthermore, the diffusion model is trained for fitting using the second mask segmentation map, the training clothing image, and the category of the training clothing in the training clothing image, so that the diffusion model replaces the clothing of the object in the second training image with the training clothing, including: adding noise to the second mask segmentation map through a diffusion process of the diffusion model based on prompt words composed of the category of the training clothing in the training clothing image and the training clothing image, and predicting and removing the noise added to the second mask segmentation map during the diffusion process through an inverse diffusion process of the diffusion model, so that the diffusion model replaces the clothing of the object in the second training image with the training clothing.
[0034] The diffusion model includes a diffusion process and an inverse diffusion process. The diffusion process adds noise to the second mask segmentation map, and the inverse diffusion process predicts and removes the noise added to the second mask segmentation map, so that the inverse diffusion process ultimately outputs an image in which the clothing of the subject in the second training image is replaced with the training clothing. The addition of noise to the second mask segmentation map during the diffusion process strengthens or does not weaken the cue words and training clothing images that are composed of the categories of the training clothing in the training clothing image (the clothing in the second mask segmentation map formed after the noise addition is the category described by the cue words, and the closer the clothing in the second mask segmentation map after the noise addition is to the training clothing image, the better). The removal of noise during the inverse diffusion process strengthens or does not weaken the cue words and training clothing images that are composed of the categories of the training clothing in the training clothing image.
[0035] The diffusion model used in this application can be any common diffusion model, such as stable diffusion.
[0036] Furthermore, the diffusion model is trained for fitting using the second mask segmentation map, the training clothing image, and the category of the training clothing in the training clothing image, so that the diffusion model replaces the clothing of the object in the second training image with the training clothing, including: using the clothing classification model to extract the category features of the clothing of the object in the training clothing image, wherein the clothing classification model has been trained to determine the category of the clothing of the object in the image and extract the category features of the clothing of the object in the image; extracting the text features of the prompt words composed of the category of the training clothing in the training clothing image through the text encoder, wherein the text encoder has been trained to extract the text features of the text serving as the prompt words; the diffusion model generates an image in which the clothing of the object in the second training image is replaced with the training clothing based on the second mask segmentation map, the training clothing image, the category features, and the text features, including: performing feature splicing on the training clothing image and the second mask segmentation map in the encoding network of the diffusion model, and aligning the category features and text features as conditions with the image information processed by the previous network layer of the Cross-Attention in the encoding network through the Cross-Attention in the encoding network.
[0037] The text encoder can use a conventional neural network model. For example, "CLIP" refers to "Contrastive Language-Image Pre-training". Cross-attention is a multi-head attention mechanism that is used to establish associations between different input sequences. This mechanism plays an important role in diffusion models, especially when processing multimodal data (such as text and images), it helps to achieve data alignment. The basic idea of cross-attention is that each element in one input sequence will pay attention to all elements in the other input sequence and calculate its representation based on the attention degree. In this way, the two sequences can transfer information to each other, thereby achieving better representation. For example, in image generation tasks, cross-attention can align text information (as a condition) with image data to generate images that match the text description.
[0038] The stable diffusion model consists of an encoding network, an intermediate network, and a decoding network, all of which can be UNet architectures. The encoding network can be viewed as multiple convolutional layers, concatenation layers, and cross-attention layers. Within the diffusion model, the multiple convolutional layers in the encoding network encode the training clothing image and the second mask segmentation map. The concatenation layer concatenates the encoded results of the training clothing image and the second mask segmentation map. Cross-attention uses category features and text features as conditions to align the cross-attention concatenation results in the encoding network. The intermediate network and decoding network sequentially process the alignment results to produce an image in which the clothing in the second training image is replaced with the training clothing.
[0039] Figure 2 FIG. 1 is a flow chart of a fitting training method provided by an embodiment of the present disclosure. Figure 2 As shown, the method includes:
[0040] S201, calculating the noise loss between the noise added by the diffusion process of the diffusion model and the noise predicted by the inverse diffusion process;
[0041] S202, calculating the generation loss between the diffusion model using the training clothing to replace the clothing of the object in the second training image and the corresponding labels of the second training image and the training clothing image;
[0042] S203: Optimize the model parameters of the diffusion model according to the noise loss and / or the generation loss to complete the fitting training.
[0043] Both the noise loss and the generation loss can be calculated using the mean square error loss function. The corresponding labels of the second training image and the training clothing image are images obtained in advance by replacing the clothing of the object in the second training image with the training clothing.
[0044] All of the above optional technical solutions can be combined in any way to form optional embodiments of the present application, and will not be described in detail here.
[0045] The following are embodiments of the apparatus disclosed herein, which can be used to implement the method embodiments disclosed herein. For details not disclosed in the apparatus embodiments disclosed herein, please refer to the method embodiments disclosed herein.
[0046] Figure 3 FIG. 1 is a schematic diagram of a device for a virtual fitting of an object provided by an embodiment of the present disclosure. Figure 3 As shown, the device for the subject to virtually try on clothes includes:
[0047] A first acquisition module 301 is configured to acquire a first training image, determine a first mask segmentation map of clothing covering an object in the first training image, and determine a category of clothing in the first training image;
[0048] A pre-training module 302 is configured to pre-train the diffusion model using the first mask segmentation map, the first training image, and the category of clothing of the object in the first training image to improve the diffusion model's clothing generation capability and its ability to understand clothing categories;
[0049] A second acquisition module 303 is configured to acquire a second training image and a training clothing image, and determine a second mask segmentation map of clothing covering an object in the second training image and a category of training clothing in the training clothing image;
[0050] a fitting training module 304 configured to perform fitting training on the diffusion model using the second mask segmentation map, the training clothing image, and the category of the training clothing in the training clothing image, so that the diffusion model replaces the clothing of the subject in the second training image with the training clothing;
[0051] The virtual fitting module 305 is configured to use the diffusion model trained for fitting as a virtual fitting model and provide a virtual fitting service for the target object using the virtual fitting model.
[0052] According to the embodiment of the present application, a technical solution is provided, which obtains a first training image, determines a first mask segmentation map of clothing covering an object in the first training image and the category of clothing of the object in the first training image; uses the first mask segmentation map, the first training image and the category of clothing of the object in the first training image to pre-train a diffusion model to improve the diffusion model's clothing generation ability and understanding of clothing categories; obtains a second training image and a training clothing image, determines a second mask segmentation map of clothing covering an object in the second training image and the category of training clothing in the training clothing image; uses the second mask segmentation map, the training clothing image and the category of training clothing in the training clothing image to train the diffusion model for fitting, so that the diffusion model replaces the clothing of the object in the second training image with the training clothing; uses the diffusion model after fitting training as a virtual fitting model, and uses the virtual fitting model to provide a virtual fitting service for the target object. The above technical means can solve the problems of virtual fitting failure and poor fitting effect in the existing technology, thereby reducing the number of fitting failures, improving fitting effect, and improving user satisfaction.
[0053] In an optional embodiment, the virtual fitting module 305 is further configured to obtain a target image of the target object and a target clothing image of the target clothing; determine a mask segmentation map of the clothing of the target object in the target image and the category of the target clothing; and replace the clothing of the target object in the target image with the target clothing using the virtual fitting model based on the mask segmentation map, the target clothing image and the category of the target clothing.
[0054] In an optional embodiment, the first acquisition module 301 is further configured to mask the clothing of the object in the first training image through an image mask segmentation method to obtain a first mask segmentation map; and use a clothing classification model to determine the category of the clothing of the object in the first training image, wherein the clothing classification model has been trained and can determine the category of the clothing of the object in the image and extract the category features of the clothing of the object in the image.
[0055] In an optional embodiment, the pre-training module 302 is further configured to use the prompt words consisting of the category of clothing of the object in the first training image and the first mask segmentation map as inputs of the diffusion model, and the first training image as output of the diffusion model, to pre-train the diffusion model to improve the diffusion model's clothing generation capability and its ability to understand clothing categories.
[0056] In an optional embodiment, the second acquisition module 303 is further configured to mask the clothing of the object in the second training image through an image mask segmentation method to obtain a second mask segmentation map; and use a clothing classification model to determine the category of the clothing of the object in the second training image, wherein the clothing classification model has been trained and can determine the category of the clothing of the object in the image and extract the category features of the clothing of the object in the image.
[0057] In an optional embodiment, the fitting training module 304 is further configured to add noise to the second mask segmentation map through a diffusion process of a diffusion model based on prompt words consisting of categories of training clothing in the training clothing image and the training clothing image, and to predict and remove the noise added to the second mask segmentation map during the diffusion process through an inverse diffusion process of the diffusion model, so that the diffusion model replaces the clothing of the object in the second training image with the training clothing.
[0058] In an optional embodiment, the fitting training module 304 is further configured to extract category features of clothing of the object in the training clothing image by using a clothing classification model, wherein the clothing classification model has been trained to determine the category of clothing of the object in the image and extract category features of clothing of the object in the image; extract text features of prompt words composed of the category of training clothing in the training clothing image by using a text encoder, wherein the text encoder has been trained to extract text features of the text as prompt words; the diffusion model generates an image in which the clothing of the object in the second training image is replaced with the training clothing based on the second mask segmentation map, the training clothing image, the category features and the text features, including: feature splicing of the training clothing image and the second mask segmentation map in the encoding network of the diffusion model, and aligning the category features and text features as conditions with the image information processed by the previous network layer of Cross-Attention in the encoding network through Cross-Attention in the encoding network.
[0059] In an optional embodiment, the fitting training module 304 is further configured to calculate the noise loss between the noise added by the diffusion process of the diffusion model and the noise predicted by the inverse diffusion process; calculate the generation loss between the image of the clothing of the object in the second training image replaced by the training clothing by the diffusion model and the corresponding labels of the second training image and the training clothing image; and optimize the model parameters of the diffusion model based on the noise loss and / or generation loss to complete the fitting training.
[0060] It should be understood that the order of execution of each step in the above embodiment does not mean the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiment of the present disclosure.
[0061] Figure 4 FIG. 4 is a schematic diagram of an electronic device 4 provided in an embodiment of the present disclosure. Figure 4 As shown, electronic device 4 in this embodiment includes: a processor 401, a memory 402, and a computer program 403 stored in memory 402 and executable by processor 401. When processor 401 executes computer program 403, steps in the aforementioned method embodiments are implemented. Alternatively, when processor 401 executes computer program 403, the functions of the modules / units in the aforementioned device embodiments are implemented.
[0062] The electronic device 4 may be a desktop computer, a notebook, a PDA, a cloud server, or other electronic device. The electronic device 4 may include but is not limited to a processor 401 and a memory 402. Those skilled in the art will appreciate that Figure 4 The electronic device 4 is merely an example and does not limit the electronic device 4 , and may include more or fewer components than shown in the figure, or different components.
[0063] The processor 401 may be a central processing unit (CPU), or other general-purpose processors, digital signal processors (DSP), application-specific integrated circuits (ASIC), field-programmable gate arrays (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc.
[0064] Memory 402 can be an internal storage unit of electronic device 4, such as a hard drive or memory of electronic device 4. Memory 402 can also be an external storage device of electronic device 4, such as a plug-in hard drive, a Smart Media Card (SMC), a Secure Digital (SD) card, a flash memory card, etc. Memory 402 can also include both internal storage units of electronic device 4 and external storage devices. Memory 402 is used to store computer programs and other programs and data required by the electronic device.
[0065] Those skilled in the art will clearly understand that for the sake of convenience and brevity of description, only the above-mentioned functional units and module divisions are used as examples for illustration. In actual applications, the above-mentioned functions can be distributed to different functional units and modules as needed, that is, the internal structure of the device can be divided into different functional units or modules to complete all or part of the functions described above. In the embodiments, each functional unit and module can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The above-mentioned integrated unit can be implemented in the form of hardware or in the form of software functional units.
[0066] If the integrated module / unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the present disclosure implements all or part of the processes in the above-mentioned embodiment method, and can also be completed by instructing the relevant hardware through a computer program. The computer program can be stored in a computer-readable storage medium, and when the computer program is executed by the processor, it can implement the steps of the above-mentioned various method embodiments. The computer program may include computer program code, which may be in source code form, object code form, executable file or some intermediate form. The computer-readable medium may include: any entity or device capable of carrying computer program code, recording medium, USB flash drive, mobile hard disk, magnetic disk, optical disk, computer memory, read-only memory (ROM), random access memory (RAM), electric carrier signal, telecommunication signal and software distribution medium, etc. It should be noted that the content contained in the computer-readable medium can be appropriately increased or decreased according to the legislation and patent practice requirements in the jurisdiction. For example, in some jurisdictions, according to legislation and patent practice, computer-readable media do not include electric carrier signals and telecommunication signals.
[0067] The above embodiments are only used to illustrate the technical solutions of the present disclosure, rather than to limit them. Although the present disclosure has been described in detail with reference to the aforementioned embodiments, it should be understood by those skilled in the art that the technical solutions described in the aforementioned embodiments may still be modified, or some of the technical features thereof may be replaced by equivalents. However, these modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the embodiments of the present disclosure, and should all be included in the protection scope of the present disclosure.
Claims
1. A method for a subject to virtually try on clothes, characterized in that: include: Acquire a first training image, and determine a first mask segmentation map of clothing covering an object in the first training image and a category of the clothing of the object in the first training image; Pre-training a diffusion model using the first mask segmentation map, the first training image, and the category of clothing of the object in the first training image to improve the diffusion model's clothing generation capability and its ability to understand clothing categories; Acquire a second training image and a training clothing image, and determine a second mask segmentation map of clothing that obscures an object in the second training image and a category of training clothing in the training clothing image; Performing fitting training on the diffusion model using the second mask segmentation map, the training clothing image, and the category of the training clothing in the training clothing image, so that the diffusion model replaces clothing of the subject in the second training image with the training clothing; Using the diffusion model trained after fitting as a virtual fitting model, and using the virtual fitting model to provide a virtual fitting service for a target object; Performing clothing fitting training on the diffusion model using the second mask segmentation map, the training clothing image, and the category of the training clothing in the training clothing image, so that the diffusion model replaces clothing of the subject in the second training image with the training clothing, comprising: Extracting category features of clothing of the subject in the training clothing image using a clothing classification model, wherein the clothing classification model has been trained to determine the category of clothing of the subject in the image and to extract the category features of clothing of the subject in the image; extracting text features of a prompt word consisting of the category of training clothing in the training clothing image by a text encoder, wherein the text encoder has been trained to extract text features of the text serving as the prompt word; The diffusion model generates an image in which the clothing of the object in the second training image is replaced with the training clothing based on the second mask segmentation map, the training clothing image, the category features, and the text features. The method includes: performing feature splicing on the training clothing image and the second mask segmentation map in the encoding network of the diffusion model, and aligning the category features and the text features as conditions with the image information processed by the previous network layer of the Cross-Attention in the encoding network through the Cross-Attention in the encoding network.
2. The method according to claim 1, characterized in that Determining a first mask segmentation map of clothing that obscures the object in the first training image and a category of the clothing of the object in the first training image includes: masking the clothing of the object in the first training image using an image mask segmentation method to obtain the first mask segmentation image; The clothing classification model is used to determine the category of clothing of the object in the first training image, wherein the clothing classification model has been trained to determine the category of clothing of the object in the image and extract category features of the clothing of the object in the image.
3. The method according to claim 1, characterized in that Pre-training a diffusion model using the first mask segmentation map, the first training image, and the category of clothing of an object in the first training image to improve the diffusion model's clothing generation capability and clothing category understanding capability, including: The diffusion model is pre-trained using a prompt word consisting of the category of clothing of the object in the first training image and the first mask segmentation map as input, and the first training image as output of the diffusion model to improve the diffusion model's clothing generation capability and its ability to understand clothing categories.
4. The method according to claim 1, characterized in that Determining a second mask segmentation map of clothing that obscures the object in the second training image and a category of training clothing in the training clothing image includes: masking the clothing of the object in the second training image using an image mask segmentation method to obtain a second mask segmentation image; The clothing classification model is used to determine the category of clothing of the object in the second training image, wherein the clothing classification model has been trained to determine the category of clothing of the object in the image and extract category features of the clothing of the object in the image.
5. The method according to claim 1, characterized in that: Performing clothing fitting training on the diffusion model using the second mask segmentation map, the training clothing image, and the category of the training clothing in the training clothing image, so that the diffusion model replaces clothing of the subject in the second training image with the training clothing, comprising: Based on the prompt words consisting of the categories of training clothing in the training clothing image and the training clothing image, noise is added to the second mask segmentation map through a diffusion process of the diffusion model, and the noise added to the second mask segmentation map during the diffusion process is predicted and removed through an inverse diffusion process of the diffusion model, so that the diffusion model replaces the clothing of the object in the second training image with the training clothing.
6. The method according to claim 1, characterized in that Performing clothing fitting training on the diffusion model using the second mask segmentation map, the training clothing image, and the category of the training clothing in the training clothing image, so that the diffusion model replaces clothing of the subject in the second training image with the training clothing, comprising: calculating the noise loss between the noise added by the diffusion process of the diffusion model and the noise predicted by the inverse diffusion process; Calculating a generation loss between an image in which the diffusion model replaces clothing of an object in the second training image with the training clothing and corresponding labels of the second training image and the training clothing image; The model parameters of the diffusion model are optimized according to the noise loss and / or the generation loss to complete the fitting training.
7. A device for a subject to virtually try on clothes, characterized in that: include: a first acquisition module configured to acquire a first training image, determine a first mask segmentation map of clothing covering an object in the first training image, and determine a category of clothing of the object in the first training image; a pre-training module configured to pre-train a diffusion model using the first mask segmentation map, the first training image, and the category of clothing of the object in the first training image, so as to improve the clothing generation capability and clothing category understanding capability of the diffusion model; a second acquisition module configured to acquire a second training image and a training clothing image, and determine a second mask segmentation map of clothing covering an object in the second training image and a category of training clothing in the training clothing image; a fitting training module configured to perform fitting training on the diffusion model using the second mask segmentation map, the training clothing image, and the category of the training clothing in the training clothing image, so that the diffusion model replaces the clothing of the object in the second training image with the training clothing; a virtual fitting module configured to use the diffusion model trained for fitting as a virtual fitting model and provide a virtual fitting service for a target object using the virtual fitting model; Performing clothing fitting training on the diffusion model using the second mask segmentation map, the training clothing image, and the category of the training clothing in the training clothing image, so that the diffusion model replaces clothing of the subject in the second training image with the training clothing, comprising: Extracting category features of clothing of the subject in the training clothing image using a clothing classification model, wherein the clothing classification model has been trained to determine the category of clothing of the subject in the image and to extract the category features of clothing of the subject in the image; extracting text features of a prompt word consisting of the category of training clothing in the training clothing image by a text encoder, wherein the text encoder has been trained to extract text features of the text serving as the prompt word; The diffusion model generates an image in which the clothing of the object in the second training image is replaced with the training clothing based on the second mask segmentation map, the training clothing image, the category features, and the text features. The method includes: performing feature splicing on the training clothing image and the second mask segmentation map in the encoding network of the diffusion model, and aligning the category features and the text features as conditions with the image information processed by the previous network layer of the Cross-Attention in the encoding network through the Cross-Attention in the encoding network.
8. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein: When the processor executes the computer program, the steps of the method according to any one of claims 1 to 6 are implemented.
9. A computer-readable storage medium storing a computer program, characterized in that: When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 6 are implemented.
Citation Information
Patent Citations
Model training method and device, electronic equipment and storage medium
CN114972919A
Virtual fitting method and device and storage medium
CN117745990A