Virtual dress-up methods, devices, computer equipment, and storage media

CN116342981BActive Publication Date: 2026-08-14XIAMEN MEITUZHIJIA TECH
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-04-04
Publication Date
2026-08-14

AI Technical Summary

Technical Problem

但是,配对模特图片以及模特试穿图片需要保证同一模特同一姿势,可见对于虚拟换装任务来说,训练数据收集的代价是很昂贵的

Benefits of technology

[0046] The aforementioned virtual clothing-changing method, apparatus, computer equipment, storage medium, and computer program product extract first object data from a first sample image; and extract second clothing data and second object data from a second sample image. The first and second sample images do not need to guarantee the same object in the same pose; they only need to show objects wearing the target clothing in different poses, significantly reducing the difficulty and cost of data collection. The first object data and second clothing data are input into a clothing-changing model to be trained, predicting a first clothing-changing image. The first clothing-changing image shows the effect of an object in a first pose wearing the target clothing. Then, the clothing data of the target clothing and the second object data from the first clothing-changing image are input into the clothing-changing model to be trained, predicting a second clothing-changing image. The target clothing in each training iteration of the first clothing-changing image will be different. Compared to directly using the clothing data of the target clothing from the first sample image as input, this enhances the generalization of virtual clothing-changing learning, and using the first clothing-changing image for back-supervision to generate the second clothing-changing image also improves data utilization. The second clothing-changing image shows the effect of an object in a second pose wearing the target clothing. The model training not only considers the difference between the first clothing change image and the first sample image, but also the difference between the second clothing change image and the second sample image, resulting in a trained clothing change model. The clothing change model can be accurately trained using the easier-to-collect first and second sample images, thus reducing costs.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116342981B_ABST
    Figure CN116342981B_ABST
Patent Text Reader

Abstract

This application relates to a virtual clothing-changing method, apparatus, computer device, and storage medium. The method includes: extracting first object data from a first sample image including an object in a first pose wearing target clothing; extracting second clothing data and second object data from a second sample image including an object in a second pose wearing target clothing; inputting the first object data and second clothing data into a clothing-changing model to be trained to obtain a first clothing-changing image; inputting the clothing data of the target clothing and the second object data from the first clothing-changing image into the clothing-changing model to be trained to obtain a second clothing-changing image; and training the model based on the differences between the first clothing-changing image and the first sample image, and the differences between the second clothing-changing image and the second sample image, to obtain a trained clothing-changing model. This method reduces the cost of virtual clothing changing.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of artificial intelligence technology, and in particular to a virtual dress-up method, apparatus, computer equipment, and storage medium. Background Technology

[0002] With the development of artificial intelligence technology, virtual dress-up technology has emerged. In many innovative applications, especially in e-commerce, virtual dress-up tasks allow buyers to upload photos for matching, enabling virtual try-ons. On the other hand, for sellers, finding professional models to shoot clothing is time-consuming, labor-intensive, and inefficient. Therefore, the demand for virtual dress-up is becoming increasingly strong.

[0003] Traditional techniques require collecting paired clothing images, model images, and model try-on images as training data for virtual dress-up tasks. However, the paired model images and model try-on images need to be the same model in the same pose, making the collection of training data very expensive for virtual dress-up tasks. Summary of the Invention

[0004] Therefore, it is necessary to provide a virtual dress-up method, apparatus, computer equipment, computer-readable storage medium, and computer program product that can reduce costs in response to the above-mentioned technical problems.

[0005] Firstly, this application provides a virtual clothing changing method. The method includes:

[0006] Extract first object data from a first sample image that includes an object in a first pose and wearing the target clothing;

[0007] Extract second clothing data and second object data from a second sample image including an object in a second pose and wearing the target clothing;

[0008] The first object data and the second clothing data are input into the clothing-changing model to be trained to predict the first clothing-changing image; the first clothing-changing image shows the effect of the object in the first pose wearing the target clothing;

[0009] The clothing data of the target garment in the first clothing image and the data of the second object are input into the clothing model to be trained to predict the second clothing image; the second clothing image shows the effect of the object in the second pose wearing the target garment;

[0010] The model is trained based on the differences between the first clothing change image and the first sample image, and the differences between the second clothing change image and the second sample image, to obtain the trained clothing change model.

[0011] In some embodiments, the processing steps of inputting object data and clothing data for predicting a target clothing change image into a clothing change model to be trained to predict the target clothing change image include:

[0012] The object data and clothing data input to the clothing-changing model to be trained are subjected to multi-resolution encoding processing to obtain object-coded data and clothing-coded data at each resolution level;

[0013] The object coding data and clothing coding data at the same resolution level are fused to obtain the fused coding data at each resolution level;

[0014] Based on the fused coded data and object coded data at each resolution level, the clothing change prediction process is performed to predict the target clothing change image;

[0015] The target clothing change image includes either the first clothing change image or the second clothing change image.

[0016] In some embodiments, the method further includes:

[0017] The object coding data and clothing coding data at the first resolution level are further encoded to obtain further encoded object coding data and further encoded clothing coding data; the resolution of the further encoded object coding data and further encoded clothing coding data is lower than that of the object coding data and clothing coding data at the first resolution level.

[0018] The object coding data and clothing coding data of the advanced coding are initially fused to obtain preliminary fused coding data;

[0019] The process of fusing object coding data and clothing coding data at the same resolution level to obtain fused coding data at each resolution level includes:

[0020] The preliminary fusion coding data, as well as the object coding data and clothing coding data of the first resolution level, are determined as the input of the fusion module corresponding to the first resolution level, so as to obtain the fusion coding data of the first resolution level output by the fusion module corresponding to the first resolution level.

[0021] Starting from the next resolution level after the first resolution level, the current resolution level is determined sequentially. The fusion encoding data of the previous resolution level of the current resolution level, the object encoding data and clothing encoding data of the current resolution level are determined as the input of the fusion module corresponding to the current resolution level, and the fusion encoding data output by the fusion module corresponding to the current resolution level is obtained.

[0022] In some embodiments, the method further includes:

[0023] Upsample the object encoding data of the advanced encoding to obtain the upsampled encoding data;

[0024] The step of determining the preliminary fused encoding data, as well as the object encoding data and clothing encoding data of the first resolution level, as the input to the fusion module corresponding to the first resolution level, and obtaining the fused encoding data of the first resolution level output by the fusion module corresponding to the first resolution level, includes:

[0025] The preliminary fused encoding data, the upsampled encoding data, and the object encoding data and clothing encoding data of the first resolution level are determined as the input of the fusion module corresponding to the first resolution level, so as to obtain the fused encoding data of the first resolution level output by the fusion module corresponding to the first resolution level.

[0026] In some embodiments, the target clothing-changing image is the clothing-changing result at the last resolution level; the clothing-changing prediction processing based on the fused coded data and object coded data at each resolution level to predict the target clothing-changing image includes:

[0027] Based on the fused encoded data and object encoded data of the first resolution level, clothing change prediction processing is performed to predict the clothing change result of the first resolution level.

[0028] Starting from the next resolution level of the first resolution level, the current resolution level is determined sequentially. Based on the clothing change result of the previous resolution level of the current resolution level, as well as the fused encoding data and object encoding data of the current resolution level, clothing change prediction processing is performed to predict the clothing change result of the current resolution level. The clothing change result of each resolution level corresponds to a different resolution.

[0029] In some embodiments, the method further includes:

[0030] Extract the first garment data from the first sample image;

[0031] The first clothing prediction data and the second clothing prediction data output by the clothing-changing model to be trained are determined; the first clothing prediction data is used to characterize the target clothing in the first clothing-changing image; the second clothing prediction data is used to characterize the target clothing in the second clothing-changing image.

[0032] The step of training a model based on the differences between the first clothing-changing image and the first sample image, and the differences between the second clothing-changing image and the second sample image, to obtain a trained clothing-changing model includes:

[0033] The model is trained based on the first difference, the second difference, the third difference, and the fourth difference to obtain the trained clothing-changing model; the first difference refers to the difference between the first clothing-changing image and the first sample image; the second difference refers to the difference between the second clothing-changing image and the second sample image; the third difference refers to the difference between the first clothing data and the first clothing prediction data; and the fourth difference refers to the difference between the second clothing data and the second clothing prediction data.

[0034] In some embodiments, the first object data includes a first object mask; the second object data includes a second object mask; the method further includes:

[0035] Determine the first object prediction mask and the second object prediction mask output by the clothing-changing model to be trained; the first object prediction mask is the mask of the object in the first clothing-changing image; the second object prediction mask is the mask of the object in the second clothing-changing image;

[0036] The process of training the model based on the first difference, second difference, third difference, and fourth difference to obtain the trained clothing-changing model includes:

[0037] The model is trained based on the first difference, second difference, third difference, fourth difference, fifth difference, and sixth difference to obtain the trained clothing-changing model; the fifth difference refers to the difference between the first object mask and the first object prediction mask; the sixth difference refers to the difference between the second object mask and the second object prediction mask.

[0038] Secondly, this application also provides a virtual clothing changing device. The device includes:

[0039] An extraction unit is configured to extract first object data from a first sample image including an object in a first pose and wearing the target clothing; and to extract second clothing data and second object data from a second sample image including an object in a second pose and wearing the target clothing.

[0040] The first training unit is used to input the first object data and the second clothing data into the clothing-changing model to be trained, and predict the first clothing-changing image; the first clothing-changing image shows the effect of the object in the first pose wearing the target clothing;

[0041] The second training unit is used to input the clothing data of the target garment in the first clothing image and the data of the second object into the clothing model to be trained, and predict the second clothing image; the second clothing image shows the effect of the object in the second pose wearing the target garment;

[0042] The optimization unit is used to train the model based on the difference between the first clothing change image and the first sample image, and the difference between the second clothing change image and the second sample image, to obtain the trained clothing change model.

[0043] Thirdly, this application also provides a computer device. The computer device includes a memory and a processor, the memory storing a computer program, and the processor executing the computer program to implement the steps of the method described above.

[0044] Fourthly, this application also provides a computer-readable storage medium. The computer-readable storage medium stores a computer program thereon, which, when executed by a processor, implements the steps of the method described above.

[0045] Fifthly, this application also provides a computer program product. The computer program product includes a computer program that, when executed by a processor, implements the steps of the method described above.

[0046] The aforementioned virtual clothing-changing method, apparatus, computer equipment, storage medium, and computer program product extract first object data from a first sample image; and extract second clothing data and second object data from a second sample image. The first and second sample images do not need to guarantee the same object in the same pose; they only need to show objects wearing the target clothing in different poses, significantly reducing the difficulty and cost of data collection. The first object data and second clothing data are input into a clothing-changing model to be trained, predicting a first clothing-changing image. The first clothing-changing image shows the effect of an object in a first pose wearing the target clothing. Then, the clothing data of the target clothing and the second object data from the first clothing-changing image are input into the clothing-changing model to be trained, predicting a second clothing-changing image. The target clothing in each training iteration of the first clothing-changing image will be different. Compared to directly using the clothing data of the target clothing from the first sample image as input, this enhances the generalization of virtual clothing-changing learning, and using the first clothing-changing image for back-supervision to generate the second clothing-changing image also improves data utilization. The second clothing-changing image shows the effect of an object in a second pose wearing the target clothing. The model training not only considers the difference between the first clothing change image and the first sample image, but also the difference between the second clothing change image and the second sample image, resulting in a trained clothing change model. The clothing change model can be accurately trained using the easier-to-collect first and second sample images, thus reducing costs. Attached Figure Description

[0047] Figure 1 This is a flowchart illustrating a virtual clothing changing method in one embodiment;

[0048] Figure 2This is a schematic diagram of a first sample image, first object data, and first clothing data in one embodiment;

[0049] Figure 3 This is a schematic diagram illustrating the extraction of second object data and second clothing data from a second sample image in one embodiment.

[0050] Figure 4 This is a schematic diagram of the internal structure of the fusion module in one embodiment;

[0051] Figure 5 This is a schematic diagram of the internal structure of the prediction module in one embodiment;

[0052] Figure 6 This is a structural schematic diagram of a clothing-changing model in one embodiment;

[0053] Figure 7 This is a structural block diagram of a virtual changing device in one embodiment;

[0054] Figure 8 This is an internal structural diagram of a computer device in one embodiment;

[0055] Figure 9 This is a diagram of the internal structure of a computer device in another embodiment. Detailed Implementation

[0056] To make the objectives, technical solutions, and advantages of this application clearer, the following detailed description is provided in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative and not intended to limit the scope of this application.

[0057] In some embodiments, such as Figure 1 As shown, a virtual clothing changing method is provided. Taking the application of this method to a computer device as an example, it can be understood that the computer device includes at least one of a terminal or a server. The method includes the following steps:

[0058] S102, extract first object data from a first sample image including an object in a first pose and wearing target clothing.

[0059] The first object data refers to the object-related data in the first sample image. This first object data reflects the pose of the object and the wearing of its clothing in the first sample image.

[0060] For example, a computer device can extract first object data from an object in a first sample image. The computer device can further extract first clothing data based on the first object data. It is understood that the first object data reflects the situation of the object wearing the target clothing in the first sample image. The first clothing data can independently reflect the situation of the target clothing in the first sample image.

[0061] In some embodiments, such as Figure 2 As shown, a first sample image, first object data, and first clothing data are provided. The first object data may include first skeleton data, a first component segmentation mask, a first dense coordinate mapping map, a first matting, and a first matting mask. The first skeleton data represents the skeleton of the object in the first sample image. The first component segmentation mask reflects the various components of the object in the first sample image. The first dense coordinate mapping map reflects the pose of the object in the first sample image, specifically, it may be the texture coordinates of a human body. The first matting is an image from the first sample image with the clothing, torso, skin, and arm areas of the object removed. The first matting mask is the mask for the first matting. The first clothing data may include a first clothing image and a first clothing mask. The first clothing image refers to the image of the target clothing in the first sample image. The first clothing mask refers to the mask corresponding to the first clothing image. It can be understood that a computer device can further determine the first matting and the first matting mask, as well as the first clothing image and the first clothing mask, based on the first component segmentation mask and the first skeleton data. It should be noted that in this embodiment, the blurring of the eyes of the human object in the image is for privacy protection and is not part of the processing described in this embodiment.

[0062] In some embodiments, the objects in the first sample image and the second sample image may be models.

[0063] In some embodiments, the terminal may be, but is not limited to, various personal computers, laptops, smartphones, tablets, IoT devices, and portable wearable devices. IoT devices may include smart speakers, smart TVs, smart air conditioners, smart in-vehicle devices, etc. Portable wearable devices may include smartwatches, smart bracelets, head-mounted devices, etc. The server may be implemented using a standalone server or a server cluster consisting of multiple servers.

[0064] S104, extract second clothing data and second object data from a second sample image including an object in a second pose and wearing target clothing.

[0065] The second object data refers to data related to the objects in the second sample image. This second object data reflects the pose of the objects and the wearing of the target clothing in the second sample image. The second clothing data refers to data related to the clothing in the second sample image. This second clothing data reflects the shape of the target clothing in the second sample image.

[0066] For example, the computer device can extract second object data from objects in the second sample image. The computer device can further extract second clothing data from the second object data.

[0067] In some embodiments, such as Figure 3 As shown, a second sample image, second object data, and second clothing data are provided. The second object data may include second skeleton data, a second part segmentation mask, a second dense coordinate map, a second matting, and a second matting mask. The second skeleton data represents the skeleton of the object in the second sample image. The second part segmentation mask reflects the individual parts of the object in the second sample image. The second dense coordinate map reflects the pose of the object in the second sample image, specifically, the texture coordinates of a human body. The second matting is an image from the second sample image with the clothing, torso, skin, and arm areas of the object removed. The second matting mask is the mask for the second matting. The second clothing data may include a second clothing image and a second clothing mask. The second clothing image refers to the image of the target clothing in the second sample image. The second clothing mask is the mask corresponding to the second clothing image. It can be understood that a computer device can further determine the second matting and the second matting mask, as well as the second clothing image and the second clothing mask, based on the second part segmentation mask and the second skeleton data.

[0068] In some embodiments, the first sample image and the second sample image include the same object. The first sample image and the second sample image can be two images of the same object wearing target clothing in different poses.

[0069] In some embodiments, the computer device may determine multiple sets of sample images, each set including a first sample image and a second sample image. These multiple sets of sample images can serve as training sample images for a clothing-changing model to be trained.

[0070] In some embodiments, the first sample image and the second sample image include different objects. The first sample image and the second sample image can be two images acquired from different objects wearing target clothing in different poses.

[0071] In some embodiments, a computer device can extract object data and clothing data from a sample image using various preset data processing algorithms and models. The sample image is either a first sample image or a second sample image. The object data is either first object data or second object data. The clothing data is either first clothing data or second clothing data.

[0072] S106, input the first object data and the second clothing data into the clothing change model to be trained, and predict the first clothing change image.

[0073] The first clothing-changing image shows the effect of an object in a first pose wearing the target clothing.

[0074] For example, the first object data may include a first matting, a first matting mask, first skeleton data, and a first dense coordinate mapping map. The second clothing data may include a second clothing image and a second clothing mask. A computer device can encode the first object data using a training-to-be-trained clothing-changing model to obtain first object-coded data. The computer device can also encode the second clothing data using the training-to-be-trained clothing-changing model to obtain first clothing-coded data. Furthermore, the computer device can fuse the first object-coded data and the first clothing-coded data using the training-to-be-trained clothing-changing model to obtain first fused encoded data. The computer device can then perform clothing-changing prediction processing based on the first fused encoded data to obtain a first clothing-changing image. It is understood that the fused encoded data...

[0075] In some embodiments, the method further includes pre-training a clothing-changing model. The computer device can use the first object data and the second clothing data as a set of pre-training data. Alternatively, the computer device can use the second object data and the first clothing data as a set of pre-training data. The computer device can use each set of pre-training data as input to an initial clothing-changing model, and predict the clothing-changing image corresponding to that set of pre-training data using the initial clothing-changing model. The computer device can determine the sample image to which the object data in each set of pre-training data belongs, and pre-train the initial clothing-changing model in the direction that reduces the difference between the clothing-changing image corresponding to that set of pre-training data and the aforementioned sample image, to obtain a pre-trained clothing-changing model. It can be understood that the pre-trained clothing-changing model can be used as a clothing-changing model to be trained.

[0076] It should be noted that the initial clothing-changing model cannot infer an accurate first clothing-changing image. If the clothing data of the target garment in an inaccurate first clothing-changing image is used for training, it can easily lead to biases in the training process. Therefore, by pre-training the initial clothing-changing model and using the pre-trained model as the model to be trained, the accuracy of the model training can be guaranteed.

[0077] In one embodiment, the computer device can obtain a pre-trained clothing-changing model after the initial clothing-changing model has completed a preset number of pre-training rounds. It can be understood that the difference between the pre-training of the initial clothing-changing model and the training process of the clothing-changing model to be trained is that the training process of the clothing-changing model to be trained uses the clothing data of the target garment in the first clothing-changing image as input data to obtain the second clothing-changing image. One round of training for the clothing-changing model to be trained includes two rounds of gradient forward and backward propagation. In contrast, one round of pre-training for the initial clothing-changing model includes one round of gradient forward and backward propagation.

[0078] S108, input the clothing data of the target garment and the second object data in the first clothing image into the clothing model to be trained, and predict the second clothing image.

[0079] The second clothing-changing image shows the effect of an object in a second pose wearing the target clothing.

[0080] For example, the second object data may include a second matting, a second matting mask, second skeleton data, and a second dense coordinate mapping map. The clothing data of the target garment in the first clothing-changing image is the first clothing prediction data. The first clothing prediction data may include a first clothing prediction image and a first clothing prediction mask. A computer device can encode the second object data using a clothing-changing model to be trained, obtaining second object encoded data. A computer device can encode the first clothing prediction data using a clothing-changing model to be trained, obtaining second clothing encoded data. A computer device can fuse the second object encoded data and the second clothing encoded data using a clothing-changing model to be trained, obtaining second fused encoded data. The computer device can perform clothing-changing prediction processing based on the second fused encoded data to obtain a second clothing-changing image.

[0081] S110, the model is trained based on the differences between the first clothing change image and the first sample image, and the differences between the second clothing change image and the second sample image, to obtain the trained clothing change model.

[0082] For example, the computer device can use the first clothing-changing image and the first sample image as input to the discriminator to obtain the first discriminator loss. The computer device can use the second clothing-changing image and the second sample image as input to the discriminator to obtain the second discriminator loss. The computer device can train the model in the direction of decreasing the first discriminator loss and the second discriminator loss to obtain a trained clothing-changing model.

[0083] In some embodiments, the computer device can calculate a first perceptual loss based on the first clothing-changing image and the first sample image. The computer device can also calculate a second perceptual loss based on the second clothing-changing image and the second sample image. The computer device can then train the model in a direction that reduces both the first and second perceptual losses to obtain a trained clothing-changing model.

[0084] In some embodiments, the computer device can optimize the initial clothing-changing model in each training round based on the difference between the first clothing-changing image and the first sample image, obtaining the clothing-changing model optimized once in this round, and determining the second clothing-changing image output by the clothing-changing model optimized once in this round. Based on the difference between the second clothing-changing image and the second sample image, the clothing-changing model optimized once in this round is optimized a second time, obtaining the clothing-changing model optimized twice in this round. The computer device can use the clothing-changing model optimized twice in this round as the initial clothing-changing model for the next round to continue iterative optimization, thereby obtaining the trained clothing-changing model.

[0085] In some embodiments, in each training round, the clothing-changing model obtains a first output (output1) for the first input (input1), and the computer device can calculate the loss and gradient backpropagation for output1 to perform a first optimization. The clothing-changing model obtains a second output (output2) for the second input (input2), and the computer device can calculate the loss and gradient backpropagation for output2 to complete a second optimization. The first input includes first object data and second clothing data. The second input includes the second object data and the clothing data of the target garment in the first clothing-changing image.

[0086] The aforementioned virtual clothing-changing method, apparatus, computer equipment, storage medium, and computer program product extract first object data from a first sample image; and extract second clothing data and second object data from a second sample image. The first and second sample images do not need to guarantee the same object in the same pose; they only need to show objects wearing the target clothing in different poses, significantly reducing the difficulty and cost of data collection. The first object data and second clothing data are input into a clothing-changing model to be trained, predicting a first clothing-changing image. The first clothing-changing image shows the effect of an object in a first pose wearing the target clothing. Then, the clothing data of the target clothing and the second object data from the first clothing-changing image are input into the clothing-changing model to be trained, predicting a second clothing-changing image. The target clothing in each training iteration of the first clothing-changing image will be different. Compared to directly using the clothing data of the target clothing from the first sample image as input, this enhances the generalization of virtual clothing-changing learning, and using the first clothing-changing image for back-supervision to generate the second clothing-changing image also improves data utilization. The second clothing-changing image shows the effect of an object in a second pose wearing the target clothing. The model training not only considers the difference between the first clothing change image and the first sample image, but also the difference between the second clothing change image and the second sample image, resulting in a trained clothing change model. The clothing change model can be accurately trained using the easier-to-collect first and second sample images, thus reducing costs.

[0087] In some embodiments, the processing steps of inputting object data and clothing data for predicting a target clothing change image into a clothing change model to predict the target clothing change image include: performing multi-resolution encoding processing on the object data and clothing data input into the clothing change model to obtain object-encoded data and clothing-encoded data at each resolution level; performing fusion processing on the object-encoded data and clothing-encoded data at the same resolution level to obtain fused encoding data at each resolution level; and performing clothing change prediction processing based on the fused encoding data and object-encoded data at each resolution level to predict the target clothing change image; wherein the target clothing change image includes either a first clothing change image or a second clothing change image.

[0088] For example, the clothing-changing model includes an object encoding module and a clothing encoding module for each resolution level. The object encoding data for each resolution level is the output of the object encoding module for that resolution level. The clothing encoding data for each resolution level is the output of the clothing encoding module for that resolution level.

[0089] Computer equipment can encode object data using object encoding modules at each resolution level to obtain object-encoded data for that resolution level. The object encoding module at the last resolution level takes the object data as its input. The input of the object encoding module at each resolution level (excluding the last one) is the output of the next resolution level. Similarly, computer equipment can encode clothing data using clothing encoding modules at each resolution level to obtain object-encoded data for that resolution level. The clothing encoding module at the last resolution level takes the clothing data as its input. The input of the clothing encoding module at each resolution level (excluding the last one) is the output of the next resolution level.

[0090] Computer equipment can perform fusion processing on object coding data and clothing coding data at the same resolution level, as well as the fused coding data of the previous resolution level, to obtain fused coding data for each resolution level.

[0091] Computer equipment can perform clothing change prediction processing on fused coded data and object coded data at the same resolution level, as well as clothing change results at the previous resolution level, to obtain the target clothing change image.

[0092] In some embodiments, a computer device can fuse object-coded data and garment-coded data at the same resolution level based on the positional relationship between the object and the target garment. It is understood that the aforementioned positional relationship characterizes the position of the target garment on the object when it is worn on the object. Specifically, this positional relationship can be a mapping relationship between key points on the target garment and key points on the object.

[0093] In some embodiments, the clothing-changing model includes a fusion module for each resolution level. A computer device can use the fusion module at each resolution level to fuse the object encoding data and clothing encoding data at that resolution level, obtaining the fused encoding data output by the fusion module at that resolution level.

[0094] In some embodiments, the clothing-changing model includes a prediction module for each resolution level. A computer device can use the prediction module at each resolution level to perform clothing-changing prediction processing on the fused coded data and object coded data at that resolution level to obtain the target clothing-changing image.

[0095] In some embodiments, the computer device can perform multi-layer convolution processing on the object data and clothing data input into the clothing model to obtain object-coded data and clothing-coded data at each resolution level. Each layer of convolution processing is implemented through at least one convolution module.

[0096] In some embodiments, the skeleton data input into the clothing model is in the form of a heatmap.

[0097] In some embodiments, the object encoding module and the clothing encoding module are each composed of convolutional modules stacked together.

[0098] In some embodiments, the object encoding module and the clothing encoding module perform encoding through convolution processing, normalization, linear activation, downsampling pooling, and residual connections, respectively.

[0099] In some embodiments, the object encoding module and the clothing encoding module may include a combination module of convolution, normalization and linear activation (conv-BN-prelu), a downsampling pooling layer and a residual connection module.

[0100] In this embodiment, the object data and clothing data input to the clothing-changing model to be trained are subjected to multi-resolution encoding processing to obtain object-coded data and clothing-coded data at each resolution level; the object-coded data and clothing-coded data at the same resolution level are fused to obtain fused-coded data at each resolution level; clothing-changing prediction processing is performed based on the fused-coded data and object-coded data at each resolution level to predict the target clothing-changing image. Multi-resolution encoding processing can provide features at different resolutions, ensuring the accuracy of virtual clothing changing.

[0101] In some embodiments, the method further includes: performing advanced encoding on object coding data and clothing coding data at a first resolution level to obtain advanced encoded object coding data and advanced encoded clothing coding data; the resolution of the advanced encoded object coding data and the advanced encoded clothing coding data is lower than that of the object coding data and clothing coding data at the first resolution level; performing preliminary fusion on the advanced encoded object coding data and the advanced encoded clothing coding data to obtain preliminary fused coding data; performing fusion processing on object coding data and clothing coding data at the same resolution level to obtain fused coding data at each resolution level, including: determining the preliminary fused coding data, as well as the object coding data and clothing coding data at the first resolution level, as the input of the fusion module corresponding to the first resolution level, to obtain the fused coding data at the first resolution level output by the fusion module corresponding to the first resolution level; sequentially determining the current resolution level starting from the next resolution level after the first resolution level, determining the fused coding data of the previous resolution level of the current resolution level and the object coding data and clothing coding data of the current resolution level as the input of the fusion module corresponding to the current resolution level, to obtain the fused coding data output by the fusion module corresponding to the current resolution level.

[0102] For example, a computer device can use a separately established object encoding module to perform advanced encoding on the object encoding data at the first resolution level, obtaining advanced encoded object encoding data. Similarly, a computer device can use a separately established clothing encoding module to perform advanced encoding on the clothing encoding data at the first resolution level, obtaining advanced encoded clothing encoding data. It can be understood that the separately established object encoding module and clothing encoding module are independent levels; the advanced encoded clothing encoding data and object encoding data correspond to the lowest resolution, representing low-resolution features.

[0103] Computer equipment can merge and convolve the advanced-coded object data and clothing code data, followed by upsampling to obtain preliminary fused coded data. This preliminary fused coded data contains rich low-resolution features. Preliminary fused coded data = Upsampling(Convolution(Merging(Advanced-coded object data and clothing code data))).

[0104] In some embodiments, the fused coded data includes object feature data and clothing feature data. It is understood that the difference between object feature data and object coded data is that the object feature data incorporates clothing features, and this incorporated clothing feature is related to the positional relationship between the object and the clothing. The difference between clothing feature data and clothing coded data is that the clothing feature data is the feature data of the deformed target clothing, and this deformation conforms to the positional relationship between the object and the clothing. It is understood that the object feature data and clothing feature data are matched in the fused coded data at each resolution level.

[0105] In some embodiments, object feature data may be a part segmentation mask. Clothing feature data may be images of the target clothing at different resolutions after deformation.

[0106] In this embodiment, the preliminary fused encoding data, object encoding data and clothing encoding data of each resolution level are fused by the fusion module corresponding to each resolution level to obtain the fused encoding data of that resolution level output by the fusion module corresponding to each resolution level. This fully integrates the features of multiple resolutions, making virtual clothing changing more accurate.

[0107] In some embodiments, the method further includes: upsampling the object coding data of the advanced coding to obtain upsampled coding data; determining the preliminary fused coding data, as well as the object coding data and clothing coding data of the first resolution level, as the input of the fusion module corresponding to the first resolution level, to obtain the fused coding data of the first resolution level output by the fusion module corresponding to the first resolution level, including: determining the preliminary fused coding data, the upsampled coding data, as well as the object coding data and clothing coding data of the first resolution level, as the input of the fusion module corresponding to the first resolution level, to obtain the fused coding data of the first resolution level output by the fusion module corresponding to the first resolution level.

[0108] For example, the resolution corresponding to the first resolution level is greater than the resolution after advanced encoding. Therefore, the computer device can upsample the object encoding data of the advanced encoding so that the resolution corresponding to the upsampled encoding data matches the resolution corresponding to the first resolution level.

[0109] The computer device can fuse the preliminary fused encoded data, the upsampled encoded data, and the object encoded data and clothing encoded data of the first resolution level through the fusion module corresponding to the first resolution level, to obtain the fused encoded data of the first resolution level output by the fusion module corresponding to the first resolution level.

[0110] In some embodiments, such as Figure 4The diagram shows the internal structure of the fusion module. The fusion module corresponding to the i-th resolution level is shown in the figure. The input of the fusion module corresponding to the i-th resolution level includes the output of the fusion module corresponding to the (i-1)-th resolution level. The i-th fusion module can upsample the clothing feature data of the (i-1)-th resolution level to obtain upsampled clothing feature data. The i-th fusion module can perform bilinear interpolation on the upsampled clothing feature data and the clothing encoding data of the i-th resolution level to obtain interpolated clothing feature data. The specific method of bilinear interpolation can be grid sampling. The i-th fusion module can merge the object feature data of the (i-1)-th resolution level, the object encoding data of the i-th resolution level, and the interpolated clothing feature data to obtain merged feature data. The i-th fusion module can perform convolution on the merged feature data to obtain convolutional feature data. The i-th fusion module can sum the convolutional feature data and the upsampled clothing feature data element-wise to obtain the clothing feature data of the i-th resolution level. The i-th fusion module can upsample the merged feature data to obtain upsampled feature data. The i-th fusion module can then perform residual convolution on the upsampled feature data to obtain the object feature data for the i-th resolution level. It can be understood that the clothing coding data input to the i-th fusion module can be obtained by element-wise summing of the clothing coding data from the previous i-1 resolution levels; that is, the input to the fusion module corresponding to this resolution level is the superposition of the clothing coding data from all previous levels.

[0111] In this embodiment, the object encoding data of the advanced encoding is upsampled to obtain the upsampled encoding data. The preliminary fused encoding data, the upsampled encoding data, and the object encoding data and clothing encoding data of the first resolution level are determined as the input of the fusion module corresponding to the first resolution level. The fused encoding data of the first resolution level output by the fusion module corresponding to the first resolution level is obtained, which fully integrates the features of multiple resolutions, making virtual clothing changing more accurate.

[0112] In some embodiments, the target clothing-changing image is the clothing-changing result of the last resolution level; the clothing-changing prediction processing based on the fused coding data and object coding data of each resolution level to predict the target clothing-changing image includes: performing clothing-changing prediction processing based on the fused coding data and object coding data of the first resolution level to predict the clothing-changing result of the first resolution level; sequentially determining the current resolution level starting from the next resolution level of the first resolution level, and performing clothing-changing prediction processing based on the clothing-changing result of the previous resolution level of the current resolution level and the fused coding data and object coding data of the current resolution level to predict the clothing-changing result of the current resolution level; wherein the resolution corresponding to the clothing-changing result of each resolution level is different.

[0113] For example, such as Figure 5 The diagram shows the internal structure of the prediction module. The module corresponding to the i-th resolution level is shown in the figure. The input to the prediction module for the i-th resolution level includes the clothing-changing result from the (i-1)-th resolution level, as well as the object feature data, object encoding data, and clothing feature data from the i-th resolution level. The prediction module for the i-th resolution level performs bilinear interpolation on the clothing feature data from the i-th resolution level and the clothing data input to the clothing-changing model to obtain the interpolated clothing data. Bilinear interpolation can be implemented using grid sampling (grid_sample).

[0114] The prediction module corresponding to the i-th resolution level can perform convolution processing on the interpolated clothing data to obtain convolutionally processed clothing data. The prediction module corresponding to the i-th resolution level can also perform convolution processing on the object encoding data of the i-th resolution level to obtain convolutionally processed object encoding data. Finally, the prediction module corresponding to the i-th resolution level can perform convolution processing on the clothing changing results of the (i-1)-th resolution level to obtain convolutionally processed clothing changing results.

[0115] The prediction module corresponding to the i-th resolution level can merge the convolutional clothing-changing results, the convolutional object encoding data, and the convolutional clothing data to obtain the merged clothing-changing data. The prediction module corresponding to the i-th resolution level can then perform convolutional processing and normalization on the merged clothing-changing data to obtain the normalized clothing-changing data.

[0116] The prediction module corresponding to the i-th resolution level can perform convolution processing on the object feature data to obtain convolutionally processed object feature data. The prediction module corresponding to the i-th resolution level can then perform element-wise multiplication on the convolutionally processed object feature data and the normalized object feature data to obtain the multiplied result. Finally, the prediction module corresponding to the i-th resolution level can element-wise sum the multiplied result and the convolutionally processed object feature data to obtain the clothing-changing result for the i-th resolution level. The clothing-changing result for each resolution level is then processed through non-linear activation to obtain the clothing-changing image for that resolution level. The target clothing-changing image can be obtained by processing the clothing-changing result for the last resolution level through non-linear activation.

[0117] In some embodiments, such as Figure 6 The diagram shows the structural schematic of the clothing-changing model. The model includes four resolution levels: object encoding modules (Ep1-Ep4), clothing encoding modules (Ec1-Ec4), fusion modules (mask_flow1-mask_flow4), and prediction modules (tryon1-tryon4). It also includes additional object encoding modules (Ep0) and clothing encoding modules (Ec0) for advanced encoding, a preliminary fusion module (flow) for initial fusion, and an upsampling module for upsampling the object encoding data from the advanced encoding. Each training round of the clothing-changing module includes two inputs and two outputs. The two inputs are the first input (input1) and the second input (input2). The two outputs are the first output (output1) and the second output (output2).

[0118] The first input includes a first matting, a first matting mask, first skeleton data, a first dense coordinate map, a second clothing image, and a second clothing mask. The first output includes a first predicted clothing image and a first predicted clothing mask, a first clothing change image, and a first object prediction mask. The second input includes a first matting, a second matting mask, second skeleton data, a second dense coordinate map, a first predicted clothing image, and a first predicted clothing mask. The second output includes a second clothing change image, a second object prediction mask, a second predicted clothing image, and a second clothing prediction mask.

[0119] In some embodiments, the computer device can perform loss calculations for each clothing change result to obtain a loss value corresponding to each resolution level. The computer device can then train the model in the direction that reduces the loss value corresponding to each resolution level.

[0120] In this embodiment, clothing swapping prediction processing is performed based on the fused coding data and object coding data of the first resolution level to predict the clothing swapping result of the first resolution level. Starting from the next resolution level after the first resolution level, the current resolution level is determined sequentially. Clothing swapping prediction processing is performed based on the clothing swapping result of the previous resolution level of the current resolution level, as well as the fused coding data and object coding data of the current resolution level to predict the clothing swapping result of the current resolution level. This achieves clothing swapping prediction at multiple resolution levels, so that the output of the last resolution level combines multiple low-resolution and high-resolution features to obtain a more accurate target clothing swapping image.

[0121] In some embodiments, the method further includes: extracting first clothing data from a first sample image; determining first clothing prediction data and second clothing prediction data output by a clothing-changing model to be trained; the first clothing prediction data is used to characterize the target clothing in the first clothing-changing image; the second clothing prediction data is used to characterize the target clothing in the second clothing-changing image; training the model based on the difference between the first clothing-changing image and the first sample image, and the difference between the second clothing-changing image and the second sample image, to obtain a trained clothing-changing model, including: training the model based on a first difference, a second difference, a third difference, and a fourth difference to obtain a trained clothing-changing model; the first difference refers to the difference between the first clothing-changing image and the first sample image; the second difference refers to the difference between the second clothing-changing image and the second sample image; the third difference refers to the difference between the first clothing data and the first clothing prediction data; and the fourth difference refers to the difference between the second clothing data and the second clothing prediction data.

[0122] For example, the first clothing prediction data is the clothing data of the target clothing in the first clothing change image. The second clothing prediction data is the clothing data of the target clothing in the second clothing change image. The computer device can train the model in the direction of decreasing according to the first difference, second difference, third difference, and fourth difference to obtain a trained clothing change model.

[0123] In some embodiments, the computer device can use first clothing prediction data and first clothing data as input to the discriminator to obtain a third discriminator loss. The computer device can use second clothing prediction data and second clothing data as input to the discriminator to obtain a fourth discriminator loss. The computer device can tune the parameters of the clothing-changing model to be trained in a direction that reduces the first, second, third, and fourth discriminator losses to obtain a trained clothing-changing model.

[0124] In some embodiments, the computer device may perform a first optimization on the initial clothing-changing model in the current round, in the direction of reducing the first discriminator loss and the third discriminator loss, to obtain a clothing-changing model optimized once in the current round. The computer device may perform a second optimization on the clothing-changing model optimized once in the current round, in the direction of reducing the second discriminator loss and the fourth discriminator loss, to obtain a clothing-changing model optimized twice in the current round.

[0125] In some embodiments, the computer device may calculate an adjacent gradient smoothing loss for the first clothing prediction data to obtain a first smoothing loss. The computer device may also calculate an adjacent gradient smoothing loss for the second clothing prediction data to obtain a second smoothing loss. It is understood that a third difference may include at least one of a third discriminator loss or a first smoothing loss. A fourth difference may include at least one of a fourth discriminator loss or a second smoothing loss.

[0126] In some embodiments, the first difference may include at least one of a first perceptual loss or a first discriminator loss. The second difference may include at least one of a second perceptual loss or a second discriminator loss.

[0127] In this embodiment, the model is trained based on the first difference, the second difference, the third difference, and the fourth difference to obtain a trained clothing-changing model. The model is trained using multiple differences to ensure the accuracy of the model training.

[0128] In some embodiments, the first object data includes a first object mask; the second object data includes a second object mask; the method further includes: determining a first object prediction mask and a second object prediction mask output by the clothing-changing model to be trained; the first object prediction mask is a mask of an object in a first clothing-changing image; the second object prediction mask is a mask of an object in a second clothing-changing image; training the model based on a first difference, a second difference, a third difference, and a fourth difference to obtain a trained clothing-changing model includes: training the model based on a first difference, a second difference, a third difference, a fourth difference, a fifth difference, and a sixth difference to obtain a trained clothing-changing model; the fifth difference refers to the difference between the first object mask and the first object prediction mask; the sixth difference refers to the difference between the second object mask and the second object prediction mask.

[0129] For example, the first object prediction mask can be a component segmentation mask of an object in the first clothing change image. The second object prediction mask can be a component segmentation mask of an object in the second clothing change image. The first object mask can be a component segmentation mask of an object in the first sample image. The second object mask can be a component segmentation mask of an object in the second sample image. The computer device can calculate the mean absolute error of the first object prediction mask and the second object mask to obtain a fifth difference. The computer device can calculate the mean absolute error of the second object prediction mask and the second object mask to obtain a sixth difference. The computer device can train the model in a direction that decreases according to the first difference, the second difference, the third difference, the fourth difference, the fifth difference, and the sixth difference to obtain a trained clothing change model.

[0130] In some embodiments, the computer device may calculate a perceptual loss on the first object prediction mask and the first object mask to obtain a fifth difference. The computer device may also calculate a perceptual loss on the second object prediction mask and the second object mask to obtain a sixth difference.

[0131] In this embodiment, the model is trained based on the first difference, the second difference, the third difference, the fourth difference, the fifth difference, and the sixth difference to obtain a trained clothing-changing model. The model is trained using multiple differences to ensure the accuracy of the model training.

[0132] In some embodiments, the computer device can identify the object to be dressed and extract third object data of that object. The computer device can also identify third garment data of the garment to be dressed. The computer device can use the third object data and the third garment data as input to a trained dressing-changing model, and after one forward inference pass through the model, obtain a third dressing-changing image output by the model. The third dressing-changing image shows the effect of the object wearing the garment. It is understood that the garment to be dressed can be a single item of clothing or clothing worn by other objects besides the object to be dressed.

[0133] It should be understood that although the steps in the flowcharts of the embodiments described above are shown sequentially according to the arrows, these steps are not necessarily executed in the order indicated by the arrows. Unless explicitly stated herein, there is no strict order restriction on the execution of these steps, and they can be executed in other orders. Moreover, at least some steps in the flowcharts of the embodiments described above may include multiple steps or multiple stages. These steps or stages are not necessarily completed at the same time, but can be executed at different times. The execution order of these steps or stages is not necessarily sequential, but can be performed alternately or in turn with other steps or at least some of the steps or stages of other steps.

[0134] Based on the same inventive concept, this application also provides a virtual clothing changing device for implementing the virtual clothing changing method described above. The solution provided by this device is similar to the solution described in the above method; therefore, the specific limitations in one or more virtual clothing changing device embodiments provided below can be found in the limitations of the virtual clothing changing method described above, and will not be repeated here.

[0135] In some embodiments, such as Figure 7 As shown, a virtual clothing changing device 700 is provided, including: an extraction unit 702, a training unit 704, and an optimization unit 706, wherein:

[0136] Extraction unit 702 is used to extract first object data from a first sample image including an object in a first pose and wearing target clothing; and to extract second clothing data and second object data from a second sample image including an object in a second pose and wearing target clothing.

[0137] Training unit 704 is used to input first object data and second clothing data into the clothing-changing model to be trained, and predict to obtain a first clothing-changing image; the first clothing-changing image presents the effect of an object in a first pose wearing the target clothing; input the clothing data of the target clothing and the second object data in the first clothing-changing image into the clothing-changing model to be trained, and predict to obtain a second clothing-changing image; the second clothing-changing image presents the effect of an object in a second pose wearing the target clothing.

[0138] The optimization unit 706 is used to train the model based on the difference between the first clothing change image and the first sample image, and the difference between the second clothing change image and the second sample image, to obtain the trained clothing change model.

[0139] In some embodiments, the training unit 704 is configured to perform multi-resolution encoding processing on the object data and clothing data input to the clothing-changing model to be trained, to obtain object-coded data and clothing-coded data at each resolution level; perform fusion processing on the object-coded data and clothing-coded data at the same resolution level, to obtain fused encoding data at each resolution level; and perform clothing-changing prediction processing based on the fused encoding data and object-coded data at each resolution level to predict the target clothing-changing image; wherein the target clothing-changing image includes either a first clothing-changing image or a second clothing-changing image.

[0140] In some embodiments, the training unit 704 is configured to perform advanced encoding on the object encoding data and clothing encoding data of the first resolution level to obtain advanced encoded object encoding data and advanced encoded clothing encoding data; the resolution of the advanced encoded object encoding data and the advanced encoded clothing encoding data is lower than that of the object encoding data and clothing encoding data of the first resolution level; perform preliminary fusion on the advanced encoded object encoding data and the advanced encoded clothing encoding data to obtain preliminary fused encoding data; determine the preliminary fused encoding data, as well as the object encoding data and clothing encoding data of the first resolution level, as the input of the fusion module corresponding to the first resolution level to obtain the fused encoding data of the first resolution level output by the fusion module corresponding to the first resolution level; sequentially determine the current resolution level starting from the next resolution level of the first resolution level, determine the fused encoding data of the previous resolution level of the current resolution level and the object encoding data and clothing encoding data of the current resolution level as the input of the fusion module corresponding to the current resolution level to obtain the fused encoding data output by the fusion module corresponding to the current resolution level.

[0141] In some embodiments, the training unit 704 is used to upsample the object encoding data of the advanced encoding to obtain upsampled encoding data; and to determine the preliminary fused encoding data, the upsampled encoding data, and the object encoding data and clothing encoding data of the first resolution level as the input of the fusion module corresponding to the first resolution level to obtain the fused encoding data of the first resolution level output by the fusion module corresponding to the first resolution level.

[0142] In some embodiments, the target clothing-changing image is the clothing-changing result of the last resolution level; the training unit 704 is used to perform clothing-changing prediction processing based on the fusion coding data and object coding data of the first resolution level to predict the clothing-changing result of the first resolution level; starting from the next resolution level of the first resolution level, the current resolution level is determined sequentially, and clothing-changing prediction processing is performed based on the clothing-changing result of the previous resolution level of the current resolution level and the fusion coding data and object coding data of the current resolution level to predict the clothing-changing result of the current resolution level; wherein, the resolution corresponding to the clothing-changing result of each resolution level is different.

[0143] In some embodiments, the extraction unit 702 is used to extract first clothing data from the first sample image; the training unit 704 is used to determine the first clothing prediction data and the second clothing prediction data output by the clothing-changing model to be trained; the first clothing prediction data is used to represent the target clothing in the first clothing-changing image; the second clothing prediction data is used to represent the target clothing in the second clothing-changing image; the optimization module is used to train the model according to the first difference, the second difference, the third difference, and the fourth difference to obtain the trained clothing-changing model; the first difference refers to the difference between the first clothing-changing image and the first sample image; the second difference refers to the difference between the second clothing-changing image and the second sample image; the third difference refers to the difference between the first clothing data and the first clothing prediction data; the fourth difference refers to the difference between the second clothing data and the second clothing prediction data.

[0144] In some embodiments, the first object data includes a first object mask; the second object data includes a second object mask; the training unit 704 is used to determine the first object prediction mask and the second object prediction mask output by the clothing-changing model to be trained; the first object prediction mask is the mask of the object in the first clothing-changing image; the second object prediction mask is the mask of the object in the second clothing-changing image; the optimization unit 706 is used to train the model according to the first difference, the second difference, the third difference, the fourth difference, the fifth difference, and the sixth difference to obtain the trained clothing-changing model; the fifth difference refers to the difference between the first object mask and the first object prediction mask; the sixth difference refers to the difference between the second object mask and the second object prediction mask.

[0145] Each module in the aforementioned virtual changing device can be implemented entirely or partially through software, hardware, or a combination thereof. These modules can be embedded in the processor of a computer device in hardware form or independent of it, or stored in the memory of the computer device in software form, so that the processor can call and execute the operations corresponding to each module.

[0146] In some embodiments, a computer device is provided, which may be a server, and its internal structure diagram may be as follows: Figure 8As shown, the computer device includes a processor, memory, input / output (I / O) interfaces, and a communication interface. The processor, memory, and I / O interfaces are connected via a system bus, and the communication interface is also connected to the system bus via the I / O interfaces. The processor provides computational and control capabilities. The memory includes non-volatile storage media and internal memory. The non-volatile storage media stores the operating system, computer programs, and a database. The internal memory provides an environment for the operation of the operating system and computer programs stored in the non-volatile storage media. The database stores a first sample image and a second sample image. The I / O interfaces are used for exchanging information between the processor and external devices. The communication interface is used for communicating with external terminals via a network connection. When executed by the processor, the computer program implements a virtual clothing changing method.

[0147] In some embodiments, a computer device is provided, which may be a terminal, and its internal structure diagram may be as follows: Figure 9 As shown, the computer device includes a processor, memory, input / output interfaces, a communication interface, a display unit, and an input device. The processor, memory, and input / output interfaces are connected via a system bus, and the communication interface, display unit, and input device are also connected to the system bus via the input / output interfaces. The processor provides computing and control capabilities. The memory includes non-volatile storage media and internal memory. The non-volatile storage media stores the operating system and computer programs. The internal memory provides an environment for the operation of the operating system and computer programs stored in the non-volatile storage media. The input / output interfaces are used for exchanging information between the processor and external devices. The communication interface is used for wired or wireless communication with external terminals; wireless communication can be achieved through Wi-Fi, mobile cellular networks, NFC (Near Field Communication), or other technologies. When the computer program is executed by the processor, it implements a virtual clothing changing method. The display unit is used to form a visually visible image and can be a display screen, a projection device, or a virtual reality imaging device. The display screen can be an LCD screen or an e-ink screen. The input device of the computer device can be a touch layer covering the display screen, or buttons, trackballs, or touchpads set on the casing of the computer device, or external keyboards, touchpads, or mice, etc.

[0148] Those skilled in the art will understand that Figure 8 or Figure 9The structure shown is merely a block diagram of a portion of the structure related to the present application and does not constitute a limitation on the computer device to which the present application is applied. Specific computer devices may include more or fewer components than those shown in the figure, or combine certain components, or have different component arrangements.

[0149] In some embodiments, a computer device is provided, including a memory and a processor, wherein the memory stores a computer program, and the processor executes the computer program to implement the steps in the above-described method embodiments.

[0150] In some embodiments, a computer-readable storage medium is provided having a computer program stored thereon, which, when executed by a processor, implements the steps in the above method embodiments.

[0151] In some embodiments, a computer program product is provided, including a computer program that, when executed by a processor, implements the steps in the above method embodiments.

[0152] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, data stored, data displayed, etc.) involved in this application are all information and data authorized by the user or fully authorized by all parties, and the collection, use and processing of the relevant data shall comply with the relevant laws, regulations and standards of the relevant countries and regions.

[0153] Those skilled in the art will understand that all or part of the processes in the methods of the above embodiments can be implemented by a computer program instructing related hardware. The computer program can be stored in a non-volatile computer-readable storage medium, and when executed, it can include the processes of the embodiments of the above methods. Any references to memory, databases, or other media used in the embodiments provided in this application can include at least one of non-volatile and volatile memory. Non-volatile memory can include read-only memory (ROM), magnetic tape, floppy disk, flash memory, optical memory, high-density embedded non-volatile memory, resistive random access memory (ReRAM), magnetic random access memory (MRAM), ferroelectric random access memory (FRAM), phase change memory (PCM), graphene memory, etc. Volatile memory can include random access memory (RAM) or external cache memory, etc. By way of illustration and not limitation, RAM can take many forms, such as Static Random Access Memory (SRAM) or Dynamic Random Access Memory (DRAM). The databases involved in the embodiments provided in this application may include at least one type of relational database and non-relational database. Non-relational databases may include, but are not limited to, blockchain-based distributed databases. The processors involved in the embodiments provided in this application may be general-purpose processors, central processing units, graphics processing units, digital signal processors, programmable logic devices, quantum computing-based data processing logic devices, etc., and are not limited to these.

[0154] The technical features of the above embodiments can be combined in any way. For the sake of brevity, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.

[0155] The embodiments described above are merely illustrative of several implementation methods of this application, and while the descriptions are specific and detailed, they should not be construed as limiting the scope of this patent application. It should be noted that those skilled in the art can make various modifications and improvements without departing from the concept of this application, and these all fall within the protection scope of this application. Therefore, the protection scope of this application should be determined by the appended claims.

Claims

1. A virtual clothing changing method, characterized in that, The method includes: Extract first object data from a first sample image that includes an object in a first pose and wearing the target clothing; Extract second clothing data and second object data from a second sample image including an object in a second pose and wearing the target clothing; The process involves inputting first object data and second clothing data into a clothing-changing model to be trained, and predicting a first clothing-changing image. This includes: performing multi-resolution encoding on the object data and clothing data input to the clothing-changing model to be trained, obtaining object-encoded data and clothing-encoded data at each resolution level; fusing the object-encoded data and clothing-encoded data at the same resolution level, obtaining fused-encoded data at each resolution level; and performing clothing-changing prediction processing based on the fused-encoded data and object-encoded data at each resolution level to predict a target clothing-changing image. The target clothing-changing image includes the first clothing-changing image, which shows the effect of an object in the first pose wearing the target clothing. The clothing data of the target garment in the first clothing image and the data of the second object are input into the clothing model to be trained to predict the second clothing image; the second clothing image shows the effect of the object in the second pose wearing the target garment; The model is trained based on the differences between the first clothing change image and the first sample image, and the differences between the second clothing change image and the second sample image, to obtain the trained clothing change model. The target clothing-changing image is the clothing-changing result at the last resolution level; the clothing-changing prediction processing based on the fused encoded data and object encoded data at each resolution level to predict the target clothing-changing image includes: Clothing change prediction processing is performed based on the fused encoded data and object encoded data of the first resolution level to predict the clothing change result of the first resolution level; starting from the next resolution level of the first resolution level, the current resolution level is determined sequentially, and clothing change prediction processing is performed based on the clothing change result of the previous resolution level of the current resolution level and the fused encoded data and object encoded data of the current resolution level to predict the clothing change result of the current resolution level; the resolution corresponding to the clothing change result of each resolution level is different.

2. The virtual clothing changing method according to claim 1, characterized in that, The second object data is used to reflect the posture of the object in the second sample image and the wearing status of the target clothing; The second clothing data is used to reflect the shape of the target clothing shown in the second sample image.

3. The virtual clothing changing method according to claim 1, characterized in that, The method further includes: The object coding data and clothing coding data at the first resolution level are further encoded to obtain further encoded object coding data and further encoded clothing coding data; the resolution of the further encoded object coding data and further encoded clothing coding data is lower than that of the object coding data and clothing coding data at the first resolution level. The object coding data and clothing coding data of the advanced coding are initially fused to obtain preliminary fused coding data; The process of fusing object coding data and clothing coding data at the same resolution level to obtain fused coding data at each resolution level includes: The preliminary fusion coding data, as well as the object coding data and clothing coding data of the first resolution level, are determined as the input of the fusion module corresponding to the first resolution level, so as to obtain the fusion coding data of the first resolution level output by the fusion module corresponding to the first resolution level. Starting from the next resolution level after the first resolution level, the current resolution level is determined sequentially. The fusion encoding data of the previous resolution level of the current resolution level, the object encoding data and clothing encoding data of the current resolution level are determined as the input of the fusion module corresponding to the current resolution level, and the fusion encoding data output by the fusion module corresponding to the current resolution level is obtained.

4. The virtual clothing changing method according to claim 3, characterized in that, The method further includes: Upsample the object encoding data of the advanced encoding to obtain the upsampled encoding data; The step of determining the preliminary fused encoding data, as well as the object encoding data and clothing encoding data of the first resolution level, as the input to the fusion module corresponding to the first resolution level, and obtaining the fused encoding data of the first resolution level output by the fusion module corresponding to the first resolution level, includes: The preliminary fused encoding data, the upsampled encoding data, and the object encoding data and clothing encoding data of the first resolution level are determined as the input of the fusion module corresponding to the first resolution level, so as to obtain the fused encoding data of the first resolution level output by the fusion module corresponding to the first resolution level.

5. The virtual clothing changing method according to claim 1, characterized in that, The first sample image and the second sample image include the same object; The first sample image and the second sample image are two images captured when the same object wears the target clothing in different postures.

6. The virtual clothing changing method according to any one of claims 1 to 5, characterized in that, The method further includes: Extract the first garment data from the first sample image; The first clothing prediction data and the second clothing prediction data output by the clothing-changing model to be trained are determined; the first clothing prediction data is used to characterize the target clothing in the first clothing-changing image; the second clothing prediction data is used to characterize the target clothing in the second clothing-changing image. The step of training the model based on the differences between the first clothing-changing image and the first sample image, and the differences between the second clothing-changing image and the second sample image, to obtain the trained clothing-changing model includes: The model is trained based on the first difference, the second difference, the third difference, and the fourth difference to obtain the trained clothing-changing model; the first difference refers to the difference between the first clothing-changing image and the first sample image; the second difference refers to the difference between the second clothing-changing image and the second sample image; the third difference refers to the difference between the first clothing data and the first clothing prediction data; and the fourth difference refers to the difference between the second clothing data and the second clothing prediction data.

7. The virtual clothing changing method according to claim 6, characterized in that, The first object data includes a first object mask; the second object data includes a second object mask; the method further includes: Determine the first object prediction mask and the second object prediction mask output by the clothing-changing model to be trained; the first object prediction mask is the mask of the object in the first clothing-changing image; the second object prediction mask is the mask of the object in the second clothing-changing image; The process of training the model based on the first difference, second difference, third difference, and fourth difference to obtain the trained clothing-changing model includes: The model is trained based on the first difference, second difference, third difference, fourth difference, fifth difference, and sixth difference to obtain the trained clothing-changing model; the fifth difference refers to the difference between the first object mask and the first object prediction mask; the sixth difference refers to the difference between the second object mask and the second object prediction mask.

8. A virtual clothing changing device, characterized in that, The virtual clothing changing device uses a virtual clothing changing method as described in any one of claims 1 to 7.

9. A computer device comprising a memory and a processor, wherein the memory stores a computer program, characterized in that, When the processor executes the computer program, it implements the steps of the method according to any one of claims 1 to 7.

10. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by a processor, it implements the steps of the method according to any one of claims 1 to 7.

Citation Information

Patent Citations

  • Fitting model training method, virtual fitting method and related device

    CN115564871A

  • Image processing method and apparatus, device, and storage medium

    WO2021164534A1