Virtual clothing changing methods, devices, computer equipment, and storage media

CN116229216BActive Publication Date: 2026-08-14XIAMEN MEITUZHIJIA TECH
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-04-04
Publication Date
2026-08-14

AI Technical Summary

Technical Problem

这种方法过于依赖配对数据,对于除配对数据外的其他数据利用率低,导致数据资源的浪费

Benefits of technology

[0037]上述虚拟换衣方法、装置、计算机设备、存储介质和计算机程序产品,通过每组样本中的目标服装数据和换衣对象数据作为待训练的换衣模型的训练数据,换衣对象数据是从原始穿衣图像中提取的,目标服装数据所表征的目标服装并非原始穿衣图像中的原始服装,目标服装和原始服装不要求完全一致,只要同类别即可。通过待训练的换衣模型从换衣对象数据中学习对象的关键点与原始服装的关键点之间的位置关系,同类别的服装的关键点对应的语义一致,上述位置关系一定程度上能够反映目标服装的关键点与对象的关键点之间的关系,进而基于位置关系能够实现换衣对象数据和目标服装数据的虚拟换衣处理,能够得到准确的换衣图像,基于换衣图像与原始图像之间的差异确定目标损失,根据目标损失优化待训练的换衣模型,以得到训练完毕的换衣模型,只需目标服装与原始服装同一类别,无需完全一致,大大提高了数据资源的利用率。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116229216B_ABST
    Figure CN116229216B_ABST
Patent Text Reader

Abstract

This application relates to a virtual clothing-changing method, apparatus, computer device, and storage medium. The method includes: determining multiple sets of samples; each set of samples includes target clothing data and clothing-changing object data; the clothing-changing object data is extracted from an original image of the object wearing the original clothing; the original clothing and the target clothing data represent the same category; a clothing-changing model to be trained learns the positional relationship between key points of the object and key points of the original clothing from the clothing-changing object data; key points of clothing of the same category correspond to semantic consistency; virtual clothing-changing processing is performed on the clothing-changing object data and target clothing data based on the positional relationship to obtain a clothing-changing image output by the clothing-changing model to be trained; a target loss is determined based on the difference between the clothing-changing image and the original clothing image, and the clothing-changing model to be trained is optimized based on the target loss to obtain a trained clothing-changing model. This method can improve the utilization rate of data resources.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of artificial intelligence technology, and in particular to a virtual clothing changing method, apparatus, computer device, and storage medium. Background Technology

[0002] In recent years, e-commerce has developed rapidly, and more and more consumers are choosing to shop on e-commerce platforms. However, when consumers buy non-standard products such as clothing online, they can only rely on the data provided by the merchants and their personal experience. Therefore, virtual clothing-changing technology has emerged.

[0003] Traditional techniques require at least matching data, such as images of clothing items and corresponding models wearing those items, to perform virtual clothing swapping by applying deep learning to the matching data. This method relies too heavily on the matching data, resulting in low utilization of other data and a waste of data resources. Summary of the Invention

[0004] Therefore, it is necessary to provide a virtual clothing changing method, apparatus, computer equipment, computer-readable storage medium, and computer program product that can improve the utilization rate of data resources, in response to the above-mentioned technical problems.

[0005] Firstly, this application provides a virtual clothing changing method. The method includes:

[0006] Multiple sets of samples are identified; each set of samples includes target clothing data and swapping object data; the swapping object data is extracted from the original clothing image of the presented object wearing the original clothing; the original clothing and the target clothing represented by the target clothing data are of the same category.

[0007] The model to be trained learns the positional relationship between the key points of the object and the key points of the original clothing from the data of the object to be dressed; the key points of clothing in the same category have the same semantics.

[0008] Virtual clothing swapping is performed on the data of the object and the target clothing based on their positional relationship, resulting in a clothing swapping image output by the clothing swapping model to be trained; the clothing swapping image shows the effect of the object wearing the target clothing.

[0009] The target loss is determined based on the difference between the clothing image and the original clothing image, and the clothing model to be trained is optimized based on the target loss to obtain the trained clothing model.

[0010] In some embodiments, the clothing swapping object data includes object skeleton data and original keypoint data; the original keypoint data is used to characterize the keypoints of the original garment; learning the positional relationship between the keypoints of the object and the keypoints of the original garment from the clothing swapping object data through the clothing swapping model to be trained includes:

[0011] Based on the gating mechanism in the clothing-changing model to be trained, the skeleton key point data is extracted from the object skeleton data;

[0012] By associating the skeleton key point data with the original key point data, the positional relationship between the skeleton key points of the learning object and the key points of the original clothing is learned.

[0013] In some embodiments, the clothing object data includes object texture data and region mask data; the region mask data refers to the mask of the remaining region after removing the region that affects the wearing of the target clothing; based on the gating mechanism in the clothing-changing model to be trained, the skeleton key point data is extracted from the object skeleton data, including:

[0014] The object texture data and region mask data are fused to obtain the object fused data;

[0015] Wearing region data is extracted from the object fusion data based on the gating data in the clothing-changing model to be trained; the wearing region data is used to indicate the wearing region when the object wears the target clothing.

[0016] Key skeletal point data are extracted from the object's skeleton data based on the wearable area data.

[0017] In some embodiments, virtual clothing changing is performed on the clothing object data and target clothing data based on positional relationships to obtain a clothing changing image output by the clothing changing model to be trained, including:

[0018] Based on the positional relationship, the key points of the target garment are mapped to obtain the target key point data of the target garment; the target key point data is used to characterize the position of the key points of the target garment when the object is wearing the target garment;

[0019] The target clothing data is adjusted based on the target key point data to obtain deformable clothing data; the deformable clothing data represents a deformable clothing that matches the posture of the object.

[0020] Virtual clothing swapping is performed on deformable clothing data and clothing swapping object data to obtain the clothing swapping image output by the clothing swapping model to be trained.

[0021] In some embodiments, target clothing data is adjusted based on target key point data to obtain deformable clothing data, including:

[0022] Feature extraction is performed on the target key point data and the target clothing data to obtain deformation feature data;

[0023] The target clothing image in the target clothing data is deformed based on the deformation feature data to obtain deformed clothing data.

[0024] In some embodiments, the target loss includes a first loss; the first loss is determined based on the difference between the clothing image and the original clothing image; the target loss also includes at least one of a second loss or a third loss; the second loss is determined based on the difference between the target clothing key point data and the original clothing key point data; the third loss is determined based on the difference between the deformed clothing data and the original clothing data in the clothing object data.

[0025] In some embodiments, virtual clothing changing processing is performed on deformable clothing data and clothing object data to obtain a clothing changing image output by the clothing changing model to be trained, including:

[0026] The deformed clothing data and the clothing object data are fused at the first resolution to obtain the clothing change result corresponding to the first resolution.

[0027] Starting from the next resolution after the first resolution, the current resolution is determined sequentially. The deformed clothing data, the clothing object data, and the clothing change results corresponding to the previous resolution are fused together to obtain the clothing change results corresponding to the current resolution.

[0028] The clothing change result corresponding to the last resolution is subjected to nonlinear mapping processing to obtain the clothing change image.

[0029] Secondly, this application also provides a virtual fitting device. The device includes:

[0030] The determination module is used to determine multiple sets of samples; each set of samples includes target clothing data and clothing swapping object data; the clothing swapping object data is extracted from the original clothing image of the presented object wearing the original clothing; the original clothing and the target clothing represented by the target clothing data are of the same category.

[0031] The learning module is used to learn the positional relationship between the key points of the object and the key points of the original clothing from the data of the object to be trained through the clothing swapping model; the key points of clothing in the same category have semantic consistency.

[0032] The output module is used to perform virtual clothing swapping on the data of the object to be swapped and the data of the target clothing based on the positional relationship, so as to obtain the clothing swapping image output by the clothing swapping model to be trained; the clothing swapping image shows the effect of the object wearing the target clothing.

[0033] The optimization module is used to determine the target loss based on the difference between the clothing image and the original clothing image, and to optimize the clothing model to be trained based on the target loss to obtain the trained clothing model.

[0034] Thirdly, this application also provides a computer device. The computer device includes a memory and a processor, the memory storing a computer program, and the processor executing the computer program to implement the steps in the method described above.

[0035] Fourthly, this application also provides a computer-readable storage medium. The computer-readable storage medium stores a computer program thereon, which, when executed by a processor, implements the steps of the method described above.

[0036] Fifthly, this application also provides a computer program product. The computer program product includes a computer program that, when executed by a processor, implements the steps of the method described above.

[0037] The aforementioned virtual clothing-changing method, apparatus, computer equipment, storage medium, and computer program products use target clothing data and clothing object data from each set of samples as training data for the clothing-changing model to be trained. The clothing object data is extracted from the original clothing image, and the target clothing data represents a different type of clothing than the original clothing in the original clothing image. The target clothing and the original clothing do not need to be completely identical, as long as they belong to the same category. The clothing-changing model to be trained learns the positional relationship between the key points of the object and the key points of the original clothing from the clothing object data. The key points of clothing of the same category have semantic consistency. The positional relationship can reflect the relationship between the key points of the target clothing and the key points of the object to a certain extent. Based on the positional relationship, virtual clothing-changing processing of the clothing object data and target clothing data can be realized, resulting in an accurate clothing-changing image. The target loss is determined based on the difference between the clothing-changing image and the original image. The clothing-changing model to be trained is optimized based on the target loss to obtain the trained clothing-changing model. As long as the target clothing and the original clothing belong to the same category, they do not need to be completely identical, which greatly improves the utilization rate of data resources. Attached Figure Description

[0038] Figure 1 This is a flowchart illustrating a virtual clothing changing method in one embodiment;

[0039] Figure 2 This is a schematic diagram of the key point prediction unit in a clothing-changing model in one embodiment.

[0040] Figure 3 This is a schematic diagram of the structure of the clothing deformation unit in a clothing changing model in one embodiment;

[0041] Figure 4 This is a schematic diagram of the structure of the fitting unit in a clothing changing model in one embodiment;

[0042] Figure 5 This is a schematic diagram of the residual block in one embodiment;

[0043] Figure 6 This is a structural schematic diagram of a clothing-changing model in one embodiment;

[0044] Figure 7This is a structural block diagram of a virtual changing device in one embodiment;

[0045] Figure 8 This is an internal structural diagram of a computer device in one embodiment;

[0046] Figure 9 This is a diagram of the internal structure of a computer device in another embodiment. Detailed Implementation

[0047] To make the objectives, technical solutions, and advantages of this application clearer, the following detailed description is provided in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative and not intended to limit the scope of this application.

[0048] In some embodiments, such as Figure 1 As shown, a virtual clothing changing method is provided. Taking the application of this method to a computer device as an example, it can be understood that the computer device includes at least one of a terminal or a server. The method includes the following steps:

[0049] S102, determine multiple groups of samples.

[0050] Each sample set includes target clothing data and object data for changing clothes; the object data for changing clothes is extracted from the original clothing image of the presented object wearing the original clothing; the original clothing and the target clothing represented by the target clothing data are of the same category.

[0051] For example, a computer device can search for original images of an object wearing the same type of clothing, based on the category of the target garment. The computer device can use the target garment and the original images as a set of training data. It is understood that since the virtual clothing changing method provided in this embodiment uses unpaired data training—that is, there are no matching garments and corresponding images of people wearing them—it is necessary to rely on the type of garment to find original images of people wearing the same type of clothing and randomly select one of these original images to form a set of training data.

[0052] The computer equipment can acquire target clothing data of the target clothing and extract clothing object data from the original clothing image for each set of training data, thus obtaining each set of samples.

[0053] In some embodiments, the computer device may acquire a clothing-class label for a target garment to determine the category of the garment. The clothing-class label may include at least one of the following garment categories: long-sleeved, short-sleeved, dress, skirt, trousers, or shorts.

[0054] In some embodiments, the target clothing data may include target clothing keypoint data (cloth-points), a target clothing image, a target clothing mask (cloth-mask), and a target clothing classification label. The target clothing keypoint data is used to characterize the positions of individual keypoints on the initial target clothing. The target clothing keypoint data may be the keypoint coordinates of the target clothing. The target clothing mask is a mask corresponding to the target clothing image. The target clothing classification label is used to indicate the category of the target clothing.

[0055] In some embodiments, the clothing object data also includes an original clothing image and a corresponding original image mask. It is understood that the original clothing image presents the original clothing worn on the object and deformed.

[0056] In some embodiments, the clothing object data may include the original clothing image, the original clothing's classification label (person-cloth-class), object texture data (densepose-uv), object skeleton data (pose), original keypoint data (person-cloth-points), object region image (agno-person), region mask data (agno-parse), and object part segmentation mask (human parse). The original clothing's classification label indicates the category of the original clothing. The object texture data represents the object's pose in texture space. The object texture data can be the object's texture map coordinates. The object skeleton data refers to the object's central axis, preserving the object's shape and structure information, and reflecting the object's shape and structure. The original keypoint data represents the keypoint positions of the original clothing when the object wears it. The original keypoint data can be the keypoint coordinates of the original clothing. The region mask data represents the mask of the remaining region after removing the areas affecting the wearing of the target clothing. The region mask data corresponds to the object region image. The object region image is the image of the remaining region after removing the areas affecting the wearing of the target clothing. It is understandable that when both the target garment and the original garment are classified as short-sleeved, the object region image can be the image of the remaining area after removing the original garment, arm, and neck regions. A component segmentation mask refers to a mask of multiple semantically consistent regions on an object. A component segmentation mask can include at least one mask such as for hair, face, legs, arms, or clothing.

[0057] In some embodiments, the terminal may be, but is not limited to, various personal computers, laptops, smartphones, tablets, IoT devices, and portable wearable devices. IoT devices may include smart speakers, smart TVs, smart air conditioners, smart in-vehicle devices, etc. Portable wearable devices may include smartwatches, smart bracelets, head-mounted devices, etc. The server may be implemented using a standalone server or a server cluster consisting of multiple servers.

[0058] S104, learn the positional relationship between the key points of the object and the key points of the original clothing from the data of the object to be trained through the clothing-changing model.

[0059] Within the same category of clothing, keypoints must have semantic consistency. Semantic consistency here means that keypoints at corresponding locations within the same category of clothing represent the same meaning; for example, they might all represent a collar or all represent a sleeve. The number of keypoints must also be consistent across clothing categories. It's understood that different categories of clothing may have different numbers and locations of keypoints. Clothing keypoints can be used to locate various areas of the garment. For example, for short-sleeved clothing, there might be 13 keypoints. The left and right sleeves each have 4 keypoints, the hem of the sleeve has 2 keypoints, and the connection between the sleeve and the torso has 2 keypoints. The collar has 3 keypoints, located at the center and on both sides. The hem of short-sleeved clothing has 2 keypoints.

[0060] For example, the computer device can use the samples as input to a clothing-changing model to be trained. The computer device can then extract skeletal keypoint data from the object's skeletal data using the training model. This skeletal keypoint data represents key points on the object that are associated with wearing the original clothing. The computer device can learn the positional relationship between the object's keypoints and the original clothing's keypoints by associating the skeletal keypoint data with the original clothing's keypoint data.

[0061] In some embodiments, the object can be a human mannequin. The pose of the object remains unchanged during the reasoning process of the clothing-changing model.

[0062] In some embodiments, object skeleton data can be object skeleton coordinates. Original keypoint data can be the keypoint coordinates of the original garment. A computer device can use the attention mechanism in the clothing-changing model to associate the object skeleton coordinates with the keypoint coordinates of the original garment to obtain the positional relationship between the object's skeleton keypoints and the keypoints of the original garment when the object wears the original garment. It is understood that the object's keypoints include skeleton keypoints. When wearing clothing of the same category, the positional relationship between the object's keypoints and the keypoints of the worn garment is consistent; that is, the positional relationship when wearing clothing of the same category is consistent to a certain extent.

[0063] In some embodiments, object skeleton data may be data in the form of a heatmap.

[0064] S106, Based on the positional relationship, perform virtual clothing changing processing on the clothing object data and the target clothing data to obtain the clothing changing image output by the clothing changing model to be trained.

[0065] Among them, the clothing change image shows the effect of the object wearing the target clothing.

[0066] For example, a computer device can adjust target clothing data based on positional relationships to obtain deformable clothing data. The relationship between the key points of the deformable clothing, as represented by the deformable clothing data, and the key points of the object's skeleton conforms to the aforementioned positional relationships. The computer device can perform virtual clothing-changing processing on the deformable clothing data and the clothing-changing object data to obtain a clothing-changing image.

[0067] S108: Determine the target loss based on the difference between the clothing image and the original clothing image, and optimize the clothing model to be trained based on the target loss to obtain the trained clothing model.

[0068] For example, the target loss may include a clothing image discriminator loss. The computer device can use the clothing image and the original clothing image as input to the discriminator to obtain the clothing image discriminator loss as the output of the discriminator. It can be understood that the discriminator is trained based on real clothing images. The discriminator can determine whether the input image is a real clothing image. The computer device can optimize the untrained clothing model in the direction of reducing the target loss to obtain a trained clothing model.

[0069] In some embodiments, the target loss may include a clothing-changing image perceptual loss. A computer device can calculate the perceptual loss for the clothing-changing image and the original image to obtain the clothing-changing image perceptual loss. It is understood that the calculation of the perceptual loss can be implemented using a perceptual loss function based on a Visual Geometry Group Network (VGG).

[0070] In some embodiments, the computer device can determine clothing data to be processed and object data to be processed. Object data to be processed is extracted from images showing an object wearing clothing similar to the clothing to be processed. Clothing data to be processed is used to characterize the clothing to be processed. The computer device can use the clothing data to be processed and the object data to be processed as input to a trained clothing-changing model to obtain the target clothing-changing image output by the model. It can be understood that in real-world scenarios, when a user selects clothing for virtual try-on, only an image of the user wearing similar clothing is needed to achieve virtual dress-up.

[0071] In the aforementioned virtual clothing-swapping method, the target clothing data and the clothing-swapping object data in each set of samples are used as training data for the clothing-swapping model to be trained. The clothing-swapping object data is extracted from the original clothing image, while the target clothing data represents a different type of clothing than the original clothing in the original clothing image. The target clothing and the original clothing do not need to be completely identical, as long as they belong to the same category. The clothing-swapping model to be trained learns the positional relationship between the key points of the object and the key points of the original clothing from the clothing-swapping object data. The key points of clothing of the same category have semantic consistency. This positional relationship can reflect the relationship between the key points of the target clothing and the key points of the object to a certain extent. Based on the positional relationship, virtual clothing-swapping processing can be achieved using the clothing-swapping object data and the target clothing data, resulting in an accurate clothing-swapping image. The target loss is determined based on the difference between the clothing-swapping image and the original image. The clothing-swapping model to be trained is then optimized based on the target loss to obtain the trained clothing-swapping model. Since the target clothing and the original clothing only need to belong to the same category, they do not need to be completely identical, which greatly improves the utilization rate of data resources.

[0072] In some embodiments, the clothing object data includes object skeleton data and original keypoint data; the original keypoint data is used to characterize the keypoints of the original garment; the positional relationship between the keypoints of the object and the keypoints of the original garment is learned from the clothing object data by a clothing-changing model to be trained, including: extracting skeleton keypoint data from the object skeleton data based on a gating mechanism in the clothing-changing model to be trained; and learning the positional relationship between the skeleton keypoints of the object and the keypoints of the original garment by associating the skeleton keypoint data with the original keypoint data.

[0073] For example, a computer device can extract skeletal keypoint data from the object's skeleton data based on a gating mechanism in the clothing-changing model to be trained. It is understood that when an object wears clothing, the clothing often does not cover the entire object; there are areas on the object related to the clothing and areas unrelated to it. The gating mechanism in the clothing-changing model is used to restrict the unrelated areas and retain the areas that are clothing-friendly. The computer device can then fuse the skeletal keypoint data and the original clothing keypoint data to obtain positional association data. This positional association data is used to characterize the positional relationship between the skeletal keypoints and the keypoints of the original clothing.

[0074] In some embodiments, the computer device can perform preliminary fusion processing on the skeleton keypoint data and the original keypoint data to obtain preliminary fused data. The computer device can then perform advanced fusion processing on the preliminary fused data and the original keypoint data to obtain location-related data.

[0075] In some embodiments, the computer device can encode the skeleton keypoint data to obtain encoded skeleton keypoint data. The computer device can then extract skeleton keypoint data from the encoded skeleton keypoint data.

[0076] In some embodiments, the computer device can perform multi-level encoding processing on the original keypoint data to obtain encoded original keypoint data. For example, the multi-level encoding processing can be second-level encoding processing.

[0077] In some embodiments, the computer device can encode the skeleton keypoint data to obtain encoded skeleton keypoint data. The computer device can then perform a matrix multiplication operation between the encoded skeleton keypoint data and the encoded original keypoint data to obtain preliminary fused data. Finally, the computer device can perform a matrix multiplication operation between the preliminary fused data and the encoded original keypoint data to obtain location association data.

[0078] In some embodiments, the encoding process may involve convolution followed by instance regularization and finally nonlinear activation (conv-IN-prelu).

[0079] In this embodiment, the skeleton key point data is extracted from the object skeleton data based on the gating mechanism in the clothing-changing model to be trained; by associating the skeleton key point data with the original key point data, the positional relationship between the skeleton key points of the object and the key points of the original clothing is learned. The positional relationship can reflect the relationship between the key points of the target clothing and the key points of the object to a certain extent. Based on the positional relationship, the virtual clothing changing of the object and the target clothing can be realized, and an accurate clothing changing image can be obtained.

[0080] In some embodiments, the clothing object data includes object texture data and region mask data; the region mask data refers to the mask of the remaining area after removing the area that affects the wearing of the target clothing; extracting skeleton key point data from the object skeleton data based on the gating mechanism in the clothing-changing model to be trained includes: fusing the object texture data and region mask data to obtain object fused data; extracting wearing region data from the object fused data according to the gating data in the clothing-changing model to be trained; the wearing region data is used to indicate the wearing area when the object wears the target clothing; and extracting skeleton key point data from the object skeleton data according to the wearing region data.

[0081] For example, a computer device can concatenate object texture data (buv) and region mask data (bnc) to obtain object fusion data (bnc|buv). It can be understood that the region mask data lacks regions related to the clothing being worn, while the object texture data is complete; concatenating the two can serve a complementary purpose. The computer device can encode the object fusion data to obtain encoded object fusion data. The computer device can extract wearing region data from the encoded object fusion data using the first gating data in the clothing-changing model to be trained. The computer device can then perform a weighted mapping on the encoded object skeleton data using the wearing region data to obtain skeleton keypoint data. It can be understood that the wearing region data is essentially the attention weights of the encoded object skeleton data.

[0082] In some embodiments, a computer device may perform element-wise multiplication on wearable area data and encoded object skeleton data to obtain skeleton key point data.

[0083] In this embodiment, object texture data and region mask data are fused to obtain object fused data; wearing region data is extracted from the object fused data based on the gating data in the clothing-changing model to be trained; skeleton key point data is extracted from the object skeleton data based on the wearing region data, and then the skeleton key point data is associated with the original key point data to extract the positional relationship. The positional relationship can reflect the relationship between the key points of the target clothing and the key points of the object when the object wears the target clothing, and thus an accurate clothing-changing image can be obtained based on the positional relationship.

[0084] In some embodiments, virtual clothing swapping is performed on the clothing object data and target clothing data based on positional relationships to obtain a clothing swapping image output by the clothing swapping model to be trained. This includes: mapping key points of the target clothing according to positional relationships to obtain target key point data of the target clothing; using the target key point data to represent the position of the key points of the target clothing when the object wears the target clothing; adjusting the target clothing data based on the target key point data to obtain deformable clothing data; matching the deformable clothing represented by the deformable clothing data with the pose of the object; and performing virtual clothing swapping on the deformable clothing data and the clothing object data to obtain a clothing swapping image output by the clothing swapping model to be trained.

[0085] For example, a computer device can perform weighted mapping on key point data of a target garment using location-related data to obtain weighted key point data. The computer device can then overlay the weighted key point data and the location-related data to obtain the target key point data of the target garment.

[0086] Computer equipment can fuse target key point data and target clothing data to obtain deformable clothing data. It can be understood that during the fusion process, the target key point data guides the adjustment direction of the target clothing data. Computer equipment can then fuse deformable clothing data and clothing-changing object data to obtain a clothing-changing image.

[0087] In some embodiments, the computer device may perform element-wise multiplication on the location association data and the target clothing key point data to obtain weighted key point data.

[0088] In some embodiments, a computer device may perform element-wise addition on the weighted key point data and location association data to obtain the target key point data of the target garment.

[0089] In some embodiments, the computer device can encode the location-related data to obtain encoded location-related data. The computer device can then use the encoded location-related data to perform a weighted mapping on the target garment's keypoint data to obtain weighted keypoint data. The computer device can then overlay the weighted keypoint data and the encoded location-related data before performing further encoding to obtain the target garment's target keypoint data.

[0090] In some embodiments, a computer device can perform multi-resolution fusion processing on deformed clothing data and clothing object data to obtain a clothing image.

[0091] In this embodiment, the key points of the target garment are mapped according to the positional relationship to obtain the target key point data of the target garment; the target garment data is adjusted based on the target key point data to obtain deformable garment data; virtual garment changing processing is performed on the deformable garment data and the changing object data to obtain the changing image output by the changing model to be trained; the target garment is guided to deform through the positional relationship, and then the deformed garment is worn on the object to obtain an accurate changing image.

[0092] In some embodiments, such as Figure 2The diagram shows the structure of the keypoint prediction unit in the clothing-changing model. The output of the keypoint prediction unit includes target keypoint data and predicted component segmentation masks. The computer device can further extract advanced region data from the wearing region data using the second gating data in the clothing-changing model to be trained. The advanced region data is more detailed than the wearing region data, and the area represented by the wearing region data is larger than that represented by the advanced region data. To avoid incomplete extraction of positional relationships, the wearing region data is used. The computer device can encode the advanced region data to obtain encoded advanced region data. It should be noted that the blurring of the eyes of the task object in the image in this embodiment is for privacy protection and is not part of the processing described in this application.

[0093] Computer equipment can perform weighted mapping on the target clothing mask using advanced region data to obtain a weighted target clothing mask. Specifically, the computer equipment can encode the target clothing mask to obtain an encoded target clothing mask. The computer equipment can then perform element-wise multiplication on the advanced region data and the encoded target clothing mask to obtain the weighted target clothing mask.

[0094] The computer equipment can fuse target keypoint data and a weighted target clothing mask to obtain preliminary clothing fusion data. This can be understood as the computer equipment performing further encoding on the weighted target clothing mask to obtain a further encoded target clothing mask. The computer equipment can then perform a dimensionality transformation (1) on the further encoded target clothing mask and the target keypoint data to match their dimensions, resulting in a dimensionally transformed target clothing mask and target keypoint data. Finally, the computer equipment can perform matrix multiplication on the dimensionally transformed target clothing mask and target keypoint data to obtain preliminary clothing fusion data.

[0095] Computer equipment can perform advanced fusion processing on the initial clothing fusion data and the weighted target clothing mask to obtain clothing fusion data. Specifically, the computer equipment can perform a dimensionality transformation on the initial clothing fusion data, making the dimensions of the dimensionally transformed initial clothing fusion data match the dimensions of the advanced-encoded target clothing mask. The computer equipment can also perform matrix multiplication on the dimensionally transformed initial clothing fusion data and the advanced-encoded target clothing mask to obtain the clothing fusion data.

[0096] Computer equipment can perform weighted mapping on advanced region data using clothing fusion data to obtain weighted advanced region data. Specifically, the computer equipment can encode the clothing fusion data to obtain encoded clothing fusion data. The computer equipment can then perform element-wise multiplication on the encoded clothing fusion data and the encoded advanced region data to obtain weighted advanced region data.

[0097] Computer equipment can overlay clothing fusion data and weighted advanced region data to obtain a predicted component segmentation mask. Specifically, the computer equipment can perform element-wise addition on the encoded clothing fusion data and weighted advanced region data to obtain the predicted component segmentation mask. This predicted component segmentation mask is the component segmentation mask when the object is wearing the target clothing.

[0098] Computer equipment can perform dimensional transformation on the encoded skeleton keypoint data and the encoded original keypoint data, followed by matrix multiplication to obtain preliminary fused data. The computer equipment can then perform dimensional transformation on the preliminary fused data to obtain dimensionally transformed preliminary fused data. Finally, the computer equipment can perform matrix multiplication on the dimensionally transformed preliminary fused data and the encoded original keypoint data to obtain location-related data.

[0099] In some embodiments, the clothing-changing model may include multiple cascaded keypoint prediction units. The output of each keypoint prediction unit serves as the input to the next keypoint prediction unit.

[0100] In some embodiments, adjusting the target clothing data based on the target key point data to obtain deformed clothing data includes: extracting features from the target key point data and the target clothing data to obtain deformed feature data; and deforming the target clothing image in the target clothing data according to the deformed feature data to obtain deformed clothing data.

[0101] For example, a computer device can encode target keypoint data, predicted component segmentation masks, and object skeleton data to obtain object encoded data. The computer device can encode a target clothing mask and a target clothing image to obtain target clothing encoded data. The computer device can fuse the object encoded data and the target clothing encoded data to obtain deformation feature data. The computer device can interpolate the target clothing image based on the deformation feature data to obtain deformed clothing data. The deformed clothing data may include deformed clothing images and deformed clothing masks.

[0102] In some embodiments, the computer device can perform preliminary fusion processing on the object coding data and the target garment coding data to obtain initial deformation feature data. The computer device can further fuse the object coding data and the initial deformation feature data to obtain advanced deformation feature data. The computer device can then perform weighted mapping processing on the target garment coding data using the advanced deformation feature data to obtain weighted target garment coding data. The computer device can also overlay the advanced deformation feature data and the weighted target garment coding data to obtain deformation feature data. Finally, the computer device can perform bilinear interpolation processing on the target garment image based on the deformation feature data to obtain deformed garment data. This bilinear interpolation processing can be implemented using grid sampling.

[0103] In some embodiments, such as Figure 3 The diagram shows the structural schematic of the clothing deformation unit in the clothing-changing model. The computer device can use target keypoint data, predicted component segmentation masks, and object skeleton data as inputs to a multi-layer encoder, and encode the output of the multi-layer encoder to obtain object-coded data. The computer device can also use the target clothing mask and target clothing image as inputs to a multi-layer encoder, and encode the output of the multi-layer encoder to obtain target clothing-coded data. The multi-layer encoder is used to extract features from high to low resolution; as the encoding level increases, the resolution decreases, and the corresponding features become more dimensional.

[0104] The computer equipment can perform a dimensional transformation (1) on the target garment coding data and the object coding data, and then perform matrix multiplication on the transformed target garment coding data and object coding data to obtain initial deformation feature data. The computer equipment can perform a dimensional transformation (2) on the initial deformation feature data, and then perform matrix multiplication on the transformed initial deformation feature data and object coding data to obtain advanced deformation feature data. The computer equipment can encode the advanced deformation feature data to obtain encoded advanced deformation feature data. The computer equipment can perform element-wise multiplication on the encoded advanced deformation feature data and the target garment coding data to obtain weighted target garment coding data. The computer equipment can perform element-wise addition on the weighted target garment coding data and the encoded advanced deformation feature data to obtain deformation feature data. The computer equipment can then perform bilinear interpolation on the target garment image based on the deformation feature data to obtain deformed garment data.

[0105] In some embodiments, a computer device can perform bilinear interpolation on a target garment image and a target garment mask based on deformation feature data to obtain deformed garment data.

[0106] In some embodiments, deformation feature data is flow feature.

[0107] In this embodiment, feature extraction is performed on the target key point data and the target clothing data to obtain deformation feature data; the target clothing image in the target clothing data is deformed according to the deformation feature data to obtain deformed clothing data; the deformation feature data guides the deformation of the target clothing image so that the deformed clothing data matches the shape of the target clothing when the object wears the target clothing, thereby achieving accurate virtual clothing changing.

[0108] In some embodiments, the target loss includes a first loss; the first loss is determined based on the difference between the clothing image and the original clothing image; the target loss also includes at least one of a second loss or a third loss; the second loss is determined based on the difference between the target clothing key point data and the original clothing key point data; the third loss is determined based on the difference between the deformed clothing data and the original clothing data in the clothing object data.

[0109] For example, the first loss may include a clothing-changing image perception loss and a clothing-changing image discriminator loss. The computer device can input the clothing-changing image and the original clothing image into the first discriminator to obtain the clothing-changing image discriminator loss output by the first discriminator. The first loss is used to evaluate the performance of the try-on unit.

[0110] The computer device can input the predicted component segmentation mask, the object's component segmentation mask, and the object's skeleton data into a second discriminator, obtaining the segmentation mask discriminator loss output by the second discriminator. The second discriminator is trained based on real component segmentation masks and can determine whether the input component segmentation mask is a real component segmentation mask. The computer device can calculate the mean absolute error (MAO) of the target garment's keypoint data and the original garment's keypoint data, obtaining the keypoint loss. The computer device can calculate the perceptual loss of the predicted component segmentation mask and the object's component segmentation mask, obtaining the segmentation mask perceptual loss. The second loss includes the segmentation mask discriminator loss, the keypoint loss, and the segmentation mask perceptual loss. The second loss is used to evaluate the performance of the keypoint prediction unit.

[0111] The computer device can input the deformable clothing image and deformable clothing mask into a third discriminator, obtaining the deformable clothing discriminator loss output by the third discriminator. The third discriminator is trained based on real worn clothing and can determine whether the input deformable clothing is genuine. The computer device can calculate the perceptual loss on the deformable clothing mask and the original clothing mask, obtaining the deformable clothing perceptual loss. The computer device can perform classification prediction on the deformable clothing image and the original clothing image, obtaining the deformable clothing category loss. It can be understood that the deformable clothing and the original clothing should belong to the same category. The computer device can calculate the adjacent gradient smoothing loss on the deformable feature data obtained in each training round, obtaining the deformable feature loss. The third loss includes the deformable clothing discriminator loss, the deformable clothing perceptual loss, the deformable clothing category loss, and the deformable feature loss.

[0112] The computer equipment can perform a weighted summation of the first loss, the second loss, and the third loss to obtain the target loss, and optimize the clothing-changing model to be trained in the direction of reducing the target loss, thus obtaining the trained clothing-changing model.

[0113] In this embodiment, by supervising the optimization of the clothing-changing model to be trained through multiple loss methods, the accuracy of the clothing-changing model training process can be guaranteed, and the performance of the trained clothing-changing model can be guaranteed.

[0114] In some embodiments, virtual clothing swapping is performed on deformable clothing data and clothing swapping object data to obtain a clothing swapping image output by the clothing swapping model to be trained. This includes: performing a first resolution fusion process on the deformable clothing data and clothing swapping object data to obtain a clothing swapping result corresponding to the first resolution; starting from the next resolution after the first resolution, the current resolution is determined sequentially, and the deformable clothing data, clothing swapping object data, and the clothing swapping result corresponding to the previous resolution are fused to obtain a clothing swapping result corresponding to the current resolution; and a nonlinear mapping process is performed on the clothing swapping result corresponding to the last resolution to obtain a clothing swapping image.

[0115] For example, the computer device can encode the object region image, region mask data, object skeleton data, object texture data, and deformed clothing data to obtain first encoded data. The computer device can then perform a first-resolution fusion process on the target keypoint data, the predicted component segmentation mask, the first encoded data, the object region image, the region mask data, the object skeleton data, the object texture data, and the deformed clothing data to obtain a clothing-changing result corresponding to the first resolution. Starting from the next resolution after the first resolution, the computer device can sequentially determine the current resolution and perform a fusion process on the target keypoint data, the predicted component segmentation mask, the object region image, the region mask data, the object skeleton data, the object texture data, the deformed clothing data, and the clothing-changing result corresponding to the previous resolution to obtain a clothing-changing result corresponding to the current resolution. Finally, the computer device can encode the clothing-changing result corresponding to the last resolution and then perform non-linear activation to obtain a clothing-changing image. Each resolution is smaller than the next resolution.

[0116] In some embodiments, such as Figure 4 The diagram shows the structure of the try-on unit in the clothing-changing model. The computer device can use the first encoded data, deformable clothing data, and clothing object data as inputs to the residual block at the first resolution, obtaining the clothing-changing result corresponding to the first resolution output by the residual block. Starting from the next resolution, the current resolution is determined sequentially, and the deformable clothing data, clothing object data, and the clothing-changing result corresponding to the previous resolution are used as inputs to the residual block at the current resolution, obtaining the clothing-changing result corresponding to the current resolution output by the residual block. The computer device can encode the clothing-changing result corresponding to the last resolution output by the residual block at the last resolution and then perform nonlinear activation to obtain the clothing-changing image. The nonlinear activation function can be the hyperbolic tangent function.

[0117] In one embodiment, such as Figure 5 The diagram shows the structure of the residual block. Since each residual block has an output at a different resolution, the residual block at the current resolution can be upsampled using the clothing-changing result corresponding to the previous resolution. The upsampled clothing-changing result conforms to the current resolution. The computer device can perform initial convolution processing on the target keypoint data and the predicted component segmentation mask, followed by further convolution processing, to obtain the second encoded data. The residual block at the current resolution can merge the object region image, region mask data, object skeleton data, object texture data, deformed clothing data, and the upsampled clothing-changing result to obtain merged data. The computer device can perform element-wise multiplication on the merged data and the second encoded data to obtain the initial clothing-changing result. The computer device can perform element-wise addition on the second encoded data and the initial clothing-changing result to obtain the final clothing-changing result.

[0118] In some embodiments, such as Figure 6 The diagram shows the structure of the clothing-changing model. Both the target and original garments can be short-sleeved. The arrows indicate the outputs of each unit. The clothing-changing model includes a keypoint prediction unit, a clothing deformation unit, and a try-on unit. The keypoint prediction unit's inputs include target clothing keypoint data, a target clothing mask, region mask data, object skeleton data, object texture data, and original keypoint data. The keypoint prediction unit's output includes the predicted part segmentation mask and the target keypoint data. The clothing deformation unit's inputs include the predicted part segmentation mask, target keypoint data, object skeleton data, the target clothing mask, and the target clothing image. The clothing deformation unit's output includes the deformed clothing mask and the deformed clothing image. The try-on unit's inputs include the deformed clothing mask, the deformed clothing image, the object region image, region mask data, object skeleton data, and object texture data. The try-on unit's output includes the changed clothing image.

[0119] In this embodiment, the deformable clothing data and the clothing object data are fused at a first resolution to obtain the clothing change result corresponding to the first resolution. Starting from the next resolution after the first resolution, the current resolution is determined sequentially, and the deformable clothing data, the clothing object data, and the clothing change result corresponding to the previous resolution are fused to obtain the clothing change result corresponding to the current resolution. The clothing change result corresponding to the last resolution is subjected to nonlinear mapping processing to obtain the clothing change image. Through multi-resolution fusion processing and combining the features of multiple resolutions, an accurate clothing change image is obtained.

[0120] It should be understood that although the steps in the flowcharts of the embodiments described above are shown sequentially according to the arrows, these steps are not necessarily executed in the order indicated by the arrows. Unless explicitly stated herein, there is no strict order restriction on the execution of these steps, and they can be executed in other orders. Moreover, at least some steps in the flowcharts of the embodiments described above may include multiple steps or multiple stages. These steps or stages are not necessarily completed at the same time, but can be executed at different times. The execution order of these steps or stages is not necessarily sequential, but can be performed alternately or in turn with other steps or at least some of the steps or stages of other steps.

[0121] Based on the same inventive concept, this application also provides a virtual clothing changing device for implementing the virtual clothing changing method described above. The solution provided by this device is similar to the solution described in the above method; therefore, the specific limitations in one or more virtual clothing changing device embodiments provided below can be found in the limitations of the virtual clothing changing method described above, and will not be repeated here.

[0122] In some embodiments, such as Figure 7 As shown, a virtual clothing changing device 700 is provided, including: a determination module 702, a learning module 704, an output module 706, and an optimization module 708, wherein:

[0123] The determination module 702 is used to determine multiple sets of samples; each set of samples includes target clothing data and clothing swapping object data; the clothing swapping object data is extracted from the original clothing image of the presented object wearing the original clothing; the original clothing and the target clothing represented by the target clothing data are consistent in category.

[0124] Learning module 704 is used to learn the positional relationship between the key points of the object and the key points of the original clothing from the clothing object data through the clothing-changing model to be trained; the key points of clothing in the same category have semantic consistency.

[0125] The output module 706 is used to perform virtual clothing swapping processing on the clothing swapping object data and the target clothing data based on the positional relationship, so as to obtain the clothing swapping image output by the clothing swapping model to be trained; the clothing swapping image presents the effect of the object wearing the target clothing.

[0126] The optimization module 708 is used to determine the target loss based on the difference between the clothing image and the original clothing image, and to optimize the clothing model to be trained based on the target loss to obtain the trained clothing model.

[0127] In some embodiments, the clothing object data includes object skeleton data and original key point data; the original key point data is used to characterize the key points of the original garment; the learning module 704 is used to extract skeleton key point data from the object skeleton data based on the gating mechanism in the clothing change model to be trained; by associating the skeleton key point data with the original key point data, the positional relationship between the skeleton key points of the object and the key points of the original garment is learned.

[0128] In some embodiments, the clothing object data includes object texture data and region mask data; the region mask data refers to the mask of the remaining area after removing the area that affects the wearing of the target clothing; the learning module 704 is used to fuse the object texture data and the region mask data to obtain object fusion data; the wearing region data is extracted from the object fusion data according to the gating data in the clothing model to be trained; the wearing region data is used to indicate the wearing area when the object wears the target clothing; the skeleton key point data is extracted from the object skeleton data according to the wearing region data.

[0129] In some embodiments, the output module 706 is used to map the key points of the target garment according to the positional relationship to obtain the target key point data of the target garment; the target key point data is used to represent the position of the key points of the target garment when the object wears the target garment; the target garment data is adjusted based on the target key point data to obtain deformable garment data; the deformable garment represented by the deformable garment data matches the pose of the object; virtual garment changing processing is performed on the deformable garment data and the changing object data to obtain the changing image output by the changing model to be trained.

[0130] In some embodiments, the output module 706 is used to extract features from the target key point data and the target clothing data to obtain deformation feature data; and to deform the target clothing image in the target clothing data according to the deformation feature data to obtain deformed clothing data.

[0131] In some embodiments, the output module 706 is used to perform a first resolution fusion processing on the deformed clothing data and the clothing object data to obtain the clothing change result corresponding to the first resolution.

[0132] Starting from the next resolution after the first resolution, the current resolution is determined sequentially. The deformed clothing data, the clothing object data, and the clothing change results corresponding to the previous resolution are fused together to obtain the clothing change results corresponding to the current resolution.

[0133] The clothing change result corresponding to the last resolution is subjected to nonlinear mapping processing to obtain the clothing change image.

[0134] In some embodiments, the target loss includes a first loss; the first loss is determined based on the difference between the clothing image and the original clothing image; the target loss also includes at least one of a second loss or a third loss; the second loss is determined based on the difference between the target clothing key point data and the original clothing key point data; the third loss is determined based on the difference between the deformed clothing data and the original clothing data in the clothing object data.

[0135] Each module in the aforementioned virtual changing device can be implemented entirely or partially through software, hardware, or a combination thereof. These modules can be embedded in the processor of a computer device in hardware form or independent of it, or stored in the memory of the computer device in software form, so that the processor can call and execute the operations corresponding to each module.

[0136] In some embodiments, a computer device is provided, which may be a server, and its internal structure diagram may be as follows: Figure 8As shown, this computer device includes a processor, memory, input / output (I / O) interfaces, and a communication interface. The processor, memory, and I / O interfaces are connected via a system bus, and the communication interface is also connected to the system bus via the I / O interfaces. The processor provides computational and control capabilities. The memory includes non-volatile storage media and internal memory. The non-volatile storage media stores the operating system, computer programs, and a database. The internal memory provides the environment for the operating system and computer programs stored in the non-volatile storage media. The database stores multiple sets of samples. The I / O interfaces are used for exchanging information between the processor and external devices. The communication interface is used for communicating with external terminals via a network connection. When executed by the processor, the computer program implements a virtual clothing changing method.

[0137] In some embodiments, a computer device is provided, which may be a terminal, and its internal structure diagram may be as follows: Figure 9 As shown, the computer device includes a processor, memory, input / output interfaces, a communication interface, a display unit, and an input device. The processor, memory, and input / output interfaces are connected via a system bus, and the communication interface, display unit, and input device are also connected to the system bus via the input / output interfaces. The processor provides computing and control capabilities. The memory includes non-volatile storage media and internal memory. The non-volatile storage media stores the operating system and computer programs. The internal memory provides an environment for the operation of the operating system and computer programs stored in the non-volatile storage media. The input / output interfaces are used for exchanging information between the processor and external devices. The communication interface is used for wired or wireless communication with external terminals; wireless communication can be achieved through Wi-Fi, mobile cellular networks, NFC (Near Field Communication), or other technologies. When the computer program is executed by the processor, it implements a virtual clothing changing method. The display unit is used to form a visually visible image and can be a display screen, a projection device, or a virtual reality imaging device. The display screen can be an LCD screen or an e-ink screen. The input device of the computer device can be a touch layer covering the display screen, or buttons, trackballs, or touchpads set on the casing of the computer device, or external keyboards, touchpads, or mice, etc.

[0138] Those skilled in the art will understand that Figure 8 or Figure 9The structure shown is merely a block diagram of a portion of the structure related to the present application and does not constitute a limitation on the computer device to which the present application is applied. Specific computer devices may include more or fewer components than those shown in the figure, or combine certain components, or have different component arrangements.

[0139] In some embodiments, a computer device is provided, including a memory and a processor, wherein the memory stores a computer program, and the processor executes the computer program to implement the steps in the above-described method embodiments.

[0140] In some embodiments, a computer-readable storage medium is provided having a computer program stored thereon, which, when executed by a processor, implements the steps in the above method embodiments.

[0141] In some embodiments, a computer program product is provided, including a computer program that, when executed by a processor, implements the steps in the above method embodiments.

[0142] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, data stored, data displayed, etc.) involved in this application are all information and data authorized by the user or fully authorized by all parties, and the collection, use and processing of the relevant data shall comply with the relevant laws, regulations and standards of the relevant countries and regions.

[0143] Those skilled in the art will understand that all or part of the processes in the methods of the above embodiments can be implemented by a computer program instructing related hardware. The computer program can be stored in a non-volatile computer-readable storage medium, and when executed, it can include the processes of the embodiments of the above methods. Any references to memory, databases, or other media used in the embodiments provided in this application can include at least one of non-volatile and volatile memory. Non-volatile memory can include read-only memory (ROM), magnetic tape, floppy disk, flash memory, optical memory, high-density embedded non-volatile memory, resistive random access memory (ReRAM), magnetic random access memory (MRAM), ferroelectric random access memory (FRAM), phase change memory (PCM), graphene memory, etc. Volatile memory can include random access memory (RAM) or external cache memory, etc. By way of illustration and not limitation, RAM can take many forms, such as Static Random Access Memory (SRAM) or Dynamic Random Access Memory (DRAM). The databases involved in the embodiments provided in this application may include at least one type of relational database and non-relational database. Non-relational databases may include, but are not limited to, blockchain-based distributed databases. The processors involved in the embodiments provided in this application may be general-purpose processors, central processing units, graphics processing units, digital signal processors, programmable logic devices, quantum computing-based data processing logic devices, etc., and are not limited to these.

[0144] The technical features of the above embodiments can be combined in any way. For the sake of brevity, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.

[0145] The embodiments described above are merely illustrative of several implementation methods of this application, and while the descriptions are specific and detailed, they should not be construed as limiting the scope of this patent application. It should be noted that those skilled in the art can make various modifications and improvements without departing from the concept of this application, and these all fall within the protection scope of this application. Therefore, the protection scope of this application should be determined by the appended claims.

Claims

1. A virtual clothing changing method, characterized in that, The method includes: Multiple sets of samples are determined; each set of samples includes target clothing data and swapping object data; the swapping object data is extracted from the original clothing image of the presented object wearing the original clothing, and the swapping object data includes object skeleton data and original key point data; the original key point data is used to characterize the key points of the original clothing; the original clothing is consistent with the category of the target clothing characterized by the target clothing data; The method involves learning the positional relationship between key points of the object and key points of the original clothing from the object data using a clothing-changing model to be trained. This includes: extracting skeleton key point data from the object skeleton data based on a gating mechanism in the clothing-changing model to be trained; learning the positional relationship between the skeleton key points of the object and key points of the original clothing by associating the skeleton key point data with the original clothing key point data; and ensuring that the key points of clothing of the same category have consistent semantics. Based on the positional relationship, virtual clothing swapping is performed on the clothing object data and the target clothing data to obtain a clothing swapping image output by the clothing swapping model to be trained. This includes: mapping key points of the target clothing according to the positional relationship to obtain target key point data of the target clothing; the target key point data is used to characterize the position of the key points of the target clothing when the object wears the target clothing; adjusting the target clothing data based on the target key point data to obtain deformable clothing data; the deformable clothing data characterizing the deformed clothing matching the pose of the object; fusing the deformable clothing data and the clothing object data at a first resolution to obtain a clothing swapping result corresponding to the first resolution; starting from the next resolution after the first resolution, the current resolution is sequentially determined, and the deformable clothing data, the clothing object data, and the clothing swapping result corresponding to the previous resolution are fused to obtain a clothing swapping result corresponding to the current resolution; non-linear mapping is performed on the clothing swapping result corresponding to the last resolution to obtain a clothing swapping image; the clothing swapping image presents the effect of the object wearing the target clothing. The target loss is determined based on the difference between the clothing image and the original clothing image, and the clothing model to be trained is optimized based on the target loss to obtain the trained clothing model.

2. The method according to claim 1, characterized in that, The method further includes: Obtain the classification tags for the target garment to determine its category.

3. The virtual clothing changing method according to claim 1, characterized in that, The clothing object data includes object texture data and region mask data; the region mask data refers to the mask of the remaining area after removing the area that affects the wearing of the target clothing. The gating mechanism based on the clothing-changing model to be trained extracts skeletal key point data from the object skeleton data, including: The object texture data and the region mask data are fused to obtain object fused data; Wearing region data is extracted from the object fusion data based on the gating data in the clothing-changing model to be trained; the wearing region data is used to indicate the wearing region when the object wears the target clothing; Based on the wearable area data, the skeleton key point data is extracted from the object skeleton data.

4. The virtual clothing changing method according to claim 1, characterized in that, The object has areas related to wearing clothing and areas unrelated to wearing clothing. The gating mechanism in the clothing-changing model is used to restrict areas unrelated to wearing clothing.

5. The virtual clothing changing method according to claim 1, characterized in that, The process of adjusting the target clothing data based on the target key point data to obtain deformable clothing data includes: Feature extraction is performed on the target key point data and the target clothing data to obtain deformation feature data; The target clothing image in the target clothing data is deformed based on the deformation feature data to obtain deformed clothing data.

6. The virtual clothing changing method according to claim 5, characterized in that, The target loss includes a first loss; the first loss is determined based on the difference between the clothing swap image and the original clothing image; the target loss also includes at least one of a second loss or a third loss; the second loss is determined based on the difference between the target clothing key point data and the original clothing key point data; the third loss is determined based on the difference between the deformed clothing data and the original clothing data in the clothing swap object data.

7. The virtual clothing changing method according to claim 3, characterized in that, The method further includes: The object fusion data is encoded to obtain encoded object fusion data; The step of extracting wearing region data from the object fusion data based on the gating data in the clothing-changing model to be trained includes: Wearing region data is extracted from the encoded object fusion data based on the gating data in the clothing-changing model to be trained.

8. A virtual fitting device, characterized in that, The virtual fitting device uses a virtual clothing changing method as described in any one of claims 1 to 7.

9. A computer device comprising a memory and a processor, wherein the memory stores a computer program, characterized in that, When the processor executes the computer program, it implements the steps of the method according to any one of claims 1 to 7.

10. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by a processor, it implements the steps of the method according to any one of claims 1 to 7.

Citation Information

Patent Citations

  • Figure virtual clothes changing method, terminal equipment and storage medium

    CN113436058A

  • Virtual clothes changing method and system based on improved GRNet network

    CN113822986A