Image virtual fitting method with robust garment fitting and conditional skin inference

By employing a deformation-mapping-combination try-on strategy and a conditional skin inference method, the robustness problem of image-based virtual try-on in complex scenarios is solved, achieving a realistic virtual try-on effect suitable for online clothing retail try-on displays.

CN115471433BActive Publication Date: 2026-01-02ZHEJIANG UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202210979724.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-08-16
Publication Date
2026-01-02
Estimated Expiration
2042-08-16

AI Technical Summary

Technical Problem

Existing image-based virtual try-on technology suffers from low robustness in complex scenarios, resulting in blurry and unrealistic clothing try-on results and a lack of effective skin inference mechanisms, leading to poor virtual try-on performance and hindering commercial promotion.

Method used

We adopt a deformation-mapping-combination try-on strategy and a conditional skin inference method. We improve the robustness of the model by untangling the semantic category representation and improving cycle consistency. We use an instance-level segmentation inference module and a progressive clothing try-on module to achieve reasonable inference of the semantic layout of clothing try-on. We also perform skin content repair by adjusting the StyleGAN2 model.

Benefits of technology

It achieves robust and realistic virtual try-on effects in complex scenarios, enhancing the immersive experience of trying on clothes and the realism of clothing displays, making it suitable for try-on displays in online clothing retail.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115471433B_ABST
    Figure CN115471433B_ABST
Patent Text Reader

Abstract

The application discloses an image virtual fitting method with robust garment fitting and conditional skin inference. The application faces virtual fitting under complex scenes, proposes garment fitting strategy and conditional skin inference mechanism. The reasonable garment fitting semantic layout is inferred by using the disentangled representation of semantic categories and the introduction of cyclic consistency. With the inferred semantic layout as the shape guide, the garment fitting is realized through the deformation-mapping-combination strategy. To solve the universality of garment fitting, the static coverage-dynamic selection deformation strategy and the explicit task allocation based on region alignment are adopted. By adjusting the StyleGAN2 model, the skin content inference according to the target skin shape and the space-independent skin features is realized. The application is suitable for image virtual fitting, helps to solve the robustness problem of virtual fitting when fitting different types of garments and skin re-exposure, helps users to experience realistic virtual fitting effect, and helps to promote the commercial application of image virtual fitting.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application belongs to the field of image virtual fitting, and particularly relates to an image virtual fitting method with robust garment fitting and conditional skin inference. BACKGROUND

[0002] Image virtual fitting is a computer vision topic with great commercial potential. Image virtual fitting does not require professional three-dimensional modeling of the human body and clothing. It takes users and fitting clothing images as input and presents a realistic effect of trying on new clothes in the form of an image. However, there are still some challenging problems in promoting the commercial application of image virtual fitting. Due to the low robustness of image virtual fitting in complex scenes, there is an urgent need for a robust advanced garment fitting and skin inference mechanism to improve the authenticity of garment fitting results.

[0003] Existing image virtual fitting research lacks robust fitting semantic layout inference, and instance-level segmentation inference of the human body performs poorly in disentangling original clothing and arm exposure inference, resulting in changes in the style of the garment and blurred boundaries during fitting. Due to excessive reliance on single-scale deformation and non-explicit assignment of deformation inference tasks, the result image of the fitting cannot effectively inherit the texture and leads to unrealistic fitting. It lacks a dedicated skin inference mechanism and reduces clothing and skin inference to similar tasks, resulting in skin artifacts and color shifts.

[0004] In summary, existing technologies cannot solve the problem of low robustness of image virtual fitting in complex environments, and blurred and unrealistic virtual fitting effects hinder its further commercial promotion. Therefore, there is an urgent need for an image virtual fitting method with robust garment fitting and conditional skin inference. SUMMARY

[0005] The present application aims to overcome the shortcomings of the prior art and provide an image virtual fitting method with robust garment fitting and conditional skin inference. The present application is aimed at garment fitting for online clothing retail, and uses a deformation-mapping-combination fitting strategy and conditional skin inference to achieve robust and realistic virtual fitting in complex scenes. The present application disentangles the semantic categories to avoid the influence of the original clothing, and uses cycle consistency to improve the robustness of the model to non-paired data, thereby achieving reasonable inference of the fitting semantic layout. With the fitting semantic layout as the shape guide, a static overlay and dynamic selection deformation strategy and explicit task assignment based on region alignment are used to achieve a balance between the degree of deformation and texture preservation in garment fitting. To solve the problem of heavy exposed skin inference in garment fitting, the StyleGAN2 model is adjusted to infer the heavy exposed skin content based on the target skin shape and spatially independent skin features, and random erasure is introduced in self-supervised training to simulate the skin exposure in actual fitting.

[0006] The purpose of the present application is realized by the following technical solutions:

[0007] 1) form a fitting image pair composed of a user image and a fitting garment image , and obtain a corresponding human residual part segmentation map after pre-processing each fitting image pair dense human pose D and fitting garment mask

[0008] 2) input the human residual part segmentation map corresponding to each fitting image pair dense human pose D and fitting garment mask into an instance-level segmentation inference module to learn the semantic features of the garment-human and infer the fitting semantic layout, and obtain a virtual fitting instance-level segmentation map corresponding to each fitting image pair segment the virtual fitting instance-level segmentation map to obtain a corresponding virtual fitting skin segmentation map virtual fitting garment segmentation map and virtual fitting residual part segmentation map

[0009] 3) input the virtual fitting garment segmentation map and the fitting garment image together into a progressive garment fitting module to deform the garment to the human pose, and obtain a final fitting result

[0010] 4) determine whether the skin is exposed according to the virtual fitting skin segmentation map and skin content prior , if the skin is exposed, input the virtual fitting skin segmentation map and the skin content prior into a re-exposed skin inference module to repair the skin content, and obtain a re-exposed skin image otherwise, multiply the virtual fitting skin segmentation map and the skin content prior to obtain a re-exposed skin image

[0011] 5) multiply the virtual fitting residual part segmentation map and the user image to obtain a virtual fitting residual part image finally, multiply the final fitting result re-exposed skin image and virtual fitting residual part image to obtain a virtual fitting image​ Implementing virtual fitting for a user.

[0012] In the step 1), the user image After the pre-trained human segmentation network, the instance-level human segmentation map is obtained Then, the instance-level human segmentation map is further processed After segmentation, the corresponding instance-level human skin segmentation map is obtained The instance-level human clothing segmentation map is obtained And the instance-level human remaining part segmentation map is obtained The instance-level human clothing segmentation map is obtained The instance-level human skin segmentation map is obtained as the input of the progressive clothing fitting module during training The user image is obtained as the input of the re-exposed skin inference module during training After the pre-trained dense pose detection network, the dense human pose D is obtained, and the fitting clothing image is obtained After the clothing segmentation network, the fitting clothing mask is obtained

[0013] In the step 2), the upper body part of the virtual fitting instance-level segmentation map is taken as the virtual fitting skin segmentation map The clothing part of the virtual fitting instance-level segmentation map is taken as the virtual fitting clothing segmentation map The remaining part of the virtual fitting instance-level segmentation map is taken as the virtual fitting remaining part segmentation map The clothing part of the virtual fitting instance-level segmentation map is taken as the virtual fitting clothing segmentation map The remaining part of the virtual fitting instance-level segmentation map is taken as the virtual fitting remaining part segmentation map

[0014] In the step 2), the instance-level segmentation inference module is composed of the first U2-Net network architecture connected with the first SoftMax layer, and when the instance-level segmentation inference module is predicted, the human remaining part segmentation map The dense human pose D and the fitting clothing mask After the first U2-Net network architecture and the SoftMax layer, the virtual fitting instance-level segmentation map is output

[0015] When the instance-level segmentation inference module is trained, the human remaining part segmentation map The dense human pose D and the fitting clothing mask After the first U2-Net network architecture and the first SoftMax layer, the virtual fitting instance-level segmentation map is output Then the virtual fitting instance-level segmentation map The dense human pose D and the second fitting clothing mask ​The first U2-Net network architecture and the first SoftMax layer are sequentially outputted to obtain a virtual fitting pair instance-level segmentation map Second fitting clothes mask Is a user image The upper body image in the image is obtained after passing through the clothes segmentation network.

[0016] In step 3), the progressive clothes fitting module includes a coarse deformation stage, a fine mapping stage, and a combination stage.

[0017] In the prediction, in the coarse deformation stage, first, the virtual fitting clothes segmentation map And the fitting clothes mask Respectively input into the pre-trained first VGG-19 model, and the output of the'relu1_1','relu2_1','relu3_1' and'relu4_1' layers of the first VGG-19 model is fused to obtain the pyramid virtual fitting clothes segmentation feature map And the fitting clothes mask feature map Then the virtual fitting clothes segmentation feature map And the fitting clothes mask feature map After down-sampling and projection, respectively, the corresponding transformed virtual fitting clothes segmentation feature map and the fitting clothes mask feature map are obtained, and then the transformed virtual fitting clothes segmentation feature map and the fitting clothes mask feature map are combined and merged with the learnable position embedding To obtain a shape embedding sequence

[0018] The shape embedding sequence Is sequentially input into the first Transformer aggregator and the TPS regression head, and the TPS regression head outputs four deformation scales, according to which the fitting clothes image And the fitting clothes mask Respectively coarsely deform into the first-fourth coarsely deformed fitting clothes images And And the first-fourth coarsely deformed clothes masks And

[0019] In the fine mapping stage, first, according to the virtual fitting clothes segmentation map And the fitting clothes mask The first-fourth coarsely deformed clothes masks And Are screened to determine the optimal scale index The optimal coarsely deformed clothes mask is determined by the optimal scale index ​And the optimal coarse-deformed fitting garment image The specific filtering formula is as follows:

[0020]

[0021] Where i represents the scale index number, i = 3, 4, 5, 6, τ is the weighting coefficient; avg(·) is the average function along the spatial scale, |||| represents the first norm; argmin[·] represents the minimum value index function;

[0022] Then, the optimal coarse-deformed fitting clothing image is obtained. And virtual try-on clothing segmentation image The data is input into the first Restormer module, which outputs the finely mapped try-on clothing. During prediction, the combination phase is performed directly; during training, the fine-mapping of the clothing is applied. With the optimal coarse deformation of the clothing mask After merging, the images are input into the second Restormer module, which outputs the restored optimal coarse-distorted fitting image of the clothing.

[0023] In the combination stage, the optimal coarsely deformed fitting garment image is used. Fitting clothes with fine reflection The data is sequentially input into the second U2-Net network architecture and the first normalization layer. The first normalization layer outputs the combined mask. Finally, combine the mask. Optimal coarse deformation of the fitting garment image Fitting clothes with fine reflection The final fitting result is obtained after image compositing. The image synthesis formula is as follows:

[0024]

[0025] Here, ⊙ represents element-wise multiplication;

[0026] During training, the progressive clothing try-on module is input as an instance-level human body clothing segmentation map. Images of clothing being tried on and try-on clothing mask

[0027] In the step 4), the re-exposed skin inference module comprises a second VGG-19 model, a mapping network, a third VGG-19 model, a StyleGAN2 module and four convolutional pooling modules, the output content feature maps {η i} i=1,2,3,4 of the'relu1_1','relu2_1','relu3_1','relu4_1' layers of the second VGG-19 model are respectively input into the four convolutional pooling modules, and then the feature maps of the four channels of the content feature maps {η i} i=1,2,3,4 are combined to obtain a one-dimensional latent feature code z, the one-dimensional latent feature code z is input into the StyleGAN2 module through the mapping network, and the output of the'relu4_1' layer of the third VGG-19 model is also input into the StyleGAN2 module, and the StyleGAN2 module outputs a re-exposed skin image

[0028] In the training, the instance-level human skin segmentation map is randomly erased to generate a skin erasing mask The skin erasing mask is multiplied by the user image to obtain a skin erasing image The skin erasing image is input into the second VGG-19 model as an upper body skin mask is input into the third VGG-19 model as an input;

[0029] In the prediction, the skin content prior is input into the second VGG-19 model as a virtual fitting skin segmentation map is input into the third VGG-19 model as an input.

[0030] According to the re-exposed skin image and the skin content prior , the final re-exposed skin image is calculated through the following formula:

[0031]

[0032] Wherein, is the final re-exposed skin image, and is element multiplication.

[0033] The beneficial effects of the present application are:

[0034] The application is mainly an image virtual fitting method for complex scenes. The method can improve the fitting robustness in fitting different types of clothes or in complex human posture occlusion conditions, and realize robust and real virtual fitting in complex scenes. The method can be used for fitting display in online clothing retail to improve the immersion of consumers in the fitting state of clothes. BRIEF DESCRIPTION OF DRAWINGS

[0035] The specific embodiments of the application will be described in detail below with reference to the accompanying drawings:

[0036] Figure 1 is a flowchart of the application.

[0037] Figure 2 is a structure diagram of the instance-level human segmentation inference module.

[0038] Figure 3 is a structure diagram of the progressive garment fitting module.

[0039] Figure 4 is a structure diagram of the re-exposed skin inference module.

[0040] Figure 5 is a many-to-many virtual fitting result.

[0041] Figure 6 is a virtual fitting result in complex conditions. DETAILED DESCRIPTION

[0042] In order to more clearly illustrate the application, the application will be further described below in conjunction with the drawings and examples. Those skilled in the art should understand that the specific description below is illustrative rather than limiting and should not limit the scope of protection of the application.

[0043] For personalized fitting display in online clothing retail, the fitting garment image and user image are input, the shape of the fitting garment is changed to produce a wearing effect, and the re-exposed skin content is inferred to enhance the fitting reality. As shown in Figure 1 The application proposes an image virtual fitting method with robust garment fitting and conditional skin inference. By unbinding representation semantics and introducing cycle consistency, a reasonable fitting semantic layout is inferred using an instance-level human segmentation inference module. As a shape guide, a deformation-mapping-combination garment fitting strategy is used to change the fitting garment to the target shape using a progressive garment fitting module. By adjusting the StyleGAN2 model, the re-exposed skin inference module is used to infer the skin content based on the target skin shape and spatially independent skin features.

[0044] The application includes the following steps:

[0045] 1) An image dataset is composed of multiple pairs of try-on image pairs, each pair of try-on image pair includes a user image and a try-on garment image Each pair of try-on image pair is preprocessed to obtain a corresponding body residual part segmentation map a dense body pose D and a try-on garment mask

[0046] In step 1), the user image is input into a pre-trained body segmentation network to obtain an instance-level body segmentation map The instance-level body segmentation map is further segmented to obtain a corresponding instance-level body skin segmentation map The instance-level body garment segmentation map and the instance-level body residual part segmentation map Specifically, the upper body part of the instance-level body segmentation map is taken as the instance-level body skin segmentation map The garment part of the instance-level body segmentation map is taken as the instance-level body garment segmentation map The remaining part of the instance-level body segmentation map after taking the upper body part and the garment part is taken as the instance-level body residual part segmentation map Wherein the instance-level body garment segmentation map is taken as the input of the progressive garment try-on module during training, the instance-level body skin segmentation map is taken as the input of the re-exposed skin inference module during training; the user image is input into a pre-trained dense pose detection network to obtain a dense body pose D, and the try-on garment image is input into a garment segmentation network to obtain a try-on garment mask

[0047] 2) The body residual part segmentation map the dense body pose D and the try-on garment mask of each pair of try-on image pair are input into an instance-level segmentation inference module to learn the semantic features of the garment-body and infer the try-on semantic layout, to obtain a corresponding virtual fitting instance-level segmentation map of each pair of try-on image pair The virtual fitting instance-level segmentation map is further segmented to obtain a corresponding virtual fitting skin segmentation map The virtual fitting garment segmentation map and the virtual fitting residual part segmentation map

[0048] In step 2), the virtual fitting instance-level segmentation map of each pair of try-on image pair is taken as the input of the progressive garment try-on module during inference, and the virtual fitting skin segmentation map ​​​upper body part of the upper body part of the user as a virtual fitting skin segmentation map taking the virtual fitting instance-level segmentation map of the upper body part and the garment part the garment part of the upper body part as a virtual fitting garment segmentation map taking the virtual fitting instance-level segmentation map of the upper body part and the garment part the remaining part of the upper body part as a virtual fitting remaining part segmentation map

[0049] Since training the instance-level human segmentation inference module using only paired fitting image pairs faces robustness problems, the case perception of the instance-level human segmentation inference module for unpaired fitting image pairs is promoted by introducing cycle consistency. The instance-level human segmentation inference module first infers the semantic layout under the paired fitting condition according to the first fitting garment mask inference of the semantic layout under the unpaired fitting condition Since the truth of the unpaired fitting condition is lacking, only the remaining part in the semantic layout is supervised

[0050] Then, the instance-level human segmentation inference module infers the semantic layout under the paired fitting condition according to the second fitting garment mask and and uses the fitting semantic layout truth to supervise In this way, the instance-level human segmentation inference module can infer a reasonable fitting semantic layout at the test stage.

[0051] The instance-level human segmentation inference module has two sub-tasks: directly mapping the remaining part in the human instance-level segmentation map and re-entangling the upper body skin and the fitting garment feature. This requires the architecture to have very high low-level feature transmission and high-level feature understanding ability, therefore, U2-Net is adopted as the architecture of the instance-level human segmentation inference module, which can realize the perception of context information by fusing multi-scale features. All the channel inputs of the inputs are input into the U2-Net, and then the predicted fitting semantic layout is output through the SoftMax layer.

[0052] As shown in Figure 2 , in step 2), the instance-level segmentation inference module is composed of the first U2-Net network architecture connected with the first SoftMax layer, and when the instance-level segmentation inference module is predicted, the fitting image pair is generally an unpaired fitting image pair, that is, the upper garment in the fitting garment image and the user image is different, the human remaining part segmentation map dense human pose D and the fitting garment mask are sequentially output after the first U2-Net network architecture and the SoftMax layer, and the virtual fitting instance-level segmentation map is output where the fitting garment mask​ as the first fitting garment mask

[0053] The instance-level segmentation inference module, in training, the fitting image pair is a non-paired fitting image pair or a paired fitting image pair, the paired fitting image pair refers to that the upper garment in the fitting garment image and the user image is the same, and the remaining part of the human body segmentation map dense human pose D and the fitting garment mask After passing through the first U2-Net network architecture and the first SoftMax layer in turn, the virtual fitting instance-level segmentation map is output that is, the virtual fitting non-paired instance-level segmentation map wherein the fitting garment mask as the first fitting garment mask Then the virtual fitting instance-level segmentation map dense human pose D and the second fitting garment mask After passing through the first U2-Net network architecture and the first SoftMax layer in turn, the virtual fitting paired instance-level segmentation map is output the second fitting garment mask is the upper garment image in the user image obtained after passing through the garment segmentation network, and the first U2-Net network architecture and the first SoftMax layer in the two training processes are shared.

[0054] The objective function l of the instance-level human body segmentation inference module si The formula is as follows:

[0055]

[0056] wherein, l1(·) and l ce (·) are L1 loss function and cross-entropy loss function respectively; λ1-λ3 are the first-third hyperparameters.

[0057] 3) inputting the virtual fitting garment segmentation map and the fitting garment image the fitting garment mask into the progressive garment fitting module to deform the garment to the human pose, and obtaining the final fitting result

[0058] As shown in Figure 3 , in step 3), the progressive garment fitting module includes a coarse deformation stage, a fine mapping stage and a combination stage.

[0059] In prediction, in the coarse deformation stage, firstly, the virtual fitting garment segmentation map and the fitting garment mask The channels of the virtual try-on garment segmentation feature map are copied to 3 and input into the pre-trained first VGG-19 model respectively, and the'relu1_1','relu2_1','relu3_1' and'relu4_1' layer outputs of the first VGG-19 model are fused to obtain the virtual try-on garment segmentation feature map of the pyramid and the try-on garment mask feature map The virtual try-on garment segmentation feature map is then down-sampled and projected in sequence and the try-on garment mask feature map After down-sampling and projection in sequence, the corresponding transformed virtual try-on garment segmentation feature map and try-on garment mask feature map are obtained Since ViT requires a fixed character length, each to a fixed size h1xw1, and the channel length is projected to δ1, then the transformed virtual try-on garment segmentation feature map and try-on garment mask feature map are merged and combined with the learnable position embedding to satisfy to obtain the shape embedding sequence satisfies The specific formula is as follows:

[0060]

[0061] wherein, is a merging operation.

[0062] The shape embedding sequence is input into the first Transformer aggregator and the TPS regression head in sequence, and in a specific implementation, the first Transformer aggregator is three cascaded Transformer aggregators, which are used to learn the interaction relationship between the marks and the global context information, and learn the attention score through the self-attention mechanism The multi-scale grid can enrich the dimension of deformation without spending too much computing cost, and the TPS regression head outputs four deformation scales, and the multi-scale TPS deformation is set to four scales and According to the four deformation scales, the try-on garment image and the try-on garment mask are coarsely deformed into the first-fourth try-on garment images after coarse deformation and and the first-fourth garment masks after coarse deformation and

[0063] Due to the static mesh partitioning in TPS warping, the similarity between the target shape and the coarse shape at each scale varies with the fitting instance. In order to determine the unique reference for the subsequent stage, the coarse result that best fits the fitting instance is selected from the multi-scale coarse warping results according to the fitting instance. In this way, the coarse warping stage learns the warping of four static scales to cover the general clothing fitting cases, and the fine mapping stage dynamically selects the optimal coarse result as the reference according to the fitting instance.

[0064] When selecting the optimal coarse result, the following two criteria need to be considered: 1) The coarse result shape should be sufficiently consistent with the target shape. Otherwise, a large area of misalignment will pose a great challenge to the learning of the fine mapping stage; 2) The texture preservation of the coarse result should be consistent with the fitted clothing image. The degree of warping and texture preservation are in conflict, with the former tending to fine-grained mesh partitioning and the latter tending to coarse-grained mesh partitioning.

[0065] In the fine mapping stage, the virtual fitting clothing segmentation map and the fitted clothing mask are first screened according to the first-fourth fitted clothing masks after coarse warping and to determine the optimal scale index The optimal coarse warping fitted clothing mask and the optimal coarse warping fitted clothing image are determined by the optimal scale index The specific screening formula is as follows:

[0066]

[0067] Where i represents the scale index number, i = 3, 4, 5, 6, τ is the trade-off coefficient; avg(·) is the average function along the spatial scale, |||| represents the first norm; argmin[·] represents the minimum value index function;

[0068] The optimal coarse warping fitted clothing image and the virtual fitting clothing segmentation map are then input into the first Restormer module, and the first Restormer module outputs the fine-mapped fitted clothing When predicting, the combination stage is directly performed; when training, the fine-mapped fitted clothing is merged with the optimal coarse warping fitted clothing mask and then input into the second Restormer module, and the second Restormer module outputs the restored optimal coarse warping fitted clothing image The optimal coarse warping fitted clothing image is restored to the optimal coarse warping fitted clothing image Supervision is performed;

[0069] The architecture of the fine mapping stage adopts the Restormer module, which has the structure of GAN and the mechanism of Transformer. Restormer can query global semantic features using cross-channel attention mechanism to realize content mapping of local regions. Cycle consistency is also introduced into the self-supervised training of the fine mapping stage. The best fine-fitting garment image after coarse deformation and the virtual fitting garment segmentation map are input into the Restormer to predict the fine-mapped fitting garment Expectation in the aligned region has the same content as , while in the non-aligned region mapping is performed. Therefore, the best fine-fitting garment image after coarse deformation and the fitting truth are used to supervise the fine-mapped fitting garment Otherwise, only as the supervision object will cause the Restormer to perform mapping within the entire shape. Then, since the content of the non-aligned region is discarded in the above mapping process, the Restormer is used to perform its content mapping to enhance the robustness of the module. The fine-mapped fitting garment and the best coarse deformation garment mask are input to predict the recovered best coarse deformation fitting garment image This helps to indirectly supervise the conversion of the content of the aligned region in the prediction .

[0070] In the combination stage, due to the limitation of the size of the feature map, there is still information loss in the content transfer of the aligned region. Therefore, by predicting the combination mask, direct transfer of texture and details at the pixel level is realized. The best coarse deformation fitting garment image and the fine-mapped fitting garment are sequentially input into the second U2-Net network architecture and the first normalization layer (with Sigmoid as the activation function), and the first normalization layer outputs the combination mask Finally, the combination mask is combined with the best coarse deformation fitting garment image and the fine-mapped fitting garment to obtain the final fitting result after image synthesis The image synthesis formula is as follows:

[0071]

[0072] where is element-wise multiplication.

[0073] During training, the input of the progressive garment fitting module is instance-level human garment segmentation map fitted garment image and fitted garment mask all virtual fitted garment segmentation maps at the time of prediction are replaced by instance-level human garment segmentation map

[0074] For the coarse deformation stage, the instance-level human garment segmentation map is used to supervise the shape deformation of each scale. To prevent texture distortion, while regularizing the position movement of the mesh, its objective function l cw is as follows:

[0075]

[0076] where l1(·) and l2(·) are L1 and L2 loss functions, respectively; is the deformed mesh at scale i; is the original mesh at scale i; λ4 and λ5 are the fourth and fifth hyperparameters, respectively.

[0077] For the fine mapping stage, the L1 loss function l1(·) and the content loss function l vgg (·) are used to constrain the prediction results and at the pixel and perception levels, respectively. and fitted ground truth are used to supervise the fitted garment of the fine mapping simultaneously. The supervision effects of the two are balanced by the weighing coefficient ξ. The optimal coarse result is used to fully supervise the recovered optimal coarse deformed fitted garment image Thus, the objective function l fm of the fine mapping stage is as follows:

[0078]

[0079] where λ6 and λ7 are the sixth and seventh hyperparameters, respectively.

[0080] For the combination stage, to achieve the pixel direct transfer task, in the aligned region the optimal coarse deformed fitted garment image and fitted ground truth are used to jointly supervise the final fitting result In the non-aligned region, only the fitted ground truth is used to supervise the final fitting result Thus, the objective function l of the combination stage is cp As follows:

[0081]

[0082] wherein ξ represents the thickness combination balance coefficient.

[0083] 4) According to the virtual fitting skin segmentation map and the skin content prior determine whether the skin is exposed, if the skin is exposed, the virtual fitting skin segmentation map and the skin content prior are input into the re-exposed skin inference module for repairing the skin content, and a re-exposed skin image is obtained Otherwise, the virtual fitting skin segmentation map and the skin content prior are multiplied to obtain a re-exposed skin image

[0084] In step 4), as shown in Figure 4 , the re-exposed skin inference module infers the re-exposed skin according to the target skin shape and the spatially independent skin features by adjusting the structure of StyleGAN2, including a second VGG-19 model, a mapping network, a third VGG-19 model, a StyleGAN2 module and four convolution pooling modules. The'relu1_1','relu2_1','relu3_1','relu4_1' layers of the second VGG-19 model output the content feature map {η i} i=1,2,3,4 , and the feature maps of the four channels of the content feature map {η i} i=1,2,3,4 After the corresponding convolution pooling modules, the four channels of the content feature map {η i} are merged to obtain a one-dimensional latent feature code z, which satisfies The one-dimensional latent feature code z is output after the mapping network to obtain the latent feature code w and input into the StyleGAN2 module. The output of the'relu4_1' layer of the third VGG-19 model (i.e., the target shape feature map ) is also input into the StyleGAN2 module, and the StyleGAN2 module outputs the re-exposed skin image

[0085] Specifically, the Conv1x1 layer is used to reduce the channels of each feature map η i , and the AveragePooling layer is used to adjust the size of each feature map and obtain the normalized feature map All η i ' are merged and flattened into a one-dimensional latent feature code Then, the one-dimensional latent feature code z is further input into the mapping layer consisting of eight connected layers to unwrap the features and obtain the latent feature code w.

[0086] During training, instance-level human skin segmentation maps are randomly erased. Creates a skin mask Remove skin mask With user image Multiply to obtain the erased skin image Erase skin image As input to the second VGG-19 model, the upper body skin mask As input to the third VGG-19 model;

[0087] Skin content prior in prediction As input to the second VGG-19 model, the virtual try-on skin segmentation map As input to the third VGG-19 model.

[0088] During the testing phase, instance-level human skin segmentation maps Given that the conditions are known, we need to first determine them. Whether it is satisfied, among which It is a virtual try-on instance-level segmentation diagram. The skin is masked. If certain conditions are met, the skin is exposed, based on prior knowledge of the skin content. And virtual try-on skin segmentation map Perform the task of inferring heavily exposed skin. To obtain better fitting results, the predicted heavily exposed skin images can also be combined. and skin content prior as follows:

[0089]

[0090] in, This is the final image of the exposed skin.

[0091] The objective function of the exposed skin inference module is l ri as follows:

[0092]

[0093] Among them, λ8 and λ9 are the eighth and ninth hyperparameters, respectively.

[0094] In summary, the instance-level segmentation inference module, the progressive clothing try-on module, and the heavily exposed skin inference module are trained independently during training. During prediction, the three modules are concatenated, such as... Figure 1 As shown.

[0095] 5) Divide the remaining part of the virtual try-on image. With user image After multiplication, the remaining image of the virtual try-on is obtained. Finally, the final fitting results Images of exposed skin And the remaining images of virtual try-on The virtual fitting image is obtained after multiplication. Enables users to virtually try on clothes.

[0096] Example

[0097] In this embodiment of the invention, paired images of clothing being tried on are used to self-supervise the training of three modules. The training set and the test set each contain 11,565 and 1,698 pairs of images, respectively, which are cropped and uniformly sized to 256×192. In the test set, the correspondence between store clothing and tried-on clothing is randomly shuffled to simulate a real clothing try-on.

[0098] This embodiment of the invention uses the Adam optimizer to optimize network training, setting the static learning rate of the three modules to 1×e. -4 2×e -4 and 1×e -5 .

[0099] This invention's embodiments test many-to-many virtual try-on cases. For example... Figure 5 As shown, five different categories of clothing were selected as examples of try-on garments, with sleeves ranging from short to long. User images wearing mid-sleeve clothing were selected as the subjects of the try-on. This demonstrates the robustness of the invention in situations involving both re-exposed and covered skin.

[0100] This invention's embodiments tested performance under two complex fitting environments: (1) fitting different types of clothing; (2) user images with complex poses, where arms partially obscure the clothing. Figure 6 As shown, in the first complex case, the instance-level human body segmentation inference module utilizes semantic category unwrapping representation to avoid the influence of the original clothing, while the re-exposed skin inference module effectively solves the skin repair problem. Therefore, the fitting result is very natural. In the second case, the instance-level human body segmentation inference module can still infer a reasonable fitting semantic layout and uses a progressive fitting module to infer the content of the segmented clothing parts. Therefore, the fitting result is robust and realistic. Thus, this invention will benefit online clothing retail and promote the further popularization of image-based virtual fitting.

Claims

1. A method for image-based virtual try-on with robust clothing try-on and conditional skin inference, characterized in that, Includes the following steps: 1) From user images Images of trying on clothes The images were paired to form a fitting image set. Each pair of fitting images was preprocessed to obtain a segmentation image of the remaining human body. Dense human poses and try-on clothing mask Among them, user images After pre-training the human segmentation network, an instance-level human segmentation map is obtained. The instance-level human body segmentation image after extracting the upper body and clothing parts. The remaining part is used as an instance-level human body remaining part segmentation map. 2) Segment the remaining human body portion corresponding to each pair of fitting images. Dense human poses and try-on clothing mask The input is fed into the instance-level segmentation inference module to learn the semantic features of clothing and human body and infer the semantic layout of try-on, thereby obtaining the virtual try-on instance-level segmentation map corresponding to each pair of try-on images. Then segment the virtual fitting instance-level graph. After segmentation, the corresponding virtual try-on skin segmentation map is obtained. Virtual try-on clothing segmentation diagram The remaining part of the virtual try-on image is divided into sections. The remaining part of the virtual try-on diagram To retrieve the virtual try-on instance-level segmentation image of the entire upper body and clothing parts. The remaining part; 3) Divide the virtual fitting room into sections. Images of trying on clothes Try on clothing mask The data is input into the progressive clothing try-on module, where the clothing is transformed to fit the human body, resulting in the final try-on effect. 4) Based on the virtual try-on skin segmentation map and skin content prior Determine if skin is exposed; if so, segment the skin in the virtual try-on image. and skin content prior The image is input into the re-exposed skin inference module for skin content repair, resulting in a re-exposed skin image. Otherwise, the virtual try-on skin segmentation image will be used. and skin content prior After multiplication, a heavily exposed skin image is obtained. 5) Divide the remaining part of the virtual try-on image. With user image After multiplication, the remaining image of the virtual try-on is obtained. Finally, the final fitting results Images of exposed skin And the remaining images of virtual try-on The virtual fitting image is obtained after multiplication. Enables users to virtually try on clothes.

2. The image-based virtual try-on method with robust clothing try-on and conditional skin inference according to claim 1, characterized in that, In step 1), the instance-level human body segmentation map is... After segmentation, the corresponding instance-level human skin segmentation map is obtained. Instance-level human body clothing segmentation diagram and instance-level human body remaining part segmentation map Among them, the instance-level human clothing segmentation diagram Instance-level human skin segmentation map serves as the input for the progressive clothing try-on module during training. User images serve as input during training for the heavy exposed skin inference module. After pre-training a dense pose detection network, dense human pose D and clothing fitting images are obtained. The clothing mask is obtained after passing through the clothing segmentation network.

3. The image-based virtual try-on method with robust clothing try-on and conditional skin inference according to claim 1, characterized in that, In step 2), the virtual try-on instance-level segmentation image is retrieved. The upper body portion is used as a virtual try-on skin segmentation map. Take the instance-level segmentation image of virtual try-on The clothing section is used as a virtual fitting room segmentation diagram. The virtual try-on instance-level segmentation image after capturing the upper body and clothing parts. The remaining part is used as a virtual fitting room remaining part segmentation diagram.

4. The image-based virtual try-on method with robust clothing try-on and conditional skin inference according to claim 1, characterized in that, In step 2), the instance-level segmentation inference module consists of a first U2-Net network architecture connected to a first SoftMax layer. During prediction, the instance-level segmentation inference module generates a segmentation map of the remaining human body portion. Dense human poses and try-on clothing mask After passing through the first U2-Net network architecture and the SoftMax layer, the virtual try-on instance-level segmentation map is output. During training, the instance-level segmentation inference module generates segmentation images of the remaining human body parts. Dense human poses and try-on clothing mask After passing through the first U2-Net network architecture and the first SoftMax layer in sequence, a virtual try-on instance-level segmentation map is output. Then, the virtual try-on instance-level segmentation diagram. Dense human poses D and second fitting clothing mask After passing through the first U2-Net network architecture and the first SoftMax layer in sequence, the output virtual try-on paired instance-level segmentation graph is generated. Second try-on clothing mask User image The image of the top in the image was obtained after being processed by a clothing segmentation network.

5. The image-based virtual try-on method with robust clothing try-on and conditional skin inference according to claim 1, characterized in that, In step 3), the progressive clothing try-on module includes a coarse deformation stage, a fine mapping stage, and a combination stage. During the prediction process, in the coarse deformation stage, firstly, the virtual fitting outfit segmentation diagram is generated. and try-on clothing mask The inputs are fed into the pre-trained first VGG-19 model, and the outputs of the 'relu1_1', 'relu2_1', 'relu3_1', and 'relu4_1' layers of the first VGG-19 model are fused to obtain the virtual fitting room segmentation feature maps of the pyramid. and the feature map of the clothing fitting Next, the virtual try-on clothing segmentation feature map and the feature map of the clothing fitting After downsampling and projection are performed sequentially, the corresponding transformed virtual try-on clothing segmentation feature map and try-on clothing mask feature map are obtained. Then, the transformed virtual try-on clothing segmentation feature map and try-on clothing mask feature map are merged and then embedded with the learnable location. Merge to obtain shape embedding sequence Shape Embedded Sequence The images are sequentially input into the first Transformer aggregator and the TPS regression head. The TPS regression head outputs four deformation scales, and the images of the tried-on clothing are then processed according to these four deformation scales. and try-on clothing mask The images are coarsely deformed into the first to fourth fitting images. and and the first to fourth garment masks after rough deformation and In the fine mapping stage, the virtual fitting and clothing segmentation map is first used as the basis. and try-on clothing mask The first to fourth garment masks after coarse deformation and Perform filtering to determine the optimal scale index. Indexed by optimal scale Determine the optimal clothing mask after coarse deformation And the optimal coarse-deformed fitting garment image The specific filtering formula is as follows: Where i represents the scale index number, i = 3, 4, 5, 6, τ is the weighting coefficient; avg(·) is the average function along the spatial scale, ||| represents the first norm; argmin[·] represents the minimum index function; Then, the optimal coarse-deformed fitting clothing image is obtained. And virtual try-on clothing segmentation image The data is input into the first Restormer module, which outputs the finely mapped try-on clothing. During prediction, the combination phase is performed directly; during training, the fine-mapping of the clothing is applied. With the optimal coarse deformation of the clothing mask After merging, the images are input into the second Restormer module, which outputs the restored optimal coarse-distorted fitting image of the clothing. In the combination stage, the optimal coarsely deformed fitting garment image is used. Fitting clothes with fine reflection The data is sequentially input into the second U2-Net network architecture and the first normalization layer. The first normalization layer outputs the combined mask. Finally, combine the mask. Optimal coarse deformation of the fitting garment image Fitting clothes with fine reflection The final fitting result is obtained after image compositing. The image synthesis formula is as follows: Here, ⊙ represents element-wise multiplication; During training, the progressive clothing try-on module is input as an instance-level human body clothing segmentation map. Images of clothing being tried on and try-on clothing mask 6. The image-based virtual try-on method with robust clothing try-on and conditional skin inference according to claim 1, characterized in that, In step 4), the exposed skin inference module includes a second VGG-19 model, a mapping network, a third VGG-19 model, a StyleGAN2 module, and four convolutional pooling modules. The 'relu1_1', 'relu2_1', 'relu3_1', and 'relu4_1' layers of the second VGG-19 model output content feature maps {η}. i } i=1,2,3,4 Content feature graph {η i } i=1,2,3,4 The feature maps of the four channels are processed by their respective convolutional pooling modules and then merged to obtain a one-dimensional latent feature code z. This one-dimensional latent feature code z is then processed by a mapping network and input into the StyleGAN2 module. The output of the 'relu4_1' layer of the third VGG-19 model is also input into the StyleGAN2 module. The StyleGAN2 module outputs a re-exposed skin image. During training, instance-level human skin segmentation maps are randomly erased. Creates a skin mask Remove skin mask With user image Multiply to obtain the erased skin image Erase skin image Instance-level human skin segmentation map as input to the second VGG-19 model As input to the third VGG-19 model; Skin content prior in prediction As input to the second VGG-19 model, the virtual try-on skin segmentation map As input to the third VGG-19 model.

7. The image-based virtual try-on method with robust clothing try-on and conditional skin inference according to claim 6, characterized in that, Based on images of heavily exposed skin and skin content prior The final re-exposed skin image is calculated using the following formula: in, This is the final image of the exposed skin, and ⊙ represents element-wise multiplication.

Citation Information

Patent Citations

  • Semantic-based multi-posture virtual fitting method

    CN113361560A

  • Omnibearing virtual fitting method and system based on circulation three-level transformation and medium

    CN114820294A