A method of adaptive garment registration and fusion-generated virtual fitting
Patent Information
- Application Number
- CN202311053947.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-08-21
- Publication Date
- 2026-09-25
- Estimated Expiration
- 2043-08-21
AI Technical Summary
[0004]先前基于2D图像的虚拟试衣方法中,在服装的几何扭曲部分,基于TPS服装对齐的试衣方法对于复杂外观的服装和身体姿势,难以实现准确的对齐扭曲,并且局部的过度扭曲会影响服装拓扑保持效果,造成服装纹理失真
[0044]1)本发明提出的多尺度邻域共识服装扭曲模块,在多个特征尺度上对目标服装与人像信息的高级全局语义匹配信息进行相关性分析,并结合邻域共识思想使用4D卷积对密集特征描述符进行过滤,以增强服装扭曲模块的鲁棒性,识别准确可靠的特征对应关系,提高服装配准的准确率。
Smart Images

Figure CN117351175B_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of deep learning and computer vision technology, specifically relating to a method for adaptive clothing registration and fusion generation of virtual try-on. Background Technology
[0002] With the rapid development of online shopping, the demand for virtual try-on has increased rapidly in recent years. This technology can simulate the actual offline dressing scene for customers. Customers only need to input their own photos and the clothes they want to try on to get the try-on result images. While making it convenient for customers, it can also reduce the risk of product returns and exchanges and reduce transportation costs. Although previous try-on methods have achieved good generation results, they are usually difficult to effectively process details and generate reliable results during the clothing transfer process due to the complexity of the clothing and the complex human posture.
[0003] Existing virtual try-on technologies are mainly divided into 3D modeling and 2D image-based methods. 3D modeling methods are primarily based on computer graphics, calculating 3D human body information and clothing models for 3D rendering. However, 3D modeling relies on measurement equipment, resulting in high computational costs, low efficiency, and difficulty in data collection, hindering its widespread application in online e-commerce environments. Image-based virtual try-on methods, with their advantages of low data acquisition difficulty, low maintenance costs, and high computational efficiency, have become the mainstream research approach for virtual try-on technology.
[0004] Previous virtual try-on methods based on 2D images struggled to accurately align the geometric distortion of clothing, particularly for complex garments and body poses, especially with TPS-based methods. Excessive distortion in certain areas could also negatively impact topology preservation, leading to texture distortion. Appearance flow-based methods for estimating distorted clothing failed to reliably preserve texture details. Furthermore, previous methods lacked flexibility in applying the same convolution kernel to all body parts during the fusion generation process, resulting in less detailed images, inability to smoothly handle seams, and an susceptibility to artifacts and noise, ultimately failing to produce clear and realistic try-on results. Summary of the Invention
[0005] To overcome the shortcomings of the prior art, the purpose of this invention is to provide a virtual try-on method based on adaptive clothing registration and fusion generation. The preprocessing module obtains a human semantic segmentation map, a dense human pose map, an occluded reference image unrelated to clothing information, and a two-dimensional mask for the clothing by inputting a reference human image and the target clothing. A human segmentation map prediction module generates a target semantic segmentation map of the reference human image after trying on the clothing. A multi-scale neighborhood consensus distortion module achieves accurate registration and distortion between the clothing to be tried on and body parts, while distortion smoothing constraints maintain smooth and natural local distortions in the clothing. A skin reconstruction module preserves or generates skin parts to complete the try-on result for greater realism and detail. A context-adaptive fusion generation module fuses human image and clothing information, using different convolutional kernel operations for different body parts to adaptively generate a naturally coupled and detailed try-on result.
[0006] The purpose of this invention is to provide a virtual try-on method based on adaptive clothing registration and fusion generation, which is implemented based on an adaptive clothing registration and fusion generation network. The network includes a preprocessing module, a human body segmentation prediction module, a multi-scale neighborhood consensus distortion module, a skin area reconstruction module, and a context-adaptive fusion generation module. Specifically, it includes the following steps:
[0007] Step 1: The preprocessing module obtains the human semantic segmentation map, human dense pose map, occlusion reference image unrelated to clothing information, and clothing 2D mask by inputting the reference human image and the target clothing.
[0008] Step 2: Use the human body segmentation prediction module to generate a target semantic segmentation map of the reference human image after trying on the clothing;
[0009] Step 3: The multi-scale neighborhood consensus distortion module realizes the registration distortion of the clothing to be tried on through the target semantic segmentation map;
[0010] Step 4: Use the skin area reconstruction module to preserve or generate the skin areas that complete the fitting results;
[0011] Step 5: Use the context-adaptive fusion generation module to fuse the portrait and clothing information to generate the final try-on result.
[0012] Preferably, step 1 specifically includes the following steps:
[0013] Step 1.1: Use a graph transfer human segmentation map prediction network as the extraction network for human semantic segmentation maps, and extract human semantic segmentation maps S using reference human image I as input. a ;
[0014] Step 1.2: Use a dense human pose estimation network as the extraction network for the dense human pose map, and extract the dense human pose map I using the reference human image I as input. p ;
[0015] Step 1.3: Use an open pose network to obtain key human pose points, and combine them with human semantic segmentation map information to obtain an occluded reference human image I that is independent of clothing information. a ;
[0016] Step 1.4: Use a U-shaped network as the target garment mask extraction network to extract the image of the garment C to be tried on. t As input, extract the two-dimensional mask C of the clothing. m .
[0017] Preferably, in step 2, the human body segmentation prediction module uses a dense human pose map I. p Human semantic segmentation map S a Clothing C to be tried on t Clothing 2D mask C m As input, a U-shaped network is used to generate a target semantic segmentation map S after the reference human image has been tried on. p .
[0018] Preferably, step 3 specifically includes the following steps:
[0019] Step 3.1, extract the target semantic segmentation map S p The target clothing part mask I is obtained from the clothing part. pc ;
[0020] Step 3.2, using a clothing-independent occlusion reference image, clothing part mask, and dense human pose map I. a I pc I p As human body representation input, the clothing to be tried on and the clothing 2D mask C t C m As input for the target garment; E p E is a four-layer pyramid multi-scale feature extraction network for human body representation input. c For the four-layer pyramid multi-scale feature extraction network for the target clothing input, E is used respectively. p E c Multi-scale feature extraction is performed on the human body representation input and the target clothing input to obtain human body representation features. and target clothing features Where l represents the number of pyramid network scales;
[0021] Step 3.3, Calculate human representation features and target clothing features Large-scale spatial features The dense semantic correspondence between them; using the enhanced features P4 and G4 at the top of the high-level semantic information pyramid, where P4 represents the human input after passing through E pThe feature map obtained from the fourth layer output, G4 represents the target clothing input after passing through E. c The fourth layer outputs the feature map; the cosine similarity between all pixels of P4 and G4 is calculated to obtain the 4-D feature map C. f i,j,k,l , where i represents The index along the height direction, j represents The index k along the width direction represents Index along the height direction, l represents Index along the width direction; P 4 i,j This represents the feature value in P4 with height index i and width index j; G 4 k,l C represents the feature value in G4 with height index k and width index l. f i,j,k,l The calculation formula is as follows:
[0022]
[0023] Step 3.4: Refine the 4-D feature map C using 4D convolution. f i,j,k,l The filtered 4-D correlation plot was obtained. i, j, k, l have been defined in step 3.3; subsequently, soft mutual nearest neighbor filtering is used to reduce the matching scores of non-mutual nearest neighbors to obtain dense 4-D matching scores. The calculation formula is as follows:
[0024]
[0025] In the above formula i, j, k, l have been defined in step 3.3, where for The best score ratio between each dimension of P4, where a is the height index of the best score in P4 and b is the width index of the best score in P4. for The best score ratio between each dimension of G4, where c is the height index of the best G4 score and d is the width index of the best G4 score;
[0026] Step 3.5, for E p Human body representation feature output of the first three layers and E c Output of target clothing features in the first three layers calculate and Multi-scale dense correspondence between The calculation formula is Where T is the transpose operation, and l = 1, 2, 3 represents Ep and E c The pyramid scale, where i represents the height index of the human body representation feature output, and j represents the width index of the human body representation feature output. k j represents the height index of the target clothing feature output. k Indicates the width index of the feature output. E represents p The value of height index i and width index j in the feature map output from layer l. E represents c The height index of the feature map obtained from the output of the l-th layer is i k The width index is j k The value;
[0027] Step 3.6, Merge and multi-scale dense correspondence The data is fed into a regression layer to obtain the spatial transformation parameters θ of the thin-plate spline interpolation TPS. The spatial transformation function is then used to distort the target garment to obtain the distorted target garment. Where T θ The TPS spatial transformation function is calculated using the following formula:
[0028]
[0029] Step 3.7, one of the training objective functions of the multi-scale neighborhood consensus distortion module is the end-to-end smooth distortion loss, calculated as follows:
[0030]
[0031] In the above equation, E(f) is the end-to-end smoothing twist loss, α is the smoothing twist loss hyperparameter, and ω i Let φ(r) be the elastic component coefficient in the elastic variation of TPS, where i = 1...n; ij Let be the radial basis function of TPS, where i = 1...n and j = 1...n are the number of marker point sets, and i and j represent different marker points, respectively. This represents the i-th marker point with coordinates (x, y). i ,y i The Euclidean distance between the j-th marker and the transpose of the TPS radial basis function is given by φ(r). ij The calculation formula is as follows:
[0032] φ(r ij ) = r ij 2 logr ij (5).
[0033] Preferably, step 4 specifically includes the following steps:
[0034] Step 4.1, referencing the human semantic segmentation map S a Obtain the mask S of human skin area b , using S b Skin region I is obtained by performing pixel-by-pixel multiplication with the reference portrait I. b , to I b Randomly erase to obtain the erased body part I' b ;
[0035] Step 4.2, extract the target semantic segmentation map S p The target skin region is generated, and the target skin region mask is obtained.
[0036] Step 4.3, using content encoder E c Will I' b Compressed into a content vector C containing body identity information b Masked by skin area Human Body Dense Posture Diagram I p As input, use the structure encoder E s Will I p and Encoded as a feature map F that preserves human body structural information b The adaptive layer-instance normalization AdaLIN is used as the fusion module to fuse skin content information and structural information to achieve C++. b Decoding, for F b Perform instance normalization and layer normalization separately to obtain two normalization results. and ρ is a hyperparameter used to weigh the two normalization results. The calculation formula is shown below:
[0037]
[0038] In the above formula: AdaLIN(F b ) represents the adaptive layer-instance normalization, and γ represents C b The denormalized scaling parameter β, obtained through predictions from several fully connected networks, represents C. b The denormalized bias parameters are predicted through several fully connected networks; finally, the reconstructed skin area is obtained by upsampling decoder D.
[0039] Preferably, step 5 specifically includes the following steps:
[0040] Step 5.1, distort the target clothing Clothing-independent occlusion reference portrait image I a Reconstructing skin areas Human Body Dense Posture Diagram I p The input to the model is a convolutional layer for feature extraction, followed by a spatial adaptive normalization layer, and then channel-based normalization, resulting in the target semantic segmentation map S. p and reconstructing skin areas Modulation is performed by learning scale and bias through a modulation parameter prediction network with convolutional operations;
[0041] Step 5.2, segment the target semantic map S p and reconstructing skin areas A conditional weight prediction network consisting of multiple stacked convolutional layers is applied to obtain the predicted convolutional kernel weight output of each conditional convolutional layer. Conditional convolution operation is performed on the modulation activation of the i-th layer. The convolution operation introduces blueprint separable convolution, which separates the traditional convolution into 1×1 pointwise convolution and K×K depthwise convolution, where K represents the convolutional kernel size. The depthwise convolutional kernel weights are predicted by the conditional weight prediction network. Then, the dimension of the output feature map is expanded by 1×1 pointwise convolution. Steps 5.1 and 5.2 are executed twice in this manner to complete the complete process of generating residual blocks by context-adaptive fusion.
[0042] Step 5.3: Repeat steps 5.1 and 5.2 four times, and then send the residual block output generated by context adaptive fusion into the upsampling layer to obtain the fitting result I'.
[0043] Compared with the prior art, the beneficial effects of the present invention are:
[0044] 1) The multi-scale neighborhood consensus clothing distortion module proposed in this invention performs correlation analysis on the high-level global semantic matching information of target clothing and human image information at multiple feature scales, and uses 4D convolution to filter dense feature descriptors in combination with the neighborhood consensus idea, so as to enhance the robustness of the clothing distortion module, identify accurate and reliable feature correspondence, and improve the accuracy of clothing registration.
[0045] 2) The novel end-to-end fabric twist energy smoothing constraint loss proposed in this invention achieves the preservation of garment texture details and improves the ability to retain complex textures during garment transfer.
[0046] 3) This invention proposes a skin area reconstruction module, which reconstructs the exposed parts of the human figure after trying on clothes to achieve a realistic skin effect in the generated result, so as to meet the needs of the user to try on clothes of different types than the clothes they are already wearing, such as changing from long sleeves to short sleeves, or from round necks to V-necks, and generate realistic and reliable try-on results.
[0047] 4) This invention proposes an adaptive fusion generation module, which can adaptively generate different parts of the body using different convolution kernels based on the target human body segmentation map information, generating more realistic edge and texture details. Attached Figure Description
[0048] Figure 1 This is an overall framework diagram of the adaptive clothing registration and fusion generation network in the virtual try-on method of adaptive clothing registration and fusion generation in this invention;
[0049] Figure 2 This is a preprocessing framework diagram of the reference portrait and the clothing to be tried on in this invention;
[0050] Figure 3 This is a schematic diagram of the multi-scale neighborhood consensus clothing distortion module in this invention;
[0051] Figure 4 This is a schematic diagram of the skin reconstruction module in this invention;
[0052] Figure 5 This is a schematic diagram of the adaptive fusion generation module in this invention;
[0053] Figure 6 This is a schematic diagram of the residual block generated by context adaptive fusion in this invention;
[0054] Figure 7 These are comparative images of qualitative experiments on clothing registration obtained through this invention;
[0055] Figure 8 These are qualitative comparison images of clothing texture detail retention obtained through this invention;
[0056] Figure 9 This is a comparison chart of the ablation qualitative results of the virtual try-on method generated by adaptive clothing registration and fusion obtained by this invention. Detailed Implementation
[0057] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0058] The specific steps are as follows:
[0059] Step 1: The preprocessing module obtains a human semantic segmentation map, a dense human pose map, an occlusion reference image unrelated to clothing information, and a two-dimensional mask of clothing by inputting a reference human image and the target clothing.
[0060] The specific steps are as follows:
[0061] Step 1.1, as follows Figure 2 As shown in (a), a graph transfer human segmentation map prediction network is used as the extraction network for human semantic segmentation maps. The human semantic segmentation map S is extracted with reference human image I as input. a ;
[0062] Step 1.2, as follows Figure 2 As shown in (b), a dense human pose estimation network is used as the extraction network for the dense human pose map. The dense human pose map I is extracted using the reference human image I as input. p ;
[0063] Step 1.3, as follows Figure 2 As shown in (c), by combining the human body key pose points obtained using an open pose network with human body semantic segmentation map information, an occlusion reference human image I that is independent of clothing information is obtained. a ;
[0064] Step 1.4, as follows Figure 2 As shown in (d), a U-shaped network is used as the target clothing mask extraction network to extract the clothing C to be tried on. t As input, extract the two-dimensional mask C of the clothing. m .
[0065] Step 2: Use the human body segmentation map prediction module to generate a target semantic segmentation map of the reference human image after trying on the clothing, specifically using the dense human pose map I. p Human semantic segmentation map S a Clothing C to be tried on t Clothing 2D mask C m As input, such as Figure 1 As shown, U-Net is used as the segmentation map prediction module to generate the target semantic segmentation map S after the reference human image tries on the clothes. p .
[0066] Step 3: The multi-scale neighborhood consensus distortion module realizes the registration distortion of the clothing to be tried on through the target semantic segmentation map;
[0067] The specific steps are as follows:
[0068] Step 3.1, extract the target semantic segmentation map S p The target clothing part mask I is obtained from the clothing part. pc ;
[0069] Step 3.2, using a clothing-independent occlusion reference image, clothing part mask, and dense human pose map I. a I pc I p As human body representation input, the clothing to be tried on and the clothing 2D mask C tC m As input for the target garment; E p E is a four-layer pyramid multi-scale feature extraction network for human body representation input. c For the four-layer pyramid multi-scale feature extraction network for the target clothing input, E is used respectively. p E c Multi-scale feature extraction is performed on the human body representation input and the target clothing input to obtain human body representation features. and target clothing features Where l represents the number of pyramid network scales;
[0070] Step 3.3, Calculate human representation features and target clothing features Large-scale spatial features The dense semantic correspondence between them; using the enhanced features P4 and G4 at the top of the high-level semantic information pyramid, where P4 represents the human input after passing through E p The feature map obtained from the fourth layer output, G4 represents the target clothing input after passing through E. c The fourth layer outputs the feature map; the cosine similarity between all pixels of P4 and G4 is calculated to obtain the 4-D feature map C. f i,j,k,l , where i represents The index along the height direction, j represents The index k along the width direction represents Index along the height direction, l represents Index along the width direction; P 4 i,j This represents the feature value in P4 with height index i and width index j; G 4 k,l C represents the feature value in G4 with height index k and width index l. f i,j,k,l The calculation formula is as follows:
[0071]
[0072] Step 3.4: Refine the 4-D feature map C using 4D convolution. f i,j,k,l The filtered 4-D correlation plot was obtained. i, j, k, l have been defined in step 3.3; subsequently, soft mutual nearest neighbor filtering is used to reduce the matching scores of non-mutual nearest neighbors to obtain dense 4-D matching scores. The calculation formula is as follows:
[0073]
[0074] In the above formula i, j, k, l have been defined in step 3.3, where for The best score ratio between each dimension of P4, where a is the height index of the best score in P4 and b is the width index of the best score in P4. for The best score ratio between each dimension of G4, where c is the height index of the best G4 score and d is the width index of the best G4 score;
[0075] Step 3.5, for E p Human body representation feature output of the first three layers and E c Output of target clothing features in the first three layers calculate and Multi-scale dense correspondence between The calculation formula is Where T is the transpose operation, and l = 1, 2, 3 represents E p and E c The pyramid scale, where i represents the height index of the human body representation feature output, and j represents the width index of the human body representation feature output. k j represents the height index of the target clothing feature output. k Indicates the width index of the feature output. E represents p The value of height index i and width index j in the feature map output from layer l. E represents c The height index of the feature map obtained from the output of the l-th layer is i k The width index is j k The value;
[0076] Step 3.6, Merge and multi-scale dense correspondence The data is fed into a regression layer to obtain the spatial transformation parameters θ of the thin-plate spline interpolation TPS. The spatial transformation function is then used to distort the target garment to obtain the distorted target garment. Where T θ The TPS spatial transformation function is calculated using the following formula:
[0077]
[0078] Step 3.7, one of the training objective functions of the multi-scale neighborhood consensus distortion module is the end-to-end smooth distortion loss, calculated as follows:
[0079]
[0080] In the above equation, E(f) is the end-to-end smoothing twist loss, α is the smoothing twist loss hyperparameter, and ω i Let φ(r) be the elastic component coefficient in the elastic variation of TPS, where i = 1...n; ij Let be the radial basis function of TPS, where i = 1...n and j = 1...n are the number of marker point sets, and i and j represent different marker points, respectively. This represents the i-th marker point with coordinates (x, y). i ,y i The Euclidean distance between the j-th marker and the transpose of the TPS radial basis function is given by φ(r). ij The calculation formula is as follows:
[0081] φ(r ij ) = r ij 2 logr ij (5).
[0082] Step 4: Use the skin area reconstruction module to preserve or generate the skin areas that complete the fitting results;
[0083] The specific steps are as follows:
[0084] Step 4.1, referencing the human semantic segmentation map S a Obtain the mask S of human skin area b , using S b Skin region I is obtained by performing pixel-by-pixel multiplication with the reference portrait I. b , to I b Randomly erase to obtain the erased body part I' b ;
[0085] Step 4.2, extract the target semantic segmentation map S p The target skin region is generated, and the target skin region mask is obtained.
[0086] Step 4.3, using content encoder E c Will I' b Compressed into a content vector C containing body identity information b Masked by skin area Human Body Dense Posture Diagram I p As input, use the structure encoder E s Will I p and Encoded as a feature map F that preserves human body structural information b The adaptive layer-instance normalization AdaLIN is used as the fusion module to fuse skin content information and structural information to achieve C++. b Decoding, for F bPerform instance normalization and layer normalization separately to obtain two normalization results. and ρ is a hyperparameter used to weigh the two normalization results. The calculation formula is shown below:
[0087]
[0088] In the above formula: AdaLIN(F b ) represents the adaptive layer-instance normalization, and γ represents C b The denormalized scaling parameter β, obtained through predictions from several fully connected networks, represents C. b The denormalized bias parameters are predicted through several fully connected networks; finally, the reconstructed skin area is obtained by upsampling decoder D.
[0089] Step 5: Use the context-adaptive fusion generation module to fuse the portrait and clothing information to generate the final try-on result.
[0090] The specific steps are as follows:
[0091] Step 5.1, distort the target clothing Clothing-independent occlusion reference portrait image I a Reconstructing skin areas Human Body Dense Posture Diagram I p The input to the model is a convolutional layer for feature extraction, followed by a spatial adaptive normalization layer, and then channel-based normalization, resulting in the target semantic segmentation map S. p and reconstructing skin areas Modulation is performed by learning scale and bias through a modulation parameter prediction network with convolutional operations;
[0092] Step 5.2, segment the target semantic map S p and reconstructing skin areas A conditional weight prediction network consisting of multiple stacked convolutional layers is applied to obtain the predicted convolutional kernel weight output of each conditional convolutional layer. Conditional convolution operation is performed on the modulation activation of the i-th layer. The convolution operation introduces blueprint separable convolution, which separates the traditional convolution into 1×1 pointwise convolution and K×K depthwise convolution, where K represents the convolutional kernel size. The depthwise convolutional kernel weights are predicted by the conditional weight prediction network. Then, the dimension of the output feature map is expanded by 1×1 pointwise convolution. Steps 5.1 and 5.2 are executed twice in this manner to complete the complete process of generating residual blocks by context-adaptive fusion.
[0093] Step 5.3: Repeat steps 5.1 and 5.2 four times, and then send the residual block output generated by context adaptive fusion into the upsampling layer to obtain the fitting result I'.
[0094] The following experiments demonstrate that the adaptive clothing registration and fusion method for generating virtual try-on is effective and feasible, and possesses certain advantages:
[0095] Based on virtual try-on datasets and virtual try-on-high-definition datasets, the test set is input into the adaptive clothing registration and fusion generation model established in this invention. The predictive performance of this model is analyzed by comparing it with traditional virtual try-on methods and conducting ablation experiments.
[0096] The specific steps are as follows:
[0097] (1) Based on the Virtual Try-On Dataset and the Virtual Try-On-HD Dataset, the Virtual Try-On-HD dataset consists of paired frontal female images and frontal upper body clothing images, and is divided into 11647 / 2032 training / test pairs. The original image resolution is 1024×768, and the dataset resolution chosen in this experiment is 256×192. The Virtual Try-On Dataset consists of frontal female images and upper body images, and is divided into 14221 / 2032 training / test pairs, with an image resolution of 256×192. The experiment was implemented in Python using the PyTorch framework on an RTX A5000 graphics card with 24GB of video memory.
[0098] The human segmentation prediction module was trained 200,000 times with a batch size of 8, using Adam as the optimizer. The decay rate for the first moment of the gradient was 0.5, and the decay rate for the second moment of the gradient was 0.999. The multi-scale neighborhood consensus warping module and the skin reconstruction module were each trained 100,000 times with a batch size of 4, and the optimizer parameters were the same as those for the human segmentation prediction module. The context-adaptive fusion generation module was trained 60,000 times with a batch size of 4, using Adam's optimizer, with a decay rate of 0.5 for the first moment of the gradient and a decay rate of 0.9 for the second moment of the gradient. The learner and discriminator learning rates for the human segmentation prediction and skin reconstruction modules were 0.0002 and 0.0002 respectively, the learner learning rate for the multi-scale neighborhood consensus warping module was 0.0001, and the learner learning rate for the context-adaptive fusion generation module was 0.0001 and 0.0004 respectively.
[0099] (2) Figure 7This is a qualitative comparison of our adaptive clothing registration and fusion virtual try-on method with HR-VITON and FIFA-VITON on a high-definition virtual try-on dataset. As shown in the first row of the figure, when the target garment's color is close to the background, HR-VITON incorrectly identifies the garment's inner edge contour as the outer edge for registration, resulting in an incorrect registration. FIFA incorrectly registers the garment as a V-neck shape. Our method, however, possesses global multi-scale perception and semantic ambiguity removal capabilities, perceiving subtle semantic differences and correctly generating the garment's edge contour. In the second row, there is a complex garment input; the target garment is not placed facing forward, presenting a higher registration difficulty. FIFA incorrectly registers a long-sleeved target garment as a short-sleeved one, while HR-VITON produces semantically ambiguous artifacts. Our model achieves correct garment registration. In the third row, our method achieves garment alignment more accurately and unambiguously. As observed in the green boxes in the third and fourth rows, our adaptive clothing registration and fusion virtual try-on method achieves seamless and smooth edge processing.
[0100] Figure 8 This is a qualitative experimental comparison of the virtual try-on method based on adaptive clothing registration and fusion on the VITON-HD dataset regarding the preservation of clothing texture details. To better illustrate the effectiveness of our proposed distortion energy constraint, we feed a regularized mesh image into a multi-scale neighborhood consensus distortion module, applying the same deformation field as the distorted clothing to the mesh to obtain a distorted mesh. The comparison results show that our method can better achieve local smoothness constraints on clothing, effectively preserving the details of target clothing with complex embroidery textures. In the third row, our model effectively avoids texture compression during clothing edge alignment. When the smoothing distortion loss is removed, the local topology of the clothing is not preserved, resulting in texture distortion.
[0101] Table 1 compares our adaptive clothing registration and fusion method with several state-of-the-art methods (CP-VTON, ACGPN, VITON-HD, HR-VITON) on the VITON-HD dataset using Fréchet distance (FID) and structural similarity index (SSIM). As shown in Table 1, our method achieves the best results, improving SSIM from 0.864 to 0.878 (for HR-VITON) and reducing FID to 7.25 (for HR-VITON). This demonstrates that our model can generate high-quality virtual try-on results.
[0102] Table 2 presents a quantitative experimental comparison of the adaptive clothing registration and fusion generation virtual try-on method on the VITON dataset. We use FID, SSIM, and Peak Signal-to-Noise Ratio (PSNR) as evaluation metrics to evaluate our method against other state-of-the-art methods. Our method also achieves the best performance, improving the learned perceptual image patch similarity (LPIPS) from 0.108 to 0.053 (for C-VITON) and the PSNR from 25.423 to 26.746 (for VTON-HF). The PSNR results indicate that CAE-VTON can generate clearer results. Our method achieves the best performance on both datasets, demonstrating that our model has strong generalization ability.
[0103] Figure 9 This is an ablation experiment result of the adaptive clothing registration and fusion generation virtual try-on method on the VITON-HD dataset. The ablation experiment implemented four variants of the adaptive clothing registration and fusion generation virtual try-on method: (1) removal of the multi-scale neighborhood consensus distortion module (W / O multi-scale neighborhood consensus module), (2) removal of the smooth distortion loss (W / O smooth distortion loss), (3) removal of the skin reconstruction module (W / O skin reconstruction module), and (4) removal of the context-adaptive fusion generator residual block (W / O context-adaptive fusion generator residual block). Figure 9 The ablation experiment shows that (1) the multi-scale neighborhood consensus distortion module can make the texture pattern transfer correctly through accurate clothing distortion; (2) distortion energy loss can effectively constrain the degree of local distortion of clothing; (3) the skin reconstruction module can generate richer skin area details and restore the real shadow effect; (4) the context adaptive fusion generator residual block can adaptively generate more realistic detail information for different body parts based on the semantic segmentation map, retain richer details of each generated part, smooth the repair of seams, and reduce the occurrence of noise and artifacts. Table 3 shows the quantitative results of this ablation experiment. It can be seen that when any key component is ablated, the scores of LPIPS and FID will increase, and the score of SSIM will decrease. The data in the table once again verify the above points we proposed.
[0104] Through the Figures 7-9 Based on the observations of Tables 1-3 and the comprehensive analysis above, it can be clearly seen that the adaptive clothing registration and fusion generation virtual try-on method of the present invention is effective and feasible, and has certain advantages compared with other state-of-the-art 2D virtual try-on methods.
[0105] This invention presents a 2D virtual try-on method based on an adaptive clothing registration and fusion generation model. It is implemented using an adaptive clothing registration and fusion generation network, which includes a preprocessing module, a human body segmentation prediction module, a multi-scale neighborhood consensus distortion module, a skin reconstruction module, and a context-adaptive fusion generation module. The proposed multi-scale neighborhood consensus distortion module effectively solves the clothing alignment problem in complex situations, and the proposed distortion energy loss naturally constrains local clothing deformation. Furthermore, the context-adaptive fusion generation module and body reconstruction module designed in this invention generate results with richer texture details, greater clarity and realism, and natural component coupling.
[0106] Table 1 shows the model comparison results using the virtual try-on - high-definition dataset in the examples.
[0107]
[0108] Table 2 shows the model comparison results using the virtual trial dataset in the examples.
[0109]
[0110] Table 3 shows the quantitative results of the model ablation experiment using the virtual try-on-high-definition dataset in the examples.
[0111]
[0112]
[0113] The above description is merely an embodiment of the present invention and is not intended to limit the invention. Various modifications and variations can be made to the present invention by those skilled in the art. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principle of the present invention should be included within the scope of the claims of the present invention.
Claims
1. A virtual try-on method based on adaptive clothing registration and fusion, characterized in that, This is achieved based on an adaptive clothing registration and fusion generation network, which includes a preprocessing module, a human segmentation map prediction module, a multi-scale neighborhood consensus distortion module, a skin part reconstruction module, and a context-adaptive fusion generation module. Specifically, the following steps are included: Step 1: The preprocessing module obtains the human semantic segmentation map, human dense pose map, occlusion reference image unrelated to clothing information, and clothing 2D mask by inputting the reference human image and the target clothing. Step 2: Use the human body segmentation prediction module to generate a target semantic segmentation map of the reference human image after trying on the clothing; Step 3: The multi-scale neighborhood consensus distortion module realizes the registration distortion of the clothing to be tried on through the target semantic segmentation map; Step 4: Use the skin area reconstruction module to preserve or generate the skin areas that complete the fitting results; Step 5: Use the context-adaptive fusion generation module to fuse the portrait and clothing information to generate the final try-on result; Step 3, as described above, specifically involves the following steps: Step 3.1, Extract the target semantic segmentation map The middle part of the clothing section obtains the mask of the target clothing part. ; Step 3.2: Use a masking reference image (unrelated to clothing information), clothing part mask, and dense human pose map. , , As human body representation input, the garment to be tried on and the garment's two-dimensional mask. , As input for the target clothing; A four-layer pyramid multi-scale feature extraction network for human body representation input. For the four-layer pyramid multi-scale feature extraction network for the target clothing input, respectively using , Multi-scale feature extraction is performed on the human body representation input and the target clothing input to obtain human body representation features. and target clothing features ,in Indicates the number of pyramid network scales; Step 3.3, Calculate human representation features and target clothing features Large-scale spatial features Dense semantic correspondences between them; enhanced features at the top of the pyramid using high-level semantic information. , ,in Indicates the process of human body input The feature map obtained from the fourth layer output, Indicates the input process of the target clothing The feature map obtained from the fourth layer output; calculation , The cosine similarity between all pixels yields a 4D feature map. ,in express Index along the height direction, express Index along the width direction express Index along the height direction, express Index along the width direction; express Medium height index is The width index is eigenvalues; express Medium height index is The width index is eigenvalues, The calculation formula is as follows: (1) Step 3.4: Refine the 4D feature map using 4D convolution. The filtered 4-D correlation plot was obtained. i, j, k, l have been defined in step 3.3; subsequently, soft mutual nearest neighbor filtering is used to reduce the matching scores of non-mutual nearest neighbors to obtain dense 4-D matching scores. The calculation formula is as follows: (2) In the above formula , i, j, k, l have been defined in step 3.3, where for and The optimal score ratio between each dimension for The height index of the best score for The width index of the best score for and The optimal score ratio between each dimension for The height index of the best score for Width index of the best score; Step 3.5, for Human body representation feature output of the first three layers and Output of target clothing features in the first three layers ,calculate and Multi-scale dense correspondence between The calculation formula is: ,in For the transpose operation, in the formula express and The pyramid scale, This represents the height index of the human body representation feature output. This represents the width index of the human body representation feature output. This represents the height index of the target clothing feature output. Indicates the width index of the feature output. express In the The height index of the feature map obtained from the layer output is The width index is The value, express In the The height index of the feature map obtained from the layer output is The width index is The value; Step 3.6, Merge and multi-scale dense correspondence The data is fed into a regression layer to obtain the spatial transformation parameters of the thin-plate spline interpolation TPS. Distorted target clothing is obtained by distorting the target clothing using a spatial transformation function. ,in The TPS spatial transformation function is calculated using the following formula: (3) Step 3.7, one of the training objective functions of the multi-scale neighborhood consensus distortion module is the end-to-end smooth distortion loss, calculated as follows: (4) In the above formula For end-to-end smooth twist loss, To smooth the distortion loss hyperparameter, Here are the elastic component coefficients in the elastic variation of TPS, where ; Here are the radial basis functions of TPS, where , To indicate the number of points in the set, and They represent different marker points. Indicates the first There are 1 marker point with coordinates of 1. ) and the Euclidean distance between the markers For the transpose operation, the radial basis functions of TPS The calculation formula is as follows: (5)。 2. The virtual try-on method for adaptive clothing registration and fusion generation according to claim 1, characterized in that, Step 1, specifically, involves the following steps: Step 1.1: Use a graph transfer human segmentation map prediction network as the extraction network for human semantic segmentation maps, referencing human images. Extract human semantic segmentation map from input ; Step 1.2: Use a dense human pose estimation network as the extraction network for dense human pose maps, referencing human images. Extracting dense human pose maps from input ; Step 1.3: Use an open pose network to obtain key human pose points, and combine them with human semantic segmentation map information to obtain an occluded reference human image that is independent of clothing information. ; Step 1.4: Use a U-shaped network as the target garment mask extraction network for the garment to be tried on. Extract the clothing 2D mask as input. .
3. The virtual try-on method for adaptive clothing registration and fusion generation according to claim 1, characterized in that, In step 2, the human body segmentation prediction module uses dense human pose maps. Human semantic segmentation map Clothing to be tried on and clothing 2D mask As input, a U-shaped network is used to generate a target semantic segmentation map after the reference human image has been tried on. .
4. The virtual try-on method for adaptive clothing registration and fusion generation according to claim 1, characterized in that, Step 4, as described above, specifically involves the following steps: Step 4.1, referencing the human semantic segmentation map Obtain human skin area mask ,use Compared with reference portrait Skin area obtained by pixel-by-pixel multiplication ,right Randomly erase to obtain the erased body parts ; Step 4.2, extract the target semantic segmentation map The target skin region is generated, and the target skin region mask is obtained. ; Step 4.3, using the content encoder Will Compressed into a content vector containing body identity information Masked by skin area Human dense posture diagram As input, use a structure encoder Will and Encoding as feature maps that preserve human body structural information Adaptive layer-instance normalized AdaLIN is used as the fusion module to fuse skin content information and structural information. Decoding, for Perform instance normalization and layer normalization separately to obtain two normalization results. and , For hyperparameters, use To weigh the two normalization results, the calculation formula is as follows: (6) In the above formula: This indicates normalization of the adaptive layer-instance. express Denormalized scaling parameters obtained through prediction using several fully connected networks. express The denormalized bias parameters are obtained through predictions using several fully connected networks; finally, they are passed through an upsampled decoder. Reconstructed skin area .
5. The virtual try-on method for adaptive clothing registration and fusion generation according to claim 1, characterized in that, Step 5 specifically involves the following steps: Step 5.1, distort the target clothing The occlusion reference portrait image is unrelated to clothing information. Reconstructing skin areas Human body dense posture diagram The input to the model is processed by convolutional layers for feature extraction, followed by spatial adaptive normalization layers, and then channel-based normalization to obtain the target semantic segmentation map. and reconstructing skin areas Modulation is performed by learning scale and bias through a modulation parameter prediction network with convolutional operations; Step 5.2, segment the target semantic map and reconstructing skin areas The conditional weight prediction network, consisting of multiple stacked convolutional layers, is used to obtain the predicted convolutional kernel weight outputs for each conditional convolutional layer. For the ... The modulation activation of the layer is used to perform conditional convolution operations; the convolution operation introduces blueprint-based separable convolution, separating traditional convolution into 1 1. Pointwise convolution and Depth convolution, where This indicates the kernel size. The conditional weight prediction network predicts the depth and kernel weights, then uses 1...
1. Expand the dimension of the output feature map by pointwise convolution. Execute steps 5.1 and 5.2 twice in this manner to complete the complete process of generating residual blocks by context-adaptive fusion. Step 5.3: Repeat steps 5.1 and 5.2 four times, then feed the residual block output generated by context adaptive fusion into the upsampling layer to obtain the fitting result. .