High-fidelity virtual fitting method and system based on multi-modal representation fusion

Through the multimodal representation fusion method, precise alignment and semantic consistency between clothing and the human body are achieved, which solves the alignment and consistency problems under complex postures in virtual fitting, generates high-fidelity virtual fitting images, and improves the stability and quality of virtual fitting.

CN120635382AInactive Publication Date: 2025-09-12COMMUNICATION UNIVERSITY OF CHINA
View PDF 0 Cites 2 Cited by

Patent Information

Application Number
CN202510818411.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-18
Publication Date
2025-09-12
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

Existing virtual fitting methods are prone to mode collapse and structural dislocation when dealing with complex clothing designs or diverse posture changes. In addition, there is a lack of large-scale, accurately annotated cross-modal datasets, which hinders the stability and overall quality of virtual fitting in cross-posture scenarios.

Method used

A multimodal representation fusion method is adopted, including an appearance-preserving deformation alignment module, a semantic representation and understanding module, and a multimodal prior-guided appearance generation module. Through a pyramid feature extraction network, a deformable appearance flow estimation network, a semantic attribute structuring module, a dual-model collaborative text generation module, a semantically aligned multi-pose virtual fitting dataset, and a dual-condition-guided appearance generation module, precise alignment and semantic consistency between clothing and the human body are achieved.

Benefits of technology

It achieves precise alignment of clothing and the human body in complex postures, enhances semantic consistency, generates high-fidelity virtual try-on images, solves the consistency problem of clothing in the virtual fitting framework, and improves image quality and stability.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120635382A_ABST
    Figure CN120635382A_ABST
Patent Text Reader

Abstract

The invention provides a high-fidelity virtual fitting method and system based on multi-modal representation fusion, and belongs to the field of computer image generation, and the method comprises the steps: S1, inputting a figure image and a garment image into an appearance keeping deformation alignment module, and outputting an appearance flow graph # imgabs0 #, thereby achieving the geometric alignment between a garment and a figure, and keeping the appearance details of the garment; s2, constructing a semantic representation and understanding module for providing semantic representation for clothing attributes, generating a corresponding text for a clothing image, and ensuring semantic alignment consistency between clothing and a human body under various human body postures; and S3, constructing a multi-modal priori guided appearance generation module, and integrating the multi-modal features and priori knowledge of the pre-training model to optimize an appearance generation effect and realize consistency and unification of semantics and geometry. According to the method, the clothes can be accurately aligned to the human body, and meanwhile fine-grained detail information of the clothes is reserved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of computer image generation, and in particular relates to a high-fidelity virtual fitting method and system based on multimodal representation fusion. Background Art

[0002] Virtual Try-On (VTON), a core task in computer vision and virtual reality, has demonstrated tremendous potential to transform industries such as e-commerce, fashion retail, and virtual human modeling. VTON effectively addresses key challenges in personalized digital fashion visualization by synthesizing clothing onto human images with high fidelity and naturalness while meticulously preserving details such as pose, shape, and appearance. This task promises to revolutionize the online shopping experience through unprecedented interactivity and immersion, while also serving as a crucial benchmark for advancing core computer vision problems, including human parsing, pose estimation, and image synthesis.

[0003] Existing virtual try-on methods can be broadly categorized into three main paradigms, each employing a different modeling approach to address this challenging task. Generative Adversarial Network (GAN)-based methods (such as CP-VTON) use adversarial training between a generator and a discriminator to learn mappings between clothing regions, thereby enhancing the realism of synthesized images. While these methods have achieved significant success in improving visual fidelity, they are prone to issues such as mode collapse and structural misalignment when dealing with complex clothing designs or diverse pose variations. In contrast, U-Net-based methods (such as PF-AFN) employ an encoder-decoder architecture with skip connections to reconstruct images, demonstrating good performance in preserving texture and enhancing clothing details. However, these methods lack flexibility in modeling non-rigid clothing deformations, making them inadequate for handling complex clothing shapes and dynamic variations. In recent years, methods based on diffusion models (such as DM-VTON) have emerged as a powerful alternative. These methods generate high-fidelity images through an iterative denoising process, demonstrating state-of-the-art performance in terms of stability and detail preservation, setting a new benchmark for virtual try-on tasks. While these approaches have achieved numerous breakthroughs in improving image quality and realism, they still face significant challenges in generating spatially consistent images. While previous studies have proposed solutions to these issues, a unified approach that can simultaneously address the triple consistency challenge remains lacking. This limits the stability and overall quality of virtual fitting in cross-pose scenarios. Furthermore, the lack of large-scale, accurately annotated cross-modal datasets further exacerbates this problem, hindering the systematic resolution of the consistency issue and the continued improvement of model performance. Summary of the Invention

[0004] To solve the above technical problems, the present invention provides a high-fidelity virtual fitting method based on multimodal representation fusion, comprising the following steps:

[0005] Step S1: Input the person image and clothing image into the appearance-preserving deformation alignment module and output the appearance flow graph , thereby achieving geometric alignment between clothing and characters and preserving the appearance details of clothing, wherein the appearance preserving deformation alignment module includes:

[0006] Pyramid Feature Extraction Network (PFEN), which is used to extract multi-scale features layer by layer and model clothing and human body structures;

[0007] Deformable Appearance Flow Estimation Network (DFEN) for refining the appearance flow between clothing and the human body;

[0008] Step S2: Constructing a semantic representation and understanding module to provide semantic representation for clothing attributes and generate corresponding text for clothing images, ensuring semantic alignment consistency between clothing and the human body under various human postures; wherein, the semantic representation and understanding module includes: a semantic attribute structuring module SAS, a dual-model collaborative text generation module DMTG, and a semantically aligned multi-pose virtual fitting dataset SAMP-VTONS;

[0009] Step S3: Construct a multimodal prior-guided appearance generation module, including: geometric condition multimodal fusion module GCMF, cross-modal semantic and visual fusion module CSVF and dual condition-guided appearance generation module DGAG; based on the output of GCMF module and CSVF module and ,DGAG module generates the final virtual fitting image.

[0010] Beneficial effects:

[0011] This invention provides a high-fidelity virtual fitting method and system based on multimodal representation fusion. It designs an appearance-preserving deformation alignment module that achieves precise alignment of clothing and the human body in complex poses, effectively addressing the problem of non-rigid clothing deformation and achieving geometric consistency in complex poses. The semantic representation and understanding module constructs a structured clothing attribute representation framework and introduces a large language image understanding model to enhance semantic consistency. The multimodal prior-guided appearance generation module fuses set conditions with multimodal semantic information to preserve fine-grained details of clothing and generate high-fidelity virtual fitting images. This invention can address the consistency issues of clothing in the virtual fitting framework in terms of set alignment, appearance details, and semantic understanding, addressing complex clothing deformations and generating realistic virtual fitting images. BRIEF DESCRIPTION OF THE DRAWINGS

[0012] Figure 1Schematic diagram of a high-fidelity virtual fitting method based on multimodal representation fusion of the present invention;

[0013] Figure 2 Schematic diagram of the overall framework of the method of the present invention;

[0014] Figure 3 Schematic diagram of the structure of the appearance-preserving deformation alignment module;

[0015] Figure 4 It is the clothing attribute category of SAS module;

[0016] Figure 5 Schematic diagram of the text generation process for the DMTG module;

[0017] Figure 6 Schematic diagram of the process of building the SAMP-VTONS dataset;

[0018] Figure 7 This is a qualitative comparison chart of the VITON-HD dataset;

[0019] Figure 8 This is a qualitative comparison chart of the SAMP-VTONS dataset;

[0020] Figure 9 This is a structural block diagram of a high-fidelity virtual fitting system based on multimodal representation fusion of the present invention. DETAILED DESCRIPTION

[0021] In order to make the objectives, technical solutions and advantages of the present invention more clearly understood, the present invention is further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only intended to illustrate the present invention and are not intended to limit the present invention. In addition, the technical features involved in the various embodiments of the present invention described below may be combined with each other as long as they do not conflict with each other.

[0022] Example 1

[0023] like Figure 1 As shown, an embodiment of the present invention provides a high-fidelity virtual fitting method based on multimodal representation fusion, comprising the following steps:

[0024] Step S1: Input the person image and clothing image into the appearance-preserving deformation alignment module and output the appearance flow graph , thereby achieving geometric alignment between clothing and characters and preserving the appearance details of clothing. The appearance-preserving deformation alignment module includes:

[0025] Pyramid Feature Extraction Network (PFEN), which is used to extract multi-scale features layer by layer and model clothing and human body structures;

[0026] Deformable Appearance Flow Estimation Network (DFEN) for refining the appearance flow between clothing and the human body;

[0027] Step S2: Construct a semantic representation and understanding module to provide semantic representation for clothing attributes and generate corresponding text for clothing images, ensuring semantic alignment consistency between clothing and the human body in various human postures. The semantic representation and understanding module includes: a semantic attribute structuring module SAS, a dual-model collaborative text generation module DMTG, and a semantically aligned multi-pose virtual fitting dataset SAMP-VTONS.

[0028] Step S3: Construct a multimodal prior-guided appearance generation module, including: geometric condition multimodal fusion module GCMF, cross-modal semantic and visual fusion module CSVF and dual condition-guided appearance generation module DGAG; based on the output of GCMF module and CSVF module and ,DGAG module generates the final virtual fitting image.

[0029] like Figure 2 FIG. 2 is a schematic diagram of the overall framework of the method of the present invention.

[0030] In one embodiment, the above step S1: inputs the person image and clothing image into the appearance-preserving deformation alignment module, and outputs the appearance flow graph , thereby achieving geometric alignment between clothing and characters and preserving the appearance details of clothing, including:

[0031] Step S11: First, according to the character image Get 3D human body key point information And the segmentation results of the person image ,Will 、 、 and clothing images Input PFEN to extract N layers of clothing texture features and human semantic features ;

[0032] PFEN enables the model to understand clothing and the human body at different scales, while retaining fine-grained details and maintaining semantic integrity, providing accurate features for subsequent appearance flow estimation and geometric alignment;

[0033] Step S12: DFEN consists of N flow networks FN-N. First, the highest layer N features extracted by the pyramid feature extraction network ( , ) is fed into the first flow network FN-1 to generate the appearance flow graph , then and N-1 layer pyramid features ( , ) is fed into the second flow network FN-2 to obtain a more refined appearance flow map ; This process is repeated N times until the most detailed appearance flow graph is finally obtained , for each Use deformable convolution to deform the target clothing. The deformable convolution is expressed as follows:

[0034] ;

[0035] in, is the appearance flow graph, is the position of the convolution kernel, w is the convolution kernel weight, yes The offset of x() is the function of taking the coordinates of the appearance flow graph.

[0036] By refining the flow field layer by layer and introducing deformable convolution, DFEN can accurately capture clothing deformation in different postures, thereby improving the visual quality and geometric alignment effect of high-fidelity virtual fitting images.

[0037] like Figure 3 As shown, a schematic diagram of the structure of the appearance-preserving deformable alignment module is shown, wherein the left side is the pyramid feature extraction network PFEN and the right side is the deformable appearance flow estimation network DFEN. In this embodiment of the present invention, N=3.

[0038] In one embodiment, step S2 above: constructing a semantic representation and understanding module to provide semantic representation for clothing attributes and generate corresponding text for clothing images, thereby ensuring semantic alignment consistency between clothing and the human body under various human postures, specifically includes:

[0039] Step S21: the semantic attribute structuring module SAS is used to provide a systematic and structured semantic representation for clothing attributes;

[0040] The SAS module provides a systematic, structured semantic representation of upper-body clothing attributes. This module improves and expands upon the hierarchical attribute tree in the FashionAI dataset, enriching the description of clothing attributes. SAS introduces three new attribute categories—fit, pattern, and color—and further refines clothing attributes into seven main categories: fit (slim, loose, straight), pattern, color, collar design (low, medium, high), neckline design (V-neck, deep V-neck, round neck, square neck, irregular neckline), sleeve length (sleeveless, short sleeve, long sleeve), and shirt length (high waist, regular, long, extra long).

[0041] like Figure 4 As shown, the clothing attribute category of the SAS module is displayed.

[0042] Step S22: The dual-model collaborative text generation module DMTG adopts a dual-model collaborative strategy, using the Starfire model as the main generator and the CogVLM2 model as the auxiliary generator to generate corresponding text for the clothing image;

[0043] The DMTG module aims to address key challenges in text generation for virtual fitting, particularly accuracy, speed, and compliance. This paper proposes a dual-model collaborative strategy: using the Starfire model as the primary generator for rapid response; and the CogVLM2 model as an auxiliary generator to specifically handle cases containing sensitive elements and provide more detailed text descriptions. Furthermore, this paper introduces a regular expression parsing and pruning mechanism to remove API metadata and non-semantic markup, further improving the accuracy and efficiency of generated descriptions.

[0044] Figure 5 A schematic diagram of the text generation process of the DMTG module is shown.

[0045] Step S23: Construct a semantically aligned multi-pose virtual fitting dataset SAMP-VTONS, use DensePose and OpenPose technologies for high-precision pose estimation, and ensure that each clothing image is accurately aligned with human body images in different poses; use the Human Parse method to perform fine-grained segmentation of the human body and label different body parts; use the ParseAgnostic method to remove identity information and non-fitting areas to reduce the impact of background interference on the fitting process; finally, use the DMTG module to generate an accurate text description for each clothing image, thereby constructing a sample triplet with semantic alignment of image, pose, and text.

[0046] The SAMP-VTONS dataset, constructed by this paper, contains 21,104 training images and 7,487 test images. Each sample includes: human pose images, clothing images, clothing mask images, clothing attribute text descriptions, pose estimation information (DensePose and OpenPose), human body part segmentation (Human Parse), and person-independent feature maps (HumanAgnostic). By enhancing pose diversity and achieving semantic alignment across images, poses, and text, it provides better data support for the training and evaluation of multi-pose virtual fitting models.

[0047] Figure 6 A schematic diagram of the SAMP-VTONS dataset construction process is shown.

[0048] In one embodiment, the above step S3: constructs a multimodal prior guided appearance generation module, including: a geometric condition multimodal fusion module GCMF, a cross-modal semantic and visual fusion module CSVF and a dual condition guided appearance generation module DGAG; based on the output of the GCMF module and the CSVF module and , the DGAG module generates the final virtual fitting image, which specifically includes:

[0049] Step S31: Based on , clothing image to obtain clothing deformation feature map ;Will Character-independent feature maps Input GCMF module, fuse deformable clothing features, character-independent feature maps and structured posture information to obtain spatially consistent geometric priors , specifically including:

[0050] Step S311: Based on and Get clothing deformation feature map ;Will The input GCMF module uses the pre-trained Stable Diffusion encoder Perform feature encoding on it to obtain the encoded deformation features :

[0051] ;

[0052] in, is the number of channels in the latent space, , is the spatial dimension of downsampling;

[0053] Step S312: Use Character-independent feature maps The same encoding method as step S31 is used to obtain the encoded character-independent features. ;

[0054] Step S313: In order to further introduce spatial constraints, a binary mask is introduced , posture graph , a noise latent variable that follows a standard normal distribution , combined with and Construct a single representation space to provide geometric priors for the subsequent DGAG module:

[0055] .

[0056] GCMF is responsible for fusing deformable clothing features, character-independent feature maps, and structured pose information, providing spatially consistent geometric priors for the DGAG module, thereby enhancing the accuracy and consistency of the generated image in structure and details.

[0057] Step S32: Input the clothing image and text into the CSVF module, and generate the final text embedding vector by aligning and merging the text with the clothing visual features by constructing a joint embedding space. , used to guide the generation process of virtual fitting images, specifically including:

[0058] Step S321: The clothing image Via CLIP vision encoder Extracting visual features ;

[0059] Step S322: A network consisting of a Vision Transformer (ViT) and a Multilayer Perceptron (MLP) right Processing is performed to map the extracted features into the CLIP text space to generate predicted pseudo-word vector embeddings ;

[0060] Step S323: Process the text generated by the SRCM module through Tokenizer and Embedding Lookup to obtain the text embedding vector:

[0061] ;

[0062] in, Description text for clothing pattern. Description text for clothing patterns, Describes the color of clothing. Design description text for clothing collar height, Design description text for clothing neckline, Design description text for clothing sleeve length, Design description text for clothing length;

[0063] Step S324: Align and transform and The feature space of , which fuses textual and visual representations, produces a joint embedding ;in, Embedding for visual pseudo-words and text features The total dimension after splicing; Indicates splicing;

[0064] Step S325: Y is passed to the CLIP text encoder , generate the final text embedding vector , as the conditional signal of the DGAG module, is used to guide the generation process of virtual fitting images; express Dimensions;

[0065] Step S33: and Input DGAG module, through iterative fusion of image and text features, gradually synthesize the final virtual fitting image , specifically including:

[0066] The dual condition guided appearance generation DGAG module includes a denoising network based on Denoising UNet. and As input, in each step of the denoising process, As a guiding signal, it ensures that the generated image is consistent with the target text description. The network gradually synthesizes the final virtual fitting image by iteratively fusing image and text features. , construct the following loss function for network training:

[0067] ;

[0068] in, Representation encoder The processed character image, represents the diffusion time step, is the noisy latent variable at time step t, for Conditional coding, is the denoising network; is Gaussian noise; represents the square of the L2 norm, that is, the square of the Euclidean distance;

[0069] The core goal of this loss function is to minimize the input image with virtual fitting images The difference between ( ), and simultaneously fuse the feature space information of image and text.

[0070] By optimizing this loss, the present invention generates high-fidelity images that are highly consistent with the input text description. This effectively guides the model to generate images consistent with both geometric features and semantic information. The dual-conditional guidance of the DGAG module improves the accuracy of the generated virtual try-on images, ensuring that the output matches the intended clothing style, pattern, and appearance.

[0071] In order to evaluate and compare existing virtual fitting frameworks, all virtual fitting frameworks are put into VITON-HD Dataset and SAMP-VTONS Dataset for experiments.

[0072] On the SAMP-VTONS dataset, our HF-VTON was compared with five state-of-the-art methods: PF-AFN (geometric alignment), DCI-VTON (diffusion-based synthesis), LaDI-VTON (latent diffusion modeling), DM-VTON (multimodal fusion), and MV-VTON (multi-view rendering). These five methods represent different technical paradigms in virtual try-on research. Quantitative comparison. Table 1 presents a comprehensive quantitative evaluation of HF-VTON on the VITON-HD and SAMP-VTONS datasets, covering both paired and unpaired settings. In the paired setting on the VITON-HD dataset, HF-VTON achieved an FID of 5.51 and a KID of 0.024, representing a 12.4% improvement over LaDI-VTON (6.29) and an 18.5% improvement over MV-VTON (6.76). The KID was reduced by 82.1% compared to MV-VTON, demonstrating the superiority of HF-VTON in terms of image quality and consistency of generated distributions. With an LPIPS of 0.050, HF-VTON outperforms LaDI-VTON (0.103) and MV-VTON (0.060), indicating enhanced preservation of texture details. In the unpaired setting, HF-VTON maintains robust performance with an FID of 10.12 and a KID of 0.196, surpassing LaDI-VTON (FID: 11.08, KID: 0.265) and MV-VTON (FID: 12.32, KID: 0.427), highlighting its strong generalization ability to compositional changes and data distribution shifts. In the single-pose scenario of the SAMP-VTONS dataset, HF-VTON achieved the highest SSIM (0.892) and the lowest LPIPS (0.043), outperforming DCI-VTON (SSIM: 0.886, LPIPS: 0.073) and MV-VTON (SSIM: 0.869, LPIPS: 0.096), and performing well in terms of structure preservation and texture fidelity. In the unpaired setting, HF-VTON further reduced the FID and KID to 6.92 and 0.256, respectively, surpassing MV-VTON (FID: 7.38, KID: 0.263) and DCI-VTON (FID: 8.15, KID: 0.307), further validating its superior image quality, structure recovery, and detail preservation in standard single-pose tasks.

[0073] Table 1 Comprehensive quantitative evaluation of HF-VTON on VITON-HD and SAMP-VTONS datasets

[0074] Figure 7 A visualization of qualitative model comparisons is presented. The figure includes results for clothing images, person images, and six competing methods. Red boxes highlight areas of poor performance, while green boxes emphasize HF-VTON's advantage in detail recovery and alignment. Existing methods exhibit significant alignment errors, particularly in the shoulder and sleeve regions, resulting in unnatural clothing-to-body fit. In contrast, HF-VTON accurately aligns clothing to the body, ensuring smooth transitions between clothing contours and poses. Existing methods often suffer from blurring or distortion when dealing with complex textures such as stripes and patterns, especially when preserving detail. However, HF-VTON maintains clear texture preservation, particularly excelling in preserving pattern integrity and natural fabric folds. Overall, HF-VTON significantly improves image quality and structural consistency through its representation consistency framework, demonstrating significant advantages in complex clothing deformations, fine-grained detail fidelity, and accurate clothing-to-body alignment. HF-VTON also demonstrates significant advantages over existing methods in single-pose scenarios on the SAMP-VTONS dataset.

[0075] like Figure 8 As shown, LaDI-VTON exhibits significant texture distortion, resulting in a mismatch between the clothing and the original image. Other methods suffer from considerable alignment errors in key areas such as shoulders and upper arms, resulting in unnatural clothing-body fit. HF-VTON achieves precise alignment between clothing and the body, especially in challenging areas, ensuring a natural transition. While MV-VTON performs relatively well, it still suffers from differences in texture detail preservation, especially in clothing detail recovery. HF-VTON excels in preserving fine details, especially in handling complex clothing deformations and texture fidelity, surpassing current methods.

[0076] Example 2

[0077] like Figure 9 As shown, an embodiment of the present invention provides a high-fidelity virtual fitting system based on multimodal representation fusion, including the following modules:

[0078] The appearance preserving deformation alignment module 41 is used to input the character image and the clothing image into the appearance preserving deformation alignment module and output the appearance flow graph , thereby achieving geometric alignment between clothing and characters and preserving the appearance details of clothing;

[0079] Semantic representation and understanding module 42, used to provide semantic representation for clothing attributes and generate corresponding text for clothing images, ensuring semantic alignment consistency between clothing and human body in various human postures;

[0080] The multimodal prior-guided appearance generation module 43 is used to construct a multimodal prior-guided appearance generation module, including: a geometric condition multimodal fusion module GCMF, a cross-modal semantic and visual fusion module CSVF, and a dual condition-guided appearance generation module DGAG; based on the output of the GCMF module and the CSVF module and ,DGAG module generates the final virtual fitting image.

[0081] A high-fidelity virtual fitting device based on multimodal representation fusion includes one or more electronic devices, wherein the one or more electronic devices are used to implement a high-fidelity virtual fitting method based on multimodal representation fusion.

[0082] An electronic device includes: one or more processors; and a memory for storing one or more programs, wherein when the one or more programs are executed by the one or more processors, the one or more processors implement a high-fidelity virtual fitting method based on multimodal representation fusion.

[0083] The foregoing description is intended only to provide specific embodiments of the present invention, which will enable those skilled in the art to understand and implement the present invention. Various modifications to these embodiments will be readily apparent to those skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of the present invention. Therefore, the present invention is not limited to the embodiments shown herein, but is intended to conform to the widest scope consistent with the principles and novel features of the present invention.

Claims

1. A high-fidelity virtual fitting method based on multimodal representation fusion, characterized in that: include: Step S1: Input the person image and clothing image into the appearance-preserving deformation alignment module and output the appearance flow graph , thereby achieving geometric alignment between clothing and characters and preserving the appearance details of clothing, wherein the appearance preserving deformation alignment module includes: Pyramid Feature Extraction Network (PFEN), which is used to extract multi-scale features layer by layer and model clothing and human body structures; Deformable Appearance Flow Estimation Network (DFEN) for refining the appearance flow between clothing and the human body; Step S2: Constructing a semantic representation and understanding module to provide semantic representation for clothing attributes and generate corresponding text for clothing images, ensuring semantic alignment consistency between clothing and the human body under various human postures; wherein, the semantic representation and understanding module includes: a semantic attribute structuring module SAS, a dual-model collaborative text generation module DMTG, and a semantically aligned multi-pose virtual fitting dataset SAMP-VTONS; Step S3: Construct a multimodal prior-guided appearance generation module, including: geometric condition multimodal fusion module GCMF, cross-modal semantic and visual fusion module CSVF and dual condition-guided appearance generation module DGAG; based on the output of GCMF module and CSVF module and ,DGAG module generates the final virtual fitting image.

2. The high-fidelity virtual fitting method based on multimodal representation fusion according to claim 1 is characterized in that: Step S1: inputting the person image and clothing image into the appearance-preserving deformation alignment module and outputting the appearance flow graph , thereby achieving geometric alignment between clothing and characters and preserving the appearance details of clothing, including: Step S11: First, according to the character image Get 3D human body key point information And the segmentation results of the person image ,Will 、 、 and clothing images Input PFEN to extract N layers of clothing texture features and human semantic features ; Step S12: DFEN consists of N flow networks FN-N. First, the highest layer N features extracted by the pyramid feature extraction network ( , ) is fed into the first flow network FN-1 to generate the appearance flow graph , then and N-1 layer pyramid features ( , ) is fed into the second flow network FN-2 to obtain a more refined appearance flow map ; This process is repeated N times until the most detailed appearance flow graph is finally obtained , for each Use deformable convolution to deform the target clothing. The deformable convolution is expressed as follows: ; in, is the appearance flow graph, is the position of the convolution kernel, w is the convolution kernel weight, yes The offset of x() is the function of taking the coordinates of the appearance flow graph.

3. The high-fidelity virtual fitting method based on multimodal representation fusion according to claim 2 is characterized in that: Step S2: constructing a semantic representation and understanding module to provide semantic representation for clothing attributes and generate corresponding text for clothing images, ensuring semantic alignment consistency between clothing and the human body under various human postures, specifically including: Step S21: the semantic attribute structuring module SAS is used to provide a systematic and structured semantic representation for clothing attributes; Step S22: The dual-model collaborative text generation module DMTG adopts a dual-model collaborative strategy, using the Starfire model as the main generator and the CogVLM2 model as the auxiliary generator to generate corresponding text for the clothing image; Step S23: Construct a semantically aligned multi-pose virtual fitting dataset SAMP-VTONS, use DensePose and OpenPose technologies for high-precision pose estimation, and ensure that each clothing image is accurately aligned with human body images in different poses; use the Human Parse method to perform fine-grained segmentation of the human body and label different body parts; use the ParseAgnostic method to remove identity information and non-fitting areas to reduce the impact of background interference on the fitting process; finally, use the DMTG module to generate an accurate text description for each clothing image, thereby constructing a sample triplet with semantic alignment of image, pose, and text.

4. The high-fidelity virtual fitting method based on multimodal representation fusion according to claim 3 is characterized in that: The step S3: constructing a multimodal prior guided appearance generation module, including: a geometric condition multimodal fusion module GCMF, a cross-modal semantic and visual fusion module CSVF and a dual condition guided appearance generation module DGAG; based on the output of the GCMF module and the CSVF module and , the DGAG module generates the final virtual fitting image, which specifically includes: Step S31: Based on , the clothing image obtains a clothing deformation feature map ;Will Character-independent feature maps Input GCMF module, fuse deformable clothing features, character-independent feature maps and structured posture information to obtain spatially consistent geometric priors ; Step S32: Input the clothing image and the text into the CSVF module, and align and merge the text with the clothing visual features by constructing a joint embedding space to generate the final text embedding vector , used to guide the generation process of virtual fitting images; Step S33: and Input DGAG module, through iterative fusion of image and text features, gradually synthesize the final virtual fitting image .

5. The high-fidelity virtual fitting method based on multimodal representation fusion according to claim 4 is characterized in that: Step S31: Based on , the clothing image obtains a clothing deformation feature map ;Will Character-independent feature maps Input GCMF module, fuse deformable clothing features, character-independent feature maps and structured posture information to obtain spatially consistent geometric priors , specifically including: Step S311: Based on and Get clothing deformation feature map ;Will Input GCMF module using encoder Perform feature encoding on it to obtain the encoded deformation features : ; in, is the number of channels in the latent space, , is the spatial dimension of downsampling; Step S312: Use Character-independent feature maps The same encoding method as step S31 is used to obtain the encoded character-independent features. ; Step S313: In order to further introduce spatial constraints, a binary mask is introduced , posture graph , a noise latent variable that follows a standard normal distribution , combined with and Construct a single representation space to provide geometric priors for the subsequent DGAG module: 。 6. The high-fidelity virtual fitting method based on multimodal representation fusion according to claim 4 is characterized in that: Step S32: Input the clothing image and the text into the CSVF module, align and merge the text with the clothing visual features by constructing a joint embedding space, and generate the final text embedding vector , used to guide the generation process of virtual fitting images, specifically including: Step S321: The clothing image Via CLIP vision encoder Extracting visual features ; Step S322: Network consisting of visual converter and multilayer perceptron right Processing is performed to map the extracted features into the CLIP text space to generate predicted pseudo-word vector embeddings ; Step S323: Process the text generated by the SRCM module through Tokenizer and Embedding Lookup to obtain the text embedding vector: ; in, Description text for clothing pattern. Description text for clothing patterns, Describes the color of clothing. Design description text for clothing collar height, Design description text for clothing neckline, Design description text for clothing sleeve length, Design description text for clothing length; Step S324: Align and transform and The feature space of , which fuses textual and visual representations, produces a joint embedding ,in, Embedding for visual pseudo-words and text features The total dimension after splicing; Indicates splicing; Step S325: Y is passed to the CLIP text encoder , generate the final text embedding vector , as the conditional signal of the DGAG module, is used to guide the generation process of virtual fitting images, where express dimension.

7. The high-fidelity virtual fitting method based on multimodal representation fusion according to claim 6 is characterized in that: Step S33: and Input DGAG module, through iterative fusion of image and text features, gradually synthesize the final virtual fitting image , specifically including: The dual condition guided appearance generation DGAG module includes a denoising network based on Denoising UNet. and As input, in each step of the denoising process, As a guiding signal, it ensures that the generated image is consistent with the target text description. The network gradually synthesizes the final virtual fitting image by iteratively fusing image and text features. , construct the following loss function for network training: ; in, Representation encoder The processed character image, represents the diffusion time step, is the noisy latent variable at time step t, for Conditional coding, is the denoising network; is Gaussian noise; represents the square of the L2 norm, that is, the square of the Euclidean distance; The core goal of this loss function is to minimize the input image with virtual fitting images The difference between ( ), and simultaneously fuse the feature space information of image and text.

8. A high-fidelity virtual fitting system based on multimodal representation fusion, characterized by: Includes the following modules: The appearance-preserving deformation alignment module is used to input the person image and clothing image into the appearance-preserving deformation alignment module and output the appearance flow graph , thereby achieving geometric alignment between clothing and characters and preserving the appearance details of clothing; The semantic representation and understanding module is used to provide semantic representation for clothing attributes and generate corresponding text for clothing images, ensuring semantic alignment and consistency between clothing and the human body in various human postures; Multimodal prior-guided appearance generation module, which is used to construct a multimodal prior-guided appearance generation module, including: geometric condition multimodal fusion module GCMF, cross-modal semantic and visual fusion module CSVF and dual condition-guided appearance generation module DGAG; Based on the output of GCMF module and CSVF module and ,DGAG module generates the final virtual fitting image.

9. A high-fidelity virtual fitting device based on multimodal representation fusion, characterized in that: The method comprises one or more electronic devices, wherein the one or more electronic devices are used to implement the method according to any one of claims 1 to 7.

10. An electronic device, characterized in that: include: one or more processors; A memory for storing one or more programs, wherein when the one or more programs are executed by the one or more processors, the one or more processors are enabled to implement the method according to any one of claims 1 to 7.

Citation Information

Cited By

  • Virtual fitting clothes recommendation method and system based on artificial intelligence

    CN121280121A

  • Model training method and device, information processing method and device, storage medium and computer program product

    CN121527589A