Multi-Garment Virtual Try-On With Single-Stage Diffusion
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing virtual try-on technologies struggle with synthesizing high-resolution multi-garment images that accurately capture individual body proportions, pose, and personal characteristics while preserving identity, often leading to loss of detail and realism.
Innovation Solution
A single-stage denoising diffusion model that processes input sets comprising person and garment images, with a progressive training strategy and efficient fine-tuning for person identity, allowing direct synthesis of high-resolution images and incorporating text-based layout descriptions.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Productivity
If multi-stage or cascaded models with super-resolution stages are used, then the model can process images at lower resolution first, but this leads to loss of important garment details and inability to create intricate warps and occlusions at higher resolution
Solution Approach 1:
The patent changes the resolution parameter by training the diffusion model at high resolution directly, eliminating the need for multi-stage processing. The model is trained on high-resolution images and patches, allowing it to generate intricate warps and occlusions while preserving garment details without the loss associated with progressive super-resolution approaches.
2Adaptability or versatility
If clothing-agnostic representations are used in VTO, then the model can replace garments effectively, but this removes significant identity information such as body shape, pose, and distinguishing features
Solution Approach 1:
The patent segments the input into high-resolution image patches that are processed independently through the diffusion model. This allows the model to preserve fine-grained details of person identity (body shape, pose, distinguishing features) while still enabling effective garment replacement. The patch-based approach maintains the relationship between garment and person details throughout the synthesis process.
3Manufacturing precision
If single-stage diffusion models are used, then the model can directly synthesize high-resolution images, but the model capacity may be insufficient to handle intricate warps and occlusions
Solution Approach 1:
The patent applies preliminary action by pre-processing the input images into high-resolution patches and preparing them for diffusion model processing. The model is pre-trained on high-resolution data, so when it receives the patch inputs, it can directly generate the final high-resolution output without requiring additional refinement stages, thus managing complexity while maintaining quality.
Data Source
AI summary
Provided are systems and methods for multi-garment virtual try-on and editing, example implementations of which can be referred to as M&M VTO. The proposed systems allow users to visualize how various combinations of garments would look on a given person. The input for this method can include multiple garment images, an image of a person, and optionally a text description for the garment layout. The output is a high-resolution visualization of how these garments would look on the person in the desired layout. For instance, a user can input an image of a shirt, an image of a pair of pants, a description such as “rolled sleeves, shirt tucked in”, and an image of a person. The output would then be a visual representation of how the person would look wearing these garments in the specified layout.


