Multi-Garment Virtual Try-On With Single-Stage Diffusion

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing virtual try-on technologies struggle with synthesizing high-resolution multi-garment images that accurately capture individual body proportions, pose, and personal characteristics while preserving identity, often leading to loss of detail and realism.

Innovation Solution

A single-stage denoising diffusion model that processes input sets comprising person and garment images, with a progressive training strategy and efficient fine-tuning for person identity, allowing direct synthesis of high-resolution images and incorporating text-based layout descriptions.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Productivity

If multi-stage or cascaded models with super-resolution stages are used, then the model can process images at lower resolution first, but this leads to loss of important garment details and inability to create intricate warps and occlusions at higher resolution

Engineering Contradiction:
Improvecomputational efficiencyVSAvoidgarment detail preservation
Core Design Contradiction:
ProductivityVSManufacturing precision

Solution Approach 1:

The patent changes the resolution parameter by training the diffusion model at high resolution directly, eliminating the need for multi-stage processing. The model is trained on high-resolution images and patches, allowing it to generate intricate warps and occlusions while preserving garment details without the loss associated with progressive super-resolution approaches.

Inventive Principle:
Principle #35Parameter changes

2Adaptability or versatility

If clothing-agnostic representations are used in VTO, then the model can replace garments effectively, but this removes significant identity information such as body shape, pose, and distinguishing features

Engineering Contradiction:
Improvegarment replacement capabilityVSAvoidperson identity information
Core Design Contradiction:
Adaptability or versatilityVSLoss of information

Solution Approach 1:

The patent segments the input into high-resolution image patches that are processed independently through the diffusion model. This allows the model to preserve fine-grained details of person identity (body shape, pose, distinguishing features) while still enabling effective garment replacement. The patch-based approach maintains the relationship between garment and person details throughout the synthesis process.

Inventive Principle:
Principle #1Segmentation

3Manufacturing precision

If single-stage diffusion models are used, then the model can directly synthesize high-resolution images, but the model capacity may be insufficient to handle intricate warps and occlusions

Engineering Contradiction:
Improvehigh-resolution image qualityVSAvoidmodel capacity requirements
Core Design Contradiction:
Manufacturing precisionVSDevice complexity

Solution Approach 1:

The patent applies preliminary action by pre-processing the input images into high-resolution patches and preparing them for diffusion model processing. The model is pre-trained on high-resolution data, so when it receives the patch inputs, it can directly generate the final high-resolution output without requiring additional refinement stages, thus managing complexity while maintaining quality.

Inventive Principle:
Principle #10Preliminary action

Data Source

PatentUS20250299302A1Diffusion Models for Multi-Garment Virtual Try-On or Editing
Publication Date: 2025.09.25 GOOGLE LLC
  • US20250299302A1 patent drawing
  • US20250299302A1 patent drawing
  • US20250299302A1 patent drawing

AI summary

Provided are systems and methods for multi-garment virtual try-on and editing, example implementations of which can be referred to as M&M VTO. The proposed systems allow users to visualize how various combinations of garments would look on a given person. The input for this method can include multiple garment images, an image of a person, and optionally a text description for the garment layout. The output is a high-resolution visualization of how these garments would look on the person in the desired layout. For instance, a user can input an image of a shirt, an image of a pair of pants, a description such as “rolled sleeves, shirt tucked in”, and an image of a person. The output would then be a visual representation of how the person would look wearing these garments in the specified layout.