Virtual Try-On Image Generation Using Shape Keypoints

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Conventional virtual try-on technologies lack the ability to accurately represent how clothing will fit on a user due to the use of sparse keypoints that do not convey shape or size information, leading to inaccurate and inefficient image generation.

Innovation Solution

The use of shape keypoints across both user and garment images, combined with a cross-attention mechanism, to guide the generation of augmented images, ensuring accurate representation and efficient processing by focusing on relevant details.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If conventional virtual try-on technologies use sparse keypoints for image generation, then the processing speed is maintained, but the accuracy of clothing fit representation deteriorates

Engineering Contradiction:
Improveaccuracy of clothing fit representationVSAvoidcomplexity of keypoint system
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The patent segments the keypoint system into two distinct components: sparse keypoints for maintaining processing efficiency and shape keypoints for capturing clothing fit details. This segmentation allows each component to serve its specific function without compromising the other, resolving the contradiction between speed and accuracy.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent adds a new dimension to the keypoint system by introducing shape keypoints that capture curvature and fit information. This dimensional enhancement allows the system to represent clothing fit accuracy without abandoning the efficient sparse keypoint structure, effectively resolving the accuracy-complexity tradeoff.

Inventive Principle:
Principle #17Another dimension (Dimensionality change)

2Measurement precision

If conventional virtual try-on technologies use sparse keypoints, then the system complexity is reduced, but the image generation accuracy deteriorates

Engineering Contradiction:
Improveimage generation accuracyVSAvoidprocessing time
Core Design Contradiction:
Measurement precisionVSLoss of time

Solution Approach 1:

The patent divides the processing into two streams: one handling sparse keypoints for efficient pose alignment and another handling shape keypoints for accurate fit representation. This segmentation enables parallel processing that maintains speed while improving accuracy.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent introduces shape keypoints as an intermediary element that bridges the gap between sparse keypoints and detailed clothing representation. This intermediary captures the necessary fit information without requiring a complete dense keypoint system, thus improving accuracy without proportionally increasing processing time.

Inventive Principle:
Principle #24Intermediary (Mediator)

3Reliability

If shape keypoints are incorporated into the image generation process, then the representation accuracy improves, but the processing complexity increases

Engineering Contradiction:
Improverealism of augmented imagesVSAvoidcomplexity of image generation system
Core Design Contradiction:
ReliabilityVSDevice complexity

Solution Approach 1:

The patent applies local quality by using shape keypoints specifically for capturing clothing fit and curvature information where it is most needed, rather than uniformly increasing complexity across the entire system. This targeted approach improves realism without uniformly increasing system complexity.

Inventive Principle:
Principle #3Local quality

Solution Approach 2:

The patent performs preliminary extraction of shape keypoints from garment images before the main image generation process. This preliminary action prepares the fit information in advance, allowing the main generation process to use pre-processed data and reducing the computational burden during real-time processing.

Inventive Principle:
Principle #10Preliminary action

4Measurement precision

If conventional methods are used for virtual try-on, then the processing efficiency is maintained, but the image quality and accuracy deteriorate

Engineering Contradiction:
Improveclothing fit accuracyVSAvoidprocessing efficiency
Core Design Contradiction:
Measurement precisionVSProductivity

Solution Approach 1:

The patent segments the feature extraction process into sparse keypoint detection for pose estimation and shape keypoint detection for fit analysis. This segmentation allows each process to be optimized independently, maintaining processing efficiency while improving clothing fit accuracy.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent extracts shape keypoints as a separate copy of relevant fit information from the garment image, independent of the main image generation process. This copying approach allows the system to use pre-extracted fit features without duplicating the entire image processing pipeline, thus improving accuracy while maintaining efficiency.

Inventive Principle:
Principle #26Copying

Data Source

PatentUS12505629B1Model-based augmented image generation
Publication Date: 2025.12.23 AMAZON TECH INC
  • US12505629B1 patent drawing
  • US12505629B1 patent drawing
  • US12505629B1 patent drawing

AI summary

Techniques are described herein for generating a virtual try-on augmented image. An example method can include A system can access a first two-dimensional model of a target garment. The system can determine a first plurality of shape keypoints for a first part of the first plurality of parts of the image. The system can access a second two-dimensional image of a first user. The system can determine a second plurality of keypoints for a second part of the second plurality of parts of the second two-dimensional image, a second number of the second plurality of keypoints based at least in part on a second predetermined number of keypoints associated with the second part. The system can generate a rendered augmented image of the target garment on the target user.