Diffusion-Generated XR Try-On Images With Product Image Replacement

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing systems require significant user effort and resources to create high-quality images, leading to inefficiencies and missed opportunities for sharing and presenting real-world objects in ideal settings, often resulting in lower quality images that undervalue the objects.

Innovation Solution

A generative machine learning model that analyzes real-world object images and textual descriptions to automatically generate photorealistic images of the objects wearing artificial fashion items, reducing the need for manual adjustments and resource-intensive image creation.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Ease of operation

If manual image creation methods are used, then users can create customized images, but significant user effort and resources are required

Engineering Contradiction:
Improveuser effortVSAvoidimage creation efficiency
Core Design Contradiction:
Ease of operationVSProductivity

Solution Approach 1:

The system automatically generates images by having the object serve itself - the real-world object is photographed, its features are extracted, and the system autonomously generates virtual scenes with the object without requiring manual manipulation or design work from users

Inventive Principle:
Principle #25Self-service

Solution Approach 2:

Manual mechanical processes of image creation (photographing, positioning, lighting adjustments) are replaced with an automated computational system that uses machine learning to generate photorealistic images from simple input photographs

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

2Manufacturing precision

If manual image creation methods are used, then users can control image quality, but resource-intensive processes are required

Engineering Contradiction:
Improveimage qualityVSAvoidresource consumption
Core Design Contradiction:
Manufacturing precisionVSUse of energy by moving object

Solution Approach 1:

Instead of creating images through resource-intensive manual processes, the system creates accurate copies by extracting features from a single input photograph and replicating the object's appearance, lighting, and texture in virtual environments

Inventive Principle:
Principle #26Copying

Solution Approach 2:

The system changes the fundamental parameters of image creation from manual control of multiple physical variables to automated computational parameter optimization, achieving high-quality results with fewer physical resources

Inventive Principle:
Principle #35Parameter changes

3Ease of operation

If automated image generation is implemented, then user effort is reduced, but system complexity increases

Engineering Contradiction:
Improveuser interactionVSAvoidsystem complexity
Core Design Contradiction:
Ease of operationVSDevice complexity

Solution Approach 1:

The system achieves universality by creating a multi-functional platform that can handle various types of objects (products, fashion items, furniture), generate different scene types, and produce multiple image variations from a single input photograph

Inventive Principle:
Principle #6Universality (Multi-functionality)

Solution Approach 2:

The system introduces an intermediary layer of feature extraction and scene generation technology that mediates between the simple input photograph and the complex output images, managing system complexity while maintaining ease of use

Inventive Principle:
Principle #24Intermediary (Mediator)

Data Source

PatentUS12475621B2Product image generation based on diffusion model
Publication Date: 2025.11.18 SNAP INC
  • US12475621B2 patent drawing
  • US12475621B2 patent drawing
  • US12475621B2 patent drawing

AI summary

Methods and systems are disclosed for generating an extended reality (XR) try-on experience based on an image produced by a diffusion model. The system receives an image depicting a real-world object and generates a prompt comprising a textual description of a fashion item. The system analyzes the image and the textual description of the fashion item using a generative machine learning model to generate an artificial image that depicts an artificial object that resembles the real-world object wearing an artificial fashion item matching the textual description of the fashion item. The system identifies an object comprising a real-world product image that matches visual attributes of the artificial fashion item and replaces the artificial fashion item in the artificial image with the object to generate an output image.