Real-time Garment Exchange via Neural Network Segmentation

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing augmented reality (AR) systems struggle to efficiently apply visual effects to users' bodies without the need for depth sensors, which increases costs and complexity. These systems often fail to recognize whole-body users, leading to poor image quality and incorrect identification of body parts as background.

Innovation Solution

The use of machine learning techniques, such as neural networks, to simultaneously extract appearance and motion features of a person in an image, allowing for real-time application of visual effects without generating a rig or bone structure. This enables seamless addition of AR graphics to images or videos on small-scale mobile devices.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If depth sensors are used to apply visual effects to users' bodies, then the accuracy of body part recognition is improved, but the cost and complexity of the system increases

Engineering Contradiction:
Improvebody part recognition accuracyVSAvoidsystem complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The patent extracts and removes the depth sensor component from the AR system, replacing it with machine learning techniques that process standard 2D images. This eliminates the need for expensive depth sensing hardware while maintaining body recognition functionality through neural network-based pose estimation and segmentation algorithms.

Inventive Principle:
Principle #2Taking out (Extraction)

Solution Approach 2:

The patent substitutes the mechanical/optical depth sensing system with a computational approach using machine learning models. Instead of using physical depth sensors to capture 3D information, the system uses 2D image processing with trained neural networks to infer body pose, shape, and garment boundaries, replacing hardware-based measurement with software-based computation.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

2Speed

If traditional AR systems process visual effects in real-time, then the user experience is improved, but the processing power and energy consumption increases

Engineering Contradiction:
Improvevisual effects application speedVSAvoidprocessing energy consumption
Core Design Contradiction:
SpeedVSUse of energy by moving object

Solution Approach 1:

The patent implements preliminary action by pre-training machine learning models offline to extract appearance and motion features. During real-time operation, these pre-trained models perform rapid inference on incoming images, significantly reducing the computational burden and energy consumption during actual AR rendering while maintaining real-time performance.

Inventive Principle:
Principle #10Preliminary action

3Manufacturing precision

If rigid rig or bone structure generation is used to apply visual effects, then the precision of garment fitting is improved, but the design constraints and complexity increases

Engineering Contradiction:
Improvegarment fitting precisionVSAvoiddesign complexity
Core Design Contradiction:
Manufacturing precisionVSDevice complexity

Solution Approach 1:

The patent replaces static rig or bone structure models with dynamic machine learning-based pose estimation. The system uses trained neural networks to directly infer body pose, shape, and garment boundaries from images, allowing the garment to adapt dynamically to different body positions and movements without requiring complex rigid structure definitions or manual rigging.

Inventive Principle:
Principle #15Dynamics

Data Source

PatentUS20250139813A1Real-time garment exchange
Publication Date: 2025.05.01 SNAP INC
  • US20250139813A1 patent drawing
  • US20250139813A1 patent drawing
  • US20250139813A1 patent drawing

AI summary

Methods and systems are disclosed for performing operations for transferring garments in a video from one real-world object to another in real time. The operations comprise receiving a first video that includes a depiction of a first person wearing a first garment in a first pose and obtaining a second video that includes a depiction of a second person wearing a second garment in a second pose. The operations comprise modifying a pose of the second person to match the first pose of the first person depicted in the first video. The operations comprise generating a whole-body segmentation of the second garment which the second person is wearing in the second video and changing an appearance of the first person from wearing the first garment to wearing the second garment based on the whole-body segmentation of the second garment which the second person is wearing in the second video.