Real-time Garment Exchange via Neural Network Segmentation
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing augmented reality (AR) systems struggle to efficiently apply visual effects to users' bodies without the need for depth sensors, which increases costs and complexity. These systems often fail to recognize whole-body users, leading to poor image quality and incorrect identification of body parts as background.
Innovation Solution
The use of machine learning techniques, such as neural networks, to simultaneously extract appearance and motion features of a person in an image, allowing for real-time application of visual effects without generating a rig or bone structure. This enables seamless addition of AR graphics to images or videos on small-scale mobile devices.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If depth sensors are used to apply visual effects to users' bodies, then the accuracy of body part recognition is improved, but the cost and complexity of the system increases
Solution Approach 1:
The patent extracts and removes the depth sensor component from the AR system, replacing it with machine learning techniques that process standard 2D images. This eliminates the need for expensive depth sensing hardware while maintaining body recognition functionality through neural network-based pose estimation and segmentation algorithms.
Solution Approach 2:
The patent substitutes the mechanical/optical depth sensing system with a computational approach using machine learning models. Instead of using physical depth sensors to capture 3D information, the system uses 2D image processing with trained neural networks to infer body pose, shape, and garment boundaries, replacing hardware-based measurement with software-based computation.
2Speed
If traditional AR systems process visual effects in real-time, then the user experience is improved, but the processing power and energy consumption increases
Solution Approach 1:
The patent implements preliminary action by pre-training machine learning models offline to extract appearance and motion features. During real-time operation, these pre-trained models perform rapid inference on incoming images, significantly reducing the computational burden and energy consumption during actual AR rendering while maintaining real-time performance.
3Manufacturing precision
If rigid rig or bone structure generation is used to apply visual effects, then the precision of garment fitting is improved, but the design constraints and complexity increases
Solution Approach 1:
The patent replaces static rig or bone structure models with dynamic machine learning-based pose estimation. The system uses trained neural networks to directly infer body pose, shape, and garment boundaries from images, allowing the garment to adapt dynamically to different body positions and movements without requiring complex rigid structure definitions or manual rigging.
Data Source
AI summary
Methods and systems are disclosed for performing operations for transferring garments in a video from one real-world object to another in real time. The operations comprise receiving a first video that includes a depiction of a first person wearing a first garment in a first pose and obtaining a second video that includes a depiction of a second person wearing a second garment in a second pose. The operations comprise modifying a pose of the second person to match the first pose of the first person depicted in the first video. The operations comprise generating a whole-body segmentation of the second garment which the second person is wearing in the second video and changing an appearance of the first person from wearing the first garment to wearing the second garment based on the whole-body segmentation of the second garment which the second person is wearing in the second video.


