Garment Segmentation via Temporal Smoothing for AR Background Replacement
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing augmented reality systems struggle to accurately recognize and replace the background of a user's whole body in images without depth sensors, leading to poor image quality and failure in multi-user scenarios, as they are unable to distinguish between the user's body and the background effectively.
Innovation Solution
The system employs machine learning techniques to segment articles of clothing in images, allowing for the application of visual effects and background replacement without depth sensors by processing current and previous video frames to improve segmentation accuracy.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Device complexity
If existing augmented reality systems use standard background replacement techniques without depth sensors, then device complexity is reduced, but segmentation precision deteriorates leading to poor image quality and failure to distinguish user body from background
Solution Approach 1:
The patent applies segmentation by dividing the image processing task into multiple stages: initial garment segmentation, temporal consistency refinement using previous frames, and boundary optimization. This multi-stage segmentation approach improves precision without requiring depth sensors, resolving the contradiction between device complexity and segmentation precision.
Solution Approach 2:
The system performs preliminary actions by processing previous video frames to establish temporal context before analyzing the current frame. This preliminary processing creates a foundation for more accurate current frame segmentation, improving precision while maintaining simple device architecture.
2Measurement precision
If machine learning techniques are applied to segment garments in real-time video frames, then segmentation precision is improved, but computing power consumption increases
Solution Approach 1:
The patent implements continuity of useful action by maintaining temporal coherence across video frames. Instead of independently processing each frame, the system continuously refines segmentation using previous frame results, reducing redundant computations and energy consumption while maintaining high precision.
Solution Approach 2:
The system applies parameter changes by adjusting processing intensity based on detected motion and scene complexity. When motion is minimal or patterns are consistent across frames, processing parameters are reduced, lowering energy consumption while preserving segmentation precision through temporal smoothing.
3Speed
If standard image processing is used without temporal analysis, then processing speed is maintained, but segmentation reliability deteriorates in multi-user and complex background scenarios
Solution Approach 1:
The system performs preliminary analysis of temporal patterns using previous frames before final segmentation decisions. This preliminary temporal context establishment enables faster and more reliable segmentation in complex scenarios without significantly impacting processing speed.
Solution Approach 2:
The patent implements feedback mechanisms where segmentation results from previous frames inform current frame processing. This feedback loop continuously refines segmentation reliability across the video sequence, handling multi-user and complex background scenarios effectively while maintaining real-time performance.
Data Source
AI summary
Methods and systems are disclosed for performing operations comprising: receiving a monocular image that includes a depiction of a user wearing a garment; generating a segmentation of the garment worn by the user in the monocular image; accessing a video feed comprising a plurality of monocular images received prior to the monocular image; smoothing, using the video feed, the segmentation of the garment worn by the user to provide a smoothed segmentation of the garment worn by the user; and applying one or more visual effects to the monocular image based on the smoothed segmentation of the garment worn by the user.


