Voice-Controlled Garment Augmentation Without Depth Sensors
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing augmented reality systems require depth sensors to modify images, increasing device cost and complexity, and struggle to recognize and apply visual effects to a user's whole body, especially when multiple users are present or at varying distances from the camera.
Innovation Solution
The system segments articles of clothing or garments worn by a user using machine learning techniques without depth sensors, allowing for the application of augmented reality elements based on voice input, and tracks the position of these items separately from the user's body parts.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If depth sensors are used to modify images in augmented reality systems, then image modification capability is improved, but device cost and complexity increase
Solution Approach 1:
The patent extracts the depth sensing function from dedicated hardware sensors and implements it through software-based monocular depth estimation using standard RGB cameras. This removes the requirement for expensive depth sensors while maintaining the core functionality of depth-aware image modification in augmented reality systems.
Solution Approach 2:
The patent replaces the mechanical/optical depth sensing system with a computational approach using machine learning models that process standard RGB images to generate depth maps. This substitution eliminates the need for specialized depth sensing hardware while achieving similar functional outcomes.
2Measurement precision
If depth sensors are used to recognize user's whole body, then recognition accuracy is improved, but device complexity increases
Solution Approach 1:
The patent extracts the depth information needed for whole-body recognition from specialized depth sensors and obtains it instead through monocular depth estimation algorithms processing standard camera feeds, thereby reducing hardware complexity while maintaining recognition capability.
Solution Approach 2:
The patent changes the input parameters from depth sensor data to RGB image data, using machine learning models trained to infer depth and spatial relationships from color images. This parameter transformation enables whole-body recognition without requiring depth sensing hardware.
3Adaptability or versatility
If multiple users are present at varying distances, then system versatility is improved, but difficulty of detecting and measuring increases
Solution Approach 1:
The patent segments the scene into multiple user instances with individual depth estimates and spatial relationships. The monocular depth estimation model processes each user separately, determining their distances and positions independently, which enables versatile multi-user support while managing detection complexity through divide-and-conquer processing.
Solution Approach 2:
The patent transforms the measurement parameters from direct depth sensor readings to inferred depth values derived from RGB image analysis. This parameter change enables the system to handle multiple users at varying distances using standard cameras, improving versatility without proportionally increasing measurement difficulty.
Data Source
AI summary
Methods and systems are disclosed for performing operations comprising: receiving an image that includes a depiction of a person wearing a fashion item; generating a segmentation of the fashion item by the person depicted in the image; receiving voice input associated with the person depicted in the image; in response to receiving the voice input, generating one or more augmented reality elements representing the voice input; and applying the one or more augmented reality elements to the fashion item worn by the person based on the segmentation of the fashion item worn by the person.


