Voice-Controlled Garment Augmentation Without Depth Sensors

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing augmented reality systems require depth sensors to modify images, increasing device cost and complexity, and struggle to recognize and apply visual effects to a user's whole body, especially when multiple users are present or at varying distances from the camera.

Innovation Solution

The system segments articles of clothing or garments worn by a user using machine learning techniques without depth sensors, allowing for the application of augmented reality elements based on voice input, and tracks the position of these items separately from the user's body parts.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Reliability

If depth sensors are used to modify images in augmented reality systems, then image modification capability is improved, but device cost and complexity increase

Engineering Contradiction:
Improveimage modification capabilityVSAvoiddevice cost and complexity
Core Design Contradiction:
ReliabilityVSDevice complexity

Solution Approach 1:

The patent extracts the depth sensing function from dedicated hardware sensors and implements it through software-based monocular depth estimation using standard RGB cameras. This removes the requirement for expensive depth sensors while maintaining the core functionality of depth-aware image modification in augmented reality systems.

Inventive Principle:
Principle #2Taking out (Extraction)

Solution Approach 2:

The patent replaces the mechanical/optical depth sensing system with a computational approach using machine learning models that process standard RGB images to generate depth maps. This substitution eliminates the need for specialized depth sensing hardware while achieving similar functional outcomes.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

2Measurement precision

If depth sensors are used to recognize user's whole body, then recognition accuracy is improved, but device complexity increases

Engineering Contradiction:
Improverecognition accuracyVSAvoiddevice complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The patent extracts the depth information needed for whole-body recognition from specialized depth sensors and obtains it instead through monocular depth estimation algorithms processing standard camera feeds, thereby reducing hardware complexity while maintaining recognition capability.

Inventive Principle:
Principle #2Taking out (Extraction)

Solution Approach 2:

The patent changes the input parameters from depth sensor data to RGB image data, using machine learning models trained to infer depth and spatial relationships from color images. This parameter transformation enables whole-body recognition without requiring depth sensing hardware.

Inventive Principle:
Principle #35Parameter changes

3Adaptability or versatility

If multiple users are present at varying distances, then system versatility is improved, but difficulty of detecting and measuring increases

Engineering Contradiction:
Improvesystem versatilityVSAvoiddifficulty of detecting and measuring
Core Design Contradiction:
Adaptability or versatilityVSDifficulty of detecting and measuring

Solution Approach 1:

The patent segments the scene into multiple user instances with individual depth estimates and spatial relationships. The monocular depth estimation model processes each user separately, determining their distances and positions independently, which enables versatile multi-user support while managing detection complexity through divide-and-conquer processing.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent transforms the measurement parameters from direct depth sensor readings to inferred depth values derived from RGB image analysis. This parameter change enables the system to handle multiple users at varying distances using standard cameras, improving versatility without proportionally increasing measurement difficulty.

Inventive Principle:
Principle #35Parameter changes

Data Source

PatentUS12380618B2Controlling interactive fashion based on voice
Publication Date: 2025.08.05 SNAP INC
  • US12380618B2 patent drawing
  • US12380618B2 patent drawing
  • US12380618B2 patent drawing

AI summary

Methods and systems are disclosed for performing operations comprising: receiving an image that includes a depiction of a person wearing a fashion item; generating a segmentation of the fashion item by the person depicted in the image; receiving voice input associated with the person depicted in the image; in response to receiving the voice input, generating one or more augmented reality elements representing the voice input; and applying the one or more augmented reality elements to the fashion item worn by the person based on the segmentation of the fashion item worn by the person.