Voice Gesture Recognition in Augmented Reality

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current augmented reality systems lack effective methods for user interaction beyond physical gestures, limiting the ability to seamlessly integrate voice commands and gestures for controlling virtual and physical objects within the environment.

Innovation Solution

The implementation of voice gestures based on sound characteristics such as pitch, volume, and directionality, which are recognized and interpreted by the augmented reality system to execute commands like zoom, pan, and navigation, allowing users to interact with projected images and control environmental attributes through voice inputs.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Adaptability or versatility

If voice gestures are implemented in augmented reality systems, then user interaction versatility is improved, but system complexity increases

Engineering Contradiction:
Improveuser interaction versatilityVSAvoidsystem complexity
Core Design Contradiction:
Adaptability or versatilityVSDevice complexity

Solution Approach 1:

The patent introduces voice gestures as an intermediary interaction modality between the user and the augmented reality system. The voice gesture processor acts as a mediator that receives audio signals, interprets them as interaction commands, and translates them into system actions, thereby enhancing interaction versatility without requiring direct complex hardware modifications

Inventive Principle:
Principle #24Intermediary (Mediator)

Solution Approach 2:

The voice gesture system enables the augmented reality system to perform multiple interaction functions through a single unified interface. The same voice processing infrastructure handles various types of commands (navigation, object manipulation, information queries), making the system multi-functional while maintaining a consistent interaction paradigm

Inventive Principle:
Principle #6Universality (Multi-functionality)

2Ease of operation

If multiple interaction methods are integrated, then ease of operation is improved, but device complexity increases

Engineering Contradiction:
Improveease of operationVSAvoiddevice complexity
Core Design Contradiction:
Ease of operationVSDevice complexity

Solution Approach 1:

The patent merges multiple interaction methods (voice gestures, physical gestures, and other input modalities) into a unified interaction framework. The gesture processor and voice gesture processor work together as integrated components, allowing users to switch between interaction modes seamlessly while the system maintains a consistent response paradigm, thereby improving ease of operation without requiring separate independent systems for each interaction type

Inventive Principle:
Principle #5Merging (Combining)

3Adaptability or versatility

If voice-based interaction is added, then adaptability of interaction methods is improved, but difficulty of detecting and measuring increases

Engineering Contradiction:
Improveinteraction method adaptabilityVSAvoidvoice gesture detection difficulty
Core Design Contradiction:
Adaptability or versatilityVSDifficulty of detecting and measuring

Solution Approach 1:

The patent replaces traditional mechanical or manual interaction detection methods with acoustic field-based voice gesture detection. Instead of relying on physical sensors that detect mechanical movements, the system uses audio signal processing to capture, analyze, and interpret voice gestures, thereby adding interaction adaptability while leveraging well-established acoustic detection technologies to manage the complexity of voice gesture measurement

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

Data Source

PatentUS9401144B1Voice gestures
Publication Date: 2016.07.26 AMAZON TECH INC
  • US9401144B1 patent drawing
  • US9401144B1 patent drawing
  • US9401144B1 patent drawing

AI summary

A voice gesture is determined from characteristics of an audio signal based on sound uttered by a user. The voice gesture may represent a command or parameters or a command, and may be context sensitive. Upon determining a command and parameters of the command based on the received voice gesture, the command is executed in accordance with the determined parameters. The command may modify any number of attributes within an environment including, but limited to, an image projected within the environment.