AR Visual Sound Selection for Smart Speaker Voice Commands

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current AI voice assistance systems struggle to differentiate between user-submitted voice commands and additional suggestions or feedback from surrounding users, leading to confusion about which commands to execute and which to ignore.

Innovation Solution

The system employs augmented reality (AR) to visualize and select sounds from surrounding transducers, allowing users to generate an augmented voice command that includes only the desired inputs while ignoring irrelevant sounds, using AR glasses to display and manage voice commands and feedback.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Adaptability or versatility

If the AI voice assistance system accepts all surrounding sounds as voice commands, then the system can capture more user inputs and suggestions, but the system cannot differentiate between the original user's command and additional feedback from other users, leading to execution confusion

Engineering Contradiction:
Improvecapability to process multiple voice inputsVSAvoidaccuracy in identifying which command to execute
Core Design Contradiction:
Adaptability or versatilityVSReliability

Solution Approach 1:

The patent segments the mixed voice inputs by identifying and separating the original user's command from subsequent feedback sounds using transducer arrays and time-stamped audio capture. This segmentation allows the system to process multiple inputs while maintaining clear distinction between command and feedback, resolving the contradiction between capturing versatile inputs and reliably identifying the executable command.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent introduces an augmented reality interface as an intermediary that visually presents separated voice inputs to the user. This intermediary allows users to review and selectively include or exclude specific sounds from the mixed input, enabling reliable command identification while preserving the system's ability to capture diverse surrounding inputs.

Inventive Principle:
Principle #24Intermediary (Mediator)

2Ease of operation

If the system visualizes all surrounding sounds using AR, then the user can selectively identify desired inputs, but the device complexity increases due to integration of AR components and sound processing

Engineering Contradiction:
Improveuser ability to select voice inputsVSAvoidintegration of AR and sound processing components
Core Design Contradiction:
Ease of operationVSDevice complexity

Solution Approach 1:

The patent employs a smart speaker device that performs multiple functions: capturing voice commands, processing audio signals, and integrating with augmented reality interfaces. This multi-functionality reduces the need for separate dedicated devices, thereby managing complexity while enabling comprehensive sound visualization and selection capabilities.

Inventive Principle:
Principle #6Universality (Multi-functionality)

Solution Approach 2:

The system automatically processes and separates voice inputs using built-in transducer arrays and audio analysis algorithms, reducing the manual effort required for sound selection. The self-service audio processing minimizes the operational burden on users while maintaining ease of selecting desired inputs through the AR interface.

Inventive Principle:
Principle #25Self-service

3Productivity

If the system processes complex mixed voice inputs in real-time, then the user experience is enhanced, but the processing time and computational resources increase

Engineering Contradiction:
Improvereal-time voice command processingVSAvoidtime for processing and separating sounds
Core Design Contradiction:
ProductivityVSLoss of time

Solution Approach 1:

The patent captures and time-stamps voice inputs as they occur, performing preliminary separation and organization of sounds before final processing. This preliminary action reduces the computational burden during real-time execution, enabling enhanced user experience without excessive processing delays.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The system implements real-time audio processing by skipping non-essential analysis steps for sounds that are clearly not commands (based on timing and acoustic characteristics). This selective processing rushes through obvious feedback sounds quickly while applying more thorough analysis only to potential commands, maintaining productivity while minimizing time loss.

Inventive Principle:
Principle #21Skipping (Rushing through)

Data Source

PatentUS11978444B2AR (augmented reality) based selective sound inclusion from the surrounding while executing any voice command
Publication Date: 2024.05.07 INTERNATIONAL BUSINESS MACHINE CORPORATION
  • US11978444B2 patent drawing
  • US11978444B2 patent drawing
  • US11978444B2 patent drawing

AI summary

A method, system and apparatus to generate an augmented voice command, including identifying a plurality of sounds from a respective plurality of transducers to a smart speaker device, generating a visualization of the sounds using an augmented reality device, wherein one or more of the sounds can be selected using the visualization, and generating the augmented voice command for the smart speaker device, wherein the augmented voice command comprises the one or more sounds selected using the visualization of the augmented reality device.