AR Visual Sound Selection for Smart Speaker Voice Commands
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current AI voice assistance systems struggle to differentiate between user-submitted voice commands and additional suggestions or feedback from surrounding users, leading to confusion about which commands to execute and which to ignore.
Innovation Solution
The system employs augmented reality (AR) to visualize and select sounds from surrounding transducers, allowing users to generate an augmented voice command that includes only the desired inputs while ignoring irrelevant sounds, using AR glasses to display and manage voice commands and feedback.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Adaptability or versatility
If the AI voice assistance system accepts all surrounding sounds as voice commands, then the system can capture more user inputs and suggestions, but the system cannot differentiate between the original user's command and additional feedback from other users, leading to execution confusion
Solution Approach 1:
The patent segments the mixed voice inputs by identifying and separating the original user's command from subsequent feedback sounds using transducer arrays and time-stamped audio capture. This segmentation allows the system to process multiple inputs while maintaining clear distinction between command and feedback, resolving the contradiction between capturing versatile inputs and reliably identifying the executable command.
Solution Approach 2:
The patent introduces an augmented reality interface as an intermediary that visually presents separated voice inputs to the user. This intermediary allows users to review and selectively include or exclude specific sounds from the mixed input, enabling reliable command identification while preserving the system's ability to capture diverse surrounding inputs.
2Ease of operation
If the system visualizes all surrounding sounds using AR, then the user can selectively identify desired inputs, but the device complexity increases due to integration of AR components and sound processing
Solution Approach 1:
The patent employs a smart speaker device that performs multiple functions: capturing voice commands, processing audio signals, and integrating with augmented reality interfaces. This multi-functionality reduces the need for separate dedicated devices, thereby managing complexity while enabling comprehensive sound visualization and selection capabilities.
Solution Approach 2:
The system automatically processes and separates voice inputs using built-in transducer arrays and audio analysis algorithms, reducing the manual effort required for sound selection. The self-service audio processing minimizes the operational burden on users while maintaining ease of selecting desired inputs through the AR interface.
3Productivity
If the system processes complex mixed voice inputs in real-time, then the user experience is enhanced, but the processing time and computational resources increase
Solution Approach 1:
The patent captures and time-stamps voice inputs as they occur, performing preliminary separation and organization of sounds before final processing. This preliminary action reduces the computational burden during real-time execution, enabling enhanced user experience without excessive processing delays.
Solution Approach 2:
The system implements real-time audio processing by skipping non-essential analysis steps for sounds that are clearly not commands (based on timing and acoustic characteristics). This selective processing rushes through obvious feedback sounds quickly while applying more thorough analysis only to potential commands, maintaining productivity while minimizing time loss.
Data Source
AI summary
A method, system and apparatus to generate an augmented voice command, including identifying a plurality of sounds from a respective plurality of transducers to a smart speaker device, generating a visualization of the sounds using an augmented reality device, wherein one or more of the sounds can be selected using the visualization, and generating the augmented voice command for the smart speaker device, wherein the augmented voice command comprises the one or more sounds selected using the visualization of the augmented reality device.


