Voice Command Disambiguation via Background Audio Context
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing voice command processing systems face ambiguity when multiple voice-controllable services use the same voice command phrase, leading to confusion and frustration for users.
Innovation Solution
A voice command processing system that stores information on various voice command phrases used by different services and utilizes background noises captured during voice commands to interpret and process the commands, selecting the appropriate service when ambiguity occurs.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Adaptability or versatility
If multiple voice-controllable services use the same voice command phrase, then service availability and versatility increase, but command interpretation accuracy deteriorates due to ambiguity
Solution Approach 1:
The patent introduces background noise analysis as an intermediary mechanism to resolve ambiguity between multiple voice commands. When a voice command could match multiple services, the system analyzes background noises (such as TV audio, music playback, or other device sounds) to determine which service is currently active or most relevant, thereby selecting the correct target service for the voice command.
Solution Approach 2:
The patent adds a new dimension of information (background noise analysis) to the voice command processing system. Instead of relying solely on the voice command text, the system now considers acoustic environment data as an additional parameter for disambiguation, effectively moving from a one-dimensional command matching approach to a multi-dimensional decision framework.
2Measurement precision
If background noise analysis is used to resolve voice command ambiguity, then command interpretation accuracy improves, but system complexity increases
Solution Approach 1:
The system performs preliminary analysis of background noises continuously or periodically before a voice command is issued. By pre-processing and storing information about the acoustic environment (such as identifying currently playing media, active devices, or characteristic sound patterns), the system reduces the computational burden during actual voice command processing, as the disambiguation can leverage pre-analyzed context rather than performing complex analysis in real-time.
Solution Approach 2:
The voice command processing system utilizes already-available background noise data that is naturally present in the environment. Rather than requiring additional dedicated sensors or complex hardware, the system repurposes existing audio capture capabilities and leverages the inherently present acoustic information to perform self-disambiguation, reducing the need for additional system components.
3Loss of information
If background noises are captured and analyzed, then contextual information for command interpretation improves, but processing time increases
Solution Approach 1:
The system performs preliminary analysis of background noises continuously or periodically before a voice command is issued. By pre-processing and storing information about the acoustic environment (such as identifying currently playing media, active devices, or characteristic sound patterns), the system reduces the computational burden during actual voice command processing, as the disambiguation can leverage pre-analyzed context rather than performing complex analysis in real-time.
Data Source
Figure 1
Figure 2
Figure 3A
AI summary
Recorded background noises, and other contextual data, may be used to assist in resolving ambiguity in spoken voice commands. The background noises may comprise sounds from entities in a room other than the user issuing the voice commands. One such entity may be a content item being watched by the user, and the captured background noises may comprise audio of the content item. The content item may be identified based on the captured audio of the content item in the background noises, and the identification may be used to interpret the ambiguous voice command. Additional contextual information associated with the voice commands (e.g., identifications of the users in the room) and/or the content item (e.g., the video quality of the content item, a service outputting the content item, a genre of the content item, etc.) may be used to identify the content item.