Voice Command Resolution via Non-Speech Sound Analysis in IoT
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current voice assistant solutions in IoT environments fail to consider surrounding non-speech sounds for optimal resolution of voice commands, leading to incomplete or incorrect execution of user requests.
Innovation Solution
A voice command resolution method and apparatus that recognizes voice commands, analyzes non-speech sounds, and determines target IoT devices to execute the commands, even when the voice command is incomplete or mixed with ambient sounds.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If voice assistant only processes speech commands without analyzing non-speech sounds, then device complexity is reduced, but voice command resolution accuracy deteriorates
Solution Approach 1:
The audio input is segmented into speech components and non-speech sound components. The speech recognizer processes only the speech portion for command extraction, while a separate analyzer processes non-speech sounds for context information. This segmentation allows the system to improve resolution accuracy by incorporating non-speech analysis without overwhelming the speech recognition process with complex general audio analysis.
Solution Approach 2:
An intermediary processing layer is introduced that takes the raw audio input, separates speech from non-speech components, and routes them to appropriate processing modules. This intermediary structure enables the system to handle both speech commands and non-speech context simultaneously, improving overall command resolution accuracy while maintaining manageable system complexity through modular architecture.
2Measurement precision
If voice assistant analyzes non-speech sounds to determine target devices, then command execution accuracy is improved, but processing time increases
Solution Approach 1:
The system performs preliminary analysis of non-speech sounds continuously in the background before a voice command is issued. Environmental sound profiles are pre-established and stored, so when a voice command is given, the system can quickly match the current non-speech audio against pre-analyzed contexts, significantly reducing the time required to identify target devices while maintaining high accuracy.
Solution Approach 2:
The patent replaces exhaustive real-time analysis of all audio components with a more efficient approach: speech recognition handles command extraction while a lighter non-speech sound detector identifies environmental context. This substitution of heavy mechanical analysis with specialized, optimized detectors reduces processing time while maintaining accurate target device identification.
3Ease of operation
If voice assistant considers surrounding non-speech sounds, then user experience is improved, but system reliability requirements increase
Solution Approach 1:
The system applies partial analysis of non-speech sounds rather than attempting to analyze every audio component in detail. It focuses on detecting prominent environmental sounds and contexts that are most relevant to command resolution, ignoring less significant audio elements. This partial action approach improves user experience by considering relevant environmental context while avoiding the reliability issues that would arise from attempting comprehensive analysis of all sound components.
Data Source
AI summary
A voice command resolution apparatus, including a memory configured to store instructions; and a processor configured to execute the instructions to: recognize a voice command of a user in an input sound, analyze a non-speech sound included in the input sound, and determine at least one target Internet of things (IoT) device related to execution of the voice command, based on an analysis result of the non-speech sound.


