Multi-Device Voice Command Selection via Audio Energy Segmentation
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
In multi-device systems with intelligent personal assistants, voice commands intended for one device can be mistakenly received and executed by another nearby device, leading to conflicts and reduced convenience.
Innovation Solution
A method that involves receiving audio signals from multiple microphones, dividing them into temporal segments, comparing sound energy levels, and selecting the segment with the strongest signal to determine the closest microphone to the user, ensuring that only the intended device responds to voice commands without requiring specific location information in the command.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Adaptability or versatility
If multiple IPA-enabled devices are deployed in proximate locations, then device coverage and accessibility are improved, but conflicts arise when devices mistakenly receive and execute voice commands intended for other devices
Solution Approach 1:
The patent applies local quality by making each device's audio response characteristics unique to its local environment. Devices transmit audio samples captured by their microphones, and the system selects the sample with the highest energy level that matches the user's voice characteristics. This ensures that the device locally closest to the user (with the best audio capture quality) is identified as the intended target, preventing cross-device command execution while maintaining multi-device deployment.
2Ease of operation
If voice commands are made detectable by multiple devices, then user convenience is improved, but conflicts arise when multiple devices respond to the same command
Solution Approach 1:
The patent implements feedback by having each device transmit its audio sample back to the system for evaluation. The system compares audio energy levels and voice characteristics across multiple devices, then selects the single best-matching device. This feedback mechanism ensures that while multiple devices can detect the command, only one device executes it based on objective audio quality metrics, eliminating conflicting responses while preserving broad detectability.
Applied Scientific Principles
This section explains which scientific principles are used to turn an abstract innovation direction into a practical engineering solution.
Function Achieved in This Case
This approach allows users to issue voice commands detectable by multiple devices while receiving a single response from the correct device, preventing conflicts and enhancing the efficiency of multi-device systems.
Implementation Method 1
receiving a first audio signal that is generated by a first microphone in response to a verbal utterance, and a second audio signal that is generated by a second microphone in response to the verbal utterance
Data Source
Figure 1
Figure 2
Figure 3
AI summary
Performing speech recognition in a multi-device system includes receiving a first audio signal that is generated by a first microphone in response to a verbal utterance, and a second audio signal that is generated by a second microphone in response to the verbal utterance; dividing the first audio signal into a first sequence of temporal segments; dividing the second audio signal into a second sequence of temporal segments; comparing a sound energy level associated with a first temporal segment of the first sequence to a sound energy level associated with a first temporal segment of the second sequence; based on the comparing, selecting, as a first temporal segment of a speech recognition audio signal, one of the first temporal segment of the first sequence and the first temporal segment of the second sequence; and performing speech recognition on the speech recognition audio signal.