Voice Device Selection via Bifurcated Audio Processing
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
In environments with multiple voice-enabled devices, determining which device should respond to a speech utterance can be challenging, especially when devices have different capabilities and are positioned near sound-emitting secondary devices, leading to poor signal-to-noise ratios and ambiguous commands.
Innovation Solution
A remote speech-processing system performs bifurcated processing using multiple audio-signal processing pipelines and natural language understanding models tailored to specific device capabilities to select the appropriate voice-enabled device for responding to a speech utterance, considering signal-to-noise ratios, device states, and contextual data.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Adaptability or versatility
If multiple voice-enabled devices are deployed in proximity to provide comprehensive voice command coverage, then user accessibility and device coverage are improved, but device selection ambiguity and response accuracy deteriorate
Solution Approach 1:
The system segments the device selection process into multiple independent analysis dimensions: audio signal quality assessment, command relevance evaluation, and device capability matching. Each dimension independently evaluates different aspects and contributes to the final selection, allowing comprehensive coverage while maintaining precise selection through modular analysis
Solution Approach 2:
The system dynamically changes evaluation parameters based on real-time conditions, adjusting the weight of different selection criteria according to signal-to-noise ratios, device states, and command context. This allows the system to adapt to varying environmental conditions and maintain accurate device selection across diverse scenarios
2Area of stationary object
If devices are positioned near sound-emitting secondary devices to expand coverage area, then spatial coverage is improved, but signal-to-noise ratio and voice command detection quality worsen
Solution Approach 1:
The system performs preliminary audio quality assessment and device state evaluation before final device selection. By pre-analyzing signal-to-noise ratios and device capabilities, the system can identify suitable receiving devices in advance, compensating for poor audio conditions caused by proximity to sound-emitting devices
Solution Approach 2:
The system implements feedback mechanisms that continuously monitor audio signal quality and device performance. Based on real-time feedback about signal-to-noise ratios and detection quality, the system dynamically adjusts device selection decisions, allowing devices to be positioned for maximum coverage while maintaining selection accuracy through continuous quality monitoring
3Measurement precision
If bifurcated processing with multiple analysis pipelines is implemented to improve device selection accuracy, then response accuracy is improved, but system complexity and processing overhead worsen
Solution Approach 1:
The complex device selection process is segmented into distinct analysis pipelines, each responsible for specific evaluation tasks such as audio quality assessment, command relevance analysis, and device capability matching. This segmentation reduces overall system complexity by dividing the problem into manageable, specialized components
Solution Approach 2:
The system employs universal processing frameworks and shared analysis components that can be applied across multiple device types and scenarios. By creating multi-functional processing pipelines that handle various command types and device configurations through common logic, the system reduces redundancy and manages complexity while maintaining high selection accuracy
Data Source
AI summary
This disclosure describes techniques for identifying a voice-enabled device from a group of voice-enabled devices to respond to a speech utterance of a user. A speech-processing system may receive an audio signal representing the speech utterance captured in an environment of a voice-enabled device, and identify another voice-enabled device located in the environment. The system may analyze the audio signal using a different natural-language-understanding model for each of the voice-enabled devices to identify an intent for each of the voice-enabled devices to respond to the speech utterance. The system may determine confidence scores that the intents are responsive to the speech utterance, and select the intent with the highest confidence score. The system may use the selected intent to generate a command for the corresponding voice-enabled device to respond to the user.


