Media Playback Voice Routing for Complex Commands
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Conventional voice assistant services (VASes) face challenges in handling nuanced voice commands due to computational intensity and resource limitations, particularly when controlling complex smart devices like media playback systems, often requiring users to use specific phraseology and failing to support multi-zone playback or advanced features.
Innovation Solution
A media playback system that selectively uses an enhanced VAS to process voice inputs, bypassing traditional VASes for advanced features by detecting specific command keywords and invoking the enhanced VAS to control media playback systems, including multi-zone playback and device grouping, using a network microphone device to capture and process voice inputs.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Adaptability or versatility
If traditional VASes are used to process voice inputs, then resource consumption is reduced, but voice control capabilities for complex smart devices are insufficient
Solution Approach 1:
The system segments voice processing into two paths: a lightweight traditional VAS for simple commands and an enhanced VAS for complex commands. The playback device determines command complexity and routes accordingly, allowing resource-efficient handling of simple tasks while enabling advanced features when needed.
Solution Approach 2:
The playback device acts as an intermediary between the traditional VAS and enhanced VAS. It receives voice inputs, determines whether they require advanced features, and selectively forwards complex commands to the enhanced VAS while handling simple commands locally, optimizing resource usage.
2Adaptability or versatility
If enhanced VAS is used to process complex voice commands, then voice control capabilities are improved, but computational intensity increases
Solution Approach 1:
The system divides voice processing workload by segmenting commands into simple and complex categories. Traditional VAS handles simple commands with low computational requirements, while enhanced VAS processes complex commands that require advanced features, optimizing the distribution of computational intensity.
Solution Approach 2:
The system applies partial action by using the enhanced VAS only when necessary for complex commands rather than for all voice inputs. This selective approach avoids excessive computational intensity while maintaining full capability when needed.
3Adaptability or versatility
If traditional VASes are used, then resource consumption is lower, but support for multi-zone playback and advanced features is lacking
Solution Approach 1:
The playback device is designed with multi-functionality, serving both as a traditional VAS endpoint for simple commands and as a gateway to the enhanced VAS for advanced features. This universal design allows the system to support multiple operational modes without requiring separate dedicated devices.
Solution Approach 2:
The playback device functions as an intermediary that bridges the gap between resource-efficient traditional VAS and capability-rich enhanced VAS. It intelligently routes commands based on feature requirements, enabling advanced features like multi-zone playback only when necessary.
Data Source
AI summary
Example techniques involve invoking voice assistance for a media playback system. In some embodiments, a NMD stores in memory a set of command information comprising a listing of playback commands and associated command criteria. The NMD captures a voice input and detects inclusion, within the voice input, of one or more particular playback commands from among the playback commands in the listing. In response, the NMD selects a local voice assistant that supports (a) one or more additional playback commands relative to a cloud-based VAS and (b) fewer non-playback commands relative to the cloud-based VAS, determines, via the local voice assistant, an intent in the captured voice input, and performs a response to the determined intent. The NMD foregoes selection of the cloud-based VAS when the local voice assistant is selected.


