Distributed Wake Word Detection Across Playback Devices
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing voice-controlled media playback systems face challenges in managing associations between playback devices and voice assistant services (VASes), often requiring users to select a single VAS due to processing power constraints and restrictions, limiting the ability to utilize multiple VASes for enhanced functionality.
Innovation Solution
The system distributes wake word detection and voice processing functions across multiple playback devices, allowing each device to detect different wake words and communicate with various VASes, thereby leveraging existing voice processing capabilities and freeing up computational resources, enabling users to interact with multiple VASes simultaneously.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Adaptability or versatility
If a single playback device is associated with a single voice assistant service, then processing power requirements are reduced, but the system loses the ability to utilize multiple VASes for enhanced functionality
Solution Approach 1:
The patent segments the voice processing workload by distributing wake word detection across multiple playback devices, each associated with different VASes. Instead of one device handling all VAS communications, the system divides the detection task among several devices, allowing users to interact with multiple VASes while each device maintains manageable processing requirements.
Solution Approach 2:
The patent enables playback devices to serve multiple functions: they can detect wake words for different VASes, communicate with multiple VASes, and be dynamically selected based on which VAS is being invoked. This multi-functionality allows the system to leverage multiple VASes without requiring dedicated hardware for each service.
2Adaptability or versatility
If multiple playback devices are used for wake word detection, then the system can support multiple VASes, but device complexity increases
Solution Approach 1:
The patent implements self-service mechanisms where playback devices automatically register their capabilities with the system, and the system automatically determines which device should handle wake word detection for each VAS. This automated capability registration and dynamic selection process reduces the complexity of manually configuring and managing multiple devices associated with multiple VASes.
3Power
If wake word detection is distributed across multiple devices, then computational load on individual devices is reduced, but system coordination complexity increases
Solution Approach 1:
The patent employs feedback mechanisms where the system continuously monitors which VAS is being invoked and dynamically directs wake word detection to the appropriate playback device. This real-time feedback and dynamic routing system coordinates multiple devices without requiring complex manual configuration, as the system automatically adjusts based on current usage patterns.
Data Source
AI summary
Systems and methods for distributed voice processing are disclosed herein. In one example, the method includes detecting sound via a microphone array of a first playback device and analyzing, via a first wake-word engine of the first playback device, the detected sound. The first playback device may transmit data associated with the detected sound to a second playback device over a local area network. A second wake-word engine of the second playback device may analyze the transmitted data associated with the detected sound. The method may further include identifying that the detected sound contains either a first wake word or a second wake word based on the analysis via the first and second wake-word engines, respectively. Based on the identification, sound data corresponding to the detected sound may be transmitted over a wide area network to a remote computing device associated with a particular voice assistant service.


