Distributed Wake Word Detection Across Playback Devices

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing voice-controlled media playback systems face challenges in managing associations between playback devices and voice assistant services (VASes), often requiring users to select a single VAS due to processing power constraints and restrictions, limiting the ability to utilize multiple VASes for enhanced functionality.

Innovation Solution

The system distributes wake word detection and voice processing functions across multiple playback devices, allowing each device to detect different wake words and communicate with various VASes, thereby leveraging existing voice processing capabilities and freeing up computational resources, enabling users to interact with multiple VASes simultaneously.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Adaptability or versatility

If a single playback device is associated with a single voice assistant service, then processing power requirements are reduced, but the system loses the ability to utilize multiple VASes for enhanced functionality

Engineering Contradiction:
Improveability to utilize multiple VASesVSAvoidprocessing power consumption
Core Design Contradiction:
Adaptability or versatilityVSUse of energy by moving object

Solution Approach 1:

The patent segments the voice processing workload by distributing wake word detection across multiple playback devices, each associated with different VASes. Instead of one device handling all VAS communications, the system divides the detection task among several devices, allowing users to interact with multiple VASes while each device maintains manageable processing requirements.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent enables playback devices to serve multiple functions: they can detect wake words for different VASes, communicate with multiple VASes, and be dynamically selected based on which VAS is being invoked. This multi-functionality allows the system to leverage multiple VASes without requiring dedicated hardware for each service.

Inventive Principle:
Principle #6Universality (Multi-functionality)

2Adaptability or versatility

If multiple playback devices are used for wake word detection, then the system can support multiple VASes, but device complexity increases

Engineering Contradiction:
Improvesupport for multiple VASesVSAvoidsystem configuration complexity
Core Design Contradiction:
Adaptability or versatilityVSDevice complexity

Solution Approach 1:

The patent implements self-service mechanisms where playback devices automatically register their capabilities with the system, and the system automatically determines which device should handle wake word detection for each VAS. This automated capability registration and dynamic selection process reduces the complexity of manually configuring and managing multiple devices associated with multiple VASes.

Inventive Principle:
Principle #25Self-service

3Power

If wake word detection is distributed across multiple devices, then computational load on individual devices is reduced, but system coordination complexity increases

Engineering Contradiction:
Improvecomputational capacityVSAvoidsystem coordination
Core Design Contradiction:
PowerVSDevice complexity

Solution Approach 1:

The patent employs feedback mechanisms where the system continuously monitors which VAS is being invoked and dynamically directs wake word detection to the appropriate playback device. This real-time feedback and dynamic routing system coordinates multiple devices without requiring complex manual configuration, as the system automatically adjusts based on current usage patterns.

Inventive Principle:
Principle #23Feedback

Data Source

PatentUS11315556B2Devices, systems, and methods for distributed voice processing by transmitting sound data associated with a wake word to an appropriate device for identification
Publication Date: 2022.04.26 SONOS INC
  • US11315556B2 patent drawing
  • US11315556B2 patent drawing
  • US11315556B2 patent drawing

AI summary

Systems and methods for distributed voice processing are disclosed herein. In one example, the method includes detecting sound via a microphone array of a first playback device and analyzing, via a first wake-word engine of the first playback device, the detected sound. The first playback device may transmit data associated with the detected sound to a second playback device over a local area network. A second wake-word engine of the second playback device may analyze the transmitted data associated with the detected sound. The method may further include identifying that the detected sound contains either a first wake word or a second wake word based on the analysis via the first and second wake-word engines, respectively. Based on the identification, sound data corresponding to the detected sound may be transmitted over a wide area network to a remote computing device associated with a particular voice assistant service.