Media Playback Arbitration for Multiple Voice Assistant Services

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Managing associations between multiple voice assistant services (VASes) in media playback systems is challenging due to interruptions and asynchronous responses, which can compromise user experience and privacy.

Innovation Solution

Playback devices arbitrate content playback and activation-word detection from multiple VASes, dynamically suppressing or delaying content from one VAS to avoid interrupting another based on content characteristics and user engagement, ensuring seamless and privacy-preserving interactions.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Adaptability or versatility

If multiple voice assistant services are integrated into media playback systems, then service versatility is improved, but system complexity increases due to content arbitration and interruption management

Engineering Contradiction:
Improvevoice assistant service integrationVSAvoidcontent arbitration management
Core Design Contradiction:
Adaptability or versatilityVSDevice complexity

Solution Approach 1:

The media playback device acts as an intermediary between multiple voice assistant services, receiving content from various VASes and arbitrating their playback. This mediator approach allows the system to integrate multiple services while centralizing the complexity management in a dedicated arbitration mechanism that coordinates content delivery and prevents conflicts.

Inventive Principle:
Principle #24Intermediary (Mediator)

Solution Approach 2:

The system dynamically adjusts content playback based on real-time conditions, including suppressing or delaying content from one VAS when content from another VAS is detected. This dynamic arbitration allows the system to adapt to changing user interactions and prioritize relevant content while maintaining overall system manageability.

Inventive Principle:
Principle #15Dynamics

2Ease of operation

If voice assistant content is played back continuously, then user engagement is improved, but interruptions from multiple services degrade user experience

Engineering Contradiction:
Improveuser interaction smoothnessVSAvoidcontent playback consistency
Core Design Contradiction:
Ease of operationVSReliability

Solution Approach 1:

The system monitors user interactions and content playback conditions in real-time, using feedback to determine when to suppress or delay content from one voice assistant service when content from another service is detected. This feedback mechanism ensures smooth user experience by preventing interruptions while maintaining reliable content delivery based on actual usage patterns.

Inventive Principle:
Principle #23Feedback

Solution Approach 2:

The content playback is dynamically adjusted based on detected conditions, including suppressing content from one VAS when content from another VAS is detected. This dynamic approach maintains user engagement by preventing interruptions while ensuring consistent and reliable content delivery according to real-time system state.

Inventive Principle:
Principle #15Dynamics

3Speed

If activation-word detection is performed for all voice assistant services simultaneously, then service responsiveness is improved, but processing time increases due to asynchronous responses

Engineering Contradiction:
Improvevoice assistant response speedVSAvoidactivation-word detection time
Core Design Contradiction:
SpeedVSLoss of time

Solution Approach 1:

The system performs activation-word detection for multiple voice assistant services simultaneously but selectively processes only the relevant responses based on detected content conditions. This partial action approach maintains service responsiveness by detecting all possible VAS activations while reducing actual processing time by filtering out non-relevant asynchronous responses.

Inventive Principle:
Principle #16Partial or excessive action

Solution Approach 2:

The system preliminarily detects activation words from multiple voice assistant services in parallel, then uses this preliminary information to determine which content to suppress or delay. This preliminary detection maintains fast response speeds while managing processing time efficiently by preparing multiple detection results simultaneously before executing the arbitration decision.

Inventive Principle:
Principle #10Preliminary action

Data Source

PatentUS20250239260A1Systems and methods of operating media playback systems having multiple voice assistant services
Publication Date: 2025.07.24 SONOS INC
  • US20250239260A1 patent drawing
  • US20250239260A1 patent drawing
  • US20250239260A1 patent drawing

AI summary

Systems and methods for managing multiple voice assistants are disclosed. Audio input is received via one or more microphones of a playback device. A first activation word is detected in the audio input via the playback device. After detecting the first activation word, the playback device transmits a voice utterance of the audio input to a first voice assistant service (VAS). The playback device receives, from the first VAS, first content to be played back via the playback device. The playback device also receives, from a second VAS, second content to be played back via the playback device. The playback device plays back the first content while suppressing the second content. Such suppression can include delaying or canceling playback of the second content.