Networked Microphone Voice Service Selection for Media Playback

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing digital audio playback systems lack advanced voice control capabilities for networked devices, limiting user interaction and customization with multiple voice services.

Innovation Solution

Networked microphone devices (NMDs) interface with multiple voice services, allowing voice inputs to be processed and commands to be executed across a media playback system, with the ability to identify and select a specific voice service based on context or wake-word, and evaluate results to provide the best response.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Adaptability or versatility

If a single voice service is integrated into the media playback system, then the system structure remains simple and easy to manage, but the adaptability and versatility of voice control capabilities are limited

Engineering Contradiction:
Improvevoice control capabilitiesVSAvoidsystem structure
Core Design Contradiction:
Adaptability or versatilityVSDevice complexity

Solution Approach 1:

The patent implements multi-functionality by enabling the media playback system to interface with multiple voice services (e.g., Siri, Google Assistant, Alexa) through a unified architecture. The system can select and switch between different voice services based on user preferences, device capabilities, and service availability, allowing a single system to perform multiple voice control functions without requiring separate dedicated systems for each service.

Inventive Principle:
Principle #6Universality (Multi-functionality)

2Adaptability or versatility

If multiple voice services are integrated into the media playback system, then adaptability and versatility of voice control are enhanced, but the device complexity and difficulty of management increase

Engineering Contradiction:
Improvevoice control capabilitiesVSAvoidsystem management
Core Design Contradiction:
Adaptability or versatilityVSEase of operation

Solution Approach 1:

The patent implements self-service by enabling the media playback system to automatically manage multiple voice services through automated service discovery, selection, and switching mechanisms. The system can autonomously determine which voice service to use based on pre-configured user preferences, device compatibility, and service availability, eliminating the need for manual intervention when switching between services. Users can set their preferences once, and the system handles the complexity of managing multiple services automatically.

Inventive Principle:
Principle #25Self-service

3Measurement precision

If voice inputs are transmitted to multiple voice services for processing, then the system can evaluate and select the best response, but the loss of time due to multiple processing requests increases

Engineering Contradiction:
Improveresponse qualityVSAvoidprocessing time
Core Design Contradiction:
Measurement precisionVSLoss of time

Solution Approach 1:

The patent implements preliminary action by pre-configuring user preferences and service priorities before actual voice processing occurs. The system stores user-defined preferences regarding which voice services to use for different types of commands, the order of service preferences, and selection criteria. When a voice input is received, the system can quickly determine which service to use based on these pre-established rules, avoiding the need to evaluate all services in real-time and thus reducing processing time while still maintaining response quality.

Inventive Principle:
Principle #10Preliminary action

Data Source

PatentEP3618064B1Multiple voice services
Publication Date: 2025.10.01 SONOS INC
  • EP3618064B1 patent drawingFigure 1
  • EP3618064B1 patent drawingFigure 2~3
  • EP3618064B1 patent drawingFigure 4

AI summary

Disclosed herein are example techniques to identify a voice service to process a voice input. An example implementation may involve an NMD receiving, via a microphone, voice data indicating a voice input. The NMD may identify, from among multiple voice services registered to a media playback system, a voice service to process the voice input and cause, via a network interface, the identified voice service to process the voice input.