Network Microphone Coordination for Persistent Voice Assistant Handover

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing voice-controlled media playback systems struggle with seamless voice interaction continuity as users move between multiple network microphone devices, leading to abrupt interruptions or dropped conversations due to varying sound quality and volume detection.

Innovation Solution

A system is implemented to coordinate sound detection, data transmission, and response output between multiple network microphone devices, selecting the nearest device to output responses based on user location, ensuring continuous voice control interactions.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Adaptability or versatility

If multiple network microphone devices are deployed to improve voice interaction coverage, then voice detection capability is improved, but voice interaction continuity deteriorates due to abrupt interruptions when users move between devices

Engineering Contradiction:
Improvevoice detection coverageVSAvoidvoice interaction continuity
Core Design Contradiction:
Adaptability or versatilityVSReliability

Solution Approach 1:

The patent merges the functionality of multiple network microphone devices into a unified voice interaction system. When a user moves between devices, the system combines detection capabilities across devices and maintains a single persistent voice assistant instance that follows the user, preventing interruptions and ensuring continuous interaction despite physical movement through the environment.

Inventive Principle:
Principle #5Merging (Combining)

2Device complexity

If each network microphone device operates independently to simplify device management, then device complexity is reduced, but voice interaction quality deteriorates due to varying sound detection quality across devices

Engineering Contradiction:
Improvedevice coordination complexityVSAvoidsound detection quality
Core Design Contradiction:
Device complexityVSMeasurement precision

Solution Approach 1:

The patent introduces a coordinator component that acts as an intermediary between multiple network microphone devices and the voice assistant. This coordinator manages device selection, monitors sound quality metrics, and ensures seamless transitions between devices based on user location and detection quality, thereby maintaining high measurement precision without requiring complex direct peer-to-peer coordination between devices.

Inventive Principle:
Principle #24Intermediary (Mediator)

3Measurement precision

If the system selects the nearest device to output responses to improve response quality, then sound detection accuracy is improved, but system complexity increases due to device selection and coordination requirements

Engineering Contradiction:
Improvesound detection accuracyVSAvoiddevice selection mechanism
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The patent implements a self-service mechanism where network microphone devices autonomously monitor their own detection quality metrics and user proximity. Each device can independently determine when it is the optimal device for current interactions based on pre-established criteria, reducing the need for complex centralized selection logic while maintaining high detection accuracy through distributed intelligence.

Inventive Principle:
Principle #25Self-service

Data Source

PatentUS12518756B2Voice assistant persistence across multiple network microphone devices
Publication Date: 2026.01.06 SONOS INC
  • US12518756B2 patent drawing
  • US12518756B2 patent drawing
  • US12518756B2 patent drawing

AI summary

Systems and methods for maintaining voice assistant persistence across multiple network microphone devices are described. In one example, first and second NMDs each identify a wake word based on detected sound, and are each transitioned from an inactive state to an active state in which the NMD captures and transmits sound data over a network interface. The first NMD is selected over the second NMD to output a first response, and both NMDs remain in the active state to further capture and transmit sound data. After further capturing and transmitting of sound data, the second NMD is selected over the first NMD to output a second response. After a predetermined time, one or both of the NMDs are transitioned back to the inactive state. The selection of one NMD over another for outputting a response can be based at least in part on user location information.