Network Microphone Device Voice Detection Optimization

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing voice-controllable media playback systems face challenges in accurately detecting voice inputs amidst background noise and environmental factors, leading to suboptimal performance in smart home settings.

Innovation Solution

The system employs network microphone devices (NMDs) equipped with wake-word engines and spatial processing algorithms, which filter background noise and adapt to environmental conditions to improve voice detection and processing.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If network microphone devices use wake-word engines and spatial processing algorithms to filter background noise, then voice detection accuracy is improved, but device complexity increases

Engineering Contradiction:
Improvevoice detection accuracyVSAvoiddevice complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The voice detection system is divided into separate functional modules: wake-word engines for keyword detection, spatial processing algorithms for noise filtering, and voice detection components. Each module handles a specific aspect of voice processing, allowing the system to achieve high detection accuracy while maintaining manageable complexity through modular architecture.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

Spatial processing algorithms act as intermediaries between the microphones and the voice detection engine. These algorithms process the raw audio signals first, filtering out background noise and enhancing voice signals before passing them to the wake-word engines, thereby improving overall detection accuracy without requiring the final detection component to handle raw noisy signals directly.

Inventive Principle:
Principle #24Intermediary (Mediator)

2Reliability

If the system adapts to environmental conditions and adjusts processing, then reliability is improved, but computational requirements and energy consumption increase

Engineering Contradiction:
Improveperformance reliabilityVSAvoidenergy consumption
Core Design Contradiction:
ReliabilityVSUse of energy by moving object

Solution Approach 1:

The system dynamically adapts its processing based on environmental conditions. The spatial processing algorithms and wake-word engines adjust their parameters and processing intensity according to the detected noise levels and environmental factors, allowing the system to maintain high reliability in varying conditions while optimizing energy consumption by not always operating at maximum processing capacity.

Inventive Principle:
Principle #15Dynamics

Solution Approach 2:

The system changes operational parameters such as noise threshold levels, processing gain, and algorithm sensitivity based on environmental conditions. By dynamically adjusting these parameters, the system maintains reliable voice detection across different environments while avoiding excessive energy consumption that would result from always using the most intensive processing settings.

Inventive Principle:
Principle #35Parameter changes

Data Source

PatentEP4297438B1Voice detection optimization based on selected voice assistant service
Publication Date: 2025.05.21 SONOS INC
  • EP4297438B1 patent drawingFigure 1A
  • EP4297438B1 patent drawingFigure 1B
  • EP4297438B1 patent drawingFigure 2A~2B

AI summary

Systems and methods for optimizing voice detection via a network microphone device (NMD) based on a selected voice-assistant service (VAS) are disclosed herein. In one example, the NMD detects sound via individual microphones and selects a first VAS to communicate with the NMD. The NMD produces a first sound-data stream based on the detected sound using a spatial processor in a first configuration. Once the NMD determines that a second VAS is to be selected over the first VAS, the spatial processor assumes a second configuration for producing a second sound-data stream based on the detected sound. The second sound-data stream is then transmitted to one or more remote computing devices associated with the second VAS.