Network Microphone Device Voice Detection Optimization
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing voice-controllable media playback systems face challenges in accurately detecting voice inputs amidst background noise and environmental factors, leading to suboptimal performance in smart home settings.
Innovation Solution
The system employs network microphone devices (NMDs) equipped with wake-word engines and spatial processing algorithms, which filter background noise and adapt to environmental conditions to improve voice detection and processing.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If network microphone devices use wake-word engines and spatial processing algorithms to filter background noise, then voice detection accuracy is improved, but device complexity increases
Solution Approach 1:
The voice detection system is divided into separate functional modules: wake-word engines for keyword detection, spatial processing algorithms for noise filtering, and voice detection components. Each module handles a specific aspect of voice processing, allowing the system to achieve high detection accuracy while maintaining manageable complexity through modular architecture.
Solution Approach 2:
Spatial processing algorithms act as intermediaries between the microphones and the voice detection engine. These algorithms process the raw audio signals first, filtering out background noise and enhancing voice signals before passing them to the wake-word engines, thereby improving overall detection accuracy without requiring the final detection component to handle raw noisy signals directly.
2Reliability
If the system adapts to environmental conditions and adjusts processing, then reliability is improved, but computational requirements and energy consumption increase
Solution Approach 1:
The system dynamically adapts its processing based on environmental conditions. The spatial processing algorithms and wake-word engines adjust their parameters and processing intensity according to the detected noise levels and environmental factors, allowing the system to maintain high reliability in varying conditions while optimizing energy consumption by not always operating at maximum processing capacity.
Solution Approach 2:
The system changes operational parameters such as noise threshold levels, processing gain, and algorithm sensitivity based on environmental conditions. By dynamically adjusting these parameters, the system maintains reliable voice detection across different environments while avoiding excessive energy consumption that would result from always using the most intensive processing settings.
Data Source
Figure 1A
Figure 1B
Figure 2A~2B
AI summary
Systems and methods for optimizing voice detection via a network microphone device (NMD) based on a selected voice-assistant service (VAS) are disclosed herein. In one example, the NMD detects sound via individual microphones and selects a first VAS to communicate with the NMD. The NMD produces a first sound-data stream based on the detected sound using a spatial processor in a first configuration. Once the NMD determines that a second VAS is to be selected over the first VAS, the spatial processor assumes a second configuration for producing a second sound-data stream based on the detected sound. The second sound-data stream is then transmitted to one or more remote computing devices associated with the second VAS.