Multi-source TDOA Tracking and Voice Activity Detection for Planar Microphone Arrays

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing audio signal processing systems for smart speakers face challenges in efficiently isolating target audio from noise and other active speakers in multi-stream environments, as current methods like blind source separation and beamforming are either delayed or overly dependent on noise estimation, making them unsuitable for real-time applications with generic microphone arrays.

Innovation Solution

A combined multi-source Time Difference of Arrival (TDOA) tracking and Voice Activity Detection (VAD) mechanism is introduced, applicable to generic planar microphone arrays, which reduces computational complexity by performing TDOA searches in separate dimensions and avoids ghost TDOAs, allowing for effective noise suppression and target audio enhancement.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If blind source separation methods are used for speech enhancement, then spatial information of sources can be exploited, but large response delays occur due to batch processing

Engineering Contradiction:
Improvespatial information exploitationVSAvoidresponse delay
Core Design Contradiction:
Measurement precisionVSLoss of time

Solution Approach 1:

The patent transitions from static batch processing to dynamic real-time processing by implementing iterative update mechanisms that continuously refine source separation results as new audio data arrives, enabling the system to adapt to changing acoustic environments without requiring complete reprocessing of historical data

Inventive Principle:
Principle #15Dynamics

Solution Approach 2:

The system performs preliminary actions by precomputing and storing covariance matrices and other statistical parameters during noise-only segments, which are then rapidly applied during speech segments without requiring real-time computation, thus reducing response delays while maintaining separation quality

Inventive Principle:
Principle #10Preliminary action

2Measurement precision

If spatial filtering or beamforming methods are used, then target audio can be enhanced, but the system becomes overly dependent on noise estimation requiring supervision under voice activity detection

Engineering Contradiction:
Improvetarget audio enhancementVSAvoidsupervision dependency
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The patent implements self-service mechanisms where the system automatically identifies noise-only and speech segments through unsupervised statistical analysis of the audio signal itself, eliminating the need for external supervision or manual voice activity detection labels while still enabling effective noise estimation and target enhancement

Inventive Principle:
Principle #25Self-service

Solution Approach 2:

The system employs feedback loops where estimation results from previous time frames are continuously refined using current observations, allowing the noise and speech components to be dynamically re-estimated and separated through iterative optimization without requiring external supervision signals

Inventive Principle:
Principle #23Feedback

3Productivity

If conventional TDOA methods are used for source localization, then multiple sources can be tracked, but ghost TDOAs require post-processing removal

Engineering Contradiction:
Improvemulti-source trackingVSAvoidpost-processing requirement
Core Design Contradiction:
ProductivityVSEase of manufacture

Solution Approach 1:

The patent applies preliminary anti-action by incorporating constraints and validation checks during the TDOA estimation process itself that prevent ghost TDOA formation in the first place, rather than requiring post-processing to remove them. This is achieved through geometric consistency checks and physical plausibility constraints applied during source localization

Inventive Principle:
Principle #9Preliminary anti-action

4Adaptability or versatility

If generic planar microphone arrays are used, then system versatility is improved, but computational complexity increases

Engineering Contradiction:
Improvemicrophone array compatibilityVSAvoidcomputational complexity
Core Design Contradiction:
Adaptability or versatilityVSDevice complexity

Solution Approach 1:

The patent applies segmentation by dividing the computational task into independent frequency bins and time frames, allowing parallel processing of each segment. This modular approach enables the system to handle generic planar microphone arrays of various configurations without requiring monolithic complex computations, as each frequency bin can be processed independently with optimized algorithms

Inventive Principle:
Principle #1Segmentation

Data Source

PatentUS11937054B2Multiple-source tracking and voice activity detections for planar microphone arrays
Publication Date: 2024.03.19 SYNAPTICS INC
  • US11937054B2 patent drawing
  • US11937054B2 patent drawing
  • US11937054B2 patent drawing

AI summary

Embodiments described herein provide a combined multi-source time difference of arrival (TDOA) tracking and voice activity detection (VAD) mechanism that is applicable for generic array geometries, e.g., a microphone array that lies on a plane. The combined multi-source TDOA tracking and VAD mechanism scans the azimuth and elevation angles of the microphone array in microphone pairs, based on which a planar locus of physically admissible TDOAs can be formed in the multi-dimensional TDOA space of multiple microphone pairs. In this way, the multi-dimensional TDOA tracking reduces the number of calculations that was usually involved in traditional TDOA by performing the TDOA search for each dimension separately.