Multi-Source Audio Location Detection Using Sub-Band Segmentation
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Conventional audio systems face challenges in accurately detecting and tracking dominant audio sources in noisy environments, especially when multiple sources are active simultaneously, due to limitations in spatial localization and computational complexity, leading to inefficiencies in speech enhancement and noise suppression.
Innovation Solution
The system employs a 360-degree multi-source location detection and tracking method using a planar microphone array, which includes an audio sensor array, a sub-band frequency analyzer, a target activity detector, and a source tracker to identify and enhance dominant audio sources by computing discrete spatial maps and applying repulsion algorithms to separate sources, thereby improving spatial resolution and adaptability.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If conventional source localization methods (SRP-PHAT, time-delay estimation) are used, then the system can identify audio sources, but the accuracy deteriorates when multiple sources are active simultaneously due to multidimensional ghost source issues
Solution Approach 1:
The patent segments the audio signal into multiple frequency sub-bands and processes each sub-band independently to identify instantaneous dominant locations. This segmentation approach allows the system to handle multiple simultaneous sources by analyzing different frequency regions separately, avoiding the ghost source problems that occur when processing the full spectrum as a single unit.
Solution Approach 2:
The patent introduces a temporal dimension by tracking dominant locations across multiple frames and applying smoothing constraints. Instead of relying solely on spatial information from a single frame, the system incorporates temporal continuity to disambiguate overlapping sources and maintain accurate tracking even when sources are close in spatial frequency bins.
2Reliability
If MIMO system identification or state-space methods are used for tracking, then the system can track source positions, but the computational complexity and memory requirements increase significantly
Solution Approach 1:
The patent implements a self-service tracking mechanism where each detected dominant location automatically generates a tracking hypothesis that is smoothed and updated using only local temporal information from previous frames. This eliminates the need for complex centralized state-space solvers or MIMO system identification algorithms, reducing computational complexity while maintaining reliable tracking through simple, localized temporal smoothing.
3Measurement precision
If a larger number of microphones are used in the array, then the spatial resolution and source separation improve, but the device complexity and cost increase
Solution Approach 1:
The patent changes the processing parameters by analyzing multiple frequency sub-bands independently and identifying instantaneous dominant locations in each sub-band. This parameter transformation allows the system to achieve better effective spatial resolution through frequency-domain separation rather than relying solely on increasing the physical number of microphones, thus improving source separation while controlling device complexity.
Data Source
AI summary
Audio processing systems and methods comprise an audio sensor array configured to receive a multichannel audio input and generate a corresponding multichannel audio signal and a target activity detector configured to identify audio target sources in the multichannel audio signal. The target activity detector includes a VAD, an instantaneous locations component configured to detect a location of a plurality of audio sources, a dominant locations component configured to selectively buffer a subset of the plurality of audio sources comprising dominant audio sources, a source tracker configured to track locations of the dominant audio sources over time, and a dominance selection component configured to select the dominant target sources for further audio processing. The instantaneous location component computes a discrete spatial map comprising the location of the plurality of audio sources, and the dominant location component selects N of the dominant sources from the discrete spatial map for source tracking.


