Broadened Beamwidth Beamforming with Spatial Activity Detection
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Conventional beamformer-postfilter systems for speech recognition in distant talk scenarios suffer from limited spatial coverage and degradation due to reverberation and interference, assuming a known speaker position, which is often not the case, and rely on suboptimal adaptation or camera-based localization.
Innovation Solution
The method involves forming multiple beams steered apart, mixing their outputs with directional and non-directional power spectral density signals using spatial activity detection, and applying dynamic postfiltering to enhance speech recognition accuracy across a broader spatial zone, including between beams, thereby suppressing reverberation and interference.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If a narrow beam (twenty-five degree width) is used to enhance ASR performance in a specific direction, then speech recognition accuracy is improved within the beam, but speech from speakers outside the beam is suppressed
Solution Approach 1:
The patent divides the spatial coverage into multiple overlapping beams (first beam and second beam) that cover different angular sectors. Each beam is processed independently with its own beamforming and postfiltering, allowing the system to maintain high ASR accuracy in each directional sector while collectively providing broader spatial coverage through the overlap of adjacent beams.
Solution Approach 2:
The patent transitions from a single narrow beam approach to a multi-dimensional spatial coverage by creating multiple beams at different orientations. The system processes signals from multiple directional sectors simultaneously and combines them, effectively expanding the coverage from one-dimensional (single direction) to two-dimensional (multiple angular sectors) spatial coverage.
2Adaptability or versatility
If acoustic speaker localization is used to steer the beam to the actual speaker position, then the beam can adapt to speaker movement, but the system fails to work robustly in scenarios with reverberation and interference
Solution Approach 1:
Instead of relying on a single localization estimate that may be corrupted by reverberation and interference, the patent segments the spatial coverage into multiple overlapping beams. Each beam processes signals independently, so if one beam's localization is affected by interference, other beams may still capture the speaker's signal reliably, providing robustness through spatial diversity.
Solution Approach 2:
The patent changes the parameter of beam orientation by creating multiple beams at different angular positions rather than relying on dynamic steering of a single beam. This static multi-beam approach is more robust to reverberation and interference because it doesn't depend on accurate real-time speaker localization, which degrades in such environments.
3Adaptability or versatility
If adaptive beamforming is used to adjust to the true speaker position, then some tracking capability is achieved, but the performance is suboptimal compared to precise localization
Solution Approach 1:
The patent applies segmentation by creating multiple fixed beams that collectively cover the entire spatial sector. Each beam maintains its own beamforming and postfiltering processing, ensuring that speech from any direction falls within at least one beam's coverage area with optimal processing, thereby maintaining high ASR accuracy without requiring adaptive tracking.
Solution Approach 2:
Instead of using a single adaptive beam that attempts to track the speaker, the patent employs multiple beams with overlapping coverage areas. This excessive action of creating redundant beam coverage ensures that speaker speech is always captured by at least one beam at optimal conditions, compensating for the lack of adaptive tracking capability.
Data Source
AI summary
Methods and apparatus for broadening the beamwidth of beamforming and postfiltering using a plurality of beamformers and signal and power spectral density mixing, and controlling a postfilter based on spatial activity detection such that de-reverberation or noise reduction is performed when a speech source is between the first and second beams.


