Multi-stream target speech detection via weighted fusion
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current audio processing systems face challenges in detecting and enhancing target speech in noisy environments, particularly in far-field VoIP applications, due to limitations in defining the target speech and adapting to changing speaker positions, leading to degraded signal quality.
Innovation Solution
A system architecture that combines multichannel source enhancement and separation with parallel pre-trained detectors to generate multiple streams, which are then weighted and fused to produce an enhanced output signal, using techniques like adaptive spatial filtering, blind source separation, and neural networks to improve keyword spotting and speech recognition.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Device complexity
If single-channel speech enhancement is used, then device complexity is reduced, but detection precision of target speech in noisy environments deteriorates
Solution Approach 1:
The system segments the speech enhancement task into multiple parallel streams, each processed by a dedicated detector engine. Multiple detector engines analyze different aspects or channels of the audio signal independently, then their results are fused to achieve superior detection precision while maintaining manageable complexity through modular processing
Solution Approach 2:
The system transitions from single-channel to multi-channel processing by introducing additional processing dimensions. Multiple detector engines operate in parallel on different signal representations or frequency channels, effectively adding dimensional complexity that improves detection precision without exponentially increasing overall system complexity
2Speed
If fixed target definition is used, then processing speed is improved, but adaptability to changing speaker positions deteriorates
Solution Approach 1:
The system implements dynamic target definition through multiple detector engines that can adaptively adjust their detection parameters and target definitions in real-time. Each detector engine can dynamically redefine what constitutes the target speech based on current acoustic conditions, speaker positions, and environmental factors, maintaining both speed and adaptability
Solution Approach 2:
The system incorporates feedback mechanisms where detector engines continuously monitor detection results and environmental conditions, then adjust their target definitions and processing parameters accordingly. This feedback loop enables rapid adaptation to changing speaker positions while maintaining processing efficiency through optimized parameter adjustment
3Reliability
If noise suppression is aggressively applied, then signal-to-noise ratio is improved, but speech quality and naturalness deteriorate
Solution Approach 1:
The system applies noise suppression selectively and locally rather than uniformly across the entire signal. Different detector engines apply different levels of noise suppression to different frequency bands, time segments, or spatial channels based on local noise characteristics, preserving speech quality in clean regions while suppressing noise in contaminated regions
Solution Approach 2:
The system dynamically adjusts noise suppression parameters based on local signal characteristics. Instead of applying fixed aggressive suppression, the detector engines modify suppression strength, frequency selectivity, and temporal smoothing parameters according to the instantaneous signal-to-noise ratio and speech activity detection, maintaining speech naturalness while improving overall signal quality
Data Source
AI summary
Audio processing systems and methods include an audio sensor array configured to receive a multichannel audio input and generate a corresponding multichannel audio signal and target-speech detection logic and an automatic speech recognition engine or VoIP application. An audio processing device includes a target speech enhancement engine configured to analyze a multichannel audio input signal and generate a plurality of enhanced target streams, a multi-stream target-speech detection generator comprising a plurality of target-speech detector engines each configured to determine a probability of detecting a specific target-speech of interest in the stream, wherein the multi-stream target-speech detection generator is configured to determine a plurality of weights associated with the enhanced target streams, and a fusion subsystem configured to apply the plurality of weights to the enhanced target streams to generate an enhancement output signal.


