Multi-stream target speech detection via weighted fusion

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current audio processing systems face challenges in detecting and enhancing target speech in noisy environments, particularly in far-field VoIP applications, due to limitations in defining the target speech and adapting to changing speaker positions, leading to degraded signal quality.

Innovation Solution

A system architecture that combines multichannel source enhancement and separation with parallel pre-trained detectors to generate multiple streams, which are then weighted and fused to produce an enhanced output signal, using techniques like adaptive spatial filtering, blind source separation, and neural networks to improve keyword spotting and speech recognition.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Device complexity

If single-channel speech enhancement is used, then device complexity is reduced, but detection precision of target speech in noisy environments deteriorates

Engineering Contradiction:
Improveprocessing complexityVSAvoiddetection precision
Core Design Contradiction:
Device complexityVSMeasurement precision

Solution Approach 1:

The system segments the speech enhancement task into multiple parallel streams, each processed by a dedicated detector engine. Multiple detector engines analyze different aspects or channels of the audio signal independently, then their results are fused to achieve superior detection precision while maintaining manageable complexity through modular processing

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The system transitions from single-channel to multi-channel processing by introducing additional processing dimensions. Multiple detector engines operate in parallel on different signal representations or frequency channels, effectively adding dimensional complexity that improves detection precision without exponentially increasing overall system complexity

Inventive Principle:
Principle #17Another dimension (Dimensionality change)

2Speed

If fixed target definition is used, then processing speed is improved, but adaptability to changing speaker positions deteriorates

Engineering Contradiction:
Improveprocessing speedVSAvoidadaptability
Core Design Contradiction:
SpeedVSAdaptability or versatility

Solution Approach 1:

The system implements dynamic target definition through multiple detector engines that can adaptively adjust their detection parameters and target definitions in real-time. Each detector engine can dynamically redefine what constitutes the target speech based on current acoustic conditions, speaker positions, and environmental factors, maintaining both speed and adaptability

Inventive Principle:
Principle #15Dynamics

Solution Approach 2:

The system incorporates feedback mechanisms where detector engines continuously monitor detection results and environmental conditions, then adjust their target definitions and processing parameters accordingly. This feedback loop enables rapid adaptation to changing speaker positions while maintaining processing efficiency through optimized parameter adjustment

Inventive Principle:
Principle #23Feedback

3Reliability

If noise suppression is aggressively applied, then signal-to-noise ratio is improved, but speech quality and naturalness deteriorate

Engineering Contradiction:
Improvesignal-to-noise ratioVSAvoidspeech quality
Core Design Contradiction:
ReliabilityVSLoss of information

Solution Approach 1:

The system applies noise suppression selectively and locally rather than uniformly across the entire signal. Different detector engines apply different levels of noise suppression to different frequency bands, time segments, or spatial channels based on local noise characteristics, preserving speech quality in clean regions while suppressing noise in contaminated regions

Inventive Principle:
Principle #3Local quality

Solution Approach 2:

The system dynamically adjusts noise suppression parameters based on local signal characteristics. Instead of applying fixed aggressive suppression, the detector engines modify suppression strength, frequency selectivity, and temporal smoothing parameters according to the instantaneous signal-to-noise ratio and speech activity detection, maintaining speech naturalness while improving overall signal quality

Inventive Principle:
Principle #35Parameter changes

Data Source

PatentUS11694710B2Multi-stream target-speech detection and channel fusion
Publication Date: 2023.07.04 SYNAPTICS INC
  • US11694710B2 patent drawing
  • US11694710B2 patent drawing
  • US11694710B2 patent drawing

AI summary

Audio processing systems and methods include an audio sensor array configured to receive a multichannel audio input and generate a corresponding multichannel audio signal and target-speech detection logic and an automatic speech recognition engine or VoIP application. An audio processing device includes a target speech enhancement engine configured to analyze a multichannel audio input signal and generate a plurality of enhanced target streams, a multi-stream target-speech detection generator comprising a plurality of target-speech detector engines each configured to determine a probability of detecting a specific target-speech of interest in the stream, wherein the multi-stream target-speech detection generator is configured to determine a plurality of weights associated with the enhanced target streams, and a fusion subsystem configured to apply the plurality of weights to the enhanced target streams to generate an enhancement output signal.