Deep Learning Multi-Channel Speech Filtering for Low Latency Enhancement
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing digital signal processing techniques for multi-channel speech signals in consumer and automotive electronics fail to accurately enhance speech signals in noisy environments, leading to high word error rates and false triggering in voice trigger detection and automatic speech recognition systems.
Innovation Solution
A deep neural network (DNN) driven multi-channel filtering process that extracts features from current frames of multi-channel speech pickup signals, including side information like echo estimates and noise, to produce a speech presence probability value, which configures a multi-channel filter to suppress undesired components and enhance the target speech signal, reducing latency and improving accuracy.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If traditional digital signal processing techniques are used for multi-channel speech enhancement, then the system complexity is low, but the speech enhancement accuracy is insufficient leading to high word error rates and false triggering
Solution Approach 1:
The patent replaces traditional mechanical signal processing techniques with a deep neural network-based system. The DNN extracts features from multi-channel speech signals and produces speech presence probability values that drive adaptive filtering, substituting conventional algorithmic approaches with a learned model that achieves superior enhancement accuracy while managing computational complexity through efficient architecture design.
2Measurement precision
If deep neural network based multi-channel filtering is applied, then speech enhancement accuracy improves, but computational complexity and processing time increase
Solution Approach 1:
The patent segments the speech enhancement task into distinct processing stages: feature extraction from multi-channel inputs, DNN-based speech presence probability estimation, and adaptive filter configuration. This segmentation allows each component to be optimized independently, with the DNN processing features in parallel and the filter applying enhancements in real-time, thereby maintaining processing efficiency while achieving high accuracy.
Solution Approach 2:
The DNN performs preliminary analysis of the multi-channel speech signals by extracting features and producing speech presence probability values before the actual filtering operation. This preliminary action prepares the system by identifying speech regions and characteristics in advance, allowing the subsequent filtering stage to operate more efficiently with pre-computed guidance information.
3Measurement precision
If deep neural network processing is used, then speech presence detection accuracy improves, but processing latency increases
Solution Approach 1:
The patent implements periodic processing where the DNN analyzes speech frames at regular intervals and updates filter configurations accordingly. By processing speech in discrete frames with periodic DNN evaluation and filter updates, the system achieves accurate speech presence detection while maintaining low latency through efficient frame-based processing rather than continuous analysis.
Data Source
AI summary
A number of features are extracted from a current frame of a multi-channel speech pickup and from side information that is a linear echo estimate, a diffuse signal component, or a noise estimate of the multi-channel speech pickup. A DNN-based speech presence probability is produced for the current frame, where the SPP value is produced in response to the extracted features being input to the DNN. The DNN-based SPP value is applied to configure a multi-channel filter whose input is the multi-channel speech pickup and whose output is a single audio signal. In one aspect, the system is designed to run online, at low enough latency for real time applications such voice trigger detection. Other aspects are also described and claimed.


