Deep Learning Multi-Channel Speech Filtering for Low Latency Enhancement

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing digital signal processing techniques for multi-channel speech signals in consumer and automotive electronics fail to accurately enhance speech signals in noisy environments, leading to high word error rates and false triggering in voice trigger detection and automatic speech recognition systems.

Innovation Solution

A deep neural network (DNN) driven multi-channel filtering process that extracts features from current frames of multi-channel speech pickup signals, including side information like echo estimates and noise, to produce a speech presence probability value, which configures a multi-channel filter to suppress undesired components and enhance the target speech signal, reducing latency and improving accuracy.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If traditional digital signal processing techniques are used for multi-channel speech enhancement, then the system complexity is low, but the speech enhancement accuracy is insufficient leading to high word error rates and false triggering

Engineering Contradiction:
Improvespeech enhancement accuracyVSAvoidsystem complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The patent replaces traditional mechanical signal processing techniques with a deep neural network-based system. The DNN extracts features from multi-channel speech signals and produces speech presence probability values that drive adaptive filtering, substituting conventional algorithmic approaches with a learned model that achieves superior enhancement accuracy while managing computational complexity through efficient architecture design.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

2Measurement precision

If deep neural network based multi-channel filtering is applied, then speech enhancement accuracy improves, but computational complexity and processing time increase

Engineering Contradiction:
Improvespeech enhancement accuracyVSAvoidreal-time processing efficiency
Core Design Contradiction:
Measurement precisionVSProductivity

Solution Approach 1:

The patent segments the speech enhancement task into distinct processing stages: feature extraction from multi-channel inputs, DNN-based speech presence probability estimation, and adaptive filter configuration. This segmentation allows each component to be optimized independently, with the DNN processing features in parallel and the filter applying enhancements in real-time, thereby maintaining processing efficiency while achieving high accuracy.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The DNN performs preliminary analysis of the multi-channel speech signals by extracting features and producing speech presence probability values before the actual filtering operation. This preliminary action prepares the system by identifying speech regions and characteristics in advance, allowing the subsequent filtering stage to operate more efficiently with pre-computed guidance information.

Inventive Principle:
Principle #10Preliminary action

3Measurement precision

If deep neural network processing is used, then speech presence detection accuracy improves, but processing latency increases

Engineering Contradiction:
Improvespeech presence detection accuracyVSAvoidprocessing latency
Core Design Contradiction:
Measurement precisionVSLoss of time

Solution Approach 1:

The patent implements periodic processing where the DNN analyzes speech frames at regular intervals and updates filter configurations accordingly. By processing speech in discrete frames with periodic DNN evaluation and filter updates, the system achieves accurate speech presence detection while maintaining low latency through efficient frame-based processing rather than continuous analysis.

Inventive Principle:
Principle #19Periodic action

Data Source

PatentUS10546593B2Deep learning driven multi-channel filtering for speech enhancement
Publication Date: 2020.01.28 APPLE INC
  • US10546593B2 patent drawing
  • US10546593B2 patent drawing
  • US10546593B2 patent drawing

AI summary

A number of features are extracted from a current frame of a multi-channel speech pickup and from side information that is a linear echo estimate, a diffuse signal component, or a noise estimate of the multi-channel speech pickup. A DNN-based speech presence probability is produced for the current frame, where the SPP value is produced in response to the extracted features being input to the DNN. The DNN-based SPP value is applied to configure a multi-channel filter whose input is the multi-channel speech pickup and whose output is a single audio signal. In one aspect, the system is designed to run online, at low enough latency for real time applications such voice trigger detection. Other aspects are also described and claimed.