DNN Acoustic Model Front-End for Multi-Channel Audio Processing

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Conventional speech recognition systems are inefficient in processing multi-channel audio data, requiring separate components for beamforming and feature extraction, which are optimized individually for signal enhancement rather than jointly for speech recognition, and are bandwidth-intensive.

Innovation Solution

A deep neural network (DNN)-based acoustic model front-end is introduced, initialized with data corresponding to multiple microphone array geometries, which mimics beamforming and feature extraction, allowing for joint optimization of speech recognition components and reducing bandwidth requirements by processing multi-channel audio data directly.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Reliability

If separate components for beamforming and feature extraction are used, then signal enhancement is optimized individually, but speech recognition performance and bandwidth efficiency deteriorate

Engineering Contradiction:
Improvesignal enhancementVSAvoidspeech recognition performance
Core Design Contradiction:
ReliabilityVSProductivity

Solution Approach 1:

The patent combines separate beamforming and feature extraction components into a unified deep neural network-based acoustic model front-end. This integration allows joint optimization of both functions, where the DNN simultaneously performs spatial filtering (beamforming) and feature extraction from multi-channel audio data, eliminating the need for separate processing stages and enabling end-to-end optimization for speech recognition performance.

Inventive Principle:
Principle #5Merging (Combining)

Solution Approach 2:

The DNN-based acoustic model front-end serves multiple functions simultaneously: it performs beamforming (spatial filtering), feature extraction, and acoustic modeling. This multi-functional approach replaces the traditional pipeline of separate components, allowing the system to optimize all functions jointly within a single unified architecture that processes multi-channel audio data end-to-end.

Inventive Principle:
Principle #6Universality (Multi-functionality)

2Reliability

If multiple microphones are used for multi-channel processing, then speech recognition robustness improves, but bandwidth requirements increase

Engineering Contradiction:
Improvespeech recognition robustnessVSAvoidbandwidth usage
Core Design Contradiction:
ReliabilityVSQuantity of substance

Solution Approach 1:

The patent extracts essential speech features directly from multi-channel audio data using the DNN-based acoustic model front-end, processing the data locally before transmission. This extraction approach sends only processed acoustic features rather than raw multi-channel audio data, significantly reducing bandwidth requirements while maintaining the robustness benefits of multi-microphone processing for noise suppression and speech enhancement.

Inventive Principle:
Principle #2Taking out (Extraction)

3Reliability

If conventional beamforming and feature extraction components are used, then signal enhancement is achieved, but device complexity increases

Engineering Contradiction:
Improvesignal enhancementVSAvoidprocessing components
Core Design Contradiction:
ReliabilityVSDevice complexity

Solution Approach 1:

The patent merges multiple processing components (beamformer, feature extractor, acoustic model) into a single integrated DNN-based acoustic model front-end. This consolidation reduces device complexity by eliminating the need for separate processing stages and their associated control logic, while maintaining signal enhancement capabilities through the DNN's joint optimization of spatial filtering and feature extraction.

Inventive Principle:
Principle #5Merging (Combining)

Data Source

PatentUS11574628B1Deep multi-channel acoustic modeling using multiple microphone array geometries
Publication Date: 2023.02.07 AMAZON TECH INC
  • US11574628B1 patent drawing
  • US11574628B1 patent drawing
  • US11574628B1 patent drawing

AI summary

Techniques for speech processing using a deep neural network (DNN) based acoustic model front-end are described. A new modeling approach directly models multi-channel audio data received from a microphone array using a first model (e.g., multi-geometry/multi-channel DNN) that is trained using a plurality of microphone array geometries. Thus, the first model may receive a variable number of microphone channels, generate multiple outputs using multiple microphone array geometries, and select the best output as a first feature vector that may be used similarly to beamformed features generated by an acoustic beamformer. A second model (e.g., feature extraction DNN) processes the first feature vector and transforms it to a second feature vector having a lower dimensional representation. A third model (e.g., classification DNN) processes the second feature vector to perform acoustic unit classification and generate text data. The DNN front-end enables improved performance despite a reduction in microphones.