Frequency Aligned Network for Multi-Channel Speech Recognition

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Conventional speech recognition systems are inefficient in processing multi-channel audio data, requiring separate optimization of components for signal enhancement and often needing more microphones, which increases bandwidth requirements and noise sensitivity.

Innovation Solution

A deep neural network (DNN)-based acoustic model front-end with a frequency aligned network (FAN) architecture performs spatial filtering and feature extraction, jointly optimizing components for speech recognition and reducing the number of microphones required, while minimizing bandwidth by uploading low-dimensional feature vectors.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Adaptability or versatility

If conventional speech recognition systems use separate components for signal enhancement, then speech recognition can be performed, but the system requires more microphones and increases bandwidth requirements

Engineering Contradiction:
Improvespeech recognition capabilityVSAvoidnumber of microphones
Core Design Contradiction:
Adaptability or versatilityVSQuantity of substance

Solution Approach 1:

The patent combines spatial filtering and feature extraction into a single integrated DNN module. The multi-channel audio data is processed through the DNN which simultaneously performs spatial filtering to separate speech from noise and extracts acoustic features for recognition, eliminating the need for separate signal enhancement components and reducing microphone requirements

Inventive Principle:
Principle #5Merging (Combining)

Solution Approach 2:

The DNN-based acoustic model front-end serves multiple functions: it performs spatial filtering for noise reduction, feature extraction for speech recognition, and adaptive processing for different acoustic environments. This multi-functional approach replaces multiple specialized components with a single universal processor

Inventive Principle:
Principle #6Universality (Multi-functionality)

2Adaptability or versatility

If conventional speech recognition systems use separate components for signal enhancement, then speech recognition can be performed, but bandwidth requirements increase

Engineering Contradiction:
Improvespeech recognition capabilityVSAvoidbandwidth
Core Design Contradiction:
Adaptability or versatilityVSQuantity of substance

Solution Approach 1:

The system extracts only the essential acoustic features from multi-channel audio data using the DNN, rather than transmitting or processing the complete audio signals. The DNN extracts low-dimensional feature vectors that capture the critical speech information while filtering out redundant data, significantly reducing bandwidth requirements

Inventive Principle:
Principle #2Taking out (Extraction)

3Adaptability or versatility

If conventional speech recognition systems process multi-channel audio data, then speech recognition can be performed, but noise sensitivity increases

Engineering Contradiction:
Improvespeech recognition capabilityVSAvoidnoise sensitivity
Core Design Contradiction:
Adaptability or versatilityVSObject-affected harmful factors

Solution Approach 1:

The patent integrates spatial filtering and feature extraction in the DNN architecture, allowing the system to perform noise reduction and feature extraction simultaneously. This joint optimization enables the system to filter out noise while extracting speech features more effectively than sequential processing

Inventive Principle:
Principle #5Merging (Combining)

4Ease of manufacture

If conventional speech recognition systems use separate optimization for components, then signal enhancement can be performed, but system complexity increases

Engineering Contradiction:
Improvecomponent optimizationVSAvoidsystem architecture
Core Design Contradiction:
Ease of manufactureVSDevice complexity

Solution Approach 1:

The patent merges spatial filtering and feature extraction into a single DNN module with unified optimization. The system uses a joint objective function that optimizes both spatial filtering performance and feature extraction quality simultaneously, eliminating the need for separate component optimization and reducing system complexity

Inventive Principle:
Principle #5Merging (Combining)

Data Source

PatentUS11495215B1Deep multi-channel acoustic modeling using frequency aligned network
Publication Date: 2022.11.08 AMAZON TECH INC
  • US11495215B1 patent drawing
  • US11495215B1 patent drawing
  • US11495215B1 patent drawing

AI summary

Techniques for speech processing using a deep neural network (DNN) based acoustic model front-end are described. A new modeling approach directly models multi-channel audio data received from a microphone array using a first model (e.g., multi-geometry/multi-channel DNN) that includes a frequency aligned network (FAN) architecture. Thus, the first model may perform spatial filtering to generate a first feature vector by processing individual frequency bins separately, such that multiple frequency bins are not combined. The first feature vector may be used similarly to beamformed features generated by an acoustic beamformer. A second model (e.g., feature extraction DNN) processes the first feature vector and transforms it to a second feature vector having a lower dimensional representation. A third model (e.g., classification DNN) processes the second feature vector to perform acoustic unit classification and generate text data. The DNN front-end enables improved performance despite a reduction in microphones.