CLDNN Joint Beamforming Acoustic Modeling

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing speech recognition systems face challenges in far-field conditions due to reverberation and noise, and traditional multi-channel ASR systems often use separate modules for beamforming and acoustic modeling, which can lead to suboptimal performance and require iterative parameter optimization.

Innovation Solution

A deep neural network (DNN) framework that performs joint beamforming and acoustic modeling using a convolutional long short-term memory (CLDNN) model, which processes raw audio waveforms from multiple microphones to learn filters that are robust to varying microphone spacings and noise conditions, thereby improving speech recognition accuracy.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Device complexity

If separate modules are used for beamforming and acoustic modeling, then system modularity is improved, but speech recognition accuracy deteriorates due to suboptimal performance and iterative parameter optimization requirements

Engineering Contradiction:
Improvesystem modularityVSAvoidspeech recognition accuracy
Core Design Contradiction:
Device complexityVSMeasurement precision

Solution Approach 1:

The patent combines beamforming and acoustic modeling into a unified deep neural network framework where spatial filtering and acoustic feature extraction are performed jointly. The CLDNN architecture integrates convolutional layers for spatial filtering with LSTM layers for temporal modeling, eliminating the need for separate beamforming and acoustic modeling modules. This unified approach allows end-to-end training that optimizes both spatial and acoustic parameters simultaneously, resolving the contradiction between modularity and accuracy.

Inventive Principle:
Principle #5Merging (Combining)

2Ease of manufacture

If traditional beamforming techniques are used, then computational simplicity is improved, but robustness to varying microphone spacings and noise conditions deteriorates

Engineering Contradiction:
Improvecomputational simplicityVSAvoidrobustness to microphone configurations
Core Design Contradiction:
Ease of manufactureVSReliability

Solution Approach 1:

The patent employs a dynamic approach where the deep neural network learns optimal filtering parameters adaptively during training rather than using fixed beamforming parameters. The CLDNN model adjusts its convolutional and recurrent parameters based on the specific acoustic environment and microphone configuration, enabling robust performance across varying microphone spacings and noise conditions while maintaining computational efficiency through differentiable operations.

Inventive Principle:
Principle #15Dynamics

3Speed

If manually defined features are extracted from audio waveforms, then processing speed is improved, but adaptability to varying acoustic conditions deteriorates

Engineering Contradiction:
Improveprocessing speedVSAvoidadaptability to acoustic conditions
Core Design Contradiction:
SpeedVSAdaptability or versatility

Solution Approach 1:

The patent implements self-service feature extraction where the deep neural network automatically learns the optimal acoustic features directly from raw audio waveforms during training. The CLDNN architecture performs feature learning end-to-end, eliminating the need for manual feature engineering while maintaining processing efficiency. The network adapts its feature extraction capabilities to different acoustic conditions through supervised training, resolving the contradiction between processing speed and adaptability.

Inventive Principle:
Principle #25Self-service

Data Source

PatentUS9697826B2Processing multi-channel audio waveforms
Publication Date: 2017.07.04 GOOGLE LLC
  • US9697826B2 patent drawing
  • US9697826B2 patent drawing
  • US9697826B2 patent drawing

AI summary

Methods, including computer programs encoded on a computer storage medium, for enhancing the processing of audio waveforms for speech recognition using various neural network processing techniques. In one aspect, a method includes: receiving multiple channels of audio data corresponding to an utterance; convolving each of multiple filters, in a time domain, with each of the multiple channels of audio waveform data to generate convolution outputs, wherein the multiple filters have parameters that have been learned during a training process that jointly trains the multiple filters and trains a deep neural network as an acoustic model; combining, for each of the multiple filters, the convolution outputs for the filter for the multiple channels of audio waveform data; inputting the combined convolution outputs to the deep neural network trained jointly with the multiple filters; and providing a transcription for the utterance that is determined.