Frequency Aligned Network for Multi-Channel Speech Recognition
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Conventional speech recognition systems are inefficient in processing multi-channel audio data, requiring separate optimization of components for signal enhancement and often needing more microphones, which increases bandwidth requirements and noise sensitivity.
Innovation Solution
A deep neural network (DNN)-based acoustic model front-end with a frequency aligned network (FAN) architecture performs spatial filtering and feature extraction, jointly optimizing components for speech recognition and reducing the number of microphones required, while minimizing bandwidth by uploading low-dimensional feature vectors.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Adaptability or versatility
If conventional speech recognition systems use separate components for signal enhancement, then speech recognition can be performed, but the system requires more microphones and increases bandwidth requirements
Solution Approach 1:
The patent combines spatial filtering and feature extraction into a single integrated DNN module. The multi-channel audio data is processed through the DNN which simultaneously performs spatial filtering to separate speech from noise and extracts acoustic features for recognition, eliminating the need for separate signal enhancement components and reducing microphone requirements
Solution Approach 2:
The DNN-based acoustic model front-end serves multiple functions: it performs spatial filtering for noise reduction, feature extraction for speech recognition, and adaptive processing for different acoustic environments. This multi-functional approach replaces multiple specialized components with a single universal processor
2Adaptability or versatility
If conventional speech recognition systems use separate components for signal enhancement, then speech recognition can be performed, but bandwidth requirements increase
Solution Approach 1:
The system extracts only the essential acoustic features from multi-channel audio data using the DNN, rather than transmitting or processing the complete audio signals. The DNN extracts low-dimensional feature vectors that capture the critical speech information while filtering out redundant data, significantly reducing bandwidth requirements
3Adaptability or versatility
If conventional speech recognition systems process multi-channel audio data, then speech recognition can be performed, but noise sensitivity increases
Solution Approach 1:
The patent integrates spatial filtering and feature extraction in the DNN architecture, allowing the system to perform noise reduction and feature extraction simultaneously. This joint optimization enables the system to filter out noise while extracting speech features more effectively than sequential processing
4Ease of manufacture
If conventional speech recognition systems use separate optimization for components, then signal enhancement can be performed, but system complexity increases
Solution Approach 1:
The patent merges spatial filtering and feature extraction into a single DNN module with unified optimization. The system uses a joint objective function that optimizes both spatial filtering performance and feature extraction quality simultaneously, eliminating the need for separate component optimization and reducing system complexity
Data Source
AI summary
Techniques for speech processing using a deep neural network (DNN) based acoustic model front-end are described. A new modeling approach directly models multi-channel audio data received from a microphone array using a first model (e.g., multi-geometry/multi-channel DNN) that includes a frequency aligned network (FAN) architecture. Thus, the first model may perform spatial filtering to generate a first feature vector by processing individual frequency bins separately, such that multiple frequency bins are not combined. The first feature vector may be used similarly to beamformed features generated by an acoustic beamformer. A second model (e.g., feature extraction DNN) processes the first feature vector and transforms it to a second feature vector having a lower dimensional representation. A third model (e.g., classification DNN) processes the second feature vector to perform acoustic unit classification and generate text data. The DNN front-end enables improved performance despite a reduction in microphones.


