Speech Recognition Acoustic Model Training via Initial Reflection Extraction
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing speech recognition systems face challenges in achieving high recognition rates due to the influence of reverberations in environments, which are difficult to remove, especially when using a single microphone, as real-time processing of large data sets is complex and inefficient.
Innovation Solution
A processing unit and system that extracts initial reflection components from the reverberation pattern of an impulse response, using these components to enhance the acoustic model for speech recognition, and creates a subtraction filter to remove diffuse reverberation components, improving recognition accuracy.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If real-time inverse-filtering is performed on frequency-domain signals to remove reverberation influences, then speech recognition rate is improved, but processing complexity and data volume become enormous making real-time execution difficult
Solution Approach 1:
The patent segments the reverberation removal task into two distinct phases: offline acoustic model learning that incorporates reverberation characteristics, and online speech recognition that uses the pre-trained model. This segmentation transfers the computational burden from real-time processing to offline preparation, resolving the contradiction between recognition accuracy and processing complexity.
Solution Approach 2:
The patent performs preliminary action by training the acoustic model offline with speech data that has been convolved with impulse response to simulate reverberation effects. This preliminary training enables the model to inherently understand and compensate for reverberation, eliminating the need for complex real-time inverse-filtering during actual speech recognition.
2Measurement precision
If multiple microphones are used to remove reflection waves reflected by walls, then reverberation removal effectiveness is improved, but device complexity and cost increase
Solution Approach 1:
The patent creates a virtual copy of the acoustic environment by convolving speech data with measured or simulated impulse response to generate training data that mimics reverberation conditions. This copying approach allows the system to learn reverberation characteristics without requiring multiple physical microphones to capture and separate reflection waves.
Solution Approach 2:
The patent replaces the mechanical approach of using multiple microphones to physically separate and remove reflection waves with a computational approach where a single microphone captures the mixed signal and an acoustic model learned from convolved training data performs the separation and recognition.
3Measurement precision
If inverse filter is estimated for input acoustic signals to remove reverberation components, then speech recognition rate is improved, but processing time increases due to enormous data to be processed
Solution Approach 1:
The patent performs the computationally intensive inverse-filtering operation in advance during offline acoustic model training rather than in real-time during speech recognition. The acoustic model is trained on speech data that has been pre-processed with convolution to simulate reverberation, enabling the model to learn optimal filtering characteristics without real-time processing delays.
Solution Approach 2:
The patent divides the processing into offline model training phase where extensive data processing occurs, and online recognition phase where pre-trained models are applied. This segmentation moves the time-consuming operations to offline execution, making real-time speech recognition feasible.
Data Source
AI summary
A processing unit is provided which executes speech recognition on speech signals captured by a microphone for capturing sounds uttered in an environment. The processing unit has: an initial reflection component extraction portion that extracts initial reflection components by removing diffuse reverberation components from a reverberation pattern of an impulse response generated in the environment; and an acoustic model learning portion that learns an acoustic model for the speech recognition by reflecting the initial reflection components to speech data for learning.


