Speech Recognition Acoustic Model Training via Initial Reflection Extraction

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing speech recognition systems face challenges in achieving high recognition rates due to the influence of reverberations in environments, which are difficult to remove, especially when using a single microphone, as real-time processing of large data sets is complex and inefficient.

Innovation Solution

A processing unit and system that extracts initial reflection components from the reverberation pattern of an impulse response, using these components to enhance the acoustic model for speech recognition, and creates a subtraction filter to remove diffuse reverberation components, improving recognition accuracy.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If real-time inverse-filtering is performed on frequency-domain signals to remove reverberation influences, then speech recognition rate is improved, but processing complexity and data volume become enormous making real-time execution difficult

Engineering Contradiction:
Improvespeech recognition rateVSAvoidprocessing complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The patent segments the reverberation removal task into two distinct phases: offline acoustic model learning that incorporates reverberation characteristics, and online speech recognition that uses the pre-trained model. This segmentation transfers the computational burden from real-time processing to offline preparation, resolving the contradiction between recognition accuracy and processing complexity.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent performs preliminary action by training the acoustic model offline with speech data that has been convolved with impulse response to simulate reverberation effects. This preliminary training enables the model to inherently understand and compensate for reverberation, eliminating the need for complex real-time inverse-filtering during actual speech recognition.

Inventive Principle:
Principle #10Preliminary action

2Measurement precision

If multiple microphones are used to remove reflection waves reflected by walls, then reverberation removal effectiveness is improved, but device complexity and cost increase

Engineering Contradiction:
Improvereverberation removal effectivenessVSAvoidnumber of microphones
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The patent creates a virtual copy of the acoustic environment by convolving speech data with measured or simulated impulse response to generate training data that mimics reverberation conditions. This copying approach allows the system to learn reverberation characteristics without requiring multiple physical microphones to capture and separate reflection waves.

Inventive Principle:
Principle #26Copying

Solution Approach 2:

The patent replaces the mechanical approach of using multiple microphones to physically separate and remove reflection waves with a computational approach where a single microphone captures the mixed signal and an acoustic model learned from convolved training data performs the separation and recognition.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

3Measurement precision

If inverse filter is estimated for input acoustic signals to remove reverberation components, then speech recognition rate is improved, but processing time increases due to enormous data to be processed

Engineering Contradiction:
Improvespeech recognition rateVSAvoidprocessing time
Core Design Contradiction:
Measurement precisionVSLoss of time

Solution Approach 1:

The patent performs the computationally intensive inverse-filtering operation in advance during offline acoustic model training rather than in real-time during speech recognition. The acoustic model is trained on speech data that has been pre-processed with convolution to simulate reverberation, enabling the model to learn optimal filtering characteristics without real-time processing delays.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The patent divides the processing into offline model training phase where extensive data processing occurs, and online recognition phase where pre-trained models are applied. This segmentation moves the time-consuming operations to offline execution, making real-time speech recognition feasible.

Inventive Principle:
Principle #1Segmentation

Data Source

PatentUS8645130B2Processing unit, speech recognition apparatus, speech recognition system, speech recognition method, storage medium storing speech recognition program
Publication Date: 2014.02.04 TOYOTA JIDOSHA KK
  • US8645130B2 patent drawing
  • US8645130B2 patent drawing
  • US8645130B2 patent drawing

AI summary

A processing unit is provided which executes speech recognition on speech signals captured by a microphone for capturing sounds uttered in an environment. The processing unit has: an initial reflection component extraction portion that extracts initial reflection components by removing diffuse reverberation components from a reverberation pattern of an impulse response generated in the environment; and an acoustic model learning portion that learns an acoustic model for the speech recognition by reflecting the initial reflection components to speech data for learning.