Front-end Processor Speech Recognition Noise Adaptation
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Speech recognition systems face reduced recognition rates due to noise and distortion in test environments that differ from the conditions under which the acoustic model was recorded, leading to a mismatch between the basic and test environments.
Innovation Solution
A front-end processor and method that convert speech from a test environment to a format similar to the basic environment using a linear dynamic system, specifically through feature vector-sequence conversion, clustering, and applying conversion rules to improve recognition rates by removing noise and distortion, employing techniques like Vector Quantization, Gaussian Mixture Models, and Expectation Maximization algorithms.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Adaptability or versatility
If speech is recorded in a test environment with noise and distortion, then the speech recognizing apparatus can process real-world speech, but the recognition rate deteriorates due to mismatch with the acoustic model trained in a basic environment
Solution Approach 1:
The patent introduces a front-end processor as an intermediary component that converts speech features from the test environment into a format compatible with the acoustic model trained in the basic environment. This processor acts as a mediator between the noisy test environment and the clean basic environment, translating feature vectors through conversion rules learned from paired data, thereby enabling the system to maintain high recognition rates across different environments without requiring the acoustic model itself to be environment-specific
Solution Approach 2:
The patent changes the parameter representation of speech by transforming feature vectors from the test environment into a different feature space that matches the acoustic model's expectations. Through conversion rules derived from comparing test environment speech with basic environment speech, the system modifies acoustic parameters (such as spectral characteristics, temporal features, or mel-frequency characteristics) to bridge the environmental gap, allowing the same acoustic model to perform accurately on speech from diverse recording conditions
2Reliability
If the acoustic model is trained using high quality equipment in a favorable environment, then the recognition accuracy is high, but the system cannot handle speech recorded in noisy or distorted environments
Solution Approach 1:
The patent segments the speech processing system into distinct functional components: a front-end processor that handles environment-specific conversions and a core speech recognizing apparatus that maintains a fixed acoustic model. This segmentation allows the acoustic model to be trained once on clean speech data while the front-end processor adapts to different recording environments, separating the universal recognition function from the environment-specific adaptation function
Solution Approach 2:
The patent creates a virtual copy of the basic environment speech characteristics through the front-end processor. By learning conversion rules from paired test and basic environment speech data, the processor generates transformed feature vectors that replicate the acoustic characteristics of basic environment speech, even when the input comes from noisy test environments. This copying approach allows the acoustic model to operate as if it received clean speech input
Data Source
AI summary
A method of recognizing speech is provided. The method includes the operations of (a) dividing first speech that is input to a speech recognizing apparatus into frames; (b) converting the frames of the first speech into frames of second speech by applying conversion rules to the divided frames, respectively; and (c) recognizing, by the speech recognizing apparatus, the frames of the second speech, wherein (b) comprises converting the frames of the first speech into the frames of the second speech by reflecting at least one frame from among the frames that are previously positioned with respect to a frame of the first speech.


