Acoustic Model Feature Adaptation for Speech Recognition
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing speech recognition systems face challenges in accurately adapting acoustic models to different environments due to discrepancies in training data conditions, leading to suboptimal performance when using small field data in conjunction with larger in-house data, especially when confidential information restricts the availability of large field data.
Innovation Solution
The method involves calculating standard deviation values from both small field and in-house data features, adjusting the in-house data features to match the distribution of small field data, and reconstructing the acoustic model using these modified features to improve accuracy in speech recognition within the original environment.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If acoustic model is trained using only small field data, then the model adapts well to the target environment, but the amount of training data is insufficient leading to poor recognition accuracy
Solution Approach 1:
The patent combines small field data and large in-house data into a unified training dataset. By merging these two data sources and applying distribution alignment, the system achieves both the quantity advantage of large data and the environment-specific adaptability of field data, resolving the contradiction between data volume and environmental adaptation.
Solution Approach 2:
The patent transforms the feature distribution parameters of in-house data to match those of field data using standard deviation ratios. This parameter transformation allows the model to learn from large in-house data while adapting to field environment characteristics, overcoming the limitation of insufficient field data volume.
2Quantity of substance
If acoustic model is trained using large in-house data, then the model has sufficient training data volume, but the feature distribution mismatch with field environment reduces recognition accuracy
Solution Approach 1:
The patent modifies the statistical parameters (standard deviation) of in-house data features to align with field data distribution. This parameter adjustment enables the model to utilize large in-house data effectively while maintaining compatibility with field environment characteristics, resolving the distribution mismatch problem.
Solution Approach 2:
The patent introduces a feature transformation process as an intermediary between in-house data and the acoustic model. This intermediary step aligns the feature distributions, allowing the model to benefit from large in-house data without suffering from environment mismatch, thus improving both data utilization and recognition accuracy.
3Adaptability or versatility
If standard field data is used for training, then the acoustic model matches the target environment, but confidential information restrictions limit data availability
Solution Approach 1:
The patent creates a statistical copy of field data distribution characteristics using in-house data. By copying the essential distribution parameters (standard deviation) of field data and applying them to in-house data, the system achieves environmental adaptation without needing access to confidential field data, thus maintaining adaptability while working within data availability constraints.
Data Source
AI summary
Embodiments include methods and systems for improving an acoustic model. Aspects include acquiring a first standard deviation value by calculating standard deviation of a feature from first training data and acquiring a second standard deviation value by calculating standard deviation of a feature from second training data acquired in a different environment from an environment of the first training data. Aspects also include creating a feature adapted to an environment where the first training data is recorded, by multiplying the feature acquired from the second training data by a ratio obtained by dividing the first standard deviation value by the second standard deviation value. Aspects further include reconstructing an acoustic model constructed using training data acquired in the same environment as the environment of the first training data using the feature adapted to the environment where the first training data is recorded.


