Acoustic Model Restructuring for Accent Adaptation
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current speech recognition systems face difficulties in accurately recognizing speech from speakers with strong regional accents or foreign accents due to the use of a single generic acoustic model, leading to reduced accuracy and usability for minority accent groups.
Innovation Solution
The system adapts automatic speech recognition by restructuring the acoustic model using a weighted sum of phonemes from a typical native speech model, creating a custom speech model that represents each phoneme as a combination of plausible phonemes, allowing for improved recognition of regional and foreign accents without altering the pronouncing dictionary.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Adaptability or versatility
If a single generic acoustic model is used to represent all speakers, then the system maintains simplicity and universality, but speech recognition accuracy deteriorates for speakers with strong regional accents or foreign accents
Solution Approach 1:
The patent segments the acoustic model into multiple dialect-specific acoustic models, each trained on speech data from speakers of a particular dialect or accent group. Instead of using a single generic model, the system divides the population into segments (e.g., General American, Southern American, British English, foreign accents) and creates specialized models for each segment, thereby improving recognition accuracy for minority accent groups while maintaining manageable complexity through modular organization
Solution Approach 2:
The patent applies local quality by tailoring the acoustic model characteristics to match specific dialect properties. Each dialect-specific model is trained with speech data that reflects the local phonetic patterns, pronunciation variations, and acoustic features of that particular accent group. This allows the system to adapt the acoustic properties locally for each dialect rather than applying a uniform generic model, thereby improving recognition accuracy for speakers with strong regional or foreign accents
2Adaptability or versatility
If pronunciation dictionaries are expanded to include all possible dialect variations, then coverage for minority accent groups improves, but model confusion increases and overall recognition accuracy deteriorates
Solution Approach 1:
The patent segments the pronunciation handling into dialect-specific acoustic models rather than expanding a single pronunciation dictionary. Each dialect-specific model is trained to recognize the phonetic patterns and pronunciation variations characteristic of that dialect, allowing the system to handle dialect diversity through specialized models rather than through a confused universal dictionary with all possible variations
Solution Approach 2:
The patent inverts the traditional approach by keeping the pronunciation dictionary relatively stable and instead adapting the acoustic models to match different dialects. Rather than modifying the dictionary to accommodate all dialect variations (which causes confusion), the system inverts the problem by training multiple acoustic models on dialect-specific data, allowing each model to interpret the stable dictionary through the lens of its dialect's acoustic characteristics
Data Source
AI summary
Disclosed herein are systems, computer-implemented methods, and computer-readable storage media for recognizing speech by adapting automatic speech recognition pronunciation by acoustic model restructuring. The method identifies an acoustic model and a matching pronouncing dictionary trained on typical native speech in a target dialect. The method collects speech from a new speaker resulting in collected speech and transcribes the collected speech to generate a lattice of plausible phonemes. Then the method creates a custom speech model for representing each phoneme used in the pronouncing dictionary by a weighted sum of acoustic models for all the plausible phonemes, wherein the pronouncing dictionary does not change, but the model of the acoustic space for each phoneme in the dictionary becomes a weighted sum of the acoustic models of phonemes of the typical native speech. Finally the method includes recognizing via a processor additional speech from the target speaker using the custom speech model.


