Text to Speech Accent Adaptation via Region-Specific Phoneme Selection
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current Speech to Text (STT) and Text to Speech (TTS) systems require separate, lengthy training processes and struggle to capture user pronunciations of domain terminology effectively, leading to a one-size-fits-all approach in TTS systems, which can result in unfamiliar pronunciation patterns for users with varying dialects and accents.
Innovation Solution
The system determines region-specific phoneme sequences for domain terminology from evaluated STT data, adapting TTS outputs to match the user's accent by using a region-specific pronunciation dictionary, allowing for customized text-to-speech responses that reflect the user's local pronunciation differences.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If separate training processes are used for STT and TTS systems, then each system can be optimized independently, but the training time and complexity increase significantly
Solution Approach 1:
The patent combines STT and TTS training processes into a unified framework where phoneme sequence data from STT training directly informs TTS training. This integration allows both systems to be optimized together rather than separately, reducing total training time while maintaining or improving optimization quality through shared phoneme representations.
2Ease of manufacture
If default phoneme sequences are used in TTS systems, then the system is simpler to implement, but it cannot adapt to users with varying dialects and accents
Solution Approach 1:
The patent implements dynamic phoneme sequence selection in TTS systems. Instead of using fixed default phoneme sequences, the system dynamically selects phoneme sequences based on the user's detected accent or dialect, allowing the TTS output to adapt to different users while maintaining a unified system architecture.
Solution Approach 2:
The patent applies local quality by using different phoneme sequences for different regions or accent groups. The system maintains multiple phoneme sequence variants and selects the appropriate local variant based on user characteristics, allowing customization for specific dialects without redesigning the entire system.
3Ease of operation
If region-specific phoneme sequences are implemented, then user familiarity and comfort improve, but the system complexity and data requirements increase
Solution Approach 1:
The patent performs preliminary action by pre-processing and organizing phoneme sequence data according to different regions and accents during the training phase. This pre-organization allows the TTS system to quickly select appropriate phoneme sequences at runtime without complex real-time analysis, reducing operational complexity while maintaining user familiarity.
Data Source
AI summary
A system and method for providing a text to speech output by receiving user audio data, determining a user region-specific-pronunciation classification according to the audio data, determining text for a response to the user according to the audio data, identifying a portion from the text, where a region specific-pronunciation dictionary includes the portion, and using a phoneme string, from the dictionary selected according to the user region-specific pronunciation classification, for the word in a text to speech output to the user.


