Multi-lingual Text-to-Speech Acoustic Model Merging
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current multi-lingual text-to-speech systems face challenges in synthesizing speech with natural prosody when processing texts in multiple languages, often resulting in interrupted speech due to the lack of consideration for acoustic-prosodic models and accent weights between languages.
Innovation Solution
A multi-lingual text-to-speech system that employs an acoustic-prosodic model selection module and a mergence module to combine language models with controllable accent weighting, allowing for the transformation of non-native language text into speech with a native language accent, thereby adjusting pronunciation and prosody to match contextual scenarios.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Adaptability or versatility
If multiple synthesizers are used to switch for different languages, then language coverage is improved, but speech prosody continuity deteriorates
Solution Approach 1:
The patent merges acoustic-prosodic models from multiple languages into a unified model structure. Instead of switching between separate synthesizers, the system combines L1 (native language) and L2 (non-native language) acoustic-prosodic models into a single integrated model that maintains consistent prosody across language boundaries, eliminating the interrupted speech effect while supporting multiple languages.
2Productivity
If non-native language phonetics are mapped directly to native language phonetics, then synthesis speed is improved, but accent accuracy deteriorates
Solution Approach 1:
The patent introduces controllable accent weighting parameters that adjust the degree of phonetic transformation from L2 to L1. By varying these parameters, the system can control the accent intensity in the synthesized speech, allowing users to balance between naturalness (higher L1 weighting) and accent preservation (lower L1 weighting), thus achieving both speed and accuracy.
3Device complexity
If acoustic-prosodic models of different languages are merged without considering accent weights, then model complexity is reduced, but speech naturalness deteriorates
Solution Approach 1:
The patent implements dynamic accent weighting that adjusts the contribution of L1 and L2 acoustic-prosodic models based on the specific phonetic context. The system dynamically determines the optimal mixing ratio of native and non-native language features for each phonetic unit, allowing the accent weight to vary adaptively throughout the speech synthesis process, thereby maintaining speech naturalness while managing model complexity.
Data Source
AI summary
A multi-lingual text-to-speech system and method processes a text to be synthesized via an acoustic-prosodic model selection module and an acoustic-prosodic model mergence module, and obtains a phonetic unit transformation table. In an online phase, the acoustic-prosodic model selection module, according to the text and a phonetic unit transcription corresponding to the text, uses at least a set controllable accent weighting parameter to select a transformation combination and find a second and a first acoustic-prosodic models. The acoustic-prosodic model mergence module merges the two acoustic-prosodic models into a merged acoustic-prosodic model, according to the at least a controllable accent weighting parameter, processes all transformations in the transformation combination and generates a merged acoustic-prosodic model sequence. A speech synthesizer and the merged acoustic-prosodic model sequence are further applied to synthesize the text into an L1-accent L2 speech.


