Speech Synthesis System with Language-Specific Acoustic Models
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current voice dubbing and text-to-speech systems face challenges in producing voices in multiple languages due to the scarcity of voice actors who can record in multiple languages, making it difficult to create products that require specific voices in various languages.
Innovation Solution
A speech synthesis system with an operating interface, storage unit, and processor that allows users to select output languages and access acoustic models corresponding to different languages, enabling the generation of speech data that mimics a specific vocal, even if the voice actor has only recorded in one language.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Device complexity
If a single voice actor records corpus in one language, then the production cost and complexity are reduced, but the ability to produce speech in multiple languages is lost
Solution Approach 1:
The system segments the speech production process into separate language-specific acoustic models. Each acoustic model is trained on corpus data for a specific language, allowing the system to process different languages independently while sharing the same voice synthesis architecture. This segmentation enables multi-language support without requiring a single actor to record all languages.
Solution Approach 2:
The speech synthesis system achieves universality by designing a common synthesis framework that can process multiple languages through different acoustic models. The system structure remains universal while adapting to different languages by selecting the appropriate acoustic model based on the input language, eliminating the need for separate synthesis systems for each language.
2Reliability
If multiple voice actors are used to record corpus in different languages, then the speech quality and authenticity in each language are improved, but the production cost and coordination complexity increase
Solution Approach 1:
The system creates acoustic models that capture and replicate the vocal characteristics of a specific voice actor for different languages. Instead of using multiple actors, the system copies the target voice's timbre, tone, and stylistic features into language-specific acoustic models, allowing a single actor's voice to be synthesized in multiple languages with consistent quality.
3Measurement precision
If acoustic models are trained with language-specific corpus data, then the speech synthesis accuracy for each language is improved, but the data processing and storage requirements increase
Solution Approach 1:
The system applies local quality by training each acoustic model with language-specific corpus data that reflects the phonetic, phonological, and prosodic characteristics of that particular language. This localized training approach ensures high synthesis accuracy for each language while maintaining distinct data requirements for each model, rather than using a universal corpus that would compromise language-specific accuracy.
Data Source
AI summary
A speech synthesis system includes an operating interface, a storage unit and a processor. The operating interface provides a plurality of language options for a user to select one output language option therefrom. The storage unit stores a plurality of acoustic models. Each acoustic model corresponds to one of the language options and includes a plurality of phoneme labels corresponding to a specific vocal. The processor receives a text file and generates output speech data corresponding to the specific vocal according to the text file, a speech synthesizer, and one of the acoustic models which corresponds to the output language option.

