Multi-lingual Text-to-Speech Synthesizer with Inter-Lingual Tone
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Conventional Text-To-Speech (TTS) methods are ineffective in processing multi-lingual text messages, as they lack corresponding voice message databases for combinations of two or more languages, leading to inadequate conversion of mixed language text into voice.
Innovation Solution
A multi-lingual text-to-speech method that separates a mixed language text message into sections, converts each section into phoneme labels, and combines these with cognate connection tone information from respective language databases to produce a coherent multi-lingual voice message, using a processor and storage device to assemble phoneme sequences and generate inter-lingual connection tones.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Adaptability or versatility
If conventional TTS methods are used for single language processing, then the conversion accuracy for that language is maintained, but the system cannot effectively process multi-lingual text messages
Solution Approach 1:
The multi-lingual text message is segmented into separate language sections, with each section processed through its corresponding language database. This allows the system to maintain high conversion accuracy for each language while achieving multi-lingual versatility.
Solution Approach 2:
The TTS system is designed to handle multiple languages through a universal processing framework that can adapt to different language databases, achieving both multi-lingual versatility and maintained conversion accuracy through standardized processing steps.
2Manufacturing precision
If separate language databases are used for different languages, then the pronunciation accuracy for each language is preserved, but the system complexity increases
Solution Approach 1:
The system divides the text processing into separate language sections, each utilizing dedicated language databases. This segmentation preserves pronunciation accuracy for each language while organizing complexity into manageable, modular components.
Solution Approach 2:
A language detection and routing mechanism acts as an intermediary, directing different language sections to appropriate databases. This mediator manages the complexity of multiple databases while ensuring each language receives accurate processing.
3Ease of operation
If multi-lingual phoneme sequences are assembled from different language databases, then the fluency of inter-lingual transitions is improved, but the processing time increases
Solution Approach 1:
Connection tone information is pre-calculated and stored for phoneme boundaries between different languages. This preliminary preparation enables fluent inter-lingual transitions during assembly without requiring time-consuming real-time calculations.
Solution Approach 2:
Phoneme sequences from different language databases are merged into a unified multi-lingual phoneme sequence, with pre-computed connection tones ensuring smooth transitions. This combining approach maintains fluency while reducing processing time through efficient integration.
Data Source
AI summary
A text-to-speech method and a multi-lingual speech synthesizer using the method are disclosed. The multi-lingual speech synthesizer and the method executed by a processor are applied for processing a multi-lingual text message in a mixture of a first language and a second language into a multi-lingual voice message. The multi-lingual speech synthesizer comprises a storage device configured to store a first language model database, a second language model database, a broadcasting device configured to broadcast the multi-lingual voice message, and a processor, connected to the storage device and the broadcasting device, configured to execute the method disclosed herein.


