Multi-lingual Text-to-Speech Acoustic Model Merging

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current multi-lingual text-to-speech systems face challenges in synthesizing speech with natural prosody when processing texts in multiple languages, often resulting in interrupted speech due to the lack of consideration for acoustic-prosodic models and accent weights between languages.

Innovation Solution

A multi-lingual text-to-speech system that employs an acoustic-prosodic model selection module and a mergence module to combine language models with controllable accent weighting, allowing for the transformation of non-native language text into speech with a native language accent, thereby adjusting pronunciation and prosody to match contextual scenarios.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Adaptability or versatility

If multiple synthesizers are used to switch for different languages, then language coverage is improved, but speech prosody continuity deteriorates

Engineering Contradiction:
Improvelanguage coverageVSAvoidspeech prosody continuity
Core Design Contradiction:
Adaptability or versatilityVSStability of the object's composition

Solution Approach 1:

The patent merges acoustic-prosodic models from multiple languages into a unified model structure. Instead of switching between separate synthesizers, the system combines L1 (native language) and L2 (non-native language) acoustic-prosodic models into a single integrated model that maintains consistent prosody across language boundaries, eliminating the interrupted speech effect while supporting multiple languages.

Inventive Principle:
Principle #5Merging (Combining)

2Productivity

If non-native language phonetics are mapped directly to native language phonetics, then synthesis speed is improved, but accent accuracy deteriorates

Engineering Contradiction:
Improvesynthesis speedVSAvoidaccent accuracy
Core Design Contradiction:
ProductivityVSManufacturing precision

Solution Approach 1:

The patent introduces controllable accent weighting parameters that adjust the degree of phonetic transformation from L2 to L1. By varying these parameters, the system can control the accent intensity in the synthesized speech, allowing users to balance between naturalness (higher L1 weighting) and accent preservation (lower L1 weighting), thus achieving both speed and accuracy.

Inventive Principle:
Principle #35Parameter changes

3Device complexity

If acoustic-prosodic models of different languages are merged without considering accent weights, then model complexity is reduced, but speech naturalness deteriorates

Engineering Contradiction:
Improvemodel complexityVSAvoidspeech naturalness
Core Design Contradiction:
Device complexityVSManufacturing precision

Solution Approach 1:

The patent implements dynamic accent weighting that adjusts the contribution of L1 and L2 acoustic-prosodic models based on the specific phonetic context. The system dynamically determines the optimal mixing ratio of native and non-native language features for each phonetic unit, allowing the accent weight to vary adaptively throughout the speech synthesis process, thereby maintaining speech naturalness while managing model complexity.

Inventive Principle:
Principle #15Dynamics

Data Source

PatentUS8898066B2Multi-lingual text-to-speech system and method
Publication Date: 2014.11.25 IND TECH RES INST
  • US8898066B2 patent drawing
  • US8898066B2 patent drawing
  • US8898066B2 patent drawing

AI summary

A multi-lingual text-to-speech system and method processes a text to be synthesized via an acoustic-prosodic model selection module and an acoustic-prosodic model mergence module, and obtains a phonetic unit transformation table. In an online phase, the acoustic-prosodic model selection module, according to the text and a phonetic unit transcription corresponding to the text, uses at least a set controllable accent weighting parameter to select a transformation combination and find a second and a first acoustic-prosodic models. The acoustic-prosodic model mergence module merges the two acoustic-prosodic models into a merged acoustic-prosodic model, according to the at least a controllable accent weighting parameter, processes all transformations in the transformation combination and generates a merged acoustic-prosodic model sequence. A speech synthesizer and the merged acoustic-prosodic model sequence are further applied to synthesize the text into an L1-accent L2 speech.