Cross-Lingual Speech Synthesis Dictionary Creation Using Mapping Tables
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing cross-lingual speaker adaptation techniques require large quantities of high-quality bilingual data to improve the quality of synthetic speech, which is a significant disadvantage.
Innovation Solution
A speech synthesis dictionary creation device that includes a mapping table creator, an estimator, and a dictionary creator, which creates a mapping table based on the similarity between speech synthesis dictionaries of a specific speaker in one language and another, and estimates a transformation matrix to adapt the dictionary for a target speaker in the same language, allowing for the creation of a speech synthesis dictionary for a target speaker in a second language using limited data.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Manufacturing precision
If cross-lingual speaker adaptation is conducted using existing techniques, then the quality of synthetic speech can be improved, but large quantities of high-quality bilingual data are required
Solution Approach 1:
The patent segments the speaker adaptation process into two independent stages: first adapting a speech synthesis dictionary to a specific speaker in a first language, then adapting it to a target speaker in a second language. This segmentation allows each adaptation to use only monolingual data from the respective language, eliminating the need for large quantities of bilingual data while maintaining high-quality synthetic speech output.
2Reliability
If cross-lingual speaker adaptation is conducted using existing techniques, then the quality of synthetic speech can be improved, but extensive bilingual data are required
Solution Approach 1:
The patent introduces a mapping table as an intermediary component that connects the speech synthesis dictionary of the first language to the second language. This mapping table, created by the mapping table creator, enables the system to transfer speaker characteristics across languages without requiring extensive bilingual data, thereby reducing data requirement complexity while maintaining reliable synthetic speech quality.
Data Source
AI summary
According to an embodiment, a device includes a table creator, an estimator, and a dictionary creator. The table creator is configured to create a table based on similarity between distributions of nodes of speech synthesis dictionaries of a specific speaker in respective first and second languages. The estimator is configured to estimate a matrix to transform the speech synthesis dictionary of the specific speaker in the first language to a speech synthesis dictionary of a target speaker in the first language, based on speech and a recorded text of the target speaker in the first language and the speech synthesis dictionary of the specific speaker in the first language. The dictionary creator is configured to create a speech synthesis dictionary of the target speaker in the second language, based on the table, the matrix, and the speech synthesis dictionary of the specific speaker in the second language.


