Multilingual Speech Synthesis Using Virtual Speaker Timbre
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing multilingual speech synthesis technologies face challenges in producing natural, accurate, and uniform timbre, particularly due to the reliance on data from different mother tongue speakers or professional multilingual speakers, which increases costs and reduces user experience.
Innovation Solution
A speech synthesis method that determines language types, adapts spectrum and fundamental frequency parameter models based on a target timbre, and adjusts parameters to generate uniform timbre for multilingual speech, using monolingual data to reduce dependence on professional multilingual speaker data.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Manufacturing precision
If data from different mother tongue speakers or professional multilingual speakers is used for speech synthesis, then pronunciation accuracy and timbre uniformity can be improved, but data collection costs and implementation difficulty increase
Solution Approach 1:
The patent uses voice print technology to create a virtual copy of the target speaker's timbre characteristics. Instead of collecting actual speech data from professional multilingual speakers, the system synthesizes a virtual speaker model that replicates the desired timbre, thereby avoiding the high costs and difficulties of recruiting and recording data from native speakers of multiple languages.
Solution Approach 2:
The system adjusts spectral parameters and fundamental frequency parameters through adaptive transformation to match the target timbre. By modifying these acoustic parameters, the patent achieves uniform timbre across different languages without requiring data from speakers with matching natural timbres, thus resolving the contradiction between timbre quality and data collection feasibility.
2Adaptability or versatility
If data from different mother tongue speakers is used for multilingual speech synthesis, then language coverage can be improved, but timbre uniformity deteriorates
Solution Approach 1:
The patent creates a unified virtual speaker model that serves as a consistent timbre source for all languages. Rather than using actual speakers for each language (which would provide language coverage but inconsistent timbre), the system copies the target speaker's timbre characteristics into a single virtual model that generates speech in multiple languages with uniform timbre.
Solution Approach 2:
The system separates the language-specific features from the timbre features. Language coverage is achieved by training on multiple languages while timbre uniformity is maintained by applying voice print transformation separately to each language's speech data, allowing independent optimization of both aspects.
3Reliability
If professional multilingual speaker data is collected, then speech naturalness can be improved, but implementation complexity increases
Solution Approach 1:
The patent replaces the mechanical process of recruiting, recording, and managing data from professional multilingual speakers with an automated voice print synthesis system. The mechanical complexity of data collection from multiple speakers is substituted by an algorithmic system that generates synthetic speech with consistent timbre, significantly reducing implementation complexity while maintaining speech naturalness.
Data Source
AI summary
A speech synthesis method and device. The method comprises: determining language types of a statement to be synthesized; determining base models corresponding to the language types; determining a target timbre, performing adaptive transformation on the spectrum parameter models based on the target timbre, and training the statement to be synthesized based on the spectrum parameter models subjected to adaptive transformation to generate spectrum parameters; training the statement to be synthesized based on the fundamental frequency parameters to generate fundamental frequency parameters, and adjusting the fundamental frequency parameters based on the target timbre; and synthesizing the statement to be synthesized into a target speech based on the spectrum parameters, and the fundamental frequency parameters after adjusting.


