Speech Synthesis Model with Decoupled Style and Tone Encoders
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current speech synthesis models are limited to single-tone and single-style speech synthesis, failing to effectively handle multi-style and multi-tone synthesis, which restricts their application in diverse speech interaction scenarios and user experience.
Innovation Solution
A method and apparatus for synthesizing speech that acquires style and tone information, along with content, to generate acoustic features using a pre-trained speech synthesis model, enabling cross-language, cross-style, and cross-tone speech synthesis by decoupling content, style, and tone encoders within the Tacotron structure and utilizing a neural vocoder for audio generation.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Adaptability or versatility
If a speech synthesis model is trained to perform single-tone and single-style synthesis, then the model structure remains simple and training is easier, but the model cannot handle multi-style and multi-tone synthesis requirements
Solution Approach 1:
The patent divides the speech synthesis model into separate encoders for content, style, and tone information. Each encoder processes specific aspects independently, allowing the model to handle multi-style and multi-tone synthesis by combining these separate representations rather than requiring a completely new model architecture.
Solution Approach 2:
The patent creates a universal speech synthesis model that can perform multiple functions: single-tone single-style synthesis, multi-style synthesis, multi-tone synthesis, and cross-language synthesis. The model achieves this universality through the combination of content, style, and tone encoders that can be activated in different combinations for different synthesis tasks.
2Adaptability or versatility
If training data is collected for various styles and speakers to enable multi-style synthesis, then the synthesis versatility improves, but the data collection and processing complexity increases
Solution Approach 1:
The patent segments the training data into separate categories for content, style, and tone information. By organizing data this way, the model can learn each aspect independently and combine them during synthesis, reducing the complexity of processing and utilizing diverse training data for multiple styles and speakers.
3Adaptability or versatility
If the speech synthesis model is enhanced to support cross-language and cross-style synthesis, then the application scope expands, but the computational requirements and processing time increase
Solution Approach 1:
The patent segments the computational processing into separate encoder modules that handle content, style, and tone information independently. This segmentation allows for more efficient computation compared to processing all information simultaneously, as each module can operate in parallel and the results are combined at the output stage, reducing overall computational resource consumption.
Data Source
AI summary
The present disclosure provides a method and apparatus of synthesizing a speech, a method and apparatus of training a speech synthesis model, an electronic device, and a storage medium. The method of synthesizing a speech includes acquiring a style information of a speech to be synthesized, a tone information of the speech to be synthesized, and a content information of a text to be processed; generating an acoustic feature information of the text to be processed, by using a pre-trained speech synthesis model, based on the style information, the tone information, and the content information of the text to be processed; and synthesizing the speech for the text to be processed, based on the acoustic feature information of the text to be processed.


