Decoupled Speech Synthesis Model for Target Timbre Control
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current speech synthesis technologies face challenges in converting text into audio with specific timbre, limiting their ability to provide individualized speech synthesis for applications like video creation and customer service.
Innovation Solution
A speech synthesis model decoupled into two sub-models for feature extraction, where the first sub-model outputs bottleneck features and the second sub-model generates Mel spectrum features, allowing independent control of timbre and other features, enabling the synthesis of audio with target timbre.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Adaptability or versatility
If traditional speech synthesis models are used, then the synthesis process can be completed, but the ability to generate audio with specific timbre is limited
Solution Approach 1:
The speech synthesis model is divided into two independent sub-models: a first sub-model that extracts acoustic features (including bottleneck features) from text, and a second sub-model that extracts Mel spectrum features from acoustic features. This segmentation allows independent control and optimization of timbre characteristics through the second sub-model while maintaining overall synthesis quality.
Solution Approach 2:
The patent applies local quality by making the second sub-model specifically responsible for timbre-related Mel spectrum feature extraction, while the first sub-model handles general acoustic feature extraction. This allows different parts of the system to have specialized functions, with the second sub-model optimized locally for timbre control.
2Manufacturing precision
If more labeled sample audios are used to improve timbre accuracy, then the synthesis quality improves, but the complexity and cost of data preparation increases
Solution Approach 1:
The patent extracts timbre characteristics into a separate second sub-model that independently processes Mel spectrum features. This extraction allows the system to focus computational resources on timbre accuracy without requiring proportional increases in labeled data, as the second sub-model can be trained separately with targeted timbre data.
Solution Approach 2:
The patent changes the parameter representation by introducing Mel spectrum features as an intermediate representation between acoustic features and final audio output. This parameter transformation enables more efficient control of timbre characteristics and reduces the amount of labeled data needed compared to direct end-to-end training.
3Adaptability or versatility
If a unified speech synthesis model is used, then the model structure is simple, but independent control of timbre and other features is not achieved
Solution Approach 1:
The unified model is segmented into two sub-models that work sequentially, with the first sub-model outputting acoustic features and the second sub-model processing these to generate Mel spectrum features. This segmentation provides independent control interfaces for different feature types while maintaining a relatively simple overall architecture.
Solution Approach 2:
The first sub-model serves a universal function by extracting general acoustic features that are then fed into the second sub-model. This multi-functional design allows the first sub-model to handle various text inputs while the second sub-model specializes in timbre control, achieving both versatility and specialized control.
Data Source
AI summary
Provided are an audio synthesis method and apparatus, an electronic device, and a readable storage medium. In the present solution, conversion from a text to an audio having a target timbre is achieved by means of a pre-trained voice synthesis model, the voice synthesis model comprising a first feature extraction sub-model and a second feature extraction sub-model, wherein the first feature extraction sub-model outputs, according to an inputted text to be processed, an acoustic feature comprising a bottleneck feature; the second feature extraction sub-model outputs, according to the inputted first acoustic features, a Mel spectrum feature corresponding to the text to be processed; according to the Mel spectrum feature corresponding to the text to be processed, the target audio corresponding to the text to be processed is obtained, and the target audio has the target timbre.


