Speech Synthesis Using Unified Decoding and Attention Models
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current speech synthesis technologies face issues with rhythm inconsistency and poor speech quality due to independent processes for generating recordings and synthesized speech, leading to discrepancies in speech speed and tone.
Innovation Solution
A method and apparatus that utilize a preset decoding model, attention model, and feature vector set to predict acoustic features for synthesized speech, incorporating both recorded and query result statements, ensuring consistent rhythm and improved quality by performing feature conversion and synthesis on predicted acoustic features.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Productivity
If independent processes are used for generating recording and synthesized speech, then the processes can be processed separately and efficiently, but the speech speed and tone become inconsistent leading to rhythm inconsistency
Solution Approach 1:
The patent merges the previously independent recording generation process and synthesized speech generation process into a unified speech synthesis system. The decoding model integrates both processes, using the same acoustic feature predictions for both the recorded statement portion and the query result statement portion, ensuring consistent speech speed and tone throughout the complete speech.
2Device complexity
If recording and synthesized speech are generated independently, then processing can be simplified, but transition time uncertainty increases and speech quality deteriorates
Solution Approach 1:
The patent implements feedback mechanisms where the decoding model uses attention models to dynamically adjust acoustic feature predictions based on the input symbol sequence. The model continuously refines its predictions by comparing expected outputs with actual outputs, ensuring consistent speech characteristics across the transition from recorded to synthesized portions, thereby improving speech quality and reducing transition time uncertainty.
Data Source
AI summary
Disclosed are a speech synthesis method and apparatus, and a storage medium. The method comprises: acquiring a symbol sequency of a statement to be synthesized, wherein the statement to be synthesized comprises a recorded statement characterizing a target object and a query result statement for the target object; encoding the symbol sequence by using a pre-set encoding model, in order to obtain a feature vector set; acquiring recording acoustic features corresponding to the recorded statement; predicting, according to a pre-set decoding model, the feature vector set, a pre-set attention model and the recording acoustic features, acoustic features corresponding to the statement to be synthesized, in order to obtain predicted acoustic features corresponding to the statement to be synthesized, wherein the pre-set attention model is a model that uses the feature vector set to generate a context vector used for decoding, and the predicted acoustic features are composed of at least one associated acoustic feature; and performing feature conversion and synthesis on the predicted acoustic features to obtain a speech corresponding to the sentence to be synthesized.


