Speech Synthesis Curriculum Learning for Long-Sentence Accuracy
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Conventional speech synthesis systems require extensive speech data and time for training, and struggle with errors when processing long-sentence texts.
Innovation Solution
A speech synthesis system employing curriculum learning to concatenate and train short-sentence texts and speeches, using text tokens and mel spectrogram-tokens to distinguish and connect them, with error rate-based initialization and addition to the model.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If conventional speech synthesis training is performed with extensive speech data, then speech quality is improved, but training time and cost increase significantly
Solution Approach 1:
The patent applies curriculum learning by pre-training the speech synthesis model on short-sentence texts first, then progressively training on longer texts. This preliminary action on simpler data enables the model to learn basic speech synthesis patterns before handling more complex long-sentence inputs, thereby reducing the overall training time and cost while maintaining speech quality
Solution Approach 2:
The patent changes the parameter of text length used for training, progressing from short sentences to long sentences in a structured curriculum. This parameter change approach allows the model to gradually adapt to different text lengths, achieving reliable speech synthesis for long texts without requiring extensive training data and time
2Adaptability or versatility
If speech synthesis model processes long-sentence texts, then text coverage is improved, but error rate increases
Solution Approach 1:
The patent performs preliminary training on short-sentence texts before training on long-sentence texts. This preliminary action establishes a solid foundation for the model, enabling it to process long-sentence texts with improved accuracy and reduced error rates while maintaining high text coverage
Solution Approach 2:
The patent implements a curriculum learning framework where the model progressively learns from texts of increasing length. This feedback mechanism allows the model to adjust its processing capabilities based on performance on shorter texts, thereby improving synthesis accuracy for long-sentence inputs while maintaining versatility
Data Source
AI summary
The present disclosure provides an operating method of a speech synthesis system, which includes, inputting a first text and a first speech for the first text, and a second text and a second speech for the second text; generating a speech synthesis model trained by applying the first and second texts and the first and second speeches to curriculum learning; and outputting a target synthesis speech corresponding to a target text based on the speech synthesis model when inputting the target text for speech output, and the generating of the speech synthesis model includes generating a concatenation text in which the first and second texts are concatenated and a concatenation speech in which the first and second speeches are concatenated, and adding the concatenation text and the concatenation speech to the speech synthesis model when an error rate is smaller than a set reference rate when learning-concatenating the concatenation text and the concatenation speech.


