Speech Synthesis Model with Phoneme-Level Stress Labels
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Traditional speech synthesis systems fail to accurately incorporate stress in synthesized speech, resulting in flat pronunciation and lack of expressiveness, often randomly selecting stress words which leads to incorrect pronunciations.
Innovation Solution
A method and apparatus for speech synthesis that involves training a speech synthesis model using sample texts marked with stress words and corresponding audios, allowing for the determination of phoneme-level stress labels and generation of audio information with accurate stressed pronunciations, conforming to actual pronunciation habits.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If traditional speech synthesis systems randomly select stress words, then the synthesis process is simple, but the pronunciation accuracy and expressiveness deteriorate
Solution Approach 1:
The patent applies preliminary action by pre-training the speech synthesis model with stress word information before actual speech synthesis. The model is trained using sample texts marked with stress words and corresponding sample audios, so that when real text is synthesized, the model already has the capability to accurately identify and pronounce stress words without requiring complex real-time analysis. This resolves the contradiction by preparing the system in advance, achieving high accuracy without increasing operational complexity.
Solution Approach 2:
The patent implements self-service by enabling the speech synthesis model to automatically identify and process stress words within the input text without external intervention. The model internally determines which words should be stressed based on the training it has received, eliminating the need for separate stress word selection modules or manual marking. This allows the system to maintain simplicity while achieving accurate stress pronunciation through the model's inherent capabilities.
2Adaptability or versatility
If speech synthesis systems do not incorporate stress information, then the system complexity is low, but the expressiveness and naturalness of synthesized speech deteriorate
Solution Approach 1:
The patent applies parameter changes by modifying the training parameters of the speech synthesis model to include stress word information. Instead of training the model with plain text only, the patent trains it with text that has stress word annotations, thereby changing the model's internal parameters to recognize and reproduce stress patterns. This enables the model to generate expressive speech with proper stress while maintaining a relatively simple architecture, as the expressiveness is achieved through parameter adjustment rather than structural complexity.
3Manufacturing precision
If stress words are marked in sample texts during training, then the synthesized speech accuracy improves, but the data preparation complexity increases
Solution Approach 1:
The patent implements self-service by enabling the speech synthesis model to automatically learn stress word patterns from the training data without requiring manual stress word marking during inference. The model internalizes the stress patterns during training and applies them automatically when synthesizing speech from new text inputs. This resolves the contradiction by making the system self-sufficient - the initial data preparation effort pays off by enabling automatic, accurate stress word selection during deployment, eliminating the need for ongoing manual annotation.
Data Source
AI summary
The present disclosure relates to a method, apparatus, storage medium and electronic device for speech synthesis. The present disclosure enables: acquiring a text to be synthesized marked with stress words; inputting the text to be synthesized into a speech synthesis model to obtain audio information corresponding to the text to be synthesized, the speech synthesis model being obtained by training sample texts marked with stress words and sample audios corresponding to the sample texts, the speech synthesis model being used to process the text to be synthesized in the following manner: determining a sequence of phonemes corresponding to the text to be synthesized; determining phoneme level stress labels according to the stress words marked in the text to be synthesized; generating audio information corresponding to the text to be synthesized according to the sequence of phonemes and the stress labels.


