Speech Synthesis Model with Phoneme-Level Stress Labels

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Traditional speech synthesis systems fail to accurately incorporate stress in synthesized speech, resulting in flat pronunciation and lack of expressiveness, often randomly selecting stress words which leads to incorrect pronunciations.

Innovation Solution

A method and apparatus for speech synthesis that involves training a speech synthesis model using sample texts marked with stress words and corresponding audios, allowing for the determination of phoneme-level stress labels and generation of audio information with accurate stressed pronunciations, conforming to actual pronunciation habits.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If traditional speech synthesis systems randomly select stress words, then the synthesis process is simple, but the pronunciation accuracy and expressiveness deteriorate

Engineering Contradiction:
Improvestress pronunciation accuracyVSAvoidsynthesis system complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The patent applies preliminary action by pre-training the speech synthesis model with stress word information before actual speech synthesis. The model is trained using sample texts marked with stress words and corresponding sample audios, so that when real text is synthesized, the model already has the capability to accurately identify and pronounce stress words without requiring complex real-time analysis. This resolves the contradiction by preparing the system in advance, achieving high accuracy without increasing operational complexity.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The patent implements self-service by enabling the speech synthesis model to automatically identify and process stress words within the input text without external intervention. The model internally determines which words should be stressed based on the training it has received, eliminating the need for separate stress word selection modules or manual marking. This allows the system to maintain simplicity while achieving accurate stress pronunciation through the model's inherent capabilities.

Inventive Principle:
Principle #25Self-service

2Adaptability or versatility

If speech synthesis systems do not incorporate stress information, then the system complexity is low, but the expressiveness and naturalness of synthesized speech deteriorate

Engineering Contradiction:
Improvespeech expressivenessVSAvoidsynthesis model complexity
Core Design Contradiction:
Adaptability or versatilityVSDevice complexity

Solution Approach 1:

The patent applies parameter changes by modifying the training parameters of the speech synthesis model to include stress word information. Instead of training the model with plain text only, the patent trains it with text that has stress word annotations, thereby changing the model's internal parameters to recognize and reproduce stress patterns. This enables the model to generate expressive speech with proper stress while maintaining a relatively simple architecture, as the expressiveness is achieved through parameter adjustment rather than structural complexity.

Inventive Principle:
Principle #35Parameter changes

3Manufacturing precision

If stress words are marked in sample texts during training, then the synthesized speech accuracy improves, but the data preparation complexity increases

Engineering Contradiction:
Improvesynthesized speech qualityVSAvoiddata preparation ease
Core Design Contradiction:
Manufacturing precisionVSEase of manufacture

Solution Approach 1:

The patent implements self-service by enabling the speech synthesis model to automatically learn stress word patterns from the training data without requiring manual stress word marking during inference. The model internalizes the stress patterns during training and applies them automatically when synthesizing speech from new text inputs. This resolves the contradiction by making the system self-sufficient - the initial data preparation effort pays off by enabling automatic, accurate stress word selection during deployment, eliminating the need for ongoing manual annotation.

Inventive Principle:
Principle #25Self-service

Data Source

PatentUS20230326446A1Method, apparatus, storage medium, and electronic device for speech synthesis
Publication Date: 2023.10.12 BEIJING YOUZHUJU NETWORK TECH CO LTD
  • US20230326446A1 patent drawing
  • US20230326446A1 patent drawing
  • US20230326446A1 patent drawing

AI summary

The present disclosure relates to a method, apparatus, storage medium and electronic device for speech synthesis. The present disclosure enables: acquiring a text to be synthesized marked with stress words; inputting the text to be synthesized into a speech synthesis model to obtain audio information corresponding to the text to be synthesized, the speech synthesis model being obtained by training sample texts marked with stress words and sample audios corresponding to the sample texts, the speech synthesis model being used to process the text to be synthesized in the following manner: determining a sequence of phonemes corresponding to the text to be synthesized; determining phoneme level stress labels according to the stress words marked in the text to be synthesized; generating audio information corresponding to the text to be synthesized according to the sequence of phonemes and the stress labels.