Speech Synthesis Model with Decoupled Style and Tone Encoders

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current speech synthesis models are limited to single-tone and single-style speech synthesis, failing to effectively handle multi-style and multi-tone synthesis, which restricts their application in diverse speech interaction scenarios and user experience.

Innovation Solution

A method and apparatus for synthesizing speech that acquires style and tone information, along with content, to generate acoustic features using a pre-trained speech synthesis model, enabling cross-language, cross-style, and cross-tone speech synthesis by decoupling content, style, and tone encoders within the Tacotron structure and utilizing a neural vocoder for audio generation.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Adaptability or versatility

If a speech synthesis model is trained to perform single-tone and single-style synthesis, then the model structure remains simple and training is easier, but the model cannot handle multi-style and multi-tone synthesis requirements

Engineering Contradiction:
Improvemulti-style and multi-tone synthesis capabilityVSAvoidmodel structure complexity
Core Design Contradiction:
Adaptability or versatilityVSDevice complexity

Solution Approach 1:

The patent divides the speech synthesis model into separate encoders for content, style, and tone information. Each encoder processes specific aspects independently, allowing the model to handle multi-style and multi-tone synthesis by combining these separate representations rather than requiring a completely new model architecture.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent creates a universal speech synthesis model that can perform multiple functions: single-tone single-style synthesis, multi-style synthesis, multi-tone synthesis, and cross-language synthesis. The model achieves this universality through the combination of content, style, and tone encoders that can be activated in different combinations for different synthesis tasks.

Inventive Principle:
Principle #6Universality (Multi-functionality)

2Adaptability or versatility

If training data is collected for various styles and speakers to enable multi-style synthesis, then the synthesis versatility improves, but the data collection and processing complexity increases

Engineering Contradiction:
Improvemulti-style synthesis capabilityVSAvoiddata collection and processing difficulty
Core Design Contradiction:
Adaptability or versatilityVSEase of manufacture

Solution Approach 1:

The patent segments the training data into separate categories for content, style, and tone information. By organizing data this way, the model can learn each aspect independently and combine them during synthesis, reducing the complexity of processing and utilizing diverse training data for multiple styles and speakers.

Inventive Principle:
Principle #1Segmentation

3Adaptability or versatility

If the speech synthesis model is enhanced to support cross-language and cross-style synthesis, then the application scope expands, but the computational requirements and processing time increase

Engineering Contradiction:
Improvecross-language and cross-style synthesis capabilityVSAvoidcomputational resource consumption
Core Design Contradiction:
Adaptability or versatilityVSUse of energy by moving object

Solution Approach 1:

The patent segments the computational processing into separate encoder modules that handle content, style, and tone information independently. This segmentation allows for more efficient computation compared to processing all information simultaneously, as each module can operate in parallel and the results are combined at the output stage, reducing overall computational resource consumption.

Inventive Principle:
Principle #1Segmentation

Data Source

PatentUS11769482B2Method and apparatus of synthesizing speech, method and apparatus of training speech synthesis model, electronic device, and storage medium
Publication Date: 2023.09.26 BEIJING BAIDU NETCOM SCI & TECH CO LTD
  • US11769482B2 patent drawing
  • US11769482B2 patent drawing
  • US11769482B2 patent drawing

AI summary

The present disclosure provides a method and apparatus of synthesizing a speech, a method and apparatus of training a speech synthesis model, an electronic device, and a storage medium. The method of synthesizing a speech includes acquiring a style information of a speech to be synthesized, a tone information of the speech to be synthesized, and a content information of a text to be processed; generating an acoustic feature information of the text to be processed, by using a pre-trained speech synthesis model, based on the style information, the tone information, and the content information of the text to be processed; and synthesizing the speech for the text to be processed, based on the acoustic feature information of the text to be processed.