Decoupled Speech Synthesis Model for Target Timbre Control

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current speech synthesis technologies face challenges in converting text into audio with specific timbre, limiting their ability to provide individualized speech synthesis for applications like video creation and customer service.

Innovation Solution

A speech synthesis model decoupled into two sub-models for feature extraction, where the first sub-model outputs bottleneck features and the second sub-model generates Mel spectrum features, allowing independent control of timbre and other features, enabling the synthesis of audio with target timbre.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Adaptability or versatility

If traditional speech synthesis models are used, then the synthesis process can be completed, but the ability to generate audio with specific timbre is limited

Engineering Contradiction:
Improvetimbre control capabilityVSAvoidsynthesis quality
Core Design Contradiction:
Adaptability or versatilityVSManufacturing precision

Solution Approach 1:

The speech synthesis model is divided into two independent sub-models: a first sub-model that extracts acoustic features (including bottleneck features) from text, and a second sub-model that extracts Mel spectrum features from acoustic features. This segmentation allows independent control and optimization of timbre characteristics through the second sub-model while maintaining overall synthesis quality.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent applies local quality by making the second sub-model specifically responsible for timbre-related Mel spectrum feature extraction, while the first sub-model handles general acoustic feature extraction. This allows different parts of the system to have specialized functions, with the second sub-model optimized locally for timbre control.

Inventive Principle:
Principle #3Local quality

2Manufacturing precision

If more labeled sample audios are used to improve timbre accuracy, then the synthesis quality improves, but the complexity and cost of data preparation increases

Engineering Contradiction:
Improvetimbre accuracyVSAvoiddata preparation complexity
Core Design Contradiction:
Manufacturing precisionVSDevice complexity

Solution Approach 1:

The patent extracts timbre characteristics into a separate second sub-model that independently processes Mel spectrum features. This extraction allows the system to focus computational resources on timbre accuracy without requiring proportional increases in labeled data, as the second sub-model can be trained separately with targeted timbre data.

Inventive Principle:
Principle #2Taking out (Extraction)

Solution Approach 2:

The patent changes the parameter representation by introducing Mel spectrum features as an intermediate representation between acoustic features and final audio output. This parameter transformation enables more efficient control of timbre characteristics and reduces the amount of labeled data needed compared to direct end-to-end training.

Inventive Principle:
Principle #35Parameter changes

3Adaptability or versatility

If a unified speech synthesis model is used, then the model structure is simple, but independent control of timbre and other features is not achieved

Engineering Contradiction:
Improveindependent feature controlVSAvoidmodel structure complexity
Core Design Contradiction:
Adaptability or versatilityVSDevice complexity

Solution Approach 1:

The unified model is segmented into two sub-models that work sequentially, with the first sub-model outputting acoustic features and the second sub-model processing these to generate Mel spectrum features. This segmentation provides independent control interfaces for different feature types while maintaining a relatively simple overall architecture.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The first sub-model serves a universal function by extracting general acoustic features that are then fed into the second sub-model. This multi-functional design allows the first sub-model to handle various text inputs while the second sub-model specializes in timbre control, achieving both versatility and specialized control.

Inventive Principle:
Principle #6Universality (Multi-functionality)

Data Source

PatentUS20240274120A1Speech synthesis method and apparatus, electronic device, and readable storage medium
Publication Date: 2024.08.15 LEMON INC(GB)
  • US20240274120A1 patent drawing
  • US20240274120A1 patent drawing
  • US20240274120A1 patent drawing

AI summary

Provided are an audio synthesis method and apparatus, an electronic device, and a readable storage medium. In the present solution, conversion from a text to an audio having a target timbre is achieved by means of a pre-trained voice synthesis model, the voice synthesis model comprising a first feature extraction sub-model and a second feature extraction sub-model, wherein the first feature extraction sub-model outputs, according to an inputted text to be processed, an acoustic feature comprising a bottleneck feature; the second feature extraction sub-model outputs, according to the inputted first acoustic features, a Mel spectrum feature corresponding to the text to be processed; according to the Mel spectrum feature corresponding to the text to be processed, the target audio corresponding to the text to be processed is obtained, and the target audio has the target timbre.