Speech Synthesis Model Attribute Registration

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing speech synthesis models are limited to synthesizing styles and tones within their training data sets, failing to meet the diverse needs of users for cross-style and cross-tone speech broadcasting, especially for ordinary users who cannot utilize their own styles and tones.

Innovation Solution

A method to register styles and tones using a small amount of user data, fine-tuning a pre-trained speech synthesis model to recognize and utilize user-specific attributes, enabling personalized speech synthesis across various scenarios.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Adaptability or versatility

If a speech synthesis model is trained with a fixed training data set, then the model can achieve stable synthesis quality for styles and tones within the training data, but it cannot synthesize cross-style and cross-tone speech that users need

Engineering Contradiction:
Improvecross-style and cross-tone synthesis capabilityVSAvoidsynthesis quality stability
Core Design Contradiction:
Adaptability or versatilityVSReliability

Solution Approach 1:

The system performs preliminary training with a fixed training data set to establish a stable base model, then prepares for subsequent adaptation by users registering their own styles and tones through attribute registration functionality

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The system transitions from a static pre-trained model to a dynamic adaptable model that allows users to register and customize their own speech attributes, enabling the model to adapt to individual user needs while maintaining core synthesis capabilities

Inventive Principle:
Principle #15Dynamics

2Adaptability or versatility

If training data from multiple speakers in various styles is collected to enable multi-style and multi-tone synthesis, then the speech synthesis model can achieve diverse synthesis capabilities, but the complexity of data collection and model training increases significantly

Engineering Contradiction:
Improvemulti-style and multi-tone synthesis capabilityVSAvoiddata collection and training complexity
Core Design Contradiction:
Adaptability or versatilityVSDevice complexity

Solution Approach 1:

The system performs preliminary training with a fixed training data set to establish a stable base model, then prepares for subsequent adaptation by users registering their own styles and tones through attribute registration functionality

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

Individual users can independently register their own speech attributes by providing their own training data, eliminating the need for the system to collect and process large amounts of diverse data from multiple speakers

Inventive Principle:
Principle #25Self-service

Data Source

PatentUS12062357B2Method of registering attribute in speech synthesis model, apparatus of registering attribute in speech synthesis model, electronic device, and medium
Publication Date: 2024.08.13 BEIJING BAIDU NETCOM SCI & TECH CO LTD
  • US12062357B2 patent drawing
  • US12062357B2 patent drawing
  • US12062357B2 patent drawing

AI summary

A method of registering an attribute in a speech synthesis model, an apparatus of registering an attribute in a speech synthesis model, an electronic device, and a medium are provided, which relate to a field of an artificial intelligence technology such as a deep learning and intelligent speech technology. The method includes: acquiring a plurality of data associated with an attribute to be registered; and registering the attribute in the speech synthesis model by using the plurality of data associated with the attribute, wherein the speech synthesis model is trained in advance by using a training data in a training data set.