Speech Synthesis Model Attribute Registration
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing speech synthesis models are limited to synthesizing styles and tones within their training data sets, failing to meet the diverse needs of users for cross-style and cross-tone speech broadcasting, especially for ordinary users who cannot utilize their own styles and tones.
Innovation Solution
A method to register styles and tones using a small amount of user data, fine-tuning a pre-trained speech synthesis model to recognize and utilize user-specific attributes, enabling personalized speech synthesis across various scenarios.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Adaptability or versatility
If a speech synthesis model is trained with a fixed training data set, then the model can achieve stable synthesis quality for styles and tones within the training data, but it cannot synthesize cross-style and cross-tone speech that users need
Solution Approach 1:
The system performs preliminary training with a fixed training data set to establish a stable base model, then prepares for subsequent adaptation by users registering their own styles and tones through attribute registration functionality
Solution Approach 2:
The system transitions from a static pre-trained model to a dynamic adaptable model that allows users to register and customize their own speech attributes, enabling the model to adapt to individual user needs while maintaining core synthesis capabilities
2Adaptability or versatility
If training data from multiple speakers in various styles is collected to enable multi-style and multi-tone synthesis, then the speech synthesis model can achieve diverse synthesis capabilities, but the complexity of data collection and model training increases significantly
Solution Approach 1:
The system performs preliminary training with a fixed training data set to establish a stable base model, then prepares for subsequent adaptation by users registering their own styles and tones through attribute registration functionality
Solution Approach 2:
Individual users can independently register their own speech attributes by providing their own training data, eliminating the need for the system to collect and process large amounts of diverse data from multiple speakers
Data Source
AI summary
A method of registering an attribute in a speech synthesis model, an apparatus of registering an attribute in a speech synthesis model, an electronic device, and a medium are provided, which relate to a field of an artificial intelligence technology such as a deep learning and intelligent speech technology. The method includes: acquiring a plurality of data associated with an attribute to be registered; and registering the attribute in the speech synthesis model by using the plurality of data associated with the attribute, wherein the speech synthesis model is trained in advance by using a training data in a training data set.


