Low-Data Custom Vocal Synthesis With Speaker Embeddings
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing vocal synthesis technologies require extensive voice data for corpus building, leading to low synthesis efficiency and cumbersome timbre customization processes.
Innovation Solution
A method involving training a first neural network with speaker records to create a speaker recognition model, followed by training a second neural network with unaccompanied singing vocal samples to generate a customized timbre vocal, using a small amount of linguistic data to adjust rhythm and pitch.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If extensive voice data is collected to build a corpus, then the vocal synthesis can be performed, but the synthesis efficiency becomes low and the process becomes time-consuming
Solution Approach 1:
The patent extracts only the essential speaker identity information from the full voice data through speaker verification, separating this from the full corpus building process. By taking out only the necessary speaker characteristics rather than using complete voice data, the system achieves synthesis without requiring extensive data collection, thus improving efficiency while maintaining quality.
Solution Approach 2:
The patent performs preliminary speaker verification and feature extraction before the actual synthesis process. By pre-determining speaker identity and characteristics through the verification model, the system prepares all necessary information in advance, eliminating the need for time-consuming corpus building during synthesis and significantly improving productivity.
2Reliability
If a large corpus is used for vocal synthesis, then the synthesis can be performed, but the timbre customization process becomes cumbersome and time-consuming
Solution Approach 1:
The patent extracts speaker-specific characteristics through verification and uses only these extracted features for customization. By taking out only the essential speaker identity and characteristic information rather than working with entire corpora, the system enables easy timbre customization through simple speaker verification inputs, making the process intuitive and efficient.
Solution Approach 2:
The patent creates a simplified representation (copy) of the speaker's voice characteristics through the verification model and speaker embedding. This compressed representation captures the essential timbre features without requiring the full original corpus, allowing for easy customization and manipulation of vocal characteristics while maintaining synthesis quality.
3Reliability
If traditional vocal synthesis methods are used, then voice data can be processed, but the process requires replacing the whole corpus which is cumbersome
Solution Approach 1:
The patent segments the voice processing task into two independent parts: speaker verification (identity determination) and vocal synthesis (sound generation). By segmenting the process this way, the system can handle speaker-specific customization through verification without requiring replacement of the entire corpus, simplifying the overall process while maintaining synthesis quality.
Solution Approach 2:
The patent introduces speaker embedding and verification features as intermediary representations between the raw voice data and the synthesis process. These intermediaries serve as a bridge that translates speaker identity into usable features for synthesis, eliminating the need for direct corpus replacement and simplifying the customization process.
Data Source
AI summary
A custom tone and vocal synthesis method and apparatus, an electronic device, and a storage medium. The synthesis method comprises: training a first neural network by means of a speaker record sample to obtain a speaker recognition model, the output training result of the first neural network being a speaker vector sample (S102); training a second neural network by means of an unaccompanied vocal singing sample and the speaker vector sample to obtain an unaccompanied singing synthesis model (S104); inputting a speaker record to be synthesized into the speaker recognition model to obtain speaker information output by the intermediate hidden layer of the speaker recognition model (S106); and inputting unaccompanied singing music information to be synthesized and the speaker information into the unaccompanied singing synthesis model to obtain a synthesized custom tone and vocal (S108).


