Low-Data Custom Vocal Synthesis With Speaker Embeddings

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing vocal synthesis technologies require extensive voice data for corpus building, leading to low synthesis efficiency and cumbersome timbre customization processes.

Innovation Solution

A method involving training a first neural network with speaker records to create a speaker recognition model, followed by training a second neural network with unaccompanied singing vocal samples to generate a customized timbre vocal, using a small amount of linguistic data to adjust rhythm and pitch.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Reliability

If extensive voice data is collected to build a corpus, then the vocal synthesis can be performed, but the synthesis efficiency becomes low and the process becomes time-consuming

Engineering Contradiction:
Improvevocal synthesis qualityVSAvoidsynthesis efficiency
Core Design Contradiction:
ReliabilityVSProductivity

Solution Approach 1:

The patent extracts only the essential speaker identity information from the full voice data through speaker verification, separating this from the full corpus building process. By taking out only the necessary speaker characteristics rather than using complete voice data, the system achieves synthesis without requiring extensive data collection, thus improving efficiency while maintaining quality.

Inventive Principle:
Principle #2Taking out (Extraction)

Solution Approach 2:

The patent performs preliminary speaker verification and feature extraction before the actual synthesis process. By pre-determining speaker identity and characteristics through the verification model, the system prepares all necessary information in advance, eliminating the need for time-consuming corpus building during synthesis and significantly improving productivity.

Inventive Principle:
Principle #10Preliminary action

2Reliability

If a large corpus is used for vocal synthesis, then the synthesis can be performed, but the timbre customization process becomes cumbersome and time-consuming

Engineering Contradiction:
Improvevocal synthesis qualityVSAvoidtimbre customization process
Core Design Contradiction:
ReliabilityVSEase of operation

Solution Approach 1:

The patent extracts speaker-specific characteristics through verification and uses only these extracted features for customization. By taking out only the essential speaker identity and characteristic information rather than working with entire corpora, the system enables easy timbre customization through simple speaker verification inputs, making the process intuitive and efficient.

Inventive Principle:
Principle #2Taking out (Extraction)

Solution Approach 2:

The patent creates a simplified representation (copy) of the speaker's voice characteristics through the verification model and speaker embedding. This compressed representation captures the essential timbre features without requiring the full original corpus, allowing for easy customization and manipulation of vocal characteristics while maintaining synthesis quality.

Inventive Principle:
Principle #26Copying

3Reliability

If traditional vocal synthesis methods are used, then voice data can be processed, but the process requires replacing the whole corpus which is cumbersome

Engineering Contradiction:
Improvevocal synthesis qualityVSAvoidcorpus replacement process
Core Design Contradiction:
ReliabilityVSDevice complexity

Solution Approach 1:

The patent segments the voice processing task into two independent parts: speaker verification (identity determination) and vocal synthesis (sound generation). By segmenting the process this way, the system can handle speaker-specific customization through verification without requiring replacement of the entire corpus, simplifying the overall process while maintaining synthesis quality.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent introduces speaker embedding and verification features as intermediary representations between the raw voice data and the synthesis process. These intermediaries serve as a bridge that translates speaker identity into usable features for synthesis, eliminating the need for direct corpus replacement and simplifying the customization process.

Inventive Principle:
Principle #24Intermediary (Mediator)

Data Source

PatentUS12424197B2Custom tone and vocal synthesis method and apparatus, electronic device, and storage medium
Publication Date: 2025.09.23 BEIJING WODONG TIANJUN INFORMATION TECH CO LTD
  • US12424197B2 patent drawing
  • US12424197B2 patent drawing
  • US12424197B2 patent drawing

AI summary

A custom tone and vocal synthesis method and apparatus, an electronic device, and a storage medium. The synthesis method comprises: training a first neural network by means of a speaker record sample to obtain a speaker recognition model, the output training result of the first neural network being a speaker vector sample (S102); training a second neural network by means of an unaccompanied vocal singing sample and the speaker vector sample to obtain an unaccompanied singing synthesis model (S104); inputting a speaker record to be synthesized into the speaker recognition model to obtain speaker information output by the intermediate hidden layer of the speaker recognition model (S106); and inputting unaccompanied singing music information to be synthesized and the speaker information into the unaccompanied singing synthesis model to obtain a synthesized custom tone and vocal (S108).