Voice Conversion Model Training Using Synthesized Timbre Pairs

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing voice conversion models face challenges in accurately converting voice timbre due to high recording costs and limited data, leading to inaccurate conversions between different speakers.

Innovation Solution

A method involving the use of synthesized voice samples to create a voice feature set, fusing initial voice features with matching content to obtain target voice features, and adjusting model parameters until convergence, enabling accurate voice conversion between different timbres.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If real voices of multiple speakers are directly acquired for training, then the voice conversion model can learn timbre conversion, but the recording costs are high and the data volume is limited

Engineering Contradiction:
Improvevoice conversion accuracyVSAvoiddata volume
Core Design Contradiction:
Measurement precisionVSQuantity of substance

Solution Approach 1:

The patent uses voice synthesis technology to generate synthetic voice samples that copy the characteristics of real voices without requiring actual recording. A pre-trained voice synthesis model generates target speaker voices from source speaker audio, creating training data that mimics real voice pairs. This copying approach solves the contradiction by providing unlimited synthetic training data without incurring recording costs, while maintaining the timbre conversion learning capability.

Inventive Principle:
Principle #26Copying

Solution Approach 2:

The patent performs preliminary voice synthesis to generate training data before model training. By pre-generating synthetic voice samples using a voice synthesis model, the system prepares sufficient training data in advance, eliminating the need for expensive real-world recording sessions. This preliminary action ensures adequate data volume is available for training the voice conversion model.

Inventive Principle:
Principle #10Preliminary action

2Reliability

If real voices are directly acquired, then training data can be obtained, but the recording costs are high

Engineering Contradiction:
Improvetraining data qualityVSAvoidrecording cost
Core Design Contradiction:
ReliabilityVSLoss of energy

Solution Approach 1:

The patent replaces expensive real voice recording with synthetic voice generation. The voice synthesis model creates reliable training samples by copying voice characteristics from real speakers, maintaining training data quality while eliminating recording costs. The synthetic voices preserve the essential timbre and acoustic properties needed for effective model training.

Inventive Principle:
Principle #26Copying

Solution Approach 2:

The system uses automated voice synthesis to generate its own training data without external recording resources. The voice conversion model training process is self-sufficient, generating all necessary training pairs through synthesis rather than requiring external recording sessions, thereby eliminating the loss of energy in the form of recording costs.

Inventive Principle:
Principle #25Self-service

3Loss of energy

If limited real voice data is used, then recording costs are reduced, but the model cannot learn conversion between different timbres accurately

Engineering Contradiction:
Improverecording costVSAvoidtimbre conversion accuracy
Core Design Contradiction:
Loss of energyVSMeasurement precision

Solution Approach 1:

The patent uses synthetic voice copying to overcome the limitation of limited real data. By generating numerous synthetic voice pairs through the voice synthesis model, the system achieves abundant training data that enables accurate timbre conversion learning without requiring extensive real voice recordings, thus maintaining low recording costs while improving conversion accuracy.

Inventive Principle:
Principle #26Copying

Solution Approach 2:

The patent changes the parameter of data source from real recorded voices to synthesized voices. This parameter change transforms the training data generation process, allowing the model to learn from synthetic samples with varied timbre characteristics generated through parameter manipulation in the synthesis model, thereby achieving accurate timbre conversion without limited real data.

Inventive Principle:
Principle #35Parameter changes

Data Source

PatentUS12603082B2Method for training voice conversion model, electronic device, and storage medium
Publication Date: 2026.04.14 TENCENT TECHNOLOGY (SHENZHEN) CO LTD
  • US12603082B2 patent drawing
  • US12603082B2 patent drawing
  • US12603082B2 patent drawing

AI summary

A method for training a voice conversion model includes acquiring a voice sample set that includes training voice samples corresponding to training object identifiers, a training voice sample corresponding to a piece of voice content, and at least one training voice sample being a synthesized voice sample; extracting voice features of the training voice samples to obtain initial voice features; fusing initial voice features that belong to different training object identifiers and have corresponding voice content matching each other to obtain target voice features; obtaining a voice feature set; and inputting a voice feature in the voice feature set into an initial voice conversion model to obtain predicted voice data, and adjusting a model parameter of the initial voice conversion model based on a difference between the predicted voice data and target voice data until a first convergence condition is met, to obtain a target voice conversion model.