Voice Conversion Model Training Using Synthesized Timbre Pairs
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing voice conversion models face challenges in accurately converting voice timbre due to high recording costs and limited data, leading to inaccurate conversions between different speakers.
Innovation Solution
A method involving the use of synthesized voice samples to create a voice feature set, fusing initial voice features with matching content to obtain target voice features, and adjusting model parameters until convergence, enabling accurate voice conversion between different timbres.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If real voices of multiple speakers are directly acquired for training, then the voice conversion model can learn timbre conversion, but the recording costs are high and the data volume is limited
Solution Approach 1:
The patent uses voice synthesis technology to generate synthetic voice samples that copy the characteristics of real voices without requiring actual recording. A pre-trained voice synthesis model generates target speaker voices from source speaker audio, creating training data that mimics real voice pairs. This copying approach solves the contradiction by providing unlimited synthetic training data without incurring recording costs, while maintaining the timbre conversion learning capability.
Solution Approach 2:
The patent performs preliminary voice synthesis to generate training data before model training. By pre-generating synthetic voice samples using a voice synthesis model, the system prepares sufficient training data in advance, eliminating the need for expensive real-world recording sessions. This preliminary action ensures adequate data volume is available for training the voice conversion model.
2Reliability
If real voices are directly acquired, then training data can be obtained, but the recording costs are high
Solution Approach 1:
The patent replaces expensive real voice recording with synthetic voice generation. The voice synthesis model creates reliable training samples by copying voice characteristics from real speakers, maintaining training data quality while eliminating recording costs. The synthetic voices preserve the essential timbre and acoustic properties needed for effective model training.
Solution Approach 2:
The system uses automated voice synthesis to generate its own training data without external recording resources. The voice conversion model training process is self-sufficient, generating all necessary training pairs through synthesis rather than requiring external recording sessions, thereby eliminating the loss of energy in the form of recording costs.
3Loss of energy
If limited real voice data is used, then recording costs are reduced, but the model cannot learn conversion between different timbres accurately
Solution Approach 1:
The patent uses synthetic voice copying to overcome the limitation of limited real data. By generating numerous synthetic voice pairs through the voice synthesis model, the system achieves abundant training data that enables accurate timbre conversion learning without requiring extensive real voice recordings, thus maintaining low recording costs while improving conversion accuracy.
Solution Approach 2:
The patent changes the parameter of data source from real recorded voices to synthesized voices. This parameter change transforms the training data generation process, allowing the model to learn from synthetic samples with varied timbre characteristics generated through parameter manipulation in the synthesis model, thereby achieving accurate timbre conversion without limited real data.
Data Source
AI summary
A method for training a voice conversion model includes acquiring a voice sample set that includes training voice samples corresponding to training object identifiers, a training voice sample corresponding to a piece of voice content, and at least one training voice sample being a synthesized voice sample; extracting voice features of the training voice samples to obtain initial voice features; fusing initial voice features that belong to different training object identifiers and have corresponding voice content matching each other to obtain target voice features; obtaining a voice feature set; and inputting a voice feature in the voice feature set into an initial voice conversion model to obtain predicted voice data, and adjusting a model parameter of the initial voice conversion model based on a difference between the predicted voice data and target voice data until a first convergence condition is met, to obtain a target voice conversion model.


