Singing Voice Conversion Using Speech Samples and DurIAN
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing singing voice conversion methods require minimal singing voice samples from target speakers, limiting their applications, and often rely on parallel data for training, which can be restrictive.
Innovation Solution
The proposed method uses a Duration Informed Attention Network (DurIAN) to convert a singing voice of a first person to a singing voice of a second person by encoding phonemes, aligning them with target acoustic frames, and recursively generating mel-spectrogram features based on the speaking voice of the second person.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Adaptability or versatility
If existing singing voice conversion methods are used, then singing voice conversion can be achieved, but minimal singing voice samples from target speakers are required which limits applicability
Solution Approach 1:
The patent uses speaking voice samples as a substitute (copy) for singing voice samples. By training the model on speaking voice data from target speakers and transferring this learned representation to singing voice conversion, the system eliminates the requirement for target speaker singing samples while maintaining conversion capability. This copying approach allows the model to generalize from speech to singing domains.
Solution Approach 2:
The patent creates a universal model that can handle both speech and singing voice processing. The DurIAN architecture is designed to work with general speech data while being adaptable to singing synthesis, making the system multi-functional. This universality allows the same model framework to serve multiple purposes: speech-to-speech conversion, speech-to-singing conversion, and singing synthesis from scratch.
2Reliability
If parallel data is used for training, then model training can be performed, but the requirement for parallel data is restrictive
Solution Approach 1:
The patent copies knowledge from speech data to singing synthesis tasks. By pre-training on abundant speech parallel data and then fine-tuning or adapting to singing tasks, the model transfers learned representations without requiring extensive singing-specific parallel data. This copying of knowledge across domains reduces the restrictive requirement for singing parallel training data.
Solution Approach 2:
The patent performs preliminary training on speech data before applying the model to singing conversion. The DurIAN model is first trained on speech synthesis tasks to learn fundamental voice conversion capabilities, then this pre-trained model serves as the foundation for singing voice conversion. This preliminary action on speech data prepares the model for singing tasks without requiring complete retraining on singing parallel data.
Data Source
AI summary
A method, computer program, and computer system is provided for converting a singing voice of a first person associated with a first speaker to a singing voice of a second person using a speaking voice of the second person associated with a second speaker. A context associated with one or more phonemes corresponding to the singing voice of a first person is encoded, and the one or more phonemes are aligned to one or more target acoustic frames based on the encoded context. One or more mel-spectrogram features are recursively generated from the aligned phonemes, the target acoustic frames, and a sample of the speaking voice of the second person. A sample corresponding to the singing voice of a first person is converted to a sample corresponding to the second singing voice using the generated mel-spectrogram features.


