Singing Voice Conversion Using Speech Samples and DurIAN

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing singing voice conversion methods require minimal singing voice samples from target speakers, limiting their applications, and often rely on parallel data for training, which can be restrictive.

Innovation Solution

The proposed method uses a Duration Informed Attention Network (DurIAN) to convert a singing voice of a first person to a singing voice of a second person by encoding phonemes, aligning them with target acoustic frames, and recursively generating mel-spectrogram features based on the speaking voice of the second person.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Adaptability or versatility

If existing singing voice conversion methods are used, then singing voice conversion can be achieved, but minimal singing voice samples from target speakers are required which limits applicability

Engineering Contradiction:
Improveapplicability of singing voice conversionVSAvoidsinging voice samples from target speakers
Core Design Contradiction:
Adaptability or versatilityVSQuantity of substance

Solution Approach 1:

The patent uses speaking voice samples as a substitute (copy) for singing voice samples. By training the model on speaking voice data from target speakers and transferring this learned representation to singing voice conversion, the system eliminates the requirement for target speaker singing samples while maintaining conversion capability. This copying approach allows the model to generalize from speech to singing domains.

Inventive Principle:
Principle #26Copying

Solution Approach 2:

The patent creates a universal model that can handle both speech and singing voice processing. The DurIAN architecture is designed to work with general speech data while being adaptable to singing synthesis, making the system multi-functional. This universality allows the same model framework to serve multiple purposes: speech-to-speech conversion, speech-to-singing conversion, and singing synthesis from scratch.

Inventive Principle:
Principle #6Universality (Multi-functionality)

2Reliability

If parallel data is used for training, then model training can be performed, but the requirement for parallel data is restrictive

Engineering Contradiction:
Improvemodel training capabilityVSAvoidtraining data flexibility
Core Design Contradiction:
ReliabilityVSAdaptability or versatility

Solution Approach 1:

The patent copies knowledge from speech data to singing synthesis tasks. By pre-training on abundant speech parallel data and then fine-tuning or adapting to singing tasks, the model transfers learned representations without requiring extensive singing-specific parallel data. This copying of knowledge across domains reduces the restrictive requirement for singing parallel training data.

Inventive Principle:
Principle #26Copying

Solution Approach 2:

The patent performs preliminary training on speech data before applying the model to singing conversion. The DurIAN model is first trained on speech synthesis tasks to learn fundamental voice conversion capabilities, then this pre-trained model serves as the foundation for singing voice conversion. This preliminary action on speech data prepares the model for singing tasks without requiring complete retraining on singing parallel data.

Inventive Principle:
Principle #10Preliminary action

Data Source

PatentUS12308019B2Learning singing from speech
Publication Date: 2025.05.20 TENCENT AMERICA LLC
  • US12308019B2 patent drawing
  • US12308019B2 patent drawing
  • US12308019B2 patent drawing

AI summary

A method, computer program, and computer system is provided for converting a singing voice of a first person associated with a first speaker to a singing voice of a second person using a speaking voice of the second person associated with a second speaker. A context associated with one or more phonemes corresponding to the singing voice of a first person is encoded, and the one or more phonemes are aligned to one or more target acoustic frames based on the encoded context. One or more mel-spectrogram features are recursively generated from the aligned phonemes, the target acoustic frames, and a sample of the speaking voice of the second person. A sample corresponding to the singing voice of a first person is converted to a sample corresponding to the second singing voice using the generated mel-spectrogram features.