Unsupervised Singing Voice Conversion via Adversarial Networks

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Traditional singing voice conversion systems require parallel data for training and struggle with processing the wide range of frequency variations and sharp changes in volume and pitch, limiting their ability to convert singing voices naturally and on-key without altering the timbre.

Innovation Solution

The use of adversarial neural networks for extracting features and pitch data allows for unsupervised singing voice conversion, enabling the system to learn singer-invariant and pitch-invariant representations, thereby converting the timbre of singing voices without parallel data, by switching the speaker between embeddings.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Reliability

If traditional singing voice conversion systems are used, then parallel data is required for training, but this increases data requirements and system complexity

Engineering Contradiction:
Improveconversion qualityVSAvoiddata requirements
Core Design Contradiction:
ReliabilityVSQuantity of substance

Solution Approach 1:

The patent extracts and removes the requirement for parallel training data by using unsupervised learning approaches. The system separates the voice conversion task into independent feature extraction (using VAE) and pitch manipulation (using adversarial networks), eliminating the need for paired source-target singing data while maintaining conversion quality

Inventive Principle:
Principle #2Taking out (Extraction)

Solution Approach 2:

The system uses self-supervised learning where the model learns from unpaired data by automatically discovering voice characteristics and pitch patterns without external supervision or parallel data. The adversarial networks and VAE perform self-training to achieve reliable voice conversion

Inventive Principle:
Principle #25Self-service

2Reliability

If traditional systems process wide frequency variations and sharp volume/pitch changes, then conversion accuracy decreases, but this limits natural-sounding output

Engineering Contradiction:
Improvepitch accuracyVSAvoidprocessing capability
Core Design Contradiction:
ReliabilityVSEase of operation

Solution Approach 1:

The patent segments the voice conversion process into distinct components: timbre extraction via VAE, pitch detection via adversarial networks, and separate pitch manipulation. This segmentation allows each component to specialize in handling specific aspects like sharp pitch changes or volume variations, improving overall accuracy and naturalness

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The system dynamically adjusts pitch parameters and processing strategies based on the input characteristics. The adversarial networks continuously adapt to handle wide frequency variations and sharp transitions, maintaining high pitch accuracy across diverse singing styles and ranges

Inventive Principle:
Principle #15Dynamics

3Adaptability or versatility

If adversarial neural networks are used for feature extraction, then unsupervised conversion is enabled, but this increases computational complexity

Engineering Contradiction:
Improveconversion flexibilityVSAvoidsystem complexity
Core Design Contradiction:
Adaptability or versatilityVSDevice complexity

Solution Approach 1:

The adversarial neural networks serve multiple functions simultaneously: they extract pitch features, enforce pitch invariance, and enable unsupervised learning. This multi-functionality reduces the need for separate processing modules, managing system complexity while enhancing conversion flexibility and adaptability

Inventive Principle:
Principle #6Universality (Multi-functionality)

4Adaptability or versatility

If parallel data is required for training, then training data availability decreases, but this limits system applicability

Engineering Contradiction:
Improvesystem applicabilityVSAvoidtraining data
Core Design Contradiction:
Adaptability or versatilityVSQuantity of substance

Solution Approach 1:

The patent removes the dependency on parallel training data by extracting only the necessary unpaired data requirements. The system can train with separate source and target voice datasets without requiring aligned pairs, dramatically expanding applicability to real-world scenarios where parallel data is unavailable

Inventive Principle:
Principle #2Taking out (Extraction)

Data Source

PatentUS11257480B2Unsupervised singing voice conversion with pitch adversarial network
Publication Date: 2022.02.22 TENCENT AMERICA LLC
  • US11257480B2 patent drawing
  • US11257480B2 patent drawing
  • US11257480B2 patent drawing

AI summary

A method, a computer readable medium, and a computer system are provided for singing voice conversion. Data corresponding to a singing voice is received. One or more features and pitch data are extracted from the received data using one or more adversarial neural networks. One or more audio samples are generated based on the extracted pitch data and the one or more features.