Voice Conversion Using CVAE Latent Vectors Without Parallel Data

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Conventional voice conversion methods struggle to effectively convert voice features while reflecting the context and dynamics of the utterance content, often requiring parallel data and time alignment, which can be challenging and inefficient.

Innovation Solution

A voice conversion learning system that uses a Conditional Variational Autoencoder (CVAE) to learn a conversion function by estimating a latent vector series from input sound feature vectors and attribution labels, allowing for the reconfiguration of sound feature vectors to achieve desired voice attributes without needing parallel data, using convolutional or gated convolutional networks to model time-series data.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If conventional voice conversion methods use parallel data and time alignment, then conversion accuracy is improved, but data preparation complexity and time consumption increase

Engineering Contradiction:
Improveconversion accuracyVSAvoiddata preparation complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The patent extracts and removes the requirement for parallel data and manual time alignment from the voice conversion process. By using independent conversion of sound feature vectors without relying on paired data, the method eliminates the complex data preparation stage while maintaining conversion accuracy through the use of sound feature vectors that inherently preserve temporal relationships.

Inventive Principle:
Principle #2Taking out (Extraction)

Solution Approach 2:

The voice conversion system performs self-alignment by processing sound feature vectors that inherently contain temporal and contextual information. The method does not require external time alignment tools or manual intervention, as the conversion process automatically preserves the temporal structure through its design of processing feature vectors in their original sequence without requiring paired target data.

Inventive Principle:
Principle #25Self-service

2Reliability

If voice recognition is used to construct parallel data, then conversion effectiveness is improved, but the need for large voice corpus and high recognition accuracy increases system complexity

Engineering Contradiction:
Improveconversion effectivenessVSAvoidsystem complexity
Core Design Contradiction:
ReliabilityVSDevice complexity

Solution Approach 1:

The patent removes the voice recognition module and large corpus requirement from the system. Instead of using voice recognition to construct parallel data, the method directly processes sound feature vectors from the input voice, extracting only the necessary acoustic information without requiring phoneme-level alignment or recognition accuracy.

Inventive Principle:
Principle #2Taking out (Extraction)

Solution Approach 2:

The patent introduces sound feature vectors as an intermediary representation that bridges the input voice and conversion output without requiring voice recognition as an intermediate step. These feature vectors capture the essential acoustic characteristics needed for conversion while avoiding the complexity of phoneme recognition and alignment.

Inventive Principle:
Principle #24Intermediary (Mediator)

3Ease of operation

If independent conversion of each sound feature amount is performed, then processing simplicity is improved, but ability to reflect voice context and dynamics deteriorates

Engineering Contradiction:
Improveprocessing simplicityVSAvoidcontext reflection ability
Core Design Contradiction:
Ease of operationVSReliability

Solution Approach 1:

The patent merges multiple sound feature vectors into a comprehensive representation that preserves temporal and contextual relationships. By processing a sequence of feature vectors that collectively represent the voice signal, the method maintains contextual information and dynamics while keeping the conversion process relatively simple through vector-based operations.

Inventive Principle:
Principle #5Merging (Combining)

Solution Approach 2:

The patent implements dynamic processing by handling sequences of sound feature vectors that capture temporal variations in the voice signal. The conversion process adapts to the dynamic characteristics of speech by processing feature vectors in their temporal sequence, preserving local time dependence and utterance dynamics without requiring complex alignment procedures.

Inventive Principle:
Principle #15Dynamics

Data Source

PatentUS11450332B2Audio conversion learning device, audio conversion device, method, and program
Publication Date: 2022.09.20 NIPPON TELEGRAPH & TELEPHONE CORP
  • US11450332B2 patent drawing
  • US11450332B2 patent drawing
  • US11450332B2 patent drawing

AI summary

To be able to convert to a voice of the desired attribution. Learning an encoder for, on the basis of parallel data of a sound feature vector series in a conversion-source voice signal and a latent vector series in the conversion-source voice signal, and an attribution label indicating attribution of the conversion-source voice signal, estimating a latent vector series from input of a sound feature vector series and an attribution label, and a decoder for reconfiguring the sound feature vector series from input of the latent vector series and the attribution label.