Voice Conversion Using CVAE Latent Vectors Without Parallel Data
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Conventional voice conversion methods struggle to effectively convert voice features while reflecting the context and dynamics of the utterance content, often requiring parallel data and time alignment, which can be challenging and inefficient.
Innovation Solution
A voice conversion learning system that uses a Conditional Variational Autoencoder (CVAE) to learn a conversion function by estimating a latent vector series from input sound feature vectors and attribution labels, allowing for the reconfiguration of sound feature vectors to achieve desired voice attributes without needing parallel data, using convolutional or gated convolutional networks to model time-series data.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If conventional voice conversion methods use parallel data and time alignment, then conversion accuracy is improved, but data preparation complexity and time consumption increase
Solution Approach 1:
The patent extracts and removes the requirement for parallel data and manual time alignment from the voice conversion process. By using independent conversion of sound feature vectors without relying on paired data, the method eliminates the complex data preparation stage while maintaining conversion accuracy through the use of sound feature vectors that inherently preserve temporal relationships.
Solution Approach 2:
The voice conversion system performs self-alignment by processing sound feature vectors that inherently contain temporal and contextual information. The method does not require external time alignment tools or manual intervention, as the conversion process automatically preserves the temporal structure through its design of processing feature vectors in their original sequence without requiring paired target data.
2Reliability
If voice recognition is used to construct parallel data, then conversion effectiveness is improved, but the need for large voice corpus and high recognition accuracy increases system complexity
Solution Approach 1:
The patent removes the voice recognition module and large corpus requirement from the system. Instead of using voice recognition to construct parallel data, the method directly processes sound feature vectors from the input voice, extracting only the necessary acoustic information without requiring phoneme-level alignment or recognition accuracy.
Solution Approach 2:
The patent introduces sound feature vectors as an intermediary representation that bridges the input voice and conversion output without requiring voice recognition as an intermediate step. These feature vectors capture the essential acoustic characteristics needed for conversion while avoiding the complexity of phoneme recognition and alignment.
3Ease of operation
If independent conversion of each sound feature amount is performed, then processing simplicity is improved, but ability to reflect voice context and dynamics deteriorates
Solution Approach 1:
The patent merges multiple sound feature vectors into a comprehensive representation that preserves temporal and contextual relationships. By processing a sequence of feature vectors that collectively represent the voice signal, the method maintains contextual information and dynamics while keeping the conversion process relatively simple through vector-based operations.
Solution Approach 2:
The patent implements dynamic processing by handling sequences of sound feature vectors that capture temporal variations in the voice signal. The conversion process adapts to the dynamic characteristics of speech by processing feature vectors in their temporal sequence, preserving local time dependence and utterance dynamics without requiring complex alignment procedures.
Data Source
AI summary
To be able to convert to a voice of the desired attribution. Learning an encoder for, on the basis of parallel data of a sound feature vector series in a conversion-source voice signal and a latent vector series in the conversion-source voice signal, and an attribution label indicating attribution of the conversion-source voice signal, estimating a latent vector series from input of a sound feature vector series and an attribution label, and a decoder for reconfiguring the sound feature vector series from input of the latent vector series and the attribution label.


