Voice Conversion Using Deep Learning Model Segmentation
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current voice conversion methods using deep learning require extensive time and data collection, as they need a conversion source voice to read multiple sentences for training, and there is a demand for avatars to speak with matching voices and for anyone's voice to be converted into various voices.
Innovation Solution
A voice conversion apparatus and method that inputs a conversion destination voice, extracts time-series data including phonemes and pitch from a conversion source voice, adjusts the pitch to match the destination voice, and uses a deep learning model trained on multiple voices to synthesize the desired voice signal, allowing anyone's voice to be converted into various voices without requiring pair data of the source and destination voices.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Manufacturing precision
If deep learning is performed with pair data of conversion source voice and conversion destination voice, then voice conversion quality is improved, but a lot of time is taken for data collection and training
Solution Approach 1:
The patent segments the voice conversion task into two independent parts: (1) extracting phoneme and pitch features from the conversion source voice, and (2) synthesizing the conversion destination voice using a pre-trained deep learning model. This segmentation eliminates the need to collect and train on paired voice data, as the model is trained separately on destination voice characteristics and then applied to convert any source voice.
Solution Approach 2:
The deep learning model is pre-trained in advance on voice data of the conversion destination without requiring conversion source voice data. This preliminary training allows the model to learn the characteristics of the destination voice beforehand, so that during actual conversion, only feature extraction and synthesis are needed, significantly reducing the time required for each conversion task.
2Adaptability or versatility
If pair data of conversion source voice and conversion destination voice is used for training, then direct end-to-end voice conversion is achieved, but the method cannot convert anyone's voice into various voices
Solution Approach 1:
The patent creates a universal voice conversion system where a single deep learning model trained on destination voice data can convert any conversion source voice into the target voice. The model is not tied to specific source-destination pairs but rather learns the characteristics of the destination voice and applies them universally to any input source, greatly enhancing adaptability.
Solution Approach 2:
The patent extracts only the essential features (phoneme and pitch) from the conversion source voice, separating these from the specific characteristics of the source speaker. This extraction allows the conversion to focus on reproducing the destination voice characteristics without being constrained by the source voice identity, enabling conversion of anyone's voice.
3Productivity
If conversion source voice data is required for deep learning training, then voice conversion can be performed, but extensive data collection is necessary
Solution Approach 1:
The patent extracts only the necessary phoneme and pitch features from the conversion source voice, discarding the need to collect and store extensive paired voice data. By taking out only the essential features needed for conversion and relying on pre-trained model knowledge, the system dramatically reduces the quantity of data required while maintaining high conversion efficiency.
Data Source
AI summary
A voice conversion apparatus includes: an input unit that inputs designation of a conversion destination voice; an extraction unit that analyzes a voice signal of a conversion source voice and extracts time series data including a phoneme and a pitch; an adjustment unit that matches a height of the pitch to a height of the designated conversion destination voice; and a generation unit that inputs the phoneme and the pitch to a deep learning model that learns voice data of many people and is capable of synthesizing a designated person's voice in time-series order, and generates a voice signal obtained by synthesizing the designated conversion destination voice.


