Voice Conversion Using Deep Learning Model Segmentation

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current voice conversion methods using deep learning require extensive time and data collection, as they need a conversion source voice to read multiple sentences for training, and there is a demand for avatars to speak with matching voices and for anyone's voice to be converted into various voices.

Innovation Solution

A voice conversion apparatus and method that inputs a conversion destination voice, extracts time-series data including phonemes and pitch from a conversion source voice, adjusts the pitch to match the destination voice, and uses a deep learning model trained on multiple voices to synthesize the desired voice signal, allowing anyone's voice to be converted into various voices without requiring pair data of the source and destination voices.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Manufacturing precision

If deep learning is performed with pair data of conversion source voice and conversion destination voice, then voice conversion quality is improved, but a lot of time is taken for data collection and training

Engineering Contradiction:
Improvevoice conversion qualityVSAvoidtime for data collection and training
Core Design Contradiction:
Manufacturing precisionVSLoss of time

Solution Approach 1:

The patent segments the voice conversion task into two independent parts: (1) extracting phoneme and pitch features from the conversion source voice, and (2) synthesizing the conversion destination voice using a pre-trained deep learning model. This segmentation eliminates the need to collect and train on paired voice data, as the model is trained separately on destination voice characteristics and then applied to convert any source voice.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The deep learning model is pre-trained in advance on voice data of the conversion destination without requiring conversion source voice data. This preliminary training allows the model to learn the characteristics of the destination voice beforehand, so that during actual conversion, only feature extraction and synthesis are needed, significantly reducing the time required for each conversion task.

Inventive Principle:
Principle #10Preliminary action

2Adaptability or versatility

If pair data of conversion source voice and conversion destination voice is used for training, then direct end-to-end voice conversion is achieved, but the method cannot convert anyone's voice into various voices

Engineering Contradiction:
Improveability to convert anyone's voice into various voicesVSAvoidcomplexity of requiring pair data collection system
Core Design Contradiction:
Adaptability or versatilityVSDevice complexity

Solution Approach 1:

The patent creates a universal voice conversion system where a single deep learning model trained on destination voice data can convert any conversion source voice into the target voice. The model is not tied to specific source-destination pairs but rather learns the characteristics of the destination voice and applies them universally to any input source, greatly enhancing adaptability.

Inventive Principle:
Principle #6Universality (Multi-functionality)

Solution Approach 2:

The patent extracts only the essential features (phoneme and pitch) from the conversion source voice, separating these from the specific characteristics of the source speaker. This extraction allows the conversion to focus on reproducing the destination voice characteristics without being constrained by the source voice identity, enabling conversion of anyone's voice.

Inventive Principle:
Principle #2Taking out (Extraction)

3Productivity

If conversion source voice data is required for deep learning training, then voice conversion can be performed, but extensive data collection is necessary

Engineering Contradiction:
Improveefficiency of voice conversion processVSAvoidamount of voice data required
Core Design Contradiction:
ProductivityVSQuantity of substance

Solution Approach 1:

The patent extracts only the necessary phoneme and pitch features from the conversion source voice, discarding the need to collect and store extensive paired voice data. By taking out only the essential features needed for conversion and relying on pre-trained model knowledge, the system dramatically reduces the quantity of data required while maintaining high conversion efficiency.

Inventive Principle:
Principle #2Taking out (Extraction)

Data Source

PatentUS20230317090A1Voice conversion device, voice conversion method, program, and recording medium
Publication Date: 2023.10.05 DOWANGO KK
  • US20230317090A1 patent drawing
  • US20230317090A1 patent drawing
  • US20230317090A1 patent drawing

AI summary

A voice conversion apparatus includes: an input unit that inputs designation of a conversion destination voice; an extraction unit that analyzes a voice signal of a conversion source voice and extracts time series data including a phoneme and a pitch; an adjustment unit that matches a height of the pitch to a height of the designated conversion destination voice; and a generation unit that inputs the phoneme and the pitch to a deep learning model that learns voice data of many people and is capable of synthesizing a designated person's voice in time-series order, and generates a voice signal obtained by synthesizing the designated conversion destination voice.