Voice Conversion Neural Network Training via Parallel Data Segmentation

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Voice conversion technologies based on signal processing and traditional machine learning methods produce unnatural and non-fluent voices, while deep learning methods require a large amount of voice data for training, making them inefficient in terms of time and storage.

Innovation Solution

A voice conversion training method that forms multiple training data sets using parallel voice data groups, allowing for efficient training of a voice conversion neural network with a reduced amount of personalized voice data, utilizing dynamic time warping and bidirectional LSTM neural networks to improve accuracy and fluency.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Manufacturing precision

If deep learning methods are used for voice conversion, then voice fluency and naturalness are improved, but the amount of training data required increases

Engineering Contradiction:
Improvevoice conversion qualityVSAvoidtraining data volume
Core Design Contradiction:
Manufacturing precisionVSQuantity of substance

Solution Approach 1:

The patent applies preliminary action by pre-training the neural network on large-scale parallel voice data groups to establish a generalized probability distribution before fine-tuning with personalized data. This preliminary training phase prepares the model to learn effective voice conversion patterns without requiring extensive personalized training data, thereby reducing the overall data requirement while maintaining high conversion quality

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The training process is segmented into multiple stages: first training on large parallel voice data groups to learn general conversion patterns, then fine-tuning with smaller personalized data sets. This segmentation allows the model to acquire fundamental voice conversion capabilities from abundant parallel data while requiring minimal personalized data for adaptation, resolving the contradiction between data volume and conversion quality

Inventive Principle:
Principle #1Segmentation

2Measurement precision

If large amounts of voice data are collected for training, then conversion accuracy is improved, but storage space and time consumption increase

Engineering Contradiction:
Improveconversion accuracyVSAvoidtraining time
Core Design Contradiction:
Measurement precisionVSLoss of time

Solution Approach 1:

The system performs preliminary training on comprehensive parallel voice data to establish a robust baseline model before fine-tuning with personalized data. This preliminary action ensures the model already possesses strong voice conversion capabilities, reducing the time needed for subsequent personalized adaptation while maintaining high accuracy

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The training process is divided into segments: initial training on large parallel data sets for general accuracy, followed by targeted fine-tuning on smaller personalized data sets. This segmentation achieves high conversion accuracy through the first phase, reducing the time and computational resources required in later phases

Inventive Principle:
Principle #1Segmentation

3Quantity of substance

If traditional machine learning methods are used, then training data requirements are reduced, but voice naturalness and fluency deteriorate

Engineering Contradiction:
Improvetraining data volumeVSAvoidvoice naturalness
Core Design Contradiction:
Quantity of substanceVSManufacturing precision

Solution Approach 1:

The patent introduces parallel voice data groups as an intermediary resource that bridges traditional machine learning efficiency and deep learning quality. These parallel data groups serve as a mediating training corpus that enables the neural network to achieve deep learning-level voice naturalness while using a more manageable data volume compared to conventional deep learning approaches

Inventive Principle:
Principle #24Intermediary (Mediator)

Solution Approach 2:

The training approach combines elements of traditional machine learning (using structured parallel data groups) with deep learning (neural network architecture). This composite methodology leverages the efficiency of traditional methods in data organization while achieving the superior voice naturalness of deep learning, resolving the contradiction between data volume and voice quality

Inventive Principle:
Principle #40Composite materials

Data Source

PatentUS11282503B2Voice conversion training method and server and computer readable storage medium
Publication Date: 2022.03.22 UBTECH ROBOTICS CORP LTD
  • US11282503B2 patent drawing
  • US11282503B2 patent drawing
  • US11282503B2 patent drawing

AI summary

The present disclosure discloses a voice conversion training method. The method includes: forming a first training data set including a plurality of training voice data groups; selecting two of the training voice data groups from the first training data set to input into a voice conversion neural network for training; forming a second training data set including the first training data set and a first source speaker voice data group; inputting one of the training voice data groups selected from the first training data set and the first source speaker voice data group into the network for training; forming the third training data set including the second source speaker voice data group and the personalized voice data group that are parallel corpus with respect to each other; and inputting the second source speaker voice data group and the personalized voice data group into the network for training.