Voice Conversion Neural Network Training via Parallel Data Segmentation
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Voice conversion technologies based on signal processing and traditional machine learning methods produce unnatural and non-fluent voices, while deep learning methods require a large amount of voice data for training, making them inefficient in terms of time and storage.
Innovation Solution
A voice conversion training method that forms multiple training data sets using parallel voice data groups, allowing for efficient training of a voice conversion neural network with a reduced amount of personalized voice data, utilizing dynamic time warping and bidirectional LSTM neural networks to improve accuracy and fluency.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Manufacturing precision
If deep learning methods are used for voice conversion, then voice fluency and naturalness are improved, but the amount of training data required increases
Solution Approach 1:
The patent applies preliminary action by pre-training the neural network on large-scale parallel voice data groups to establish a generalized probability distribution before fine-tuning with personalized data. This preliminary training phase prepares the model to learn effective voice conversion patterns without requiring extensive personalized training data, thereby reducing the overall data requirement while maintaining high conversion quality
Solution Approach 2:
The training process is segmented into multiple stages: first training on large parallel voice data groups to learn general conversion patterns, then fine-tuning with smaller personalized data sets. This segmentation allows the model to acquire fundamental voice conversion capabilities from abundant parallel data while requiring minimal personalized data for adaptation, resolving the contradiction between data volume and conversion quality
2Measurement precision
If large amounts of voice data are collected for training, then conversion accuracy is improved, but storage space and time consumption increase
Solution Approach 1:
The system performs preliminary training on comprehensive parallel voice data to establish a robust baseline model before fine-tuning with personalized data. This preliminary action ensures the model already possesses strong voice conversion capabilities, reducing the time needed for subsequent personalized adaptation while maintaining high accuracy
Solution Approach 2:
The training process is divided into segments: initial training on large parallel data sets for general accuracy, followed by targeted fine-tuning on smaller personalized data sets. This segmentation achieves high conversion accuracy through the first phase, reducing the time and computational resources required in later phases
3Quantity of substance
If traditional machine learning methods are used, then training data requirements are reduced, but voice naturalness and fluency deteriorate
Solution Approach 1:
The patent introduces parallel voice data groups as an intermediary resource that bridges traditional machine learning efficiency and deep learning quality. These parallel data groups serve as a mediating training corpus that enables the neural network to achieve deep learning-level voice naturalness while using a more manageable data volume compared to conventional deep learning approaches
Solution Approach 2:
The training approach combines elements of traditional machine learning (using structured parallel data groups) with deep learning (neural network architecture). This composite methodology leverages the efficiency of traditional methods in data organization while achieving the superior voice naturalness of deep learning, resolving the contradiction between data volume and voice quality
Data Source
AI summary
The present disclosure discloses a voice conversion training method. The method includes: forming a first training data set including a plurality of training voice data groups; selecting two of the training voice data groups from the first training data set to input into a voice conversion neural network for training; forming a second training data set including the first training data set and a first source speaker voice data group; inputting one of the training voice data groups selected from the first training data set and the first source speaker voice data group into the network for training; forming the third training data set including the second source speaker voice data group and the personalized voice data group that are parallel corpus with respect to each other; and inputting the second source speaker voice data group and the personalized voice data group into the network for training.


