Text-Audio Pair Splicing for Neural Network Training Data
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
The end-to-end speech synthesis method based on neural networks requires high-quality speech data, which is costly and time-consuming to prepare, often resulting in a small amount of sample data with uneven audio lengths, leading to poor synthesis effects.
Innovation Solution
A sample generation method that acquires text-audio pairs, calculates audio features, screens and splices suitable pairs, and detects them to meet preset conditions for writing high-quality sample data into a training database, reducing resource consumption and improving data quantity and quality.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Quantity of substance
If the end-to-end speech synthesis method based on neural networks is used, then the speech synthesis can be achieved with smaller amount of data compared to other methods, but the quality requirement of speech data is greatly increased, leading to increased cost and time of data preparation
Solution Approach 1:
The patent changes the parameters of speech data by applying various processing techniques including noise addition, pitch modification, speed variation, and volume adjustment to existing high-quality speech data, thereby generating diverse training samples that meet the high quality requirements while reducing the need for collecting large amounts of new data
Solution Approach 2:
The patent creates copies of existing high-quality speech data through processing and transformation, generating multiple variations of the same speech content. This allows the system to expand the training dataset without needing to collect equivalent amounts of new high-quality speech data, thus reducing data preparation cost and time
2Manufacturing precision
If manual preparation of high-quality speech data is performed, then the quality of training data is improved, but the cost and time of data preparation are greatly increased
Solution Approach 1:
The patent performs preliminary processing on high-quality speech data by pre-calculating audio features, pre-segmenting speech content, and pre-applying various transformations. This preliminary action prepares the data in advance for the end-to-end model training, reducing the time required during the actual training process while maintaining high data quality
Solution Approach 2:
The patent replaces manual mechanical processing of speech data with automated computational methods. Algorithms automatically perform noise addition, pitch modification, segmentation, and feature extraction, substituting manual labor with efficient computational processes that maintain quality while significantly reducing preparation time
3Quantity of substance
If a small amount of sample data with uneven audio lengths is used, then the data preparation cost is reduced, but the synthesis effect becomes poor
Solution Approach 1:
The patent segments speech data into smaller units such as phonemes or syllables, then recombines them in various ways to create diverse training samples. This segmentation approach allows the system to generate sufficient training data from limited source material, improving the amount and diversity of training data while maintaining audio quality consistency through proper segmentation and recombination techniques
Data Source
AI summary
Provided are a sample generation method and apparatus. The sample generation method comprises: acquiring a plurality of text-audio pairs, wherein each text-audio pair contains a text segment and an audio segment; calculating an audio feature of an audio segment of each of the plurality of text-audio pairs, and selecting, by means of screening and according to the audio feature, a target text-audio pair and a splicing text-audio pair corresponding to the target text-audio pair from among the plurality of text-audio pairs; splicing the target text-audio pair and the splicing text-audio pair into a text-audio pair to be tested, and testing the text-audio pair to be tested; and when the text-audio pair to be tested meets a preset test condition, writing the text-audio pair to be tested into a training database.


