Text-Audio Pair Splicing for Neural Network Training Data

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

The end-to-end speech synthesis method based on neural networks requires high-quality speech data, which is costly and time-consuming to prepare, often resulting in a small amount of sample data with uneven audio lengths, leading to poor synthesis effects.

Innovation Solution

A sample generation method that acquires text-audio pairs, calculates audio features, screens and splices suitable pairs, and detects them to meet preset conditions for writing high-quality sample data into a training database, reducing resource consumption and improving data quantity and quality.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Quantity of substance

If the end-to-end speech synthesis method based on neural networks is used, then the speech synthesis can be achieved with smaller amount of data compared to other methods, but the quality requirement of speech data is greatly increased, leading to increased cost and time of data preparation

Engineering Contradiction:
Improveamount of dataVSAvoidquality of speech data
Core Design Contradiction:
Quantity of substanceVSManufacturing precision

Solution Approach 1:

The patent changes the parameters of speech data by applying various processing techniques including noise addition, pitch modification, speed variation, and volume adjustment to existing high-quality speech data, thereby generating diverse training samples that meet the high quality requirements while reducing the need for collecting large amounts of new data

Inventive Principle:
Principle #35Parameter changes

Solution Approach 2:

The patent creates copies of existing high-quality speech data through processing and transformation, generating multiple variations of the same speech content. This allows the system to expand the training dataset without needing to collect equivalent amounts of new high-quality speech data, thus reducing data preparation cost and time

Inventive Principle:
Principle #26Copying

2Manufacturing precision

If manual preparation of high-quality speech data is performed, then the quality of training data is improved, but the cost and time of data preparation are greatly increased

Engineering Contradiction:
Improvequality of training dataVSAvoidtime of data preparation
Core Design Contradiction:
Manufacturing precisionVSLoss of time

Solution Approach 1:

The patent performs preliminary processing on high-quality speech data by pre-calculating audio features, pre-segmenting speech content, and pre-applying various transformations. This preliminary action prepares the data in advance for the end-to-end model training, reducing the time required during the actual training process while maintaining high data quality

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The patent replaces manual mechanical processing of speech data with automated computational methods. Algorithms automatically perform noise addition, pitch modification, segmentation, and feature extraction, substituting manual labor with efficient computational processes that maintain quality while significantly reducing preparation time

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

3Quantity of substance

If a small amount of sample data with uneven audio lengths is used, then the data preparation cost is reduced, but the synthesis effect becomes poor

Engineering Contradiction:
Improveamount of sample dataVSAvoidsynthesis effect
Core Design Contradiction:
Quantity of substanceVSReliability

Solution Approach 1:

The patent segments speech data into smaller units such as phonemes or syllables, then recombines them in various ways to create diverse training samples. This segmentation approach allows the system to generate sufficient training data from limited source material, improving the amount and diversity of training data while maintaining audio quality consistency through proper segmentation and recombination techniques

Inventive Principle:
Principle #1Segmentation

Data Source

PatentUS11810546B2Sample generation method and apparatus
Publication Date: 2023.11.07 BEIJING YUANLI WEILAI SCI & TECH CO LTD
  • US11810546B2 patent drawing
  • US11810546B2 patent drawing
  • US11810546B2 patent drawing

AI summary

Provided are a sample generation method and apparatus. The sample generation method comprises: acquiring a plurality of text-audio pairs, wherein each text-audio pair contains a text segment and an audio segment; calculating an audio feature of an audio segment of each of the plurality of text-audio pairs, and selecting, by means of screening and according to the audio feature, a target text-audio pair and a splicing text-audio pair corresponding to the target text-audio pair from among the plurality of text-audio pairs; splicing the target text-audio pair and the splicing text-audio pair into a text-audio pair to be tested, and testing the text-audio pair to be tested; and when the text-audio pair to be tested meets a preset test condition, writing the text-audio pair to be tested into a training database.