Deep Generative Models for Synthetic Speech Data Augmentation
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
The limited availability of rich, feature-augmented speech datasets hinders the effective training of deep neural networks for speech recognition and natural language understanding, as existing datasets are costly and of questionable quality, limiting the success of neural networks in these applications.
Innovation Solution
The development of systems, methods, and devices that generate synthetic speech utterances using deep generative models and convolutional neural networks to produce large quantities of high-quality, feature-rich datasets for training neural networks, particularly for Automatic Speech Recognition (ASR) and Natural Language Understanding (NLU).
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Quantity of substance
If synthetic speech data is generated using deep generative models, then the quantity and quality of training data is improved, but the device complexity and computational resources required increase
Solution Approach 1:
The system segments the speech data generation process into distinct components: a generative model for creating synthetic speech, a quality assessment module for evaluating generated samples, and a training data curation system for selecting and preparing final training datasets. This segmentation allows each component to be optimized independently while managing overall system complexity.
Solution Approach 2:
The system performs preliminary quality assessment and filtering of synthetic speech data before it is used for training neural networks. By pre-evaluating generated samples using quality metrics and automated assessment, the system ensures that only high-quality synthetic data proceeds to the training stage, improving data quality without requiring complete re-generation.
2Reliability
If deep generative models are used to create feature-rich speech datasets, then the quality and diversity of training data is improved, but the training time and computational cost increase
Solution Approach 1:
The system dynamically adjusts generation parameters such as noise levels, speech characteristics, and data augmentation factors based on the specific training needs and quality assessment results. By optimizing these parameters, the system generates high-quality diverse speech data while minimizing unnecessary computational iterations and reducing overall training time.
3Ease of manufacture
If existing speech datasets are used for training, then the ease of obtaining training data is improved, but the quality and feature richness of the data deteriorates
Solution Approach 1:
The system creates synthetic copies of speech data with enhanced features and controlled characteristics by training a generative model on existing datasets. These synthetic copies preserve the statistical properties and linguistic patterns of real speech while adding diversity and feature richness, effectively copying and improving upon the source data.
Solution Approach 2:
The generative model acts as an intermediary between existing speech datasets and the final training data. It transforms and augments the source data, adding feature richness and diversity while maintaining the essential speech characteristics, thereby mediating between the ease of obtaining existing data and the need for high-quality training data.
Data Source
AI summary
Systems, methods, and devices for speech transformation and generating synthetic speech using deep generative models are disclosed. A method of the disclosure includes receiving input audio data comprising a plurality of iterations of a speech utterance from a plurality of speakers. The method includes generating an input spectrogram based on the input audio data and transmitting the input spectrogram to a neural network configured to generate an output spectrogram. The method includes receiving the output spectrogram from the neural network and, based on the output spectrogram, generating synthetic audio data comprising the speech utterance.


