Deep Generative Models for Synthetic Speech Data Augmentation

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

The limited availability of rich, feature-augmented speech datasets hinders the effective training of deep neural networks for speech recognition and natural language understanding, as existing datasets are costly and of questionable quality, limiting the success of neural networks in these applications.

Innovation Solution

The development of systems, methods, and devices that generate synthetic speech utterances using deep generative models and convolutional neural networks to produce large quantities of high-quality, feature-rich datasets for training neural networks, particularly for Automatic Speech Recognition (ASR) and Natural Language Understanding (NLU).

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Quantity of substance

If synthetic speech data is generated using deep generative models, then the quantity and quality of training data is improved, but the device complexity and computational resources required increase

Engineering Contradiction:
Improvequantity of training dataVSAvoidcomplexity of generative model system
Core Design Contradiction:
Quantity of substanceVSDevice complexity

Solution Approach 1:

The system segments the speech data generation process into distinct components: a generative model for creating synthetic speech, a quality assessment module for evaluating generated samples, and a training data curation system for selecting and preparing final training datasets. This segmentation allows each component to be optimized independently while managing overall system complexity.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The system performs preliminary quality assessment and filtering of synthetic speech data before it is used for training neural networks. By pre-evaluating generated samples using quality metrics and automated assessment, the system ensures that only high-quality synthetic data proceeds to the training stage, improving data quality without requiring complete re-generation.

Inventive Principle:
Principle #10Preliminary action

2Reliability

If deep generative models are used to create feature-rich speech datasets, then the quality and diversity of training data is improved, but the training time and computational cost increase

Engineering Contradiction:
Improvequality of training dataVSAvoidtraining time
Core Design Contradiction:
ReliabilityVSLoss of time

Solution Approach 1:

The system dynamically adjusts generation parameters such as noise levels, speech characteristics, and data augmentation factors based on the specific training needs and quality assessment results. By optimizing these parameters, the system generates high-quality diverse speech data while minimizing unnecessary computational iterations and reducing overall training time.

Inventive Principle:
Principle #35Parameter changes

3Ease of manufacture

If existing speech datasets are used for training, then the ease of obtaining training data is improved, but the quality and feature richness of the data deteriorates

Engineering Contradiction:
Improveease of obtaining training dataVSAvoidquality of speech data
Core Design Contradiction:
Ease of manufactureVSReliability

Solution Approach 1:

The system creates synthetic copies of speech data with enhanced features and controlled characteristics by training a generative model on existing datasets. These synthetic copies preserve the statistical properties and linguistic patterns of real speech while adding diversity and feature richness, effectively copying and improving upon the source data.

Inventive Principle:
Principle #26Copying

Solution Approach 2:

The generative model acts as an intermediary between existing speech datasets and the final training data. It transforms and augments the source data, adding feature richness and diversity while maintaining the essential speech characteristics, thereby mediating between the ease of obtaining existing data and the need for high-quality training data.

Inventive Principle:
Principle #24Intermediary (Mediator)

Data Source

PatentUS10937438B2Neural network generative modeling to transform speech utterances and augment training data
Publication Date: 2021.03.02 FORD GLOBAL TECH LLC
  • US10937438B2 patent drawing
  • US10937438B2 patent drawing
  • US10937438B2 patent drawing

AI summary

Systems, methods, and devices for speech transformation and generating synthetic speech using deep generative models are disclosed. A method of the disclosure includes receiving input audio data comprising a plurality of iterations of a speech utterance from a plurality of speakers. The method includes generating an input spectrogram based on the input audio data and transmitting the input spectrogram to a neural network configured to generate an output spectrogram. The method includes receiving the output spectrogram from the neural network and, based on the output spectrogram, generating synthetic audio data comprising the speech utterance.