Voice Sample Augmentation via Spectral Class Warping
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing methods for augmenting speech data for acoustic modeling in speech recognition often result in unrealistic voice transformations, leading to poor model performance in distinguishing similar sounds and vowels.
Innovation Solution
The method involves grouping voice samples into spectral classes based on speaker types, determining spectral change ratios from warp distributions, and applying these transformations to generate realistic augmented voice samples that can be used to expand the training dataset.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Quantity of substance
If data augmentation methods are applied to increase the amount of training data, then the quantity of training samples increases, but the quality of voice transformations deteriorates resulting in unrealistic voice samples
Solution Approach 1:
The patent applies parameter changes by transforming voice samples through spectral warping operations that adjust frequency and temporal parameters. Specifically, the system applies spectral changes based on Gaussian distributions of spectral warps for different speaker classes, transforming voice samples from one class to another while maintaining realistic characteristics. This resolves the contradiction by ensuring that the augmentation process modifies parameters in a controlled, distribution-based manner rather than applying arbitrary transformations.
2Adaptability or versatility
If aggressive voice transformation is applied to create diverse voice samples, then the variety of voice characteristics increases, but the realism of the transformed voices deteriorates
Solution Approach 1:
The patent implements feedback mechanisms by using the original voice sample's spectral characteristics and the target speaker class distribution to guide the transformation process. The system calculates spectral warps based on the difference between source and target class distributions, then applies these warps to generate transformed samples. This feedback loop ensures that transformations remain within realistic boundaries while achieving the desired variety, preventing over-transformation that would compromise realism.
3Device complexity
If simple augmentation transformations are applied to expand training data, then the implementation complexity decreases, but the model performance deteriorates due to poor distinction of similar sounds
Solution Approach 1:
The patent applies segmentation by dividing the voice transformation process into distinct stages: (1) grouping voice samples into speaker classes, (2) calculating spectral warps for each class, (3) determining transformation parameters based on class differences, and (4) applying transformations to generate augmented samples. This segmented approach maintains implementation clarity while achieving high model performance by ensuring each transformation step is optimized for preserving acoustic characteristics that distinguish similar sounds.
Data Source
AI summary
A method of augmenting a training dataset of voice samples is provided. An audio processing system obtains voice samples and groups the voice samples into classes of spectral representations. The system obtains warp distributions associated with the classes of spectral representations and determines spectral change ratios based on a comparison of the warp distributions. The system determines transformations based at least in part on the spectral change ratios and applies the transformations to the voice samples grouped into the classes of spectral representations to generate a set of augmented voice samples. The system compiles the training dataset using at least the set of augmented voice samples. A recognition model is trained using the training dataset.


