Voice Sample Augmentation via Spectral Class Warping

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing methods for augmenting speech data for acoustic modeling in speech recognition often result in unrealistic voice transformations, leading to poor model performance in distinguishing similar sounds and vowels.

Innovation Solution

The method involves grouping voice samples into spectral classes based on speaker types, determining spectral change ratios from warp distributions, and applying these transformations to generate realistic augmented voice samples that can be used to expand the training dataset.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Quantity of substance

If data augmentation methods are applied to increase the amount of training data, then the quantity of training samples increases, but the quality of voice transformations deteriorates resulting in unrealistic voice samples

Engineering Contradiction:
Improvequantity of training samplesVSAvoidquality of voice transformations
Core Design Contradiction:
Quantity of substanceVSManufacturing precision

Solution Approach 1:

The patent applies parameter changes by transforming voice samples through spectral warping operations that adjust frequency and temporal parameters. Specifically, the system applies spectral changes based on Gaussian distributions of spectral warps for different speaker classes, transforming voice samples from one class to another while maintaining realistic characteristics. This resolves the contradiction by ensuring that the augmentation process modifies parameters in a controlled, distribution-based manner rather than applying arbitrary transformations.

Inventive Principle:
Principle #35Parameter changes

2Adaptability or versatility

If aggressive voice transformation is applied to create diverse voice samples, then the variety of voice characteristics increases, but the realism of the transformed voices deteriorates

Engineering Contradiction:
Improvevariety of voice characteristicsVSAvoidrealism of transformed voices
Core Design Contradiction:
Adaptability or versatilityVSReliability

Solution Approach 1:

The patent implements feedback mechanisms by using the original voice sample's spectral characteristics and the target speaker class distribution to guide the transformation process. The system calculates spectral warps based on the difference between source and target class distributions, then applies these warps to generate transformed samples. This feedback loop ensures that transformations remain within realistic boundaries while achieving the desired variety, preventing over-transformation that would compromise realism.

Inventive Principle:
Principle #23Feedback

3Device complexity

If simple augmentation transformations are applied to expand training data, then the implementation complexity decreases, but the model performance deteriorates due to poor distinction of similar sounds

Engineering Contradiction:
Improveimplementation complexityVSAvoidmodel performance in distinguishing similar sounds
Core Design Contradiction:
Device complexityVSMeasurement precision

Solution Approach 1:

The patent applies segmentation by dividing the voice transformation process into distinct stages: (1) grouping voice samples into speaker classes, (2) calculating spectral warps for each class, (3) determining transformation parameters based on class differences, and (4) applying transformations to generate augmented samples. This segmented approach maintains implementation clarity while achieving high model performance by ensuring each transformation step is optimized for preserving acoustic characteristics that distinguish similar sounds.

Inventive Principle:
Principle #1Segmentation

Data Source

PatentUS12340793B1Augmenting voice samples based on distributions of speaker classes
Publication Date: 2025.06.24 INTERACTIONS LLC (US)
  • US12340793B1 patent drawing
  • US12340793B1 patent drawing
  • US12340793B1 patent drawing

AI summary

A method of augmenting a training dataset of voice samples is provided. An audio processing system obtains voice samples and groups the voice samples into classes of spectral representations. The system obtains warp distributions associated with the classes of spectral representations and determines spectral change ratios based on a comparison of the warp distributions. The system determines transformations based at least in part on the spectral change ratios and applies the transformations to the voice samples grouped into the classes of spectral representations to generate a set of augmented voice samples. The system compiles the training dataset using at least the set of augmented voice samples. A recognition model is trained using the training dataset.