Synthetic Voice Sample Generation for Model Training

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Voice recognition systems face challenges in handling a wide variety of human voice types due to the difficulty in obtaining diverse and large numbers of training samples, especially with new words and pronunciations, which affects their performance and vocabulary expansion.

Innovation Solution

The development of systems and methods to generate simulated synthetic voice samples that mimic human inputs, including spoken and textual articulations, to train and evaluate language processing models efficiently, allowing for diversified and large quantities of samples to be created, thereby improving model performance across various conditions.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Adaptability or versatility

If real human voice samples are collected to train voice recognition models, then the diversity of voice types and commands is improved, but the cost and complexity of data collection increase significantly

Engineering Contradiction:
Improvevoice type diversityVSAvoiddata collection complexity
Core Design Contradiction:
Adaptability or versatilityVSDevice complexity

Solution Approach 1:

The patent uses text-to-speech synthesis to generate synthetic voice samples that copy the characteristics of real human speech without requiring actual human recording. This allows creation of diverse voice samples across different accents, ages, and speaking styles through algorithmic generation rather than physical collection

Inventive Principle:
Principle #26Copying

Solution Approach 2:

The patent introduces text as an intermediary medium between the desired voice output and the synthesis process. By starting with text commands and using TTS to convert them to speech, the system avoids direct human recording while maintaining natural speech characteristics through the text intermediary

Inventive Principle:
Principle #24Intermediary (Mediator)

2Reliability

If large numbers of diversified voice samples are obtained to improve model performance, then the recognition accuracy is improved, but the time and resources required for sample collection increase

Engineering Contradiction:
Improverecognition accuracyVSAvoidsample collection time
Core Design Contradiction:
ReliabilityVSLoss of time

Solution Approach 1:

The patent performs preliminary voice sample generation through text-to-speech synthesis before actual model training begins. By pre-generating diverse synthetic samples covering various voice types, accents, and commands, the system eliminates the time-consuming process of collecting real samples during the training phase

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The patent changes the fundamental parameter of voice sample generation from physical recording to algorithmic synthesis. By adjusting TTS parameters such as pitch, speed, accent, and speaking style, the system can rapidly generate diverse voice samples without the time constraints of human recording sessions

Inventive Principle:
Principle #35Parameter changes

3Adaptability or versatility

If new words and destinations are added to the system's vocabulary to expand capabilities, then the system versatility is improved, but the difficulty of obtaining accurate pronunciation samples increases

Engineering Contradiction:
Improvevocabulary expansionVSAvoidpronunciation sample acquisition
Core Design Contradiction:
Adaptability or versatilityVSDifficulty of detecting and measuring

Solution Approach 1:

The patent uses text-to-speech synthesis to copy and generate pronunciation samples for new words and proper nouns without requiring native speakers to record them. The TTS system can accurately pronounce newly added vocabulary by converting text to speech based on linguistic rules and phonetic algorithms

Inventive Principle:
Principle #26Copying

Solution Approach 2:

The patent enables the system to self-generate pronunciation samples for new vocabulary entries through automated text-to-speech conversion. Instead of relying on external human resources to record new words, the system serves itself by synthesizing accurate pronunciations from text inputs using built-in TTS capabilities

Inventive Principle:
Principle #25Self-service

Applied Scientific Principles

This section explains which scientific principles are used to turn an abstract innovation direction into a practical engineering solution.

Function Achieved in This Case

This approach enables cost-effective evaluation and retraining of voice recognition models to handle diverse voice types and commands, enhancing their performance in realistic scenarios by providing a comprehensive set of simulated samples that mimic human interactions.

Implementation Method 1

text inputs that are converted to speech waveforms via text-to-speech synthesis

Methodology Applied
Scientific EffectText-to-speech synthesis:

Data Source

PatentUS20240290318A1Training a Voice Recognition Model Using Simulated Voice Samples
Publication Date: 2024.08.29 COMCAST CABLE COMM LLC
  • US20240290318A1 patent drawing
  • US20240290318A1 patent drawing
  • US20240290318A1 patent drawing

AI summary

Systems, apparatuses, and methods are described for generating simulated synthetic voice samples for use in training voice recognition models. The simulated synthetic voice samples may be diversified and in large quantities, in order to train the voice recognition models to handle a large variety of possible voice types and commands. These voice samples may be generated based on simulated user profiles indicating different types of speaker characteristics and words. The generated synthetic voice samples mimic realist human inputs and voice traffic, and may be used to test these voice recognition models for their performance in various situations. Based on the testing, these models may be efficiently retrained to improve their performance in a wide variety of conditions.