ASR Training via TTS Morphing for Speech Pattern Adaptation

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current speech recognition systems face challenges in achieving high accuracy due to variations in human speech patterns, particularly in pitch, duration, and prosody, which existing Automatic Speech Recognition (ASR) engines struggle to adapt to effectively.

Innovation Solution

The proposed solution involves training an optimized ASR engine using a text-to-speech (TTS) engine, where input speech is morphed to match the audio output of the TTS engine, and utilizing speaker-dependent training to fine-tune recognition, combining morphing of speech units with TTS training to create a near-perfect acoustic model for improved recognition.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Adaptability or versatility

If traditional ASR engines are used to recognize varied human speech patterns, then the system can process diverse speech inputs, but recognition accuracy deteriorates due to variations in pitch, duration, and prosody

Engineering Contradiction:
Improveability to process diverse speech inputsVSAvoidrecognition accuracy
Core Design Contradiction:
Adaptability or versatilityVSMeasurement precision

Solution Approach 1:

The patent creates synthetic speech samples by copying and transforming text-to-speech (TTS) audio outputs to generate training data. The ASR engine is trained on these synthesized samples where the ground truth is known, enabling the system to learn optimal speech patterns without requiring extensive manual annotation of diverse real-world speech data.

Inventive Principle:
Principle #26Copying

Solution Approach 2:

The system applies parameter transformations to TTS audio signals, including pitch shifting, duration adjustment, and prosody modification. These parameter changes create varied training samples from a single TTS source, enabling the ASR engine to learn robust speech recognition across different speech characteristics while maintaining high accuracy.

Inventive Principle:
Principle #35Parameter changes

2Measurement precision

If speaker-dependent training is implemented to improve accuracy for specific speakers, then recognition precision improves, but system complexity increases due to additional training requirements

Engineering Contradiction:
Improverecognition precisionVSAvoidtraining system complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The system performs self-training by automatically generating training data from TTS outputs without requiring manual speech collection or annotation. The ASR engine trains on synthetically generated samples where the correct transcription is inherently known, eliminating the need for external linguists or manual data preparation processes.

Inventive Principle:
Principle #25Self-service

Solution Approach 2:

The system performs preliminary training using synthesized speech data before deploying the ASR engine for actual speech recognition. This pre-training phase establishes a strong baseline model that can be further fine-tuned with speaker-specific data if needed, reducing the complexity of on-the-fly adaptation.

Inventive Principle:
Principle #10Preliminary action

3Measurement precision

If more training data is collected to improve recognition accuracy, then transcription precision improves, but time and resource consumption increase

Engineering Contradiction:
Improvetranscription precisionVSAvoidtraining time
Core Design Contradiction:
Measurement precisionVSLoss of time

Solution Approach 1:

Instead of collecting extensive real-world speech data, the system copies and transforms TTS audio outputs to generate large volumes of training samples. This synthetic data generation approach creates unlimited training data from a single TTS source, eliminating the time-consuming process of manual speech collection and annotation while maintaining high transcription precision.

Inventive Principle:
Principle #26Copying

Data Source

PatentUS10068565B2Method and apparatus for an exemplary automatic speech recognition system
Publication Date: 2018.09.04 YASSA FATHY
  • US10068565B2 patent drawing
  • US10068565B2 patent drawing
  • US10068565B2 patent drawing

AI summary

An exemplary computer system configured to train an ASR using the output from a TTS engine.