Speech-Text Prompting With Voice-Preserving Word Replacement
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current automatic speech recognition (ASR) models struggle with reduced robustness and accuracy due to reliance on limited and biased training data, particularly from TTS-generated data, which fails to capture the full spectrum of natural human speech variability, leading to performance issues in real-world scenarios.
Innovation Solution
A speech resynthesis process that generates synthetic speech by replacing words in a reference utterance with new words while preserving the voice and speaking style of the reference speaker, using a text-to-speech (TTS) model conditioned on a speaker embedding, to create diverse and representative training data for ASR models.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Quantity of substance
If ASR models are trained primarily on TTS-generated data, then the availability of training data is improved, but the robustness and accuracy of the model deteriorates
Solution Approach 1:
The patent uses TTS models to generate synthetic speech data that copies the structural and acoustic characteristics of real human speech. By creating artificial utterances that replicate natural speech patterns, the system expands training data availability while maintaining quality that reflects real-world speech variability.
Solution Approach 2:
The patent applies parameter changes by modifying acoustic features, speaker characteristics, and speech parameters in the generated data. By varying these parameters to create diverse synthetic utterances, the system improves both data quantity and the representativeness of the training set, addressing the reliability issue.
2Productivity
If current TTS technologies are used to generate synthetic speech data, then the production of training data is improved, but the capture of natural human speech variability deteriorates
Solution Approach 1:
The patent implements dynamics by enabling the TTS system to generate varied and adaptive speech outputs. The model dynamically adjusts acoustic parameters, intonation patterns, and speech characteristics to match the diversity found in natural human speech, thereby improving adaptability while maintaining high productivity in data generation.
3Device complexity
If ASR models are trained on limited training data, then the training process is simplified, but the performance in real-world scenarios deteriorates
Solution Approach 1:
The patent applies self-service by using the TTS system to automatically generate and create training data without requiring manual collection and annotation of real speech data. This self-generating approach simplifies the training process while producing diverse, high-quality data that improves real-world performance.
Data Source
AI summary
A method includes receiving a reference utterance and an input text utterance. The reference utterance includes a plurality of terms spoken by a reference speaker and the input text sequence includes a corresponding transcript for each of the plurality of terms spoken by the reference speaker. The method includes obtaining a speaker embedding characterizing speaker characteristics of the reference speaker that spoke a plurality of terms. The method includes generating a replacement input text sequence by replacing the corresponding transcript of a respective one of the plurality of terms with a replacement transcript corresponding to a different term not included in the reference utterance. The method includes generating, using a text-to-speech (TTS) model conditioned on the reference utterance and the speaker embedding, resynthesized speech based on the replacement input text sequence in a voice of the reference speaker.


