Speech-Text Prompting With Voice-Preserving Word Replacement

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current automatic speech recognition (ASR) models struggle with reduced robustness and accuracy due to reliance on limited and biased training data, particularly from TTS-generated data, which fails to capture the full spectrum of natural human speech variability, leading to performance issues in real-world scenarios.

Innovation Solution

A speech resynthesis process that generates synthetic speech by replacing words in a reference utterance with new words while preserving the voice and speaking style of the reference speaker, using a text-to-speech (TTS) model conditioned on a speaker embedding, to create diverse and representative training data for ASR models.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Quantity of substance

If ASR models are trained primarily on TTS-generated data, then the availability of training data is improved, but the robustness and accuracy of the model deteriorates

Engineering Contradiction:
Improvetraining data availabilityVSAvoidmodel robustness and accuracy
Core Design Contradiction:
Quantity of substanceVSReliability

Solution Approach 1:

The patent uses TTS models to generate synthetic speech data that copies the structural and acoustic characteristics of real human speech. By creating artificial utterances that replicate natural speech patterns, the system expands training data availability while maintaining quality that reflects real-world speech variability.

Inventive Principle:
Principle #26Copying

Solution Approach 2:

The patent applies parameter changes by modifying acoustic features, speaker characteristics, and speech parameters in the generated data. By varying these parameters to create diverse synthetic utterances, the system improves both data quantity and the representativeness of the training set, addressing the reliability issue.

Inventive Principle:
Principle #35Parameter changes

2Productivity

If current TTS technologies are used to generate synthetic speech data, then the production of training data is improved, but the capture of natural human speech variability deteriorates

Engineering Contradiction:
Improvetraining data productionVSAvoidspeech variability capture
Core Design Contradiction:
ProductivityVSAdaptability or versatility

Solution Approach 1:

The patent implements dynamics by enabling the TTS system to generate varied and adaptive speech outputs. The model dynamically adjusts acoustic parameters, intonation patterns, and speech characteristics to match the diversity found in natural human speech, thereby improving adaptability while maintaining high productivity in data generation.

Inventive Principle:
Principle #15Dynamics

3Device complexity

If ASR models are trained on limited training data, then the training process is simplified, but the performance in real-world scenarios deteriorates

Engineering Contradiction:
Improvetraining process complexityVSAvoidreal-world performance
Core Design Contradiction:
Device complexityVSReliability

Solution Approach 1:

The patent applies self-service by using the TTS system to automatically generate and create training data without requiring manual collection and annotation of real speech data. This self-generating approach simplifies the training process while producing diverse, high-quality data that improves real-world performance.

Inventive Principle:
Principle #25Self-service

Data Source

PatentUS20250279087A1Speech-text prompting for speech tasks
Publication Date: 2025.09.04 GOOGLE LLC
  • US20250279087A1 patent drawing
  • US20250279087A1 patent drawing
  • US20250279087A1 patent drawing

AI summary

A method includes receiving a reference utterance and an input text utterance. The reference utterance includes a plurality of terms spoken by a reference speaker and the input text sequence includes a corresponding transcript for each of the plurality of terms spoken by the reference speaker. The method includes obtaining a speaker embedding characterizing speaker characteristics of the reference speaker that spoke a plurality of terms. The method includes generating a replacement input text sequence by replacing the corresponding transcript of a respective one of the plurality of terms with a replacement transcript corresponding to a different term not included in the reference utterance. The method includes generating, using a text-to-speech (TTS) model conditioned on the reference utterance and the speaker embedding, resynthesized speech based on the replacement input text sequence in a voice of the reference speaker.