Text-Conditioned Speech Inpainting for Gaps Over 250 ms

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing speech inpainting technologies struggle to effectively fill gaps in speech samples longer than 250 ms, and require significant parallel training data for each target speaker, making them impractical for many settings.

Innovation Solution

A machine-learned audio inpainting model that uses audio spectrograms and associated textual transcripts to generate replacement audio content for missing portions, maintaining speaker identity, prosody, and recording environment conditions, and generalizing to unseen speakers.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Duration of action of moving object

If existing speech inpainting technologies are used to fill gaps in speech samples, then shorter gaps (up to 250 ms) can be filled, but gaps longer than 250 ms cannot be effectively filled

Engineering Contradiction:
Improvegap durationVSAvoidinpainting effectiveness
Core Design Contradiction:
Duration of action of moving objectVSReliability

Solution Approach 1:

The patent introduces text transcripts as an intermediary modality to bridge the gap between audio input and output. The text-conditioned model uses the transcript as a mediator to guide the inpainting process, enabling the model to generate coherent speech content for gaps longer than 250ms by leveraging textual information about the missing content.

Inventive Principle:
Principle #24Intermediary (Mediator)

Solution Approach 2:

The patent changes the conditioning parameters of the inpainting model by incorporating text embeddings alongside audio features. This parameter change allows the model to access semantic information from the transcript, enabling effective inpainting of longer gaps where traditional audio-only methods fail.

Inventive Principle:
Principle #35Parameter changes

2Adaptability or versatility

If speech inpainting models are trained to handle different speakers, then speaker generalization is achieved, but significant parallel training data for each target speaker is required

Engineering Contradiction:
Improvespeaker generalizationVSAvoidtraining data quantity
Core Design Contradiction:
Adaptability or versatilityVSQuantity of substance

Solution Approach 1:

The patent creates a universal inpainting model that can handle multiple speakers and domains through text conditioning. The model architecture is designed to be speaker-agnostic, using the text transcript as the primary guide rather than speaker-specific audio patterns. This allows a single model to generalize across unseen speakers without requiring extensive parallel training data for each speaker.

Inventive Principle:
Principle #6Universality (Multi-functionality)

Solution Approach 2:

The text transcript serves as a universal intermediary that bridges different speakers and domains. By conditioning on text rather than speaker-specific audio features, the model achieves cross-speaker generalization with minimal training data, as the text provides domain and content information that is speaker-independent.

Inventive Principle:
Principle #24Intermediary (Mediator)

Data Source

PatentUS20250149022A1Text-Conditioned Speech Inpainting
Publication Date: 2025.05.08 GOOGLE LLC
  • US20250149022A1 patent drawing
  • US20250149022A1 patent drawing
  • US20250149022A1 patent drawing

AI summary

Provided are systems, methods, and machine learning models for filling in gaps (e.g., of up to one second) in speech samples by leveraging an auxiliary textual input. Example machine learning models described herein can perform speech inpainting with the appropriate content, while maintaining speaker identity, prosody and recording environment conditions, and generalizing to unseen speakers. This approach significantly outperforms baselines constructed using adaptive TTS, as judged by human raters in side-by-side preference and MOS tests.