Text-Based Audio Insertion Using Voice Conversion

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current audio editing systems lack efficient methods for interactive text-based insertion and replacement in audio streams, particularly in synthesizing new words or phrases that match the voice quality of existing narration, often resulting in noticeable artifacts due to limited training sets and re-synthesis issues.

Innovation Solution

An optimized voice conversion algorithm using range selection and exchangeable triphones, which allows for dynamic programming and phoneme sequence selection, enabling seamless integration of new audio segments into existing narrations with improved quality and interactive performance.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Productivity

If parametric voice conversion methods are used to synthesize new words or phrases, then the editing operation becomes simple and fast, but the synthesized audio produces noticeable artifacts and has a muffled effect

Engineering Contradiction:
Improveediting speedVSAvoidaudio quality
Core Design Contradiction:
ProductivityVSManufacturing precision

Solution Approach 1:

The patent segments the voice waveform into phoneme-level units and organizes them into a database with associated features (pitch, energy, phoneme type, context). This segmentation allows the system to select and concatenate pre-recorded phoneme segments to synthesize new words and phrases, avoiding the artifacts of parametric re-synthesis while maintaining high editing speed through efficient database querying and concatenation.

Inventive Principle:
Principle #1Segmentation

2Manufacturing precision

If unit selection technique is used with large training sets to achieve high quality speech, then the voice conversion quality improves, but the processing time increases and interactivity is reduced

Engineering Contradiction:
Improvevoice conversion qualityVSAvoidprocessing time
Core Design Contradiction:
Manufacturing precisionVSLoss of time

Solution Approach 1:

The patent performs preliminary action by pre-segmenting and storing phoneme units from the training voice in an organized database structure before editing operations are needed. The system pre-extracts features (pitch contours, energy, phoneme identities, contextual information) and indexes them for rapid retrieval. This preliminary organization enables fast, interactive editing operations without requiring real-time processing of large training sets.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The patent applies local quality by selecting phoneme segments with specific local characteristics (matching pitch contours, energy levels, and contextual features) rather than using global average parameters. This allows the synthesized speech to maintain the natural variations and individuality of the target voice at each phoneme level, achieving high quality conversion with reduced processing requirements.

Inventive Principle:
Principle #3Local quality

3Manufacturing precision

If new audio is recorded to replace missing words or phrases, then the voice quality matches the original narration, but the process requires access to original voice talent and recording equipment

Engineering Contradiction:
Improvevoice quality matchingVSAvoidrecording equipment requirements
Core Design Contradiction:
Manufacturing precisionVSDevice complexity

Solution Approach 1:

The patent uses copying by extracting and replicating phoneme segments from an existing target voice recording that already exists in the database. Instead of requiring new recording sessions with original voice talent and equipment, the system copies appropriate phoneme segments from the available training data, synthesizing the missing or replaced words using these copied segments. This eliminates the need for additional recording equipment and voice talent while maintaining consistent voice quality.

Inventive Principle:
Principle #26Copying

Data Source

PatentUS10347238B2Text-based insertion and replacement in audio narration
Publication Date: 2019.07.09 ADOBE INC
  • US10347238B2 patent drawing
  • US10347238B2 patent drawing
  • US10347238B2 patent drawing

AI summary

Systems and techniques are disclosed for synthesizing a new word or short phrase such that it blends seamlessly in the context of insertion or replacement in an existing narration. In one such embodiment, a text-to-speech synthesizer is utilized to say the word or phrase in a generic voice. Voice conversion is then performed on the generic voice to convert it into a voice that matches the narration. An editor and interface are described that support fully automatic synthesis, selection among a candidate set of alternative pronunciations, fine control over edit placements and pitch profiles, and guidance by the editors own voice.