Text-Based Audio Insertion Using Voice Conversion
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current audio editing systems lack efficient methods for interactive text-based insertion and replacement in audio streams, particularly in synthesizing new words or phrases that match the voice quality of existing narration, often resulting in noticeable artifacts due to limited training sets and re-synthesis issues.
Innovation Solution
An optimized voice conversion algorithm using range selection and exchangeable triphones, which allows for dynamic programming and phoneme sequence selection, enabling seamless integration of new audio segments into existing narrations with improved quality and interactive performance.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Productivity
If parametric voice conversion methods are used to synthesize new words or phrases, then the editing operation becomes simple and fast, but the synthesized audio produces noticeable artifacts and has a muffled effect
Solution Approach 1:
The patent segments the voice waveform into phoneme-level units and organizes them into a database with associated features (pitch, energy, phoneme type, context). This segmentation allows the system to select and concatenate pre-recorded phoneme segments to synthesize new words and phrases, avoiding the artifacts of parametric re-synthesis while maintaining high editing speed through efficient database querying and concatenation.
2Manufacturing precision
If unit selection technique is used with large training sets to achieve high quality speech, then the voice conversion quality improves, but the processing time increases and interactivity is reduced
Solution Approach 1:
The patent performs preliminary action by pre-segmenting and storing phoneme units from the training voice in an organized database structure before editing operations are needed. The system pre-extracts features (pitch contours, energy, phoneme identities, contextual information) and indexes them for rapid retrieval. This preliminary organization enables fast, interactive editing operations without requiring real-time processing of large training sets.
Solution Approach 2:
The patent applies local quality by selecting phoneme segments with specific local characteristics (matching pitch contours, energy levels, and contextual features) rather than using global average parameters. This allows the synthesized speech to maintain the natural variations and individuality of the target voice at each phoneme level, achieving high quality conversion with reduced processing requirements.
3Manufacturing precision
If new audio is recorded to replace missing words or phrases, then the voice quality matches the original narration, but the process requires access to original voice talent and recording equipment
Solution Approach 1:
The patent uses copying by extracting and replicating phoneme segments from an existing target voice recording that already exists in the database. Instead of requiring new recording sessions with original voice talent and equipment, the system copies appropriate phoneme segments from the available training data, synthesizing the missing or replaced words using these copied segments. This eliminates the need for additional recording equipment and voice talent while maintaining consistent voice quality.
Data Source
AI summary
Systems and techniques are disclosed for synthesizing a new word or short phrase such that it blends seamlessly in the context of insertion or replacement in an existing narration. In one such embodiment, a text-to-speech synthesizer is utilized to say the word or phrase in a generic voice. Voice conversion is then performed on the generic voice to convert it into a voice that matches the narration. An editor and interface are described that support fully automatic synthesis, selection among a candidate set of alternative pronunciations, fine control over edit placements and pitch profiles, and guidance by the editors own voice.


