Audio Editing Model Synthesizing New Words
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Traditional text-based audio content editing technologies face challenges in synthesizing new words and maintaining natural speech, as they rely on manual copying and pasting, leading to discontinuity in fundamental frequency and unnatural sound, and lack the ability to edit specific words in synthesized speech.
Innovation Solution
A method and apparatus for editing audio using a pre-trained audio editing model that predicts the duration of audio corresponding to modified text, adjusts the audio region, and integrates acoustic features through coarse and fine decoders to produce edited audio that supports the synthesis of new words, ensuring natural context.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Ease of operation
If manual copying and pasting is used for text-based audio content editing, then the editing operation can be performed, but the fundamental frequency becomes discontinuous and the sound becomes unnatural
Solution Approach 1:
The patent replaces manual mechanical operations (copying and pasting) with an automated neural network-based system. The audio editing model automatically processes text modifications and generates corresponding audio edits, eliminating the need for manual frame-by-frame manipulation while maintaining natural speech characteristics through learned patterns from training data.
Solution Approach 2:
The patent changes the fundamental frequency parameter continuously rather than through discrete manual edits. The neural network model adjusts the fundamental frequency and other acoustic parameters smoothly across the audio timeline, ensuring continuity and natural transitions that manual editing cannot achieve.
2Adaptability or versatility
If traditional text-based editing is used, then existing text can be modified, but new words not in the transcribed text cannot be synthesized
Solution Approach 1:
The patent creates a universal audio editing system that handles multiple types of text modifications (additions, deletions, replacements) and can synthesize both existing and new words. The neural network model is trained on diverse data and can generalize to generate audio for any text input, not just modifications to original audio content.
Solution Approach 2:
The patent performs preliminary training of the audio editing model on extensive text-audio pairs before actual editing operations. This pre-training enables the model to learn how to synthesize various words and phrases, so when new words need to be synthesized during editing, the model already has the capability to generate natural-sounding audio for them.
3Device complexity
If audio editing is performed without duration prediction, then the editing process is simpler, but the alignment between modified text and audio region is inaccurate
Solution Approach 1:
The patent performs duration prediction as a preliminary step before actual audio editing. The model first predicts how long the modified text will take to speak, then uses this information to accurately identify and isolate the relevant audio region for editing, ensuring precise alignment between text and audio without adding significant complexity to the overall process.
Solution Approach 2:
The patent uses duration prediction as a feedback mechanism to adjust the audio editing process. The predicted duration informs the editing model about the expected timing of the modified text, allowing it to make accurate adjustments to the audio region and ensure proper synchronization between the edited audio and the modified text.
Data Source
AI summary
Disclosed are a method and an apparatus for editing audio, an electronic device and a storage medium. The method includes: acquiring a modified text obtained by modifying a known original text of an audio to be edited according to a known text for modification; predicting a duration of an audio corresponding to the text for modification; adjusting a region to be edited of the audio to be edited according to the duration of the audio corresponding to the text for modification, to obtain an adjusted audio to be edited; obtaining, based on a pre-trained audio editing model, an edited audio according to the adjusted audio to be edited and the modified text. In the present disclosure, the edited audio obtained by the audio editing model sounds natural in the context, and supports the function of synthesizing new words that do not appear in the corpus.


