Audio Editing Model Synthesizing New Words

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Traditional text-based audio content editing technologies face challenges in synthesizing new words and maintaining natural speech, as they rely on manual copying and pasting, leading to discontinuity in fundamental frequency and unnatural sound, and lack the ability to edit specific words in synthesized speech.

Innovation Solution

A method and apparatus for editing audio using a pre-trained audio editing model that predicts the duration of audio corresponding to modified text, adjusts the audio region, and integrates acoustic features through coarse and fine decoders to produce edited audio that supports the synthesis of new words, ensuring natural context.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Ease of operation

If manual copying and pasting is used for text-based audio content editing, then the editing operation can be performed, but the fundamental frequency becomes discontinuous and the sound becomes unnatural

Engineering Contradiction:
Improvetext-based editing capabilityVSAvoidnaturalness of synthesized speech
Core Design Contradiction:
Ease of operationVSReliability

Solution Approach 1:

The patent replaces manual mechanical operations (copying and pasting) with an automated neural network-based system. The audio editing model automatically processes text modifications and generates corresponding audio edits, eliminating the need for manual frame-by-frame manipulation while maintaining natural speech characteristics through learned patterns from training data.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

Solution Approach 2:

The patent changes the fundamental frequency parameter continuously rather than through discrete manual edits. The neural network model adjusts the fundamental frequency and other acoustic parameters smoothly across the audio timeline, ensuring continuity and natural transitions that manual editing cannot achieve.

Inventive Principle:
Principle #35Parameter changes

2Adaptability or versatility

If traditional text-based editing is used, then existing text can be modified, but new words not in the transcribed text cannot be synthesized

Engineering Contradiction:
Improveflexibility in text modificationVSAvoidability to synthesize new content
Core Design Contradiction:
Adaptability or versatilityVSReliability

Solution Approach 1:

The patent creates a universal audio editing system that handles multiple types of text modifications (additions, deletions, replacements) and can synthesize both existing and new words. The neural network model is trained on diverse data and can generalize to generate audio for any text input, not just modifications to original audio content.

Inventive Principle:
Principle #6Universality (Multi-functionality)

Solution Approach 2:

The patent performs preliminary training of the audio editing model on extensive text-audio pairs before actual editing operations. This pre-training enables the model to learn how to synthesize various words and phrases, so when new words need to be synthesized during editing, the model already has the capability to generate natural-sounding audio for them.

Inventive Principle:
Principle #10Preliminary action

3Device complexity

If audio editing is performed without duration prediction, then the editing process is simpler, but the alignment between modified text and audio region is inaccurate

Engineering Contradiction:
Improveediting process complexityVSAvoidalignment precision between text and audio
Core Design Contradiction:
Device complexityVSManufacturing precision

Solution Approach 1:

The patent performs duration prediction as a preliminary step before actual audio editing. The model first predicts how long the modified text will take to speak, then uses this information to accurately identify and isolate the relevant audio region for editing, ensuring precise alignment between text and audio without adding significant complexity to the overall process.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The patent uses duration prediction as a feedback mechanism to adjust the audio editing process. The predicted duration informs the editing model about the expected timing of the modified text, allowing it to make accurate adjustments to the audio region and ensure proper synchronization between the edited audio and the modified text.

Inventive Principle:
Principle #23Feedback

Data Source

PatentUS11462207B1Method and apparatus for editing audio, electronic device and storage medium
Publication Date: 2022.10.04 INST OF AUTOMATION CHINESE ACAD OF SCI
  • US11462207B1 patent drawing
  • US11462207B1 patent drawing
  • US11462207B1 patent drawing

AI summary

Disclosed are a method and an apparatus for editing audio, an electronic device and a storage medium. The method includes: acquiring a modified text obtained by modifying a known original text of an audio to be edited according to a known text for modification; predicting a duration of an audio corresponding to the text for modification; adjusting a region to be edited of the audio to be edited according to the duration of the audio corresponding to the text for modification, to obtain an adjusted audio to be edited; obtaining, based on a pre-trained audio editing model, an edited audio according to the adjusted audio to be edited and the modified text. In the present disclosure, the edited audio obtained by the audio editing model sounds natural in the context, and supports the function of synthesizing new words that do not appear in the corpus.