Context-Aware Prosody Correction for Edited Speech Audio
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Conventional speech editors fail to maintain prosody continuity when editing speech audio recordings, resulting in unnatural sound due to prosody mismatches at cut points, despite allowing intuitive cut, copy, and paste operations.
Innovation Solution
A context-aware correction system predicts phoneme durations and pitch contours for edited audio data, applying a target prosody through digital signal processing and removing artifacts to ensure seamless integration with unedited audio.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Ease of operation
If text-based speech editors are used to cut and paste speech audio recordings, then editing ease is improved, but prosody continuity deteriorates
Solution Approach 1:
The system changes prosody parameters (pitch contour, phoneme durations, energy) of the inserted speech segment to match the surrounding unedited audio. This is achieved through predicting target prosody parameters from the unedited audio context and applying them to the edited region, thereby resolving the prosody discontinuity while preserving editing ease
Solution Approach 2:
The system introduces an intermediary processing stage between cutting and pasting speech segments. This intermediary step involves predicting phoneme durations and pitch contours based on the unedited audio context, then applying these predictions as intermediate parameters to ensure smooth prosodic transitions at the edit boundaries
2Adaptability or versatility
If speech audio is edited by cutting and pasting, then editing capability is improved, but audio naturalness deteriorates
Solution Approach 1:
The system modifies key audio parameters including pitch contour, phoneme durations, and energy levels of the edited region to match the natural characteristics of the surrounding unedited audio. This parameter adjustment process eliminates the artificial 'cut and paste' sound while preserving the editing capability
Solution Approach 2:
The system uses the unedited audio as a reference to predict the appropriate prosody parameters for the edited region. This feedback mechanism ensures that the edited speech segment adapts to the acoustic context, maintaining naturalness while allowing flexible editing operations
Data Source
AI summary
Methods are performed by one or more processing devices for correcting prosody in audio data. A method includes operations for accessing subject audio data in an audio edit region of the audio data. The subject audio data in the audio edit region potentially lacks prosodic continuity with unedited audio data in an unedited audio portion of the audio data. The operations further include predicting, based on a context of the unedited audio data, phoneme durations including a respective phoneme duration of each phoneme in the unedited audio data. The operations further include predicting, based on the context of the unedited audio data, a pitch contour comprising at least one respective pitch value of each phoneme in the unedited audio data. Additionally, the operations include correcting prosody of the subject audio data in the audio edit region by applying the phoneme durations and the pitch contour to the subject audio data.


