Seamless Audio Melding via Spectrogram Cross-Correlation
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing speech synthesis systems face challenges in seamlessly merging new audio segments with existing ones, often resulting in perceptible transitions due to mismatched tempo and volume, and disregard of energy differences between segments.
Innovation Solution
A method that identifies a sequence of audio items in a playlist, modifies the end portion of the first audio item and the beginning portion of the second audio item by generating spectrograms, identifying cross-correlation time windows, and adjusting amplitudes to ensure a smooth transition, with the goal of matching tempo and volume.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Adaptability or versatility
If pre-recorded audio segments are joined to broaden output phrases, then the range of output phrases is changed or broadened, but the transition between segments is perceptible due to mismatched tempo and volume
Solution Approach 1:
The patent applies preliminary action by pre-processing audio segments to extract tempo and volume characteristics before joining them. The system analyzes each segment's temporal and amplitude properties in advance, then uses this information to adjust and synchronize segments during the concatenation process, ensuring smooth transitions while maintaining the ability to broaden output phrase ranges.
Solution Approach 2:
The patent employs parameter changes by dynamically adjusting tempo and volume parameters of audio segments during the joining process. The system modifies temporal characteristics (speed, duration) and amplitude characteristics (volume, energy) of segments to match their neighbors, resolving the contradiction between versatility and transition smoothness through continuous parameter optimization.
2Ease of manufacture
If new audio segments are added to amend or replace existing segments, then the speech synthesis system is updated, but the tempo and volume of new segments do not match existing segments
Solution Approach 1:
The patent applies feedback by continuously monitoring the tempo and volume characteristics of both existing and new audio segments. The system extracts temporal and amplitude features from existing segments, compares them with new segments, and uses this feedback information to automatically adjust the new segments' parameters, ensuring seamless integration while simplifying the system update process.
Solution Approach 2:
The patent uses preliminary action by pre-analyzing the characteristics of existing audio segments before introducing new ones. The system establishes reference tempo and volume profiles from existing segments in advance, then uses these profiles as templates for adjusting new segments during the update process, maintaining precision while improving ease of manufacture.
3Productivity
If audio segments are concatenated to produce output audio phrases, then speech synthesis output is generated, but energy information differences between segments are disregarded leading to perceptible transitions
Solution Approach 1:
The patent employs parameter changes by extracting and matching energy information parameters (amplitude, power, spectral energy) between adjacent audio segments. The system dynamically adjusts energy-related parameters during concatenation to ensure continuous energy flow across segment boundaries, eliminating perceptible transitions while maintaining high productivity in output audio generation.
Solution Approach 2:
The patent applies preliminary action by pre-extracting energy information characteristics from each audio segment before concatenation. The system calculates temporal and spectral energy profiles in advance, then uses this pre-computed energy information to guide the joining process, ensuring energy continuity is maintained throughout the synthesized output.
Data Source
Figure 1
Figure 2
Figure 3
AI summary
In accordance with some embodiments of the disclosed subject matter, mechanisms for seamless audio melding between audio items in a playlist are provided. In some embodiments, a method for transitioning between audio items in playlists is provided, comprising: identifying a sequence of audio items in a playlist of audio items, wherein the sequence of audio items includes a first audio item and a second audio item that is to be played subsequent to the first audio item; and modifying an end portion of the first audio item and a beginning portion of the second audio item, where the end portion of the first audio item and the beginning portion of the second audio item are to be played concurrently to transition between the first audio item and the second audio item, wherein the end portion of the first audio item and the beginning portion of the second audio item have an overlap duration, and wherein modifying the end portion of the first audio item and the beginning portion of the second audio item comprises: generating a first spectrogram corresponding to the end portion of the first audio item and a second spectrogram corresponding to the beginning portion of the second audio item; identifying, for each frequency band in a series of frequency bands, a window over which the first spectrogram within the end portion of the first audio item and the second spectrogram within the beginning portion of the second audio item have a particular cross-correlation; modifying, for each frequency band in the series of frequency bands, the end portion of the first spectrogram and the beginning portion of the second spectrogram such that amplitudes of frequencies within the frequency band decrease within the first spectrogram over the end portion of the first spectrogram and that amplitudes of frequencies within the frequency band increase within the second spectrogram over the beginning portion of the second spectrogram; and generating a modified version of the first audio item the includes the modified end portion of the first audio item based on the modified end portion of the first spectrogram and generating a modified version of the second audio item that includes the modified beginning portion of the second audio item based on the modified beginning portion of the second spectrogram.