Seamless Audio Melding via Spectrogram Cross-Correlation

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing speech synthesis systems face challenges in seamlessly merging new audio segments with existing ones, often resulting in perceptible transitions due to mismatched tempo and volume, and disregard of energy differences between segments.

Innovation Solution

A method that identifies a sequence of audio items in a playlist, modifies the end portion of the first audio item and the beginning portion of the second audio item by generating spectrograms, identifying cross-correlation time windows, and adjusting amplitudes to ensure a smooth transition, with the goal of matching tempo and volume.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Adaptability or versatility

If pre-recorded audio segments are joined to broaden output phrases, then the range of output phrases is changed or broadened, but the transition between segments is perceptible due to mismatched tempo and volume

Engineering Contradiction:
Improverange of output phrasesVSAvoidtransition smoothness
Core Design Contradiction:
Adaptability or versatilityVSManufacturing precision

Solution Approach 1:

The patent applies preliminary action by pre-processing audio segments to extract tempo and volume characteristics before joining them. The system analyzes each segment's temporal and amplitude properties in advance, then uses this information to adjust and synchronize segments during the concatenation process, ensuring smooth transitions while maintaining the ability to broaden output phrase ranges.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The patent employs parameter changes by dynamically adjusting tempo and volume parameters of audio segments during the joining process. The system modifies temporal characteristics (speed, duration) and amplitude characteristics (volume, energy) of segments to match their neighbors, resolving the contradiction between versatility and transition smoothness through continuous parameter optimization.

Inventive Principle:
Principle #35Parameter changes

2Ease of manufacture

If new audio segments are added to amend or replace existing segments, then the speech synthesis system is updated, but the tempo and volume of new segments do not match existing segments

Engineering Contradiction:
Improvesystem update capabilityVSAvoidtempo and volume matching
Core Design Contradiction:
Ease of manufactureVSManufacturing precision

Solution Approach 1:

The patent applies feedback by continuously monitoring the tempo and volume characteristics of both existing and new audio segments. The system extracts temporal and amplitude features from existing segments, compares them with new segments, and uses this feedback information to automatically adjust the new segments' parameters, ensuring seamless integration while simplifying the system update process.

Inventive Principle:
Principle #23Feedback

Solution Approach 2:

The patent uses preliminary action by pre-analyzing the characteristics of existing audio segments before introducing new ones. The system establishes reference tempo and volume profiles from existing segments in advance, then uses these profiles as templates for adjusting new segments during the update process, maintaining precision while improving ease of manufacture.

Inventive Principle:
Principle #10Preliminary action

3Productivity

If audio segments are concatenated to produce output audio phrases, then speech synthesis output is generated, but energy information differences between segments are disregarded leading to perceptible transitions

Engineering Contradiction:
Improveoutput audio generationVSAvoidenergy information matching
Core Design Contradiction:
ProductivityVSManufacturing precision

Solution Approach 1:

The patent employs parameter changes by extracting and matching energy information parameters (amplitude, power, spectral energy) between adjacent audio segments. The system dynamically adjusts energy-related parameters during concatenation to ensure continuous energy flow across segment boundaries, eliminating perceptible transitions while maintaining high productivity in output audio generation.

Inventive Principle:
Principle #35Parameter changes

Solution Approach 2:

The patent applies preliminary action by pre-extracting energy information characteristics from each audio segment before concatenation. The system calculates temporal and spectral energy profiles in advance, then uses this pre-computed energy information to guide the joining process, ensuring energy continuity is maintained throughout the synthesized output.

Inventive Principle:
Principle #10Preliminary action

Data Source

PatentEP4038610B1Methods, systems, and media for seamless audio melding
Publication Date: 2025.02.12 GOOGLE LLC
  • EP4038610B1 patent drawingFigure 1
  • EP4038610B1 patent drawingFigure 2
  • EP4038610B1 patent drawingFigure 3

AI summary

In accordance with some embodiments of the disclosed subject matter, mechanisms for seamless audio melding between audio items in a playlist are provided. In some embodiments, a method for transitioning between audio items in playlists is provided, comprising: identifying a sequence of audio items in a playlist of audio items, wherein the sequence of audio items includes a first audio item and a second audio item that is to be played subsequent to the first audio item; and modifying an end portion of the first audio item and a beginning portion of the second audio item, where the end portion of the first audio item and the beginning portion of the second audio item are to be played concurrently to transition between the first audio item and the second audio item, wherein the end portion of the first audio item and the beginning portion of the second audio item have an overlap duration, and wherein modifying the end portion of the first audio item and the beginning portion of the second audio item comprises: generating a first spectrogram corresponding to the end portion of the first audio item and a second spectrogram corresponding to the beginning portion of the second audio item; identifying, for each frequency band in a series of frequency bands, a window over which the first spectrogram within the end portion of the first audio item and the second spectrogram within the beginning portion of the second audio item have a particular cross-correlation; modifying, for each frequency band in the series of frequency bands, the end portion of the first spectrogram and the beginning portion of the second spectrogram such that amplitudes of frequencies within the frequency band decrease within the first spectrogram over the end portion of the first spectrogram and that amplitudes of frequencies within the frequency band increase within the second spectrogram over the beginning portion of the second spectrogram; and generating a modified version of the first audio item the includes the modified end portion of the first audio item based on the modified end portion of the first spectrogram and generating a modified version of the second audio item that includes the modified beginning portion of the second audio item based on the modified beginning portion of the second spectrogram.