Time-Warped Transform Coding With Overlap-Add Audio Reconstruction

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing audio coding techniques face inefficiencies in encoding signals with pitch variations, leading to audible discontinuities and increased bit rate demands due to uncontrollable duration and bandwidth limitations, especially in time-warped signals.

Innovation Solution

Estimating a common time warp for consecutive audio frames to enable efficient block transform coding, integrating time-warp operations with window functions and resampling, thereby reducing audible discontinuities and bit rate overhead by transmitting warp parameters instead of pitch information.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Stability of the object's composition

If simple time warping is applied to achieve constant pitch over voiced segments, then pitch stability is improved, but duration control deteriorates and bandwidth limitations are violated

Engineering Contradiction:
Improvepitch stabilityVSAvoidduration control
Core Design Contradiction:
Stability of the object's compositionVSDuration of action of moving object

Solution Approach 1:

The patent applies time warping by changing the time parameter of the audio signal to achieve constant pitch over voiced segments. A time warp function is applied to the input signal to transform the time axis, making the pitch stable during voiced portions while allowing flexible duration control through parameter adjustment

Inventive Principle:
Principle #35Parameter changes

2Productivity

If frame-based transform coding is used for stationary signals, then coding efficiency is improved, but handling of non-stationary signals deteriorates

Engineering Contradiction:
Improvecoding efficiencyVSAvoidhandling of non-stationary signals
Core Design Contradiction:
ProductivityVSAdaptability or versatility

Solution Approach 1:

The patent makes the transform adaptive by dynamically adjusting the transform basis functions according to the signal's stationarity characteristics. The system switches between different transform types (DCT for stationary, DFT for transient) based on signal analysis, enabling both high coding efficiency for stationary segments and proper handling of non-stationary portions

Inventive Principle:
Principle #15Dynamics

3Measurement precision

If pitch information is transmitted to control time warping, then reconstruction accuracy is improved, but bit rate overhead increases substantially

Engineering Contradiction:
Improvereconstruction accuracyVSAvoidbit rate overhead
Core Design Contradiction:
Measurement precisionVSQuantity of substance

Solution Approach 1:

The patent extracts only the essential time warp parameters needed for reconstruction rather than transmitting full pitch information. By taking out only the critical timing transformation parameters and discarding redundant pitch data, the system achieves accurate reconstruction with minimal bit rate overhead

Inventive Principle:
Principle #2Taking out (Extraction)

4Productivity

If large transform sizes are used for stationary signals, then coding gain is improved, but resolution of transient events deteriorates

Engineering Contradiction:
Improvecoding gainVSAvoidtransient event resolution
Core Design Contradiction:
ProductivityVSManufacturing precision

Solution Approach 1:

The patent segments the audio signal into stationary and transient portions, applying different transform sizes to each segment. Large transform sizes are used for stationary segments to maximize coding gain, while small transform sizes are used for transient segments to maintain temporal resolution, with smooth transitions between segments

Inventive Principle:
Principle #1Segmentation

Data Source

PatentEP4290513B1Time warped modified transform coding of audio signals
Publication Date: 2024.12.25 DOLBY INTERNATIONAL AB
  • EP4290513B1 patent drawingFigure 1
  • EP4290513B1 patent drawingFigure 2~2B
  • EP4290513B1 patent drawingFigure 3A~3B

AI summary

A spectral representation of an audio signal having consecutive audio frames can be derived more efficiently, when a common time warp is estimated for any two neighbouring frames, such that a following block transform can additionally use the warp information. Thus, window functions required for successful application of an overlap and add procedure during reconstruction can be derived and applied, the window functions already anticipating the re-sampling of the signal due to the time warping. Therefore, the increased efficiency of block-based transform coding of time-warped signals can be used without introducing audible discontinuities.