Time-Warped Transform Coding for Continuous Pitch Audio Frames

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing audio coding techniques face inefficiencies in encoding signals with pitch variations, leading to audible discontinuities and increased bit rates due to uncontrollable duration and bandwidth limitations in time-warped signals, particularly in handling non-stationary audio signals like voiced speech and musical instruments.

Innovation Solution

A method that estimates a common time warp for consecutive audio frames to enable efficient block transform coding, integrating time-warp operations with window functions and resampling, allowing for continuous time transforms with perfect reconstruction capabilities and reduced bit rate demand by transmitting warp parameters instead of pitch information.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If simple time warping is applied to adjust pitch dynamically, then pitch constancy is achieved, but duration control is lost and bandwidth limitations are violated

Engineering Contradiction:
Improvepitch constancyVSAvoidduration control
Core Design Contradiction:
Measurement precisionVSDuration of action of moving object

Solution Approach 1:

The audio signal is divided into overlapping frames, with each frame processed independently through time-warping operations. This segmentation allows duration control at the frame level while maintaining pitch constancy within each segment, resolving the contradiction between global duration control and local pitch correction.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The time-warping factor is made dynamic and adaptive, adjusting locally for each frame based on pitch detection rather than applying a global transformation. This allows the system to maintain pitch constancy where needed while preserving overall duration through adaptive local transformations.

Inventive Principle:
Principle #15Dynamics

2Measurement precision

If tape speed is adjusted dynamically to achieve constant pitch, then pitch variation is corrected, but additional side information must be transmitted increasing bit rate

Engineering Contradiction:
Improvepitch correctionVSAvoidbit rate overhead
Core Design Contradiction:
Measurement precisionVSLoss of information

Solution Approach 1:

The pitch correction information is extracted and represented compactly as time-warping factors for each frame. Instead of transmitting full pitch sequences or complex correction data, only the essential warping parameters are transmitted, minimizing bit rate overhead while maintaining correction effectiveness.

Inventive Principle:
Principle #2Taking out (Extraction)

Solution Approach 2:

The system changes the parameter representation from detailed pitch sequences to compact time-warping factors. This parameter transformation reduces the amount of side information that must be transmitted while preserving the essential pitch correction functionality.

Inventive Principle:
Principle #35Parameter changes

3Productivity

If large transform sizes are used for stationary signals, then coding gain is maximized, but transient events cause scrambling of frequency analysis

Engineering Contradiction:
Improvecoding gainVSAvoidfrequency analysis accuracy
Core Design Contradiction:
ProductivityVSReliability

Solution Approach 1:

The transform approach is made dynamic by applying time-warping preprocessing that adapts to signal characteristics. For transient events, the warping factor adjusts locally to maintain frequency analysis accuracy, while for stationary segments, larger transforms can be used to maximize coding gain.

Inventive Principle:
Principle #15Dynamics

Solution Approach 2:

Time-warping is applied as a preliminary action before transform coding. This preprocessing step corrects pitch variations and transient effects in advance, ensuring that subsequent transform operations work on pre-conditioned data that maintains frequency analysis accuracy regardless of transform size.

Inventive Principle:
Principle #10Preliminary action

4Reliability

If transform size is reduced for transient events, then frequency analysis accuracy is maintained, but coding gain decreases rapidly

Engineering Contradiction:
Improvefrequency analysis accuracyVSAvoidcoding gain
Core Design Contradiction:
ReliabilityVSProductivity

Solution Approach 1:

Time-warping is applied as a preliminary action that corrects transient effects before transform coding. This allows larger transform sizes to be used even for transient-containing segments, because the warping preprocessing has already aligned the frequency content, thereby maintaining both accuracy and coding gain.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The system changes the approach by introducing time-warping parameters that compensate for transient effects. This allows the transform size parameter to be increased for better coding gain while the warping parameters compensate for any transient-related inaccuracies, effectively decoupling the trade-off between transform size and accuracy.

Inventive Principle:
Principle #35Parameter changes

Data Source

PatentUS8838441B2Time warped modified transform coding of audio signals
Publication Date: 2014.09.16 DOLBY INTERNATIONAL AB
  • US8838441B2 patent drawing
  • US8838441B2 patent drawing
  • US8838441B2 patent drawing

AI summary

A representation of an audio signal having a first, a second and a third frame is derived by estimating first warp information for the first and second frames and second warp information for the second and third frames, the warp information describing pitch information of the audio signal. First or second spectral coefficients for first and second frames or second and third frames are derived using first or second warp information and a first or second weighted representation of the first and second frames or second and third frames, the first or second weighted representation derived by applying a first or second window function to the first and second frames or second and third frames, wherein the first or second window function depends on the first or second warp information. The representation of the audio signal is generated including the first and the second spectral coefficients.