Modified Transform Coding With Common Time Warp for Audio Frames
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing audio coding methods using time warping introduce discontinuities at frame borders and require significant bit-rate overhead for pitch variation parameters, leading to inefficient coding and audible artifacts.
Innovation Solution
Estimate a common time warp for consecutive audio frames to derive spectral representations, integrating time-warp operations with window functions for block transforms, allowing for efficient coding without discontinuities and reducing the need for pitch parameter transmission.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If simple time warping is applied to achieve constant pitch over voiced segments, then pitch variation is corrected, but absolute tape speed becomes uncontrollable leading to duration violation and bandwidth limitations
Solution Approach 1:
The patent changes the parameter representation from absolute pitch values to relative pitch differences (pitch deviations). By encoding the difference between actual pitch and reference pitch rather than absolute pitch, the system maintains pitch correction capability while controlling the overall duration through reference pitch sequences, thus resolving the contradiction between pitch constancy and duration control.
2Duration of action of stationary object
If consecutive non-overlapping segments are processed independently by time warp, then duration of each segment is preserved, but pitch exhibits jumps at segment boundaries leading to loss of coding efficiency and audible discontinuities
Solution Approach 1:
The patent merges consecutive pitch deviation sequences by identifying and eliminating redundant pitch information at segment boundaries. The pitch deviation of the first sample of a segment is set to match the last sample of the previous segment, creating a continuous pitch trajectory across segment boundaries while maintaining the duration-preserving benefit of independent segment processing.
3Productivity
If large transform sizes are used for stationary signals, then coding gain is maximized, but transform size must be reduced for non-stationary signals causing coding gain to decrease rapidly
Solution Approach 1:
The patent applies dynamic transform size switching based on signal stationarity detection. For stationary segments, large transform sizes are used to maximize coding gain. For non-stationary segments containing transients or pitch changes, the transform size is reduced adaptively. This dynamic adaptation allows the system to maintain high coding gain across varying signal conditions by optimizing transform size locally rather than using a fixed size throughout.
4Measurement precision
If pitch variation parameters are transmitted to correct for frequency changes, then reconstruction accuracy is improved, but substantial bit-rate overhead is introduced especially at low bit-rates
Solution Approach 1:
The patent extracts and transmits only the essential pitch variation information in the form of pitch deviations from a reference pitch sequence, rather than transmitting complete pitch contours or absolute frequency data. This selective extraction of only the necessary correction parameters reduces the amount of data that needs to be transmitted while still providing sufficient information for accurate reconstruction of pitch variations.
Solution Approach 2:
The patent introduces a reference pitch sequence as an intermediary that serves as a common baseline for both encoder and decoder. Instead of transmitting absolute pitch values or complex frequency correction data, the system transmits relative deviations from this reference, using the reference as a mediator that enables accurate reconstruction with minimal transmitted data.
Data Source
Figure 1
Figure 2~2B
Figure 3A~3B
AI summary
A spectral representation of an audio signal having consecutive audio frames can be derived more efficiently, when a common time warp is estimated for any two neighbouring frames, such that a following block transform can additionally use the warp information. Thus, window functions required for successful application of an overlap and add procedure during reconstruction can be derived and applied, the window functions already anticipating the re-sampling of the signal due to the time warping. Therefore, the increased efficiency of block-based transform coding of time-warped signals can be used without introducing audible discontinuities.