Frequency-Domain Audio Coding with Transform Length Switching
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing frequency-domain audio coding systems struggle to achieve satisfactory quality at low bitrates for audio signals with prominent transients, such as rain or applause, due to issues like time-smearing and increased data overhead when using either long or short block coding.
Innovation Solution
A frequency-domain audio codec that supports transform length switching, allowing for both long and short transforms within a single frame, with interleaved coefficient transmission to maintain backward compatibility and minimize coding efficiency penalties.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Loss of information
If long block coding is used, then spectral efficiency is improved, but time-smearing and pre-echo artifacts increase
Solution Approach 1:
The audio signal is divided into multiple sub-blocks within a frame, with each sub-block processed by a separate transform. This segmentation allows the system to capture transient events with short transforms while maintaining spectral efficiency through the combined long-transform equivalent, thereby reducing pre-echo artifacts while preserving spectral information.
2Object-generated harmful factors
If short block coding is used, then pre-echo artifacts are reduced, but data overhead increases and spectral holes appear
Solution Approach 1:
Multiple short transforms are combined within a single frame structure, where the frequency-domain coefficients from multiple sub-blocks are merged and processed together. This combining approach recovers spectral efficiency by filling spectral holes while maintaining the transient-capturing advantage of short transforms, thus avoiding the data overhead penalty of purely short-block coding.
3Loss of information
If a new frequency-domain audio codec supporting additional transform lengths is created, then coding quality for specific signal types is improved, but backward compatibility with existing codecs is lost
Solution Approach 1:
The codec is designed to handle multiple transform lengths (both long and short transforms) within a unified framework that maintains compatibility with existing frequency-domain audio codecs. The system can adaptively select appropriate transform lengths based on signal characteristics while preserving the ability to decode streams from legacy codecs, thus achieving both improved coding quality and backward compatibility.
Data Source
AI summary
A frequency-domain audio codec is provided with the ability to additionally support a certain transform length in a backward-compatible manner, by the following: the frequency-domain coefficients of a respective frame are transmitted in an interleaved manner irrespective of the signalization signaling for the frames as to which transform length actually applies, and additionally the frequency-domain coefficient extraction and the scale factor extraction operate independent from the signalization. By this measure, old-fashioned frequency-domain audio coders/decoders, insensitive for the signalization, would be able to nevertheless operate without faults and with reproducing a reasonable quality. Concurrently, frequency-domain audio coders/decoders able to support the additional transform length would offer even better quality despite the backward compatibility. As far as coding efficiency penalties due to the coding of the frequency domain coefficients in a manner transparent for older decoders are concerned, same are of comparatively minor nature due to the interleaving.


