Frequency-Domain Audio Coding With Backward-Compatible Transform Switching
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing frequency-domain audio coding systems face challenges in achieving satisfactory quality for audio signals with prominent transients, such as rain or applause, due to issues with pre-echo and inefficiencies at low bitrates, as they either suffer from time-smearing with long blocks or increased data overhead with short blocks.
Innovation Solution
A frequency-domain audio codec is enhanced to support additional transform lengths in a backward-compatible manner by interleaving frequency-domain coefficients and operating independently of signaling, allowing for switchable spectro-temporal resolutions without affecting compatibility with existing decoders.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If long blocks are used for frequency-domain audio coding, then spectral resolution is improved, but time-smearing of coding error (pre-echo) occurs for transient signals
Solution Approach 1:
The system dynamically switches between long and short transform blocks based on the temporal stationarity of the signal. For transient signals, short blocks are used to reduce pre-echo, while for stationary signals, long blocks are used to maintain spectral resolution. This is achieved through signal analysis that detects transient characteristics and adaptively selects the appropriate block length.
Solution Approach 2:
The audio signal is divided into separate transform blocks that can be independently processed. By segmenting the signal into short blocks for transient portions and long blocks for stationary portions, the system avoids applying long-block transformation to transient signals, thereby reducing pre-echo while maintaining spectral resolution where appropriate.
2Object-generated harmful factors
If short blocks are used for frequency-domain audio coding, then pre-echo is reduced for transient signals, but data overhead increases leading to spectral holes
Solution Approach 1:
The system uses dynamic block switching to apply short blocks only where necessary (during transients) and long blocks where sufficient (during stationary periods). This adaptive approach minimizes the total number of short blocks used, thereby reducing data overhead and preventing spectral holes while still effectively reducing pre-echo during transient segments.
Solution Approach 2:
Different transform block lengths are applied to different portions of the audio signal based on local signal characteristics. Short blocks are applied locally to transient segments where pre-echo occurs, while long blocks are used for stationary segments to maintain spectral continuity and avoid spectral holes. This localized application optimizes both pre-echo reduction and spectral quality.
3Manufacturing precision
If a new frequency-domain audio codec supporting additional transform lengths is developed, then coding quality for transient signals is improved, but backward compatibility with existing decoders is lost
Solution Approach 1:
The codec is designed with multi-functionality to support both traditional long/short block modes and the new transform length modes. The system can operate in multiple modes depending on the capabilities of the decoder, allowing it to maintain backward compatibility with existing decoders while providing enhanced quality when used with newer decoders that support the additional transform lengths.
Solution Approach 2:
The codec dynamically adapts its operation based on decoder capabilities. It can switch between different transform block lengths and coding modes depending on whether the decoder supports the new features. This dynamic adaptation ensures that the codec maintains backward compatibility with legacy decoders while utilizing advanced transform lengths for improved quality when appropriate decoders are available.
4Adaptability or versatility
If frequency-domain coefficients are transmitted in an interleaved manner to support transform length switching, then backward compatibility is maintained, but coding efficiency is reduced
Solution Approach 1:
The system uses dynamic signaling to indicate when interleaved transmission is required versus when standard sequential transmission can be used. By analyzing the signal characteristics and decoder capabilities, the system dynamically selects the most efficient transmission mode, minimizing the use of interleaved transmission to only when necessary for supporting transform length switching, thereby reducing the overall impact on coding efficiency.
Data Source
Figure 1
Figure 2
Figure 3~4
AI summary
A frequency-domain audio codec is provided with the ability to additionally support a certain transform length in a backward-compatible manner, by the following: the frequency-domain coefficients of a respective frame are transmitted in an interleaved manner irrespective of the signalization signaling for the frames as to which transform length actually applies, and additionally the frequency-domain coefficient extraction and the scale factor extraction operate independent from the signalization. By this measure, old-fashioned frequency-domain audio coders/decoders, insensitive for the signalization, would be able to nevertheless operate without faults and with reproducing a reasonable quality. Concurrently, frequency-domain audio coders/decoders able to support the additional transform length would offer even better quality despite the backward compatibility. As far as coding efficiency penalties due to the coding of the frequency domain coefficients in a manner transparent for older decoders are concerned, same are of comparatively minor nature due to the interleaving.