Interleaved Audio Codec Switching for Transient Block Lengths
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing frequency-domain audio coding systems face challenges in achieving satisfactory quality for audio signals with prominent transients, such as rain or applause, at low bitrates, due to issues like time-smearing and inefficient data overhead when using either long or short transform blocks.
Innovation Solution
A frequency-domain audio codec is enhanced to support multiple transform lengths, allowing switching between them, with frequency-domain coefficients transmitted in an interleaved manner to maintain backward compatibility with existing codecs, enabling better quality and efficiency.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If long transform blocks are used for coding audio frames, then spectral resolution is improved, but time-smearing of coding error (pre-echo) occurs frequently for signals with prominent transients
Solution Approach 1:
The audio frame is divided into multiple sub-frames, each processed with a short transform. This segmentation allows the system to achieve good temporal resolution for transient signals while maintaining the ability to represent spectral information, thereby reducing pre-echo artifacts that occur with long transforms.
2Object-generated harmful factors
If short transform blocks are used for coding audio frames, then time-smearing is reduced, but data overhead increases leading to spectral holes
Solution Approach 1:
Multiple short transforms are combined into a unified coding structure where sub-frames are processed together. This merging approach allows efficient packing of transform coefficients, reducing data overhead while maintaining the temporal resolution benefits of short transforms and avoiding spectral holes.
3Measurement precision
If a new frequency-domain audio codec supporting additional transform lengths is developed, then coding quality for specific audio signals is improved, but backward compatibility with existing codecs is lost
Solution Approach 1:
The codec is designed with multi-functionality to handle both traditional long/short transform modes and the new transform length mode. By incorporating universal processing structures that can adapt to different transform types, the system achieves improved coding quality for specific signals while maintaining backward compatibility with existing decoders.
Solution Approach 2:
The codec dynamically adapts its processing mode based on the transform length being used. The system can switch between different operational modes (long transform, short transform, and new transform length) to optimize performance for different signal types while ensuring that legacy decoders can still process the bitstream in traditional modes.
Data Source
Figure 1
Figure 2
Figure 3~4
AI summary
A frequency-domain audio codec is provided with the ability to additionally support a certain transform length in a backward-compatible manner, by the following: the frequency-domain coefficients of a respective frame are transmitted in an interleaved manner irrespective of the signalization signaling for the frames as to which transform length actually applies, and additionally the frequency-domain coefficient extraction and the scale factor extraction operate independent from the signalization. By this measure, old-fashioned frequency-domain audio coders/decoders, insensitive for the signalization, would be able to nevertheless operate without faults and with reproducing a reasonable quality. Concurrently, frequency-domain audio coders/decoders able to support the additional transform length would offer even better quality despite the backward compatibility. As far as coding efficiency penalties due to the coding of the frequency domain coefficients in a manner transparent for older decoders are concerned, same are of comparatively minor nature due to the interleaving.