Interleaved Audio Codec Switching for Transient Block Lengths

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing frequency-domain audio coding systems face challenges in achieving satisfactory quality for audio signals with prominent transients, such as rain or applause, at low bitrates, due to issues like time-smearing and inefficient data overhead when using either long or short transform blocks.

Innovation Solution

A frequency-domain audio codec is enhanced to support multiple transform lengths, allowing switching between them, with frequency-domain coefficients transmitted in an interleaved manner to maintain backward compatibility with existing codecs, enabling better quality and efficiency.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If long transform blocks are used for coding audio frames, then spectral resolution is improved, but time-smearing of coding error (pre-echo) occurs frequently for signals with prominent transients

Engineering Contradiction:
Improvespectral resolutionVSAvoidtime-smearing (pre-echo)
Core Design Contradiction:
Measurement precisionVSObject-generated harmful factors

Solution Approach 1:

The audio frame is divided into multiple sub-frames, each processed with a short transform. This segmentation allows the system to achieve good temporal resolution for transient signals while maintaining the ability to represent spectral information, thereby reducing pre-echo artifacts that occur with long transforms.

Inventive Principle:
Principle #1Segmentation

2Object-generated harmful factors

If short transform blocks are used for coding audio frames, then time-smearing is reduced, but data overhead increases leading to spectral holes

Engineering Contradiction:
Improvetime-smearing (pre-echo)VSAvoidspectral holes
Core Design Contradiction:
Object-generated harmful factorsVSLoss of information

Solution Approach 1:

Multiple short transforms are combined into a unified coding structure where sub-frames are processed together. This merging approach allows efficient packing of transform coefficients, reducing data overhead while maintaining the temporal resolution benefits of short transforms and avoiding spectral holes.

Inventive Principle:
Principle #5Merging (Combining)

3Measurement precision

If a new frequency-domain audio codec supporting additional transform lengths is developed, then coding quality for specific audio signals is improved, but backward compatibility with existing codecs is lost

Engineering Contradiction:
Improvecoding qualityVSAvoidbackward compatibility
Core Design Contradiction:
Measurement precisionVSAdaptability or versatility

Solution Approach 1:

The codec is designed with multi-functionality to handle both traditional long/short transform modes and the new transform length mode. By incorporating universal processing structures that can adapt to different transform types, the system achieves improved coding quality for specific signals while maintaining backward compatibility with existing decoders.

Inventive Principle:
Principle #6Universality (Multi-functionality)

Solution Approach 2:

The codec dynamically adapts its processing mode based on the transform length being used. The system can switch between different operational modes (long transform, short transform, and new transform length) to optimize performance for different signal types while ensuring that legacy decoders can still process the bitstream in traditional modes.

Inventive Principle:
Principle #15Dynamics

Data Source

PatentEP3312836B1Frequency-domain audio coding supporting transform length switching
Publication Date: 2021.10.27 FRAUNHOFER GESELLSCHAFT ZUR FORDERUNG DER ANGEWANDTEN FORSCHUNG EV
  • EP3312836B1 patent drawingFigure 1
  • EP3312836B1 patent drawingFigure 2
  • EP3312836B1 patent drawingFigure 3~4

AI summary

A frequency-domain audio codec is provided with the ability to additionally support a certain transform length in a backward-compatible manner, by the following: the frequency-domain coefficients of a respective frame are transmitted in an interleaved manner irrespective of the signalization signaling for the frames as to which transform length actually applies, and additionally the frequency-domain coefficient extraction and the scale factor extraction operate independent from the signalization. By this measure, old-fashioned frequency-domain audio coders/decoders, insensitive for the signalization, would be able to nevertheless operate without faults and with reproducing a reasonable quality. Concurrently, frequency-domain audio coders/decoders able to support the additional transform length would offer even better quality despite the backward compatibility. As far as coding efficiency penalties due to the coding of the frequency domain coefficients in a manner transparent for older decoders are concerned, same are of comparatively minor nature due to the interleaving.