Frequency-Domain Audio Coding With Backward-Compatible Transform Switching

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Modern frequency-domain audio coding systems face challenges in achieving satisfactory quality for audio signals with prominent transients, such as rain or applause, at low bitrates, as existing long or short block coding methods result in time-smearing or inefficiencies due to increased data overhead.

Innovation Solution

A frequency-domain audio codec that supports transform length switching, allowing for the use of multiple transform lengths within a frame, enabling better adaptation to signal characteristics while maintaining backward compatibility with existing codecs through interleaving of frequency-domain coefficients and independent extraction of scale factors.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If long block coding is used, then temporal resolution is improved, but pre-echo artifacts increase

Engineering Contradiction:
Improvetemporal resolutionVSAvoidpre-echo artifacts
Core Design Contradiction:
Measurement precisionVSObject-generated harmful factors

Solution Approach 1:

The frame is divided into multiple sub-frames, each processed with a separate transform. This segmentation allows the system to achieve fine temporal resolution in regions with transients while maintaining good spectral resolution in stationary regions, thereby reducing pre-echo artifacts while preserving temporal precision where needed.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The transform length is dynamically adjusted based on the signal characteristics detected in different parts of the frame. By switching between short and long transforms adaptively, the system optimizes temporal resolution locally without incurring pre-echo artifacts in transient regions.

Inventive Principle:
Principle #15Dynamics

2Object-generated harmful factors

If short block coding is used, then pre-echo artifacts are reduced, but coding efficiency deteriorates

Engineering Contradiction:
Improvepre-echo artifactsVSAvoidcoding efficiency
Core Design Contradiction:
Object-generated harmful factorsVSProductivity

Solution Approach 1:

Different transform lengths are applied to different sub-frames based on local signal characteristics. Regions with transients use short transforms to avoid pre-echo, while stationary regions use long transforms to maintain coding efficiency, achieving local optimization of both quality and efficiency.

Inventive Principle:
Principle #3Local quality

Solution Approach 2:

The transform length parameter is changed adaptively across different sub-frames within a frame. This parameter variation allows the system to optimize coding efficiency in stationary regions while using short transforms in transient regions to eliminate pre-echo artifacts.

Inventive Principle:
Principle #35Parameter changes

3Adaptability or versatility

If multiple transform lengths are supported, then adaptability to signal characteristics is improved, but device complexity increases

Engineering Contradiction:
Improveadaptability to signal characteristicsVSAvoiddevice complexity
Core Design Contradiction:
Adaptability or versatilityVSDevice complexity

Solution Approach 1:

The frame is segmented into sub-frames that can be processed with different transform lengths. This segmentation approach enables multi-transform-length support with manageable complexity, as each sub-frame can be independently processed with the appropriate transform length based on its characteristics.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The system dynamically selects transform lengths based on signal characteristics detected in each sub-frame. This dynamic adaptation improves versatility while keeping device complexity manageable through algorithmic decision-making rather than hardware complexity.

Inventive Principle:
Principle #15Dynamics

4Object-generated harmful factors

If transform length switching is implemented, then audio quality for transient signals is improved, but backward compatibility challenges arise

Engineering Contradiction:
Improvepre-echo artifactsVSAvoidbackward compatibility
Core Design Contradiction:
Object-generated harmful factorsVSAdaptability or versatility

Solution Approach 1:

The decoder is designed to handle both single-transform-length and multi-transform-length modes universally. The frequency-domain coefficient extractor and scale factor extractor operate independently of the transform length configuration, allowing legacy decoders to process the bitstream correctly while enhanced decoders can exploit transform length switching for improved quality.

Inventive Principle:
Principle #6Universality (Multi-functionality)

Solution Approach 2:

The patent uses an intermediary signaling mechanism that indicates transform length configuration in the bitstream. This intermediary allows enhanced decoders to interpret the signaling and apply appropriate inverse transforms, while legacy decoders ignore the signaling and use default processing, thereby maintaining backward compatibility.

Inventive Principle:
Principle #24Intermediary (Mediator)

Data Source

PatentUS10242682B2Frequency-domain audio coding supporting transform length switching
Publication Date: 2019.03.26 FRAUNHOFER GESELLSCHAFT ZUR FORDERUNG DER ANGEWANDTEN FORSCHUNG EV
  • US10242682B2 patent drawing
  • US10242682B2 patent drawing
  • US10242682B2 patent drawing

AI summary

A frequency-domain audio codec is provided with the ability to additionally support a certain transform length in a backward-compatible manner, by the following: the frequency-domain coefficients of a respective frame are transmitted in an interleaved manner irrespective of the signalization signaling for the frames as to which transform length actually applies, and additionally the frequency-domain coefficient extraction and the scale factor extraction operate independent from the signalization. By this measure, old-fashioned frequency-domain audio coders/decoders, insensitive for the signalization, would be able to nevertheless operate without faults and with reproducing a reasonable quality. Concurrently, frequency-domain audio coders/decoders able to support the additional transform length would offer even better quality despite the backward compatibility. As far as coding efficiency penalties due to the coding of the frequency domain coefficients in a manner transparent for older decoders are concerned, same are of comparatively minor nature due to the interleaving.