Audio Decoder With Frequency-Band Convolution for Lower Complexity

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing audio signal processing methods for headphone playback, particularly for object-based content, face high computational complexity and memory requirements due to HRIR/BRIR convolution, and require impractical bit rates for delivery, especially on battery-powered devices.

Innovation Solution

A method using multi-tap convolution matrices for low frequencies and high frequency resolution in the encoder, combined with stateless matrices for higher frequencies, reduces computational complexity and memory usage by compensating for reduced decoder frequency resolution.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If HRIR/BRIR convolution is applied for headphone playback, then audio spatialization and sound source localization are improved, but computational complexity and memory requirements increase significantly

Engineering Contradiction:
Improveaudio spatialization accuracyVSAvoidcomputational complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The frequency spectrum is divided into multiple bands (low, mid, high frequencies), with different processing methods applied to each band. Multi-tap convolution matrices are used for low frequencies while stateless matrices are used for higher frequencies, reducing overall computational complexity while maintaining spatialization accuracy.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

Different processing techniques are applied to different frequency regions based on their specific characteristics. The patent applies multi-tap convolution for low frequencies where spatial cues are critical, while using simpler stateless matrix processing for higher frequencies, optimizing the balance between quality and complexity.

Inventive Principle:
Principle #3Local quality

2Measurement precision

If high frequency resolution is used in the encoder, then audio quality is improved, but bit rate requirements increase to impractical levels

Engineering Contradiction:
Improvefrequency resolutionVSAvoidbit rate
Core Design Contradiction:
Measurement precisionVSQuantity of substance

Solution Approach 1:

The patent changes the frequency resolution parameter dynamically based on the frequency band. High frequency resolution is maintained in the encoder for low frequencies, while the decoder uses reduced frequency resolution for higher frequencies, compensated by the stateless matrix processing, thereby reducing bit rate requirements while maintaining perceived audio quality.

Inventive Principle:
Principle #35Parameter changes

3Device complexity

If multi-tap convolution matrices are used for low frequencies, then computational complexity is reduced, but decoder frequency resolution is lowered

Engineering Contradiction:
Improvecomputational complexityVSAvoidfrequency resolution
Core Design Contradiction:
Device complexityVSMeasurement precision

Solution Approach 1:

The frequency spectrum is segmented into low and high frequency bands, with multi-tap convolution applied to low frequencies and stateless matrices to high frequencies. This segmentation allows the system to accept reduced frequency resolution at the decoder for low frequencies while maintaining it for high frequencies, balancing complexity and quality.

Inventive Principle:
Principle #1Segmentation

Data Source

PatentUS12400669B2Audio decoder and decoding method
Publication Date: 2025.08.26 DOLBY LABORATORIES LICENSING CORP
  • US12400669B2 patent drawing
  • US12400669B2 patent drawing
  • US12400669B2 patent drawing

AI summary

A method for representing a second presentation of audio channels or objects as a data stream, the method comprising the steps of: (a) providing a set of base signals, the base signals representing a first presentation of the audio channels or objects; (b) providing a set of transformation parameters, the transformation parameters intended to transform the first presentation into the second presentation; the transformation parameters further being specified for at least two frequency bands and including a set of multi-tap convolution matrix parameters for at least one of the frequency bands.