Encoded Audio Scene Bandwidth Extension for Low-Delay Stereo

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing audio processing systems for immersive communication, such as DirAC, face high delays and complexity when converting mono recordings to stereo output due to the use of filter banks, which is suboptimal for low-bitrate scenarios.

Innovation Solution

A method involving a parameter converter that converts DirAC side parameters into stereo parameters using a Short-Time Fourier Transform (STFT) filterbank, allowing parallel processing to achieve low-delay upmixing without additional delay, and combining this with ACELP speech coder processing to maintain the same overall delay as EVS codecs.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If DirAC uses filter-bank domain processing for spatial restoration, then spatial information can be restored from mono channel, but processing delay increases significantly

Engineering Contradiction:
Improvespatial information restorationVSAvoidprocessing delay
Core Design Contradiction:
Measurement precisionVSLoss of time

Solution Approach 1:

The patent replaces the traditional filter-bank domain spatial restoration with a time-domain approach using impulse response convolution. Instead of using complex filter banks to restore spatial information, the system uses a simplified impulse response-based method that operates in the time domain, significantly reducing processing delay while maintaining spatial restoration capability.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

Solution Approach 2:

The patent changes the processing domain from frequency domain (filter-bank) to time domain (impulse response). By transforming the spatial restoration mechanism from filter-bank domain to time-domain impulse response convolution, the system achieves the same spatial restoration function with much lower delay, as the impulse response method requires no additional filter-bank analysis/synthesis steps.

Inventive Principle:
Principle #35Parameter changes

2Adaptability or versatility

If FOA upmix with L/R conversion is used for stereo output, then stereo playback is achieved, but overall delay increases compared to direct stereo extraction

Engineering Contradiction:
Improvestereo playback capabilityVSAvoidoverall delay
Core Design Contradiction:
Adaptability or versatilityVSLoss of time

Solution Approach 1:

The patent extracts and removes the unnecessary FOA upmix and L/R conversion steps from the traditional DirAC processing chain. By directly extracting stereo information from the encoded audio scene using a simplified time-domain method, the system eliminates the intermediate FOA representation step that causes additional delay, while still achieving high-quality stereo playback.

Inventive Principle:
Principle #2Taking out (Extraction)

3Adaptability or versatility

If CLDFB filter-bank analysis/synthesis is used for spatial rendering, then immersive audio is achieved, but processing complexity and delay increase

Engineering Contradiction:
Improveimmersive audio renderingVSAvoidfilter-bank processing complexity
Core Design Contradiction:
Adaptability or versatilityVSDevice complexity

Solution Approach 1:

The patent substitutes the complex CLDFB filter-bank analysis/synthesis system with a simple time-domain impulse response convolution system. Instead of using sophisticated filter banks to achieve immersive audio rendering, the system uses a computationally efficient impulse response-based approach that achieves the same rendering goal with much lower complexity and delay.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

Data Source

PatentUS12425793B2Apparatus, method, or computer program for processing an encoded audio scene using a bandwidth extension
Publication Date: 2025.09.23 FRAUNHOFER GESELLSCHAFT ZUR FORDERUNG DER ANGEWANDTEN FORSCHUNG EV
  • US12425793B2 patent drawing
  • US12425793B2 patent drawing
  • US12425793B2 patent drawing

AI summary

Apparatus for processing an audio scene representing a sound field, the audio scene comprising information on a transport signal and a set of parameters. The apparatus comprising an output interface for generating a processed audio scene using the set of parameters and the information on the transport signal, wherein the output interface is configured to generate a raw representation of two or more channels using the set of parameters and the transport signal and a multichannel enhancer for generating an enhancement representation of the two or more channels using the transport signal, and a signal combiner for combining the raw representation of the two or more channels and the enhancement representation of the two or more channels to obtain the processed audio scene.