Adaptive Downmix Encoding for Immersive Audio Reconstruction

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing immersive voice and audio services (IVAS) codecs face challenges in reconstructing immersive audio scenes accurately due to imperfect decorrelation and quantization limitations in passive downmix schemes, leading to audio artifacts and spatial distortion.

Innovation Solution

Implementing adaptive downmix strategies that combine passive and active downmix techniques, adjusting input downmixing gains and scaling factors to minimize prediction errors and ensure accurate reconstruction of immersive audio scenes, using signal-dependent computations and adaptive operation based on voice activity detection.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Ease of manufacture

If passive downmix scheme is used, then implementation is simple, but reconstruction accuracy deteriorates due to imperfect decorrelation and quantization limitations

Engineering Contradiction:
Improveimplementation simplicityVSAvoidreconstruction accuracy
Core Design Contradiction:
Ease of manufactureVSMeasurement precision

Solution Approach 1:

The system dynamically adapts the downmix scheme based on signal characteristics. Voice activity detection determines when to switch between passive and active downmixing, allowing the system to optimize between simplicity and accuracy depending on the audio content being processed.

Inventive Principle:
Principle #15Dynamics

Solution Approach 2:

The patent changes the downmixing parameters (gains, scaling factors) based on signal-dependent computations. By adjusting these parameters adaptively, the system improves reconstruction accuracy while maintaining implementation feasibility through controlled complexity.

Inventive Principle:
Principle #35Parameter changes

2Measurement precision

If active downmix scheme is used, then reconstruction accuracy improves, but computational complexity increases

Engineering Contradiction:
Improvereconstruction accuracyVSAvoidcomputational complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The system uses voice activity detection to dynamically control when active downmixing is applied. This reduces computational complexity by limiting complex processing to only when necessary (during voice activity), while maintaining high reconstruction accuracy when needed.

Inventive Principle:
Principle #15Dynamics

Solution Approach 2:

Instead of applying complex active downmixing continuously, the system applies it partially - only during voice activity periods. This reduces overall computational complexity while maintaining reconstruction accuracy for the most important signal portions.

Inventive Principle:
Principle #16Partial or excessive action

3Productivity

If downmix channels are reduced at low bitrate, then bandwidth efficiency improves, but audio quality deteriorates

Engineering Contradiction:
Improvebandwidth efficiencyVSAvoidaudio quality
Core Design Contradiction:
ProductivityVSMeasurement precision

Solution Approach 1:

The system changes downmixing parameters (gains, scaling factors) based on signal characteristics to optimize the representation of the audio scene in fewer channels. This allows better audio quality preservation at low bitrates by intelligently allocating the limited channel resources.

Inventive Principle:
Principle #35Parameter changes

Solution Approach 2:

The system performs preliminary signal analysis (voice activity detection, signal-dependent computations) before downmixing to prepare optimal parameters. This preliminary action enables the reduced channel configuration to maintain better audio quality by pre-planning the downmix strategy.

Inventive Principle:
Principle #10Preliminary action

Data Source

PatentUS12431145B2Immersive voice and audio services (IVAS) with adaptive downmix strategies
Publication Date: 2025.09.30 DOLBY LABORATORIES LICENSING CORP
  • US12431145B2 patent drawing
  • US12431145B2 patent drawing
  • US12431145B2 patent drawing

AI summary

Disclosed is an audio signal encoding/decoding method that uses an encoding downmix strategy applied at an encoder that is different than a decoding re-mix/upmix strategy applied at a decoder. Based on the type of downmix coding scheme, the method comprises: computing input downmixing gains to be applied to the input audio signal to construct a primary downmix channel; determining downmix scaling gains to scale the primary downmix channel; generating prediction gains based on the input audio signal, the input downmixing gains and the downmix scaling gains; determining residual channel(s) from the side channels by using the primary downmix channel and the prediction gains to generate side channel predictions and subtracting the side channel predictions from the side channels; determining decorrelation gains based on energy in the residual channels; encoding the primary downmix channel, the residual channel(s), the prediction gains and the decorrelation gains; and sending the bitstream to a decoder.