Adaptive Downmix Encoding for Immersive Audio Reconstruction
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing immersive voice and audio services (IVAS) codecs face challenges in reconstructing immersive audio scenes accurately due to imperfect decorrelation and quantization limitations in passive downmix schemes, leading to audio artifacts and spatial distortion.
Innovation Solution
Implementing adaptive downmix strategies that combine passive and active downmix techniques, adjusting input downmixing gains and scaling factors to minimize prediction errors and ensure accurate reconstruction of immersive audio scenes, using signal-dependent computations and adaptive operation based on voice activity detection.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Ease of manufacture
If passive downmix scheme is used, then implementation is simple, but reconstruction accuracy deteriorates due to imperfect decorrelation and quantization limitations
Solution Approach 1:
The system dynamically adapts the downmix scheme based on signal characteristics. Voice activity detection determines when to switch between passive and active downmixing, allowing the system to optimize between simplicity and accuracy depending on the audio content being processed.
Solution Approach 2:
The patent changes the downmixing parameters (gains, scaling factors) based on signal-dependent computations. By adjusting these parameters adaptively, the system improves reconstruction accuracy while maintaining implementation feasibility through controlled complexity.
2Measurement precision
If active downmix scheme is used, then reconstruction accuracy improves, but computational complexity increases
Solution Approach 1:
The system uses voice activity detection to dynamically control when active downmixing is applied. This reduces computational complexity by limiting complex processing to only when necessary (during voice activity), while maintaining high reconstruction accuracy when needed.
Solution Approach 2:
Instead of applying complex active downmixing continuously, the system applies it partially - only during voice activity periods. This reduces overall computational complexity while maintaining reconstruction accuracy for the most important signal portions.
3Productivity
If downmix channels are reduced at low bitrate, then bandwidth efficiency improves, but audio quality deteriorates
Solution Approach 1:
The system changes downmixing parameters (gains, scaling factors) based on signal characteristics to optimize the representation of the audio scene in fewer channels. This allows better audio quality preservation at low bitrates by intelligently allocating the limited channel resources.
Solution Approach 2:
The system performs preliminary signal analysis (voice activity detection, signal-dependent computations) before downmixing to prepare optimal parameters. This preliminary action enables the reduced channel configuration to maintain better audio quality by pre-planning the downmix strategy.
Data Source
AI summary
Disclosed is an audio signal encoding/decoding method that uses an encoding downmix strategy applied at an encoder that is different than a decoding re-mix/upmix strategy applied at a decoder. Based on the type of downmix coding scheme, the method comprises: computing input downmixing gains to be applied to the input audio signal to construct a primary downmix channel; determining downmix scaling gains to scale the primary downmix channel; generating prediction gains based on the input audio signal, the input downmixing gains and the downmix scaling gains; determining residual channel(s) from the side channels by using the primary downmix channel and the prediction gains to generate side channel predictions and subtracting the side channel predictions from the side channels; determining decorrelation gains based on energy in the residual channels; encoding the primary downmix channel, the residual channel(s), the prediction gains and the decorrelation gains; and sending the bitstream to a decoder.


