Parametric Spatial Audio Rendering Mixing Values
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing rendering techniques for immersive audio codecs, such as IVAS, fail to accurately replicate the energy decrease of side channels when downmixing a multichannel audio signal to stereo, leading to unwanted equal loudness perception and increased processing requirements, which can degrade the audio experience.
Innovation Solution
The proposed solution involves generating mixing values based on spatial metadata and predefined parameters to create direct and ambient sound values for each channel, using a predefined matrix or vector of channel gains to maintain the desired spectral effects without introducing decorrelation, thus preserving the energy decrease of side channels.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If existing rendering techniques are used to downmix multichannel audio to stereo, then the audio signal can be rendered, but the energy decrease of side channels is not accurately replicated, leading to unwanted equal loudness perception
Solution Approach 1:
The patent changes the rendering parameters by introducing mixing values that are specifically calculated to replicate the energy decrease of side channels. Instead of using conventional rendering parameters that treat all channels equally, the invention modifies the energy parameters for side channels (e.g., applying 0.7 energy values to side channels versus 1.0 for front channels) to accurately reproduce the perceptual energy distribution of the original multichannel signal.
2Productivity
If conventional rendering methods are used, then processing can be performed, but processing requirements increase and decorrelation artifacts are introduced
Solution Approach 1:
The patent extracts only the essential spectral characteristics needed for accurate rendering, specifically the energy distribution pattern across channels. Instead of performing complex full-spectrum analysis and processing, the invention extracts and applies pre-calculated mixing values that capture the essential energy relationships, thereby reducing processing complexity while maintaining accuracy.
3Device complexity
If side channel energy is not properly decreased in downmixing, then the rendering process is simpler, but the spectral effects of downmixing are lost
Solution Approach 1:
The patent applies preliminary action by pre-calculating the mixing values that encode the desired energy relationships before the actual rendering process. The mixing values are prepared in advance based on the target spectral characteristics, so that during rendering, the system simply applies these pre-determined values rather than computing energy relationships in real-time, thus preserving spectral effects while maintaining simplicity.
Data Source
AI summary
An apparatus (317) comprising means configured to: receive a spatial audio signal, the spatial audio signal comprising at least one audio signal and spatial metadata (122) associated with the at least one audio signal; generate a mixing value (320) based on the spatial metadata (122) and a predefined parameter (322) which imparts effects of a rendering of a multichannel audio signal having a multichannel configuration to a further multichannel audio signal having a further multichannel configuration on generated output signals; and generating the output audio signals having the further multichannel configuration based on the mixing value (320) and the spatial audio signal.


