Parametric Spatial Audio Rendering Mixing Values

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing rendering techniques for immersive audio codecs, such as IVAS, fail to accurately replicate the energy decrease of side channels when downmixing a multichannel audio signal to stereo, leading to unwanted equal loudness perception and increased processing requirements, which can degrade the audio experience.

Innovation Solution

The proposed solution involves generating mixing values based on spatial metadata and predefined parameters to create direct and ambient sound values for each channel, using a predefined matrix or vector of channel gains to maintain the desired spectral effects without introducing decorrelation, thus preserving the energy decrease of side channels.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If existing rendering techniques are used to downmix multichannel audio to stereo, then the audio signal can be rendered, but the energy decrease of side channels is not accurately replicated, leading to unwanted equal loudness perception

Engineering Contradiction:
Improveenergy replication accuracyVSAvoidloudness perception accuracy
Core Design Contradiction:
Measurement precisionVSReliability

Solution Approach 1:

The patent changes the rendering parameters by introducing mixing values that are specifically calculated to replicate the energy decrease of side channels. Instead of using conventional rendering parameters that treat all channels equally, the invention modifies the energy parameters for side channels (e.g., applying 0.7 energy values to side channels versus 1.0 for front channels) to accurately reproduce the perceptual energy distribution of the original multichannel signal.

Inventive Principle:
Principle #35Parameter changes

2Productivity

If conventional rendering methods are used, then processing can be performed, but processing requirements increase and decorrelation artifacts are introduced

Engineering Contradiction:
Improveprocessing efficiencyVSAvoidprocessing complexity
Core Design Contradiction:
ProductivityVSDevice complexity

Solution Approach 1:

The patent extracts only the essential spectral characteristics needed for accurate rendering, specifically the energy distribution pattern across channels. Instead of performing complex full-spectrum analysis and processing, the invention extracts and applies pre-calculated mixing values that capture the essential energy relationships, thereby reducing processing complexity while maintaining accuracy.

Inventive Principle:
Principle #2Taking out (Extraction)

3Device complexity

If side channel energy is not properly decreased in downmixing, then the rendering process is simpler, but the spectral effects of downmixing are lost

Engineering Contradiction:
Improverendering process complexityVSAvoidspectral effects
Core Design Contradiction:
Device complexityVSLoss of information

Solution Approach 1:

The patent applies preliminary action by pre-calculating the mixing values that encode the desired energy relationships before the actual rendering process. The mixing values are prepared in advance based on the target spectral characteristics, so that during rendering, the system simply applies these pre-determined values rather than computing energy relationships in real-time, thus preserving spectral effects while maintaining simplicity.

Inventive Principle:
Principle #10Preliminary action

Data Source

PatentUS20240274137A1Parametric spatial audio rendering
Publication Date: 2024.08.15 NOKIA TECHNOLOGIES OY
  • US20240274137A1 patent drawing
  • US20240274137A1 patent drawing
  • US20240274137A1 patent drawing

AI summary

An apparatus (317) comprising means configured to: receive a spatial audio signal, the spatial audio signal comprising at least one audio signal and spatial metadata (122) associated with the at least one audio signal; generate a mixing value (320) based on the spatial metadata (122) and a predefined parameter (322) which imparts effects of a rendering of a multichannel audio signal having a multichannel configuration to a further multichannel audio signal having a further multichannel configuration on generated output signals; and generating the output audio signals having the further multichannel configuration based on the mixing value (320) and the spatial audio signal.