Spatial Audio Mixing with Metadata Expansion

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current methods for transmitting and rendering spatial audio in virtual reality (VR) and augmented reality (AR) environments are inefficient, requiring a high number of audio channels and computational resources, which can lead to reduced spatial fidelity and increased computational loading/storage and transmission capacity requirements.

Innovation Solution

The proposed solution involves an apparatus and method for mixing spatial audio capture microphone channels with external audio channels, using a processor to generate combined parameter outputs and a mixer to produce a combined audio signal with fewer channels, while expanding spatial metadata to maintain high spatial fidelity, allowing for efficient transmission and rendering.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If traditional multi-channel audio coding techniques are used to transmit spatial audio, then spatial fidelity can be maintained, but the number of transmitted channels and computational resources increase significantly

Engineering Contradiction:
Improvespatial fidelityVSAvoidnumber of transmitted channels
Core Design Contradiction:
Measurement precisionVSQuantity of substance

Solution Approach 1:

The patent combines multiple audio channels into a reduced set of channels by merging spatially correlated channels. This is achieved through parametric audio coding where the essential spatial information is extracted and represented with fewer channels, reducing transmission requirements while preserving perceptual spatial fidelity.

Inventive Principle:
Principle #5Merging (Combining)

Solution Approach 2:

The patent transforms the audio signal representation from time-domain waveforms to frequency-domain parameters including spectral information and spatial metadata. By encoding audio in frequency bands with parametric spatial information (such as inter-channel level differences, inter-channel time differences, and spatial position data), the system achieves efficient compression while maintaining spatial accuracy.

Inventive Principle:
Principle #35Parameter changes

2Quantity of substance

If parametric multi-channel audio coding is used to reduce the number of channels, then transmission efficiency improves, but spatial fidelity may be compromised

Engineering Contradiction:
Improvenumber of transmitted channelsVSAvoidspatial fidelity
Core Design Contradiction:
Quantity of substanceVSMeasurement precision

Solution Approach 1:

The patent adds spatial metadata dimensions to the audio coding system, including three-dimensional position information (azimuth, elevation, distance), spatial spread, and temporal coherence parameters. This multi-dimensional parameter space allows the system to represent complex spatial audio scenes with fewer audio channels while maintaining high spatial fidelity through accurate parametric description.

Inventive Principle:
Principle #17Another dimension (Dimensionality change)

3Measurement precision

If high-resolution HRTF rendering is applied for headphone binaural output, then spatial accuracy improves, but computational loading increases

Engineering Contradiction:
Improvespatial accuracyVSAvoidcomputational loading
Core Design Contradiction:
Measurement precisionVSUse of energy by moving object

Solution Approach 1:

The patent performs preliminary spatial analysis and parameter extraction during the encoding phase, pre-computing spatial metadata including direction of arrival, spatial position, and spectral characteristics. This preliminary processing allows the decoder to efficiently render binaural audio using stored HRTF filters without requiring complex real-time spatial analysis, significantly reducing computational loading while maintaining high spatial accuracy.

Inventive Principle:
Principle #10Preliminary action

Data Source

PatentUS10674262B2Merging audio signals with spatial metadata
Publication Date: 2020.06.02 NOKIA TECHNOLOGIES OY
  • US10674262B2 patent drawing
  • US10674262B2 patent drawing
  • US10674262B2 patent drawing

AI summary

Apparatus for mixing at least two audio signals, at least one audio signal associated with at least one parameter, and at least one second audio signal further associated with at least one second parameter, wherein the at least one audio signal and the at least one second audio signal are associated with a sound scene and wherein the at least one audio signal represent spatial audio capture microphone channels and the at least one second audio signal represents an external audio channel separate from the spatial audio capture microphone channels, the apparatus comprising: a processor configured to generate a combined parameter output based on the at least one second parameter and the at least one parameter; and a mixer configured to generate a combined audio signal with a same number or fewer number of channels as the at least one audio signal based on the at least one audio signal and the at least one second audio signal, wherein the combined audio signal is associated with the combined parameter.