Spatial Audio Mixing with Metadata Expansion
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current methods for transmitting and rendering spatial audio in virtual reality (VR) and augmented reality (AR) environments are inefficient, requiring a high number of audio channels and computational resources, which can lead to reduced spatial fidelity and increased computational loading/storage and transmission capacity requirements.
Innovation Solution
The proposed solution involves an apparatus and method for mixing spatial audio capture microphone channels with external audio channels, using a processor to generate combined parameter outputs and a mixer to produce a combined audio signal with fewer channels, while expanding spatial metadata to maintain high spatial fidelity, allowing for efficient transmission and rendering.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If traditional multi-channel audio coding techniques are used to transmit spatial audio, then spatial fidelity can be maintained, but the number of transmitted channels and computational resources increase significantly
Solution Approach 1:
The patent combines multiple audio channels into a reduced set of channels by merging spatially correlated channels. This is achieved through parametric audio coding where the essential spatial information is extracted and represented with fewer channels, reducing transmission requirements while preserving perceptual spatial fidelity.
Solution Approach 2:
The patent transforms the audio signal representation from time-domain waveforms to frequency-domain parameters including spectral information and spatial metadata. By encoding audio in frequency bands with parametric spatial information (such as inter-channel level differences, inter-channel time differences, and spatial position data), the system achieves efficient compression while maintaining spatial accuracy.
2Quantity of substance
If parametric multi-channel audio coding is used to reduce the number of channels, then transmission efficiency improves, but spatial fidelity may be compromised
Solution Approach 1:
The patent adds spatial metadata dimensions to the audio coding system, including three-dimensional position information (azimuth, elevation, distance), spatial spread, and temporal coherence parameters. This multi-dimensional parameter space allows the system to represent complex spatial audio scenes with fewer audio channels while maintaining high spatial fidelity through accurate parametric description.
3Measurement precision
If high-resolution HRTF rendering is applied for headphone binaural output, then spatial accuracy improves, but computational loading increases
Solution Approach 1:
The patent performs preliminary spatial analysis and parameter extraction during the encoding phase, pre-computing spatial metadata including direction of arrival, spatial position, and spectral characteristics. This preliminary processing allows the decoder to efficiently render binaural audio using stored HRTF filters without requiring complex real-time spatial analysis, significantly reducing computational loading while maintaining high spatial accuracy.
Data Source
AI summary
Apparatus for mixing at least two audio signals, at least one audio signal associated with at least one parameter, and at least one second audio signal further associated with at least one second parameter, wherein the at least one audio signal and the at least one second audio signal are associated with a sound scene and wherein the at least one audio signal represent spatial audio capture microphone channels and the at least one second audio signal represents an external audio channel separate from the spatial audio capture microphone channels, the apparatus comprising: a processor configured to generate a combined parameter output based on the at least one second parameter and the at least one parameter; and a mixer configured to generate a combined audio signal with a same number or fewer number of channels as the at least one audio signal based on the at least one audio signal and the at least one second audio signal, wherein the combined audio signal is associated with the combined parameter.


