Spatial Audio Encoding with Metadata for Resolution
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current systems for spatial audio in video games struggle with mixing audio sources effectively, leading to reduced spatial resolution and difficulty in working with object-based audio.
Innovation Solution
The proposed system encodes audio data with metadata for sound sources into a multichannel soundfield, allowing for optimized spatial audio decoding and compression, which can be later rendered through binaural or multichannel loudspeakers.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Ease of operation
If object-based audio streams are mixed together, then the audio processing becomes simpler, but spatial resolution is reduced and difficulty in working with object-based audio increases
Solution Approach 1:
The patent segments the audio processing into two distinct stages: first encoding individual object audio streams with their spatial metadata separately, then mixing the encoded streams. This segmentation preserves the spatial information of each object while enabling simplified mixing operations, resolving the contradiction between ease of operation and spatial resolution.
Solution Approach 2:
The patent applies preliminary encoding to object audio streams before mixing, where spatial metadata is embedded in each object stream during the encoding phase. This preliminary action ensures spatial resolution is maintained before the mixing operation occurs, allowing simple mixing without sacrificing spatial precision.
2Device complexity
If object metadata is discarded during encoding, then the encoding process becomes simpler, but spatial audio quality and rendering accuracy deteriorate
Solution Approach 1:
The patent extracts spatial metadata from object audio streams and embeds it within the encoded streams themselves. This extraction and embedding process ensures that spatial information is preserved and available during decoding and rendering, maintaining spatial audio quality without significantly increasing encoding complexity.
Solution Approach 2:
The patent uses spatial metadata as an intermediary element that bridges the object-based audio representation and the final spatial rendering. By embedding this metadata in the encoded streams, it mediates between the simple encoding process and the high-quality spatial output, ensuring both simplicity and reliability.
3Quantity of substance
If audio is encoded to multichannel soundfield without object metadata, then compression is achieved, but spatial rendering optimization is lost
Solution Approach 1:
The patent changes the representation parameters by encoding object audio streams with embedded spatial metadata into a multichannel soundfield format. This parameter transformation achieves compression while preserving spatial information, allowing both compression and spatial rendering optimization to coexist.
Solution Approach 2:
The patent creates a composite audio representation that combines compressed multichannel soundfield data with embedded spatial metadata. This composite structure enables both compression benefits and spatial rendering optimization, as the metadata provides spatial guidance while the compressed soundfield provides efficient storage and transmission.
Data Source
AI summary
Systems and methods for modifying spatial audio are described. One of the methods includes obtaining a first set of metadata for a first set of audio data and a second set of metadata for a second set of audio data. The first and second sets of metadata and the first and second sets of audio data are associated with a display of a virtual scene. The method further includes encoding the first set of audio data to output a first soundfield and the second set of audio data to output a second soundfield. The method also includes mixing the first and second soundfields to output a mixed soundfield, decoding the mixed soundfield based on at least one of the first set of metadata and the second set of metadata to provide mixed audio data, and outputting the mixed audio data as an audio output.


