Spatial Audio Encoding with Metadata for Resolution

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current systems for spatial audio in video games struggle with mixing audio sources effectively, leading to reduced spatial resolution and difficulty in working with object-based audio.

Innovation Solution

The proposed system encodes audio data with metadata for sound sources into a multichannel soundfield, allowing for optimized spatial audio decoding and compression, which can be later rendered through binaural or multichannel loudspeakers.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Ease of operation

If object-based audio streams are mixed together, then the audio processing becomes simpler, but spatial resolution is reduced and difficulty in working with object-based audio increases

Engineering Contradiction:
Improveease of working with object-based audioVSAvoidspatial resolution
Core Design Contradiction:
Ease of operationVSMeasurement precision

Solution Approach 1:

The patent segments the audio processing into two distinct stages: first encoding individual object audio streams with their spatial metadata separately, then mixing the encoded streams. This segmentation preserves the spatial information of each object while enabling simplified mixing operations, resolving the contradiction between ease of operation and spatial resolution.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent applies preliminary encoding to object audio streams before mixing, where spatial metadata is embedded in each object stream during the encoding phase. This preliminary action ensures spatial resolution is maintained before the mixing operation occurs, allowing simple mixing without sacrificing spatial precision.

Inventive Principle:
Principle #10Preliminary action

2Device complexity

If object metadata is discarded during encoding, then the encoding process becomes simpler, but spatial audio quality and rendering accuracy deteriorate

Engineering Contradiction:
Improveencoding process complexityVSAvoidspatial audio quality
Core Design Contradiction:
Device complexityVSReliability

Solution Approach 1:

The patent extracts spatial metadata from object audio streams and embeds it within the encoded streams themselves. This extraction and embedding process ensures that spatial information is preserved and available during decoding and rendering, maintaining spatial audio quality without significantly increasing encoding complexity.

Inventive Principle:
Principle #2Taking out (Extraction)

Solution Approach 2:

The patent uses spatial metadata as an intermediary element that bridges the object-based audio representation and the final spatial rendering. By embedding this metadata in the encoded streams, it mediates between the simple encoding process and the high-quality spatial output, ensuring both simplicity and reliability.

Inventive Principle:
Principle #24Intermediary (Mediator)

3Quantity of substance

If audio is encoded to multichannel soundfield without object metadata, then compression is achieved, but spatial rendering optimization is lost

Engineering Contradiction:
Improveaudio data compressionVSAvoidspatial rendering optimization
Core Design Contradiction:
Quantity of substanceVSAdaptability or versatility

Solution Approach 1:

The patent changes the representation parameters by encoding object audio streams with embedded spatial metadata into a multichannel soundfield format. This parameter transformation achieves compression while preserving spatial information, allowing both compression and spatial rendering optimization to coexist.

Inventive Principle:
Principle #35Parameter changes

Solution Approach 2:

The patent creates a composite audio representation that combines compressed multichannel soundfield data with embedded spatial metadata. This composite structure enables both compression benefits and spatial rendering optimization, as the metadata provides spatial guidance while the compressed soundfield provides efficient storage and transmission.

Inventive Principle:
Principle #40Composite materials

Data Source

PatentUS12317058B2Systems and methods for modifying spatial audio
Publication Date: 2025.05.27 SONY INTERACTIVE ENTERTAINMENT LLC
  • US12317058B2 patent drawing
  • US12317058B2 patent drawing
  • US12317058B2 patent drawing

AI summary

Systems and methods for modifying spatial audio are described. One of the methods includes obtaining a first set of metadata for a first set of audio data and a second set of metadata for a second set of audio data. The first and second sets of metadata and the first and second sets of audio data are associated with a display of a virtual scene. The method further includes encoding the first set of audio data to output a first soundfield and the second set of audio data to output a second soundfield. The method also includes mixing the first and second soundfields to output a mixed soundfield, decoding the mixed soundfield based on at least one of the first set of metadata and the second set of metadata to provide mixed audio data, and outputting the mixed audio data as an audio output.