3D Audio Scene Space Representation with Selective Syntax

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing methods fail to effectively represent and process the space of interest in an audio scene, which is crucial for audio coding, processing, and rendering, particularly in applications like medical imaging and geographical information systems.

Innovation Solution

The method involves defining and representing the space of interest in an audio scene using syntax elements to indicate the type, number, and identification indices of listener spaces, audio channel configurations, and audio object configurations, along with their subtypes, enabling efficient audio and video correlation.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Device complexity

If a region of interest (ROI) concept is directly applied to audio scene, then the representation becomes simple, but it fails to capture the three-dimensional spatial characteristics of audio objects

Engineering Contradiction:
Improverepresentation complexityVSAvoidspatial representation accuracy
Core Design Contradiction:
Device complexityVSMeasurement precision

Solution Approach 1:

The patent transitions from two-dimensional ROI concepts to three-dimensional audio scene representation by introducing elevation angles and vertical position parameters. Audio objects are represented with spherical coordinates including azimuth, elevation, and distance, enabling accurate spatial localization in 3D space rather than flat 2D regions.

Inventive Principle:
Principle #17Another dimension (Dimensionality change)

Solution Approach 2:

The audio scene is segmented into multiple audio objects, each with independent spatial parameters. The scene is divided into listener space, sound source space, and transition space, with each segment represented by specific syntax elements indicating position, distance, and spatial characteristics.

Inventive Principle:
Principle #1Segmentation

2Measurement precision

If detailed syntax elements are used to represent audio scene space, then the precision of space of interest representation is improved, but the data complexity and processing overhead increase

Engineering Contradiction:
Improvespace of interest representation precisionVSAvoiddata structure complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The patent creates a universal syntax element structure that can represent multiple types of spatial information (listener space, sound source space, transition space) using a unified framework. The same syntax elements serve multiple functions by indicating different spatial relationships depending on the context and type values.

Inventive Principle:
Principle #6Universality (Multi-functionality)

Solution Approach 2:

The representation uses variable parameter structures where syntax elements are selectively included or excluded based on the specific spatial configuration. Parameters such as elevation angles and vertical positions are introduced only when needed to represent three-dimensional space, avoiding unnecessary data complexity for simpler scenarios.

Inventive Principle:
Principle #35Parameter changes

3Manufacturing precision

If audio scene data includes multiple syntax elements for space representation, then the audio coding accuracy is improved, but the encoding and decoding complexity increases

Engineering Contradiction:
Improveaudio coding accuracyVSAvoidencoding complexity
Core Design Contradiction:
Manufacturing precisionVSDevice complexity

Solution Approach 1:

The patent performs preliminary spatial analysis during encoding to determine which syntax elements are necessary for accurate representation. The encoder pre-processes the audio scene to identify important spatial relationships and selects only the essential syntax elements, reducing the overall complexity while maintaining coding accuracy.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The patent extracts only the critical spatial parameters needed for accurate audio scene representation. By selectively extracting and encoding only the most important syntax elements (such as primary listener space and dominant sound sources), the system achieves high coding accuracy without processing all possible spatial parameters.

Inventive Principle:
Principle #2Taking out (Extraction)

Data Source

PatentEP4122225B1Method and apparatus for representing space of interest of audio scene
Publication Date: 2025.08.20 TENCENT AMERICA LLC
  • EP4122225B1 patent drawingFigure 1~2
  • EP4122225B1 patent drawingFigure 3~4
  • EP4122225B1 patent drawingFigure 5

AI summary

Aspects of the disclosure include methods, apparatuses, and non-transitory computer-readable storage mediums for representing a space of interest of an audio scene. One apparatus includes processing circuitry that decodes audio scene data for the audio scene. The audio scene data includes (i) audio content for a plurality of items representing the audio scene and (ii) a first syntax element indicating a type of a subset of the plurality of items. The subset of the plurality of items represents the space of interest of the audio scene. The processing circuitry determines a part of the audio content for the subset of the plurality of items based on the type of the subset of the plurality of items indicated in the first syntax element. The processing circuitry renders the determined part of the audio content.