AAC Audio Decoding with Height Metadata for Spatial Sound

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current multi-channel audio encoding standards, such as MPEG-2AAC and MPEG-4AAC, struggle to effectively extend channels beyond the traditional 5.1 configuration, particularly in the vertical direction, leading to difficulties in reproducing high-quality realistic sound with extended channel arrangements.

Innovation Solution

The proposed technique involves encoding and decoding audio data by including sound source position information about the height of sound sources, allowing for the storage and retrieval of this information in the encoded bit stream, which enables the reproduction of sound images in both the horizontal and vertical planes, using specific encoding and decoding processes that include synchronous words and CRC check codes for accurate data identification.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Productivity

If multi-channel encoding based on MPEG-2AAC or MPEG-4AAC standards is used, then encoding efficiency is improved, but the ability to extend channels in vertical direction and reproduce realistic sound is insufficient

Engineering Contradiction:
Improveencoding efficiencyVSAvoidchannel extension capability
Core Design Contradiction:
ProductivityVSAdaptability or versatility

Solution Approach 1:

The patent extends the traditional horizontal plane channel arrangement by introducing vertical dimension information through height parameters (front_element_height_info, side_element_height_info, back_element_height_info). This allows speakers to be positioned in three-dimensional space, enabling channel extension beyond the conventional 5.1 configuration into immersive spatial audio formats while maintaining compatibility with existing encoding standards.

Inventive Principle:
Principle #17Another dimension (Dimensionality change)

2Device complexity

If only horizontal speaker arrangement information is stored, then data structure simplicity is maintained, but vertical direction sound positioning accuracy is lost

Engineering Contradiction:
Improvedata structure complexityVSAvoidsound source position accuracy
Core Design Contradiction:
Device complexityVSMeasurement precision

Solution Approach 1:

The patent nests height information data structures within the existing MPEG-AAC encoding framework. The height_extension_element is embedded in the comment field of the ProgramConfigElement, and height information is nested within speaker arrangement data structures. This nested approach allows vertical positioning information to be integrated into the existing horizontal arrangement structure without creating a completely separate data system, thus maintaining relative simplicity while enhancing precision.

Inventive Principle:
Principle #7Nested doll (Nesting)

3Adaptability or versatility

If height information is added to speaker arrangement data, then vertical direction sound reproduction is enabled, but data transmission volume increases

Engineering Contradiction:
Improvespatial audio reproduction capabilityVSAvoiddata transmission volume
Core Design Contradiction:
Adaptability or versatilityVSQuantity of substance

Solution Approach 1:

The patent extracts only the essential height information needed for spatial positioning and stores it in a compact format within the height_extension_element. Rather than transmitting complete three-dimensional coordinates for all speakers, the system extracts and transmits only the vertical position parameters (height info values) that are necessary for reproducing the spatial audio effect, thereby minimizing the increase in data transmission volume.

Inventive Principle:
Principle #2Taking out (Extraction)

Data Source

PatentEP2741284B1Decoding device and method, encoding device and method, and program
Publication Date: 2020.04.22 SONY GROUP CORP
  • EP2741284B1 patent drawingFigure 1
  • EP2741284B1 patent drawingFigure 2
  • EP2741284B1 patent drawingFigure 3

AI summary

The present technique relates to a decoding device, a decoding method, an encoding device, an encoding method, and a program which can obtain a high-quality realistic sound. The encoding device stores speaker arrangement information in a comment region in a PCE of an encoded bit stream and stores a synchronous word and identification information in the comment region such that other public comments and the speaker arrangement information stored in the comment region can be distinguished from each other. When an encoded bit stream is decoded, it is determined whether the speaker arrangement information is stored on the basis of the synchronous word and the identification information stored in the comment region. Audio data included in the encoded bit stream is output according to the arrangement of the speakers corresponding to the determination result. The present technique can be applied to an encoding device.