VBAP Gain Calculation via Multi-Metadata Encoding

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing audio encoding and decoding techniques, such as MPEG-H 3D Audio standards, face challenges in achieving high-quality sound reproduction due to interpolation-based VBAP gain calculations, which can lead to unstable sound image localization and inaccurate movement rendering, especially in scenes with discontinuous changes.

Innovation Solution

The proposed solution involves encoding and decoding multiple metadata per frame, allowing for more precise calculation of VBAP gains by distributing metadata across samples within a frame using methods like count designation, sample designation, and automatic switching, thereby reducing the segment length for interpolation and improving sound quality.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Loss of information

If only one metadata is encoded per frame (representative sample only), then the data transmission load is reduced, but the sound image localization stability deteriorates due to long interpolation segments

Engineering Contradiction:
Improvedata transmission loadVSAvoidsound image localization stability
Core Design Contradiction:
Loss of informationVSReliability

Solution Approach 1:

The frame is divided into multiple segments, with metadata being encoded at multiple segmentation points within the frame rather than only at the representative sample. This segmentation approach reduces the interpolation segment length while maintaining efficient data transmission by strategically placing metadata at key positions (e.g., every N samples or at scene change points).

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

Metadata is prepared and encoded at specific sample points within the frame in advance, particularly at positions that will minimize interpolation requirements. This preliminary placement of metadata at strategic points ensures that when decoding occurs, the interpolation segments are naturally shorter and more stable.

Inventive Principle:
Principle #10Preliminary action

2Device complexity

If linear interpolation is used to calculate VBAP gains between frames, then the calculation complexity is reduced, but the accuracy of audio object movement rendering deteriorates in discontinuous scenes

Engineering Contradiction:
Improvecalculation complexityVSAvoidmovement rendering accuracy
Core Design Contradiction:
Device complexityVSMeasurement precision

Solution Approach 1:

The metadata encoding strategy dynamically adapts to scene characteristics. In discontinuous scenes or when scene changes are detected, additional metadata is encoded at more frequent intervals within frames to capture the abrupt changes. In stable scenes, the normal reduced metadata frequency is maintained, keeping calculation complexity low while ensuring accuracy when needed.

Inventive Principle:
Principle #15Dynamics

Solution Approach 2:

The system changes the metadata encoding frequency and positioning parameters based on scene analysis. When discontinuous movements or scene changes are detected, the parameter for metadata density increases, providing more reference points for accurate VBAP gain calculation without always maintaining high complexity.

Inventive Principle:
Principle #35Parameter changes

3Ease of manufacture

If metadata is encoded only at the last sample of each frame, then the encoding process is simplified, but the VBAP gain calculation accuracy for intermediate samples deteriorates

Engineering Contradiction:
Improveencoding process simplicityVSAvoidVBAP gain calculation accuracy
Core Design Contradiction:
Ease of manufactureVSMeasurement precision

Solution Approach 1:

Instead of treating the frame as a single unit with one representative sample, the frame is segmented into multiple sections, each with its own metadata encoding point. This segmentation maintains relative encoding simplicity while significantly improving VBAP gain accuracy for intermediate samples by reducing the maximum interpolation distance.

Inventive Principle:
Principle #1Segmentation

Data Source

PatentUS11170796B2Multiple metadata part-based encoding apparatus, encoding method, decoding apparatus, decoding method, and program
Publication Date: 2021.11.09 SONY GROUP CORP
  • US11170796B2 patent drawing
  • US11170796B2 patent drawing
  • US11170796B2 patent drawing

AI summary

The present technology relates to an encoding apparatus, an encoding method, a decoding apparatus, a decoding method, and a program for obtaining sound of higher quality. An audio signal decoding section decodes encoded audio data to acquire an audio signal of each object. A metadata decoding section decodes encoded metadata to acquire a plurality of metadata about each object in each frame of the audio signal. A gain calculating section calculates VBAP gains of each object in the audio signal for each speaker based on the metadata. An audio signal generating section generates an audio signal to be fed to each speaker by having the audio signal of each object multiplied by the corresponding VBAP gain and by adding up the multiplied audio signals. The present technology may be applied to decoding apparatuses.