VBAP Gain Calculation via Multi-Metadata Encoding
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing audio encoding and decoding techniques, such as MPEG-H 3D Audio standards, face challenges in achieving high-quality sound reproduction due to interpolation-based VBAP gain calculations, which can lead to unstable sound image localization and inaccurate movement rendering, especially in scenes with discontinuous changes.
Innovation Solution
The proposed solution involves encoding and decoding multiple metadata per frame, allowing for more precise calculation of VBAP gains by distributing metadata across samples within a frame using methods like count designation, sample designation, and automatic switching, thereby reducing the segment length for interpolation and improving sound quality.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Loss of information
If only one metadata is encoded per frame (representative sample only), then the data transmission load is reduced, but the sound image localization stability deteriorates due to long interpolation segments
Solution Approach 1:
The frame is divided into multiple segments, with metadata being encoded at multiple segmentation points within the frame rather than only at the representative sample. This segmentation approach reduces the interpolation segment length while maintaining efficient data transmission by strategically placing metadata at key positions (e.g., every N samples or at scene change points).
Solution Approach 2:
Metadata is prepared and encoded at specific sample points within the frame in advance, particularly at positions that will minimize interpolation requirements. This preliminary placement of metadata at strategic points ensures that when decoding occurs, the interpolation segments are naturally shorter and more stable.
2Device complexity
If linear interpolation is used to calculate VBAP gains between frames, then the calculation complexity is reduced, but the accuracy of audio object movement rendering deteriorates in discontinuous scenes
Solution Approach 1:
The metadata encoding strategy dynamically adapts to scene characteristics. In discontinuous scenes or when scene changes are detected, additional metadata is encoded at more frequent intervals within frames to capture the abrupt changes. In stable scenes, the normal reduced metadata frequency is maintained, keeping calculation complexity low while ensuring accuracy when needed.
Solution Approach 2:
The system changes the metadata encoding frequency and positioning parameters based on scene analysis. When discontinuous movements or scene changes are detected, the parameter for metadata density increases, providing more reference points for accurate VBAP gain calculation without always maintaining high complexity.
3Ease of manufacture
If metadata is encoded only at the last sample of each frame, then the encoding process is simplified, but the VBAP gain calculation accuracy for intermediate samples deteriorates
Solution Approach 1:
Instead of treating the frame as a single unit with one representative sample, the frame is segmented into multiple sections, each with its own metadata encoding point. This segmentation maintains relative encoding simplicity while significantly improving VBAP gain accuracy for intermediate samples by reducing the maximum interpolation distance.
Data Source
AI summary
The present technology relates to an encoding apparatus, an encoding method, a decoding apparatus, a decoding method, and a program for obtaining sound of higher quality. An audio signal decoding section decodes encoded audio data to acquire an audio signal of each object. A metadata decoding section decodes encoded metadata to acquire a plurality of metadata about each object in each frame of the audio signal. A gain calculating section calculates VBAP gains of each object in the audio signal for each speaker based on the metadata. An audio signal generating section generates an audio signal to be fed to each speaker by having the audio signal of each object multiplied by the corresponding VBAP gain and by adding up the multiplied audio signals. The present technology may be applied to decoding apparatuses.


