Ambisonic Audio Coding via Spatial Segmentation and Gain-Shape Analysis
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current psychoacoustic audio coding techniques face challenges in efficiently encoding and decoding ambisonic audio data, particularly in effectively representing spatial characteristics and energy distribution within soundfields, which affects the quality and efficiency of audio transmission and storage.
Innovation Solution
The proposed solution involves a device and method for encoding and decoding ambisonic audio data using spatial audio encoding to separate foreground and background components, followed by gain and shape analysis to encode these components, and specifying them in a bitstream, allowing for efficient transmission and reconstruction of the audio data.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Loss of information
If traditional psychoacoustic audio coding is applied to ambisonic audio data, then compression is achieved, but spatial characteristics and energy distribution are not effectively represented
Solution Approach 1:
The ambisonic audio data is segmented into foreground audio signals and background audio signals using spatial audio encoding. The foreground signals are further processed through gain and shape analysis to separate amplitude information from spectral information. This segmentation allows different processing strategies for different components, improving both spatial representation and encoding efficiency.
Solution Approach 2:
The patent transforms the audio data from the traditional time-domain representation into a multi-dimensional representation that includes spatial components, gain components, and shape components. This dimensional transformation enables more effective compression while preserving spatial characteristics by exploiting the structure of the transformed domain.
2Productivity
If ambisonic audio data is compressed to reduce storage and transmission requirements, then efficiency improves, but audio quality may deteriorate
Solution Approach 1:
The patent applies parameter changes by transforming the audio data through spatial audio encoding, gain analysis, and shape analysis. These transformations recode the audio in terms of new parameters (spatial components, gain, shape) that are more suitable for compression. The inverse transformations during decoding restore the original audio quality, achieving both compression efficiency and quality preservation.
3Measurement precision
If spatial audio encoding is performed to separate foreground and background components, then spatial characteristics are improved, but computational resources increase
Solution Approach 1:
The patent applies partial action by performing gain and shape analysis only on the foreground audio signals rather than processing all audio data uniformly. This selective processing reduces computational complexity while maintaining high spatial characteristics representation for the most important audio components.
Data Source
AI summary
In general, techniques are described for psychoacoustic audio coding of ambisonic audio data. A device comprising a memory and one or more processors may be configured to perform the techniques. The memory may store the bitstream that includes an encoded audio object and a corresponding spatial component that defines spatial characteristics of the encoded foreground audio signal. The encoded foreground audio signal may include a coded gain and a coded shape. The one or more processors may perform a gain and shape synthesis with respect to the coded gain and the coded shape to obtain a foreground audio signal, and reconstruct, based on the foreground audio signal and the spatial component, the ambisonic audio data.


