Ambisonic Audio Coding via Spatial Segmentation and Gain-Shape Analysis

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current psychoacoustic audio coding techniques face challenges in efficiently encoding and decoding ambisonic audio data, particularly in effectively representing spatial characteristics and energy distribution within soundfields, which affects the quality and efficiency of audio transmission and storage.

Innovation Solution

The proposed solution involves a device and method for encoding and decoding ambisonic audio data using spatial audio encoding to separate foreground and background components, followed by gain and shape analysis to encode these components, and specifying them in a bitstream, allowing for efficient transmission and reconstruction of the audio data.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Loss of information

If traditional psychoacoustic audio coding is applied to ambisonic audio data, then compression is achieved, but spatial characteristics and energy distribution are not effectively represented

Engineering Contradiction:
Improvespatial characteristics representationVSAvoidencoding efficiency
Core Design Contradiction:
Loss of informationVSProductivity

Solution Approach 1:

The ambisonic audio data is segmented into foreground audio signals and background audio signals using spatial audio encoding. The foreground signals are further processed through gain and shape analysis to separate amplitude information from spectral information. This segmentation allows different processing strategies for different components, improving both spatial representation and encoding efficiency.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent transforms the audio data from the traditional time-domain representation into a multi-dimensional representation that includes spatial components, gain components, and shape components. This dimensional transformation enables more effective compression while preserving spatial characteristics by exploiting the structure of the transformed domain.

Inventive Principle:
Principle #17Another dimension (Dimensionality change)

2Productivity

If ambisonic audio data is compressed to reduce storage and transmission requirements, then efficiency improves, but audio quality may deteriorate

Engineering Contradiction:
Improvecompression efficiencyVSAvoidaudio quality
Core Design Contradiction:
ProductivityVSReliability

Solution Approach 1:

The patent applies parameter changes by transforming the audio data through spatial audio encoding, gain analysis, and shape analysis. These transformations recode the audio in terms of new parameters (spatial components, gain, shape) that are more suitable for compression. The inverse transformations during decoding restore the original audio quality, achieving both compression efficiency and quality preservation.

Inventive Principle:
Principle #35Parameter changes

3Measurement precision

If spatial audio encoding is performed to separate foreground and background components, then spatial characteristics are improved, but computational resources increase

Engineering Contradiction:
Improvespatial characteristics representationVSAvoidcomputational resources
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The patent applies partial action by performing gain and shape analysis only on the foreground audio signals rather than processing all audio data uniformly. This selective processing reduces computational complexity while maintaining high spatial characteristics representation for the most important audio components.

Inventive Principle:
Principle #16Partial or excessive action

Data Source

PatentUS12073842B2Psychoacoustic audio coding of ambisonic audio data
Publication Date: 2024.08.27 QUALCOMM INC
  • US12073842B2 patent drawing
  • US12073842B2 patent drawing
  • US12073842B2 patent drawing

AI summary

In general, techniques are described for psychoacoustic audio coding of ambisonic audio data. A device comprising a memory and one or more processors may be configured to perform the techniques. The memory may store the bitstream that includes an encoded audio object and a corresponding spatial component that defines spatial characteristics of the encoded foreground audio signal. The encoded foreground audio signal may include a coded gain and a coded shape. The one or more processors may perform a gain and shape synthesis with respect to the coded gain and the coded shape to obtain a foreground audio signal, and reconstruct, based on the foreground audio signal and the spatial component, the ambisonic audio data.