Joint Audio Source Coding Using Statistical Spatial Cues

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing audio coding methods require high bitrate for transmitting multiple source signals, especially in complex scenes, and are limited by predefined playback formats, which restrict user flexibility and audio quality.

Innovation Solution

A joint-coding scheme that encodes multiple audio source signals by transmitting their sum and statistical side information, allowing for flexible playback formats and improved audio quality by synthesizing cues at the decoder, such as inter-channel time difference, level difference, and coherence, to generate independent audio channels.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If separate mono audio coders are used for each source signal, then audio quality is maintained, but bitrate becomes high and scales with the number of sources

Engineering Contradiction:
Improveaudio qualityVSAvoidbitrate
Core Design Contradiction:
Measurement precisionVSQuantity of substance

Solution Approach 1:

The patent combines multiple source signals into a single mixed audio signal for encoding, rather than encoding each source separately. This merging approach allows the system to transmit one encoded audio stream while using parametric side information to represent multiple sources, thereby reducing overall bitrate while preserving the ability to reconstruct individual sources at the decoder

Inventive Principle:
Principle #5Merging (Combining)

Solution Approach 2:

The patent transforms the representation of multiple audio sources from time-domain waveforms to parametric descriptions. By encoding statistical parameters (mean, covariance, eigenvalues) that characterize the relationships between sources, the system achieves efficient compression while maintaining the ability to reconstruct source signals through parameter-based synthesis at the decoder

Inventive Principle:
Principle #35Parameter changes

2Quantity of substance

If Binaural Cue Coding is used to reduce bitrate, then bandwidth is reduced, but source signals cannot be recovered and audio quality decreases with more sources

Engineering Contradiction:
ImprovebitrateVSAvoidaudio quality
Core Design Contradiction:
Quantity of substanceVSMeasurement precision

Solution Approach 1:

The patent incorporates feedback mechanisms where the decoder uses transmitted parametric information about source relationships (covariance matrices, eigenvalues) to adaptively reconstruct individual source signals from the mixed audio. This feedback loop allows the system to maintain audio quality by dynamically adjusting source separation based on the statistical parameters received, preventing quality degradation as the number of sources increases

Inventive Principle:
Principle #23Feedback

Solution Approach 2:

The patent performs preliminary encoding of statistical parameters at the encoder side that describe the relationships between multiple sources. By pre-computing and transmitting these parametric descriptions (mean vectors, covariance matrices, eigenvalue decompositions), the decoder is equipped with the necessary information to accurately reconstruct sources, preventing quality loss before the reconstruction process begins

Inventive Principle:
Principle #10Preliminary action

3Device complexity

If predefined playback formats are used, then coding is simplified, but user flexibility is restricted

Engineering Contradiction:
Improvecoding complexityVSAvoidplayback flexibility
Core Design Contradiction:
Device complexityVSAdaptability or versatility

Solution Approach 1:

The patent creates a universal coding framework that can represent multiple audio sources in a format-independent manner. By encoding parametric information about source relationships rather than format-specific mixing data, the system enables flexible rendering to various playback configurations (stereo, surround, spatial audio) without requiring format-specific encoding, thus achieving both simplicity and versatility

Inventive Principle:
Principle #6Universality (Multi-functionality)

Solution Approach 2:

The patent introduces dynamic adaptability by allowing the decoded source signals to be rendered in different playback formats based on user preference and system capabilities. The parametric representation enables real-time adjustment of source positions, panning, and spatial distribution, transforming a static coding approach into a dynamic system that adapts to various playback scenarios

Inventive Principle:
Principle #15Dynamics

Data Source

PatentUS11495239B2Parametric joint-coding of audio sources
Publication Date: 2022.11.08 FRAUNHOFER GESELLSCHAFT ZUR FORDERUNG DER ANGEWANDTEN FORSCHUNG EV
  • US11495239B2 patent drawing
  • US11495239B2 patent drawing
  • US11495239B2 patent drawing

AI summary

The following coding scenario is addressed: A number of audio source signals need to be transmitted or stored for the purpose of mixing wave field synthesis, multi-channel surround, or stereo signals after decoding the source signals. The proposed technique offers significant coding gain when jointly coding the source signals, compared to separately coding them, even when no redundancy is present between the source signals. This is possible by considering statistical properties of the source signals, the properties of mixing techniques, and spatial hearing. The sum of the source signals is transmitted plus the statistical properties of the source signals, which mostly determine the perceptually important spatial cues of the final mixed audio channels. Source signals are recovered at the receiver such that their statistical properties approximate the corresponding properties of the original source signals. Subjective evaluations indicate that high audio quality is achieved by the proposed scheme.