Joint Audio Source Coding Using Statistical Spatial Cues
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing audio coding methods require high bitrate for transmitting multiple source signals, especially in complex scenes, and are limited by predefined playback formats, which restrict user flexibility and audio quality.
Innovation Solution
A joint-coding scheme that encodes multiple audio source signals by transmitting their sum and statistical side information, allowing for flexible playback formats and improved audio quality by synthesizing cues at the decoder, such as inter-channel time difference, level difference, and coherence, to generate independent audio channels.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If separate mono audio coders are used for each source signal, then audio quality is maintained, but bitrate becomes high and scales with the number of sources
Solution Approach 1:
The patent combines multiple source signals into a single mixed audio signal for encoding, rather than encoding each source separately. This merging approach allows the system to transmit one encoded audio stream while using parametric side information to represent multiple sources, thereby reducing overall bitrate while preserving the ability to reconstruct individual sources at the decoder
Solution Approach 2:
The patent transforms the representation of multiple audio sources from time-domain waveforms to parametric descriptions. By encoding statistical parameters (mean, covariance, eigenvalues) that characterize the relationships between sources, the system achieves efficient compression while maintaining the ability to reconstruct source signals through parameter-based synthesis at the decoder
2Quantity of substance
If Binaural Cue Coding is used to reduce bitrate, then bandwidth is reduced, but source signals cannot be recovered and audio quality decreases with more sources
Solution Approach 1:
The patent incorporates feedback mechanisms where the decoder uses transmitted parametric information about source relationships (covariance matrices, eigenvalues) to adaptively reconstruct individual source signals from the mixed audio. This feedback loop allows the system to maintain audio quality by dynamically adjusting source separation based on the statistical parameters received, preventing quality degradation as the number of sources increases
Solution Approach 2:
The patent performs preliminary encoding of statistical parameters at the encoder side that describe the relationships between multiple sources. By pre-computing and transmitting these parametric descriptions (mean vectors, covariance matrices, eigenvalue decompositions), the decoder is equipped with the necessary information to accurately reconstruct sources, preventing quality loss before the reconstruction process begins
3Device complexity
If predefined playback formats are used, then coding is simplified, but user flexibility is restricted
Solution Approach 1:
The patent creates a universal coding framework that can represent multiple audio sources in a format-independent manner. By encoding parametric information about source relationships rather than format-specific mixing data, the system enables flexible rendering to various playback configurations (stereo, surround, spatial audio) without requiring format-specific encoding, thus achieving both simplicity and versatility
Solution Approach 2:
The patent introduces dynamic adaptability by allowing the decoded source signals to be rendered in different playback formats based on user preference and system capabilities. The parametric representation enables real-time adjustment of source positions, panning, and spatial distribution, transforming a static coding approach into a dynamic system that adapts to various playback scenarios
Data Source
AI summary
The following coding scenario is addressed: A number of audio source signals need to be transmitted or stored for the purpose of mixing wave field synthesis, multi-channel surround, or stereo signals after decoding the source signals. The proposed technique offers significant coding gain when jointly coding the source signals, compared to separately coding them, even when no redundancy is present between the source signals. This is possible by considering statistical properties of the source signals, the properties of mixing techniques, and spatial hearing. The sum of the source signals is transmitted plus the statistical properties of the source signals, which mostly determine the perceptually important spatial cues of the final mixed audio channels. Source signals are recovered at the receiver such that their statistical properties approximate the corresponding properties of the original source signals. Subjective evaluations indicate that high audio quality is achieved by the proposed scheme.


