Joint Audio Source Coding with Statistical Cue Reconstruction
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing audio coding methods require high bitrate for transmitting multiple source signals, especially in complex scenes, and are limited by predefined playback formats, which restrict user flexibility and audio quality.
Innovation Solution
A joint-coding scheme that encodes multiple audio source signals by transmitting only their sum and low-bitrate side information, using statistical parameters to synthesize cues for decoding into stereo or multi-channel formats, allowing flexible playback and improved audio quality.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If separate mono audio coders are used for each source signal as in ISO/IEC MPEG-4, then high audio quality is achieved for each source, but the bitrate becomes very high when multiple sources are transmitted
Solution Approach 1:
The patent merges multiple source signals into a single mixed signal for transmission, rather than transmitting each source signal separately. This combining approach reduces the total bitrate required while maintaining the ability to reproduce individual sources at the decoder through parametric information about source positions and mixing coefficients
Solution Approach 2:
The patent transmits parametric information (mixing coefficients, source positions, spatial parameters) instead of complete waveform data for each source. This parameter-based representation dramatically reduces bitrate while preserving the ability to reconstruct the spatial audio scene at the decoder
2Quantity of substance
If Binaural Cue Coding is used to reduce bitrate by transmitting only the sum of source signals plus side information, then low bitrate is achieved, but the source signals cannot be recovered at the decoder and audio quality decreases as the number of sources increases
Solution Approach 1:
The patent incorporates feedback mechanisms where the decoder uses transmitted parametric information about source signals and mixing coefficients to reconstruct individual source signals from the mixed signal. This feedback loop enables accurate source recovery despite transmitting only the sum signal plus parameters
Solution Approach 2:
The patent performs preliminary encoding of parametric information (source characteristics, spatial parameters, mixing coefficients) at the encoder side before transmission. This preliminary preparation enables the decoder to accurately reconstruct sources without needing to transmit complete source waveforms
3Device complexity
If predefined playback formats are used as in most known methods, then the coding scenario is simplified, but user flexibility is restricted and binding to a predefined playback scenario occurs
Solution Approach 1:
The patent implements dynamic playback format adaptation where the system can switch between different playback configurations (stereo, multi-channel, binaural) based on the capabilities of the playback system. The parametric representation allows flexible adaptation to various output formats without requiring predefined format constraints
Data Source
AI summary
The following coding scenario is addressed: A number of audio source signals need to be transmitted or stored for the purpose of mixing wave field synthesis, multi-channel surround, or stereo signals after decoding the source signals. The proposed technique offers significant coding gain when jointly coding the source signals, compared to separately coding them, even when no redundancy is present between the source signals. This is possible by considering statistical properties of the source signals, the properties of mixing techniques, and spatial hearing. The sum of the source signals is transmitted plus the statistical properties of the source signals which mostly determine the perceptually important spatial cues of the final mixed audio channels. Source signals are recovered at the receiver such that their statistical properties approximate the corresponding properties of the original source signals. Subjective evaluations indicate that high audio quality is achieved by the proposed scheme.


