Multi-channel Audio Coding via Frequency Band Segmentation
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing parametric coding/decoding techniques for multi-channel audio signals face challenges in maintaining spatial coherence and introducing artifacts when converting between time-frequency and temporal domains, particularly in reconstructing the original sound scene with accurate sound source positioning.
Innovation Solution
A method that decomposes multi-channel audio signals into frequency bands, adapts sound source direction data to ensure minimum separation, determines a mixing matrix, and codes direction data separately to form a binary stream, allowing for improved spatial coherence and reduced artifacts during reconstruction.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Productivity
If multi-channel audio signals are converted to time-frequency space for matrixing, then compression rate is improved, but spatial coherence and sound source positioning accuracy deteriorate
Solution Approach 1:
The patent segments the multi-channel audio signal into multiple frequency bands before processing. By decomposing the signal in the frequency domain and processing each band separately, the method preserves spatial coherence information while enabling efficient compression. The segmentation allows independent handling of different frequency components, maintaining their spatial characteristics through separate mixing matrices.
Solution Approach 2:
The patent changes the representation parameters by extracting spatial parameters (inter-channel level differences, inter-channel time differences, and spatial covariance matrix parameters) from the time-frequency domain signal. These parameters are then used to reconstruct the multi-channel signal, allowing compression while preserving spatial coherence and positioning accuracy through parametric representation.
2Ease of operation
If signal switching from time-frequency space to temporal space is performed, then temporal signal transmission is achieved, but artifacts and defects are introduced
Solution Approach 1:
The patent performs preliminary action by determining separate mixing matrices for each frequency band before the final temporal signal reconstruction. By pre-calculating the mixing matrices in the frequency domain and applying them to the downmixed signal, the method avoids artifacts that would arise from direct time-domain processing, ensuring smooth transition back to temporal space without defects.
Solution Approach 2:
The patent introduces the frequency domain as an intermediary between the original multi-channel signal and the final temporal downmixed signal. By processing through frequency band decomposition, extracting spatial parameters, and reconstructing via inverse transform, the frequency domain acts as a mediator that preserves spatial coherence while enabling efficient compression and transmission without introducing artifacts.
3Device complexity
If conventional matrixing is applied without frequency band decomposition, then processing complexity is reduced, but spatial coherence between sum signal and multi-channel signal deteriorates
Solution Approach 1:
The patent applies segmentation by dividing the frequency spectrum into multiple bands and processing each band independently with its own mixing matrix. This segmentation preserves the spatial coherence characteristics of different frequency components, which have different spatial properties, while keeping the overall processing manageable through systematic decomposition and reconstruction.
Data Source
AI summary
A method is provided for coding a multi-channel audio signal representing a sound scene comprising a plurality of sound sources. The method comprises decomposing the multi-channel signal into frequency bands and the following performed per frequency band: obtaining data representative of the direction of the sound sources of the sound scene, selecting a set of sound sources constituting principal sources, adapting the data representative of the direction of the selected principal sources, as a function of restitution characteristics of the multi-channel signal, determining a matrix for mixing the principal sources as a function of the adapted data, matrixing the principal sources by the matrix determined so as to obtain a sum signal with a reduced number of channels and coding the data representative of the direction of the sound sources and forming a binary stream comprising the coded data, the binary stream being transmittable in parallel with the sum signal.


