Multichannel Audio Encoding via Frequency Band Directivity
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current parametric coding/decoding techniques for multichannel audio signals, such as Binaural Cue Coding, are inadequate for complex signals with phase oppositions and fail to fully exploit interchannel redundancies, leading to unsatisfactory quality, especially for surround sound signals.
Innovation Solution
A method that decomposes multichannel audio signals into frequency bands, obtains directivity information for each source, selects main sources, and codes this information separately to form a binary stream, allowing for better reconstruction of complex sound scenes and compatibility with existing decoders.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If conventional parametric coding techniques (Binaural Cue Coding) are used, then backward compatibility with existing broadcasting systems is ensured, but quality of complex multichannel signals with phase oppositions deteriorates
Solution Approach 1:
The patent segments the multichannel signal processing into separate frequency bands and identifies individual sound sources within each band. By decomposing the complex signal into manageable components (frequency bands × sound sources), the system can apply different processing strategies to each segment, preserving compatibility while improving quality for complex signals.
Solution Approach 2:
The patent changes the parameter representation from traditional spatial cues to directivity information per sound source. This includes parameters like directivity vectors, angular width, and spectral shape, which better characterize complex sound sources with phase oppositions while maintaining compatibility with existing decoding frameworks.
2Device complexity
If tree structure coding is used, then device complexity is reduced, but exploitation of interchannel redundancy is insufficient
Solution Approach 1:
The patent moves from a hierarchical tree structure to a matrix-based approach that operates across the dimension of frequency bands and sound sources simultaneously. This allows parallel processing of multiple channels and frequency bands, fully exploiting interchannel redundancies without increasing device complexity, as the matrix operations can be efficiently implemented.
3Manufacturing precision
If directivity information is coded separately per frequency band, then quality of sound scene reconstruction is improved, but bit rate increases
Solution Approach 1:
The patent applies partial coding by selecting only the most significant sound sources and frequency bands for detailed directivity coding. Less important sources or bands can use simplified or omitted directivity information, achieving good reconstruction quality while controlling bit rate through selective application of the full coding approach.
Data Source
Figure 1
Figure 2
Figure 3a~3b
AI summary
The invention relates to a method for encoding a multichannel audio signal representing a sound scene including a plurality of sound sources. Said method is characterised in that it comprises a step of decomposing (T) the multichannel signal into a frequency band, and the following steps of, by frequency band, obtaining (OBT) directivity information for each sound source of the sound scene, the information being representative of the space distribution of the sound source of the sound scene, selecting (Select) a set of sound sources in the sound scene defining the main sources, mastering (M) the selected main sources in order to obtain a sum signal with a reduced number of channels, encoding (Cod.Di) the directivity information, and forming (Con.Fb) a binary flow including the encoded directivity information, wherein the binary flow can be transmitted in parallel with the sum signal. The invention also relates to a decoding method for decoding the sum signal and the directivity information in order to obtain a multichannel signal, and to a related encoder and decoder.