Multi-Channel Audio Encoding for Binaural Spatial Signals
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing audio encoding and decoding methods for multi-channel signals, such as 5.1 surround sound, require significant bit rate increases and complex processing, leading to increased computational demands and reduced user experience due to the need for dedicated decoders and complex signal processing, especially when attempting to provide a binaural virtual spatial experience.
Innovation Solution
An audio encoder and decoder system that down-mixes an M-channel audio signal to a stereo signal with associated parametric data, modifies the stereo signal using spatial parameter data for a binaural perceptual transfer function to generate a binaural virtual spatial signal, and encodes this signal with reduced complexity, allowing legacy decoders to provide a high-quality multi-channel decoding experience.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If multi-channel audio signals are encoded using traditional methods (e.g., MPEG2, Dolby Digital), then the audio quality and spatial experience are improved, but the bit rate increases significantly and device complexity increases
Solution Approach 1:
The patent extracts only the essential spatial parameters (inter-channel cross-correlation, power ratio) from the multi-channel signal and transmits them as side information, while the main audio content is transmitted as a down-mixed stereo signal. This separation allows legacy decoders to ignore the parameter data and process only the stereo signal, while advanced decoders can utilize the parameters for multi-channel reconstruction.
Solution Approach 2:
The patent introduces an intermediary representation - the down-mixed stereo signal with embedded spatial parameters - that serves as a bridge between legacy stereo decoders and advanced multi-channel decoders. This intermediary format ensures backward compatibility while enabling enhanced functionality for capable devices.
2Adaptability or versatility
If additional signals are encoded to extend stereo to multi-channel audio (e.g., MPEG2 method), then the spatial audio capability is improved, but the additional bit rate is significant (same order of magnitude as stereo signal)
Solution Approach 1:
The patent changes the representation from transmitting full multi-channel audio signals to transmitting compact spatial parameters (inter-channel cross-correlation and power ratio) that describe the spatial relationships. This parameter-based approach reduces the data quantity from full-channel audio to concise statistical descriptors, achieving multi-channel capability with minimal bit rate overhead.
3Adaptability or versatility
If matrixed-surround methods are used for backwards-compatible multi-channel transmission, then the compatibility is improved, but the spatial accuracy and binaural experience are reduced
Solution Approach 1:
The patent incorporates feedback from the decoded stereo signal by calculating spatial parameters (inter-channel cross-correlation, power ratio) from the actual decoded output rather than relying on pre-computed matrix coefficients. This feedback mechanism allows the system to adapt to the specific characteristics of the decoded signal and reconstruct the multi-channel signal with higher spatial accuracy.
4Device complexity
If legacy decoders are used for multi-channel signals, then the device complexity is reduced, but the spatial audio experience and user experience are degraded
Solution Approach 1:
The patent creates a universal bit stream format that can be processed by decoders of any capability level. Legacy decoders can process the down-mixed stereo signal for acceptable audio playback, while advanced decoders can utilize the embedded spatial parameters to reconstruct multi-channel signals for enhanced spatial audio experience. The same encoded signal serves multiple functions and decoder types.
Data Source
AI summary
An audio encoder comprises a multi-channel receiver which receives an M-channel audio signal where M>2. A down-mix processor down-mixes the M-channel audio signal to a first stereo signal and associated parametric data and a spatial processor modifies the first stereo signal to generate a second stereo signal in response to the associated parametric data and spatial parameter data for a binaural perceptual transfer function, such as a Head Related Transfer Function (HRTF). The second stereo signal is a binaural signal and may specifically be a (3D) virtual spatial signal. An output data stream comprising the encoded data and the associated parametric data is generated by an encode processor and an output processor. The HRTF processing may allow the generation of a (3D) virtual spatial signal by conventional stereo decoders. A multi-channel decoder may reverse the process of the spatial processor to generate an improved quality multi-channel signal.


