Multisignal Audio Encoding with Whitening and Adaptive Joint Processing
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing multi-channel audio coding technologies are inflexible and inefficient, particularly for immersive 3D audio formats, as they fail to adaptively process arbitrary channel setups and require significant bit allocation for prediction coefficients, leading to suboptimal encoding efficiency and perceptual quality.
Innovation Solution
A multi-signal encoding and decoding system that preprocesses audio signals to be perceptually whitened and energy normalized, followed by adaptive joint signal processing, including band-wise M/S transform decisions based on estimated bitrate, to enhance encoding efficiency and quality.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If signal whitening and energy normalization are applied before joint signal processing, then encoding efficiency and perceptual quality are improved, but device complexity and processing overhead increase
Solution Approach 1:
The patent applies signal whitening and energy normalization as preliminary preprocessing steps before joint signal processing. This prepares the input signals in advance to have flattened spectral characteristics and uniform energy distribution, which optimizes the performance of subsequent adaptive joint signal processing operations without requiring complex real-time adjustments during encoding.
2Adaptability or versatility
If adaptive joint signal processing is performed on whitened signals with arbitrary channel configurations, then flexibility and encoding efficiency are improved, but bit allocation complexity increases
Solution Approach 1:
The patent implements adaptive joint signal processing that dynamically adjusts processing parameters based on the specific characteristics of whitened input signals and arbitrary channel configurations. The system adapts its processing strategy frame-by-frame and channel-pair-by-channel-pair, optimizing encoding efficiency for diverse immersive audio formats while managing bit allocation through systematic evaluation of signal correlations and processing requirements.
3Productivity
If band-wise M/S transform decisions are made based on estimated bitrate, then encoding efficiency is optimized, but processing complexity and computational load increase
Solution Approach 1:
The patent employs band-wise M/S transform decisions that dynamically select between different transform types (M/S or L/R) for each frequency band based on estimated bitrate requirements and signal characteristics. This parameter-based adaptation allows the encoder to optimize encoding efficiency by applying the most suitable transform strategy to each band while systematically managing computational complexity through structured decision-making.
4Loss of information
If residual signals are transmitted individually without joint stereo coding, then signal independence is maintained, but transmission data volume and encoding efficiency worsen
Solution Approach 1:
The patent applies joint stereo coding to combine correlated channel pairs into mid and side signals, thereby reducing the overall transmission data volume while preserving essential signal information. By identifying and processing correlated channels together through M/S transforms, the system achieves more efficient encoding compared to transmitting each channel independently, particularly for immersive audio formats with multiple correlated channels.
Data Source
Figure 1
Figure 2
Figure 3
AI summary
A multisignal encoder for encoding at least three audio signals, comprises: a signal preprocessor (100) for individually preprocessing each audio signal to obtain at least three preprocessed audio signals, wherein the preprocessing is performed so that a preprocessed audio signal is whitened with respect to the signal before preprocessing; an adaptive joint signal processor (200) for performing a processing of the at least three preprocessed audio signals to obtain at least three jointly processed signals or at least two jointly processed signals and an unprocessed signal; a signal encoder (300) for encoding each signal to obtain one or more encoded signals; and an output interface (400) for transmitting or storing an encoded multisignal audio signal comprising the one or more encoded signals, side information relating to the preprocessing and side information relating to the processing.