Audio Encoder Block Switching for Pre-Echo Reduction
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing audio codec technologies fail to effectively control or eliminate pre-echo artifacts, especially in audio with sharp impulses and transient signals, due to inaccuracies during time domain to frequency domain transformation and back, leading to suboptimal sound quality in IP-based multimedia applications.
Innovation Solution
An encoder system that identifies pre-echo events using timing information and PQMF output thresholds, encodes transformed and quantized error signals, and adjusts MDCT block sizes and scale factors to minimize or eliminate pre-echo effects, while maintaining a perceptually lossless or near-lossless audio encoding and decoding process.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Productivity
If MDCT transform is used to convert time domain signal to frequency domain and back, then audio compression is achieved, but pre-echo artifacts are introduced due to error spreading across block size
Solution Approach 1:
The audio signal is divided into multiple blocks of different sizes (e.g., 256, 512, 1024, 2048 samples). By segmenting the signal into variable block sizes rather than using a fixed block size, the patent enables more precise localization of transient signals and reduces the spreading of quantization errors across the entire block, thereby minimizing pre-echo artifacts while maintaining compression efficiency.
Solution Approach 2:
The patent dynamically adjusts MDCT block sizes based on the characteristics of the audio signal being processed. For transient signals with sharp impulses, smaller block sizes are used to localize errors; for steady-state signals, larger block sizes are used to maintain compression efficiency. This dynamic adaptation resolves the contradiction between compression efficiency and pre-echo reduction.
2Ease of operation
If fixed block size MDCT is used for encoding, then processing simplicity is maintained, but pre-echo control is ineffective especially for transient signals
Solution Approach 1:
The system transitions from fixed block size to dynamic block size selection, where the encoder automatically chooses appropriate block sizes based on signal characteristics. This maintains ease of operation through automated decision-making while effectively controlling pre-echo artifacts by adapting block sizes to match the temporal characteristics of the audio signal.
Solution Approach 2:
The patent changes the parameter of block size from a fixed value to a variable parameter that adapts to the input signal. By modifying this key parameter based on signal analysis, the system achieves effective pre-echo control for transient signals while maintaining processing simplicity through algorithmic automation.
3Quantity of substance
If quantization is applied during MDCT transformation, then bit rate reduction is achieved, but audio quality deteriorates due to error spreading across the block
Solution Approach 1:
By segmenting the audio signal into variable-sized blocks, the patent limits the spread of quantization errors to smaller regions. This allows aggressive quantization (lower bit rate) in regions where it matters less while maintaining higher quality in regions with transient signals, thus resolving the contradiction between bit rate reduction and audio quality preservation.
Data Source
Figure 1A~1B
Figure 2
Figure 3
AI summary
An encoder operable to filter audio signals into a plurality of frequency band components, generate quantized digital components for each band, identify a potential for pre-echo events within the generated quantized digital components, generate an approximate signal by decoding the quantized digital components using inverse pulse code modulation, generate an error signal by comparing the approximate signal with the sampled audio signal, and process the error signal and quantized digital components. The encoder operable to process the error signal by processing delayed audio signals and Q band values, determining the potential for pre-echo events from the Q band values, and determining scale factors and MDCT block sizes for the potential for pre-echo events. The encoder operable to transform the error signal into high resolution frequency components using the MDCT block sizes, quantize the scale factors and frequency components, and encode the quantized lines, block sizes, and quantized scale factors for inclusion in the bitstream.