Sub-Band Audio Coding with Dynamic Blocks for Pre-Echo Control
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing audio codec technologies fail to effectively control or eliminate pre-echo artifacts, particularly noticeable in audio with sharp impulses and transient signals, due to inaccuracies during time-domain to frequency-domain transformations and back, leading to suboptimal sound quality in dynamic and distributed IP-based multimedia systems.
Innovation Solution
A computer-implemented system and method that encodes sampled audio signals by identifying potential pre-echo events, generating an error signal, and encoding this information into a bitstream, allowing for the removal of pre-echo artifacts during downstream decoding, using techniques such as pulse code modulation, adaptive pulse code modulation, and modified discrete cosine transform (MDCT) block sizes.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Productivity
If MDCT transform is used to convert time domain signal to frequency domain and back, then audio compression is achieved, but pre-echo artifacts are introduced due to error spreading across block size
Solution Approach 1:
The audio signal is divided into multiple blocks for MDCT processing, and the patent identifies specific blocks containing transient signals. By segmenting the processing approach - using different block sizes (smaller blocks) for transient-containing segments and larger blocks for steady-state segments - the patent prevents error spreading across entire large blocks, thereby reducing pre-echo artifacts while maintaining compression efficiency.
Solution Approach 2:
The patent dynamically adjusts MDCT block sizes based on the content of the audio signal. When transient signals are detected, smaller block sizes are used to limit the temporal spread of quantization errors. This dynamic adaptation allows the system to maintain high compression efficiency for steady-state signals while minimizing pre-echo artifacts during transient portions.
2Quantity of substance
If quantization is applied during MDCT transformation, then bit rate reduction is achieved, but pre-echo artifacts become noticeable especially in transient signals
Solution Approach 1:
The patent applies different quantization strategies to different portions of the audio signal. For blocks containing transient signals, finer quantization (higher precision) is applied locally to prevent the amplification of quantization errors that would cause pre-echo. For steady-state portions, coarser quantization can be used to achieve bit rate reduction. This local differentiation allows bit rate optimization without introducing noticeable pre-echo artifacts.
3Device complexity
If fixed block size MDCT is used, then encoding simplicity is maintained, but quality of service deteriorates in distributed IP networks with latency issues
Solution Approach 1:
The patent implements dynamic block size selection based on signal analysis, which increases encoding complexity but significantly improves quality of service in distributed networks. By adapting block sizes to signal characteristics, the system can better handle latency variations and packet loss in IP networks, ensuring more reliable audio delivery.
Solution Approach 2:
The patent changes the MDCT block size parameter dynamically based on the detected signal characteristics. This parameter adaptation allows the encoding system to optimize performance for different network conditions and signal types, improving reliability in distributed IP networks while maintaining reasonable encoding complexity through automated detection algorithms.
4Object-affected harmful factors
If error signal processing is added to identify and remove pre-echo events, then sound quality is enhanced, but device complexity increases
Solution Approach 1:
The patent performs preliminary analysis of the audio signal to identify blocks containing transient signals before applying MDCT transformation and quantization. By detecting potential pre-echo problems in advance, the system can pre-adjust block sizes and quantization parameters to prevent artifact formation, rather than requiring complex post-processing to remove artifacts after encoding.
Solution Approach 2:
The patent incorporates feedback mechanisms where the encoded audio is analyzed to detect pre-echo artifacts, and this information is used to adjust encoding parameters for subsequent processing. This feedback loop enables automatic optimization of encoding quality without requiring manual intervention, enhancing sound quality while keeping the system relatively simple through automated control.
Data Source
AI summary
A sub-band coder operable to process audio samples for use in a digital encoder. The sub-band coder comprising application code instructions executable on a processor configured to cause the coder to filter the audio samples into a plurality of frequency band components using at least one Pseudo-Quadrature Mirror Filter (PQMF) and modulate the plurality of frequency band components into a plurality of quantized band values using a pulse code modulation technique. The application code instructions further operable to cause the coder to decode the plurality of quantized band values into an approximation signal using an inverse pulse code modulation technique and at least one Inverse Pseudo-Quadrature Mirror Filter (IPQMF). The application code instructions operable to cause the coder generates an output for use by the digital encoder that includes the plurality of quantized band values, the approximation signal, and a plurality of encoded quantized band values.


