Audio Codec Companding with Spectral Extension for Lower Coding Noise

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing audio codecs introduce noticeable distortion in the form of coding noise due to lossy compression techniques, particularly evident during low-intensity segments of audio signals, which can manifest as pre-echo artifacts, and current solutions like filters introduce phase distortion or reduce frequency resolution.

Innovation Solution

A companding technique is employed to process audio signals by dividing them into short time segments, calculating and applying wideband gains in the frequency domain to amplify low-intensity segments and attenuate high-intensity segments during compression, and inversely doing so during expansion to restore the original dynamic range, effectively reducing quantization noise.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Productivity

If lossy data compression techniques are used to reduce storage or data rate requirements, then the data rate is reduced, but the fidelity of source content deteriorates and coding noise is introduced

Engineering Contradiction:
Improvedata rateVSAvoidfidelity of source content
Core Design Contradiction:
ProductivityVSManufacturing precision

Solution Approach 1:

The audio signal is divided into multiple frames, and each frame is further divided into frequency subbands using a filter bank. This segmentation allows different quantization strategies to be applied to different frequency regions, enabling better quality at lower bitrates by preserving important frequency components while compressing less critical ones.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

Different quantization precision is applied to different frequency subbands based on their perceptual importance. High-frequency subbands use coarser quantization while low-frequency subbands use finer quantization, optimizing the trade-off between bitrate and perceived audio quality.

Inventive Principle:
Principle #3Local quality

2Object-affected harmful factors

If coding noise is shaped in the frequency domain over long frames, then the noise becomes less audible through masking effects, but pre-echo distortion occurs during low intensity parts of the frame

Engineering Contradiction:
Improveaudibility of coding noiseVSAvoidpre-echo distortion
Core Design Contradiction:
Object-affected harmful factorsVSObject-generated harmful factors

Solution Approach 1:

Each frame is divided into multiple frequency subbands using a filter bank. This allows independent noise shaping control for each subband, enabling the system to reduce pre-echo in transient regions while maintaining noise masking benefits in stationary regions.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The noise shaping filter characteristics are made adaptive and time-varying within each frame. The system dynamically adjusts the noise shaping parameters based on the local signal characteristics, allowing it to respond to transients and prevent pre-echo while maintaining effective noise masking during steady-state portions.

Inventive Principle:
Principle #15Dynamics

3Object-generated harmful factors

If filters are used to avoid pre-echo artifacts, then pre-echo is reduced, but phase distortion and temporal smearing are introduced

Engineering Contradiction:
Improvepre-echo artifactsVSAvoidphase accuracy
Core Design Contradiction:
Object-generated harmful factorsVSManufacturing precision

Solution Approach 1:

Different noise shaping strategies are applied to different frequency subbands. By operating in the frequency domain with subband-specific parameters, the system can reduce pre-echo in transient regions without applying broad spectral filtering that would cause phase distortion and temporal smearing across the entire signal.

Inventive Principle:
Principle #3Local quality

4Object-generated harmful factors

If smaller transform windows are used to reduce pre-echo, then transient response is improved, but frequency resolution is significantly reduced

Engineering Contradiction:
Improvetransient distortionVSAvoidfrequency resolution
Core Design Contradiction:
Object-generated harmful factorsVSMeasurement precision

Solution Approach 1:

The signal is segmented into frequency subbands using a filter bank with longer transform windows. This segmentation allows the system to achieve good frequency resolution in each subband while using adaptive noise shaping to handle transients, avoiding the need to reduce the overall window size and lose frequency resolution.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The problem is solved by moving from a single-dimensional time-domain approach to a two-dimensional frequency-subband domain. By transforming the problem into the frequency domain and applying subband-specific processing, the system can maintain long window lengths for frequency resolution while addressing transient issues through frequency-selective noise shaping.

Inventive Principle:
Principle #17Another dimension (Dimensionality change)

Data Source

PatentUS12175994B2Companding system and method to reduce quantization noise using advanced spectral extension
Publication Date: 2024.12.24 DOLBY LABORATORIES LICENSING CORP
  • US12175994B2 patent drawing
  • US12175994B2 patent drawing
  • US12175994B2 patent drawing

AI summary

Embodiments are directed to a companding method and system for reducing coding noise in an audio codec. A compression process reduces an original dynamic range of an initial audio signal through a compression process that divides the initial audio signal into a plurality of segments using a defined window shape, calculates a wideband gain in the frequency domain using a non-energy based average of frequency domain samples of the initial audio signal, and applies individual gain values to amplify segments of relatively low intensity and attenuate segments of relatively high intensity. The compressed audio signal is then expanded back to the substantially the original dynamic range that applies inverse gain values to amplify segments of relatively high intensity and attenuating segments of relatively low intensity. A QMF filterbank is used to analyze the initial audio signal to obtain a frequency domain representation.