Audio Scale Parameter Downsampling and Interpolation

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing audio codecs face challenges in achieving low bitrate without significant quality loss, particularly due to high bit requirements for scale factor encoding and high computational complexity, and lack flexibility in noise shaping.

Innovation Solution

An apparatus and method for encoding and decoding audio signals using downsampling and interpolation of scale parameters, which involves converting audio signals to a spectral representation, calculating a first set of scale parameters, downsampling to a second set with fewer parameters, and interpolating back to a higher number on the decoder side for fine scaling, utilizing non-linear frequency scaling and amplitude-related measures like the Bark scale.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Manufacturing precision

If a high number of scale factors are used for spectral noise shaping, then the quality of audio processing is improved, but the bitrate increases significantly

Engineering Contradiction:
Improvespectral processing qualityVSAvoidbitrate
Core Design Contradiction:
Manufacturing precisionVSQuantity of substance

Solution Approach 1:

The scale factors are segmented into two groups: a first set of scale factors is downsampled to create a second set with fewer parameters for transmission, while a third set of scale factors is generated through interpolation at the decoder. This segmentation allows the system to transmit fewer parameters while maintaining the ability to reconstruct a complete set for high-quality spectral processing.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

An interpolation process acts as an intermediary between the transmitted downsampled scale factors and the full set of scale factors needed for spectral processing. The interpolator reconstructs the third set of scale factors from the second set, enabling high-fidelity noise shaping without transmitting all necessary parameters explicitly.

Inventive Principle:
Principle #24Intermediary (Mediator)

2Quantity of substance

If LPC-based perceptual filtering is used for spectral noise shaping, then the bitrate is reduced, but the computational complexity increases

Engineering Contradiction:
ImprovebitrateVSAvoidcomputational complexity
Core Design Contradiction:
Quantity of substanceVSDevice complexity

Solution Approach 1:

The patent extracts the essential perceptual filtering function from the complex LPC-based approach and implements it directly in the spectral domain using simple scaling operations. By working with MDCT coefficients and applying scale factors directly in the frequency domain, the system eliminates the need for time-domain LPC estimation, autocorrelation computation, and LPC-to-LSF conversion, thereby reducing computational complexity while maintaining bitrate efficiency.

Inventive Principle:
Principle #2Taking out (Extraction)

Solution Approach 2:

The patent replaces the mechanical LPC processing chain (time-domain filtering, autocorrelation, Levinson-Durbin recursion, LSF conversion) with a simpler spectral domain approach using MDCT-based filtering and direct scale factor application. This substitution maintains the perceptual noise shaping effect while dramatically reducing the computational operations required.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

3Quantity of substance

If LPC-based perceptual filtering is used for spectral noise shaping, then the bitrate is reduced, but the flexibility in noise shaping is reduced

Engineering Contradiction:
ImprovebitrateVSAvoidnoise shaping flexibility
Core Design Contradiction:
Quantity of substanceVSAdaptability or versatility

Solution Approach 1:

The patent implements dynamic noise shaping by allowing independent modification of scale factors in different frequency regions. The non-uniform scaling of MDCT coefficients and the ability to apply different scale factors to different scale factor bands enable adaptive noise shaping that can be tuned for specific audio content types, replacing the fixed LPC-based perceptual filter with a flexible spectral domain approach.

Inventive Principle:
Principle #15Dynamics

4Quantity of substance

If downsampled scale parameters are transmitted, then the bitrate is reduced, but the measurement precision of spectral representation is reduced

Engineering Contradiction:
ImprovebitrateVSAvoidspectral parameter precision
Core Design Contradiction:
Quantity of substanceVSMeasurement precision

Solution Approach 1:

The patent performs preliminary downsampling of the scale factors before transmission, creating a compressed representation that is then expanded back to full resolution through interpolation at the decoder. This preliminary action allows the system to transmit fewer bits while preserving the ability to reconstruct precise spectral parameters for high-quality audio processing.

Inventive Principle:
Principle #10Preliminary action

Data Source

PatentEP4375995B1Apparatus and method for encoding and decoding an audio signal using downsampling or interpolation of scale parameters
Publication Date: 2025.06.25 FRAUNHOFER GESELLSCHAFT ZUR FORDERUNG DER ANGEWANDTEN FORSCHUNG EV
  • EP4375995B1 patent drawingFigure 1
  • EP4375995B1 patent drawingFigure 2
  • EP4375995B1 patent drawingFigure 3~4

AI summary

An apparatus for encoding an audio signal (160), comprises: a converter (100) for converting the audio signal into a spectral representation; a scale parameter calculator (110) for calculating a first set of scale parameters from the spectral representation: a downsampler (130) for downsampling the first set of scale parameters to obtain a second set of scale parameters, wherein a second number of scale parameters in the second set of scale parameters is lower than a first number of scale parameters in the first set of scale parameters; a scale parameter encoder (140) for generating an encoded representation of the second set of scale parameters; a spectral processor (120) for processing the spectral representation using a third set of scale parameters, the third set of scale parameters having a third number of scale parameters being greater than the second number of scale parameters, wherein the spectral processor (120) is configured to use the first set of scale parameters or to derive the third set of scale parameters from the second set of scale parameters or from the encoded representation of the second set of scale parameters using an interpolation operation; and an output interface (150) for generating an encoded output signal (170) comprising information on the encoded representation of the spectral representation and information on the encoded representation of the second set of scale parameters.