Perceptual Audio Coding Scale Factor Adaptation

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing perceptual coders are unreliable at low bit rates and have high complexity, leading to suboptimal audio quality and increased bit usage due to the inability to accurately model masking thresholds.

Innovation Solution

A method for perceptual transform coding that determines transform coefficients, masking thresholds, and scale factors for audio signals, adapting these factors to prevent energy loss in perceptually relevant sub-bands, enabling high-quality low bit rate coding with low complexity.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Reliability

If existing perceptual coders use complex auditory models to compute masking thresholds, then audio quality is improved, but device complexity and computational load increase significantly

Engineering Contradiction:
Improveaudio qualityVSAvoidcomplexity of psychoacoustical model
Core Design Contradiction:
ReliabilityVSDevice complexity

Solution Approach 1:

The patent segments the audio spectrum into multiple sub-bands and processes each sub-band independently with simplified masking threshold computations. This divides the complex global psychoacoustical model into smaller, more manageable local models that are computationally less intensive while maintaining overall audio quality.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent changes the parameters of the psychoacoustical model by using simplified masking threshold computation methods that rely on fewer and less complex parameters compared to traditional models. This reduces computational complexity while preserving the essential perceptual characteristics needed for high-quality audio coding.

Inventive Principle:
Principle #35Parameter changes

2Device complexity

If perceptual coders use simplified models to reduce complexity, then device complexity is reduced, but reliability and audio quality deteriorate at low bit rates

Engineering Contradiction:
Improvecomplexity of psychoacoustical modelVSAvoidaudio quality at low bit rates
Core Design Contradiction:
Device complexityVSReliability

Solution Approach 1:

The patent applies different levels of modeling complexity to different spectral regions (sub-bands) based on their perceptual importance. Perceptually critical sub-bands receive more sophisticated processing while less critical regions use simpler models, optimizing the trade-off between complexity and quality locally rather than globally.

Inventive Principle:
Principle #3Local quality

Solution Approach 2:

The patent introduces dynamic adaptation of coding parameters and bit allocation based on the local characteristics of each sub-band. This allows the system to dynamically adjust the level of detail and complexity applied to different parts of the spectrum, maintaining high quality where needed while reducing complexity where permissible.

Inventive Principle:
Principle #15Dynamics

3Reliability

If traditional perceptual coders allocate bits based on instantaneous masking thresholds, then audio quality is maintained, but bit rate efficiency decreases at low bit rates

Engineering Contradiction:
Improveaudio qualityVSAvoidbit rate efficiency
Core Design Contradiction:
ReliabilityVSProductivity

Solution Approach 1:

The patent performs preliminary analysis and planning of bit allocation across sub-bands before actual encoding, using simplified predictions of masking thresholds and perceptual importance. This preliminary action allows for more efficient global optimization of bit distribution, ensuring that bits are allocated to the most perceptually critical regions first, improving bit rate efficiency.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The patent uses simplified models and approximations (copies) of the complex psychoacoustical processes to guide bit allocation decisions. These simplified copies are computationally efficient and provide sufficient accuracy for making optimal bit allocation decisions, improving productivity without significantly compromising audio quality.

Inventive Principle:
Principle #26Copying

Data Source

PatentUS9153240B2Transform coding of speech and audio signals
Publication Date: 2015.10.06 TELEFONAKTIEBOLAGET LM ERICSSON (PUBL)
  • US9153240B2 patent drawing
  • US9153240B2 patent drawing
  • US9153240B2 patent drawing

AI summary

In a method of perceptual transform coding of audio signals in a telecommunication system, performing the steps of determining transform coefficients representative of a time to frequency transformation of a time segmented input audio signal; determining a spectrum of perceptual sub-bands for said input audio signal based on said determined transform coefficients; determining masking thresholds for each said sub-band based on said determined spectrum; computing scale factors for each said sub-band based on said determined masking thresholds, and finally adapting said computed scale factors for each said sub-band to prevent energy loss for perceptually relevant sub-bands.