Audio Encoding Bandwidth Extension Using Spectral Correlation

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing audio compression techniques face challenges in reducing transmission bandwidth without degrading the quality of reconstructed audio signals, particularly in accurately modeling wideband audio signals and accounting for phenomena like Comodulation Release of Masking (CMR).

Innovation Solution

The method involves transforming audio signals into basic and extended transform coefficients, correlating them using primary frequency scaling and translation parameters to increase correlation, and employing algorithms like Accurate Spectral Replacement (ASR) and Fractal Self-Similarity Model (FSSM) for bandwidth extension, along with a psychoacoustic model that accounts for CMR.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Quantity of substance

If bandwidth extension techniques are used to restore high-frequency components from base band samples, then transmission bandwidth is reduced, but the quality of reconstructed signal may be degraded

Engineering Contradiction:
Improvetransmission bandwidthVSAvoidquality of reconstructed signal
Core Design Contradiction:
Quantity of substanceVSManufacturing precision

Solution Approach 1:

The patent creates a copy of the base band spectral components and applies frequency scaling to generate extended high-frequency components. The decoded base band signal is transformed and scaled to reconstruct high-frequency content, effectively creating a copied and transformed version of the original signal to populate the extended frequency range.

Inventive Principle:
Principle #26Copying

Solution Approach 2:

The patent changes the frequency parameter of the base band components by applying a frequency scaling factor. The decoded spectral components are transformed with a scaling operation that maps base band frequencies to extended high-frequency ranges, thereby changing the frequency parameter to reconstruct the full bandwidth signal.

Inventive Principle:
Principle #35Parameter changes

2Productivity

If traditional psychoacoustic modeling is used for audio coding, then coding efficiency is improved, but accuracy in modeling wideband signals and CMR phenomenon is insufficient

Engineering Contradiction:
Improvecoding efficiencyVSAvoidaccuracy of psychoacoustic model
Core Design Contradiction:
ProductivityVSMeasurement precision

Solution Approach 1:

The patent applies different masking models for different frequency regions. Separate masking thresholds are calculated for base band and extended high-frequency bands, with the extended band masking model specifically designed to account for CMR effects. This local differentiation improves modeling accuracy for wideband signals while maintaining coding efficiency.

Inventive Principle:
Principle #3Local quality

Solution Approach 2:

The patent incorporates feedback by using the decoded base band signal to generate extended components, which are then combined and re-encoded. The process includes iterative refinement where the reconstructed signal quality informs the psychoacoustic model adjustments, particularly in modeling CMR effects based on the actual decoded content.

Inventive Principle:
Principle #23Feedback

Data Source

PatentUS7953605B2Method and apparatus for audio encoding and decoding using wideband psychoacoustic modeling and bandwidth extension
Publication Date: 2011.05.31 AUDIO TECH & CODECS
  • US7953605B2 patent drawing
  • US7953605B2 patent drawing
  • US7953605B2 patent drawing

AI summary

A novel bandwidth extension technique allows information to be encoded and decoded using a fractal self similarity model or an accurate spectral replacement model, or both. Also a multi-band temporal amplitude coding technique, useful as an enhancement to any coding/decoding technique, helps with accurate reconstruction of the temporal envelope and employs a utility filterbank. A perceptual coder using a comodulation masking release model, operating typically with more conventional perceptual coders, makes the perceptual model more accurate and hence increases the efficiency of the overall perceptual coder.