Harmonic Audio Transform Encoding via Spectral Peak Quantization

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Conventional transform encoding methods fail to produce audio signals of acceptable quality for harmonic audio signals, as they do not effectively handle the normalization and residual encoding of such signals, especially at lower bitrates.

Innovation Solution

A transform encoding/decoding scheme that locates spectral peaks and encodes peak regions, low-frequency coefficients outside peak regions, and noise-floor gains for high-frequency coefficients, replacing the conventional spectrum envelope and residual model with a spectral peaks and noise-floor model.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Device complexity

If conventional transform encoding with spectrum envelope and residual model is used, then encoding complexity is reduced, but audio quality for harmonic signals deteriorates

Engineering Contradiction:
Improveencoding complexityVSAvoidaudio quality
Core Design Contradiction:
Device complexityVSManufacturing precision

Solution Approach 1:

The patent changes the encoding parameters from conventional spectrum envelope and residual coefficients to spectral peak parameters (position, magnitude, width) and noise-floor gain. This parameter transformation allows harmonic signals to be represented more accurately by focusing on dominant spectral features rather than attempting to encode the entire spectrum uniformly, thereby improving audio quality while maintaining manageable encoding complexity.

Inventive Principle:
Principle #35Parameter changes

Solution Approach 2:

The patent applies different encoding strategies to different parts of the spectrum: spectral peaks are encoded with high precision (position, magnitude, width parameters) while non-peak regions are represented by a single noise-floor gain value. This localized approach allocates encoding resources efficiently, providing high quality where it matters most (harmonic regions) while reducing complexity in less critical areas.

Inventive Principle:
Principle #3Local quality

2Measurement precision

If spectral peaks are directly quantized with position, magnitude and width parameters, then representation accuracy for harmonic signals is improved, but encoding complexity increases

Engineering Contradiction:
Improvespectral representation accuracyVSAvoidencoding complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The patent extracts only the most significant spectral features (peaks) from the full spectrum for encoding, rather than attempting to encode all frequency components. By identifying and encoding only peak positions, magnitudes, and widths along with noise-floor gain, the method captures the essential characteristics of harmonic signals while dramatically reducing the number of parameters that need to be encoded compared to conventional approaches.

Inventive Principle:
Principle #2Taking out (Extraction)

Solution Approach 2:

The patent segments the spectrum into peak regions and non-peak regions, applying different encoding methods to each segment. Peak regions are encoded with detailed parameters (position, magnitude, width) while non-peak regions use a simplified noise-floor gain representation. This segmentation allows the system to achieve high accuracy for harmonic content without the complexity of encoding the entire spectrum in detail.

Inventive Principle:
Principle #1Segmentation

Data Source

PatentUS20240321283A1Transform Encoding/Decoding of Harmonic Audio Signals
Publication Date: 2024.09.26 TELEFONAKTIEBOLAGET LM ERICSSON (PUBL)
  • US20240321283A1 patent drawing
  • US20240321283A1 patent drawing
  • US20240321283A1 patent drawing

AI summary

An encoder for encoding frequency transform coefficients of a harmonic audio signal include the following elements: A peak locator configured to locate spectral peaks having magnitudes exceeding a predetermined frequency dependent threshold. A peak region encoder configured to encode peak regions including and surrounding the located peaks. A low-frequency set encoder configured to encode at least one low-frequency set of coefficients outside the peak regions and below a crossover frequency that depends on the number of bits used to encode the peak regions. A noise-floor gain encoder configured to encode a noise-floor gain of at least one high-frequency set of not yet encoded coefficients outside the peak regions.