Neural Audio Encoding with Whitened Spectrum Shaping

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Conventional audio encoding technologies suffer from low encoding quality due to large variations in statistical average energy distribution of whitened spectra, leading to loss of spectral features, particularly in frequency bands with low energy, during processing by neural networks.

Innovation Solution

The method involves shaping the whitened spectrum to increase spectral amplitudes in target frequency bands, reducing the dynamic range of energy distribution, thereby retaining more spectral features for improved encoding quality. This is achieved by determining and adjusting the target frequency band based on factors like sampling rate, channel quantity, encoding rate, and encoding mode, and using gain adjustment factors to modify spectral values.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Manufacturing precision

If conventional audio encoding is performed using whitened spectrum processing through neural networks, then encoding speed and basic functionality are maintained, but spectral features are lost due to large variations in energy distribution across frequency bands

Engineering Contradiction:
Improveencoding qualityVSAvoidspectral feature loss
Core Design Contradiction:
Manufacturing precisionVSLoss of information

Solution Approach 1:

The patent applies spectral shaping to modify the energy distribution parameters of the whitened spectrum. By adjusting the spectral amplitude in target frequency bands through shaping filters, the energy distribution is transformed to reduce dynamic range variations, thereby preventing spectral feature loss during neural network processing while maintaining encoding efficiency

Inventive Principle:
Principle #35Parameter changes

2Manufacturing precision

If the whitened spectrum is processed directly by the encoding neural network without spectral shaping, then processing complexity is reduced, but spectral lines in low energy regions are lost leading to poor encoding quality

Engineering Contradiction:
Improveencoding qualityVSAvoidprocessing complexity
Core Design Contradiction:
Manufacturing precisionVSDevice complexity

Solution Approach 1:

The patent performs spectral shaping as a preliminary action before the neural network encoding process. By pre-adjusting the energy distribution in the whitened spectrum to reduce dynamic range variations, the spectral features are preserved and made more suitable for subsequent neural network processing, thereby improving encoding quality without significantly increasing overall system complexity

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

Spectral shaping modifies the energy distribution parameters of the whitened spectrum by adjusting amplitudes in target frequency bands. This parameter transformation reduces the dynamic range of energy variations, ensuring that spectral lines in previously low-energy regions are enhanced and preserved during encoding

Inventive Principle:
Principle #35Parameter changes

Data Source

PatentEP4339945B1Encoding method and device, decoding method and device, and storage medium
Publication Date: 2026.04.22 HUAWEI TECH CO LTD
  • EP4339945B1 patent drawingFigure 1
  • EP4339945B1 patent drawingFigure 2~3
  • EP4339945B1 patent drawingFigure 4

AI summary

Embodiments of this application disclose an encoding method and apparatus, a decoding method and apparatus, a device, a storage medium, and a computer program, and belong to the field of encoding and decoding technologies. In embodiments of this application, a first whitened spectrum for media data is whitened to obtain a second whitened spectrum, and then encoding is performed based on the second whitened spectrum. A spectral amplitude of the second whitened spectrum in a target frequency band is greater than or equal to a spectral amplitude of the first whitened spectrum in the target frequency band. It can be learned that, in this solution, the spectral amplitude of the first whitened spectrum in the target frequency band is increased, so that a difference between statistical average energy of spectral lines for different frequencies in the obtained second whitened spectrum is small. In this way, in a process of processing the second whitened spectrum by using an encoding neural network model, more spectral lines in the second whitened spectrum can be retained. To be specific, in this solution, more spectral lines can be encoded, so that more spectral features are retained, and encoding quality is improved.