Neural Network High Frequency Audio Reconstruction
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing high frequency reconstruction (HFR) systems for audio coding face challenges in accurately reconstructing high-frequency bands, especially at low bitrates, leading to single sideband distortion and a metallic, synthetic sound quality.
Innovation Solution
The use of a generative deep neural network operating in a filter bank domain to reconstruct high-frequency bands, where the neural network is trained to predict high-band audio signal samples given decoded low-band samples and high frequency reconstruction parameters, thereby synthesizing a time-domain audio signal.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If traditional HFR methods (copy-up or harmonic transposer) are used to reconstruct high-frequency bands at low bitrates, then computational complexity is reduced and robustness is improved, but single sideband distortion occurs and sound quality becomes metallic and synthetic
Solution Approach 1:
The patent replaces traditional mechanical signal processing methods (copy-up transposition, harmonic transposer with phase vocoder) with a neural network-based system. The neural network learns optimal high-frequency reconstruction by processing low-band filter bank domain samples and HFR parameters, substituting deterministic mechanical algorithms with adaptive learned models that avoid SSB distortion and metallic artifacts while maintaining robustness.
Solution Approach 2:
The patent changes the domain in which HFR is performed from time-domain or traditional frequency-domain methods to filter bank domain. By operating in the filter bank domain and using neural networks to predict high-band filter bank samples, the system transforms the reconstruction process to achieve better quality without the harmful effects of traditional methods.
2Manufacturing precision
If more bitrate is allocated to encode full bandwidth audio, then audio quality is improved, but transmission efficiency is reduced
Solution Approach 1:
The patent applies partial action by allocating limited bitrate resources to encode only the essential low-frequency bands with high quality, while the high-frequency bands are reconstructed using a neural network. This selective encoding approach achieves perceptually full bandwidth audio quality without transmitting all frequency components explicitly, thereby maintaining transmission efficiency.
Solution Approach 2:
The neural network acts as an intermediary that reconstructs the missing high-frequency information from the encoded low-band data and HFR parameters. Instead of directly transmitting full bandwidth audio, the system uses the neural network mediator to generate the high-frequency content, achieving quality improvement without proportional bitrate increase.
3Manufacturing precision
If a neural network system is used to reconstruct high-frequency bands in filter bank domain, then sound quality is improved and distortion is reduced, but computational complexity increases
Solution Approach 1:
The neural network is pre-trained offline on large datasets to learn optimal high-frequency reconstruction patterns. During actual audio decoding, the pre-trained network performs inference with reduced computational burden compared to training. This preliminary action separates the heavy computational work (training) from the real-time operation (inference), achieving high reconstruction accuracy with manageable runtime complexity.
Data Source
AI summary
A method for reconstructing an audio signal, comprising receiving a bitstream including an encoded low-band audio signal representation and a set of high frequency reconstruction, HFR, parameters, decoding the low-band audio signal representation to provide a low-band audio signal in a filter bank domain, reconstructing a filter bank domain high-band audio signal using a neural network system trained to predict samples of the high-band audio signal in the filter bank domain given samples of the filter bank domain low-band signal and the HFR parameters, and synthesizing a time domain output audio signal from the filter bank domain low-band signal and the reconstructed filter bank domain high-band signal. By using a generative model in the form of a neural network system to reconstruct the high-frequency range, a perceptually improved audio output can be achieved.


