Generative Waveform Coding for Low-Bitrate Audio Reconstruction
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing audio coding technologies struggle with perceptual artifacts due to low-rate quantization, and deep generative models have not been effectively applied to improve perceptual quality in audio coding, particularly in reconstructing plausible signal structures.
Innovation Solution
A method and system that utilizes a generative model to decode a finite bitrate representation of a source signal, implementing a probability density function to generate a reconstructed signal, combining waveform and parametric coding for improved perceptual performance.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Loss of energy
If low-rate quantization is used for waveform coding, then bitrate is reduced, but coding artifacts and perceptual quality deteriorate
Solution Approach 1:
The patent introduces a generative model as an intermediary component between the quantizer and the final reconstructed signal. This generative model takes the quantized waveform approximation and transforms it into a more perceptually accurate reconstruction by learning the underlying probability distribution of audio signals, thereby mediating the harmful effects of quantization artifacts while maintaining low bitrate
Solution Approach 2:
The patent changes the parameter representation from direct waveform samples to probability distribution parameters. By modeling the audio signal as a probability distribution and using generative sampling, the system transforms deterministic quantization into a probabilistic framework that can better represent perceptually relevant features even at low bitrates
2Device complexity
If simple quantizers are used in transform coding, then device complexity is reduced, but manufacturing precision and perceptual quality worsen
Solution Approach 1:
The patent replaces traditional mechanical/mathematical quantization systems with a data-driven generative model. Instead of using fixed quantization tables or simple scalar quantizers, the system employs a neural network-based generative model that has been trained to capture perceptually relevant patterns, substituting complex mathematical transformations with learned representations that achieve better perceptual quality
3Object-affected harmful factors
If deep generative models are applied to audio coding, then perceptual quality is improved, but device complexity increases
Solution Approach 1:
The patent applies preliminary action by training the generative model offline before actual audio coding operations. The model is pre-trained on large datasets to learn audio signal characteristics, and this pre-learned knowledge is then used during coding without requiring complex real-time computations. The bitstream includes only essential parameters that guide the pre-trained model, reducing online complexity
Solution Approach 2:
The patent uses copying by training the generative model to learn and replicate the statistical properties and perceptual characteristics of audio signals from training data. The model creates a copy of the underlying probability distribution of audio signals, which can then be sampled to generate realistic waveforms without needing to process and transmit the actual original signal details
Data Source
AI summary
Described herein is a method of waveform decoding, the method including the steps of: (a) receiving, by a waveform decoder, a bitstream including a finite bitrate representation of a source signal; (b) waveform decoding the finite bitrate representation of the source signal to obtain a waveform approximation of the source signal; (c) providing the waveform approximation of the source signal to a generative model that implements a probability density function, to obtain a probability distribution for a reconstructed signal of the source signal; and (d) generating the reconstructed signal of the source signal based on the probability distribution. Described are further a method and system for waveform coding and a method of training a generative model.


