Audio Dequantization via Diffusion Model Latent Vector
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing technologies face challenges in effectively decoding compressed speech signals from low bit rate discrete latent vectors to high bit rate continuous latent vectors, which affects the quality of restored speech signals.
Innovation Solution
A method and device that utilize a generative model, specifically a diffusion model, to estimate a high bit rate continuous latent vector from a low bit rate discrete latent vector by gradually up-sampling the discrete latent vector through multiple layers of a neural network, removing noise from each layer to dequantize the vector.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Loss of energy
If a low bit rate discrete latent vector is used for compressed speech signal encoding, then transmission efficiency and storage are improved, but the quality and fidelity of the restored speech signal deteriorates
Solution Approach 1:
The patent transforms the discrete latent vector parameters into continuous parameters through a diffusion model-based dequantization process. This parameter transformation allows the system to convert low-bitrate discrete representations into high-bitrate continuous representations, thereby improving speech quality while maintaining transmission efficiency benefits from the original compressed discrete form.
Solution Approach 2:
The patent introduces a continuous latent vector as an intermediary between the discrete latent vector and the final speech signal reconstruction. This intermediary continuous representation serves as a bridge that enables high-quality speech reconstruction without requiring direct transmission of high-bitrate continuous data, thus resolving the contradiction between compression efficiency and quality.
2Quantity of substance
If quantization is applied to compress the speech signal, then data rate is reduced, but measurement precision of the latent vector deteriorates
Solution Approach 1:
The patent applies parameter transformation by converting discrete quantized parameters into continuous parameters through the diffusion model. This transformation recovers the precision information that was lost during quantization, enabling high-precision speech signal reconstruction from low-precision discrete latent vectors.
Solution Approach 2:
The patent performs dequantization as a preliminary action before speech signal reconstruction. By converting the discrete latent vector to a continuous latent vector beforehand, the system prepares high-precision input for the subsequent speech reconstruction process, ensuring optimal quality without requiring high-bitrate transmission.
3Device complexity
If direct reconstruction from discrete latent vector is performed, then processing complexity is reduced, but speech signal fidelity deteriorates
Solution Approach 1:
The patent introduces a continuous latent vector as an intermediary representation that enables high-fidelity speech reconstruction. This intermediate continuous form allows the reconstruction process to operate on high-precision data without requiring complex high-bitrate encoding/decoding systems, thus achieving high fidelity with manageable processing complexity.
Solution Approach 2:
The patent segments the reconstruction process into two distinct stages: first converting discrete latent vectors to continuous latent vectors through diffusion modeling, then reconstructing the speech signal from the continuous representation. This segmentation allows each stage to be optimized independently, maintaining high fidelity while controlling overall processing complexity.
Data Source
AI summary
A method and device for encoding/decoding an audio signal based on dequantization through potential diffusion are provided. The method of decoding an audio signal includes obtaining a discrete latent vector in which a speech signal is quantized and based on the discrete latent vector, outputting a continuous latent vector in which the discrete latent vector is dequantized.


