Audio Dequantization via Diffusion Model Latent Vector

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing technologies face challenges in effectively decoding compressed speech signals from low bit rate discrete latent vectors to high bit rate continuous latent vectors, which affects the quality of restored speech signals.

Innovation Solution

A method and device that utilize a generative model, specifically a diffusion model, to estimate a high bit rate continuous latent vector from a low bit rate discrete latent vector by gradually up-sampling the discrete latent vector through multiple layers of a neural network, removing noise from each layer to dequantize the vector.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Loss of energy

If a low bit rate discrete latent vector is used for compressed speech signal encoding, then transmission efficiency and storage are improved, but the quality and fidelity of the restored speech signal deteriorates

Engineering Contradiction:
Improvetransmission efficiencyVSAvoidspeech signal quality
Core Design Contradiction:
Loss of energyVSManufacturing precision

Solution Approach 1:

The patent transforms the discrete latent vector parameters into continuous parameters through a diffusion model-based dequantization process. This parameter transformation allows the system to convert low-bitrate discrete representations into high-bitrate continuous representations, thereby improving speech quality while maintaining transmission efficiency benefits from the original compressed discrete form.

Inventive Principle:
Principle #35Parameter changes

Solution Approach 2:

The patent introduces a continuous latent vector as an intermediary between the discrete latent vector and the final speech signal reconstruction. This intermediary continuous representation serves as a bridge that enables high-quality speech reconstruction without requiring direct transmission of high-bitrate continuous data, thus resolving the contradiction between compression efficiency and quality.

Inventive Principle:
Principle #24Intermediary (Mediator)

2Quantity of substance

If quantization is applied to compress the speech signal, then data rate is reduced, but measurement precision of the latent vector deteriorates

Engineering Contradiction:
Improvedata rateVSAvoidlatent vector precision
Core Design Contradiction:
Quantity of substanceVSMeasurement precision

Solution Approach 1:

The patent applies parameter transformation by converting discrete quantized parameters into continuous parameters through the diffusion model. This transformation recovers the precision information that was lost during quantization, enabling high-precision speech signal reconstruction from low-precision discrete latent vectors.

Inventive Principle:
Principle #35Parameter changes

Solution Approach 2:

The patent performs dequantization as a preliminary action before speech signal reconstruction. By converting the discrete latent vector to a continuous latent vector beforehand, the system prepares high-precision input for the subsequent speech reconstruction process, ensuring optimal quality without requiring high-bitrate transmission.

Inventive Principle:
Principle #10Preliminary action

3Device complexity

If direct reconstruction from discrete latent vector is performed, then processing complexity is reduced, but speech signal fidelity deteriorates

Engineering Contradiction:
Improveprocessing complexityVSAvoidspeech signal fidelity
Core Design Contradiction:
Device complexityVSManufacturing precision

Solution Approach 1:

The patent introduces a continuous latent vector as an intermediary representation that enables high-fidelity speech reconstruction. This intermediate continuous form allows the reconstruction process to operate on high-precision data without requiring complex high-bitrate encoding/decoding systems, thus achieving high fidelity with manageable processing complexity.

Inventive Principle:
Principle #24Intermediary (Mediator)

Solution Approach 2:

The patent segments the reconstruction process into two distinct stages: first converting discrete latent vectors to continuous latent vectors through diffusion modeling, then reconstructing the speech signal from the continuous representation. This segmentation allows each stage to be optimized independently, maintaining high fidelity while controlling overall processing complexity.

Inventive Principle:
Principle #1Segmentation

Data Source

PatentUS20250104722A1Method and device for encoding/decoding audio signal based on dequantization through potential diffusion
Publication Date: 2025.03.27 ELECTRONICS & TELECOMM RES INST
  • US20250104722A1 patent drawing
  • US20250104722A1 patent drawing
  • US20250104722A1 patent drawing

AI summary

A method and device for encoding/decoding an audio signal based on dequantization through potential diffusion are provided. The method of decoding an audio signal includes obtaining a discrete latent vector in which a speech signal is quantized and based on the discrete latent vector, outputting a continuous latent vector in which the discrete latent vector is dequantized.