GAN-Based Audio Compression for High-Fidelity File Reduction

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing audio compression technologies are limited in their ability to achieve high levels of compression while maintaining audio fidelity, leading to significant storage, transmission, and processing overhead.

Innovation Solution

The method employs AI transformer-based compression using a deep learning-enabled Generative Adversarial Network (GAN) for upsampling audio content, applying AI-assisted mapping of dynamics and harmonics, increasing dimensional complexity, and utilizing a transcoding transform to facilitate spectral transformation, along with GAN-assisted attention transformer methods for selective encoding and optimization.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Quantity of substance

If traditional audio compression methods are used, then compression ratio is limited, but storage and transmission overhead remains high

Engineering Contradiction:
Improveaudio data sizeVSAvoidaudio fidelity
Core Design Contradiction:
Quantity of substanceVSReliability

Solution Approach 1:

The patent transforms audio data from the traditional time-domain representation into a latent space representation using a neural network encoder. This dimensional transformation allows the system to capture essential audio features in a compressed latent code space, achieving high compression ratios while maintaining audio fidelity through the learned mapping between original and latent dimensions

Inventive Principle:
Principle #17Another dimension (Dimensionality change)

Solution Approach 2:

The system changes the representation parameters of audio data by learning an optimal mapping from the original audio signal to a compressed latent code space. The neural network automatically adjusts the dimensionality and feature representation parameters to achieve the best compression-fidelity tradeoff, transforming the audio into a more efficient representation format

Inventive Principle:
Principle #35Parameter changes

2Quantity of substance

If higher compression ratios are achieved, then storage and transmission overhead is reduced, but audio quality deteriorates

Engineering Contradiction:
Improvecompressed data sizeVSAvoidaudio quality
Core Design Contradiction:
Quantity of substanceVSManufacturing precision

Solution Approach 1:

The patent incorporates a discriminator component that provides feedback on the quality of reconstructed audio. The discriminator evaluates the generated audio samples and provides gradients to the generator network, enabling iterative improvement of the compression-reconstruction process to maintain high audio quality even at high compression ratios

Inventive Principle:
Principle #23Feedback

Solution Approach 2:

The system replaces traditional mechanical compression algorithms with a neural network-based approach that learns optimal feature representations. The neural network encoder-decoder framework substitutes for conventional lossy compression methods, achieving superior quality preservation through data-driven learning of audio characteristics

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

3Productivity

If AI-based compression is applied, then compression efficiency improves, but computational complexity increases

Engineering Contradiction:
Improvecompression efficiencyVSAvoidcomputational requirements
Core Design Contradiction:
ProductivityVSDevice complexity

Solution Approach 1:

The patent pre-trains a neural network encoder on a large dataset of audio data to learn optimal compression representations. This preliminary training phase captures essential audio patterns and features, allowing the model to achieve high compression efficiency during actual use without requiring complex real-time computations, as the heavy lifting is done during the pre-training phase

Inventive Principle:
Principle #10Preliminary action

Data Source

PatentUS12387736B2Audio compression with generative adversarial networks
Publication Date: 2025.08.12 ZON GLOBAL IP INC
  • US12387736B2 patent drawing
  • US12387736B2 patent drawing
  • US12387736B2 patent drawing

AI summary

An AI-based audio compression method for use with audio formats, alone or in combination with other audio compression and enhancement approaches. A combination of audio pre-processing, sound to visual transcoding of audio, and a sequence of AI-enabled methods enabling maximal entropy order extraction applied within the sound and dimensionally extended visual domain projection of the audio significantly increases the degree of compression of audio files, thereby reducing storage, transmission and processing overhead associated with audio. A unique AI-driven domain conversion is leveraged together with domain-specific AI processing stages to reduce file size, while supporting optional use of standard and proprietary audio encoding, decoding, compression, and other methods. Support for native mode photonic computer processing of the higher-dimensional order representation of media enables further optimization via photonic computing methods that would not be possible if the audio was not extended into higher order visual domain space.