AI Transformer Audio Compression Using Spectral Order Extraction

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current audio compression technologies are limited in their ability to achieve high levels of compression while maintaining audio fidelity, leading to significant storage, transmission, and processing overhead.

Innovation Solution

The method employs AI transformer-based file compression using a deep learning enabled temporal Generative Adversarial Network (GAN) for upsampling, AI-assisted mapping of dynamics and harmonics, noise reduction, and deep learning-based audio compression with attention transformers to increase dimensional complexity and optimize sampling rate and depth, resulting in enhanced compression efficiency.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Quantity of substance

If traditional audio compression methods are used, then compression ratio is limited, but storage and processing overhead remains high

Engineering Contradiction:
Improvecompression ratioVSAvoidstorage and processing overhead
Core Design Contradiction:
Quantity of substanceVSLoss of energy

Solution Approach 1:

The patent applies parameter changes by transforming audio from the time domain to the frequency domain using Fast Fourier Transform (FFT), then applying compression in the frequency domain where spectral characteristics reveal hidden patterns and correlations. This dimensional transformation enables significantly higher compression ratios while maintaining audio fidelity, directly resolving the contradiction between compression ratio and storage overhead.

Inventive Principle:
Principle #35Parameter changes

Solution Approach 2:

The patent replaces traditional mechanical compression systems with an AI transformer-based system that uses deep learning models trained on audio data. The transformer architecture processes audio spectrograms and generates compressed representations through learned patterns, achieving superior compression efficiency compared to conventional algorithms, thereby reducing both storage requirements and processing overhead.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

2Quantity of substance

If compression level is increased, then storage overhead is reduced, but audio fidelity deteriorates

Engineering Contradiction:
Improvecompression levelVSAvoidaudio fidelity
Core Design Contradiction:
Quantity of substanceVSManufacturing precision

Solution Approach 1:

The patent applies dimensionality change by converting audio signals into spectrogram representations in the frequency-time domain. This transformation reveals spectral correlations and patterns that are not apparent in the time domain, enabling the AI transformer to identify and compress redundant information while preserving essential audio characteristics. The result is high compression levels maintained with excellent audio fidelity.

Inventive Principle:
Principle #17Another dimension (Dimensionality change)

Solution Approach 2:

The patent implements feedback mechanisms through the transformer architecture's attention mechanisms and residual connections, which continuously refine the compression process. The model learns from the input audio spectrum and adjusts its compression strategy in real-time, ensuring that compressed representations maintain high fidelity while achieving maximum compression, thus resolving the contradiction between compression level and audio quality.

Inventive Principle:
Principle #23Feedback

Data Source

PatentUS20250364003A1Transformer sequenced order-extracted ensemble compression
Publication Date: 2025.11.27 ZON GLOBAL IP INC
  • US20250364003A1 patent drawing
  • US20250364003A1 patent drawing
  • US20250364003A1 patent drawing

AI summary

An AI-based audio compression method for use with audio formats, alone or in combination with other audio compression and enhancement approaches. A combination of audio pre-processing, sound to visual transcoding of audio, and a sequence of AI-enabled methods enabling maximal entropy order extraction applied within the sound and dimensionally extended visual domain projection of the audio significantly increases the degree of compression of audio files, thereby reducing storage, transmission and processing overhead associated with audio. A unique AI-driven domain conversion is leveraged together with domain-specific AI processing stages to reduce file size, while supporting optional use of standard and proprietary audio encoding, decoding, compression, and other methods. Support for native mode photonic computer processing of the higher-dimensional order representation of media enables further optimization via photonic computing methods that would not be possible if the audio was not extended into higher order visual domain space.