Hierarchical Audio/Video Compression for High-Bit-Rate Efficiency
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing AI-based audio/video or picture compression technologies face limitations at high bit rates, with model compression quality stabilizing and requiring excessive computing power, while conventional signal processing methods struggle to achieve significant breakthroughs in compression efficiency.
Innovation Solution
A hierarchical audio/video or picture compression method utilizing a neural network that processes signals at different layers, employing entropy encoding and decoding with probability estimation to improve compression performance and reduce computing power by capturing high and low-frequency information separately, allowing for dynamic adjustment of compression layers and channels.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Manufacturing precision
If variational autoencoder (AE) method is used for AI-based picture compression, then compression quality is improved, but computing power requirement increases excessively at high bit rates
Solution Approach 1:
The patent segments the compression process into multiple hierarchical layers, where different layers process different frequency components of the image. This segmentation allows the system to achieve high compression quality without requiring excessive computing power at any single layer, thereby resolving the contradiction between compression quality and computing power requirements.
Solution Approach 2:
The patent introduces a hierarchical layering dimension to the compression architecture, transforming a single-stage compression process into a multi-stage process. This dimensional change enables the system to distribute computational load across layers, achieving high compression quality while controlling computing power consumption through progressive refinement from coarse to fine details.
2Productivity
If conventional signal processing methods are used, then computing power is reduced, but compression efficiency cannot achieve significant breakthroughs
Solution Approach 1:
The patent employs dynamic adaptive processing where the hierarchical layers and channels are configured based on the specific characteristics of the input image and desired compression ratio. This dynamic approach allows the system to optimize compression efficiency for different scenarios while adapting computing power requirements, achieving breakthroughs in compression efficiency beyond conventional static signal processing methods.
Solution Approach 2:
The patent changes key parameters of the compression system by introducing hierarchical layering and selective entropy encoding. By adjusting the number of layers, channels, and encoding precision at different levels, the system achieves significant improvements in compression efficiency while maintaining controllable computing power requirements through parameter optimization.
3Manufacturing precision
If fixed model is used for compression, then model complexity is reduced, but compression quality tends to be stable at high bit rate (AE limit problem)
Solution Approach 1:
The patent segments the compression model into hierarchical layers with different functions - coarse layers for overall structure and fine layers for detailed refinement. This segmentation allows the model to progressively improve compression quality without requiring a single overly complex fixed model, thereby breaking through the AE limit while controlling overall model complexity through modular layer design.
Solution Approach 2:
The patent introduces dynamic adaptability to the compression model through hierarchical layers that can be selectively activated based on the desired compression ratio and image characteristics. This dynamic approach allows the model to adjust its complexity and processing depth, achieving high compression quality at various bit rates without being constrained by a fixed model structure.
Data Source
AI summary
This application provides an audio/video or picture compression method and apparatus, which relates to the field of artificial intelligence (AI)-based audio/video or picture compression technologies, and to the field of neural network-based audio/video or picture compression technologies. The method includes: transforming a raw audio/video or picture to feature space through a multilayer convolution operation, extracting features of different layers in the feature space, outputting rounded feature signals of the different layers, predicting probability distribution of shallow feature signals by using deep feature signals or entropy estimation results, and performing entropy encoding on the rounded feature signals. In this application, signal correlation between different layers is utilized. In this way, audio/video or picture compression performance can be improved.


