Neural Network Video Compression Auxiliary Information Positioning
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current video compression technologies face challenges in achieving high efficiency and quality due to limited bandwidth and storage resources, with existing methods often sacrificing compression ratio for picture quality.
Innovation Solution
A neural network-based method that processes picture feature data in multiple stages, utilizing prediction error and prediction signals from different stages to generate input for the neural network, allowing for configurable architecture and improved encoding and decoding efficiency.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Quantity of substance
If video data is compressed to reduce data size for transmission and storage, then the quantity of data is reduced, but picture quality deteriorates
Solution Approach 1:
The video encoding process is divided into multiple stages including prediction, residual calculation, transform, quantization, and entropy coding. Each stage processes specific aspects of the video data independently, allowing optimized compression at each step while maintaining overall quality through selective processing of different frequency components and prediction errors
Solution Approach 2:
The system dynamically adjusts quantization parameters, prediction modes, and transform block sizes based on content characteristics and desired quality levels. By changing these parameters adaptively, the encoder achieves optimal compression ratios while maintaining picture quality within acceptable ranges for different transmission conditions
2Productivity
If complex neural network processing is applied to improve encoding and decoding performance, then picture quality and compression efficiency are improved, but computational complexity increases
Solution Approach 1:
Prediction signals are generated in advance using motion estimation and compensation techniques before the main encoding process. This preliminary action reduces the amount of data that needs to be processed by subsequent neural network stages, lowering computational complexity while maintaining encoding efficiency
Solution Approach 2:
The neural network processes only the most critical components of the video data, such as prediction errors and high-frequency residuals, rather than processing all data uniformly. This partial processing approach achieves significant compression efficiency improvements with reduced computational burden compared to full-frame processing
3Productivity
If multiple stages of neural network processing are used to enhance compression ratio, then compression efficiency is improved, but processing latency increases
Solution Approach 1:
For real-time or near-real-time applications, the system can skip certain non-critical processing stages or use simplified models in the neural network pipeline. This allows maintaining acceptable compression ratios while significantly reducing processing latency to meet timing constraints for live streaming or interactive applications
Data Source
AI summary
This application provides methods and apparatuses for processing of picture data or picture feature data using a neural network with two or more layers. The present disclosure may be applied in the field of artificial intelligence (AI)-based video or picture compression technologies, and in particular, to the field of neural network-based video compression technologies. According to some embodiments, two kinds of data are combined during the processing including processing by the neural network. The two kinds of data are obtained from different stages of processing by the network. Some of the advantages may include greater scalability and a more flexible design of the neural network architecture which may further lead to better encoding/decoding performance.


