Multistage Context Modeling for Faster Visual Data Compression
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing neural network-based video compression methods face challenges in coding efficiency and complexity due to the inherent difficulty of modeling high-dimensional visual data, particularly in high-resolution images and videos, leading to suboptimal performance and high decoding times.
Innovation Solution
A multistage context model is introduced, which simplifies the prediction fusion network by limiting the number of convolutional layers to improve coding effectiveness and efficiency, utilizing a multistage context module for visual data processing.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Manufacturing precision
If neural network-based video compression methods are used to model high-dimensional visual data, then coding effectiveness can be improved, but device complexity and decoding time increase significantly
Solution Approach 1:
The patent segments the visual data processing into multiple stages, where different processing operations are applied at different stages. This segmentation allows the system to achieve high coding effectiveness through careful modeling at each stage while avoiding the need for a single complex all-encompassing model, thus reducing overall decoding complexity.
Solution Approach 2:
The patent employs dynamic adaptive modeling where the modeling approach adapts based on the specific characteristics of the visual data being processed. This allows the system to optimize coding effectiveness for different types of content while maintaining manageable decoding complexity through adaptive rather than static complex modeling.
2Productivity
If complex neural network models are used for visual data compression, then coding efficiency may improve, but decoding time increases
Solution Approach 1:
The patent performs preliminary processing and modeling operations during the encoding phase, preparing data in advance to facilitate faster decoding. By doing the heavy computational lifting during encoding when time is less critical, the system achieves high coding efficiency while reducing the time required for decoding operations.
Solution Approach 2:
The multi-stage processing approach segments computational operations between encoding and decoding phases, allowing complex operations to be performed during encoding to improve coding efficiency, while simpler operations remain for decoding to reduce decoding time.
Data Source
AI summary
Embodiments of the present disclosure provide a solution for visual data processing. A method for visual data processing is proposed. In the method, for a conversion between a current visual unit of visual data and a bitstream of the visual data, a probability representation of the current visual unit is determined based on a multistage context module. The conversion is performed based on the probability representation. The multistage context module at least comprises at least one prediction fusion network. The number of convolutional layers in the at least one prediction fusion network is less than or equal to a first threshold number.


