Visual Data Processing Residual Representation Gain Units
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current neural network-based image/video compression methods require training multiple models for different rate points, which is inefficient in terms of time and storage, and can only achieve discretized rate points.
Innovation Solution
The proposed method uses a single model with gain units to achieve continuously variable rate adaptation by adjusting the value range of the residual representation based on a gain parameter, allowing for improved coding efficiency and effectiveness.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Adaptability or versatility
If multiple models are trained for different rate points, then rate adaptation capability is improved, but training time and storage requirements increase
Solution Approach 1:
The patent applies universality by designing a single neural network model that can adapt to multiple rate points through gain units. Instead of training separate specialized models for different rate points, the universal model with gain parameters can continuously adjust its behavior to achieve different compression rates, eliminating the need for multiple models and reducing both training time and storage requirements.
Solution Approach 2:
The patent utilizes parameter changes by introducing gain parameters that control the value range of residual representation. By dynamically adjusting these gain parameters, the model can adapt to different rate points without retraining. This parameter-based adaptation allows continuous rate adjustment while maintaining a single model structure, directly addressing the contradiction between adaptability and training efficiency.
2Adaptability or versatility
If multiple models are trained for different rate points, then rate adaptation capability is improved, but storage requirements increase
Solution Approach 1:
The universal model with gain units serves multiple rate adaptation functions within a single model structure. This eliminates the need to store multiple separate model files for different rate points, significantly reducing storage requirements while maintaining the capability to adapt to various compression rates through parameter adjustment.
Solution Approach 2:
The patent merges the functionality of multiple rate-specific models into a single unified model. By combining different rate adaptation capabilities into one model with adjustable gain parameters, the system reduces storage requirements by storing only one model instead of multiple models, while still achieving comprehensive rate adaptation.
3Ease of manufacture
If discretized rate points are used, then model training is simplified, but rate adjustment flexibility is reduced
Solution Approach 1:
The patent applies dynamics by transforming the static, discretized rate points into a dynamic, continuous rate adjustment mechanism. The gain units enable the model to continuously vary its behavior based on the desired rate, moving from fixed discrete rate options to flexible continuous rate control while maintaining training simplicity through the unified model structure.
Solution Approach 2:
The patent uses parameter changes to enable continuous rate adjustment. By introducing gain parameters that can take continuous values, the system transitions from discretized rate points to continuous rate control. This parameter-based approach maintains training simplicity while dramatically improving rate adjustment flexibility, as the model can adapt to any rate within its operational range.
Data Source
AI summary
Embodiments of the present disclosure provide a solution for visual data processing. A method for visual data processing is proposed. The method comprises: determining, for a conversion between at least one bitstream of visual data and the visual data, a residual representation of the visual data at least based on a first probability distribution parameter of the visual data and a gain parameter, the residual representation representing a residual value compared to a second probability distribution representation of the visual data, the gain parameter adjusting a value range of the residual representation; and performing the conversion based on the residual representation.


