Scalable Video Compression for Hybrid Machine Human Vision
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current video compression technologies fail to efficiently compress videos for both machine vision and human vision, as they do not adequately account for the distinct characteristics required for each domain, leading to suboptimal performance in machine-to-machine communication applications.
Innovation Solution
A scalable video compression structure is proposed, which uses an adaptive loop filter with coefficients derived through feature domain minimum error and task error minimum error methods, classifying coding tree units into significant and insignificant groups based on pixel importance and edge direction, and performing filtering accordingly to enhance encoding efficiency.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Productivity
If conventional video compression is used, then encoding speed is maintained, but compression efficiency for both machine vision and human vision deteriorates
Solution Approach 1:
The image is divided into coding tree units that are further segmented into sub-blocks, allowing different filtering operations to be applied to different segments based on their specific characteristics and importance for machine vision tasks
Solution Approach 2:
Different filter coefficients are derived and applied to different coding tree units based on their importance for machine vision tasks, enabling optimized compression performance for each local region rather than using a uniform approach
2Reliability
If adaptive loop filtering is applied to all coding tree units, then compression performance improves, but processing time increases
Solution Approach 1:
Instead of applying complex adaptive loop filtering to all coding tree units, the method applies filtering selectively to only those coding tree units that are important for machine vision tasks, reducing processing time while maintaining compression performance for critical regions
Solution Approach 2:
Coding tree units are classified into important and less important groups before filtering is applied, allowing the system to prioritize processing time for regions that matter most for machine vision while using faster processing for other regions
Data Source
AI summary
The present invention proposes a scalable-based video compression structure in a video compression technology for supporting a hybrid task. In an adaptive loop filter step of an encoder of a layer for a machine task, a coding tree unit may be classified into a coding tree unit significant group and a coding tree unit insignificant group, and, for the coding tree unit significant group, filter coefficients may be derived by a feature domain minimum error method and a task error minimum error method.


