Video Coding for Machines Prediction Encoding
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing video coding schemes are optimized for human vision and fail to meet the latency and scale requirements of machine vision applications, such as connected vehicles and IoT devices, necessitating a new paradigm for compressing machine vision data.
Innovation Solution
A Video Coding for Machines (VCM) apparatus and method that uses prediction to generate prediction data and residual data based on correlation, improving encoding efficiency by setting reference data and encoding residual data for machine vision systems.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If existing video coding schemes are used, then image quality for human vision is improved, but encoding efficiency for machine vision data deteriorates
Solution Approach 1:
The patent changes the fundamental parameters of video coding by switching from human vision-oriented quality metrics to machine vision-oriented feature representation metrics. This involves transforming the coding objective from preserving visual quality to preserving task-relevant features, thereby improving encoding efficiency for machine vision applications without sacrificing the necessary information for machine tasks.
Solution Approach 2:
The patent applies local quality by encoding different regions of the video data with different levels of detail based on their importance for machine vision tasks. Rather than uniformly encoding all pixels, the system identifies and prioritizes encoding of task-relevant regions, improving overall encoding efficiency while maintaining necessary quality for machine analysis.
2Loss of energy
If video data is compressed for transmission, then transmission costs are reduced, but latency increases
Solution Approach 1:
The patent extracts only the essential feature information needed for machine vision tasks from the full video data, discarding redundant information. This extraction approach reduces the amount of data requiring transmission, thereby lowering transmission costs and bandwidth requirements while maintaining the latency performance needed for real-time machine vision applications.
3Measurement precision
If feature maps from deep learning models are transmitted, then machine vision accuracy is maintained, but data volume increases
Solution Approach 1:
The patent performs preliminary feature extraction and selection before transmission, preprocessing the data to retain only the most task-relevant features. This preliminary action reduces the subsequent data volume requiring transmission while ensuring that the essential information for maintaining machine vision accuracy is preserved.
Solution Approach 2:
The patent segments the feature maps into task-relevant and redundant components, transmitting only the essential segments. This segmentation approach maintains machine vision accuracy by preserving critical feature information while reducing overall data volume by excluding unnecessary features.
Data Source
AI summary
The present disclosure relates to an apparatus for and a method of coding machine vision data by using prediction, and for improving the efficiency of encoding the data used for machine vision, provides an apparatus for Video Coding for Machines (VCM) which sets reference data according to a correlation between the data, generates, based on the reference data, prediction data for original data having a high correlation with the reference data, and generates residual data between the prediction data and the original data, and provides a coding method performed by the apparatus for VCM.


