Video Coding for Machines Prediction Encoding

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing video coding schemes are optimized for human vision and fail to meet the latency and scale requirements of machine vision applications, such as connected vehicles and IoT devices, necessitating a new paradigm for compressing machine vision data.

Innovation Solution

A Video Coding for Machines (VCM) apparatus and method that uses prediction to generate prediction data and residual data based on correlation, improving encoding efficiency by setting reference data and encoding residual data for machine vision systems.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If existing video coding schemes are used, then image quality for human vision is improved, but encoding efficiency for machine vision data deteriorates

Engineering Contradiction:
Improveimage qualityVSAvoidencoding efficiency
Core Design Contradiction:
Measurement precisionVSProductivity

Solution Approach 1:

The patent changes the fundamental parameters of video coding by switching from human vision-oriented quality metrics to machine vision-oriented feature representation metrics. This involves transforming the coding objective from preserving visual quality to preserving task-relevant features, thereby improving encoding efficiency for machine vision applications without sacrificing the necessary information for machine tasks.

Inventive Principle:
Principle #35Parameter changes

Solution Approach 2:

The patent applies local quality by encoding different regions of the video data with different levels of detail based on their importance for machine vision tasks. Rather than uniformly encoding all pixels, the system identifies and prioritizes encoding of task-relevant regions, improving overall encoding efficiency while maintaining necessary quality for machine analysis.

Inventive Principle:
Principle #3Local quality

2Loss of energy

If video data is compressed for transmission, then transmission costs are reduced, but latency increases

Engineering Contradiction:
Improvetransmission costsVSAvoidlatency
Core Design Contradiction:
Loss of energyVSLoss of time

Solution Approach 1:

The patent extracts only the essential feature information needed for machine vision tasks from the full video data, discarding redundant information. This extraction approach reduces the amount of data requiring transmission, thereby lowering transmission costs and bandwidth requirements while maintaining the latency performance needed for real-time machine vision applications.

Inventive Principle:
Principle #2Taking out (Extraction)

3Measurement precision

If feature maps from deep learning models are transmitted, then machine vision accuracy is maintained, but data volume increases

Engineering Contradiction:
Improvemachine vision accuracyVSAvoiddata volume
Core Design Contradiction:
Measurement precisionVSQuantity of substance

Solution Approach 1:

The patent performs preliminary feature extraction and selection before transmission, preprocessing the data to retain only the most task-relevant features. This preliminary action reduces the subsequent data volume requiring transmission while ensuring that the essential information for maintaining machine vision accuracy is preserved.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The patent segments the feature maps into task-relevant and redundant components, transmitting only the essential segments. This segmentation approach maintains machine vision accuracy by preserving critical feature information while reducing overall data volume by excluding unnecessary features.

Inventive Principle:
Principle #1Segmentation

Data Source

PatentUS11516478B2Method and apparatus for coding machine vision data using prediction
Publication Date: 2022.11.29 HYUNDAI MOTOR CO LTD
  • US11516478B2 patent drawing
  • US11516478B2 patent drawing
  • US11516478B2 patent drawing

AI summary

The present disclosure relates to an apparatus for and a method of coding machine vision data by using prediction, and for improving the efficiency of encoding the data used for machine vision, provides an apparatus for Video Coding for Machines (VCM) which sets reference data according to a correlation between the data, generates, based on the reference data, prediction data for original data having a high correlation with the reference data, and generates residual data between the prediction data and the original data, and provides a coding method performed by the apparatus for VCM.