Feature Compression for Video Coding Machines

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current video coding technologies are inadequate for efficiently compressing videos for machine vision tasks, as they do not effectively utilize the characteristics of feature channels in a feature map, leading to suboptimal compression efficiency and interoperability issues among devices.

Innovation Solution

A method and device for performing feature compression by obtaining an input video, generating a feature map with multiple channels, reordering these channels based on characteristics such as mean, variance, or correlation coefficients, and compressing them to generate an encoded bitstream, which can be decoded and used for machine vision tasks.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Productivity

If conventional video coding technologies are used, then compatibility with human visual systems is maintained, but compression efficiency for machine vision tasks deteriorates

Engineering Contradiction:
Improvecompression efficiencyVSAvoidadaptability to machine vision tasks
Core Design Contradiction:
ProductivityVSAdaptability or versatility

Solution Approach 1:

The patent applies local quality by treating different feature channels with different importance weights based on their correlation characteristics. Highly correlated channels are compressed with higher compression ratios while maintaining quality for less correlated channels, optimizing compression efficiency for machine vision tasks without uniformly compromising all channels.

Inventive Principle:
Principle #3Local quality

Solution Approach 2:

The patent changes the parameter organization by reordering feature channels based on correlation coefficients before compression. This parameter transformation enables the compression algorithm to exploit channel correlations more effectively, improving compression efficiency specifically for machine vision applications where feature correlations are more predictable than in natural video.

Inventive Principle:
Principle #35Parameter changes

2Quantity of substance

If feature channels are compressed without reordering, then processing speed is maintained, but compression ratio deteriorates

Engineering Contradiction:
Improvecompression ratioVSAvoidprocessing complexity
Core Design Contradiction:
Quantity of substanceVSDevice complexity

Solution Approach 1:

The patent applies preliminary action by reordering feature channels based on correlation coefficients before the compression process. This preprocessing step organizes the data in a way that maximizes compression efficiency, allowing the subsequent compression algorithm to achieve better ratios without significantly increasing overall processing complexity.

Inventive Principle:
Principle #10Preliminary action

3Productivity

If feature channels are reordered based on correlation, then compression efficiency improves, but computational overhead increases

Engineering Contradiction:
Improvecompression efficiencyVSAvoidcomputational power
Core Design Contradiction:
ProductivityVSPower

Solution Approach 1:

The patent extracts and utilizes the correlation structure inherent in feature channels by computing correlation coefficients and using them to guide reordering. This extraction of structural information enables more efficient compression without requiring excessive computational resources, as the correlation-based ordering exploits natural redundancies in the data.

Inventive Principle:
Principle #2Taking out (Extraction)

Data Source

PatentUS12126841B2Feature compression for video coding for machines
Publication Date: 2024.10.22 TENCENT AMERICA LLC
  • US12126841B2 patent drawing
  • US12126841B2 patent drawing
  • US12126841B2 patent drawing

AI summary

Systems, devices, and methods for performing feature compression, including obtaining an input video; obtaining a feature map corresponding to the input video, the feature map including a plurality of feature channels; reordering the plurality of feature channels based on at least one characteristic of the plurality of feature channels; compressing the reordered plurality of feature channels; and generating an encoded bitstream based on the compressed and reordered plurality of feature channels.