Machine Vision Compression with Semantic Post-Processing

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing video coding technologies prioritize human visual system characteristics for compression, which may not be optimal for machine vision tasks, leading to inefficiencies in machine task performance and computational challenges.

Innovation Solution

A post-processing network is applied to enhance semantic-related information in reconstructed visual signals without altering existing codecs, focusing on improving machine vision performance through methods like QP adaptive visual signal enhancement.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Productivity

If existing video coding technologies are used for compression, then compression efficiency is improved, but machine vision task performance deteriorates

Engineering Contradiction:
Improvecompression efficiencyVSAvoidmachine vision task performance
Core Design Contradiction:
ProductivityVSReliability

Solution Approach 1:

The processing pipeline is segmented into three independent stages: (1) standard video coding for compression, (2) post-processing network for semantic enhancement, and (3) machine vision tasks. This allows each stage to optimize for its specific function without compromising the others.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

A post-processing network is introduced as an intermediary component between the compressed video and machine vision tasks. This intermediary enhances semantic information in the reconstructed visual signal, bridging the gap between compression efficiency and machine vision performance requirements.

Inventive Principle:
Principle #24Intermediary (Mediator)

2Reliability

If post-processing network is added to enhance semantic information, then machine vision performance is improved, but device complexity increases

Engineering Contradiction:
Improvemachine vision task performanceVSAvoidprocessing system complexity
Core Design Contradiction:
ReliabilityVSDevice complexity

Solution Approach 1:

The post-processing network performs preliminary enhancement of semantic information in the reconstructed visual signal before it is fed to machine vision tasks. This preliminary action prepares the data in advance, reducing the computational burden on subsequent processing stages.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The system dynamically adjusts processing parameters based on the compression ratio and visual signal characteristics. The post-processing network adapts its operation based on the input signal quality, optimizing performance while managing computational resources efficiently.

Inventive Principle:
Principle #35Parameter changes

Data Source

PatentUS20250227311A1Method and compression framework with post-processing for machine vision
Publication Date: 2025.07.10 ALIBABA (CHINA) CO LTD
  • US20250227311A1 patent drawing
  • US20250227311A1 patent drawing
  • US20250227311A1 patent drawing

AI summary

A video processing method includes compressing and reconstructing an original visual signal to obtain a reconstructed visual signal; processing the reconstructed visual signal to obtain a post-processed visual signal; and feeding the post-processed visual signal to a machine task network