Adaptive Spatial Resampling for Machine Vision Encoding

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Conventional image/video compression techniques are not optimized for machine vision tasks, as they focus on human perception rather than machine understanding of visual data, leading to inefficient data transmission and storage.

Innovation Solution

The method involves encoding a video sequence into a bitstream by determining spatial information features at various resolutions, performing adaptive spatial resampling using a pre-trained cluster model, and encoding the resampled video sequence.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Illumination intensity

If conventional image/video compression techniques are used, then human perception quality is improved, but machine vision task performance deteriorates

Engineering Contradiction:
Improveimage qualityVSAvoidmachine vision task performance
Core Design Contradiction:
Illumination intensityVSProductivity

Solution Approach 1:

The patent changes the compression parameters from human-perception-based metrics to machine-vision-based metrics. Specifically, it uses spatial information features (such as edge density, texture complexity, and frequency content) as the basis for compression, rather than traditional human visual system models. This parameter change enables the compression algorithm to preserve features that are critical for machine vision tasks like object detection and recognition, while still achieving efficient compression.

Inventive Principle:
Principle #35Parameter changes

2Productivity

If spatial resampling is performed without adaptation, then compression efficiency is improved, but performance on diverse video spatial complexities deteriorates

Engineering Contradiction:
Improvecompression efficiencyVSAvoidperformance across diverse video spatial complexities
Core Design Contradiction:
ProductivityVSAdaptability or versatility

Solution Approach 1:

The patent implements dynamic spatial resampling by computing spatial information features for each video block or region and adapting the resampling strategy accordingly. Regions with high spatial information (complex textures, edges, or details) are resampled at lower rates, while regions with low spatial information (smooth areas) are resampled at higher rates. This dynamic adaptation ensures that compression efficiency is maintained across diverse video content with varying spatial complexities, preventing performance degradation on any particular type of video.

Inventive Principle:
Principle #15Dynamics

Data Source

PatentUS20250113037A1Methods and non-transitory computer readable storage medium for adaptive spatial resampling towards machine vision
Publication Date: 2025.04.03 ALIBABA (CHINA) CO LTD
  • US20250113037A1 patent drawing
  • US20250113037A1 patent drawing
  • US20250113037A1 patent drawing

AI summary

A method of encoding a video sequence into a bitstream, the method includes receiving a video sequence comprising a plurality of pictures; determining one or more spatial information features of the video sequence corresponding to various resolution; performing spatial resampling on the video sequence according to the one or more spatial information features and a pre-trained cluster model; and encoding the spatial resampled video sequence into the bitstream.