Adaptive Spatial Resampling for Machine Vision Encoding
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Conventional image/video compression techniques are not optimized for machine vision tasks, as they focus on human perception rather than machine understanding of visual data, leading to inefficient data transmission and storage.
Innovation Solution
The method involves encoding a video sequence into a bitstream by determining spatial information features at various resolutions, performing adaptive spatial resampling using a pre-trained cluster model, and encoding the resampled video sequence.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Illumination intensity
If conventional image/video compression techniques are used, then human perception quality is improved, but machine vision task performance deteriorates
Solution Approach 1:
The patent changes the compression parameters from human-perception-based metrics to machine-vision-based metrics. Specifically, it uses spatial information features (such as edge density, texture complexity, and frequency content) as the basis for compression, rather than traditional human visual system models. This parameter change enables the compression algorithm to preserve features that are critical for machine vision tasks like object detection and recognition, while still achieving efficient compression.
2Productivity
If spatial resampling is performed without adaptation, then compression efficiency is improved, but performance on diverse video spatial complexities deteriorates
Solution Approach 1:
The patent implements dynamic spatial resampling by computing spatial information features for each video block or region and adapting the resampling strategy accordingly. Regions with high spatial information (complex textures, edges, or details) are resampled at lower rates, while regions with low spatial information (smooth areas) are resampled at higher rates. This dynamic adaptation ensures that compression efficiency is maintained across diverse video content with varying spatial complexities, preventing performance degradation on any particular type of video.
Data Source
AI summary
A method of encoding a video sequence into a bitstream, the method includes receiving a video sequence comprising a plurality of pictures; determining one or more spatial information features of the video sequence corresponding to various resolution; performing spatial resampling on the video sequence according to the one or more spatial information features and a pre-trained cluster model; and encoding the spatial resampled video sequence into the bitstream.


