Lightweight Spatial Upsampling for Machine Vision Coding
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Conventional image/video compression techniques focus on human perception, failing to optimize performance for machine vision tasks, which require different compression methods to efficiently process and transmit large volumes of video data for applications like autonomous driving and intelligent transportation.
Innovation Solution
The development of lightweight spatial upsampling models for decoding and encoding video streams, reducing the number of coding bits in the spatial upsampling model parameters to less than a predetermined threshold, enabling efficient compression and transmission while maintaining desired video quality for machine vision applications.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If conventional image/video compression techniques are used, then human perception quality is improved, but machine vision task performance deteriorates
Solution Approach 1:
The patent applies local quality by using different compression strategies for different parts of the image data. Specifically, it uses a spatial upsampling model with limited coding bits for regions important to machine vision tasks while maintaining higher quality in regions critical for human perception, thereby resolving the contradiction between human perception quality and machine vision performance
Solution Approach 2:
The patent changes the parameter of coding bits allocation for spatial upsampling model parameters, constraining it to be less than a predetermined threshold. This parameter change enables the system to achieve machine vision task performance while controlling compression efficiency, resolving the contradiction between human perception quality and machine vision reliability
2Measurement precision
If spatial upsampling model parameters are compressed with high precision, then reconstructed picture quality is improved, but coding bit length increases
Solution Approach 1:
The patent directly applies parameter changes by setting a constraint on the total length of coding bits for spatial upsampling model parameters to be less than a predetermined threshold. This threshold is determined based on desired reconstructed picture quality, thereby achieving the optimal balance between coding efficiency and reconstruction quality
Solution Approach 2:
The patent uses partial action by applying compression to the spatial upsampling model parameters rather than compressing all image data at maximum precision. This selective compression approach maintains sufficient quality for machine vision tasks while significantly reducing the total coding bit length
3Loss of energy
If video data is compressed for efficient transmission, then bandwidth cost is reduced, but machine vision task accuracy deteriorates
Solution Approach 1:
The patent applies local quality by focusing compression efforts on aspects of video data that are less critical for machine vision tasks while preserving quality in regions and features important for tasks like object detection and tracking. This selective approach reduces bandwidth consumption while maintaining task accuracy
Solution Approach 2:
The patent uses copying by creating a reconstructed version of the video data that is optimized for machine vision tasks rather than attempting to perfectly replicate the original. The spatial upsampling model generates a sufficient copy that maintains the essential features needed for machine vision accuracy while using fewer bits for transmission
Data Source
AI summary
The present disclosure provides spatial upsampling models used for processing video data suitable for machine vision tasks. An exemplary decoding method includes: receiving a bitstream; and decoding, using coded information of the bitstream, one or more pictures, wherein the decoding includes: generating one or more decompressed pictures by decompressing one or more compressed pictures included in the bitstream; and performing spatial upsampling on the one or more decompressed pictures by a spatial upsampling model to obtain one or more reconstructed pictures, respectively, wherein a total length of coding bits of parameters of the spatial upsampling model is less than a threshold that is pre-determined based on a desired quality of the reconstructed pictures.


