Temporal Resampling and Restoration for Machine Vision Video
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing video coding and decoding technologies face challenges in efficiently managing bitrate requirements and adapting to different video consumption scenarios, particularly for machine vision tasks where video quality and resolution needs differ between human and machine applications.
Innovation Solution
Implementing temporal resampling and restoration methods in video coding and decoding systems, including determining sequence-level temporal restoration flags and indexes for resampling ratios, to optimize video data for specific use cases.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Quantity of substance
If video data is compressed using traditional coding techniques, then bitrate requirements are reduced, but video quality and temporal resolution are degraded
Solution Approach 1:
The patent segments video processing into two distinct paths: one for human consumption with full temporal resolution and one for machine consumption with reduced temporal resolution. This segmentation allows each path to be optimized independently, enabling aggressive compression for machine tasks without compromising human viewing quality.
Solution Approach 2:
The patent introduces a new dimension of temporal resampling by applying different temporal downsampling ratios specifically for machine vision tasks. This allows the system to reduce the temporal dimension selectively for machine consumption while maintaining full temporal resolution for human consumption, thereby reducing overall data volume without sacrificing quality where needed.
2Productivity
If temporal resolution is reduced for machine vision tasks, then data volume and transmission bandwidth are reduced, but video quality for human consumption is compromised
Solution Approach 1:
The patent applies local quality by providing different quality levels to different consumption contexts. Machine vision tasks receive video with reduced temporal resolution optimized for their specific requirements, while human consumption receives full-quality video. This localized optimization ensures each consumer type gets appropriate quality without unnecessary data overhead.
Solution Approach 2:
The patent implements dynamic temporal resampling where the temporal downsampling ratio can be adjusted based on the specific machine vision task requirements. This dynamic approach allows the system to adapt the compression level to match the actual needs of machine algorithms, improving transmission efficiency while maintaining sufficient quality for the intended application.
3Adaptability or versatility
If traditional video coding is used, then compatibility with existing systems is maintained, but adaptability to different consumption scenarios is limited
Solution Approach 1:
The patent creates a universal video coding framework that can serve multiple consumption scenarios simultaneously. By incorporating temporal resampling capabilities that can be selectively applied based on the target consumer type, the system achieves multi-functionality - it can produce video optimized for human consumption, machine consumption, or both concurrently, all through a single coding infrastructure.
Data Source
AI summary
This disclosure relates generally to video coding/decoding and particularly for signaling in temporal resampling and restoration in video coding and/or decoding systems. One method includes obtaining, by a device, a coded video bitstream; determining, by the device from the coded video bitstream, a sequence-level temporal restoration flag for a picture sequence; when the sequence-level temporal restoration flag indicates that temporal restoration is enabled for the picture sequence, determining, by the device from the coded video bitstream, an index indicating a temporal resampling ratio; and decoding, by the device, the coded video bitstream by generating temporal resampling data based on the temporal resampling ratio.


