Adaptive Neural Network Video Down-sampling for Compression Efficiency

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing video coding techniques face challenges in efficiently managing down-sampling and up-sampling processes, particularly in determining optimal down-sampling ratios for each frame and component, which affects bandwidth usage and compression efficiency.

Innovation Solution

The proposed method involves down-sampling video units using neural network-based techniques, such as convolutional neural networks, on a frame-by-frame basis, allowing for different down-sampling ratios for various video units and components, and selecting the best performing down-sampling method based on quality metrics like PSNR or MS-SSIM.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Productivity

If conventional video coding techniques are used for down-sampling and up-sampling, then the processing is simpler and faster, but the bandwidth usage is higher and compression efficiency is lower

Engineering Contradiction:
Improvecompression efficiencyVSAvoidprocessing complexity
Core Design Contradiction:
ProductivityVSDevice complexity

Solution Approach 1:

The patent applies neural network-based down-sampling and up-sampling with adaptive down-sampling ratios to transform the processing parameters, achieving higher compression efficiency (BD-rate savings) while managing complexity through learned optimization rather than exhaustive search

Inventive Principle:
Principle #35Parameter changes

Solution Approach 2:

The patent replaces conventional mechanical interpolation filters with neural network-based processing, substituting traditional signal processing mechanics with learned patterns that achieve superior compression efficiency without linearly increasing complexity

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

2Manufacturing precision

If a fixed down-sampling ratio is used for all video units, then the processing is simpler and faster, but the compression quality is suboptimal

Engineering Contradiction:
Improvecompression qualityVSAvoidprocessing speed
Core Design Contradiction:
Manufacturing precisionVSProductivity

Solution Approach 1:

The patent implements dynamic down-sampling ratios that adapt to different video units, frames, and components rather than using fixed ratios, allowing the system to optimize compression quality for each specific case while the neural network learns to process these variations efficiently

Inventive Principle:
Principle #15Dynamics

Solution Approach 2:

The patent applies different down-sampling ratios to different video units, frames, and color components based on their specific characteristics, enabling localized optimization of compression quality without uniformly increasing processing complexity across the entire video stream

Inventive Principle:
Principle #3Local quality

3Productivity

If neural network-based down-sampling is applied to all video units, then the compression efficiency is maximized, but the processing complexity and computational load increase

Engineering Contradiction:
Improvecompression efficiencyVSAvoidcomputational energy
Core Design Contradiction:
ProductivityVSUse of energy by stationary object

Solution Approach 1:

The patent applies neural network-based processing selectively to different video units, frames, and components rather than uniformly to all content, achieving significant compression efficiency gains while reducing unnecessary computational energy expenditure on segments where simpler methods suffice

Inventive Principle:
Principle #16Partial or excessive action

Data Source

PatentUS20240236325A9Super resolution downsampling
Publication Date: 2024.07.11 DOUYIN VISION CO LTD
  • US20240236325A9 patent drawing
  • US20240236325A9 patent drawing
  • US20240236325A9 patent drawing

AI summary

A method of processing video data. The method includes down-sampling a video unit of a video prior to application of a super resolution (SR) process and performing a conversion between the video including the video unit and a bitstream of the video based on the video unit as down-sampled. A corresponding video coding apparatus and non-transitory computer-readable recording medium are also disclosed.