Geometric Transform in Neural Network Video Coding
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
The increasing demand for digital video bandwidth due to the growing number of connected user devices poses a challenge for efficient video compression and processing technologies.
Innovation Solution
A method and apparatus for processing video data that involves determining to modify a video unit by applying a video compression function and performing a conversion between visual media data and a bitstream based on the modified video unit, utilizing neural network-based coding tools and geometric transformations to enhance video compression efficiency.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Loss of energy
If traditional video compression methods are used, then bandwidth consumption is high, but compression efficiency is insufficient
Solution Approach 1:
The patent replaces traditional mechanical video compression algorithms with a neural network-based system. The neural network learns optimal compression patterns from training data and applies geometric transformations (rotation, flipping, cropping) to enhance compression efficiency, thereby reducing bandwidth consumption while improving productivity.
Solution Approach 2:
The patent dynamically adjusts geometric transformation parameters (rotation angles, flip directions, crop regions) based on the content characteristics of video blocks. By changing these parameters adaptively, the system optimizes compression efficiency for different video content types, achieving better bandwidth reduction without sacrificing quality.
2Measurement precision
If video resolution and quality are maintained, then bandwidth requirement increases, but compression ratio decreases
Solution Approach 1:
The patent applies geometric transformations preliminarily to video blocks before encoding. By pre-processing the video content with rotations, flips, and crops that optimize for compression, the system maintains visual quality while reducing the bandwidth required for transmission. The transformations are designed to preserve essential visual information.
Solution Approach 2:
The patent introduces asymmetric geometric transformations that are tailored to specific video content characteristics. Rather than applying uniform compression, the system uses asymmetric crop regions, selective rotations, and directional flips that adapt to the asymmetric nature of different video scenes, maintaining quality while reducing bandwidth.
3Productivity
If complex compression algorithms are applied, then compression efficiency improves, but computational complexity increases
Solution Approach 1:
The patent divides the video stream into smaller blocks and applies geometric transformations and neural network processing to each block independently or in small groups. This segmentation allows the complex compression task to be broken down into manageable units, improving compression efficiency while distributing computational complexity across multiple processing stages.
Solution Approach 2:
The patent applies geometric transformations selectively to only those video blocks that benefit most from them, rather than processing every block with the full neural network pipeline. This partial action approach maintains compression efficiency for complex regions while reducing computational complexity for simpler regions.
Data Source
AI summary
A mechanism for processing video data is disclosed. The mechanism determines to modify a video unit attendant to applying a video compression function. The modification may include applying a geometric conversion to the video unit. A conversion is performed between a visual media data and a bitstream based on the modified video unit.


