Video Compression Patch Metadata for Encoder-Decoder Consistency
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing video coding solutions using neural networks face challenges in ensuring replicability and flexibility of the inference process across encoders and decoders with different hardware and software constraints, particularly in handling patch sizes and receptive fields during neural network-based image processing.
Innovation Solution
The method involves obtaining metadata associated with the video stream that includes syntax elements defining the receptive field and allowable margins for the neural network-based image processing tool, allowing flexibility in processing patches larger or smaller than the original patch size, and signaling these margins in the form of metadata.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Adaptability or versatility
If neural network inference is applied with fixed patch sizes and receptive fields, then processing consistency is maintained, but flexibility to handle different hardware constraints and varying patch sizes is lost
Solution Approach 1:
The patent introduces dynamic parameters (receptive field size and margin values) that can be adjusted based on hardware constraints and patch sizes. The encoder and decoder both adapt their inference processes using these dynamic parameters signaled in the bitstream, allowing flexible handling of different patch sizes while maintaining output consistency through synchronized parameter usage.
Solution Approach 2:
The patent changes the parameters of the neural network inference process by signaling receptive field size and margin values in the bitstream. These parameter changes allow the system to adapt to different hardware constraints and patch sizes while ensuring that both encoder and decoder use identical parameters to maintain replicability and output consistency.
2Reliability
If encoder and decoder use identical neural network inference processes, then output replicability is ensured, but flexibility to accommodate different hardware constraints is reduced
Solution Approach 1:
The patent creates a universal inference framework where both encoder and decoder use the same neural network model and inference process. By signaling the necessary parameters (receptive field size, margins) in the bitstream, the system achieves multi-functionality that accommodates different hardware constraints while ensuring output replicability through identical processing on both sides.
Solution Approach 2:
The patent introduces metadata (receptive field size and margin values) as an intermediary that mediates between the neural network inference process and hardware constraints. These intermediaries are signaled in the bitstream and enable both encoder and decoder to adapt their inference processes to different hardware while maintaining identical output results.
3Adaptability or versatility
If fixed margins are used around patches for neural network inference, then processing simplicity is maintained, but ability to handle patches of varying sizes is limited
Solution Approach 1:
The patent performs preliminary action by signaling the margin values and receptive field size in the bitstream before the actual inference process. This allows the decoder to pre-configure its inference process with the correct parameters, enabling it to handle patches of varying sizes without increasing the complexity of the inference algorithm itself.
Data Source
AI summary
A method comprising obtaining a video stream; obtaining metadata associated with the video stream representative of allowable margins around a patch for an inference process of a neural network based image processing tool; and, decoding the video stream applying the neural network based image processing tool.


