Chroma Upsampling for Resolution-Matched Neural Video Coding
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing video coding technologies face challenges in efficiently compressing and decompressing video data with subsampled chroma formats, leading to inefficiencies and potential quality loss due to mismatched resolutions between luma and chroma components.
Innovation Solution
The method involves up-sampling the chroma component to match the resolution of the luma component and encoding both within a trained network, preserving spatial correlation and enabling efficient encoding and decoding by utilizing cross-color correlation.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Loss of substance
If chroma component is subsampled to reduce data size, then compression ratio is improved, but picture quality deteriorates due to resolution mismatch between luma and chroma components
Solution Approach 1:
The patent changes the resolution parameter of the chroma component by up-sampling it to match the luma component's resolution. This allows the trained network to process both components at the same resolution, improving picture quality while maintaining compression efficiency through the learned representation.
Solution Approach 2:
The trained network acts as an intermediary that processes both luma and chroma components together at matched resolutions. This intermediary processing step enables efficient handling of the resolution mismatch while preserving picture quality through learned cross-color correlations.
2Device complexity
If chroma component resolution is reduced for efficient encoding, then encoding complexity is reduced, but spatial correlation between luma and chroma is lost
Solution Approach 1:
The patent changes the chroma resolution parameter to match the luma resolution, enabling the trained network to process both components at the same resolution. This preserves spatial correlation information while the trained network efficiently handles the processing complexity through learned representations.
Solution Approach 2:
The trained network provides a universal processing framework that handles both luma and chroma components together at matched resolutions. This multi-functional approach preserves spatial correlations while efficiently managing encoding complexity through a single integrated model.
3Productivity
If different resolution handling is used for luma and chroma components, then encoding efficiency is improved, but trained network performance deteriorates due to mismatched input dimensions
Solution Approach 1:
The patent changes the chroma component resolution parameter to match the luma component's resolution before input to the trained network. This ensures matched input dimensions for the network while maintaining encoding efficiency through the learned compressed representation of both components.
Solution Approach 2:
The up-sampling process acts as an intermediary step that prepares the chroma component with matched resolution before network input. This intermediary processing ensures the trained network receives properly dimensioned inputs, maintaining reliability while preserving encoding efficiency through the learned representation.
Data Source
AI summary
The present disclosure relates to video encoding and decoding, and in particular to handling of chroma subsampled formats in machine-learning-based video coding. Corresponding apparatuses and methods enable the processing for encoding and decoding of a respective picture portion that includes a luma component and a chroma component with a resolution lower than the luma component. In order to handle such different sized luma-chroma channels, the chroma component is up-sampled such that the obtained up-sampled chroma component has a resolution matching the one of the luma component. The luma and the up-sampled chroma component are then encoded into a bitstream. To reconstruct the picture portion, the luma component and an intermediate chroma component matching the resolution of the luma component are decoded from the bitstream, followed by down-sampling the intermediate chroma component. Thus, sub-sampled chroma formats may be handled by an autoencoder/autodecoder framework, while preserving the luma channel.


