Neural Network Front-End Architecture for Video Coding
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing video coding techniques struggle to efficiently compress high-quality video data while maintaining visual quality, due to the large amounts of data required and the burden this places on communication networks and devices.
Innovation Solution
An end-to-end machine learning-based image and video coding system, specifically designed for YUV 4:2:0 input formats, which uses a front-end architecture with separate neural network layers for the luminance and chrominance channels, combined using a 1×1 convolutional layer for improved coding performance.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Quantity of substance
If traditional video coding techniques are used to compress video data, then the data volume is reduced, but video quality degradation occurs
Solution Approach 1:
The video data is segmented into luminance (Y) and chrominance (UV) channels, with the luminance channel processed at full resolution and the chrominance channel subsampled at half resolution. This segmentation allows independent optimization of each channel's compression strategy, reducing overall data volume while preserving visual quality through selective detail retention in the more important luminance channel.
2Manufacturing precision
If high quality video with high resolution and frame rate is provided, then consumer demand is met, but the burden on communication networks and devices increases
Solution Approach 1:
The patent changes the resolution parameter of the chrominance channel from full resolution to half resolution (subsampling), while maintaining the luminance channel at full resolution. This parameter change reduces the total data volume by approximately 50% for the chrominance portion, significantly lowering network transmission burden and device processing requirements while preserving perceived video quality, as the human visual system is less sensitive to chrominance details.
Data Source
AI summary
Techniques are described herein for processing video data using a neural network system. For instance, a process can include generating, by a first convolutional layer of an encoder sub-network of the neural network system, output values associated with a luminance channel of a frame. The process can include generating, by a second convolutional layer of the encoder sub-network, output values associated with at least one chrominance channel of the frame. The process can include generating, by a third convolutional layer based on the output values associated with the luminance channel of the frame and the output values associated with the at least one chrominance channel of the frame, a combined representation of the frame. The process can further include generating encoded video data based on the combined representation of the frame.


