Neural Network Multi-Layer Video Coding for Adaptive Bit-Rate
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
The increasing demand for high-resolution video content, such as UltraHD and 8K, necessitates a reduction in video transmission bit-rate without compromising visual presentation quality, while also managing computational complexity and facilitating transitions between different video codecs.
Innovation Solution
A content-adaptive multi-layer coding approach using neural networks for downscaling and upscaling video content, optimizing encoding parameters to minimize residual signals and enhance coding gain, applicable to various video codecs like AVC, HEVC, VVC, and EVC, and delivery methods including DASH and HLS.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If video resolution is increased to UltraHD or 8K, then visual presentation quality is improved, but bandwidth requirements increase significantly
Solution Approach 1:
The patent segments video content into multiple layers with different resolutions (base layer and enhancement layers). Each layer is encoded separately and can be transmitted independently, allowing receivers to select appropriate layers based on available bandwidth while maintaining visual quality through progressive refinement.
Solution Approach 2:
The patent changes the resolution parameter across different video layers, creating a hierarchy from low to high resolution. This allows the system to adapt transmission quality to available bandwidth by selecting which layers to transmit, thereby maintaining visual quality when bandwidth permits while reducing bandwidth consumption when it is limited.
2Device complexity
If downscaling and upscaling operations are performed using traditional methods, then computational complexity is reduced, but coding gain and visual quality deteriorate
Solution Approach 1:
The patent replaces traditional mechanical downscaling and upscaling operations with neural network-based processing. The neural networks learn optimal transformation parameters from training data and apply content-adaptive processing that preserves visual quality while achieving the required resolution transformations, thereby reducing information loss compared to conventional methods.
Solution Approach 2:
The patent changes the approach to resolution transformation by using learned parameters from neural networks instead of fixed algorithms. The neural networks determine optimal downscaling and upscaling parameters based on content characteristics, achieving better coding gain while managing computational complexity through efficient network architectures.
3Adaptability or versatility
If multiple video layers are encoded and transmitted, then adaptability to different bandwidth conditions is improved, but device complexity increases
Solution Approach 1:
The patent segments video content into multiple resolution layers that can be independently encoded and transmitted. This segmentation enables receivers to select appropriate layers based on available bandwidth, improving adaptability while allowing encoding to be performed once at the source, thereby distributing complexity rather than concentrating it.
Solution Approach 2:
The patent creates a multi-layer video structure that serves multiple functions: it provides adaptability to different bandwidth conditions, enables progressive quality enhancement, and allows selective transmission of layers. This universal structure improves system versatility without proportionally increasing encoding complexity, as the same base layer can serve multiple enhancement purposes.
Data Source
AI summary
Methods, systems, and apparatuses are described for encoding video. Video content to be encoded and sent to a computing device may be downscaled into one or more layers. The one or more layers may represent one or more versions of the video content such as one or more versions encoded at different resolutions. The residuals between each layer and the base layer may be upscaled so that one or more parameters associated with optimizing the encoding of the one or more layers may be determined by one or more neural networks based on the downscaling and upscaling process. The residuals between each layer, the one or more parameters, and the base layer may be encoded and sent to a computing device for decoding and playback of the video content using any of the versions of the video content.


