Output Layer Sets Signaling for Adaptive Video Resolution
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current video coding and decoding technologies face challenges in efficiently signaling and managing adaptive resolution changes within coded video streams, particularly in scenarios requiring separate resolution settings for multiple semantically independent source pictures, such as 360-degree video or surveillance applications, where different parts of a scene may require distinct adaptive resolution settings.
Innovation Solution
The method involves signaling output layer sets in coded video data, allowing for the specification and decoding of multiple layers with adaptive resolution changes, enabling flexible resampling of reference pictures and improved encoding, decoding, and display of video layers by using syntax elements to indicate output layers and their corresponding resolutions.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Productivity
If adaptive resolution changes are implemented for multiple semantically independent source pictures, then video encoding and decoding efficiency is improved, but device complexity increases due to the need to manage multiple layers with different resolution settings
Solution Approach 1:
The video stream is divided into multiple layers, each representing a different resolution setting for semantically independent source pictures. Each layer can be independently decoded and output, allowing the system to handle complex adaptive resolution scenarios by breaking them down into manageable segments.
Solution Approach 2:
The output layer set syntax elements are designed to be universal and applicable to multiple different coding scenarios, including 360-degree video, surveillance applications, and standard video sequences. The same mechanism handles various resolution change patterns without requiring scenario-specific implementations.
2Adaptability or versatility
If multiple output layers with different resolutions are signaled, then adaptability for different scene activities is improved, but loss of information increases due to the overhead of signaling syntax elements for each layer
Solution Approach 1:
Multiple output layer configurations are merged into a single unified syntax structure. The output layer set syntax elements define a comprehensive set of layers that can accommodate different resolution requirements, allowing the system to select appropriate layers based on scene activity without requiring separate signaling for each scenario.
Solution Approach 2:
The syntax elements allow dynamic parameter changes between layers, enabling the system to adapt resolution settings based on scene activity. By encoding parameter differences between layers rather than absolute values for each layer, the bitstream overhead is reduced while maintaining full adaptability.
3Manufacturing precision
If reference pictures are resampled to different resolutions, then manufacturing precision of video output is improved, but loss of energy increases due to the computational requirements of resampling operations
Solution Approach 1:
Reference pictures are resampled in advance during the encoding process and stored at multiple resolutions. This preliminary resampling allows the decoder to directly use pre-prepared reference pictures at the required resolution without performing computationally intensive resampling operations during decoding, significantly reducing decoder energy consumption.
Solution Approach 2:
Multiple copies of reference pictures at different resolutions are created and stored. Instead of resampling a single reference picture multiple times during decoding, the system maintains pre-generated copies at various resolutions, allowing the decoder to select the appropriate copy based on the current output layer requirements without additional computational overhead.
Data Source
AI summary
A method, computer program, and computer system is provided for signaling output layer sets in a coded video stream. Video data having multiple layers is received. One or more syntax elements are identified. The syntax elements specify one or more output layer sets corresponding to output layers from among the multiple layers of the received video data. The one or more output layers corresponding to the specified output layer sets are decoded and displayed.


