Video Encoder Motion Vector Wrap-Around for Omnidirectional Video
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing video codecs, such as H.265/HEVC, while supporting parallel processing, do not efficiently leverage this capability for omnidirectional video content, leading to visual artifacts and inefficiencies in decoding and encoding processes.
Innovation Solution
The proposed solution involves adapting the motion vector wrap-around mechanism to accommodate rotations and non-traditional tile arrangements, and optimizing affine prediction modes to reduce runtime costs and improve chroma handling, while also enhancing bitstream signaling to support diverse video profiles and levels.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Productivity
If motion vector wrap-around mechanism is adapted to accommodate rotations and non-traditional tile arrangements, then parallel processing efficiency is improved, but device complexity increases
Solution Approach 1:
The motion vector wrap-around mechanism is made adaptive to different tile arrangements and rotation angles. The decoder dynamically determines wrap-around behavior based on the specific tile configuration and rotation parameters, allowing efficient parallel processing without requiring a completely new decoding architecture for each scenario.
Solution Approach 2:
The system changes parameters such as tile width, tile height, and rotation angle to accommodate different omnidirectional video configurations. By parameterizing the wrap-around mechanism, the same decoder can handle various tile arrangements and rotation scenarios without fundamental structural changes.
2Speed
If affine prediction modes are optimized to reduce runtime costs, then processing speed is improved, but manufacturing precision deteriorates
Solution Approach 1:
The affine prediction is applied selectively to specific regions or blocks within the video content. By determining whether affine prediction is beneficial for each local region, the system achieves processing speed improvements in complex regions while maintaining standard prediction methods in simpler areas, thus preserving overall coding precision.
Solution Approach 2:
The optimization applies affine prediction modes partially - only when the computational benefit outweighs the precision cost. The system evaluates conditions under which affine prediction should be used, applying it selectively rather than universally, thus balancing speed improvement with precision maintenance.
3Productivity
If chroma handling is optimized in affine prediction modes, then coding efficiency is improved, but device complexity increases
Solution Approach 1:
The chroma handling is segmented into distinct processing stages within the affine prediction framework. By separating chroma subsampling, motion compensation, and prediction steps, the system achieves better coding efficiency through targeted optimizations while organizing the complexity into manageable, modular components that can be implemented systematically.
4Adaptability or versatility
If bitstream signaling is enhanced to support diverse video profiles and levels, then adaptability is improved, but loss of information increases
Solution Approach 1:
The bitstream signaling mechanism is designed to be universal, supporting multiple video profiles and levels through a unified parameter structure. By creating a multi-functional signaling framework that can represent different tile arrangements, rotation angles, and profile requirements in a single coherent syntax, the system achieves broad adaptability without requiring separate signaling paths for each profile type.
Data Source
AI summary
A video encoder according to embodiments is provided. The video encoder is configured for encoding a plurality of pictures of a video by generating an encoded video signal, wherein each of the plurality of pictures includes original picture data. The video encoder includes a data encoder configured for generating the encoded video signal including encoded picture data, wherein the data encoder is configured to encode the plurality of pictures of the video into the encoded picture data. Moreover, the video encoder includes an output interface configured for outputting the encoded picture data of each of the plurality of pictures. Furthermore, a video decoders, systems, methods for encoding and decoding, computer programs and encoded video signals according to embodiments are provided.


