Joint Spatial Temporal Block Merge Mode for HEVC
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
In video compression, existing technologies face inefficiencies in representing object motion, particularly for irregularly shaped objects, leading to high bit usage and reduced coding efficiency due to the need to describe motion parameters for each individual block, which is challenging in high efficiency video coding (HEVC).
Innovation Solution
A method that determines a merge mode for a current block by analyzing motion vector differences between spatial and temporal neighboring blocks, allowing the current block to share motion parameters with either a spatially or temporally located neighboring block, eliminating the need to code and transmit individual motion parameters for each block, and using a candidate list that includes both spatial and temporal neighbors without requiring additional bits, flags, or indexes to indicate the merge mode.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Adaptability or versatility
If motion parameters are coded for each individual block, then motion representation flexibility is improved, but bit usage increases and coding efficiency decreases
Solution Approach 1:
The patent merges motion parameters of multiple blocks by identifying a current block and one or more neighboring blocks that share the same motion parameters. Instead of coding motion parameters for each block individually, the encoder groups blocks with identical motion characteristics and transmits a single set of motion parameters for the entire group, thereby reducing bit usage while maintaining motion representation flexibility.
Solution Approach 2:
The patent creates a universal motion parameter set that can be applied to multiple blocks simultaneously. By determining that a current block and neighboring blocks share the same motion parameters, the system establishes a multi-functional motion parameter that serves multiple blocks, eliminating redundant coding and reducing overall bit consumption.
2Measurement precision
If motion parameters are coded for each individual block, then motion accuracy is improved, but coding efficiency decreases
Solution Approach 1:
The patent combines multiple blocks into a single coding unit when they share identical motion parameters. By merging blocks with the same motion characteristics, the system maintains accurate motion representation for each block while improving coding efficiency through reduced parameter transmission. The encoder identifies neighboring blocks with matching motion parameters and groups them together, ensuring motion accuracy is preserved while productivity increases.
3Productivity
If spatial and temporal redundancy are leveraged, then coding efficiency is improved, but additional signaling overhead is required
Solution Approach 1:
The patent extracts and removes the need for additional signaling by utilizing implicit information. Instead of adding new signaling mechanisms to indicate merge mode or neighboring block selection, the system extracts motion parameters from already-decoded neighboring blocks and reference pictures. The encoder and decoder both independently determine the same current block based on identical criteria, eliminating the need for explicit signaling and avoiding additional overhead.
Solution Approach 2:
The patent implements a self-service mechanism where the decoder independently determines the current block using the same logic as the encoder. By leveraging already-decoded neighboring blocks and reference pictures, the system allows the decoder to autonomously reconstruct motion parameters without requiring additional signaling from the encoder. This self-service approach maintains coding efficiency while avoiding increased signaling overhead.
Data Source
AI summary
In one embodiment, a spatial merge mode or a temporal merge mode for a block of video content may be used in merging motion parameters. Both spatial and temporal merge parameters are considered concurrently and do not require utilization of bits or flags or indexing to signal a decoder. If the spatial merge mode is determined, the method merges the block of video content with a spatially-located block, where merging shares motion parameters between the spatially-located block and the block of video content. If the temporal merge mode is determined, the method merges the block of video content with a temporally-located block, where merging shares motion parameters between the temporally-located block and the block of video content.


