Video Encoding Mode Selection for Motion Vector Optimization
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Conventional image encoding devices cannot switch between temporal direct mode and spatial direct mode on a per macroblock basis, leading to suboptimal encoding and increased code amount due to unnecessary motion vector encoding.
Innovation Solution
A moving image encoding device that determines the maximum block size and hierarchy depth for encoding, allowing selection of optimal direct modes for each block unit through a motion-compensated prediction process using a block dividing unit and encoding controlling unit, which selects and variable-length encodes the appropriate motion vector.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Ease of manufacture
If temporal direct mode is used for all macroblocks in a slice, then encoding simplicity is maintained, but encoding efficiency deteriorates due to inability to select optimal mode per macroblock
Solution Approach 1:
The slice is divided into multiple macroblocks, and each macroblock is independently evaluated to select the most appropriate direct mode (temporal or spatial). This segmentation allows different encoding strategies to be applied to different regions based on their specific characteristics, resolving the contradiction between maintaining simplicity and improving efficiency.
Solution Approach 2:
The encoding mode is made dynamic by allowing selection between temporal direct mode and spatial direct mode for each macroblock based on motion characteristics. The encoder dynamically determines which mode to use by comparing motion vector differences, enabling adaptability while maintaining a relatively simple overall structure.
2Productivity
If spatial direct mode is used for all macroblocks in a slice, then encoding efficiency improves for some cases, but code amount increases due to encoding of motion vector differences
Solution Approach 1:
The invention changes the parameter selection criterion by introducing a threshold-based comparison of motion vector differences. When the motion vector difference exceeds a predetermined threshold, spatial direct mode is selected; otherwise, temporal direct mode is used. This parameter-based decision mechanism optimizes the balance between encoding efficiency and code amount.
Solution Approach 2:
Different macroblocks are assigned different direct modes based on their local motion characteristics. Macroblocks with significant motion changes use spatial direct mode, while those with minimal changes use temporal direct mode. This local optimization prevents unnecessary encoding of motion vector differences, reducing overall code amount while maintaining efficiency where needed.
3Productivity
If direct mode selection is performed per macroblock basis, then encoding efficiency improves, but device complexity increases
Solution Approach 1:
Each macroblock performs self-evaluation by comparing its own motion vector characteristics against a predetermined threshold to determine the appropriate direct mode. This self-service approach eliminates the need for complex centralized control logic, allowing per-macroblock optimization while keeping device complexity relatively low.
Solution Approach 2:
The invention uses a simple predetermined threshold as a disposable criterion for mode selection, rather than employing complex, reusable decision-making structures. This cheap decision mechanism can be applied independently to each macroblock without requiring sophisticated device architecture, thus improving efficiency while limiting complexity increase.
Data Source
AI summary
When an encoding mode corresponding to one of blocks to be encoded into which an image is divided by a block dividing part 2 is an inter encoding mode which is a direct mode, a motion-compensated prediction part 5 selects a motion vector suitable for generation of a prediction image from one or more selectable motion vectors and also carries out a motion-compensated prediction process on the block to be encoded to generate a prediction image by using the motion vector, and outputs index information showing the motion vector to a variable length encoding part 13, and the variable length encoding unit 13 variable-length-encodes the index information.


