Video Encoding Geometric Reference Transformations
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
High Efficiency Video Coding (HEVC) and its successors face challenges in achieving accurate prediction, particularly for videos with fast and complex picture changes, leading to suboptimal quality at a given bitrate, especially when scaling factors need to be signaled in the bitstream, which is costly in terms of bits.
Innovation Solution
Applying geometric transformations such as scaling, rotation, shearing, reflection, and projection to reference pictures before matching procedures in both the encoder and decoder, allowing for improved prediction accuracy by selecting prediction areas with the lowest matching error, thereby reducing the need for extra signaling.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If traditional template matching or bilateral matching is used for prediction, then the encoding process is simple, but the prediction accuracy is insufficient for videos with fast and complex picture changes
Solution Approach 1:
The patent applies geometric transformations (scaling, rotation, shearing, reflection, projection) to reference pictures before performing matching procedures. This preliminary transformation prepares the reference pictures to better match the current picture geometry, improving prediction accuracy for fast and complex picture changes without significantly increasing encoding complexity
Solution Approach 2:
The patent transforms reference pictures using geometric parameters (scaling factors, rotation angles, shearing factors, reflection flags, projection parameters) to adapt them to the current picture's geometric characteristics. By changing these transformation parameters, the matching accuracy is improved while maintaining a manageable encoding complexity through selective application
2Measurement precision
If geometric transformations are applied to reference pictures before matching, then prediction accuracy is improved, but the bitrate increases due to signaling requirements
Solution Approach 1:
The patent extracts and applies only the necessary geometric transformations (scaling, rotation, shearing, reflection, projection) that are most beneficial for the current picture content. By selectively applying transformations rather than always applying all possible transformations, the bitrate overhead is reduced while maintaining improved matching accuracy where needed
Solution Approach 2:
The patent applies geometric transformations partially - only to the extent necessary to improve matching accuracy. The transformations are applied selectively based on the content characteristics, avoiding unnecessary transformation signaling and reducing bitrate overhead while still achieving improved prediction for complex picture changes
3Measurement precision
If scaling factors are signaled in the bitstream to improve prediction accuracy, then the prediction quality improves, but the bitrate increases due to signaling overhead
Solution Approach 1:
The patent merges the geometric transformation parameters (including scaling factors) into a unified transformation framework. By combining multiple transformation operations (scaling, rotation, shearing, reflection, projection) into a single integrated process, the signaling overhead is reduced compared to separately signaling each parameter, while maintaining improved prediction quality
Data Source
AI summary
A method (20) is disclosed performed in an encoder (40) for encoding video pictures into a video bit stream, the method (20) comprising: obtaining (21) a transformed version (2′; 12′, 13′) of a reference picture (2; 12, 13), by using a geometric transformation comprising at least one of: scaling, rotation, shearing, reflection, and projection; performing (22) a matching procedure at least once, the matching procedure comprising matching a reference matching area (6; 15, 16) of the reference picture (2; 12, 13) to a matching area (4; 16, 15) of a second picture (1; 13, 12) and matching a reference matching area (6′; 15′, 16′) of the transformed version (2′; 12′, 13′) to the matching area (4; 16, 15) of the second picture (1; 13, 12); and encoding (23) a block (3; 14) of the current picture (1; 11) by selecting for the block (3; 14) a first prediction area (5; 15, 16) based on the reference matching area (6; 15, 16) or a second prediction area (5′; 15′, 16′) based on the transformed reference matching area (6′; 15′, 16′), wherein the first and second prediction areas at least partly overlap the respective reference matching areas (6; 6′; 15, 16, 15′, 16′) and wherein the prediction area having lowest matching error to a corresponding matching area (4; 15, 16) of the second picture (1; 13, 12) is selected as prediction for the block. A corresponding method (30) in a decoder (50) is disclosed, and encoder (40), decoder (50), computer programs and computer program products.


