Steerable Discrete Cosine Transform for Image Coding
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Conventional 2D Discrete Cosine Transform (DCT) is inefficient for digital images with arbitrarily shaped discontinuities, leading to higher bitrates and reconstruction artifacts due to its inability to effectively capture directional information.
Innovation Solution
The Steerable Discrete Cosine Transform (SDCT) is introduced, which modifies the orientation of the 2D DCT basis using graph-transform theory to adapt to the semantic of each image block, optimizing coding performance by determining a rotation angle for precise directional matching.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Productivity
If conventional 2D DCT is used for image compression, then the transform is simple and computationally efficient, but it produces higher bitrate and reconstruction artifacts when images contain arbitrarily shaped discontinuities
Solution Approach 1:
The patent applies dynamics by making the transform basis rotatable and adaptable to different orientations. The SDCT allows the basis vectors to be dynamically adjusted according to the dominant orientation of discontinuities in each image block, transforming the static conventional DCT into a dynamic, context-aware transform that adapts to local image characteristics.
Solution Approach 2:
The patent changes the parameter of basis vector orientation by introducing a rotation angle parameter. By varying this angle parameter to match the orientation of image discontinuities, the transform can optimize its alignment with the image content, thereby reducing artifacts and improving compression efficiency for directional features.
2Ease of operation
If conventional 2D DCT with fixed horizontal/vertical basis is used, then the implementation is straightforward, but it cannot effectively capture directional information for arbitrarily shaped discontinuities
Solution Approach 1:
The patent introduces dynamics by enabling the transform basis to rotate and adapt to different orientations. The SDCT computes transformed coefficients using basis vectors that can be rotated by an angle parameter, allowing the transform to dynamically align with the dominant orientation of discontinuities in each image block rather than being fixed to horizontal/vertical directions.
Solution Approach 2:
The patent changes the orientation parameter of the transform basis by introducing a rotation angle. This parameter allows the basis vectors to be rotated to match the orientation of image features, providing adaptability to different directional patterns while maintaining a systematic approach to implementation.
3Productivity
If the DCT basis is rotated to match image discontinuities, then coding efficiency improves, but the transform becomes more complex and requires determination of rotation angles
Solution Approach 1:
The patent applies dynamics by making the transform basis rotatable and adaptable to different orientations. The SDCT allows the basis vectors to be dynamically adjusted according to the dominant orientation of discontinuities in each image block, transforming the static conventional DCT into a dynamic, context-aware transform that adapts to local image characteristics.
Solution Approach 2:
The patent changes the parameter of basis vector orientation by introducing a rotation angle parameter. By varying this angle parameter to match the orientation of image discontinuities, the transform can optimize its alignment with the image content, thereby reducing artifacts and improving compression efficiency for directional features.
Data Source
Figure 1
Figure 2(a)~2(b)
Figure 3(a)~3(c)
AI summary
The present invention relates to a method and an apparatus for encoding and/or decoding digital images or video streams, wherein the encoding apparatus (11) comprises processing means (1100) configured for reading at least a portion of an image (f), determining a rotation angle (θ) on the basis of said portion of said image (f), determining rotated transform matrix (V ') which is the result of a rotation of at least one basis vector of a Discrete Cosine Transform matrix (V) by said rotation angle (θ), computing transformed coefficients (f^) on the basis of the pixel values contained in said portion of the image (f) and said rotated transform information (V'), outputting the transformed coefficients (f^) to said destination (1110).