Reduced Secondary Transform Coding for High-Resolution Video Compression
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
The increasing demand for high-resolution and high-quality images/videos, particularly in immersive media like VR and AR, necessitates a highly efficient image/video compression technique to reduce transmission and storage costs.
Innovation Solution
An image coding method and apparatus utilizing a reduced secondary transform (RST) with optimized transform kernel matrices and adaptive transform coefficient arrays based on intra prediction modes to enhance coding efficiency.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If high-resolution and high-quality images/videos are transmitted or stored using conventional methods, then image quality is maintained, but transmission cost and storage cost increase
Solution Approach 1:
The patent applies parameter changes by transforming image data from spatial domain to frequency domain using transform techniques (DCT, DST, KLT, etc.), and by applying quantization with different precision levels to different frequency coefficients. This transforms the representation parameters of image data to achieve compression while maintaining perceptual quality.
Solution Approach 2:
The patent extracts and removes redundant information from high-resolution images/videos through transform coding and quantization. By converting spatial correlations into frequency domain representations and selectively discarding less important high-frequency coefficients, the patent reduces data quantity while preserving essential visual information.
2Productivity
If conventional transform methods are used for image coding, then coding process is simple, but coding efficiency is insufficient
Solution Approach 1:
The patent implements dynamic adaptability by allowing the transform type and kernel to be selected based on prediction mode, block size, and other coding conditions. The coding system dynamically switches between different transform methods (DCT, DST, KLT) and applies adaptive secondary transforms to optimize coding efficiency for different content characteristics.
Solution Approach 2:
The patent segments the transform process into multiple stages: primary transform, quantization, and optional secondary transform. This segmentation allows each stage to be independently optimized and controlled, enabling flexible application of different transform methods at different stages based on coding requirements.
3Productivity
If transform coefficients are not optimized according to intra prediction mode, then processing is simpler, but transform efficiency decreases
Solution Approach 1:
The patent applies local quality by using different transform kernels and secondary transform parameters for different intra prediction modes. Each prediction mode (planar, DC, angular directions) has optimized transform coefficients that match the directional characteristics of the prediction, improving transform efficiency for each local case.
Solution Approach 2:
The patent changes transform parameters based on intra prediction mode by selecting different secondary transform kernels and applying mode-dependent coefficient scaling. This parameter adaptation optimizes the transform for the specific directional characteristics of each prediction mode.
Data Source
AI summary
A video decoding method according to the present document is characterized by comprising: a step for deriving transform coefficients through inverse quantization on the basis of quantized transform coefficients for a target block; a step for deriving modified transform coefficients on the basis of an inverse reduced secondary transform (RST) of the transform coefficients; and a step for generating a reconstructed picture on the basis of residual samples for the target block on the basis of an inverse primary transform of the modified transform coefficients, wherein the inverse RST using a transform kernel matrix is performed on transform coefficients of the upper-left 4×4 region of an 8×8 region of the target block, and the modified transform coefficients of the upper-left 4×4 region, upper-right 4×4 region, and lower-left 4×4 region of the 8×8 region are derived through the inverse RST.


