Integer Transform Video Coding Relaxing Orthogonality
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Conventional video coding methods do not effectively balance vector length, block size, transform element size, and inner products, leading to inefficiencies in data representation and picture quality.
Innovation Solution
A method for video coding that employs a transform matrix formed by basis vectors derived from discrete cosine transform (DCT) or Karhunen-Loève transform (KLT), with relaxed orthogonality and norm requirements, to generate transform coefficients that are close to but not exactly orthogonal and have similar norms, while using integer values to improve coding efficiency and reduce complexity.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If conventional video coding methods use strict orthogonality and norm equality requirements for transform matrices, then transform accuracy is improved, but device complexity and computational complexity increase
Solution Approach 1:
The patent changes the parameters of the transform matrix by relaxing the strict orthogonality and norm equality requirements. Instead of requiring exact orthogonality (inner products equal to zero) and equal norms, the patent allows for approximate orthogonality and similar norms, which simplifies the computational requirements while maintaining adequate transform accuracy for video compression
Solution Approach 2:
The patent employs integer-valued basis vectors that are computationally cheaper to process than floating-point or complex orthogonal vectors. These simplified basis vectors can be derived from standard transforms like DCT or KLT but with relaxed precision requirements, reducing the computational burden while maintaining sufficient accuracy for practical video coding applications
2Productivity
If conventional video coding methods use large block sizes for transformation, then data representation efficiency is improved, but picture quality deteriorates due to loss of detail
Solution Approach 1:
The patent segments the video data into multiple blocks of different sizes (e.g., 4x4, 8x8, 16x16 pixels) and applies the transform selectively to different blocks. This allows finer blocks to preserve local details and textures while larger blocks efficiently represent smooth areas, thereby maintaining both compression efficiency and picture quality
Solution Approach 2:
The patent applies different transform configurations to different regions of the image based on local characteristics. By using smaller block sizes in areas requiring detail preservation and larger block sizes in areas where compression is more effective, the system achieves local optimization of both quality and efficiency
3Measurement precision
If conventional video coding methods use floating-point transform coefficients, then transform accuracy is improved, but bandwidth requirements increase
Solution Approach 1:
The patent uses integer-valued transform coefficients instead of floating-point coefficients. This substitution reduces the precision requirements and allows for more compact representation of the transform results, thereby reducing the bandwidth required for transmitting compressed video data while maintaining adequate quality
Solution Approach 2:
The patent changes the parameter precision from floating-point to integer values in the transform coefficients. This parameter change reduces the bit depth required to represent each coefficient, leading to lower overall bandwidth requirements while maintaining sufficient accuracy for practical video compression applications
Data Source
Figure 1
Figure 2
Figure 3
AI summary
A video coding/decoding system, method and computer program product employ an integer transform matrix for transforming to/from transform coefficients and residual pixel data in moving pictures by a set of semi- orthonormal basis vectors. The basis vectors are derived from conventional DCT or KTL matrixes, but relaxes to some extent the requirements for orthogonality, norm equality and element size limitation. In this way improved coding efficiency and lower complexity compared to previously used integer transforms are possible.