Transform-Based Image Coding With Adaptive Contexts for Compression

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

The increasing demand for high-resolution and high-quality images/videos, particularly in immersive media like VR and AR, leads to higher transmission and storage costs due to increased bit amounts, necessitating a more efficient image/video compression technique.

Innovation Solution

An image coding method involving inverse primary and non-separable transforms, using a transform index based on multiple transform selection (MTS) and context information to derive syntax element bin strings, enhancing coding efficiency.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Manufacturing precision

If high-resolution and high-quality images/videos are transmitted or stored, then image quality and resolution are improved, but transmission cost and storage cost increase due to increased bit amount

Engineering Contradiction:
Improveimage qualityVSAvoidbit amount
Core Design Contradiction:
Manufacturing precisionVSQuantity of substance

Solution Approach 1:

The patent applies segmentation by dividing the transform process into multiple stages: primary transform, secondary transform, and tertiary transform. Each stage processes the residual signal with different transform kernels, allowing progressive refinement of compression efficiency while managing bit amount for high-quality images

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent implements dynamic transform selection through multiple transform selection (MTS) and multiple secondary transform selection (MSSTS). The system dynamically chooses appropriate transform kernels based on block characteristics, prediction modes, and content features, optimizing compression efficiency adaptively for different image regions and qualities

Inventive Principle:
Principle #15Dynamics

2Productivity

If conventional transform methods are used for image coding, then device complexity is reduced, but coding efficiency decreases

Engineering Contradiction:
Improvecoding efficiencyVSAvoidtransform processing complexity
Core Design Contradiction:
ProductivityVSDevice complexity

Solution Approach 1:

The transform process is segmented into primary, secondary, and tertiary transforms, each handling different frequency components and block characteristics. This segmentation allows the system to achieve higher coding efficiency by selecting appropriate transform types for different regions without uniformly increasing complexity across the entire image

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent applies local quality by using different transform kernels for different blocks based on their characteristics. The multiple transform selection mechanism chooses from various kernel types (e.g., DCT, DST, KLT) depending on local image features, block size, and prediction mode, optimizing compression efficiency locally while managing overall complexity

Inventive Principle:
Principle #3Local quality

3Productivity

If transform index coding is performed without context adaptation, then device complexity is reduced, but coding efficiency decreases

Engineering Contradiction:
Improvetransform index coding efficiencyVSAvoidcontext modeling complexity
Core Design Contradiction:
ProductivityVSDevice complexity

Solution Approach 1:

The patent implements dynamic context adaptation for transform index coding. The context model is dynamically selected and updated based on previously decoded transform indices, block characteristics, and prediction modes. This dynamic approach improves coding efficiency by adapting to local statistics while managing complexity through selective context model usage

Inventive Principle:
Principle #15Dynamics

Data Source

PatentUS12382042B2Transform-based image coding method and device therefor
Publication Date: 2025.08.05 BEIJING XIAOMI MOBILE SOFTWARE CO LTD
  • US12382042B2 patent drawing
  • US12382042B2 patent drawing
  • US12382042B2 patent drawing

AI summary

An image decoding method according to the present document comprises a step for performing inverse first-order transform and inverse non-separable transform on a residual sample. The inverse non-separable transform is performed on the basis of a transform index indicating a predetermined transform kernel matrix, the inverse first-order transform is performed on the basis of a multiple transform selection (MTS) index indicating MTS for a horizontal transform kernel and a vertical transform kernel, and a syntax element bin string for the transform index is derived on the basis of first context information when a tree type for a split structure of a target block is not a single tree type and is derived on the basis of second context information when the tree type is the single tree type.