Matrix-Based Intra Prediction with Boundary Downsampling for Video Coding

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing video coding technologies face challenges in efficiently handling high-resolution video data due to increased bandwidth demands and complexity in encoding and decoding processes, particularly in intra prediction methods.

Innovation Solution

Implementing matrix-based intra prediction (MIP) techniques that include boundary downsampling, matrix vector multiplication, and optional upsampling operations to enhance video coding efficiency, along with arithmetic coding for syntax elements, and adapting affine linear weighted intra prediction (ALWIP) modes for improved encoding and decoding processes.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If traditional intra prediction methods are used for high-resolution video, then video quality can be maintained, but bandwidth consumption and computational complexity increase significantly

Engineering Contradiction:
Improvevideo qualityVSAvoidbandwidth consumption
Core Design Contradiction:
Measurement precisionVSLoss of energy

Solution Approach 1:

The patent extracts and processes only the boundary samples of the current block rather than the entire block. By performing boundary downsampling on selected boundary samples and using these reduced samples for prediction, the method reduces the amount of data that needs to be transmitted and processed, thereby reducing bandwidth consumption while maintaining prediction accuracy for high-resolution video

Inventive Principle:
Principle #2Taking out (Extraction)

Solution Approach 2:

The patent segments the boundary samples into different groups (e.g., left boundary samples and top boundary samples) and applies different processing and prediction strategies to each segment. This segmentation allows for more efficient processing of boundary information, reducing computational complexity while preserving the essential features needed for accurate intra prediction

Inventive Principle:
Principle #1Segmentation

2Measurement precision

If traditional intra prediction methods are used for high-resolution video, then video quality can be maintained, but encoding and decoding complexity increase

Engineering Contradiction:
Improvevideo qualityVSAvoidencoding and decoding complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The patent extracts only the necessary boundary samples from the current block and performs downsampling on these extracted samples. This approach reduces the computational load for both encoding and decoding operations compared to processing the entire high-resolution block, while still maintaining the quality needed for accurate prediction

Inventive Principle:
Principle #2Taking out (Extraction)

Solution Approach 2:

The patent performs boundary downsampling as a preliminary step before the main prediction process. By pre-processing the boundary samples to create a reduced representation, the method simplifies subsequent prediction operations and reduces the overall computational complexity of the encoding and decoding processes

Inventive Principle:
Principle #10Preliminary action

3Device complexity

If boundary downsampling is performed on all boundary samples, then computational complexity is reduced, but prediction accuracy deteriorates

Engineering Contradiction:
Improvecomputational complexityVSAvoidprediction accuracy
Core Design Contradiction:
Device complexityVSMeasurement precision

Solution Approach 1:

The patent applies different processing qualities to different parts of the boundary samples. Instead of uniformly downsampling all boundary samples, the method selectively processes specific boundary samples with appropriate levels of detail, preserving local features that are most important for prediction accuracy while reducing complexity in less critical areas

Inventive Principle:
Principle #3Local quality

Solution Approach 2:

The patent performs downsampling on only a selected subset of boundary samples rather than all boundary samples. This partial action approach reduces computational complexity while maintaining sufficient prediction accuracy by focusing processing resources on the most influential boundary samples

Inventive Principle:
Principle #16Partial or excessive action

4Device complexity

If reduced boundary samples are used in upsampling operation, then computational complexity is reduced, but prediction quality deteriorates

Engineering Contradiction:
Improvecomputational complexityVSAvoidprediction quality
Core Design Contradiction:
Device complexityVSMeasurement precision

Solution Approach 1:

The patent introduces reduced boundary samples as an intermediary representation that bridges the original boundary samples and the final prediction block. These reduced samples serve as a compressed intermediate form that reduces computational complexity in the upsampling operation while still containing sufficient information to generate accurate predictions through the matrix-based prediction process

Inventive Principle:
Principle #24Intermediary (Mediator)

Data Source

PatentUS12375714B2Context coding for matrix-based intra prediction
Publication Date: 2025.07.29 DOUYIN VISION CO LTD
  • US12375714B2 patent drawing
  • US12375714B2 patent drawing
  • US12375714B2 patent drawing

AI summary

Devices, systems, and methods for digital video coding, which includes matrix-based intra prediction methods for video coding, are described. In a representative aspect, a method for video processing includes encoding a current video block of a video using a matrix intra prediction (MIP) mode in which a prediction block of the current video block is determined by performing, on previously coded samples of the video, a boundary downsampling operation, followed by a matrix vector multiplication operation, and selectively followed by an upsampling operation; and adding, to a coded representation of the current video block, a syntax element indicative of applicability of the MIP mode to the current video block using arithmetic coding in which a context for the syntax element is derived based on a rule.