Video Decoding with Hybrid Reference Filtering for Inter-Frame Prediction

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing video processing techniques using neural networks (NNs) suffer from local distortions and error propagation due to the block-by-block selection and correction of filtering modes, leading to poor inter-frame prediction performance and local distortion in reference frames.

Innovation Solution

A video decoding method that enhances video picture quality by using a hybrid coding framework to process reference frames with various loop filtering methods, increasing the diversity of reference frames through selective application of traditional and neural network-based filtering.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Manufacturing precision

If block-by-block selection and correction of filtering modes is performed, then overall video quality is improved, but local distortion occurs and error propagation is caused

Engineering Contradiction:
Improvevideo qualityVSAvoidlocal distortion and error propagation
Core Design Contradiction:
Manufacturing precisionVSObject-affected harmful factors

Solution Approach 1:

The patent segments the video frame into multiple blocks and applies different filtering modes to different blocks based on their characteristics. By dividing the frame into blocks and selectively applying filtering operations, the system achieves overall quality improvement while controlling local distortions through targeted processing rather than uniform application.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent implements local quality enhancement by analyzing the characteristics of each block and applying appropriate filtering modes accordingly. Different blocks receive different levels and types of filtering based on their specific properties, allowing optimal quality improvement in each region while preventing harmful effects from propagating across the entire frame.

Inventive Principle:
Principle #3Local quality

2Device complexity

If reference frames are selected without diversity, then the reference picture list is simple, but inter-frame prediction performance deteriorates

Engineering Contradiction:
Improvereference picture list complexityVSAvoidinter-frame prediction performance
Core Design Contradiction:
Device complexityVSManufacturing precision

Solution Approach 1:

The patent dynamically adjusts the reference picture list by inserting multiple reference frames with different characteristics and temporal layers. Instead of using a static, simple reference list, the system adaptively selects and inserts reference frames based on the current frame's characteristics, thereby improving inter-frame prediction performance while managing complexity through dynamic adjustment rather than static complexity.

Inventive Principle:
Principle #15Dynamics

Solution Approach 2:

The patent introduces temporal layer diversity into the reference picture list, adding a new dimension to reference frame selection. By incorporating reference frames from different temporal layers alongside traditional spatial references, the system enhances prediction capability without simply increasing the number of references, thus improving performance while controlling complexity through dimensional enrichment.

Inventive Principle:
Principle #17Another dimension (Dimensionality change)

Data Source

PatentUS20260019598A1Video decoding method, video processing device, medium, and product
Publication Date: 2026.01.15 ZTE CORP
  • US20260019598A1 patent drawing
  • US20260019598A1 patent drawing
  • US20260019598A1 patent drawing

AI summary

Provided are a video decoding method, a video processing device, a computer-readable storage medium, and a computer program product. The video decoding method includes acquiring reference frame information of a to-be-decoded video frame; determining a reference picture list of the to-be-decoded video frame based on the reference frame information and supplementary frame information; and decoding the to-be-decoded video frame based on the reference picture list to obtain a first reconstructed picture and a decoded picture of the to-be-decoded video frame.