Deep Neural Network Residual Processing in Video Coding

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current video coding standards, such as HEVC, face challenges in improving video quality and coding efficiency due to limitations in processing residual signals and artifacts, particularly in the prediction, reconstruction, and filtering processes.

Innovation Solution

The integration of Deep Neural Networks (DNNs) in the video coding system, where DNNs process reconstructed residuals, prediction outputs, and filtering processes to enhance video encoding and decoding, allowing for pixel value restoration and residual error reduction by signaling the absolute value of residual pixels and their signs in the bitstream.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Manufacturing precision

If traditional video coding standards (HEVC) are used for processing residual signals, then coding efficiency is maintained at current levels, but video quality improvement is limited due to inability to effectively reduce residual errors and artifacts

Engineering Contradiction:
Improvevideo qualityVSAvoidresidual errors
Core Design Contradiction:
Manufacturing precisionVSLoss of energy

Solution Approach 1:

The patent replaces traditional mechanical/video coding processing systems with a neural network-based system. The neural network learns optimal residual signal processing patterns from training data, substituting conventional block-based transforms and quantization with adaptive, data-driven processing that reduces residual errors more effectively while maintaining coding efficiency

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

Solution Approach 2:

The patent changes the processing parameters from fixed, predefined transform coefficients and quantization steps to dynamic, learned parameters from the neural network. The network outputs adaptive residual corrections that adjust to local image characteristics, enabling better quality reconstruction with reduced energy loss in residual signals

Inventive Principle:
Principle #35Parameter changes

2Manufacturing precision

If Deep Neural Networks are integrated into video coding to reduce residual errors and improve video quality, then manufacturing precision (video quality) improves, but device complexity increases due to additional processing layers

Engineering Contradiction:
Improvevideo qualityVSAvoidprocessing complexity
Core Design Contradiction:
Manufacturing precisionVSDevice complexity

Solution Approach 1:

The patent segments the video coding process into distinct neural network processing stages: residual error prediction networks, reconstruction networks, and refinement networks. Each segment handles specific aspects of quality improvement, allowing modular implementation and reducing overall system complexity while achieving high video quality through coordinated processing

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent introduces intermediary neural network components that bridge traditional coding stages. These intermediary networks process residual signals and generate correction terms that are integrated into the conventional coding framework, enabling quality improvement without completely replacing the existing system architecture

Inventive Principle:
Principle #24Intermediary (Mediator)

3Manufacturing precision

If neural networks process reconstructed residuals and prediction outputs to reduce residual errors, then video quality improves, but processing time increases due to additional computational steps

Engineering Contradiction:
Improvevideo qualityVSAvoidprocessing time
Core Design Contradiction:
Manufacturing precisionVSLoss of time

Solution Approach 1:

The patent performs preliminary neural network processing on residual signals during the encoding phase to pre-correct prediction errors. By applying the neural network residual correction before final reconstruction, the system reduces the computational burden on subsequent decoding stages while achieving quality improvement, effectively preparing the data in advance to minimize later processing time

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The patent merges neural network processing with traditional video coding operations by integrating residual correction networks into the existing transform and quantization pipeline. This merging allows simultaneous execution of conventional coding steps and neural network processing, reducing overall processing time through operational fusion rather than sequential execution

Inventive Principle:
Principle #5Merging (Combining)

4Adaptability or versatility

If DNN parameters are pre-defined and control flags are used to enable/disable processing, then adaptability to different video characteristics improves, but device complexity increases due to parameter management and control mechanisms

Engineering Contradiction:
Improvevideo characteristic adaptationVSAvoidparameter management complexity
Core Design Contradiction:
Adaptability or versatilityVSDevice complexity

Solution Approach 1:

The patent designs universal neural network architectures that can process multiple types of video content (different resolutions, formats, and characteristics) using the same core network structure. Pre-defined parameters are configured to handle various video types, and control flags enable selective activation of appropriate processing modes, providing broad adaptability without requiring separate specialized systems for each video type

Inventive Principle:
Principle #6Universality (Multi-functionality)

Data Source

PatentUS11589041B2Method and apparatus of neural network based processing in video coding
Publication Date: 2023.02.21 MEDIATEK INC
  • US11589041B2 patent drawing
  • US11589041B2 patent drawing
  • US11589041B2 patent drawing

AI summary

A method and apparatus of video coding incorporating Deep Neural Network are disclosed. A target signal is processed using DNN (Deep Neural Network), where the target signal provided to DNN input corresponds to the reconstructed residual, output from the prediction process, the reconstruction process, one or more filtering processes, or a combination of them. The output data from DNN output is provided for the encoding process or the decoding process. The DNN can be used to restore pixel values of the target signal or to predict a sign of one or more residual pixels between the target signal and an original signal. An absolute value of one or more residual pixels can be signalled in the video bitstream and used with the sign to reduce residual error of the target signal.