Neural Network Sign Prediction for Video Coding Efficiency

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current video encoding and decoding technologies, such as AOMedia Video 1 (AV1), do not effectively utilize deep learning methods for predicting sign values of transform coefficients, limiting coding efficiency.

Innovation Solution

A method and system that uses neural networks to identify reference samples and predict sign values of transform coefficients for video data, enabling improved encoding and decoding by leveraging deep learning architectures like convolutional neural networks and dense residual networks.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Productivity

If traditional video encoding methods are used, then device complexity is low, but coding efficiency is insufficient

Engineering Contradiction:
Improvecoding efficiencyVSAvoiddevice complexity
Core Design Contradiction:
ProductivityVSDevice complexity

Solution Approach 1:

The patent replaces traditional mechanical video encoding algorithms with a neural network-based deep learning system. The neural network model processes video data through learned patterns and transformations, substituting conventional signal processing methods with intelligent computational approaches that achieve superior coding efficiency while managing complexity through specialized hardware acceleration.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

2Productivity

If neural networks are introduced for sign value prediction, then coding efficiency improves, but device complexity increases

Engineering Contradiction:
Improvecoding efficiencyVSAvoiddevice complexity
Core Design Contradiction:
ProductivityVSDevice complexity

Solution Approach 1:

The patent segments the video encoding process into distinct functional modules, with the neural network specifically dedicated to sign value prediction of transform coefficients. This segmentation allows the complex neural network operation to be isolated and optimized independently, integrating deep learning capabilities without overwhelming the entire encoding system with complexity.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The neural network model is designed to perform multiple functions within the video encoding pipeline, including sign prediction, feature extraction, and pattern recognition. This multi-functionality consolidates several processing tasks into a single versatile component, improving coding efficiency while reducing the overall system complexity compared to having separate specialized modules for each function.

Inventive Principle:
Principle #6Universality (Multi-functionality)

Data Source

PatentUS12143573B2Neural network based coefficient sign prediction field
Publication Date: 2024.11.12 TENCENT AMERICA LLC
  • US12143573B2 patent drawing
  • US12143573B2 patent drawing
  • US12143573B2 patent drawing

AI summary

A method, computer program, and computer system is provided for coding video data. Reference samples and magnitudes of transform coefficients corresponding to a current block of video data from an input to a neural network are identified. Sign values associated with the transform coefficients are predicted using neural networks. The video data is encoded/decoded based on the predicted sign values.