AI Video Encoding Prediction for Latency Reduction

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing video encoding methods face challenges in achieving low latency and high video quality while minimizing bit rate, as they often require lookahead techniques that introduce latency and limit coding efficiency in real-time broadcasting scenarios.

Innovation Solution

A frame-encoding method that predicts characteristics of current blocks based on previously encoded frames, allowing for encoding without the need for lookahead buffering, by using artificial intelligence algorithms like supervised learning and neural networks to determine optimal coding modes and reduce latency.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Manufacturing precision

If lookahead buffering is used to improve video quality and optimize bit rate, then coding efficiency is improved, but latency increases

Engineering Contradiction:
Improvevideo qualityVSAvoidlatency
Core Design Contradiction:
Manufacturing precisionVSLoss of time

Solution Approach 1:

The patent replaces the mechanical lookahead buffering system with an AI-based prediction system. Instead of storing and analyzing future frames to optimize current encoding decisions, the system uses neural networks to predict optimal encoding parameters based on current and past frame information, eliminating the latency introduced by buffering while maintaining coding efficiency

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

Solution Approach 2:

The patent changes the fundamental parameter of how encoding decisions are made - transitioning from offline optimization based on future frame analysis to real-time prediction based on learned patterns from training data. This allows the system to make accurate encoding decisions without requiring lookahead buffering

Inventive Principle:
Principle #35Parameter changes

2Manufacturing precision

If complex encoding modes are used to improve video quality, then distortion is reduced, but computational resources increase

Engineering Contradiction:
Improvevideo qualityVSAvoidcomputational resources
Core Design Contradiction:
Manufacturing precisionVSUse of energy by moving object

Solution Approach 1:

The patent performs encoding optimization in advance during an offline training phase. The AI model learns optimal encoding strategies from training data, so that during real-time encoding, pre-computed predictions guide the selection of encoding modes. This preliminary action transfers computational burden from real-time processing to offline training, reducing real-time computational requirements while maintaining high video quality

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The system uses learned patterns from training data to automatically make encoding decisions without requiring complex real-time analysis. The AI model serves itself by having already learned the optimal encoding strategies during training, eliminating the need for computationally intensive real-time optimization

Inventive Principle:
Principle #25Self-service

Data Source

PatentUS20220345715A1Ai prediction for video compression
Publication Date: 2022.10.27 ATEME
  • US20220345715A1 patent drawing
  • US20220345715A1 patent drawing
  • US20220345715A1 patent drawing

AI summary

A method for encoding a first image within a first set of images, in which the first image is cut into blocks, each block being encoded according to one among a plurality of coding modes, is proposed, which comprises, for a current block of the first image, the determination, on the basis of at least one second image distinct from the first image and previously encoded according to an encoding sequence of the images of the first set of images, of a prediction of a feature of the current block in one or more third images from the first set of images distinct from the first image and not yet encoded according to the encoding sequence, and the use of the prediction to encode the current block while minimizing a flow-distortion criterion.