Video Coding Tool Selection Using Machine Learning Models

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Conventional video coding techniques, such as MPEG-2, MPEG-4, ITU-T.263, ITU-T.264/AVC, HEVC, and VVC, suffer from low coding efficiency, which is undesirable for digital video applications.

Innovation Solution

A method and apparatus for video processing that utilizes a machine learning model to determine the optimal coding tool for a target video block, improving coding effectiveness and efficiency by selecting a more appropriate coding tool during the conversion between a video block and a bitstream.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Productivity

If conventional video coding techniques are used, then device complexity is reduced, but coding efficiency deteriorates

Engineering Contradiction:
Improvecoding efficiencyVSAvoiddevice complexity
Core Design Contradiction:
ProductivityVSDevice complexity

Solution Approach 1:

A machine learning model is introduced as an intermediary component between the video block and coding tool selection process. The model takes video block characteristics as input and outputs optimized coding tool selections, thereby improving coding efficiency without requiring fundamental changes to the underlying video coding standards and algorithms.

Inventive Principle:
Principle #24Intermediary (Mediator)

Solution Approach 2:

The invention changes the parameter selection approach by using machine learning to dynamically select coding tools based on video content characteristics. Instead of using fixed or rule-based parameter selection, the system adapts parameter choices (coding tools) according to learned patterns from training data, improving coding efficiency for different video types.

Inventive Principle:
Principle #35Parameter changes

2Productivity

If machine learning model is introduced for coding tool selection, then coding efficiency is improved, but device complexity increases

Engineering Contradiction:
Improvecoding efficiencyVSAvoiddevice complexity
Core Design Contradiction:
ProductivityVSDevice complexity

Solution Approach 1:

The video processing system is segmented into distinct functional components: the machine learning model for tool selection and the conventional video coding pipeline for actual encoding. This segmentation allows the complex ML functionality to be isolated and managed separately from the core coding operations, making the overall system more manageable despite increased complexity.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The machine learning model performs preliminary analysis of video block characteristics before the actual coding process. By pre-determining the optimal coding tools based on content analysis, the system avoids trial-and-error approaches during encoding, improving efficiency while keeping the actual coding phase simple and fast.

Inventive Principle:
Principle #10Preliminary action

Data Source

PatentUS20240244239A1Method, apparatus, and medium for video processing
Publication Date: 2024.07.18 DOUYIN VISION CO LTD
  • US20240244239A1 patent drawing
  • US20240244239A1 patent drawing
  • US20240244239A1 patent drawing

AI summary

Embodiments of the present disclosure provide a solution for video processing. A method for video processing is proposed. The method comprises: determining, during a conversion between a target video block of a video and a bitstream of the video, a target coding tool for the target video block by using a machine learning model; and performing the conversion by using the target coding tool. By taking the machine learning model into consideration in selecting the coding tool, a more proper coding tool can be selected. In this way, the coding performance can be enhanced.