Adaptive GOP Length Determination for HEVC Video Encoding

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

The fixed Group of Pictures (GOP) length in high efficiency video coding (HEVC) encoders can lead to reduced encoding quality in certain scenarios.

Innovation Solution

A method and apparatus for adaptively determining the GOP length based on target characteristics of a video frame sequence, using a trained model to predict the optimal GOP length for each video frame.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Reliability

If a fixed GOP length is used in HEVC encoder, then the encoding process is simple and stable, but the encoding quality deteriorates in scenarios with large fluctuations

Engineering Contradiction:
Improveencoding qualityVSAvoidGOP length determination complexity
Core Design Contradiction:
ReliabilityVSDevice complexity

Solution Approach 1:

The patent transforms the static fixed GOP length into a dynamic adaptive GOP length that changes according to video content characteristics. The system dynamically adjusts GOP length based on motion complexity, texture complexity, and other video frame features, allowing the encoding system to adapt to different scenarios and maintain high encoding quality while managing complexity through automated analysis.

Inventive Principle:
Principle #15Dynamics

Solution Approach 2:

The patent changes the GOP length parameter from a fixed value to a variable that is determined by analyzing video frame characteristics. By computing complexity metrics and using these as inputs to determine GOP length, the system optimizes the parameter based on actual video content rather than using a predetermined fixed value, thereby improving encoding quality in fluctuating scenarios.

Inventive Principle:
Principle #35Parameter changes

2Adaptability or versatility

If a fixed GOP length is used in HEVC encoder, then the encoding process is stable, but the adaptability to different video scenarios deteriorates

Engineering Contradiction:
Improveadaptability to video scenariosVSAvoidmodel training and inference complexity
Core Design Contradiction:
Adaptability or versatilityVSDevice complexity

Solution Approach 1:

The patent performs preliminary analysis of video frame characteristics before encoding to determine the appropriate GOP length. By pre-computing motion complexity, texture complexity, and other features, and using these to select GOP length before the actual encoding process, the system achieves adaptability to different video scenarios while maintaining encoding stability through structured preprocessing.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The patent introduces an intermediary analysis layer that sits between the video input and the encoding process. This intermediary computes video frame characteristics and uses them to determine GOP length, acting as a mediator that translates raw video content into encoding parameters. This approach enhances adaptability by analyzing actual video content while managing complexity through a structured intermediate processing stage.

Inventive Principle:
Principle #24Intermediary (Mediator)

Data Source

PatentUS12309379B2Method for processing video frame, method for training model, electronic device, and storage medium
Publication Date: 2025.05.20 BEIJING BAIDU NETCOM SCI & TECH CO LTD
  • US12309379B2 patent drawing
  • US12309379B2 patent drawing
  • US12309379B2 patent drawing

AI summary

Provided are a video frame processing method, a model training method, a device and a storage medium, relates to a field of artificial intelligence, and in particular, to cloud computing, video processing, and medium cloud technology, and may be applied in an intelligent cloud scenario. The video frame processing method includes: acquiring a target characteristic corresponding to a current video frame in a video frame sequence to be encoded, in the case of the video frame sequence to be encoded satisfies a preset condition; inputting the target characteristic corresponding to the current video frame to a first target model, to obtain a first output result corresponding to the current video frame; and determining a first target group of pictures (GOP) length corresponding to the current video frame, based on the first output result corresponding to the current video frame.