Adaptive Video Encoding Using Machine Learning for Object Detection

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

In video surveillance applications, determining optimal coding parameters for video encoders to effectively detect and differentiate between moving objects and background in video images is challenging, as existing methods often rely on preset parameters or limited information, leading to suboptimal performance in detecting interesting objects like persons or vehicles.

Innovation Solution

An apparatus and method using a machine-learning-based model to determine coding parameters by inputting motion information and video image samples, which classifies regions of interest and non-interest, allowing for adaptive encoding that enhances object detection and background separation.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Productivity

If preset coding parameters are used for video encoding, then the encoding process is simple and fast, but the object detection accuracy and background separation performance deteriorate

Engineering Contradiction:
Improveencoding speedVSAvoidobject detection accuracy
Core Design Contradiction:
ProductivityVSMeasurement precision

Solution Approach 1:

The patent applies dynamics by transitioning from static preset coding parameters to dynamic adaptive coding parameters. The system continuously analyzes motion information and texture features of video frames, then adjusts coding parameters in real-time based on detected objects and regions of interest. This allows the encoding process to adapt to changing scene conditions, improving both detection accuracy and encoding efficiency across different video content.

Inventive Principle:
Principle #15Dynamics

Solution Approach 2:

The patent implements parameter changes by modifying coding parameters such as quantization parameter (QP), block size, and transformation type based on analyzed video content. The system changes these parameters dynamically according to motion magnitude, object type, and region importance, thereby optimizing both compression efficiency and surveillance performance for different scene conditions.

Inventive Principle:
Principle #35Parameter changes

2Measurement precision

If coding parameters are adapted dynamically based on motion and texture information, then object detection accuracy improves, but the device complexity and processing requirements increase

Engineering Contradiction:
Improveobject detection accuracyVSAvoidprocessing complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The patent applies segmentation by dividing the video frame into multiple regions and processing them differently. The system segments the image into regions of interest (containing detected objects) and background regions, then applies different coding parameters to each segment. This allows complex adaptive processing only where necessary (in regions with objects), while using simpler preset parameters for background areas, thereby reducing overall processing complexity.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent implements local quality by applying different coding parameter settings to different spatial regions of the video frame. High-quality encoding with fine detail preservation is applied to regions containing detected objects, while lower-quality encoding is used for uniform background regions. This localized approach maintains detection accuracy for important regions while reducing overall processing complexity.

Inventive Principle:
Principle #3Local quality

3Ease of manufacture

If preset coding parameters are used, then the encoding process is simple, but the adaptability to different surveillance scenarios deteriorates

Engineering Contradiction:
Improveencoding simplicityVSAvoidsurveillance scenario adaptability
Core Design Contradiction:
Ease of manufactureVSAdaptability or versatility

Solution Approach 1:

The patent applies dynamics by enabling the encoding system to adapt to different surveillance scenarios automatically. The system dynamically analyzes video content to detect objects, classify them by type (person, vehicle, animal), and adjust coding parameters accordingly. This dynamic adaptation allows the same encoding system to effectively handle diverse surveillance scenarios without manual reconfiguration.

Inventive Principle:
Principle #15Dynamics

Solution Approach 2:

The patent implements self-service by enabling the encoding system to automatically analyze video content, detect objects, and adjust coding parameters without external intervention. The system serves itself by using its own output (motion information, texture analysis) to control its input (coding parameters), creating a self-adaptive encoding process that automatically optimizes for different surveillance scenarios.

Inventive Principle:
Principle #25Self-service

Data Source

PatentEP3808086B1Machine-learning-based adaptation of coding parameters for video encoding using motion and object detection
Publication Date: 2024.10.09 HUAWEI TECH CO LTD
  • EP3808086B1 patent drawingFigure 1
  • EP3808086B1 patent drawingFigure 2
  • EP3808086B1 patent drawingFigure 3

AI summary

The present disclosure relates to encoding of a video image using coding parameters, adapted on basis of motion of the video image and of an output of a machine-learning based model, which is fed with samples of a block of the video image and motion information of the samples. With this input along with texture, the machine-learning model segments the video image into regions based on the strength of motion determined from the motion information. An object is detected within the video based on motion and texture, and the spatial-time coding parameters are determined based on strength of the motion, and whether or not the detected objects moves. The use of the machine-learning model, fed with motion information and block samples, combined with texture information of the object allows for a more accurate image segmentation, and thus optimization of coding parameters depending on the importance of the image content in terms of less relevant background and dynamic image content, including fast and slow moving objects of different sizes.