Video Region of Interest Detection via Object Segmentation
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing video processing methods face challenges in accurately determining region of interest (ROI) in video frames, especially when multiple targets are present, leading to unstable performance and limited detection accuracy.
Innovation Solution
A method and apparatus that acquire object regions through object detection, determine non-ROIs based on preset conditions such as area ratios and priorities, and adjust quantization parameters for each ROI to improve encoding quality, allowing for comprehensive ROI determination and resource optimization.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If image distinctiveness detection is used to perform ROI division, then ROI regions can be identified, but the detection accuracy is limited and performance is unstable
Solution Approach 1:
The patent segments the video frame into multiple object regions using object detection technology, where each detected object becomes a separate region. This segmentation approach allows for more precise ROI identification by treating each object independently rather than using subjective image distinctiveness detection, thereby improving both accuracy and stability.
Solution Approach 2:
The patent changes the detection parameters from subjective image distinctiveness metrics to objective object detection parameters such as object boundaries, types, and attributes. By using predefined object detection models with specific detection thresholds and confidence levels, the system achieves more stable and accurate ROI determination.
2Measurement precision
If all object regions are treated as ROIs, then comprehensive coverage is achieved, but computational resources are wasted on encoding non-important regions
Solution Approach 1:
The patent applies different quality levels to different regions by setting specific quantization parameters for each ROI based on its importance. Non-ROI regions use default quantization parameters while ROI regions use adjusted parameters, thereby optimizing computational resources by focusing encoding efforts only on important regions rather than uniformly processing the entire frame.
Solution Approach 2:
The patent performs partial action by selectively applying detailed object detection and ROI processing only to regions that meet specific criteria (such as object type, size, or position), rather than processing the entire video frame uniformly. This reduces computational overhead while maintaining accuracy for critical regions.
3Adaptability or versatility
If multiple types of targets are present in the screen, then comprehensive object detection is achieved, but it becomes difficult to accurately determine which regions belong to ROI
Solution Approach 1:
The patent assigns different weights, priorities, or importance levels to different object types based on their relevance to the specific application scenario. For example, in a video conference scenario, human faces and bodies are assigned higher priority than background objects. This local quality differentiation enables accurate ROI determination even when multiple target types are present.
Solution Approach 2:
The patent changes detection parameters dynamically based on object types, such as adjusting detection thresholds, confidence levels, or region selection criteria according to the specific object category detected. This allows the system to maintain high ROI determination accuracy across diverse multi-target scenarios by adapting parameters to each object type's characteristics.
Data Source
AI summary
Embodiments of the present disclosure relate to a method and apparatus for processing a video. The method may include: acquiring object regions obtained by performing object detection on a target video frame, a type of an object in each of the object regions being a preset type; determining, for an object region in the acquired object regions, in response to determining that the object region satisfies a preset condition, that the object region is a non-ROI; using an object region other than the non-ROI in the object regions of the target video frame as a ROI; and acquiring a quantization parameter change corresponding to each ROI, and encoding the target video frame based on the quantization parameter change.


