Motion-Based ROI Video Processing for Limited-Bandwidth Clarity

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing video transmission systems face challenges in efficiently utilizing transmission resources while ensuring clear visualization of the region of interest (ROI) under limited bandwidth conditions.

Innovation Solution

A video processing method that divides captured video into regions of interest (ROI) and non-regions of interest (non-ROI) based on global motion states between frames, applying different image processing to achieve varying clarity levels.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Reliability

If a uniform encoding strategy is applied to all areas within a single image frame, then the encoding process is simple, but the transmission bandwidth is insufficient to provide clear view of the region of interest under limited transmission resources

Engineering Contradiction:
Improvevisual quality experienceVSAvoidtransmission bandwidth
Core Design Contradiction:
ReliabilityVSQuantity of substance

Solution Approach 1:

The video frame is segmented into multiple regions based on motion analysis: stationary regions, moving regions, and regions of interest. Different encoding strategies are applied to each segment, with higher quality encoding for ROI regions and lower quality for non-ROI regions, thereby optimizing bandwidth utilization while maintaining visual quality experience.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

Different quality levels are assigned to different spatial regions within the video frame. The ROI regions receive high-quality encoding with finer detail preservation, while non-ROI regions use coarser encoding. This local quality differentiation ensures that transmission bandwidth is allocated efficiently according to visual importance.

Inventive Principle:
Principle #3Local quality

2Loss of energy

If different image processing is applied to ROI and non-ROI to achieve varying clarity levels, then transmission resource consumption is reduced, but the device complexity increases

Engineering Contradiction:
Improvetransmission resource consumptionVSAvoidprocessing complexity
Core Design Contradiction:
Loss of energyVSDevice complexity

Solution Approach 1:

Motion analysis and region classification are performed in advance before the encoding process. The video frame is pre-analyzed to identify stationary and moving regions, and ROI areas are determined based on motion characteristics. This preliminary classification enables subsequent encoding stages to apply appropriate quality levels without complex real-time decision-making during encoding.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The system automatically identifies ROI regions and determines appropriate encoding parameters based on motion analysis results. The encoding device uses the motion information to self-determine which regions require high-quality encoding and which can use lower quality, reducing the need for manual configuration or complex external control systems.

Inventive Principle:
Principle #25Self-service

Data Source

PatentEP4580184A1Video processing method and apparatus, device, and computer storage medium
Publication Date: 2025.07.02 SZ DJI TECH CO LTD
  • EP4580184A1 patent drawingFigure 1~2
  • EP4580184A1 patent drawingFigure 3~4
  • EP4580184A1 patent drawingFigure 5~7

AI summary

A video processing method and device, an apparatus and a computer storage medium are provided. The method includes: obtaining a video captured by a photographing device; dividing the video into a plurality of regions based on information associated with a global motion state between frames of the video, wherein the plurality of regions includes a region of interest (ROI) and a non-region of interest (non-ROI); and performing different image processing on the ROI and the non-ROI to achieve different levels of clarity for the ROI and the non-ROI. This application can effectively save transmission resource usage while ensuring the user's subjective visual experience.