Video Capture Metadata Segmentation for Bandwidth Reduction

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

The high computational complexity and bandwidth requirements of semantical analysis in video processing lead to energy consumption and bandwidth issues, particularly in transmitting and processing video data between capture apparatus and target devices.

Innovation Solution

A method where the capture apparatus performs object segmentation and generates metadata representing semantical annotation maps, transmitting only this metadata instead of the full video, reducing bandwidth and energy consumption by eliminating the need for video encoding and transmission.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If video data is transmitted from capture apparatus to target device for semantical analysis, then object segmentation can be performed, but bandwidth consumption increases significantly

Engineering Contradiction:
Improveobject segmentation accuracyVSAvoidbandwidth consumption
Core Design Contradiction:
Measurement precisionVSQuantity of substance

Solution Approach 1:

The patent extracts only the essential segmentation results (masks and metadata) from the video processing pipeline, transmitting only these extracted elements rather than the full video data. This allows the target device to perform segmentation analysis while minimizing transmitted data volume, directly resolving the bandwidth consumption issue.

Inventive Principle:
Principle #2Taking out (Extraction)

Solution Approach 2:

The patent segments the video data transmission into two parts: essential segmentation results (masks) and optional original video data. By transmitting only the segmented masks and associated metadata, the system achieves object segmentation functionality with dramatically reduced bandwidth requirements compared to transmitting complete video streams.

Inventive Principle:
Principle #1Segmentation

2Measurement precision

If video encoding and transmission is performed, then segmented objects can be obtained at target device, but energy consumption increases

Engineering Contradiction:
Improvesegmented object extractionVSAvoidenergy consumption
Core Design Contradiction:
Measurement precisionVSUse of energy by moving object

Solution Approach 1:

The patent performs object segmentation and generates masks at the capture apparatus before transmission. By completing the computationally intensive segmentation process in advance and transmitting only the resulting masks, the system eliminates the need for video encoding and re-processing at the target device, significantly reducing energy consumption.

Inventive Principle:
Principle #10Preliminary action

3Productivity

If full video data is transmitted, then complete processing can be performed, but bandwidth cost increases

Engineering Contradiction:
Improveprocessing completenessVSAvoidbandwidth cost
Core Design Contradiction:
ProductivityVSQuantity of substance

Solution Approach 1:

The patent creates a simplified copy of the video data in the form of segmentation masks that retain the essential information needed for object identification and analysis. These mask copies enable complete processing of segmented objects without requiring transmission of the original full-resolution video data, thereby reducing bandwidth costs while maintaining processing effectiveness.

Inventive Principle:
Principle #26Copying

Data Source

PatentEP4447003A1Method of determining segmentation mask, device, system data, data structure and non-transitory storage medium
Publication Date: 2024.10.16 BEIJING XIAOMI MOBILE SOFTWARE CO LTD
  • EP4447003A1 patent drawingFigure 1
  • EP4447003A1 patent drawingFigure 2
  • EP4447003A1 patent drawingFigure 3a~3b

AI summary

The present application relates to a method of obtaining, by a capture apparatus, at least one segmented object extracted from a video . The method comprises the following steps executed by the capture apparatus: obtaining (210) said at least one segmented object from the video or at least a part of the video by applying an object segmentation process on the video; generating (220) first metadata representative of a semantical annotation map comprising said at least one mask, each mask being representative of at least one segmented object; writing (230) the first metadata into a container; transmitting (240) the container to the target device over a communication network.