Video Object Segmentation Using Shape Probability and Graph Cut

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing methods for video object segmentation under weakly-labeled conditions are inaccurate due to the lack of location information and require multiple videos for segmentation, making them unsuitable for single input videos.

Innovation Solution

A method using an object bounding box detector and an object contour detector to estimate the initial object segment sequence, followed by a joint assignment model and graph cut algorithm to optimize the segmentation, ensuring accurate and single-video capable segmentation.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Adaptability or versatility

If weakly-supervised learning method is used for video object segmentation, then the method can handle weakly-labeled videos, but the classification accuracy becomes inaccurate due to lack of location information

Engineering Contradiction:
Improveability to handle weakly-labeled videosVSAvoidclassification accuracy
Core Design Contradiction:
Adaptability or versatilityVSMeasurement precision

Solution Approach 1:

The patent applies preliminary action by first detecting candidate object regions using detectors before performing classification. The method detects candidate objects in each video frame, then uses these detected regions as basis for subsequent classification and segmentation, avoiding the accuracy problems of direct weakly-supervised classification

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The patent introduces an intermediary mechanism by using detected candidate object regions as a bridge between weak labels and final segmentation. The candidate regions serve as intermediate representations that connect the weak semantic labels to the final object segmentation results, improving classification accuracy

Inventive Principle:
Principle #24Intermediary (Mediator)

2Adaptability or versatility

If weakly-supervised learning method is used for video object segmentation, then the method can process videos with semantic tags, but multiple videos are required as input making it unsuitable for single video segmentation

Engineering Contradiction:
Improveability to process videos with semantic tagsVSAvoidapplicability to single input video
Core Design Contradiction:
Adaptability or versatilityVSEase of operation

Solution Approach 1:

The patent applies segmentation by processing each video frame independently to detect candidate objects, then performing temporal segmentation to identify object instances across frames. This frame-level independent processing allows the method to work with single videos while maintaining the ability to handle semantic tags

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent implements self-service by enabling the system to process single videos autonomously without requiring multiple video inputs. The method uses the semantic tag of the single input video to guide the detection and segmentation process, making the system self-sufficient for single-video applications

Inventive Principle:
Principle #25Self-service

3Adaptability or versatility

If weakly-supervised learning method is used for video object segmentation, then the method can segment objects with semantic categories, but the segmentation accuracy becomes inaccurate due to wrong classification

Engineering Contradiction:
Improveability to segment objects with semantic categoriesVSAvoidsegmentation accuracy
Core Design Contradiction:
Adaptability or versatilityVSManufacturing precision

Solution Approach 1:

The patent applies feedback by using detection results to guide classification and segmentation decisions. The detected candidate object regions provide feedback that improves the accuracy of subsequent classification and segmentation steps, creating a closed-loop system that reduces wrong classifications

Inventive Principle:
Principle #23Feedback

Solution Approach 2:

The patent applies preliminary action by performing object detection before classification and segmentation. This preliminary detection step provides accurate location information that improves the precision of subsequent segmentation operations, reducing segmentation errors

Inventive Principle:
Principle #10Preliminary action

Data Source

PatentUS9740956B2Method for object segmentation in videos tagged with semantic labels
Publication Date: 2017.08.22 BEIHANG UNIV
  • US9740956B2 patent drawing
  • US9740956B2 patent drawing
  • US9740956B2 patent drawing

AI summary

The present invention provides a method for object segmentation in videos tagged with semantic labels, including: detecting each frame of a video sequence with an object bounding box detector from a given semantic category and an object contour detector, and obtaining a candidate object bounding box set and a candidate object contour set for each frame of the input video; building a joint assignment model for the candidate object bounding box set and the candidate object contour set and solving the model to obtain the initial object segment sequence; processing the initial object segment, to estimate a probability distribution of the object shapes; and optimizing the initial object segment sequence with a variant of graph cut algorithm that integrates the shape probability distribution, to obtain an optimal segment sequence.