Video Object Segmentation Using Spatio-Temporal Refinement

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing video object segmentation methods face challenges in accurately separating locally moving objects from globally moving backgrounds, often introducing artifacts and losing background details, especially when dealing with complex scenarios like video surveillance and object-based video coding.

Innovation Solution

A method and apparatus for video object segmentation that employs four stages: frame alignment, pixel alignment, consensus filtering, and spatio-temporal refinement, using temporal and spatial contexts to generate accurate binary foreground masks, which are then incorporated into a sampling-based super-resolution framework to enhance compression efficiency and preserve background details.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If traditional background subtraction methods are used to segment foreground objects, then segmentation can be achieved, but artifacts are introduced and background details are lost

Engineering Contradiction:
Improvesegmentation accuracyVSAvoidartifacts and background detail loss
Core Design Contradiction:
Measurement precisionVSObject-generated harmful factors

Solution Approach 1:

The patent divides the segmentation process into four distinct stages: frame alignment, pixel alignment, consensus filtering, and spatio-temporal refinement. Each stage processes the foreground mask separately with specific operations tailored to that stage's requirements, preventing artifact propagation and preserving background details through staged refinement

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent introduces an intermediate foreground mask as a mediator between initial segmentation and final output. This intermediate mask undergoes consensus filtering that combines multiple candidate masks, acting as a buffer that reduces artifacts while preserving essential foreground information before final spatio-temporal refinement

Inventive Principle:
Principle #24Intermediary (Mediator)

2Productivity

If compressed domain techniques are used for object segmentation, then processing efficiency is improved, but block resolution artifacts occur and integration with spatial domain equipment is difficult

Engineering Contradiction:
Improveprocessing efficiencyVSAvoidsegmentation precision and compatibility
Core Design Contradiction:
ProductivityVSManufacturing precision

Solution Approach 1:

The patent replaces compressed domain block-based processing with spatial domain pixel-level processing. By operating on individual pixels rather than compressed blocks, the method achieves higher precision segmentation and seamless integration with spatial domain imaging equipment while maintaining processing efficiency through the staged approach

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

3Ease of manufacture

If running average or Gaussian distribution methods are used to model background, then background estimation is achieved, but reliability decreases with globally moving backgrounds

Engineering Contradiction:
Improvebackground modeling simplicityVSAvoidbackground model reliability
Core Design Contradiction:
Ease of manufactureVSReliability

Solution Approach 1:

The patent implements dynamic frame alignment that adapts to globally moving backgrounds by computing alignment transformations for each frame based on its specific characteristics. This dynamic adaptation allows the background model to remain reliable even when the background itself is moving, overcoming the limitation of static background assumptions in traditional methods

Inventive Principle:
Principle #15Dynamics

Data Source

PatentUS8781253B2Method and apparatus for video object segmentation
Publication Date: 2014.07.15 INTERDIGITAL MADISON PATENT HLDG
  • US8781253B2 patent drawing
  • US8781253B2 patent drawing
  • US8781253B2 patent drawing

AI summary

Methods and apparatus for video object segmentation are provided, suitable for use in a super-resolution system. The method comprises alignment of frames of a video sequence, pixel alignment to generate initial foreground masks using a similarity metric, consensus filtering to generate an intermediate foreground mask, and refinement of the mask using spatio-temporal information from the video sequence. In various embodiments, the similarity metric is computed using a sum of squared differences approach, a correlation, or a modified normalized correlation metric. Soft thresholding of the similarity metric is also used in one embodiment of the present principles. Weighting factors are also applied to certain critical frames in the consensus filtering stage in one embodiment using the present principles.