Video Object Segmentation Using Spatio-Temporal Refinement
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing video object segmentation methods face challenges in accurately separating locally moving objects from globally moving backgrounds, often introducing artifacts and losing background details, especially when dealing with complex scenarios like video surveillance and object-based video coding.
Innovation Solution
A method and apparatus for video object segmentation that employs four stages: frame alignment, pixel alignment, consensus filtering, and spatio-temporal refinement, using temporal and spatial contexts to generate accurate binary foreground masks, which are then incorporated into a sampling-based super-resolution framework to enhance compression efficiency and preserve background details.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If traditional background subtraction methods are used to segment foreground objects, then segmentation can be achieved, but artifacts are introduced and background details are lost
Solution Approach 1:
The patent divides the segmentation process into four distinct stages: frame alignment, pixel alignment, consensus filtering, and spatio-temporal refinement. Each stage processes the foreground mask separately with specific operations tailored to that stage's requirements, preventing artifact propagation and preserving background details through staged refinement
Solution Approach 2:
The patent introduces an intermediate foreground mask as a mediator between initial segmentation and final output. This intermediate mask undergoes consensus filtering that combines multiple candidate masks, acting as a buffer that reduces artifacts while preserving essential foreground information before final spatio-temporal refinement
2Productivity
If compressed domain techniques are used for object segmentation, then processing efficiency is improved, but block resolution artifacts occur and integration with spatial domain equipment is difficult
Solution Approach 1:
The patent replaces compressed domain block-based processing with spatial domain pixel-level processing. By operating on individual pixels rather than compressed blocks, the method achieves higher precision segmentation and seamless integration with spatial domain imaging equipment while maintaining processing efficiency through the staged approach
3Ease of manufacture
If running average or Gaussian distribution methods are used to model background, then background estimation is achieved, but reliability decreases with globally moving backgrounds
Solution Approach 1:
The patent implements dynamic frame alignment that adapts to globally moving backgrounds by computing alignment transformations for each frame based on its specific characteristics. This dynamic adaptation allows the background model to remain reliable even when the background itself is moving, overcoming the limitation of static background assumptions in traditional methods
Data Source
AI summary
Methods and apparatus for video object segmentation are provided, suitable for use in a super-resolution system. The method comprises alignment of frames of a video sequence, pixel alignment to generate initial foreground masks using a similarity metric, consensus filtering to generate an intermediate foreground mask, and refinement of the mask using spatio-temporal information from the video sequence. In various embodiments, the similarity metric is computed using a sum of squared differences approach, a correlation, or a modified normalized correlation metric. Soft thresholding of the similarity metric is also used in one embodiment of the present principles. Weighting factors are also applied to certain critical frames in the consensus filtering stage in one embodiment using the present principles.


