Video Object-Related Effect Extraction via Self-Supervised Matte Generation

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current computer vision systems fail to effectively identify and associate scene effects such as shadows, reflections, and smoke with the objects that produce them in videos, limiting their ability to improve visual scene understanding and applications like object removal and enhancement.

Innovation Solution

A computer-implemented method using a machine-learned matte generation model, trained on video data, generates binary object masks and optical flows to separate objects from their effects, allowing for the extraction of object-related effects like shadows, reflections, and smoke, and composites them back into the video frames.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If traditional computer vision segmentation is used to identify objects in videos, then object detection capability is achieved, but scene effects related to objects (shadows, reflections, smoke) are overlooked and not associated with the objects

Engineering Contradiction:
Improveobject detection accuracyVSAvoidscene effects information
Core Design Contradiction:
Measurement precisionVSLoss of information

Solution Approach 1:

The patent applies segmentation by dividing the video scene into distinct layers: object layers and background layer. The machine-learned matte generation model separates objects from their effects by generating alpha mattes that define object boundaries, enabling independent processing of objects and their associated scene effects like shadows and reflections.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent introduces an intermediary mechanism (the machine-learned matte generation model) that acts as a mediator between object detection and scene effect identification. This model takes object masks as input and generates corresponding scene effects, bridging the gap between traditional object segmentation and effect association.

Inventive Principle:
Principle #24Intermediary (Mediator)

2Measurement precision

If manual methods are used to associate objects with their effects, then accurate association can be achieved, but the process is time-consuming and computationally expensive

Engineering Contradiction:
Improveobject-effect association accuracyVSAvoidprocessing speed
Core Design Contradiction:
Measurement precisionVSProductivity

Solution Approach 1:

The system implements self-service by enabling automatic object-effect association through the machine-learned matte generation model. The model autonomously processes video frames, generates object masks, and associates scene effects with objects without requiring manual intervention, thereby achieving both accuracy and efficiency.

Inventive Principle:
Principle #25Self-service

Solution Approach 2:

The patent replaces manual mechanical processes with an automated machine-learned model. The matte generation model substitutes human operators by automatically analyzing video data, generating alpha mattes, and associating effects with objects, significantly improving processing speed while maintaining accuracy.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

3Productivity

If prior automatic methods are used to separate objects from effects, then some level of automation is achieved, but the quality of separation and association is insufficient

Engineering Contradiction:
Improveautomation levelVSAvoidseparation quality
Core Design Contradiction:
ProductivityVSManufacturing precision

Solution Approach 1:

The patent applies parameter changes by utilizing the machine-learned model's ability to dynamically adjust alpha matte parameters to optimize the separation between objects and scene effects. The model learns optimal parameters during training to accurately delineate object boundaries and associate effects, improving separation quality over fixed-parameter methods.

Inventive Principle:
Principle #35Parameter changes

Solution Approach 2:

The patent employs composite materials by combining multiple processing components: object detection algorithms, machine-learned matte generation, and effect association mechanisms. This composite approach integrates different technical elements to achieve superior separation quality and effect association compared to single-method approaches.

Inventive Principle:
Principle #40Composite materials

Data Source

PatentUS20240249523A1Systems and Methods for Identifying and Extracting Object-Related Effects in Videos
Publication Date: 2024.07.25 GOOGLE LLC
  • US20240249523A1 patent drawing
  • US20240249523A1 patent drawing
  • US20240249523A1 patent drawing

AI summary

The present disclosure provides systems and methods for identifying and extracting object-related effects in videos. Given an ordinary video and a rough segmentation mask overtime of one or more subjects of interest, example systems proposed herein can estimate an omnimatte for each subject—an alpha matte and color image that includes the subject along with all its related time-varying scene elements. Example implementations of the proposed models can be trained only on the input video in a self-supervised manner, without any manual labels, and are generic. For example, the models can produce omnimattes automatically for arbitrary objects and a variety of effects.