Stereo Imaging Alpha Matting via Depth-Aware Neural Networks

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current image segmentation methods, particularly deep learning-based approaches, lack depth information, leading to errors in alpha matting, such as misclassifying body parts or accessories due to similarities in appearance with the background, resulting in inaccurate foreground-background separation.

Innovation Solution

Incorporating depth perception into image segmentation by using stereo imaging with two or more image sensors to capture and rectify images, feeding the rectified images into a deep neural network that includes depth information for more accurate alpha matting mask determination.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Extent of automation

If deep learning networks are used for image segmentation, then automation and processing speed are improved, but accuracy deteriorates due to lack of depth information

Engineering Contradiction:
Improveautomation of image segmentationVSAvoidaccuracy of alpha matting
Core Design Contradiction:
Extent of automationVSMeasurement precision

Solution Approach 1:

The patent transitions from 2D image data to 3D spatial understanding by incorporating depth maps generated from stereo imaging. The depth information adds a new dimension to the input data, enabling the neural network to distinguish foreground from background more accurately based on spatial depth rather than just color and texture similarities.

Inventive Principle:
Principle #17Another dimension (Dimensionality change)

Solution Approach 2:

The patent combines multiple types of data (color images, depth maps, and their gradients) into a composite input tensor for the neural network. This composite approach integrates information from different sources and modalities, similar to how composite materials combine different substances to achieve superior properties.

Inventive Principle:
Principle #40Composite materials

2Measurement precision

If stereo imaging is used to add depth information, then accuracy of alpha matting is improved, but device complexity increases

Engineering Contradiction:
Improveaccuracy of alpha mattingVSAvoidcomplexity of imaging system
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The patent makes the existing neural network architecture multi-functional by enabling it to process both 2D color image data and 3D depth information. The same network structure that processes color images can now also ingest depth maps and their gradients, eliminating the need for separate processing pipelines and reducing overall system complexity despite the added imaging capability.

Inventive Principle:
Principle #6Universality (Multi-functionality)

3Reliability

If depth information is incorporated into the neural network, then foreground-background separation is improved, but computational requirements increase

Engineering Contradiction:
Improvereliability of segmentationVSAvoidcomputational energy consumption
Core Design Contradiction:
ReliabilityVSUse of energy by moving object

Solution Approach 1:

The patent performs preliminary processing of depth information by computing depth gradients and combining them with color gradients before feeding data to the neural network. This pre-computation of derivative information prepares the data in advance, reducing the computational burden during the main segmentation process and enabling more efficient training and inference.

Inventive Principle:
Principle #10Preliminary action

Data Source

PatentUS20240331099A1Background replacement with depth information generated by stereo imaging
Publication Date: 2024.10.03 VISIONARY AI VISION LTD
  • US20240331099A1 patent drawing
  • US20240331099A1 patent drawing
  • US20240331099A1 patent drawing

AI summary

A method of background replacement includes: receiving a main image of a scene from a main image sensor and a secondary image of a scene from a secondary image sensor, wherein the main image sensor and secondary image sensor are displaced relative to each other in at least one dimension; performing stereo rectification on the main image and secondary image; inputting the rectified images into a deep neural network, and applying the deep neural network on the rectified images to generate an alpha matting mask for the main image.