Multi-Camera Foreground Extraction Through Depth-Guided Image Synthesis

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing camera technologies struggle to accurately extract foreground obstacles from video images captured by multiple camera sensors in real-time, especially in complex environments with challenging lighting conditions and occlusions, often requiring significant computational resources and leading to inaccuracies and incomplete coverage.

Innovation Solution

A system utilizing multiple camera sensors with overlapping fields of view, employing deep learning techniques and image processing algorithms to estimate spatial depths and render obstructing objects transparent by synthesizing images from secondary cameras, filling in unknown areas with background information.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If background subtraction or segmentation techniques are used to extract foreground obstacles, then foreground detection can be achieved, but inaccuracies occur due to lighting changes, camera motion, and occlusions

Engineering Contradiction:
Improveforeground detection accuracyVSAvoiddetection reliability under varying conditions
Core Design Contradiction:
Measurement precisionVSReliability

Solution Approach 1:

The patent transitions from 2D image processing to 3D spatial reasoning by introducing depth estimation. Multiple cameras capture images from different viewpoints, and depth information is used to distinguish foreground obstacles from background elements. This dimensional addition resolves ambiguities that plague 2D segmentation methods under varying lighting and motion conditions.

Inventive Principle:
Principle #17Another dimension (Dimensionality change)

Solution Approach 2:

The patent introduces depth estimation as an intermediary between image capture and foreground extraction. Rather than directly segmenting foreground from background in 2D space, the system first estimates depth maps from multiple camera views, then uses this depth information as a mediator to accurately identify and extract foreground obstacles, improving reliability under challenging conditions.

Inventive Principle:
Principle #24Intermediary (Mediator)

2Measurement precision

If deep learning techniques are used for foreground extraction, then accuracy improves, but large amounts of training data are required and overfitting occurs

Engineering Contradiction:
Improveforeground detection accuracyVSAvoidtraining data requirements and model generalization
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The patent replaces complex deep learning models with a geometric approach based on multiple camera views and depth estimation. Instead of using data-heavy neural networks to learn foreground patterns, the system uses physics-based triangulation and spatial reasoning to directly compute depth and identify obstacles, achieving high accuracy without extensive training data.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

Solution Approach 2:

The patent changes the fundamental parameters used for foreground detection from pixel intensity values (used in traditional segmentation) to depth values derived from multiple camera viewpoints. This parameter transformation enables accurate foreground extraction using simple geometric relationships rather than complex learned models.

Inventive Principle:
Principle #35Parameter changes

3Device complexity

If a single camera or limited number of cameras are used, then device complexity is reduced, but incomplete or inaccurate foreground extraction occurs

Engineering Contradiction:
Improvenumber of camerasVSAvoidforeground extraction completeness
Core Design Contradiction:
Device complexityVSMeasurement precision

Solution Approach 1:

The patent segments the imaging task across multiple cameras, with each camera capturing a specific viewpoint. By dividing the scene capture function across multiple sensors and combining their depth estimates, the system achieves complete and accurate foreground extraction that would be impossible with a single camera, while keeping each individual camera unit relatively simple.

Inventive Principle:
Principle #1Segmentation

4Area of stationary object

If replacement pixels are used to fill missing coverage areas, then coverage gaps are filled, but disturbing artifacts are introduced

Engineering Contradiction:
Improveimage coverage areaVSAvoidpixel accuracy and artifact freedom
Core Design Contradiction:
Area of stationary objectVSMeasurement precision

Solution Approach 1:

The patent uses depth information as an intermediary to selectively fill only those pixels that are actually occluded by foreground obstacles. By using depth maps to identify true coverage gaps versus areas blocked by foreground objects, the system fills missing background pixels without introducing artifacts, as the depth information provides accurate spatial context for pixel replacement.

Inventive Principle:
Principle #24Intermediary (Mediator)

Data Source

PatentUS12354278B2Method for extracting obstructing objects from an image captured in a multiple camera system
Publication Date: 2025.07.08 MUYBRIDGE AS
  • US12354278B2 patent drawing
  • US12354278B2 patent drawing
  • US12354278B2 patent drawing

AI summary

The present invention provides a system and method for utilizing multiple camera systems for extracting obstructing objects or foreground close to primary camera(s) and filling in the unknown area using pixel information obtained from multiple cameras, i.e. synthesized images where foreground obstructions are removed is provided. This is done by using a primary camera capturing a target scene in the background, obstructed by foreground, and secondary cameras with viewpoints different from the primary camera provide coverage of the scene behind the foreground obstructions. A processing unit will synthesize the image from the primary camera by replacing foreground obstructions with pixels from secondary cameras, or a combination of both to obtain a transparent obstructing object with the background visible through the object. The area to be replaced is identified by calculating depths in the captured area of the primary camera, estimated by comparing images from multiple cameras.