Video Background Object Removal Using Depth Maps

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current video conferencing technologies struggle to effectively remove distracting or private objects from the middle ground of video conference images, especially as wider fields of view become available, leading to increased background noise and potential privacy issues.

Innovation Solution

The use of depth maps created by cameras such as structured light, time-of-flight, or infrared cameras to identify and remove objects from the middle ground, replacing them with background content using machine learning algorithms or depth estimation from a single camera.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Area of stationary object

If the field of view of integrated cameras is increased to capture the entire physical space around the user, then more background information is captured, but more potentially private and distracting objects are included in the video conference

Engineering Contradiction:
Improvefield of view areaVSAvoidprivacy exposure and distractions
Core Design Contradiction:
Area of stationary objectVSObject-affected harmful factors

Solution Approach 1:

The patent segments the video scene into three depth-based layers: foreground (participant), middle-ground (objects to be removed), and background (static environment). This segmentation allows selective processing of the middle-ground objects while preserving the background, resolving the contradiction between capturing wide field of view and removing distracting objects.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent extracts and removes specific middle-ground objects from the video feed using depth map analysis and object detection. By taking out only the unwanted objects rather than the entire background, the system maintains the benefits of wide field of view while eliminating privacy and distraction issues.

Inventive Principle:
Principle #2Taking out (Extraction)

2Object-affected harmful factors

If background blurring or replacement is applied to remove distracting objects, then privacy and distractions are reduced, but the natural background view is lost when the attendee actually wants to show the real background

Engineering Contradiction:
Improveprivacy and distractionsVSAvoidbackground display flexibility
Core Design Contradiction:
Object-affected harmful factorsVSAdaptability or versatility

Solution Approach 1:

The patent applies different processing qualities to different spatial regions: the middle-ground objects are selectively removed or blurred, while the background remains fully visible and natural. This local differentiation allows the system to reduce distractions in specific areas while preserving the natural background view in other areas, maintaining adaptability.

Inventive Principle:
Principle #3Local quality

Solution Approach 2:

The system dynamically adjusts the level of background processing based on detected objects and user preferences. The background can transition between fully visible, partially blurred, or fully replaced states, providing versatility while maintaining natural appearance when appropriate.

Inventive Principle:
Principle #15Dynamics

3Measurement precision

If object removal solutions focused on still images are used, then object removal capability is achieved, but latency makes them unsuitable for video conferencing

Engineering Contradiction:
Improveobject detection accuracyVSAvoidprocessing latency
Core Design Contradiction:
Measurement precisionVSLoss of time

Solution Approach 1:

The patent performs preliminary depth map generation and object detection on incoming video frames before final rendering. By preparing depth information and identifying objects in advance within the video processing pipeline, the system achieves accurate object removal without excessive latency, making it suitable for real-time video conferencing.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The patent replaces traditional still-image processing methods with a video-optimized processing pipeline that uses depth maps from structured light or time-of-flight cameras. This substitution enables real-time processing by leveraging depth information for efficient object segmentation and removal, reducing latency compared to pixel-based still image methods.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

4Measurement precision

If segmentation of objects to be removed is applied, then object removal is achieved, but objects coming into the frame are not identified in time

Engineering Contradiction:
Improveobject identification accuracyVSAvoiddetection delay for entering objects
Core Design Contradiction:
Measurement precisionVSLoss of time

Solution Approach 1:

The system performs preliminary detection and tracking of objects as they enter the frame using depth map analysis. By continuously monitoring the middle-ground region and identifying objects early in their appearance, the system ensures timely detection and removal of entering objects without significant delay.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The patent implements continuous feedback loops that monitor incoming video frames and update object detection in real-time. As objects enter the frame, the depth-based segmentation system provides immediate feedback to identify and mark them for removal, ensuring timely detection and consistent object removal performance.

Inventive Principle:
Principle #23Feedback

Applied Scientific Principles

This section explains which scientific principles are used to turn an abstract innovation direction into a practical engineering solution.

Function Achieved in This Case

This solution allows for real-time removal and replacement of middle ground objects, enhancing video conferencing privacy and reducing distractions, while maintaining a natural and uninterrupted background view.

Implementation Method 1

depth maps created by cameras such as structured light

Methodology Applied
Scientific EffectStructured light:

Implementation Method 2

depth maps created by cameras such as structured light, time-of-flight

Methodology Applied
Scientific EffectTime of flight: Time of Flight

Implementation Method 3

depth maps created by cameras such as structured light, time-of-flight, or infrared cameras

Methodology Applied
Scientific EffectInfrared radiation: Infrared Radiation

Data Source

PatentUS12340489B2Object removal during video conferencing
Publication Date: 2025.06.24 LENOVO (SINGAPORE) PTE LTD
  • US12340489B2 patent drawing
  • US12340489B2 patent drawing
  • US12340489B2 patent drawing

AI summary

A video conferencing system includes an image of a participant in a video conference and a depth map of the image. The system identifies objects in the background of the image, identifies objects in the foreground of the image, and identifies objects in the middle-ground of the image. The system removes the objects from the middle-ground, and replaces the removed objects from the middle-ground with the objects from the background that are located behind the removed objects. The system then uses the image with the removed and replaced objects in a video stream of the video conference.