Video Background Object Removal Using Depth Maps
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current video conferencing technologies struggle to effectively remove distracting or private objects from the middle ground of video conference images, especially as wider fields of view become available, leading to increased background noise and potential privacy issues.
Innovation Solution
The use of depth maps created by cameras such as structured light, time-of-flight, or infrared cameras to identify and remove objects from the middle ground, replacing them with background content using machine learning algorithms or depth estimation from a single camera.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Area of stationary object
If the field of view of integrated cameras is increased to capture the entire physical space around the user, then more background information is captured, but more potentially private and distracting objects are included in the video conference
Solution Approach 1:
The patent segments the video scene into three depth-based layers: foreground (participant), middle-ground (objects to be removed), and background (static environment). This segmentation allows selective processing of the middle-ground objects while preserving the background, resolving the contradiction between capturing wide field of view and removing distracting objects.
Solution Approach 2:
The patent extracts and removes specific middle-ground objects from the video feed using depth map analysis and object detection. By taking out only the unwanted objects rather than the entire background, the system maintains the benefits of wide field of view while eliminating privacy and distraction issues.
2Object-affected harmful factors
If background blurring or replacement is applied to remove distracting objects, then privacy and distractions are reduced, but the natural background view is lost when the attendee actually wants to show the real background
Solution Approach 1:
The patent applies different processing qualities to different spatial regions: the middle-ground objects are selectively removed or blurred, while the background remains fully visible and natural. This local differentiation allows the system to reduce distractions in specific areas while preserving the natural background view in other areas, maintaining adaptability.
Solution Approach 2:
The system dynamically adjusts the level of background processing based on detected objects and user preferences. The background can transition between fully visible, partially blurred, or fully replaced states, providing versatility while maintaining natural appearance when appropriate.
3Measurement precision
If object removal solutions focused on still images are used, then object removal capability is achieved, but latency makes them unsuitable for video conferencing
Solution Approach 1:
The patent performs preliminary depth map generation and object detection on incoming video frames before final rendering. By preparing depth information and identifying objects in advance within the video processing pipeline, the system achieves accurate object removal without excessive latency, making it suitable for real-time video conferencing.
Solution Approach 2:
The patent replaces traditional still-image processing methods with a video-optimized processing pipeline that uses depth maps from structured light or time-of-flight cameras. This substitution enables real-time processing by leveraging depth information for efficient object segmentation and removal, reducing latency compared to pixel-based still image methods.
4Measurement precision
If segmentation of objects to be removed is applied, then object removal is achieved, but objects coming into the frame are not identified in time
Solution Approach 1:
The system performs preliminary detection and tracking of objects as they enter the frame using depth map analysis. By continuously monitoring the middle-ground region and identifying objects early in their appearance, the system ensures timely detection and removal of entering objects without significant delay.
Solution Approach 2:
The patent implements continuous feedback loops that monitor incoming video frames and update object detection in real-time. As objects enter the frame, the depth-based segmentation system provides immediate feedback to identify and mark them for removal, ensuring timely detection and consistent object removal performance.
Applied Scientific Principles
This section explains which scientific principles are used to turn an abstract innovation direction into a practical engineering solution.
Function Achieved in This Case
This solution allows for real-time removal and replacement of middle ground objects, enhancing video conferencing privacy and reducing distractions, while maintaining a natural and uninterrupted background view.
Implementation Method 1
depth maps created by cameras such as structured light
Implementation Method 2
depth maps created by cameras such as structured light, time-of-flight
Implementation Method 3
depth maps created by cameras such as structured light, time-of-flight, or infrared cameras
Data Source
AI summary
A video conferencing system includes an image of a participant in a video conference and a depth map of the image. The system identifies objects in the background of the image, identifies objects in the foreground of the image, and identifies objects in the middle-ground of the image. The system removes the objects from the middle-ground, and replaces the removed objects from the middle-ground with the objects from the background that are located behind the removed objects. The system then uses the image with the removed and replaced objects in a video stream of the video conference.


