Mixed-Reality Telepresence Using Depth Sensors to Avoid Segmentation Errors

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current 3D technologies in virtual environments lack realism due to image artifacts and edge effects when capturing human users for telepresence, requiring extensive computing resources and complex motion tracking.

Innovation Solution

A mixed-reality telepresence system that captures real-time 3D data using Kinect sensors or similar technology, displaying users in a virtual environment without converting them to avatars, thus avoiding segmentation errors and allowing for more natural and realistic 3D representations with omni-directional rendering.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Manufacturing precision

If 3D images of human users are extracted from real-world background and inserted into virtual environment, then realism is improved, but image artifacts and edge effects occur that negate the gain in realism

Engineering Contradiction:
Improverealism of 3D representationVSAvoidimage artifacts and edge effects
Core Design Contradiction:
Manufacturing precisionVSObject-affected harmful factors

Solution Approach 1:

The patent extracts only the necessary depth information from the real-world environment using a depth sensor (Kinect), while keeping the video feed of the user intact. This selective extraction of depth data allows the system to create a 3D effect without attempting to extract and reinsert the user into a virtual environment, thereby avoiding the image artifacts and edge effects that would result from such extraction and insertion processes.

Inventive Principle:
Principle #2Taking out (Extraction)

2Manufacturing precision

If sophisticated and high-speed inverse kinematics are used to derive skeleton and physical model of real-time captured human object, then parameterized 3D human realism is improved, but extensive computing resources are required

Engineering Contradiction:
Improveparameterized 3D human realismVSAvoidcomputing resources
Core Design Contradiction:
Manufacturing precisionVSPower

Solution Approach 1:

The patent creates a simplified 3D representation by mapping the video feed of the user onto a 3D model with depth information, rather than using sophisticated inverse kinematics to derive a parameterized 3D human model. This copying approach uses the existing video data and depth sensor information to create a 3D effect without requiring extensive computing resources for complex skeletal and physical model derivation.

Inventive Principle:
Principle #26Copying

3Manufacturing precision

If objects or items attached to the user (such as HMD or chair) are segmented away in real-time capture, then realism of user representation is improved, but segmentation errors occur

Engineering Contradiction:
Improveuser representation accuracyVSAvoidsegmentation accuracy
Core Design Contradiction:
Manufacturing precisionVSMeasurement precision

Solution Approach 1:

The patent uses the depth sensor to segment the user from the background by capturing depth information, which naturally separates the user (closer to the sensor) from the background (farther from the sensor). This depth-based segmentation avoids the need for complex image processing to remove attached objects like HMDs or chairs, as the depth data provides a clean separation based on spatial distance rather than requiring pixel-level segmentation that would introduce errors.

Inventive Principle:
Principle #1Segmentation

Data Source

PatentEP3219098B1System and method for 3D telepresence
Publication Date: 2021.10.06 PCMS HOLDINGS INC
  • EP3219098B1 patent drawingFigure 1
  • EP3219098B1 patent drawingFigure 2A~2C
  • EP3219098B1 patent drawingFigure 3A~3C

AI summary

Systems and methods are described that enable a 3D telepresence. In an exemplary method, a 3D image stream is generated of a first participant in a virtual meeting. A virtual meeting room is generated. The virtual meeting room includes a virtual window, and the 3D image stream is reconstructed in the virtual window. The first participant thus appears as a 3D presence within the virtual window. The virtual meeting room may also include virtual windows providing 3D views of other participants in the virtual meeting and may further include avatars of other meeting participants and/or of the first meeting participant.