3D Background Removal for Vision System Gesture Tracking

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Accurately distinguishing a user from a complex background in video-game applications using vision systems is challenging due to the complexity of background features, which makes it difficult to determine the user's posture and gestures effectively.

Innovation Solution

The method involves acquiring a time-resolved sequence of depth maps, fitting a geometric model to each depth map, tracking it into subsequent maps, and selecting a background section that lacks coherent motion and is far from the model's coordinates for exclusion, allowing for accurate foreground isolation and input control.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If a vision system processes the entire video frame including complex background, then comprehensive scene information is captured, but the accuracy of user posture and gesture detection deteriorates

Engineering Contradiction:
Improveuser posture and gesture detection accuracyVSAvoidbackground complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The patent segments the video frame into foreground (user) and background regions by fitting a geometric model to depth map data. This segmentation isolates the user from the complex background, allowing accurate posture and gesture detection while reducing the impact of background complexity on measurement precision.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent extracts the geometric model representing the user from the full depth map by identifying pixels that fit the model parameters. This extraction removes the complex background information while retaining essential user posture data, thereby improving detection accuracy without being overwhelmed by background complexity.

Inventive Principle:
Principle #2Taking out (Extraction)

2Loss of information

If a geometric model is fitted to all pixels in the depth map, then complete scene analysis is performed, but computational overhead increases

Engineering Contradiction:
Improvescene information completenessVSAvoidprocessing speed
Core Design Contradiction:
Loss of informationVSProductivity

Solution Approach 1:

The patent applies partial action by fitting the geometric model only to relevant depth map pixels rather than processing all pixels. By using model parameters (position, orientation, dimensions) to identify and process only foreground-related pixels, the system maintains scene information completeness while significantly reducing computational overhead for real-time processing.

Inventive Principle:
Principle #16Partial or excessive action

Solution Approach 2:

The patent performs preliminary action by first fitting the geometric model to obtain position, orientation, and dimension parameters before conducting detailed scene analysis. This preliminary modeling step organizes the data structure and identifies relevant regions, enabling more efficient subsequent processing and reducing overall computational requirements.

Inventive Principle:
Principle #10Preliminary action

3Measurement precision

If background sections are included in model fitting, then comprehensive scene understanding is achieved, but foreground isolation accuracy deteriorates

Engineering Contradiction:
Improveforeground isolation accuracyVSAvoidscene understanding completeness
Core Design Contradiction:
Measurement precisionVSAdaptability or versatility

Solution Approach 1:

The patent applies local quality by treating foreground and background regions differently in the model fitting process. The geometric model parameters (position, orientation, dimensions) are used to apply different processing qualities to different spatial regions - precise model fitting for foreground pixels and exclusion for background pixels - thereby achieving high foreground isolation accuracy while maintaining appropriate scene understanding.

Inventive Principle:
Principle #3Local quality

Solution Approach 2:

The patent introduces depth information as an additional dimension by using depth map data alongside traditional image processing. This dimensional addition enables the geometric model to distinguish foreground from background based on depth characteristics, improving foreground isolation accuracy while preserving comprehensive three-dimensional scene understanding.

Inventive Principle:
Principle #17Another dimension (Dimensionality change)

Data Source

PatentUS8526734B2Three-dimensional background removal for vision system
Publication Date: 2013.09.03 MICROSOFT TECHNOLOGY LICENSING LLC
  • US8526734B2 patent drawing
  • US8526734B2 patent drawing
  • US8526734B2 patent drawing

AI summary

A method for controlling a computer system includes acquiring video of a subject, and obtaining from the video a time-resolved sequence of depth maps. A geometric model of the subject is fit to each depth map in the sequence and tracked into a subsequent depth map in the sequence. From the subsequent depth map, a background section is selected for exclusion. The background section is one that lacks coherent motion and is located more than a threshold distance from the coordinates of the geometric model tracked in. Then, a subsequent geometric model of the subject is fit to the depth map with the background section excluded.