3D Background Removal for Vision System Gesture Tracking
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Accurately distinguishing a user from a complex background in video-game applications using vision systems is challenging due to the complexity of background features, which makes it difficult to determine the user's posture and gestures effectively.
Innovation Solution
The method involves acquiring a time-resolved sequence of depth maps, fitting a geometric model to each depth map, tracking it into subsequent maps, and selecting a background section that lacks coherent motion and is far from the model's coordinates for exclusion, allowing for accurate foreground isolation and input control.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If a vision system processes the entire video frame including complex background, then comprehensive scene information is captured, but the accuracy of user posture and gesture detection deteriorates
Solution Approach 1:
The patent segments the video frame into foreground (user) and background regions by fitting a geometric model to depth map data. This segmentation isolates the user from the complex background, allowing accurate posture and gesture detection while reducing the impact of background complexity on measurement precision.
Solution Approach 2:
The patent extracts the geometric model representing the user from the full depth map by identifying pixels that fit the model parameters. This extraction removes the complex background information while retaining essential user posture data, thereby improving detection accuracy without being overwhelmed by background complexity.
2Loss of information
If a geometric model is fitted to all pixels in the depth map, then complete scene analysis is performed, but computational overhead increases
Solution Approach 1:
The patent applies partial action by fitting the geometric model only to relevant depth map pixels rather than processing all pixels. By using model parameters (position, orientation, dimensions) to identify and process only foreground-related pixels, the system maintains scene information completeness while significantly reducing computational overhead for real-time processing.
Solution Approach 2:
The patent performs preliminary action by first fitting the geometric model to obtain position, orientation, and dimension parameters before conducting detailed scene analysis. This preliminary modeling step organizes the data structure and identifies relevant regions, enabling more efficient subsequent processing and reducing overall computational requirements.
3Measurement precision
If background sections are included in model fitting, then comprehensive scene understanding is achieved, but foreground isolation accuracy deteriorates
Solution Approach 1:
The patent applies local quality by treating foreground and background regions differently in the model fitting process. The geometric model parameters (position, orientation, dimensions) are used to apply different processing qualities to different spatial regions - precise model fitting for foreground pixels and exclusion for background pixels - thereby achieving high foreground isolation accuracy while maintaining appropriate scene understanding.
Solution Approach 2:
The patent introduces depth information as an additional dimension by using depth map data alongside traditional image processing. This dimensional addition enables the geometric model to distinguish foreground from background based on depth characteristics, improving foreground isolation accuracy while preserving comprehensive three-dimensional scene understanding.
Data Source
AI summary
A method for controlling a computer system includes acquiring video of a subject, and obtaining from the video a time-resolved sequence of depth maps. A geometric model of the subject is fit to each depth map in the sequence and tracked into a subsequent depth map in the sequence. From the subsequent depth map, a background section is selected for exclusion. The background section is one that lacks coherent motion and is located more than a threshold distance from the coordinates of the geometric model tracked in. Then, a subsequent geometric model of the subject is fit to the depth map with the background section excluded.


