Humanoid Depth Map Segmentation for Markerless Motion Tracking
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current methods for 3D modeling of humanoid forms from depth maps often require dedicated markers or multiple cameras, which can be cumbersome and inefficient, and lack the ability to process data in real-time without artifacts from background objects.
Innovation Solution
A processor-based method that segments and analyzes depth maps to identify humanoids without markers, using a single stationary imaging device to generate a contour, identify torso and limbs, and derive stick-figure representations for real-time motion analysis and control inputs.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If dedicated markers are attached to the subject's body for tracking, then motion tracking precision is improved, but device complexity and ease of operation deteriorate
Solution Approach 1:
The patent extracts and removes the markers from the system, achieving markerless 3D human form modeling. Instead of requiring physical markers attached to the body, the system uses depth map data and probabilistic models to directly infer body shape and motion, thereby eliminating the complexity of marker attachment while maintaining tracking precision
Solution Approach 2:
The patent creates a virtual 3D model (copy) of the human body from depth map images rather than using physical markers. The probabilistic model generates a digital representation of body shape and motion that can be tracked without physical attachments, replacing the need for markers with a computational model
2Measurement precision
If multiple cameras are used to provide 3D stereo image information, then measurement precision is improved, but device complexity and cost increase
Solution Approach 1:
The patent extracts and removes the requirement for multiple cameras from the system. Instead of using stereo camera setups, it employs a single depth map source combined with probabilistic modeling to achieve 3D human form reconstruction, thereby reducing hardware complexity while maintaining measurement precision
Solution Approach 2:
The patent replaces the mechanical/optical system of multiple cameras with a computational approach using probabilistic models. The system substitutes physical stereo vision hardware with algorithmic 3D reconstruction from depth maps, reducing device complexity while preserving measurement capabilities
3Loss of information
If statistical background subtraction is applied to remove background objects, then purity of depth map is improved, but processing time and computational complexity increase
Solution Approach 1:
The patent performs background subtraction and depth map purification as a preliminary step before main processing. By removing background artifacts early in the pipeline using statistical methods, the system prepares clean depth data for subsequent 3D modeling operations, preventing background interference from propagating through later stages
Solution Approach 2:
The patent segments the depth map processing into distinct stages: background subtraction, depth map refinement, and 3D model generation. This segmentation allows optimized processing at each stage, with background removal handled efficiently separately from the main modeling pipeline, reducing overall computational burden
4Productivity
If real-time processing at standard video rates is achieved, then productivity is improved, but computational complexity and processing speed requirements increase
Solution Approach 1:
The patent processes depth maps in discrete frames at standard video rates (e.g., 30 fps), using periodic frame-by-frame analysis rather than continuous processing. This approach maintains real-time performance by synchronizing with the natural frame rate of video input, reducing computational load while preserving real-time capability
Solution Approach 2:
The patent applies probabilistic models and optimization algorithms that may perform more computations than strictly necessary for basic functionality. This excessive action ensures robust real-time performance across varying input conditions, maintaining frame rate consistency even when computational complexity increases
Data Source
AI summary
A computer-implemented method includes receiving a depth map (30) of a scene containing a body of a humanoid subject (28). The depth map includes a matrix of pixels (32), each corresponding to a respective location in the scene and having a respective pixel value indicative of a distance from a reference location to the respective location. The depth map is segmented so as to find a contour (64) of the body. The contour is processed in order to identify a torso (70) and one or more limbs (76, 78, 80, 82) of the subject. An input is generated to control an application program running on a computer by analyzing a disposition of at least one of the identified limbs in the depth map.


