3D Skeleton Extraction From LIDAR FMV for Occluded Human Tracking
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing methods for 3-D human identification and tracking are limited by their reliance on warped point cloud matching and structured light sensors, primarily focusing on single human silhouettes and struggling with view invariance and occlusion, while LIDAR mapping has been used only for extracting single human silhouettes and differentiating them from non-human objects.
Innovation Solution
A machine learning system utilizing Lidar full motion video and computational skeleton extraction modules to identify, track, and predict multiple subjects by assigning nodes to joints, tracking their motion, and comparing it to a historical database of characteristic motions, with modules for occlusion completion and view-invariant skeleton extraction.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If warped point cloud matching and structured light sensor estimates are used for 3-D human identification, then single human silhouette extraction and differentiation from non-human objects is achieved, but view invariance and handling of occlusion are limited
Solution Approach 1:
The system segments the 3-D human figure into a skeleton representation with joints and limbs, separating the essential structural information from the surface geometry. This skeleton-based approach enables view-invariant identification since skeletal structure remains consistent across different viewing angles, while the original point cloud data can be processed independently for occlusion handling.
Solution Approach 2:
The patent transitions from 2-D image data to 3-D skeleton data, adding the dimension of depth and spatial structure. The skeleton representation in 3-D space provides view-invariant characteristics while the temporal dimension is added through motion tracking across multiple frames, enabling the system to handle occlusion by predicting hidden body parts based on skeletal motion patterns.
2Measurement precision
If LIDAR mapping is used to extract single human silhouettes, then differentiation from non-human objects is achieved, but multiple subject tracking and motion prediction are not supported
Solution Approach 1:
The system segments the scene into multiple independent skeleton tracks, allowing simultaneous extraction and tracking of multiple human subjects. Each subject is represented by its own skeleton with joints and limbs, enabling independent processing and identification while maintaining the ability to handle occlusion through temporal motion prediction.
Solution Approach 2:
The patent implements continuous motion tracking across multiple video frames, maintaining skeleton representations of subjects over time. This temporal continuity enables the system to track multiple subjects simultaneously and predict their future positions, addressing the limitation of single-silhouette extraction methods.
3Adaptability or versatility
If 3-D skeleton extraction is performed on long-range LIDAR data, then view-invariant human identification is achieved, but noise and low resolution degrade recognition accuracy
Solution Approach 1:
The system performs preliminary skeleton extraction and joint detection on each individual frame before temporal processing. By pre-processing the skeleton representation and identifying characteristic motions in advance, the system can filter out noise and resolve ambiguities using temporal context from previous and future frames, improving overall recognition accuracy.
Solution Approach 2:
The patent implements feedback mechanisms where the skeleton extraction and motion prediction are iteratively refined. The system compares extracted skeletons across frames, uses motion models to predict expected positions, and adjusts the skeleton estimation based on temporal consistency, thereby reducing the impact of noise and low resolution in the original LIDAR data.
Applied Scientific Principles
This section explains which scientific principles are used to turn an abstract innovation direction into a practical engineering solution.
Function Achieved in This Case
The system effectively identifies and tracks multiple subjects in varying orientations and conditions, improving recognition accuracy by handling noise and low resolution in long-range 3-D FMV, enabling view-invariant human identification and prediction of movements.
Implementation Method 1
each frame contains at least image data and depth data, which in some instances can be provided as a LIDAR sensing system
Data Source
AI summary
A machine learning system for the identifying, tracking, and prediction of multiple subjects in a series of images or frames. The system including a front-end sensing assembly, and server configured to process images having depth wherein the server is configured to: identify one or more one or more subjects from each frame utilizing the image data and the depth data; compare recognized movement with a gallery of time-series frames as contained in a historical database; determine one or more uncharacteristic motions by comparing motion of each of the relative nodes to characteristic node motions as stored in the historical database; assign nodes following uncharacteristic motions to a new subject; determine whether by assigning the nodes following uncharacteristic motions to a new subject can resolve any uncharacteristic motions of the associated nodes.


