3D Skeleton Extraction From LIDAR FMV for Occluded Human Tracking

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing methods for 3-D human identification and tracking are limited by their reliance on warped point cloud matching and structured light sensors, primarily focusing on single human silhouettes and struggling with view invariance and occlusion, while LIDAR mapping has been used only for extracting single human silhouettes and differentiating them from non-human objects.

Innovation Solution

A machine learning system utilizing Lidar full motion video and computational skeleton extraction modules to identify, track, and predict multiple subjects by assigning nodes to joints, tracking their motion, and comparing it to a historical database of characteristic motions, with modules for occlusion completion and view-invariant skeleton extraction.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If warped point cloud matching and structured light sensor estimates are used for 3-D human identification, then single human silhouette extraction and differentiation from non-human objects is achieved, but view invariance and handling of occlusion are limited

Engineering Contradiction:
Improvehuman identification accuracyVSAvoidview invariance and occlusion handling
Core Design Contradiction:
Measurement precisionVSAdaptability or versatility

Solution Approach 1:

The system segments the 3-D human figure into a skeleton representation with joints and limbs, separating the essential structural information from the surface geometry. This skeleton-based approach enables view-invariant identification since skeletal structure remains consistent across different viewing angles, while the original point cloud data can be processed independently for occlusion handling.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent transitions from 2-D image data to 3-D skeleton data, adding the dimension of depth and spatial structure. The skeleton representation in 3-D space provides view-invariant characteristics while the temporal dimension is added through motion tracking across multiple frames, enabling the system to handle occlusion by predicting hidden body parts based on skeletal motion patterns.

Inventive Principle:
Principle #17Another dimension (Dimensionality change)

2Measurement precision

If LIDAR mapping is used to extract single human silhouettes, then differentiation from non-human objects is achieved, but multiple subject tracking and motion prediction are not supported

Engineering Contradiction:
Improvesilhouette extraction accuracyVSAvoidmultiple subject processing capability
Core Design Contradiction:
Measurement precisionVSProductivity

Solution Approach 1:

The system segments the scene into multiple independent skeleton tracks, allowing simultaneous extraction and tracking of multiple human subjects. Each subject is represented by its own skeleton with joints and limbs, enabling independent processing and identification while maintaining the ability to handle occlusion through temporal motion prediction.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent implements continuous motion tracking across multiple video frames, maintaining skeleton representations of subjects over time. This temporal continuity enables the system to track multiple subjects simultaneously and predict their future positions, addressing the limitation of single-silhouette extraction methods.

Inventive Principle:
Principle #20Continuity of useful action

3Adaptability or versatility

If 3-D skeleton extraction is performed on long-range LIDAR data, then view-invariant human identification is achieved, but noise and low resolution degrade recognition accuracy

Engineering Contradiction:
Improveview-invariant identificationVSAvoidrecognition accuracy
Core Design Contradiction:
Adaptability or versatilityVSMeasurement precision

Solution Approach 1:

The system performs preliminary skeleton extraction and joint detection on each individual frame before temporal processing. By pre-processing the skeleton representation and identifying characteristic motions in advance, the system can filter out noise and resolve ambiguities using temporal context from previous and future frames, improving overall recognition accuracy.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The patent implements feedback mechanisms where the skeleton extraction and motion prediction are iteratively refined. The system compares extracted skeletons across frames, uses motion models to predict expected positions, and adjusts the skeleton estimation based on temporal consistency, thereby reducing the impact of noise and low resolution in the original LIDAR data.

Inventive Principle:
Principle #23Feedback

Applied Scientific Principles

This section explains which scientific principles are used to turn an abstract innovation direction into a practical engineering solution.

Function Achieved in This Case

The system effectively identifies and tracks multiple subjects in varying orientations and conditions, improving recognition accuracy by handling noise and low resolution in long-range 3-D FMV, enabling view-invariant human identification and prediction of movements.

Implementation Method 1

each frame contains at least image data and depth data, which in some instances can be provided as a LIDAR sensing system

Methodology Applied
Scientific EffectLIDAR: LIDAR

Data Source

PatentUS12573061B2System for view invariant 3-D skeleton estimation and human identification using LIDAR full motion video
Publication Date: 2026.03.10 OLD DOMINION UNIVERSITY
  • US12573061B2 patent drawing
  • US12573061B2 patent drawing
  • US12573061B2 patent drawing

AI summary

A machine learning system for the identifying, tracking, and prediction of multiple subjects in a series of images or frames. The system including a front-end sensing assembly, and server configured to process images having depth wherein the server is configured to: identify one or more one or more subjects from each frame utilizing the image data and the depth data; compare recognized movement with a gallery of time-series frames as contained in a historical database; determine one or more uncharacteristic motions by comparing motion of each of the relative nodes to characteristic node motions as stored in the historical database; assign nodes following uncharacteristic motions to a new subject; determine whether by assigning the nodes following uncharacteristic motions to a new subject can resolve any uncharacteristic motions of the associated nodes.