LiDARCap Markerless 3D Motion Capture with ML
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing technologies face challenges in capturing long-range three-dimensional human motion accurately, especially with sparse and noisy point clouds from LiDAR sensors, and lack a large-scale dataset with accurate annotations.
Innovation Solution
A marker-less, long-range, and data-driven motion capture method using a single LiDAR sensor, known as LiDARCap, which trains machine learning models on a large dataset like LiDARHuman26M, incorporating synchronous LiDAR point clouds and IMU-captured ground-truth motions to generate accurate three-dimensional human motions.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Length of stationary object
If LiDAR sensors are used for long-range motion capture, then capturing distance is improved, but point cloud density and accuracy deteriorate
Solution Approach 1:
The patent introduces an intermediary processing pipeline that includes point cloud filtering, feature extraction, and machine learning-based pose estimation. The system uses intermediate representations (keypoints, skeletal structure) to bridge the gap between sparse LiDAR point clouds and accurate 3D motion reconstruction, enabling long-range capture while maintaining precision through multi-stage processing
2Measurement precision
If marker-based motion capture is used, then measurement precision is improved, but device complexity and ease of operation deteriorate
Solution Approach 1:
The patent extracts and eliminates the complex marker-based tracking system, replacing it with a markerless LiDAR-based approach. By removing the need for physical markers and complex calibration procedures, the system achieves simpler operation while maintaining accuracy through machine learning models trained on synthetic data with ground truth annotations
3Measurement precision
If large-scale dataset with accurate annotations is created, then training accuracy is improved, but data collection complexity and time deteriorate
Solution Approach 1:
The patent performs preliminary actions by pre-training machine learning models on large-scale synthetic datasets with perfectly annotated ground truth before actual data collection. This pre-training enables the system to handle real-world LiDAR data more effectively, reducing the time needed for annotation during actual motion capture sessions while maintaining high training accuracy
Applied Scientific Principles
This section explains which scientific principles are used to turn an abstract innovation direction into a practical engineering solution.
Function Achieved in This Case
LiDARCap effectively captures highly accurate three-dimensional human motions in long-range scenarios, outperforming state-of-the-art image-based methods, and provides a robust solution for applications like VR/AR and action quality assessment.
Implementation Method 1
a first long-range LiDAR-based motion capture dataset, the LiDARHuman26M dataset, is provided herein
Data Source
AI summary
Described herein are systems and methods for training machine learning models to generate three-dimensional (3D) motions based on light detection and ranging (LiDAR) point clouds. In various embodiments, a computing system can encode a machine learning model representing an object in a scene. The computing system can train the machine learning model using a dataset comprising synchronous LiDAR point clouds captured by monocular LiDAR sensors and ground-truth three-dimensional motions obtained from IMU devices. The machine learning model can be configured to generate a three-dimensional motion of the object based on an input of a plurality of point cloud frames captured by a monocular LiDAR sensor.


