LiDARCap Markerless 3D Motion Capture with ML

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing technologies face challenges in capturing long-range three-dimensional human motion accurately, especially with sparse and noisy point clouds from LiDAR sensors, and lack a large-scale dataset with accurate annotations.

Innovation Solution

A marker-less, long-range, and data-driven motion capture method using a single LiDAR sensor, known as LiDARCap, which trains machine learning models on a large dataset like LiDARHuman26M, incorporating synchronous LiDAR point clouds and IMU-captured ground-truth motions to generate accurate three-dimensional human motions.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Length of stationary object

If LiDAR sensors are used for long-range motion capture, then capturing distance is improved, but point cloud density and accuracy deteriorate

Engineering Contradiction:
Improvecapturing distanceVSAvoidpoint cloud accuracy
Core Design Contradiction:
Length of stationary objectVSMeasurement precision

Solution Approach 1:

The patent introduces an intermediary processing pipeline that includes point cloud filtering, feature extraction, and machine learning-based pose estimation. The system uses intermediate representations (keypoints, skeletal structure) to bridge the gap between sparse LiDAR point clouds and accurate 3D motion reconstruction, enabling long-range capture while maintaining precision through multi-stage processing

Inventive Principle:
Principle #24Intermediary (Mediator)

2Measurement precision

If marker-based motion capture is used, then measurement precision is improved, but device complexity and ease of operation deteriorate

Engineering Contradiction:
Improvemotion capture accuracyVSAvoidsystem complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The patent extracts and eliminates the complex marker-based tracking system, replacing it with a markerless LiDAR-based approach. By removing the need for physical markers and complex calibration procedures, the system achieves simpler operation while maintaining accuracy through machine learning models trained on synthetic data with ground truth annotations

Inventive Principle:
Principle #2Taking out (Extraction)

3Measurement precision

If large-scale dataset with accurate annotations is created, then training accuracy is improved, but data collection complexity and time deteriorate

Engineering Contradiction:
Improveground truth accuracyVSAvoiddata collection time
Core Design Contradiction:
Measurement precisionVSLoss of time

Solution Approach 1:

The patent performs preliminary actions by pre-training machine learning models on large-scale synthetic datasets with perfectly annotated ground truth before actual data collection. This pre-training enables the system to handle real-world LiDAR data more effectively, reducing the time needed for annotation during actual motion capture sessions while maintaining high training accuracy

Inventive Principle:
Principle #10Preliminary action

Applied Scientific Principles

This section explains which scientific principles are used to turn an abstract innovation direction into a practical engineering solution.

Function Achieved in This Case

LiDARCap effectively captures highly accurate three-dimensional human motions in long-range scenarios, outperforming state-of-the-art image-based methods, and provides a robust solution for applications like VR/AR and action quality assessment.

Implementation Method 1

a first long-range LiDAR-based motion capture dataset, the LiDARHuman26M dataset, is provided herein

Methodology Applied
Scientific EffectLight detection and ranging (LiDAR): LIDAR

Data Source

PatentUS12270910B2System and method of capturing three-dimensional human motion capture with LiDAR
Publication Date: 2025.04.08 XIAMEN UNIV
  • US12270910B2 patent drawing
  • US12270910B2 patent drawing
  • US12270910B2 patent drawing

AI summary

Described herein are systems and methods for training machine learning models to generate three-dimensional (3D) motions based on light detection and ranging (LiDAR) point clouds. In various embodiments, a computing system can encode a machine learning model representing an object in a scene. The computing system can train the machine learning model using a dataset comprising synchronous LiDAR point clouds captured by monocular LiDAR sensors and ground-truth three-dimensional motions obtained from IMU devices. The machine learning model can be configured to generate a three-dimensional motion of the object based on an input of a plurality of point cloud frames captured by a monocular LiDAR sensor.