LiDAR-Only 3D Human Pose Estimation via Two-Stage Voxel Processing

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing methods for 3D human pose estimation using LiDAR point clouds face challenges due to the difficulty in acquiring accurate 3D annotations and the reliance on weakly supervised approaches that assume precise calibration between camera and LiDAR inputs.

Innovation Solution

A two-stage top-down 3D human pose estimation framework that uses only LiDAR point cloud data and is trained solely on 3D annotations. The first stage detects human bounding boxes and generates fine-grained voxel features, while the second stage employs a transformer-based network to regress human keypoints using attention mechanisms.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Adaptability or versatility

If weakly supervised approaches using camera and LiDAR fusion are used, then pose estimation can be performed with available data, but precise calibration between camera and LiDAR is required which increases system complexity and reliability issues

Engineering Contradiction:
Improvepose estimation capabilityVSAvoidcalibration requirement
Core Design Contradiction:
Adaptability or versatilityVSDevice complexity

Solution Approach 1:

The patent extracts and removes the camera input requirement from the system, creating a LiDAR-only pose estimation framework. This eliminates the need for camera-LiDAR calibration while maintaining pose estimation functionality, directly resolving the technical contradiction between adaptability and device complexity.

Inventive Principle:
Principle #2Taking out (Extraction)

Solution Approach 2:

The patent segments the pose estimation task into two stages: first generating 3D bounding boxes and semantic segmentation, then predicting keypoints from these intermediate results. This segmentation allows the system to achieve accurate pose estimation using only LiDAR data without requiring camera fusion or calibration.

Inventive Principle:
Principle #1Segmentation

2Measurement precision

If 3D annotations are acquired for training, then accurate pose estimation is achieved, but acquiring accurate 3D annotations is difficult and time-consuming

Engineering Contradiction:
Improvepose estimation accuracyVSAvoidannotation acquisition time
Core Design Contradiction:
Measurement precisionVSLoss of time

Solution Approach 1:

The patent performs preliminary 3D object detection and semantic segmentation to generate accurate 3D bounding boxes and voxel features before keypoint prediction. These pre-computed 3D structures serve as high-quality training targets that automatically provide precise spatial information without requiring manual 3D annotation, thus achieving measurement precision while avoiding the time-consuming annotation process.

Inventive Principle:
Principle #10Preliminary action

3Measurement precision

If image features are used in addition to LiDAR data, then pose estimation performance improves, but the system requires multiple sensor inputs increasing complexity

Engineering Contradiction:
Improvepose estimation accuracyVSAvoidsensor fusion requirement
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The patent removes the image feature input requirement from the system, demonstrating that high-accuracy pose estimation can be achieved using only LiDAR point cloud data. This extraction of the image modality eliminates sensor fusion complexity while maintaining measurement precision through advanced 3D feature processing.

Inventive Principle:
Principle #2Taking out (Extraction)

Solution Approach 2:

The patent transforms the input data into 3D voxel space and processes features in three-dimensional coordinate systems rather than relying on 2D image projections. This dimensional transformation allows the system to achieve image-level accuracy using only LiDAR data by fully exploiting the 3D spatial information inherent in point clouds.

Inventive Principle:
Principle #17Another dimension (Dimensionality change)

Data Source

PatentUS20250086828A1Human pose estimation from point cloud data
Publication Date: 2025.03.13 CREATEAI INC
  • US20250086828A1 patent drawing
  • US20250086828A1 patent drawing
  • US20250086828A1 patent drawing

AI summary

An image processing method includes performing, using images obtained from one or more sensors onboard a vehicle, a 2-dimensional (2D) feature extraction; performing, a 3-dimensional (3D) feature extraction on the images; detecting objects in the images by fusing detection results from the 2D feature extraction and the 3D feature extraction.