LiDAR-Only 3D Human Pose Estimation via Two-Stage Voxel Processing
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing methods for 3D human pose estimation using LiDAR point clouds face challenges due to the difficulty in acquiring accurate 3D annotations and the reliance on weakly supervised approaches that assume precise calibration between camera and LiDAR inputs.
Innovation Solution
A two-stage top-down 3D human pose estimation framework that uses only LiDAR point cloud data and is trained solely on 3D annotations. The first stage detects human bounding boxes and generates fine-grained voxel features, while the second stage employs a transformer-based network to regress human keypoints using attention mechanisms.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Adaptability or versatility
If weakly supervised approaches using camera and LiDAR fusion are used, then pose estimation can be performed with available data, but precise calibration between camera and LiDAR is required which increases system complexity and reliability issues
Solution Approach 1:
The patent extracts and removes the camera input requirement from the system, creating a LiDAR-only pose estimation framework. This eliminates the need for camera-LiDAR calibration while maintaining pose estimation functionality, directly resolving the technical contradiction between adaptability and device complexity.
Solution Approach 2:
The patent segments the pose estimation task into two stages: first generating 3D bounding boxes and semantic segmentation, then predicting keypoints from these intermediate results. This segmentation allows the system to achieve accurate pose estimation using only LiDAR data without requiring camera fusion or calibration.
2Measurement precision
If 3D annotations are acquired for training, then accurate pose estimation is achieved, but acquiring accurate 3D annotations is difficult and time-consuming
Solution Approach 1:
The patent performs preliminary 3D object detection and semantic segmentation to generate accurate 3D bounding boxes and voxel features before keypoint prediction. These pre-computed 3D structures serve as high-quality training targets that automatically provide precise spatial information without requiring manual 3D annotation, thus achieving measurement precision while avoiding the time-consuming annotation process.
3Measurement precision
If image features are used in addition to LiDAR data, then pose estimation performance improves, but the system requires multiple sensor inputs increasing complexity
Solution Approach 1:
The patent removes the image feature input requirement from the system, demonstrating that high-accuracy pose estimation can be achieved using only LiDAR point cloud data. This extraction of the image modality eliminates sensor fusion complexity while maintaining measurement precision through advanced 3D feature processing.
Solution Approach 2:
The patent transforms the input data into 3D voxel space and processes features in three-dimensional coordinate systems rather than relying on 2D image projections. This dimensional transformation allows the system to achieve image-level accuracy using only LiDAR data by fully exploiting the 3D spatial information inherent in point clouds.
Data Source
AI summary
An image processing method includes performing, using images obtained from one or more sensors onboard a vehicle, a 2-dimensional (2D) feature extraction; performing, a 3-dimensional (3D) feature extraction on the images; detecting objects in the images by fusing detection results from the 2D feature extraction and the 3D feature extraction.


