LiDAR Motion Mask Detection for Occlusion-Robust Object Tracking
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Conventional object detection systems for autonomous vehicles are limited by the need for extensive training data and struggle with texture information in LiDAR data due to sensor viewpoint changes and occlusion, leading to inadequate object detection and tracking.
Innovation Solution
A system using LiDAR range images and projection images, processed by a lightweight convolutional neural network, to generate motion masks and vectors by comparing depth values across frames, accounting for noise and occlusion through multiple frame inputs.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Adaptability or versatility
If conventional object detectors are used to identify objects across frames, then object detection can be performed, but extensive training data is required and detection is limited to pre-trained object types
Solution Approach 1:
The patent replaces conventional object detectors with a motion detection system that uses optical flow algorithms to detect pixel-level motion vectors between LiDAR frames. This substitution eliminates the need for trained object detectors while enabling detection of any moving object regardless of pre-training categories.
Solution Approach 2:
The patent introduces projection images as an intermediary representation that transforms 3D LiDAR point cloud data into 2D images with depth encoding. This intermediary format enables direct motion analysis without requiring object detection, bridging the gap between raw sensor data and motion information.
2Measurement precision
If optical flow approaches are used to find pixel-level flow field, then motion tracking can be achieved, but adequate texture information is required which is challenging to generate with LiDAR due to viewpoint changes and occlusion
Solution Approach 1:
The patent projects 3D LiDAR point cloud data onto a 2D plane to create projection images, adding a depth dimension encoding to the 2D representation. This dimensional transformation preserves motion information while creating a format suitable for optical flow analysis, overcoming the texture limitation of direct LiDAR data.
Solution Approach 2:
The patent changes the parameter representation of LiDAR data by encoding depth information into the projection image values. Instead of relying on surface texture, the system uses depth parameter variations across frames to enable reliable motion detection even under viewpoint changes and occlusion.
3Loss of information
If LiDAR data is used for object detection, then 3D spatial information is available, but texture information is insufficient for accurate tracking across frames due to sensor viewpoint changes and occlusion
Solution Approach 1:
The patent transforms 3D LiDAR point cloud data into 2D projection images while preserving depth information through encoding. This dimensional transformation maintains 3D spatial information in a format that enables 2D image processing techniques like optical flow to function effectively.
Data Source
AI summary
In various examples, systems and methods of the present disclosure detect and/or track objects in an environment using projection images generated from LiDAR. For example, a machine learning model—such as a deep neural network (DNN)—may be used to compute a motion mask indicative of motion corresponding to points representing objects in an environment. Various input channels may be provided as input to the machine learning model to compute a motion mask. One or more comparison images may be generated based on comparing depth values projected from a current range image to a coordinate space of a previous range image to depth values of the previous range image. The machine learning model may use the one or more projection images, the one or more comparison images, and/or the one or more range images to compute a motion mask and/or a motion vector output representation.


