LidarMultiNet 3D Voxel Network for Unified LiDAR Perception

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing LiDAR-based multi-task networks struggle to achieve state-of-the-art performance in unifying 3D semantic segmentation, 3D object detection, and panoptic segmentation tasks, often underperforming compared to single-task networks.

Innovation Solution

The proposed LidarMultiNet employs a strong 3D voxel-based encoder-decoder architecture with a Global Context Pooling (GCP) module and task-specific heads to unify the mentioned LiDAR perception tasks, achieving state-of-the-art performance by leveraging synergy among tasks.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Device complexity

If existing LiDAR-based multi-task networks are used to unify 3D semantic segmentation, 3D object detection, and panoptic segmentation tasks, then device complexity is reduced, but measurement precision deteriorates as they underperform compared to single-task networks

Engineering Contradiction:
Improvenetwork architecture complexityVSAvoidsegmentation accuracy
Core Design Contradiction:
Device complexityVSMeasurement precision

Solution Approach 1:

The network is divided into distinct task-specific heads (semantic segmentation head, object detection head, panoptic segmentation head) that process features from a shared backbone separately, allowing each task to be optimized independently while sharing computational resources

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

A unified multi-task network architecture is designed that can perform semantic segmentation, object detection, and panoptic segmentation simultaneously through a shared backbone and task-specific heads, making the system multi-functional without requiring separate networks for each task

Inventive Principle:
Principle #6Universality (Multi-functionality)

2Measurement precision

If single-task networks are used for each LiDAR perception task, then measurement precision is improved, but device complexity increases due to multiple separate networks

Engineering Contradiction:
Improvesegmentation accuracyVSAvoidnetwork architecture complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

Multiple single-task networks are merged into a single multi-task network by sharing the backbone feature extraction layers and introducing task-specific heads, reducing overall system complexity while maintaining task-specific performance through specialized processing modules

Inventive Principle:
Principle #5Merging (Combining)

3Ease of operation

If a unified multi-task network is designed to perform multiple LiDAR perception tasks, then ease of operation is improved, but measurement precision deteriorates due to performance underperformance compared to single-task networks

Engineering Contradiction:
Improvesystem operation simplicityVSAvoiddetection accuracy
Core Design Contradiction:
Ease of operationVSMeasurement precision

Solution Approach 1:

Different parts of the network have specialized functions: the backbone provides general feature extraction, while task-specific heads provide specialized processing for semantic segmentation, object detection, and panoptic segmentation, ensuring each task receives optimized processing

Inventive Principle:
Principle #3Local quality

Data Source

PatentUS20250086802A1Detection of objects in lidar point clouds
Publication Date: 2025.03.13 CREATEAI INC
  • US20250086802A1 patent drawing
  • US20250086802A1 patent drawing
  • US20250086802A1 patent drawing

AI summary

A method of processing point cloud information includes converting points in a point cloud obtained from a lidar sensor into a voxel grid, generating, from the voxel grid, sparse voxel features by applying a multi-layer perceptron and one or more max pooling layers that reduce dimension of input data; applying a cascade of an encoder that performs a N-stage sparse-to-dense feature operation, a global context pooling (GCP) module, and an M-stage decoder that performs a dense-to-sparse feature generation operation. The GCP module bridges an output of a last stage of the N-stages with an input of a first stage of the M-stages, where N and M are positive integers. The GCP module comprises a multi-scale feature extractor; and performing one or more perception operations on an output of the M-stage decoder and/or an output of the GCP module.