Semantic Label Generation from Sparse 3D Data

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

There is a lack of datasets with semantic labels for 3D object bounding boxes and masks, which hinders the training of machine learning-based computer vision systems for object detection, classification, and segmentation, particularly in autonomous vehicles, due to the time-consuming and expensive nature of manual labeling.

Innovation Solution

The method generates semantic labels for 2D data by converting sparse 3D data into dense 3D data and then assigning labels based on 3D bounding boxes, with the labels being mapped to corresponding 2D data points, reducing the need for human supervision and enhancing the quality of training datasets.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If manual creation of semantically labeled data is performed, then data accuracy and quality are improved, but time consumption and cost increase significantly

Engineering Contradiction:
Improvedata qualityVSAvoidlabeling time
Core Design Contradiction:
Measurement precisionVSLoss of time

Solution Approach 1:

The system enables self-service by using automatically generated semantic labels from sparse 3D data to create training datasets without requiring manual human annotation. The computer vision system generates its own training data through automated processing of sensor inputs, eliminating the need for external human labeling while maintaining data quality through algorithmic consistency and reproducibility.

Inventive Principle:
Principle #25Self-service

2Measurement precision

If manual creation of semantically labeled data is performed, then data accuracy and quality are improved, but cost increases significantly

Engineering Contradiction:
Improvedata qualityVSAvoidcost
Core Design Contradiction:
Measurement precisionVSQuantity of substance

Solution Approach 1:

The system eliminates external human annotation services by enabling the computer vision system to automatically generate semantically labeled training data from sparse 3D sensor inputs. This self-service approach removes the need for expensive human labor while maintaining data quality through automated processing pipelines and algorithmic label generation.

Inventive Principle:
Principle #25Self-service

Solution Approach 2:

The system creates synthetic copies of real-world sensor data by generating dense 3D data from sparse 3D inputs and automatically assigning semantic labels. These synthesized training datasets replicate the structure and characteristics of manually labeled data without requiring actual human annotation, providing a cost-effective alternative that maintains data fidelity.

Inventive Principle:
Principle #26Copying

3Speed

If sparse 3D data is directly used for training, then data processing speed is improved, but label accuracy and completeness deteriorate

Engineering Contradiction:
Improveprocessing speedVSAvoidlabel accuracy
Core Design Contradiction:
SpeedVSMeasurement precision

Solution Approach 1:

The system performs preliminary densification of sparse 3D data before semantic label assignment, converting sparse point clouds into dense 3D representations that contain sufficient detail for accurate labeling. This preliminary processing step ensures that subsequent label generation operates on complete data structures, improving label accuracy while maintaining efficient processing through automated pipelines.

Inventive Principle:
Principle #10Preliminary action

Data Source

PatentUS10983217B2Method and system for semantic label generation using sparse 3D data
Publication Date: 2021.04.20 HUAWEI TECH CO LTD
  • US10983217B2 patent drawing
  • US10983217B2 patent drawing
  • US10983217B2 patent drawing

AI summary

Methods and apparatuses for generating a frame of semantically labeled 2D data are described. A frame of sparse 3D data is generated from a frame of sparse 3D data. Semantic labels are assigned to the frame of dense 3D data, based on a set of 3D bounding boxes determined for the frame of sparse 3D data. Semantic labels are assigned to a corresponding frame of 2D data based on a mapping between the frame of sparse 3D data and the frame of 2D data. The mapping is used to map a 3D data point in the frame of dense 3D data to a mapped 2D data point in the frame of 2D data. The semantic label assigned to the 3D data point is assigned to the mapped 2D data point. The frame of semantically labeled 2D data, including the assigned semantic labels, is outputted.