Neural 2D-3D Bounding Box Association for ADAS Training Data
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current methods for generating training data for advanced driver assistance systems (ADAS) rely heavily on human labeling, which is costly and time-consuming, especially for associating 2D and 3D object detection data from camera images and LiDAR point clouds.
Innovation Solution
A computer-implemented method using a neural network to associate 2D bounding boxes from camera images with 3D bounding boxes from LiDAR data, generating training data for a second neural network to perform 2D and 3D object detection in ADAS, thereby automating the labeling process and reducing costs.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If human labeling is used to associate 2D and 3D bounding boxes, then labeling accuracy can be maintained, but cost and time consumption increase significantly
Solution Approach 1:
The system uses automatically generated 3D bounding boxes from LiDAR data to label 2D camera images, eliminating the need for manual human labeling. The neural network performs self-labeling by detecting objects in 3D space and projecting them to 2D image space, thereby maintaining accuracy while dramatically reducing time consumption and costs.
Solution Approach 2:
The patent replaces the mechanical human labeling process with an automated neural network system that uses LiDAR-based 3D object detection to generate labels for 2D images. This substitution eliminates manual intervention while maintaining or improving labeling consistency and accuracy through algorithmic processing.
2Measurement precision
If human labeling is used for training data generation, then data quality can be ensured, but costs increase significantly
Solution Approach 1:
The system generates its own training data automatically by using LiDAR-derived 3D bounding boxes to label corresponding 2D camera images. This self-service approach eliminates the need for expensive human annotators while ensuring data quality through the reliability of LiDAR technology and neural network processing.
Solution Approach 2:
The patent creates training data by copying and transforming 3D object information from LiDAR point clouds into 2D image annotations. The 3D bounding boxes serve as templates that are projected onto 2D image space, creating accurate training labels without requiring manual copying or human intervention.
3Productivity
If automated methods are used for training data generation, then cost and time are reduced, but association accuracy between 2D and 3D data may deteriorate
Solution Approach 1:
The patent solves the association problem by working in 3D space rather than directly in 2D image space. LiDAR provides accurate 3D object locations, which are then projected to 2D image coordinates using camera calibration matrices. This dimensional transformation approach maintains high association accuracy while enabling full automation.
Solution Approach 2:
The patent uses 3D bounding boxes from LiDAR data as an intermediary to associate 2D camera images with object locations. Instead of directly matching 2D features, the system introduces 3D space as an intermediate representation, which provides a common reference frame for both sensor types and improves association accuracy.
Data Source
AI summary
A computer-implemented method comprises: receiving camera images and three dimensional (3D) data captured by a first vehicle during travel; obtaining, by providing the camera images to a two-dimensional (2D) object detection algorithm, 2D bounding boxes corresponding to first objects visible in the camera images; obtaining 3D bounding boxes corresponding to second objects in the 3D data; performing association of the 2D bounding boxes with corresponding ones of the 3D bounding boxes using a first neural network; and generating, using the camera images and the 2D bounding boxes associated with corresponding ones of the 3D bounding boxes, training data for training a second neural network to perform 2D and 3D object detection in an advanced driver assistance system (ADAS) for a second vehicle.


