Vehicle Sensor Labeling via 2D-to-3D Track Aggregation
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current approaches for labeling sensor data from vehicles are labor-intensive, time-consuming, and prone to human error, particularly when dealing with large-scale data sets required for autonomous navigation and HD map creation, as they rely on manual annotation of 2D and 3D data from multiple sources.
Innovation Solution
A computer-based method that processes and aggregates 2D and 3D sensor data to generate time-aggregated 3D visualizations, allowing for automatic labeling of objects and reducing the need for capture-by-capture annotation, using techniques like semantic segmentation and motion modeling to improve data representation and labeling efficiency.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If manual annotation of 2D and 3D sensor data is used, then labeling accuracy can be maintained, but labor intensity and time consumption increase significantly
Solution Approach 1:
The patent segments the labeling process into multiple stages: automatic 2D label generation from sensor data, 3D label generation from 2D labels, and selective manual verification. This segmentation allows most labels to be generated automatically while maintaining quality through targeted manual review, significantly reducing overall time consumption while preserving accuracy.
Solution Approach 2:
The patent introduces 2D labels as an intermediary step between raw sensor data and final 3D labels. The system first generates 2D labels automatically, then uses these as input to generate 3D labels, reducing the complexity and time required for direct 3D annotation while maintaining accuracy through the structured transformation process.
2Manufacturing precision
If capture-by-capture annotation is performed, then detailed labeling can be achieved, but the process becomes extremely labor-intensive for large data sets
Solution Approach 1:
The patent merges multiple captures of sensor data into a single aggregated 3D representation. Instead of annotating each capture separately, the system combines data from multiple captures to create a unified 3D model that can be labeled once, significantly improving productivity while maintaining detailed labeling quality through the accumulated data.
Solution Approach 2:
The patent performs preliminary processing of sensor data by generating 2D labels automatically from raw captures before the main 3D labeling step. This preliminary action reduces the workload for subsequent detailed labeling by pre-identifying objects and their positions, enabling faster and more efficient detailed annotation.
3Loss of information
If multiple sources of 2D and 3D sensor data are integrated, then comprehensive data representation is achieved, but system complexity increases
Solution Approach 1:
The patent implements a universal labeling framework that can process multiple types of sensor data (2D images, 3D point clouds, LiDAR data) through a single integrated system. The same core algorithms and processing pipeline handle different data sources, reducing system complexity while achieving comprehensive data representation through multi-functional capability.
Data Source
AI summary
Examples disclosed herein may involve (i) based on an analysis of 2D data captured by a vehicle while operating in a real-world environment during a window of time, generating a 2D track for at least one object detected in the environment comprising one or more 2D labels representative of the object, (ii) for the object detected in the environment: (a) using the 2D track to identify, within a 3D point cloud representative of the environment, 3D data points associated with the object, and (b) based on the 3D data points, generating a 3D track for the object that comprises one or more 3D labels representative of the object, and (iii) based on the 3D point cloud and the 3D track, generating a time-aggregated, 3D visualization of the environment in which the vehicle was operating during the window of time that includes at least one 3D label for the object.


