3D LiDAR Object Detection Using Targeted Rare-Object Simulation
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Autonomous vehicles face challenges in acquiring sufficient real-world data to adequately train machine learning models, particularly for rare objects, due to limitations in time, cost, and opportunities for data collection, which affects their ability to detect and react to infrequent 3-D objects in real-world scenarios.
Innovation Solution
The approach involves injecting real-world data into simulated scenes to enhance the training of machine learning algorithms, using a placement system to accurately position objects in simulated environments based on their real-world probabilities, thereby creating more realistic and diverse training scenarios.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If real-world data is collected to train machine learning models for rare object detection, then detection accuracy for rare objects improves, but data collection time and cost increase significantly
Solution Approach 1:
The patent creates synthetic copies of rare objects through 3-D representations and simulated scenes. Instead of collecting extensive real-world data of rare objects, the system generates virtual copies that replicate the visual and spatial characteristics of these objects, allowing the machine learning model to learn from abundant synthetic examples without requiring prolonged real-world data collection
Solution Approach 2:
The system varies parameters in simulated scenes such as object positions, orientations, lighting conditions, and environmental contexts to generate diverse training examples. By changing these parameters systematically, the patent creates a comprehensive dataset that covers various scenarios where rare objects might appear, thereby improving detection accuracy without extending data collection time
2Measurement precision
If real-world data is collected to train machine learning models for rare object detection, then detection accuracy for rare objects improves, but training cost increases significantly
Solution Approach 1:
The patent replaces expensive real-world data collection and annotation processes with automated synthetic data generation. Virtual copies of rare objects are created through computer graphics and simulation engines, eliminating the need for costly field data collection campaigns and manual labeling efforts while maintaining high detection accuracy
Solution Approach 2:
The system automatically generates and annotates training data through the simulation environment without requiring human intervention for data collection and labeling. The simulated scenes self-generate ground truth information, reducing the need for expensive manual annotation services and lowering overall training costs
3Quantity of substance
If more real-world driving opportunities are provided to collect data, then training data quantity increases, but operational availability and safety of autonomous vehicles decrease
Solution Approach 1:
The patent generates abundant training data through synthetic copies in virtual environments, eliminating the need to take vehicles off the road for extensive data collection. This allows continuous operational deployment of autonomous vehicles while still accumulating sufficient training data through parallel synthetic data generation processes
Solution Approach 2:
The system performs preliminary data generation and model training using synthetic data before deploying vehicles to real-world operations. By preparing training datasets and refining models in advance through simulation, the patent reduces the need for vehicles to be unavailable for data collection, thereby maintaining higher operational availability and safety
Data Source
AI summary
The subject disclosure relates to techniques for improving performance of a machine learning algorithm that at least receives data descriptive of a 3-D object and provides an output, where the 3-D object occurs infrequently in a training dataset. A process of the disclosed technology can include determining that the machine learning algorithm performed below a threshold performance score when receiving data descriptive of the 3-D object in a real-world scene, wherein the 3-D object is classified as a first type of object, creating at least one 3-D representation of the first type of object for use in a simulation, modifying a plurality of simulated scenes to include the at least one 3-D representation of the first type of object, and training the machine learning algorithm with the modified simulated scenes, whereby the machine learning algorithm has greater exposure to the first type of object.


