Synthetic LiDAR Data Generation via Physics and Machine Learning

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing simulation systems for autonomous vehicles primarily focus on simulating physics and behaviors rather than sensory input, which limits the realism and effectiveness of testing perception systems in safety-critical situations.

Innovation Solution

The proposed system combines physics-based rendering with machine-learned models to generate synthetic LiDAR data that accurately mimics real-world data, including geometry and intensity, by processing initial point clouds to predict dropout probabilities and adjust the point cloud accordingly.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Reliability

If physics-based rendering is used to generate synthetic LiDAR data, then the simulation capability is improved, but the realism of sensory input is insufficient

Engineering Contradiction:
Improvesimulation capabilityVSAvoidrealism of sensory input
Core Design Contradiction:
ReliabilityVSMeasurement precision

Solution Approach 1:

The patent combines physics-based rendering with machine learning models to generate synthetic LiDAR data. The physics-based system provides the structural framework while the machine learning component adds realistic sensory characteristics, merging the advantages of both approaches to resolve the contradiction between simulation capability and sensory realism.

Inventive Principle:
Principle #5Merging (Combining)

Solution Approach 2:

The machine learning model acts as an intermediary between the physics-based rendering system and the final synthetic LiDAR data. It processes the physics-generated data and transforms it into more realistic sensory input, mediating the transition from physically accurate but less realistic data to both physically accurate and realistic data.

Inventive Principle:
Principle #24Intermediary (Mediator)

2Measurement precision

If machine learning models are used to predict dropout probabilities, then the quality of synthetic data is improved, but the computational complexity increases

Engineering Contradiction:
Improvequality of synthetic dataVSAvoidcomputational complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The machine learning model is trained in advance on real LiDAR data to learn dropout patterns and characteristics. This preliminary training allows the model to quickly predict dropout probabilities during synthetic data generation without requiring complex real-time computations, thus improving data quality while managing computational complexity.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The machine learning model learns to copy the dropout characteristics from real LiDAR data by training on actual sensor measurements. Instead of modeling the complex physical processes that cause dropouts, the system copies the observed dropout patterns from real data, achieving high quality synthetic data with reduced computational complexity.

Inventive Principle:
Principle #26Copying

Data Source

PatentUS20250130909A1Systems and Methods for Generating Synthetic Sensor Data via Machine Learning
Publication Date: 2025.04.24 AURORA OPERATIONS INC
  • US20250130909A1 patent drawing
  • US20250130909A1 patent drawing
  • US20250130909A1 patent drawing

AI summary

The present disclosure provides systems and methods that combine physics-based systems with machine learning to generate synthetic LiDAR data that accurately mimics a real-world LiDAR sensor system. In particular, aspects of the present disclosure combine physics-based rendering with machine-learned models such as deep neural networks to simulate both the geometry and intensity of the LiDAR sensor. As one example, a physics-based ray casting approach can be used on a three-dimensional map of an environment to generate an initial three-dimensional point cloud that mimics LiDAR data. According to an aspect of the present disclosure, a machine-learned model can predict one or more dropout probabilities for one or more of the points in the initial three-dimensional point cloud, thereby generating an adjusted three-dimensional point cloud which more realistically simulates real-world LiDAR data.