3D Object Perception via Multi-Modal Cost Function Optimization

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current perception techniques for autonomous vehicles are often unsuitable for real-time applications due to non-causality or non-real-time processing constraints, necessitating the development of methods that can accurately and reliably detect 3D objects in sensor data from multiple modalities for safe navigation and scenario reconstruction.

Innovation Solution

A computer-implemented method optimizing a cost function across multiple time-series of sensor data from various modalities to locate and model 3D objects, incorporating shape and motion parameters, and environmental constraints, which aggregates error across time and modalities to provide accurate pseudo-ground truth for offline and online applications.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If offline perception techniques are used to improve measurement precision, then productivity decreases due to non-real-time processing

Engineering Contradiction:
Improveperception accuracyVSAvoidreal-time processing capability
Core Design Contradiction:
Measurement precisionVSProductivity

Solution Approach 1:

The system dynamically adapts the perception processing mode based on operational context. Online perception algorithms are used during real-time autonomous vehicle operation to meet timing constraints, while offline perception techniques are applied during post-processing and simulation phases where higher accuracy is prioritized over speed. This dynamic switching resolves the contradiction by allowing both high precision and real-time performance in their respective appropriate contexts.

Inventive Principle:
Principle #15Dynamics

Solution Approach 2:

The perception system is segmented into multiple independent modules: online perception for real-time detection, offline perception for high-precision post-processing, and a hybrid approach for intermediate scenarios. Each segment handles specific tasks based on timing requirements, allowing the system to achieve both real-time responsiveness and high measurement precision without requiring a single approach to excel at both.

Inventive Principle:
Principle #1Segmentation

2Reliability

If multiple sensor modalities are integrated to improve reliability, then device complexity increases

Engineering Contradiction:
Improvedetection reliabilityVSAvoidsensor system complexity
Core Design Contradiction:
ReliabilityVSDevice complexity

Solution Approach 1:

A unified perception framework is implemented that can process multiple sensor modalities (camera, LiDAR, radar) through a common cost function optimization architecture. The same optimization engine handles different sensor types by adjusting the cost function terms accordingly, rather than requiring separate processing pipelines for each modality. This multi-functional approach improves reliability through sensor fusion while controlling system complexity by reusing core processing components.

Inventive Principle:
Principle #6Universality (Multi-functionality)

Solution Approach 2:

The patent merges multiple sensor modalities and multiple time-series data into a unified cost function that simultaneously optimizes for all inputs. Rather than processing each sensor separately and combining results, the system combines all sensor data at the cost function level, allowing joint optimization that improves reliability while avoiding the complexity of multiple separate processing chains.

Inventive Principle:
Principle #5Merging (Combining)

3Ease of manufacture

If offline processing is used to reduce human annotation effort, then loss of time increases due to non-real-time operation

Engineering Contradiction:
Improveannotation effortVSAvoidprocessing time
Core Design Contradiction:
Ease of manufactureVSLoss of time

Solution Approach 1:

The system performs preliminary automated processing of sensor data using offline perception techniques to generate initial annotations and pseudo-ground truth before any human annotation occurs. This preliminary action reduces the volume and complexity of data that requires manual annotation, significantly reducing overall annotation effort. The preliminary automated results serve as a foundation that humans can then refine rather than creating annotations from scratch.

Inventive Principle:
Principle #10Preliminary action

Data Source

PatentUS20240338916A1Perception of 3D objects in sensor data
Publication Date: 2024.10.10 FIVE AI LTD
  • US20240338916A1 patent drawing
  • US20240338916A1 patent drawing
  • US20240338916A1 patent drawing

AI summary

To locate and model a 3D object captured in multiple time-series of sensor data of multiple sensor modalities, a cost function applied to the multiple time-series of sensor data is optimized. The cost function aggregates over time and the multiple sensor modalities, and is defined over a set of variables comprising one or more shape parameters of a 3D object model and a time sequence of poses of the 3D object model. The cost function penalizes inconsistency between the multiple time-series of sensor data and the set of variables. The object belongs to a known object class, and the 3D object model or the cost function encodes expected 3D shape information associated with the known object class, whereby the 3D object is located at multiple time instants and modelled by tuning each pose and the shape parameters with the objective of optimizing the cost function.