Range-View Sensor Fusion for Autonomous Motion Prediction

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Autonomous vehicles face challenges in accurately detecting and predicting the movement of objects using separate detection and prediction systems, which can be inefficient and prone to errors due to the need for converting sensor data between different coordinate frames.

Innovation Solution

A computer-implemented method and system that fuse multiple sensor sweeps into a common coordinate frame, using machine-learned models to transform and map feature data, allowing for accurate prediction of object movement without the need for bird's eye view conversion, thereby improving detection and prediction efficiency.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Reliability

If separate detection and prediction systems are used, then system modularity is maintained, but accuracy and efficiency of object detection and prediction deteriorate due to coordinate frame conversions

Engineering Contradiction:
Improveaccuracy of object detection and predictionVSAvoidcomplexity of separate detection and prediction systems
Core Design Contradiction:
ReliabilityVSDevice complexity

Solution Approach 1:

The patent combines separate detection and prediction systems into a unified system that processes sensor sweeps through a single coordinate frame. The fused representation integrates detection features and prediction features from multiple sensor sweeps without requiring conversions between different coordinate frames, thereby improving accuracy while reducing system complexity.

Inventive Principle:
Principle #5Merging (Combining)

2Ease of operation

If bird's eye view conversion is used, then a unified perspective is achieved, but processing cycles and energy consumption increase

Engineering Contradiction:
Improveunified perspective for detectionVSAvoidenergy consumption for data processing
Core Design Contradiction:
Ease of operationVSUse of energy by moving object

Solution Approach 1:

The patent extracts and eliminates the unnecessary bird's eye view conversion step from the processing pipeline. By maintaining sensor data in its original coordinate frame and fusing multiple sweeps directly, the system achieves unified perspective without the computational overhead of conversion, reducing processing cycles and energy consumption.

Inventive Principle:
Principle #2Taking out (Extraction)

3Adaptability or versatility

If multiple coordinate frames are used, then specialized processing for each view is enabled, but data storage needs and processing overhead increase

Engineering Contradiction:
Improvespecialized processing capabilityVSAvoiddata storage requirements
Core Design Contradiction:
Adaptability or versatilityVSQuantity of substance

Solution Approach 1:

The patent creates a universal fused representation that handles multiple sensor sweeps in a single coordinate frame. This unified structure serves multiple functions: detection, prediction, and temporal fusion, eliminating the need to maintain separate data structures for different coordinate frames and reducing overall data storage requirements.

Inventive Principle:
Principle #6Universality (Multi-functionality)

Data Source

PatentUS11762094B2Systems and methods for object detection and motion prediction by fusing multiple sensor sweeps into a range view representation
Publication Date: 2023.09.19 AURORA OPERATIONS INC
  • US11762094B2 patent drawing
  • US11762094B2 patent drawing
  • US11762094B2 patent drawing

AI summary

Systems and methods for detecting objects and predicting their motion are provided. In particular, a computing system can obtain a plurality of sensor sweeps. The computing system can determine movement data associated with movement of the autonomous vehicle. For each sensor sweep, the computing system can generate an image associated with the sensor sweep. The computing system can extract, using the respective image as input to one or more machine-learned models, feature data from the respective image. The computing system can transform the feature data into a coordinate frame associated with a next time step. The computing system can generate a fused image. The computing system can generate a final fused image. The computing system can predict, based, at least in part, on the final fused representation of the plurality of sensors sweeps from the plurality of sensor sweeps, movement associated with the feature data at one or more time steps in the future.