Top-Down Image Features for Accurate Object Behavior Prediction

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Autonomous vehicles face challenges in accurately predicting the behavior of objects in their environment, particularly due to the limitations of relying solely on top-down representations generated from sensor data, which may not incorporate essential image features effectively.

Innovation Solution

The integration of image feature representations with top-down representations using machine-learned models, allowing the model to learn and weigh relevant features from image data without prior enumeration, and adjust parameters for improved prediction accuracy.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Device complexity

If autonomous vehicles rely solely on top-down representations generated from sensor data, then the system complexity is reduced, but the prediction accuracy of object behavior deteriorates

Engineering Contradiction:
Improvesystem complexityVSAvoidprediction accuracy
Core Design Contradiction:
Device complexityVSMeasurement precision

Solution Approach 1:

The patent combines top-down representations from sensor data with image feature representations from camera data to create a hybrid prediction system. This merging allows the system to maintain relatively simple architecture while achieving improved prediction accuracy by leveraging complementary information from multiple data sources.

Inventive Principle:
Principle #5Merging (Combining)

Solution Approach 2:

The system processes multiple types of data (sensor data for top-down views and image data for feature extraction) through a unified machine-learned model that performs both representation generation and behavior prediction functions, reducing overall system complexity while maintaining high prediction accuracy.

Inventive Principle:
Principle #6Universality (Multi-functionality)

2Use of energy by moving object

If autonomous vehicles use traditional sensor data processing methods, then processing resources are conserved, but the speed of decision-making deteriorates

Engineering Contradiction:
Improveprocessing resourcesVSAvoiddecision-making speed
Core Design Contradiction:
Use of energy by moving objectVSSpeed

Solution Approach 1:

The system performs preliminary processing by generating top-down representations and extracting image features in advance before behavior prediction is needed. This pre-processing allows the main prediction model to work more efficiently, achieving faster decision-making speeds while managing processing resources effectively.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The processing pipeline is segmented into distinct stages: sensor data processing for top-down views, image processing for feature extraction, and final behavior prediction. This segmentation allows each component to be optimized independently, improving overall processing efficiency and decision-making speed.

Inventive Principle:
Principle #1Segmentation

3Measurement precision

If autonomous vehicles use detailed image feature analysis, then prediction accuracy is improved, but processing latency increases

Engineering Contradiction:
Improveprediction accuracyVSAvoidprocessing latency
Core Design Contradiction:
Measurement precisionVSLoss of time

Solution Approach 1:

The system extracts only the most relevant image features needed for behavior prediction rather than processing all image data in detail. This selective extraction maintains high prediction accuracy by focusing on critical features while significantly reducing processing latency.

Inventive Principle:
Principle #2Taking out (Extraction)

Solution Approach 2:

The system applies partial processing to image data by generating top-down representations at reduced resolution or with selective feature extraction, achieving sufficient prediction accuracy for safety-critical applications while minimizing processing latency through controlled reduction of processing depth.

Inventive Principle:
Principle #16Partial or excessive action

Data Source

PatentUS11380108B1Supplementing top-down predictions with image features
Publication Date: 2022.07.05 ZOOX INC
  • US11380108B1 patent drawing
  • US11380108B1 patent drawing
  • US11380108B1 patent drawing

AI summary

The described techniques relate to predicting object behavior based on top-down representations of an environment comprising top-down representations of image features in the environment. For example, a top-down representation may comprise a multi-channel image that includes semantic map information along with additional information for a target object and/or other objects in an environment. A top-down image feature representation may also be a multi-channel image that incorporates various tensors for different image features with channels of the multi-channel image, and may be generated directly from an input image. A prediction component can generate predictions of object behavior based at least in part on the top-down image feature representation, and in some cases, can generate predictions based on the top-down image feature representation together with the additional top-down representation.