Top-Down Prediction Using Image Features

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Autonomous vehicles face challenges in accurately predicting the behavior of objects in their environment, particularly due to the limitations of relying solely on top-down representations generated from sensor data, which may not incorporate essential image features effectively.

Innovation Solution

The integration of image feature representations with top-down representations using machine-learned models, allowing the model to learn and weight features important for predicting object behavior, thereby improving the accuracy and efficiency of behavior prediction.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Device complexity

If autonomous vehicles rely solely on top-down representations generated from sensor data, then the system complexity is reduced, but the accuracy of object behavior prediction deteriorates

Engineering Contradiction:
Improvesystem complexityVSAvoidaccuracy of object behavior prediction
Core Design Contradiction:
Device complexityVSMeasurement precision

Solution Approach 1:

The patent combines top-down representations from sensor data with image feature representations from visual data to create a fused representation. This merging allows the system to maintain relatively simple processing architecture while achieving higher prediction accuracy by leveraging complementary information from multiple data sources.

Inventive Principle:
Principle #5Merging (Combining)

Solution Approach 2:

The machine-learned model is designed to process multiple types of input data (sensor data and image data) through a unified architecture. The model learns to weight and integrate different feature types automatically, making the system adaptable to various data sources while maintaining consistent prediction performance.

Inventive Principle:
Principle #6Universality (Multi-functionality)

2Measurement precision

If autonomous vehicles use machine-learned models to integrate image features with top-down representations, then the accuracy of behavior prediction is improved, but the processing resources required increase

Engineering Contradiction:
Improveaccuracy of behavior predictionVSAvoidprocessing resources
Core Design Contradiction:
Measurement precisionVSUse of energy by moving object

Solution Approach 1:

The system performs preliminary processing of image data to extract key features before feeding them into the machine-learned model. By pre-processing and filtering image data to retain only salient features, the system reduces the computational burden on the main prediction model while preserving the accuracy benefits of image feature integration.

Inventive Principle:
Principle #10Preliminary action

3Measurement precision

If autonomous vehicles use machine-learned models to integrate image features with top-down representations, then the accuracy of behavior prediction is improved, but the device complexity increases

Engineering Contradiction:
Improveaccuracy of behavior predictionVSAvoiddevice complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The patent introduces an intermediary feature extraction module that processes image data and outputs condensed feature representations. This intermediary layer acts as a bridge between raw image data and the main prediction model, reducing the complexity burden on the core system while still enabling accurate behavior prediction through enriched feature inputs.

Inventive Principle:
Principle #24Intermediary (Mediator)

Data Source

PatentUS11409304B1Supplementing top-down predictions with image features
Publication Date: 2022.08.09 ZOOX INC
  • US11409304B1 patent drawing
  • US11409304B1 patent drawing
  • US11409304B1 patent drawing

AI summary

The described techniques relate to predicting object behavior based on top-down representations of an environment comprising top-down representations of image features in the environment. For example, a top-down representation may comprise a multi-channel image that includes semantic map information along with additional information for a target object and/or other objects in an environment. A top-down image feature representation may also be a multi-channel image that incorporates various tensors for different image features with channels of the multi-channel image, and may be generated directly from an input image. A prediction component can generate predictions of object behavior based at least in part on the top-down image feature representation, and in some cases, can generate predictions based on the top-down image feature representation together with the additional top-down representation.