Top-Down Trajectory Prediction for Complex Object Interactions

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current prediction techniques for future states of entities in environments, such as autonomous vehicles, often rely on physics-based modeling or rules-of-the-road simulations, which may not accurately account for complex interactions between objects and dynamic environments, leading to inefficiencies in trajectory prediction.

Innovation Solution

A system that uses a machine learning model to process multi-channel images representing top-down views of environments, incorporating semantic information about objects and road networks, to generate trajectory templates and predicted trajectories, accounting for interactions between objects and dynamic conditions.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Reliability

If physics-based modeling or rules-of-the-road simulations are used for trajectory prediction, then the prediction process follows established physical laws, but the accuracy of predicting complex object interactions and dynamic environments deteriorates

Engineering Contradiction:
Improveprediction reliabilityVSAvoidtrajectory prediction accuracy
Core Design Contradiction:
ReliabilityVSMeasurement precision

Solution Approach 1:

The patent replaces physics-based modeling and rules-of-the-road simulations with a machine learning model that learns trajectory patterns directly from historical data. The system processes multi-channel images representing top-down views of environments and uses neural networks to predict future object positions, substituting mechanical simulation approaches with data-driven learning that better captures complex interactions.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

2Reliability

If complex physics-based modeling is used to account for all environmental factors, then comprehensive coverage of dynamic conditions is achieved, but computational complexity and processing time increase

Engineering Contradiction:
Improveenvironmental factor coverageVSAvoidprediction system complexity
Core Design Contradiction:
ReliabilityVSDevice complexity

Solution Approach 1:

The patent substitutes complex physics-based environmental modeling with a machine learning approach that implicitly learns environmental dynamics from data. The system processes multi-channel images containing semantic information about objects and road networks, allowing the neural network to capture environmental factors without explicit physical modeling, thereby reducing computational complexity while maintaining comprehensive coverage.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

3Measurement precision

If detailed semantic information about objects and road networks is processed, then the understanding of environment interactions improves, but the data processing complexity and computational load increase

Engineering Contradiction:
Improveenvironmental understanding precisionVSAvoiddata processing complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The patent merges multiple information channels (object semantics, road network data, spatial relationships) into a unified multi-channel image representation. This integrated approach allows the machine learning model to process all semantic information simultaneously through a single neural network architecture, reducing the complexity of handling separate data streams while improving environmental understanding.

Inventive Principle:
Principle #5Merging (Combining)

Data Source

PatentUS11734832B1Prediction on top-down scenes based on object motion
Publication Date: 2023.08.22 ZOOX INC
  • US11734832B1 patent drawing
  • US11734832B1 patent drawing
  • US11734832B1 patent drawing

AI summary

Techniques for determining predictions on a top-down representation of an environment based on object movement are discussed herein. Sensors of a first vehicle (such as an autonomous vehicle) may capture sensor data of an environment, which may include object(s) separate from the first vehicle (e.g., a vehicle, a pedestrian, a bicycle). A multi-channel image representing a top-down view of the object(s) and the environment may be generated based in part on the sensor data. Environmental data (object extents, velocities, lane positions, crosswalks, etc.) may also be encoded in the image. Multiple images may be generated representing the environment over time and input into a prediction system configured to output a trajectory template (e.g., general intent for future movement) and a predicted trajectory (e.g., more accurate predicted movement) associated with each object. The prediction system may include a machine learned model configured to output the trajectory template(s) and the predicted trajector(ies).