Neural Encoder Orthographic Transform for 3D Object Tracking

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current vehicle systems face challenges in reliably and precisely tracking objects using image data from cameras, particularly in predicting the position and orientation of objects in the surroundings for robust tracking.

Innovation Solution

A device employing a neural encoder network and a neural evaluation network, combined with an orthographic feature transform, processes image data from cameras to predict 3D object data on a bird's eye view plane, enabling precise and robust tracking of objects by transforming camera-based feature tensors into a grid plane and fusing them with sensor data from other surroundings sensors.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Reliability

If a vehicle uses image data from cameras to track objects in its surroundings, then the system can detect objects, but the tracking reliability and precision are insufficient

Engineering Contradiction:
Improvetracking reliabilityVSAvoidobject position and orientation precision
Core Design Contradiction:
ReliabilityVSMeasurement precision

Solution Approach 1:

The patent transforms 2D image data from the camera's image plane into a 3D bird's eye view grid plane, adding a vertical dimension to the data representation. This dimensional transformation enables more accurate tracking of objects by representing their positions and orientations in three-dimensional space, thereby improving both tracking reliability and measurement precision simultaneously

Inventive Principle:
Principle #17Another dimension (Dimensionality change)

Solution Approach 2:

The patent introduces an orthographic feature transform as an intermediary processing step between the 2D image data and the 3D object representation. This transform acts as a mediator that converts camera-based feature tensors into a standardized 3D grid format, enabling reliable and precise tracking without requiring direct complex calculations between the original image data and object properties

Inventive Principle:
Principle #24Intermediary (Mediator)

2Ease of operation

If the vehicle moves on the roadway, then the vehicle can operate in its intended environment, but the camera's viewpoint changes making consistent object tracking difficult

Engineering Contradiction:
Improvevehicle operationVSAvoidobject data consistency
Core Design Contradiction:
Ease of operationVSStability of the object's composition

Solution Approach 1:

By transforming data into a 3D bird's eye view grid plane, the system creates a viewpoint-independent representation of objects. This 3D grid structure maintains object data consistency regardless of camera viewpoint changes, allowing the vehicle to move freely while maintaining stable object tracking

Inventive Principle:
Principle #17Another dimension (Dimensionality change)

Solution Approach 2:

The patent changes the coordinate system parameters from camera-based 2D image coordinates to a vehicle-centered 3D grid coordinate system. This parameter transformation makes the object representation invariant to camera viewpoint changes, ensuring stable object data composition during vehicle movement

Inventive Principle:
Principle #35Parameter changes

Data Source

PatentUS20240212206A1Method and Device for Predicting Object Data Concerning an Object
Publication Date: 2024.06.27 BAYERISCHE MOTOREN WERKE AG
  • US20240212206A1 patent drawing
  • US20240212206A1 patent drawing
  • US20240212206A1 patent drawing

AI summary

A device for determining object data in relation to an object in the environment of at least one image camera is described. The device is configured to determine a camera-based feature tensor on the basis of at least one image from the image camera for a first point in time by means of a neural encoder network. Furthermore, the device is configured to transform and/or project the camera-based feature tensor from an image plane of the image onto a grid plane of an environment grid of the environment of the image camera in order to determine a transformed feature tensor. The device is furthermore configured to determine object data in relation to the object in the environment of the image camera on the basis of the transformed feature tensor by means of a neural evaluation network, the object data comprising one or more predicted properties of the object at a point in time succeeding the first point in time.