Distributed AI Object Tracking via Coordinate Transformation

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Determining accurate three-dimensional locations of objects within video feeds from moving robotic systems, such as drones, is challenging due to camera orientation and position uncertainties, especially when multiple objects of the same type are present, and is further complicated by factors like vibration and weather conditions.

Innovation Solution

An image processing system that utilizes machine learning models and object recognition algorithms to detect objects within images, retrieves metadata for camera orientation and position, and matches detected objects with known objects in a database using GPS coordinates and object metadata, generating indicators for accurate object identification and tracking.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Difficulty of detecting and measuring

If object detection is performed in video feeds from moving robotic systems, then object identification capability is improved, but measurement precision of three-dimensional location deteriorates due to camera orientation and position uncertainties

Engineering Contradiction:
Improveobject identification capabilityVSAvoidthree-dimensional location accuracy
Core Design Contradiction:
Difficulty of detecting and measuringVSMeasurement precision

Solution Approach 1:

The patent introduces an intermediary coordinate transformation process that mediates between the camera's local coordinate system and the global three-dimensional space. By using camera orientation (roll, pitch, yaw) and position data as intermediate parameters, the system transforms two-dimensional image coordinates into accurate three-dimensional locations, resolving the precision issue caused by camera movement and orientation changes.

Inventive Principle:
Principle #24Intermediary (Mediator)

2Productivity

If multiple objects of the same type are detected, then object detection coverage is improved, but difficulty of linking objects with known objects increases

Engineering Contradiction:
Improveobject detection coverageVSAvoidobject matching difficulty
Core Design Contradiction:
ProductivityVSDifficulty of detecting and measuring

Solution Approach 1:

The patent implements a feedback mechanism where detected objects are linked with known objects using GPS coordinates and metadata verification. The system continuously compares detected object positions with known object databases, uses GPS location feedback to confirm matches, and applies metadata feedback (object type, size, characteristics) to disambiguate between multiple objects of the same type, enabling accurate identification even when multiple similar objects are present.

Inventive Principle:
Principle #23Feedback

3Measurement precision

If camera metadata is collected for location determination, then three-dimensional location accuracy is improved, but device complexity increases

Engineering Contradiction:
Improvethree-dimensional location accuracyVSAvoidsystem complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The patent applies universality by using a multi-functional metadata collection system that gathers camera orientation (roll, pitch, yaw), position (GPS coordinates), and lens parameters (focal length, field of view) through existing camera and navigation system components. This multi-functional approach consolidates multiple measurement functions into a unified metadata framework, improving three-dimensional location accuracy without proportionally increasing device complexity.

Inventive Principle:
Principle #6Universality (Multi-functionality)

Data Source

PatentEP4343700A1Architecture for distributed artificial intelligence augmentation
Publication Date: 2024.03.27 TOMAHAWK ROBOTICS INC
  • EP4343700A1 patent drawingFigure 1
  • EP4343700A1 patent drawingFigure 2
  • EP4343700A1 patent drawingFigure 3

AI summary

Methods and systems are described herein for determining three-dimensional locations of objects within a video stream and linking those objects with known objects. An image processing system may receive an image and image metadata and detect an object and a location of the object within the image. The estimated location of each object is then determined within the three-dimensional space. In addition, the image processing system may retrieve, for a plurality of known objects, a plurality of known locations within the three-dimensional space and determine, based on estimated location and the known location data, which of the known objects matches the detected object in the image. An indicator for the object is then generated at the location of the object within the image.