UAV Visual Object Tracking With 3D Trajectory Prediction

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing autonomous vehicle navigation systems face challenges in accurately tracking and navigating around dynamic objects in complex environments, particularly due to issues with object detection, tracking, and trajectory prediction, which can lead to inefficiencies and potential collisions.

Innovation Solution

The system employs an unmanned aerial vehicle (UAV) equipped with multiple image capture devices and a tracking system that utilizes stereoscopic image capture, deep convolutional neural networks, and spatiotemporal factor graphs to track objects, incorporating semantic and 3D geometry information, and predicts future trajectories based on motion models and sensor data.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If visual sensors and image processing are used for object tracking, then navigation accuracy is improved, but computational complexity and processing time increase

Engineering Contradiction:
Improveobject detection accuracyVSAvoidcomputational complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The system performs preliminary actions by pre-processing images to extract key features and pre-calculating trajectory predictions based on motion models before full navigation decisions are made. This reduces the computational burden during real-time tracking while maintaining accuracy.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The navigation system segments the visual processing task into distinct modules: object detection, feature extraction, trajectory prediction, and navigation decision-making. Each module processes specific aspects independently, reducing overall computational complexity while improving detection accuracy through specialized processing.

Inventive Principle:
Principle #1Segmentation

2Reliability

If multiple image capture devices are used for stereoscopic tracking, then object tracking robustness is improved, but device complexity and energy consumption increase

Engineering Contradiction:
Improvetracking robustnessVSAvoidenergy consumption
Core Design Contradiction:
ReliabilityVSUse of energy by moving object

Solution Approach 1:

The system merges data from multiple image capture devices through stereoscopic processing to achieve robust 3D object tracking. By combining information from multiple sensors into a unified spatial model, the system improves tracking reliability while optimizing energy usage through coordinated sensor operation rather than independent processing.

Inventive Principle:
Principle #5Merging (Combining)

3Reliability

If trajectory prediction based on motion models is implemented, then collision avoidance capability is improved, but prediction accuracy deteriorates in complex environments

Engineering Contradiction:
Improvecollision avoidance capabilityVSAvoidtrajectory prediction accuracy
Core Design Contradiction:
ReliabilityVSMeasurement precision

Solution Approach 1:

The trajectory prediction system implements feedback by continuously comparing predicted object positions with actual sensor measurements and adjusting motion model parameters accordingly. This closed-loop approach maintains prediction accuracy in complex environments by adapting to changing conditions while preserving collision avoidance capabilities.

Inventive Principle:
Principle #23Feedback

Data Source

PatentUS12367670B2Object tracking by an unmanned aerial vehicle using visual sensors
Publication Date: 2025.07.22 SKYDIO INC
  • US12367670B2 patent drawing
  • US12367670B2 patent drawing
  • US12367670B2 patent drawing

AI summary

Systems and methods are disclosed for tracking objects in a physical environment using visual sensors onboard an autonomous unmanned aerial vehicle (UAV). In certain embodiments, images of the physical environment captured by the onboard visual sensors are processed to extract semantic information about detected objects. Processing of the captured images may involve applying machine learning techniques such as a deep convolutional neural network to extract semantic cues regarding objects detected in the images. The object tracking can be utilized, for example, to facilitate autonomous navigation by the UAV or to generate and display augmentative information regarding tracked objects to users.