Multi-Camera Agent Trajectory Derivation Without LiDAR

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current methods for collecting prior agent trajectories, such as those using LiDAR-based sensor systems, are expensive and limited in scalability due to the high cost of equipment and restricted geographic availability, making it impractical to collect accurate and large-scale data for applications like autonomous driving and transportation matching platforms.

Innovation Solution

The use of monocular cameras and stereo camera systems, integrated with telematics sensors, to capture and process image data for deriving agent trajectories, which involves identifying tracking points, estimating depth, and applying motion models to correct for occlusions and inaccuracies, allowing for more affordable and widespread data collection.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If LiDAR-based sensor systems are used to collect agent trajectories, then measurement precision is improved, but device complexity and cost increase significantly

Engineering Contradiction:
Improvetrajectory data accuracyVSAvoidsensor system complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The patent uses monocular camera images as a simplified copy or alternative representation of the complex LiDAR point cloud data. Instead of directly processing expensive LiDAR data, the system captures 2D image data that can be processed to derive trajectory information, effectively creating a cheaper substitute that maintains sufficient accuracy for the application

Inventive Principle:
Principle #26Copying

Solution Approach 2:

The patent replaces expensive, complex LiDAR sensors with inexpensive monocular cameras that are already widely available in consumer vehicles. This substitution uses cheap, readily available hardware to achieve the same functional goal of collecting trajectory data at scale, making the system economically viable for large-scale deployment

Inventive Principle:
Principle #27Cheap short-living objects (Disposable)

2Measurement precision

If LiDAR-based sensor systems are deployed, then trajectory data accuracy is improved, but scalability deteriorates due to high cost and limited geographic availability

Engineering Contradiction:
Improvetrajectory data accuracyVSAvoiddata collection scalability
Core Design Contradiction:
Measurement precisionVSProductivity

Solution Approach 1:

The patent leverages the universality of monocular cameras, which are already present in most modern vehicles for basic functions like dashcams or driver assistance. By repurposing this existing, universal component for trajectory data collection, the system achieves scalability without requiring specialized equipment deployment across different geographic regions

Inventive Principle:
Principle #6Universality (Multi-functionality)

Solution Approach 2:

By using inexpensive monocular cameras instead of expensive LiDAR systems, the patent enables widespread deployment across many vehicles in diverse geographic locations. The low cost of the camera system removes the economic barrier that limits LiDAR scalability, allowing for large-scale data collection from diverse real-world scenarios

Inventive Principle:
Principle #27Cheap short-living objects (Disposable)

3Device complexity

If monocular cameras are used instead of LiDAR, then device cost decreases and scalability improves, but measurement precision and depth estimation accuracy worsen

Engineering Contradiction:
Improvehardware costVSAvoiddepth estimation accuracy
Core Design Contradiction:
Device complexityVSMeasurement precision

Solution Approach 1:

The patent addresses the monocular camera's inherent lack of depth information by introducing temporal dimensionality. Instead of relying on single-frame depth cues, the system tracks agents across multiple frames and uses motion models to infer 3D trajectory from 2D position sequences, effectively compensating for the missing depth dimension through time-based analysis

Inventive Principle:
Principle #17Another dimension (Dimensionality change)

Solution Approach 2:

The patent introduces motion models as an intermediary computational layer that bridges the gap between 2D monocular image data and 3D trajectory estimation. These models use physical constraints and motion patterns to generate plausible 3D trajectories from 2D observations, acting as a mediator that translates limited visual information into accurate trajectory data

Inventive Principle:
Principle #24Intermediary (Mediator)

4Loss of information

If tracking is performed on partially occluded agents, then trajectory completeness is improved, but measurement precision deteriorates due to occlusion handling challenges

Engineering Contradiction:
Improvetrajectory completenessVSAvoidposition estimation accuracy
Core Design Contradiction:
Loss of informationVSMeasurement precision

Solution Approach 1:

The patent applies motion models in advance to predict agent positions during occlusion periods. By using the agent's historical motion patterns and physical constraints, the system can infer missing trajectory points before they are lost to occlusion, maintaining trajectory completeness without relying on potentially inaccurate visual data during occluded periods

Inventive Principle:
Principle #10Preliminary action

Data Source

PatentUS11961304B2Systems and methods for deriving an agent trajectory based on multiple image sources
Publication Date: 2024.04.16 LYFT INC
  • US11961304B2 patent drawing
  • US11961304B2 patent drawing
  • US11961304B2 patent drawing

AI summary

Examples disclosed herein may involve a computing system that is operable to (i) receive a first sequence of images captured by a monocular camera associated with a vehicle during a given period of operation and a second sequence of image pairs captured by a stereo camera associated with the vehicle during the given period of operation, (ii) derive, from the first sequence of images captured by the monocular camera, a first track for a given agent that comprises a first sequence of position information for the given agent, (iii) derive, from the second sequence of image pairs captured by the stereo camera, a second track for the given agent that comprises a second sequence of position information for the given agent, and (iv) determine a trajectory for the given agent based on the first and second tracks for the given agent.