Pseudo-LiDAR 3D Object Localization from Camera Image Sequences

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Autonomous vehicles face challenges in determining the three-dimensional location of objects from camera images when lidar sensors are unavailable or malfunctioning, as camera images lack direct three-dimensional information, leading to inaccurate depth estimates from dense depth maps.

Innovation Solution

A system that combines initial depth estimates with image features to generate pseudo-lidar representations, using neural networks to enhance depth predictions by incorporating both pseudo-lidar and image patch features, particularly from multiple camera images.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Adaptability or versatility

If camera images are used to determine three-dimensional location of objects, then the system can operate without lidar sensors, but the depth estimates become inaccurate because camera images lack direct three-dimensional information

Engineering Contradiction:
Improvesensor availabilityVSAvoiddepth estimation accuracy
Core Design Contradiction:
Adaptability or versatilityVSMeasurement precision

Solution Approach 1:

The patent introduces pseudo-lidar representations as an intermediary that bridges the gap between 2D camera images and 3D spatial information. These pseudo-lidar features are generated from initial depth estimates and serve as a mediator that enables the neural network to infer accurate three-dimensional locations from two-dimensional image data, effectively resolving the contradiction between sensor availability and measurement precision

Inventive Principle:
Principle #24Intermediary (Mediator)

Solution Approach 2:

The patent transforms two-dimensional camera image data into three-dimensional location predictions by generating pseudo-lidar representations that add depth information. This dimensionality change is achieved through neural networks that process initial depth estimates and image features to produce accurate 3D object locations, allowing the system to operate without direct 3D sensors

Inventive Principle:
Principle #17Another dimension (Dimensionality change)

2Loss of information

If dense depth maps are generated from camera images, then three-dimensional information can be obtained, but the depth estimates remain noisy and inaccurate

Engineering Contradiction:
Improvethree-dimensional information availabilityVSAvoiddepth map accuracy
Core Design Contradiction:
Loss of informationVSMeasurement precision

Solution Approach 1:

The patent creates pseudo-lidar representations that copy and enhance the depth information contained in initial depth estimates. Rather than directly using noisy dense depth maps, the system generates refined pseudo-lidar features through neural networks that replicate accurate three-dimensional structure from two-dimensional image data, thereby obtaining 3D information without the noise of direct depth mapping

Inventive Principle:
Principle #26Copying

Solution Approach 2:

The pseudo-lidar representations serve as an intermediary layer between raw camera images and final depth predictions. This intermediary processing through neural networks filters out noise from dense depth maps while preserving and enhancing genuine three-dimensional information, resolving the contradiction between information availability and accuracy

Inventive Principle:
Principle #24Intermediary (Mediator)

3Measurement precision

If pseudo-lidar features and image patch features are combined, then prediction accuracy improves, but the computational complexity increases

Engineering Contradiction:
Improvelocation prediction accuracyVSAvoidsystem complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The patent merges pseudo-lidar features derived from initial depth estimates with image patch features extracted directly from camera images. This combination of two different feature types provides complementary information that significantly improves location prediction accuracy. The neural network integrates these merged features to produce robust three-dimensional location predictions that leverage both geometric and visual information

Inventive Principle:
Principle #5Merging (Combining)

Solution Approach 2:

The neural network performs multiple functions simultaneously: it processes initial depth estimates to generate pseudo-lidar features, extracts image patch features from camera images, combines these features, and produces final location predictions. This multi-functionality achieves high prediction accuracy while consolidating complexity into a single integrated system rather than separate processing stages

Inventive Principle:
Principle #6Universality (Multi-functionality)

Data Source

PatentUS20250349024A1Three-dimensional location prediction from images
Publication Date: 2025.11.13 WAYMO LLC
  • US20250349024A1 patent drawing
  • US20250349024A1 patent drawing
  • US20250349024A1 patent drawing

AI summary

Methods, systems, and apparatus, including computer programs encoded on computer storage media, for predicting three-dimensional object locations from images. One of the methods includes obtaining a sequence of images that comprises, at each of a plurality of time steps, a respective image that was captured by a camera at the time step; generating, for each image in the sequence, respective pseudo-lidar features of a respective pseudo-lidar representation of a region in the image that has been determined to depict a first object; generating, for a particular image at a particular time step in the sequence, image patch features of the region in the particular image that has been determined to depict the first object; and generating, from the respective pseudo-lidar features and the image patch features, a prediction that characterizes a location of the first object in a three-dimensional coordinate system at the particular time step in the sequence.