Text-Image Joint Embeddings for Autonomous Vehicle Data Labeling

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Autonomous vehicles face challenges in efficiently and accurately labeling large volumes of onboard data, such as camera and sensor images, due to the complexity and size of the data sets, which hinders the improvement of autonomy models and feature development.

Innovation Solution

A computer-implemented system using text-image joint embeddings to automatically label images by generating an embedding dataset, determining a high probability field for query inputs, and matching input images within this field, allowing for efficient and accurate identification and labeling of relevant data.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If manual review and processing techniques are used to label onboard data, then labeling accuracy can be maintained, but the time and resources required for labeling increase significantly

Engineering Contradiction:
Improvelabeling accuracyVSAvoidlabeling time
Core Design Contradiction:
Measurement precisionVSLoss of time

Solution Approach 1:

The patent introduces text-image joint embeddings as an intermediary mechanism that bridges manual labeling and automated processing. The embeddings create a shared representation space where image data and text queries can be efficiently matched, enabling automated systems to achieve labeling accuracy previously only attainable through manual review while dramatically reducing time requirements.

Inventive Principle:
Principle #24Intermediary (Mediator)

Solution Approach 2:

The patent replaces the mechanical manual review process with an automated embedding-based search system. By substituting human reviewers with computational models that operate in embedding space, the system maintains high labeling accuracy while eliminating the time and resource constraints inherent in manual processing.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

2Reliability

If extensive manual review is performed on large pools of onboard data, then relevant data can be identified, but the process becomes difficult and intensive to scale

Engineering Contradiction:
Improvedata identification accuracyVSAvoidprocessing complexity
Core Design Contradiction:
ReliabilityVSDevice complexity

Solution Approach 1:

The patent transforms the data identification problem from pixel-space or feature-space operations to embedding-space operations. By projecting images and queries into a shared embedding dimension, the system enables efficient similarity search and filtering that scales well with data volume while maintaining identification reliability, avoiding the combinatorial complexity of traditional manual review methods.

Inventive Principle:
Principle #17Another dimension (Dimensionality change)

3Productivity

If traditional search methods are used on unlabeled image data, then basic filtering can be performed, but accurate identification of relevant samples is difficult

Engineering Contradiction:
Improvesearch efficiencyVSAvoidrelevance identification accuracy
Core Design Contradiction:
ProductivityVSMeasurement precision

Solution Approach 1:

The patent creates a universal embedding space that handles multiple functions simultaneously: it enables both efficient search operations and accurate relevance identification. The joint text-image embeddings serve as a multi-functional tool that combines the benefits of fast computational search with the precision of semantic understanding, eliminating the trade-off between search efficiency and identification accuracy.

Inventive Principle:
Principle #6Universality (Multi-functionality)

Data Source

PatentUS20240257536A1System, Method, and Computer Program Product for Streaming Data Mining with Text-Image Joint Embeddings
Publication Date: 2024.08.01 VOLKSWAGEN GROUP OF AMERICA INVESTMENTS LLC
  • US20240257536A1 patent drawing
  • US20240257536A1 patent drawing
  • US20240257536A1 patent drawing

AI summary

Provided are systems, methods, and computer program products for streaming data mining with text-image joint embeddings by obtaining a roadway dataset, having a condition that each record in the roadway dataset includes an image logged by an autonomous vehicle (AV) in a roadway, generating based on a text-image model, an embedding dataset comprising an image embedding for each record in the roadway dataset, each image embedding identifying a point in a high dimension space, determining, by the one or more processors, a parametric description of a high probability field which surrounds a query input text in the high dimension space, determining, by the one or more processors, at least one input image of an input stream of images that is in the high probability field, and automatically labeling, by the one or more processors, the at least one input image based on the query input text.