Text-Image Joint Embeddings for Autonomous Vehicle Data Labeling
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Autonomous vehicles face challenges in efficiently and accurately labeling large volumes of onboard data, such as camera and sensor images, due to the complexity and size of the data sets, which hinders the improvement of autonomy models and feature development.
Innovation Solution
A computer-implemented system using text-image joint embeddings to automatically label images by generating an embedding dataset, determining a high probability field for query inputs, and matching input images within this field, allowing for efficient and accurate identification and labeling of relevant data.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If manual review and processing techniques are used to label onboard data, then labeling accuracy can be maintained, but the time and resources required for labeling increase significantly
Solution Approach 1:
The patent introduces text-image joint embeddings as an intermediary mechanism that bridges manual labeling and automated processing. The embeddings create a shared representation space where image data and text queries can be efficiently matched, enabling automated systems to achieve labeling accuracy previously only attainable through manual review while dramatically reducing time requirements.
Solution Approach 2:
The patent replaces the mechanical manual review process with an automated embedding-based search system. By substituting human reviewers with computational models that operate in embedding space, the system maintains high labeling accuracy while eliminating the time and resource constraints inherent in manual processing.
2Reliability
If extensive manual review is performed on large pools of onboard data, then relevant data can be identified, but the process becomes difficult and intensive to scale
Solution Approach 1:
The patent transforms the data identification problem from pixel-space or feature-space operations to embedding-space operations. By projecting images and queries into a shared embedding dimension, the system enables efficient similarity search and filtering that scales well with data volume while maintaining identification reliability, avoiding the combinatorial complexity of traditional manual review methods.
3Productivity
If traditional search methods are used on unlabeled image data, then basic filtering can be performed, but accurate identification of relevant samples is difficult
Solution Approach 1:
The patent creates a universal embedding space that handles multiple functions simultaneously: it enables both efficient search operations and accurate relevance identification. The joint text-image embeddings serve as a multi-functional tool that combines the benefits of fast computational search with the precision of semantic understanding, eliminating the trade-off between search efficiency and identification accuracy.
Data Source
AI summary
Provided are systems, methods, and computer program products for streaming data mining with text-image joint embeddings by obtaining a roadway dataset, having a condition that each record in the roadway dataset includes an image logged by an autonomous vehicle (AV) in a roadway, generating based on a text-image model, an embedding dataset comprising an image embedding for each record in the roadway dataset, each image embedding identifying a point in a high dimension space, determining, by the one or more processors, a parametric description of a high probability field which surrounds a query input text in the high dimension space, determining, by the one or more processors, at least one input image of an input stream of images that is in the high probability field, and automatically labeling, by the one or more processors, the at least one input image based on the query input text.


