Unified Video and Behavior Data Retrieval via Text Queries
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing technologies fail to effectively retrieve video and vehicle behavior data pairs corresponding to a driving scene described in search text, as they either require inputting driving behavior data as a query or cannot retrieve vehicle behavior data simultaneously with video data.
Innovation Solution
A retrieval device and training device that utilize pre-trained text and video feature extraction models to compute text distances between search text and associated video and vehicle behavior data, outputting pairs of video and vehicle behavior data based on minimized text distances, enabling the retrieval of appropriate data pairs described by search text.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If driving behavior data is used as a query to retrieve similar driving behavior data, then the retrieval of vehicle behavior data is achieved, but video data corresponding to the driving scene cannot be retrieved simultaneously
Solution Approach 1:
The patent combines video data and vehicle behavior data into a unified retrieval system. The video encoder and behavior encoder both output features that are merged into a common embedding space, allowing simultaneous retrieval of both data types through a single query mechanism. This resolves the contradiction by enabling both retrieval accuracy for behavior data and versatility to retrieve video data together.
Solution Approach 2:
The retrieval system is designed to handle multiple query types universally - it can accept both driving behavior data queries and video queries, and return corresponding behavior data and video data pairs. The dual-encoder architecture with unified embedding space provides multi-functionality, allowing the system to adapt to different query types while maintaining retrieval accuracy.
2Ease of operation
If text queries are used to retrieve video data using general video retrieval technologies, then video retrieval is achieved, but vehicle behavior data corresponding to the driving scene cannot be retrieved
Solution Approach 1:
The patent merges the text-query interface with the behavior data retrieval mechanism. The text encoder processes search queries and maps them to the same embedding space as behavior data features, enabling the system to retrieve both video and behavior data simultaneously through text queries. This maintains ease of operation while achieving precise behavior data retrieval.
Solution Approach 2:
The unified embedding space acts as an intermediary that connects text queries, video features, and behavior data features. This mediator enables seamless interaction between different data types and query methods, allowing text-based queries to effectively retrieve both video and behavior data with high precision without requiring separate retrieval systems.
3Adaptability or versatility
If separate retrieval systems are used for video data and vehicle behavior data, then each data type can be retrieved independently, but the system complexity increases and coordinated retrieval becomes difficult
Solution Approach 1:
The patent combines separate video retrieval and behavior data retrieval systems into a unified dual-encoder architecture. Both video encoder and behavior encoder share a common embedding space and can be trained jointly, reducing system complexity while maintaining the ability to retrieve each data type independently or together through a single coordinated mechanism.
Data Source
AI summary
The retrieval device extracts a feature corresponding to search text by inputting the search text into a pre-trained text feature extraction model. The retrieval device then, for plural combinations stored in a database associating a text description including plural sentences, with a vehicle-view video, and with vehicle behavior data representing temporal vehicle behavior, computes a text distance represented by a difference between a feature extracted from each sentence of the text description associated with the video and vehicle behavior data, and the feature corresponding to the search text. The retrieval device outputs as the search result a prescribed number of pairs of video and vehicle behavior data pairs in sequence from the smallest text distance according to the text distances.


