Visual Citations for Verifiable Multimodal Query Answers
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Users struggle to formulate accurate text-based queries, especially when describing unfamiliar objects or expressing intent, leading to inefficiencies in search services that do not provide direct answers and lack verification of answer accuracy.
Innovation Solution
A visual search system processes multimodal queries, retrieving visually similar images and deriving textual content using machine-learned models to generate visual citations, allowing users to verify the accuracy of answers through associated source document information.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If text-based search services are used, then users can perform searches, but users struggle to formulate accurate queries and cannot verify answer accuracy
Solution Approach 1:
The patent introduces an image-based intermediary system that mediates between the user and the search service. Instead of directly formulating text queries, users submit images as queries, and the system retrieves and displays relevant images with source attributions. This intermediary image-based approach eliminates the difficulty of text formulation while maintaining search functionality and enabling verification through visual citations.
2Reliability
If traditional search services provide textual results, then information can be retrieved, but users cannot quickly verify the accuracy of answers
Solution Approach 1:
The patent uses visual copying by displaying thumbnail images that are direct copies or representations of the source images. These visual citations serve as immediate verification evidence, allowing users to quickly confirm answer accuracy without spending time searching for or analyzing textual source documents. The visual copy provides instant verification.
3Measurement precision
If multimodal queries with images are used, then query accuracy improves, but the system complexity increases
Solution Approach 1:
The system employs machine learning models that automatically process image queries, retrieve relevant results, extract source information, and generate visual citations without human intervention. This self-service automation handles the complexity internally, allowing the system to accept and process multimodal queries effectively while managing system complexity through automated processing pipelines.
Data Source
AI summary
A result image is retrieved based on a similarity between a query image and the result image. A first unit of text is obtained, wherein the first unit of text comprises at least a portion of textual content of a source document that includes the result image. A second unit of text is determined responsive to a prompt associated with the query image, wherein the second unit of text comprises one or more of (a) at least some of the first unit of text, or (b) text derived from the first unit of text. The second unit of text and the result image are provided for display within an interface.


