Visual Citations for Multimodal Query Refinement
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Users face difficulties in formulating text-based queries, especially when describing unfamiliar objects or expressing intent accurately, leading to inefficient search results, as most search services lack the ability to verify the accuracy of provided answers.
Innovation Solution
A computer-implemented method that processes multimodal queries using a machine-learned language model to generate visual citations, allowing users to refine search results by visually identifying the sources of information, enabling quick verification of accuracy through visual similarity and user feedback.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Ease of operation
If text-based search services are used, then search functionality is provided, but users struggle to formulate queries accurately leading to inefficient search results
Solution Approach 1:
The patent introduces visual citations as an intermediary element between the query image and the search results. These visual citations display source images alongside textual answers, allowing users to verify accuracy visually without requiring them to formulate precise text queries. The visual citation acts as a mediator that bridges the gap between simple image input and accurate information retrieval.
Solution Approach 2:
The patent replaces the mechanical process of text query formulation with a visual verification mechanism. Instead of requiring users to carefully craft text queries to achieve accurate results, the system provides textual answers and allows users to verify accuracy through visual inspection of citation images. This substitutes the complex text formulation task with a simpler visual verification process.
2Reliability
If search services provide textual answers, then information is retrieved, but users lack ability to verify accuracy of provided answers
Solution Approach 1:
The patent implements a feedback mechanism where users can visually inspect source images associated with each textual answer through visual citations. This allows users to provide feedback on the accuracy of the answers by comparing the textual information with the actual images. The visual citation system enables users to verify whether the textual answer accurately represents the source material, creating a feedback loop that improves information reliability.
Solution Approach 2:
Visual citations serve as an intermediary layer between the search system and the user, providing a bridge for verification. They display the source images that generated the textual answers, allowing users to verify accuracy without requiring direct access to the original source documents. This intermediary presentation enables accurate verification while simplifying the user's task.
3Reliability
If multimodal queries with visual citations are provided, then answer verification is enabled, but system complexity increases
Solution Approach 1:
The patent segments the search response into distinct components: textual answers and visual citations. Each citation is a separate unit that displays source images alongside metadata about the answer's origin. This segmentation allows the system to maintain high reliability through verification while managing complexity by organizing information into manageable, discrete elements rather than a monolithic structure.
Solution Approach 2:
The patent adds a visual dimension to the traditional text-based search response. By incorporating images of source materials alongside textual answers, the system creates a multimodal output that enables verification without fundamentally redesigning the search architecture. This dimensional addition (visual vs. textual) provides verification capability while building upon existing text-processing infrastructure.
Data Source
AI summary
Result images are retrieved based on a similarity to a query image. A set of textual inputs is processed with a machine-learned language model to obtain a language output comprising textual content, wherein the set of textual inputs comprises textual content from source documents that include the result images, and a prompt associated with the query image. The language output and the result images are provided to a user computing device. Information is received descriptive of an indication by a user that a first result image is visually dissimilar to the query image. Textual content associated with the source document that includes the first result image from the set of textual inputs is removed. The set of textual inputs is processed with the machine-learned language model to obtain a refined language output. The refined language output is provided to the user computing device.


