Visual Citations for Multimodal Query Refinement

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Users face difficulties in formulating text-based queries, especially when describing unfamiliar objects or expressing intent accurately, leading to inefficient search results, as most search services lack the ability to verify the accuracy of provided answers.

Innovation Solution

A computer-implemented method that processes multimodal queries using a machine-learned language model to generate visual citations, allowing users to refine search results by visually identifying the sources of information, enabling quick verification of accuracy through visual similarity and user feedback.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Ease of operation

If text-based search services are used, then search functionality is provided, but users struggle to formulate queries accurately leading to inefficient search results

Engineering Contradiction:
Improvequery formulation easeVSAvoidsearch accuracy
Core Design Contradiction:
Ease of operationVSMeasurement precision

Solution Approach 1:

The patent introduces visual citations as an intermediary element between the query image and the search results. These visual citations display source images alongside textual answers, allowing users to verify accuracy visually without requiring them to formulate precise text queries. The visual citation acts as a mediator that bridges the gap between simple image input and accurate information retrieval.

Inventive Principle:
Principle #24Intermediary (Mediator)

Solution Approach 2:

The patent replaces the mechanical process of text query formulation with a visual verification mechanism. Instead of requiring users to carefully craft text queries to achieve accurate results, the system provides textual answers and allows users to verify accuracy through visual inspection of citation images. This substitutes the complex text formulation task with a simpler visual verification process.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

2Reliability

If search services provide textual answers, then information is retrieved, but users lack ability to verify accuracy of provided answers

Engineering Contradiction:
Improveanswer accuracy verificationVSAvoidinformation verification capability
Core Design Contradiction:
ReliabilityVSLoss of information

Solution Approach 1:

The patent implements a feedback mechanism where users can visually inspect source images associated with each textual answer through visual citations. This allows users to provide feedback on the accuracy of the answers by comparing the textual information with the actual images. The visual citation system enables users to verify whether the textual answer accurately represents the source material, creating a feedback loop that improves information reliability.

Inventive Principle:
Principle #23Feedback

Solution Approach 2:

Visual citations serve as an intermediary layer between the search system and the user, providing a bridge for verification. They display the source images that generated the textual answers, allowing users to verify accuracy without requiring direct access to the original source documents. This intermediary presentation enables accurate verification while simplifying the user's task.

Inventive Principle:
Principle #24Intermediary (Mediator)

3Reliability

If multimodal queries with visual citations are provided, then answer verification is enabled, but system complexity increases

Engineering Contradiction:
Improveinformation reliabilityVSAvoidsystem complexity
Core Design Contradiction:
ReliabilityVSDevice complexity

Solution Approach 1:

The patent segments the search response into distinct components: textual answers and visual citations. Each citation is a separate unit that displays source images alongside metadata about the answer's origin. This segmentation allows the system to maintain high reliability through verification while managing complexity by organizing information into manageable, discrete elements rather than a monolithic structure.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent adds a visual dimension to the traditional text-based search response. By incorporating images of source materials alongside textual answers, the system creates a multimodal output that enables verification without fundamentally redesigning the search architecture. This dimensional addition (visual vs. textual) provides verification capability while building upon existing text-processing infrastructure.

Inventive Principle:
Principle #17Another dimension (Dimensionality change)

Data Source

PatentUS20240378237A1Visual Citations for Information Provided in Response to Multimodal Queries
Publication Date: 2024.11.14 GOOGLE LLC
  • US20240378237A1 patent drawing
  • US20240378237A1 patent drawing
  • US20240378237A1 patent drawing

AI summary

Result images are retrieved based on a similarity to a query image. A set of textual inputs is processed with a machine-learned language model to obtain a language output comprising textual content, wherein the set of textual inputs comprises textual content from source documents that include the result images, and a prompt associated with the query image. The language output and the result images are provided to a user computing device. Information is received descriptive of an indication by a user that a first result image is visually dissimilar to the query image. Textual content associated with the source document that includes the first result image from the set of textual inputs is removed. The set of textual inputs is processed with the machine-learned language model to obtain a refined language output. The refined language output is provided to the user computing device.