Visual Citations for Verifiable Multimodal Query Answers

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Users struggle to formulate accurate text-based queries, especially when describing unfamiliar objects or expressing intent, leading to inefficiencies in search services that do not provide direct answers and lack verification of answer accuracy.

Innovation Solution

A visual search system processes multimodal queries, retrieving visually similar images and deriving textual content using machine-learned models to generate visual citations, allowing users to verify the accuracy of answers through associated source document information.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Reliability

If text-based search services are used, then users can perform searches, but users struggle to formulate accurate queries and cannot verify answer accuracy

Engineering Contradiction:
Improveanswer accuracy verificationVSAvoidquery formulation difficulty
Core Design Contradiction:
ReliabilityVSEase of operation

Solution Approach 1:

The patent introduces an image-based intermediary system that mediates between the user and the search service. Instead of directly formulating text queries, users submit images as queries, and the system retrieves and displays relevant images with source attributions. This intermediary image-based approach eliminates the difficulty of text formulation while maintaining search functionality and enabling verification through visual citations.

Inventive Principle:
Principle #24Intermediary (Mediator)

2Reliability

If traditional search services provide textual results, then information can be retrieved, but users cannot quickly verify the accuracy of answers

Engineering Contradiction:
Improveanswer accuracy verificationVSAvoidverification time
Core Design Contradiction:
ReliabilityVSLoss of time

Solution Approach 1:

The patent uses visual copying by displaying thumbnail images that are direct copies or representations of the source images. These visual citations serve as immediate verification evidence, allowing users to quickly confirm answer accuracy without spending time searching for or analyzing textual source documents. The visual copy provides instant verification.

Inventive Principle:
Principle #26Copying

3Measurement precision

If multimodal queries with images are used, then query accuracy improves, but the system complexity increases

Engineering Contradiction:
Improvequery accuracyVSAvoidsystem complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The system employs machine learning models that automatically process image queries, retrieve relevant results, extract source information, and generate visual citations without human intervention. This self-service automation handles the complexity internally, allowing the system to accept and process multimodal queries effectively while managing system complexity through automated processing pipelines.

Inventive Principle:
Principle #25Self-service

Data Source

PatentUS20260111481A1Visual Citations for Information Provided in Response to Multimodal Queries
Publication Date: 2026.04.23 GOOGLE LLC
  • US20260111481A1 patent drawing
  • US20260111481A1 patent drawing
  • US20260111481A1 patent drawing

AI summary

A result image is retrieved based on a similarity between a query image and the result image. A first unit of text is obtained, wherein the first unit of text comprises at least a portion of textual content of a source document that includes the result image. A second unit of text is determined responsive to a prompt associated with the query image, wherein the second unit of text comprises one or more of (a) at least some of the first unit of text, or (b) text derived from the first unit of text. The second unit of text and the result image are provided for display within an interface.