Multimodal Search Refinement via Textual Query Appending

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Visual search applications struggle to accurately interpret user intent when provided with images, leading to incorrect determination of user intent and the need for users to re-capture query targets, resulting in a frustrating user experience and unnecessary resource usage.

Innovation Solution

A computer-implemented method and system that allows users to refine visual search queries using textual data, appending it to the visual search query to form a multimodal search query, thereby improving search accuracy and user experience.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If visual search applications only use image-based queries, then the system complexity remains low, but the search accuracy and user intent interpretation deteriorate

Engineering Contradiction:
Improvesearch accuracyVSAvoidsystem complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The patent combines visual search (image-based) and textual search capabilities into a unified multimodal search system. The system merges image queries and text queries into a single search framework, allowing users to provide both image and text inputs simultaneously. This integration resolves the contradiction by improving search accuracy through multiple input modalities while managing system complexity through a cohesive architecture that processes both types of queries together.

Inventive Principle:
Principle #5Merging (Combining)

Solution Approach 2:

The search system is designed to accept and process multiple types of queries (image-only, text-only, and combined image-text queries). This multi-functional capability allows the system to adapt to different user needs and input preferences, improving search accuracy across diverse scenarios while maintaining a single unified system rather than requiring separate systems for each query type.

Inventive Principle:
Principle #6Universality (Multi-functionality)

2Adaptability or versatility

If visual search applications do not accept additional input modes, then the device complexity remains low, but the adaptability to different user intents deteriorates

Engineering Contradiction:
Improveuser intent interpretationVSAvoidinput processing complexity
Core Design Contradiction:
Adaptability or versatilityVSDevice complexity

Solution Approach 1:

The system is designed to handle multiple query types (image-only, text-only, and combined queries) through a single unified interface. This multi-functional design enables the system to adapt to various user intents and input preferences without requiring separate systems for each query type, thereby improving adaptability while managing complexity through a cohesive architecture.

Inventive Principle:
Principle #6Universality (Multi-functionality)

Solution Approach 2:

The search system dynamically adapts to different user input modes and preferences. Users can flexibly choose to provide image queries, text queries, or both together, and the system adjusts its processing accordingly. This dynamic capability allows the system to accommodate evolving user needs and different search scenarios without requiring rigid pre-programming for each specific case.

Inventive Principle:
Principle #15Dynamics

3Productivity

If users must re-capture images to refine searches, then the search interface remains simple, but the productivity and user experience deteriorate

Engineering Contradiction:
Improvesearch refinement efficiencyVSAvoidinterface complexity
Core Design Contradiction:
ProductivityVSDevice complexity

Solution Approach 1:

The system merges image-based refinement and text-based refinement into a single unified refinement mechanism. Instead of requiring users to re-capture images to refine their searches, the system allows users to provide text inputs that are combined with the original image query. This integration improves search refinement efficiency by eliminating the need to re-capture images while managing interface complexity through a cohesive refinement process.

Inventive Principle:
Principle #5Merging (Combining)

4Measurement precision

If visual search applications use only image queries, then the resource usage remains low, but the search precision and user experience deteriorate

Engineering Contradiction:
Improveuser intent determination accuracyVSAvoidresource usage
Core Design Contradiction:
Measurement precisionVSUse of energy by moving object

Solution Approach 1:

The system combines image queries and text queries into a unified multimodal search approach. By merging these two input types, the system improves user intent determination accuracy because text provides additional contextual information that helps disambiguate image-based queries. The combined approach leverages the strengths of both modalities while managing resource usage through efficient integrated processing rather than requiring multiple separate processing systems.

Inventive Principle:
Principle #5Merging (Combining)

Data Source

PatentUS20240028638A1Systems and Methods for Efficient Multimodal Search Refinement
Publication Date: 2024.01.25 GOOGLE LLC
  • US20240028638A1 patent drawing
  • US20240028638A1 patent drawing
  • US20240028638A1 patent drawing

AI summary

Systems and methods of the present disclosure are directed to a computer-implemented method for multimodal search refinement. The method includes obtaining a visual search query from a user comprising one or more query images. The method includes providing a search interface for display to the user, the search interface comprising one or more result images responsive to the one or more query images and an interface element indicative of a request to the user to refine the visual search query. The method includes obtaining, from the user, textual data comprising a refinement to the visual search query. The method includes appending, by the computing system, the textual data to the visual search query to obtain a multimodal search query.