Multimodal Query Prediction via Image Object Classification

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Users face difficulties in searching for information due to challenges in describing objects in text queries, limited availability of content, and the inability to express imagined concepts clearly.

Innovation Solution

A computing system that processes image data from a live camera feed using an object classification model to generate multimodal query suggestions, which include suggested text strings to be provided with the image data to a search engine, enabling real-time image processing and text query suggestions.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If text searching is used to find information about objects, then the search can be performed using text queries, but users struggle to determine which words to use and the words may not be descriptive enough to generate desired results

Engineering Contradiction:
Improvedescriptiveness of search queryVSAvoiddifficulty of formulating text query
Core Design Contradiction:
Measurement precisionVSEase of operation

Solution Approach 1:

The patent introduces an image processing intermediary that automatically generates text queries from captured images. The system processes the image through object recognition and natural language generation to create descriptive search queries, eliminating the need for users to manually formulate text descriptions while maintaining high descriptiveness.

Inventive Principle:
Principle #24Intermediary (Mediator)

Solution Approach 2:

The patent replaces the manual text formulation process with an automated computer vision and natural language processing system. Instead of users mechanically selecting and combining words, the system automatically generates optimized text queries from image data, substituting human cognitive effort with computational processes.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

2Adaptability or versatility

If users search for content based on imagined concepts, then they can look for things they have in mind, but they lack a clear way to express the imagined concept

Engineering Contradiction:
Improveability to search for imagined conceptsVSAvoidclarity of concept expression
Core Design Contradiction:
Adaptability or versatilityVSLoss of information

Solution Approach 1:

The patent uses image capture as a direct copy of the user's visual perception or imagined concept. By allowing users to photograph or capture an image representing their imagined concept, the system preserves the complete visual information without requiring translation into potentially lossy text descriptions, thereby maintaining concept clarity while enabling versatile searching.

Inventive Principle:
Principle #26Copying

3Speed

If real-time image processing is performed to generate query suggestions, then the system can provide immediate search assistance, but the processing requires significant computational resources

Engineering Contradiction:
Improvereal-time query suggestion generationVSAvoidcomputational energy consumption
Core Design Contradiction:
SpeedVSUse of energy by moving object

Solution Approach 1:

The patent performs preliminary image processing and object recognition in real-time as the camera captures frames, generating query suggestions before the user needs to search. By preparing potential search queries in advance based on ongoing image analysis, the system enables immediate search assistance without requiring intensive processing at the moment of user interaction.

Inventive Principle:
Principle #10Preliminary action

Data Source

PatentUS12321401B1Multimodal query prediction
Publication Date: 2025.06.03 GOOGLE LLC
  • US12321401B1 patent drawing
  • US12321401B1 patent drawing
  • US12321401B1 patent drawing

AI summary

Systems and methods for multimodal query suggestion can include obtaining image data, determining text strings associated with the image data, and providing the text strings as selectable options for performing a multimodal query with the image data. The text strings can be determined based on performing object detection and classification on the image data. The object classifications can then be leveraged for determining potential text queries a user may select for obtaining additional information about the classified object.