Visual-Audio Search Term Generation for Easier Query Formulation

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Conventional search systems rely solely on text input, which can be challenging for users to formulate effective queries, especially when describing visual data or objects.

Innovation Solution

A multimodal search system that processes both visual data from a camera and audio data from a microphone to generate search terms, using machine-learned models to process and combine visual features and audio words for more accurate search results.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Ease of operation

If text input is used for search queries, then the search system can process queries, but users struggle to formulate effective queries and determine which words to use

Engineering Contradiction:
Improveease of formulating search queriesVSAvoiddescriptive accuracy of search terms
Core Design Contradiction:
Ease of operationVSLoss of information

Solution Approach 1:

The patent introduces an image as an intermediary object between the user and the search system. Instead of requiring users to directly formulate text queries, they capture an image of the target object. The system then automatically extracts visual features from the image and generates search terms, serving as a mediator that translates visual information into searchable text without requiring user linguistic skills.

Inventive Principle:
Principle #24Intermediary (Mediator)

Solution Approach 2:

The patent replaces the manual mechanical process of selecting and typing search words with an automated computer vision system. The system uses image processing algorithms to automatically identify objects in the captured image, extract their features, and generate appropriate search terms, eliminating the need for users to manually choose descriptive words.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

2Productivity

If users manually select search words, then they can control the query, but the process is time-consuming and friction-filled

Engineering Contradiction:
Improvespeed of search query formulationVSAvoidcomplexity of search formulation process
Core Design Contradiction:
ProductivityVSDevice complexity

Solution Approach 1:

The system performs self-service by automatically generating search terms from the captured image without requiring user intervention in the term selection process. The computer vision system independently analyzes the image, identifies objects, extracts features, and formulates search queries autonomously, freeing the user from the tedious task of manual query construction.

Inventive Principle:
Principle #25Self-service

Solution Approach 2:

The system performs preliminary action by pre-processing the image to extract visual features and generate candidate search terms before the user even submits the query. This advance preparation of search terms based on image analysis eliminates the need for users to spend time on query formulation during the search process.

Inventive Principle:
Principle #10Preliminary action

3Measurement precision

If text search alone is used, then the system is simple to operate, but it cannot effectively search for visual objects or determine where images were captured

Engineering Contradiction:
Improveaccuracy of object identificationVSAvoidcomplexity of search system
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The patent merges two previously separate search modalities into a unified system: traditional text-based search and image-based visual search. By combining computer vision technology with text search capabilities, the system can process both image data and text queries simultaneously, enabling users to search for objects by capturing their image while maintaining the power of text-based search for conceptual queries.

Inventive Principle:
Principle #5Merging (Combining)

Data Source

PatentUS20250291862A1Visual and Audio Multimodal Searching System
Publication Date: 2025.09.18 GOOGLE LLC
  • US20250291862A1 patent drawing
  • US20250291862A1 patent drawing
  • US20250291862A1 patent drawing

AI summary

A multimodal search system is described. The system can receive image data captured by a camera of a user device. Additionally, the system can receive audio data associated with the image data. The audio data can be captured by a microphone of the user device. Moreover, the system can process the image data to generate visual features. Furthermore, the system can process the audio data to generate a plurality of words. The system can generate a plurality of search terms based on the plurality of words and the visual features. Subsequently, the system can determine one or more search results associated with the plurality of search terms and provide the one or more search results as an output.