Multimodal Search Refinement via Combined Feature Vectors
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Users face challenges in accurately describing their search queries, leading to unsatisfactory results when using traditional search engines, especially when trying to refine initial queries.
Innovation Solution
The implementation of a multimodal search system that allows users to input queries using different modalities (e.g., images, text, audio) and processes these inputs through a trained machine learning system to generate combined feature vectors, which are then used to refine search results.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If users provide detailed verbal descriptions or explanations of their search queries, then search accuracy improves, but the complexity of the operation increases and users find it challenging to describe what they are searching for
Solution Approach 1:
The patent replaces the mechanical/linguistic system of verbal query description with a visual system using image uploads. Instead of requiring users to articulate their search intent through language, the system accepts images as input, allowing users to visually point to or show what they are searching for. This substitution resolves the contradiction by making the operation easier (simple image upload vs. complex verbal description) while maintaining or improving search accuracy through visual content analysis.
Solution Approach 2:
The patent introduces an image as an intermediary between the user's search intent and the search engine. Rather than directly translating verbal descriptions into search queries, the system uses images as a mediating representation that captures the user's intent more naturally. This intermediary approach allows users to bypass the difficulty of verbal description while enabling the system to extract meaningful search parameters from the visual content.
2Loss of information
If users provide images to guide the search engine, then visual information is provided, but there are too many visual structures within the images to accurately extract sufficient information to improve search results
Solution Approach 1:
The patent applies extraction by selectively identifying and isolating specific visual elements within images that are relevant to search queries. Rather than attempting to process all visual structures in an image, the system extracts key features, objects, or regions that directly relate to the user's search intent. This selective extraction resolves the contradiction by filtering out unnecessary visual complexity while retaining the essential information needed for accurate search results.
Solution Approach 2:
The patent applies local quality by treating different regions or elements within an image differently based on their relevance to the search task. Instead of uniformly processing the entire image, the system identifies specific local areas or features that contain the most valuable search information and focuses its analysis there. This approach allows the system to handle complex images effectively by concentrating computational resources on the most informative local regions.
Data Source
AI summary
Combined feature vectors may be generated to map features of two or more search queries to a common embedding space. A user may provide an initial input query and then provide a refinement query. Independent feature vectors may be generated for each of the initial input query and the refinement query, may be weighted, and then may be combined to form a combined feature vector. The combined feature vector aligns different search modalities within the common embedding space that may be executed against an index.


