Multimodal Product Query Ranking for Accurate Object Retrieval
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing query services often fail to meet user expectations due to the limitations of image-based or text-based methods, resulting in less comprehensive query outcomes and inaccurate results, especially in scenarios with diverse user needs.
Innovation Solution
An information processing method that integrates image-text query information, determines the information attribute type, constructs image-text fusion information, and performs object retrieval for each type of information to rank and determine a target object based on relevance, ensuring accurate alignment with user query needs.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If image-based or text-based query methods are used separately, then the query process is simple, but the query results are less comprehensive and less accurate
Solution Approach 1:
The patent combines image-based query and text-based query into a unified multi-modal query system. The system simultaneously processes both image and text inputs, fuses their respective features, and performs joint retrieval to produce comprehensive query results that leverage the strengths of both modalities, thereby improving accuracy without requiring entirely separate systems
Solution Approach 2:
The patent creates a universal query interface that accepts both image and text inputs and can handle various query scenarios through a single multi-functional system. The system dynamically adapts to different input types and combines them appropriately, eliminating the need for users to choose between separate image-based or text-based query systems
2Adaptability or versatility
If single modality query (image or text) is performed, then the system complexity is low, but the information coverage is narrow
Solution Approach 1:
The patent transitions from single-modality querying to multi-modal querying by adding another dimension of information processing. Instead of solely relying on text or image inputs, the system processes both modalities simultaneously, extracting features from each and fusing them to create a more comprehensive representation of user intent, thereby expanding information coverage
3Measurement precision
If image-text fusion information is constructed and multiple retrieval operations are performed, then query accuracy is improved, but processing time increases
Solution Approach 1:
The patent performs preliminary feature extraction from both image and text inputs before the actual retrieval process. By pre-processing and fusing features in advance, the system prepares optimized query representations that can be efficiently matched against the database, reducing the computational burden during the retrieval phase and mitigating processing time increases
Data Source
AI summary
The embodiments of the present disclosure provide an information processing method and apparatus, as well as a product query method and apparatus. The information processing method includes: obtaining image-text query information comprising image query information and text query information, and determining an information attribute type corresponding to the image-text query information; identifying the image query information and the text query information within the image-text query information, and constructing image-text fusion information based on the image query information and the text query information; performing object retrieval for the image query information, the text query information, and the image-text fusion information respectively to obtain an image-retrieved object, a text-retrieved object, and an image-text retrieved object; ranking the image-retrieved object, the text-retrieved object, and the image-text retrieved object according to the information attribute type, and determining a target object corresponding to the image-text query information based on a ranking result.


