Unified Multi-Modal Index for Efficient Retrieval
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing multi-modal retrieval methods face high retrieval delays due to the need to access multiple indexes for different modalities, leading to increased storage overheads and retrieval performance issues.
Innovation Solution
A method that constructs a single index for multiple modalities by encoding feature vectors into codes, allowing for multi-modal retrieval by accessing only one index, which reduces storage overheads and retrieval delay by calculating correlations between retrieval objects and object information using encoded data.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If multiple indexes are constructed for different modalities, then retrieval accuracy is improved, but retrieval delay increases
Solution Approach 1:
The patent merges multiple modality-specific indexes (image index, text index, audio index) into a single unified multi-modal index. This unified index stores feature vectors from different modalities together with their modality identifiers, allowing the system to perform retrieval by accessing only one index structure rather than multiple separate indexes, thereby reducing retrieval delay while maintaining the ability to accurately retrieve across different modalities
Solution Approach 2:
The unified multi-modal index serves multiple functions that were previously handled by separate indexes. It can store and retrieve feature vectors from any modality (image, text, audio) within a single index structure, making the index system universal and multi-functional. This eliminates the need to access multiple specialized indexes for different retrieval scenarios
2Measurement precision
If multiple indexes are constructed for different modalities, then modality-specific retrieval is improved, but device complexity increases
Solution Approach 1:
The patent combines multiple separate index structures into a single unified index. Instead of maintaining distinct image indexes, text indexes, and audio indexes, the system uses one unified multi-modal index that accommodates all modalities. This reduces the overall complexity of the index management system while preserving modality-specific retrieval capabilities through the use of modality identifiers and appropriate distance metric selection
3Adaptability or versatility
If multiple indexes are accessed for multi-modal retrieval, then comprehensive search coverage is improved, but retrieval efficiency decreases
Solution Approach 1:
The patent merges multiple index access operations into a single index access operation. The unified multi-modal index allows the system to search across all modalities simultaneously by accessing one index structure, rather than sequentially accessing multiple separate indexes. This dramatically improves retrieval efficiency while maintaining comprehensive search coverage across all modalities
Solution Approach 2:
The system performs preliminary organization of feature vectors from different modalities into a unified index structure during the indexing phase. By pre-organizing all modality data in a single accessible structure with proper indexing and metadata, the system eliminates the need for multiple sequential access operations during retrieval, thereby improving efficiency without sacrificing search comprehensiveness
Data Source
AI summary
A retrieval method includes obtaining first data corresponding to a retrieval object, wherein the first data indicates M feature vectors of the retrieval object, wherein each feature vector of the retrieval object corresponds to one modality of the retrieval object, and wherein M is an integer greater than 1, and obtaining a correlation between a plurality of groups of object information and the retrieval information to output at least one group of retrieved object information, wherein each group of object information corresponds to M feature vectors in an index, wherein the M feature vectors of each group of object information are indicated by one group of second data, and wherein each feature vector of the object information corresponds to one modality of the object information.


