Multi-Modal Search Model Using Data Distillation
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
The existing technologies for data search, particularly in the context of multi-modal data management, face challenges in reducing storage costs and improving the efficiency of model reproduction and data management.
Innovation Solution
A method involving data distillation to create a distilled dataset, which is then used to train a first multi-modal search model. This model encodes search inputs into dense vectors, determines corresponding distilled data items, and further encodes original data items into second dense vectors to identify search results.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Loss of energy
If traditional deduplication technology is used for data management, then storage costs are reduced slightly, but storage costs remain high and complexity issues are not significantly alleviated
Solution Approach 1:
The patent extracts and separates duplicate data identification from the main data storage system by using a dedicated distillation dataset that contains only representative samples. This extraction allows the system to identify and manage duplicates more efficiently without burdening the entire storage infrastructure, directly reducing storage costs while simplifying management complexity.
Solution Approach 2:
The patent changes the parameter of data representation by transforming raw data into a distillation dataset with reduced redundancy. This parameter change enables more efficient storage and management by representing large volumes of data through compact, representative samples, thereby reducing both storage costs and management complexity simultaneously.
2Productivity
If multi-modal deep learning neural network is trained for searching images through text, then search capability is improved, but training and maintenance costs increase significantly
Solution Approach 1:
The patent creates a distilled dataset that copies only the essential representative information from the original large-scale dataset. This copying approach enables the training of multi-modal search models with significantly reduced data volumes, maintaining search capability while dramatically reducing training and maintenance costs associated with handling full-scale datasets.
Solution Approach 2:
The patent applies partial action by using a distilled subset of data rather than the complete dataset for training purposes. This partial approach provides sufficient search capability for practical applications while avoiding the excessive costs of training and maintaining models on full-scale datasets, achieving an optimal balance between performance and cost.
3Measurement precision
If large volumes of original data are stored and processed, then search accuracy is maintained, but storage costs increase and processing efficiency decreases
Solution Approach 1:
The patent extracts the essential representative samples from the original large-volume dataset to create a distillation dataset. This extraction maintains search accuracy by preserving the most informative data points while removing redundant information, thereby reducing storage costs without sacrificing search precision.
Solution Approach 2:
The patent changes the data volume parameter by transforming the original large dataset into a compact distillation dataset. This parameter change reduces storage costs significantly while maintaining search accuracy through the strategic selection of representative samples that preserve the essential information needed for accurate searching.
Data Source
AI summary
The present disclosure relates to a method, a device, and a product for searching data. The method includes: encoding a search input into a first dense vector based on a first multi-modal search model; determining, based on the first dense vector, a distilled data item corresponding to the search input from a distilled dataset corresponding to the first multi-modal search model; encoding, based on the first multi-modal search model, an original data item in an original data subset corresponding to the distilled data item in an original dataset corresponding to the distilled dataset into a second dense vector; and determining, based on the second dense vector, an original data item from the original data subset as a search result corresponding to the search input. The method for searching data according to the present disclosure can improve the efficiency and security of data storage, model reproduction, and multi-modal data management.


