Cross-lingual Image Retrieval via Intermediary Translation
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Conventional image retrieval systems are inflexible and inaccurate in retrieving digital images across languages due to their reliance on multi-lingual datasets and inability to capture the intent and context of text queries effectively, leading to the retrieval of irrelevant images.
Innovation Solution
A cross-lingual image search system that uses a zero-shot approach to train a cross-lingual image retrieval model with a multimodal metric loss function, allowing it to generate cross-lingual-multimodal embeddings for texts in multiple languages without relying on multi-lingual data, thereby accurately identifying relevant images based on query texts.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Adaptability or versatility
If conventional systems use multi-lingual datasets to train image retrieval models, then cross-lingual retrieval capability is provided, but the systems are rigid and require data in other languages to train, reducing flexibility
Solution Approach 1:
The patent introduces an intermediary translation model that translates queries from target languages to the training language, allowing the system to handle multiple languages without requiring multi-lingual training data. This intermediary translation layer enables cross-lingual retrieval while maintaining training on monolingual datasets.
Solution Approach 2:
The system changes the parameter of training language from multiple languages to a single training language, while maintaining multi-lingual operational capability through translation. This parameter change allows the model to be trained on simpler monolingual data while still achieving cross-lingual functionality.
2Adaptability or versatility
If conventional systems retrieve images based on text queries in various languages, then cross-lingual digital image retrieval is provided, but the systems inaccurately retrieve digital images that are not relevant to the search queries
Solution Approach 1:
The translation model serves as an intermediary that accurately converts queries from target languages to the training language, preserving the semantic meaning and intent of the original query. This intermediary translation ensures that the image retrieval model can accurately process multi-lingual queries while maintaining high retrieval precision.
Solution Approach 2:
The system performs preliminary translation of the query to the training language before processing it through the image retrieval model. This preliminary action ensures that the query is in the correct language format for accurate matching with image captions, improving retrieval accuracy.
3Ease of manufacture
If the system trains on monolingual datasets, then training flexibility is improved, but the ability to handle cross-lingual queries during inference is limited
Solution Approach 1:
The translation model acts as an intermediary that bridges the gap between monolingual training data and multi-lingual inference needs. By translating target language queries to the training language, the system maintains training simplicity while achieving cross-lingual inference capability.
Solution Approach 2:
The system achieves universal cross-lingual functionality through a combination of monolingual training and translation. The image retrieval model trained on one language can universally handle queries in multiple languages through the translation intermediary, providing multi-functionality without requiring multi-lingual training.
Data Source
AI summary
The present disclosure relates to methods, systems, and non-transitory computer-readable media for retrieving digital images in response to queries. For example, in one or more embodiments, the disclosed systems receive a query comprising text and generates a cross-lingual-multimodal embedding for the text within a multimodal embedding space. The disclosed systems further identifies an image embedding for a digital image that corresponds to (e.g., is relevant to) the text from the query based on an embedding distance between the image embedding and the cross-lingual-multimodal embedding for the text within the multimodal embedding space. Accordingly, the disclosed systems retrieve the digital image associated with the image embedding for display on a client device, such as the client device that submitted the query.


