Deep Multimodal Similarity Model for Image Captioning
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current technologies face challenges in accurately associating text with images, requiring significant human effort and time, and often result in inaccurate captions, especially when determining semantic similarities between images and text.
Innovation Solution
The implementation of a deep multimodal similarity model (DMSM) that learns neural networks to map images and text fragments into vector representations, allowing for the generation of captions by measuring similarity and selecting the most relevant text based on a training set, thereby reducing human intervention and improving accuracy.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Extent of automation
If traditional caption generation methods are used, then human effort and time are required to associate text with images, but the process is slow and labor-intensive
Solution Approach 1:
The system enables automatic caption generation where the computer independently associates text with images using deep multimodal similarity models, eliminating the need for human intervention in the captioning process
Solution Approach 2:
The patent replaces manual human effort with an automated neural network-based system that uses deep multimodal similarity modeling to generate captions, substituting mechanical human labor with computational processes
2Measurement precision
If traditional caption generation methods are used, then captions can be generated, but the accuracy of caption association is low
Solution Approach 1:
The system transforms the caption association problem into a vector space representation where images and text are mapped to comparable vectors, enabling precise similarity measurement through mathematical operations on these parameterized representations
Solution Approach 2:
The patent introduces vector representations as an intermediary layer between images and text, allowing the system to measure similarity through a standardized mathematical space rather than directly comparing disparate modalities
3Measurement precision
If deep multimodal similarity model is implemented, then accuracy of automatic caption generation is improved, but the computational complexity increases
Solution Approach 1:
The system divides the complex task of caption generation into distinct modules: image encoding to generate image vectors, text encoding to generate text vectors, similarity computation between vectors, and caption selection based on similarity scores
Solution Approach 2:
The patent uses vector representations as an intermediary that simplifies the comparison between images and text, transforming a complex multimodal matching problem into a simpler vector similarity computation problem
Data Source
AI summary
Disclosed herein are technologies directed to discovering semantic similarities between images and text, which can include performing image search using a textual query, performing text search using an image as a query, and/or generating captions for images using a caption generator. A semantic similarity framework can include a caption generator and can be based on a deep multimodal similar model. The deep multimodal similarity model can receive sentences and determine the relevancy of the sentences based on similarity of text vectors generated for one or more sentences to an image vector generated for an image. The text vectors and the image vector can be mapped in a semantic space, and their relevance can be determined based at least in part on the mapping. The sentence associated with the text vector determined to be the most relevant can be output as a caption for the image.


