Deep Multimodal Similarity Model for Image Captioning

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current technologies face challenges in accurately associating text with images, requiring significant human effort and time, and often result in inaccurate captions, especially when determining semantic similarities between images and text.

Innovation Solution

The implementation of a deep multimodal similarity model (DMSM) that learns neural networks to map images and text fragments into vector representations, allowing for the generation of captions by measuring similarity and selecting the most relevant text based on a training set, thereby reducing human intervention and improving accuracy.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Extent of automation

If traditional caption generation methods are used, then human effort and time are required to associate text with images, but the process is slow and labor-intensive

Engineering Contradiction:
Improveautomation of caption generationVSAvoidtime required for caption generation
Core Design Contradiction:
Extent of automationVSLoss of time

Solution Approach 1:

The system enables automatic caption generation where the computer independently associates text with images using deep multimodal similarity models, eliminating the need for human intervention in the captioning process

Inventive Principle:
Principle #25Self-service

Solution Approach 2:

The patent replaces manual human effort with an automated neural network-based system that uses deep multimodal similarity modeling to generate captions, substituting mechanical human labor with computational processes

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

2Measurement precision

If traditional caption generation methods are used, then captions can be generated, but the accuracy of caption association is low

Engineering Contradiction:
Improveaccuracy of caption associationVSAvoidcomplexity of similarity measurement system
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The system transforms the caption association problem into a vector space representation where images and text are mapped to comparable vectors, enabling precise similarity measurement through mathematical operations on these parameterized representations

Inventive Principle:
Principle #35Parameter changes

Solution Approach 2:

The patent introduces vector representations as an intermediary layer between images and text, allowing the system to measure similarity through a standardized mathematical space rather than directly comparing disparate modalities

Inventive Principle:
Principle #24Intermediary (Mediator)

3Measurement precision

If deep multimodal similarity model is implemented, then accuracy of automatic caption generation is improved, but the computational complexity increases

Engineering Contradiction:
Improveaccuracy of semantic similarity measurementVSAvoidcomplexity of neural network model
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The system divides the complex task of caption generation into distinct modules: image encoding to generate image vectors, text encoding to generate text vectors, similarity computation between vectors, and caption selection based on similarity scores

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent uses vector representations as an intermediary that simplifies the comparison between images and text, transforming a complex multimodal matching problem into a simpler vector similarity computation problem

Inventive Principle:
Principle #24Intermediary (Mediator)

Data Source

PatentUS9836671B2Discovery of semantic similarities between images and text
Publication Date: 2017.12.05 MICROSOFT TECHNOLOGY LICENSING LLC
  • US9836671B2 patent drawing
  • US9836671B2 patent drawing
  • US9836671B2 patent drawing

AI summary

Disclosed herein are technologies directed to discovering semantic similarities between images and text, which can include performing image search using a textual query, performing text search using an image as a query, and/or generating captions for images using a caption generator. A semantic similarity framework can include a caption generator and can be based on a deep multimodal similar model. The deep multimodal similarity model can receive sentences and determine the relevancy of the sentences based on similarity of text vectors generated for one or more sentences to an image vector generated for an image. The text vectors and the image vector can be mapped in a semantic space, and their relevance can be determined based at least in part on the mapping. The sentence associated with the text vector determined to be the most relevant can be output as a caption for the image.