Discriminative Caption Generation via Retrieval Feedback

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Conventional digital captioning techniques generate generic descriptions that fail to distinguish between similar digital images, making them ineffective in replacing human-generated captions.

Innovation Solution

A discriminative caption generation system using a caption generation machine learning system and a retrieval machine learning system, which generates captions by minimizing a discriminability loss to ensure the captions can accurately and concisely differentiate between digital images.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If conventional caption generation techniques are used, then the caption generation process is simple and fast, but the captions are generic and fail to distinguish between similar digital images

Engineering Contradiction:
Improvecaption discriminabilityVSAvoidsystem complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The patent implements a feedback mechanism where a retrieval system evaluates the generated captions and provides discriminability loss signals back to the caption generation system. This feedback loop enables the generation system to iteratively improve caption quality by learning from retrieval performance metrics, resolving the contradiction between simple generation and discriminative captions.

Inventive Principle:
Principle #23Feedback

Solution Approach 2:

The patent introduces a retrieval system as an intermediary component between the caption generation system and the final caption output. This intermediary evaluates caption quality through discriminability loss calculation and feeds this information back to improve generation, allowing the system to achieve high discriminability without requiring the generation system itself to be overly complex.

Inventive Principle:
Principle #24Intermediary (Mediator)

2Reliability

If conventional caption generation techniques are used, then the system is easy to implement, but the captions are too broad and ineffective to replace human-generated captions

Engineering Contradiction:
Improvecaption effectivenessVSAvoidsystem architecture
Core Design Contradiction:
ReliabilityVSDevice complexity

Solution Approach 1:

The feedback mechanism from the retrieval system provides reliability signals to the generation system through discriminability loss. This enables the system to produce reliable, human-like captions by continuously improving based on retrieval performance evaluation, overcoming the limitation of conventional simple generation techniques.

Inventive Principle:
Principle #23Feedback

Solution Approach 2:

The patent changes the training parameters and loss functions of the caption generation system by incorporating discriminability loss from retrieval evaluations. This parameter change transforms the generation system from producing generic captions to generating reliable, discriminative captions that effectively replace human-generated captions.

Inventive Principle:
Principle #35Parameter changes

3Measurement precision

If captions are made more discriminative and specific, then image differentiation capability improves, but the caption generation time and computational cost increase

Engineering Contradiction:
Improveimage differentiation capabilityVSAvoidcaption generation time
Core Design Contradiction:
Measurement precisionVSLoss of time

Solution Approach 1:

The patent performs preliminary action by pre-training the retrieval system and pre-computing feature representations before final caption generation. This allows the generation system to leverage pre-computed information and retrieval models, reducing real-time computational costs while maintaining high image differentiation capability through discriminative captions.

Inventive Principle:
Principle #10Preliminary action

Data Source

PatentUS11514252B2Discriminative caption generation
Publication Date: 2022.11.29 ADOBE INC
  • US11514252B2 patent drawing
  • US11514252B2 patent drawing
  • US11514252B2 patent drawing

AI summary

A discriminative captioning system generates captions for digital images that can be used to tell two digital images apart. The discriminative captioning system includes a machine learning system that is trained by a discriminative captioning training system that includes a retrieval machine learning system. For training, a digital image is input to the caption generation machine learning system, which generates a caption for the digital image. The digital image and the generated caption, as well as a set of additional images, are input to the retrieval machine learning system. The retrieval machine learning system generates a discriminability loss that indicates how well the retrieval machine learning system is able to use the caption to discriminate between the digital image and each image in the set of additional digital images. This discriminability loss is used to train the caption generation machine learning system.