Text-to-Image Model Fine-Tuning With Performance-Aware Captions
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing text-to-image models prioritize visual quality over performance metrics, making them ineffective in generating images that achieve measurable goals such as high click-through rates or conversion rates in digital advertising.
Innovation Solution
A system that trains or finetunes a text-to-image generative AI model using high-performing images and associated captions, incorporating an image-to-text model to generate descriptive text, and optionally includes performance labels to enhance the model's understanding of both aesthetic and performance aspects.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Manufacturing precision
If existing text-to-image models prioritize visual quality, then aesthetic appeal is improved, but performance metrics (click-through rates, conversion rates) deteriorate
Solution Approach 1:
The patent changes the training parameters and objectives of the text-to-image model by incorporating performance metrics (click-through rates, conversion rates) alongside visual quality metrics. This allows the model to learn to generate images that satisfy both aesthetic requirements and performance goals simultaneously, resolving the contradiction between visual quality and performance metrics.
Solution Approach 2:
The patent introduces an intermediary component that bridges visual quality and performance metrics. This intermediary mechanism enables the model to understand the relationship between aesthetic attributes and performance outcomes, allowing it to generate images that achieve both visual appeal and desired performance metrics rather than treating them as opposing objectives.
2Quantity of substance
If only images with existing captions are used for training, then training data availability is limited, but using image-to-text models to generate captions expands the training pool
Solution Approach 1:
The patent implements self-service by using an image-to-text model to automatically generate captions for images in the training pool. This eliminates the need for manual captioning or pre-existing captions, allowing the system to independently create its own training data annotations and significantly expand the available training data without external assistance.
Solution Approach 2:
The patent applies preliminary action by pre-generating captions for all training images using the image-to-text model before the actual model training begins. This preliminary caption generation prepares the training data in advance, ensuring that both visual and textual data are ready for efficient model training without requiring real-time caption generation during the training process.
Data Source
AI summary
A method of generating high-performance images includes generating, by one or more processors, a first plurality of captions each corresponding to a different one of a first plurality of images. Generating the first plurality of captions includes inputting the first plurality of images into a first generative artificial intelligence (AI) model. The method also includes training or finetuning, by the one or more processors, a second generative AI model using the first plurality of images and the first plurality of captions, and generating, by the one or more processors, a second plurality of images. Generating the second plurality of images includes inputting a plurality of text prompts into the trained or finetuned second generative AI model.


