Text-to-Image Model Fine-Tuning With Performance-Aware Captions

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing text-to-image models prioritize visual quality over performance metrics, making them ineffective in generating images that achieve measurable goals such as high click-through rates or conversion rates in digital advertising.

Innovation Solution

A system that trains or finetunes a text-to-image generative AI model using high-performing images and associated captions, incorporating an image-to-text model to generate descriptive text, and optionally includes performance labels to enhance the model's understanding of both aesthetic and performance aspects.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Manufacturing precision

If existing text-to-image models prioritize visual quality, then aesthetic appeal is improved, but performance metrics (click-through rates, conversion rates) deteriorate

Engineering Contradiction:
Improvevisual qualityVSAvoidperformance metrics
Core Design Contradiction:
Manufacturing precisionVSReliability

Solution Approach 1:

The patent changes the training parameters and objectives of the text-to-image model by incorporating performance metrics (click-through rates, conversion rates) alongside visual quality metrics. This allows the model to learn to generate images that satisfy both aesthetic requirements and performance goals simultaneously, resolving the contradiction between visual quality and performance metrics.

Inventive Principle:
Principle #35Parameter changes

Solution Approach 2:

The patent introduces an intermediary component that bridges visual quality and performance metrics. This intermediary mechanism enables the model to understand the relationship between aesthetic attributes and performance outcomes, allowing it to generate images that achieve both visual appeal and desired performance metrics rather than treating them as opposing objectives.

Inventive Principle:
Principle #24Intermediary (Mediator)

2Quantity of substance

If only images with existing captions are used for training, then training data availability is limited, but using image-to-text models to generate captions expands the training pool

Engineering Contradiction:
Improvetraining data availabilityVSAvoidcaption generation requirement
Core Design Contradiction:
Quantity of substanceVSDevice complexity

Solution Approach 1:

The patent implements self-service by using an image-to-text model to automatically generate captions for images in the training pool. This eliminates the need for manual captioning or pre-existing captions, allowing the system to independently create its own training data annotations and significantly expand the available training data without external assistance.

Inventive Principle:
Principle #25Self-service

Solution Approach 2:

The patent applies preliminary action by pre-generating captions for all training images using the image-to-text model before the actual model training begins. This preliminary caption generation prepares the training data in advance, ensuring that both visual and textual data are ready for efficient model training without requiring real-time caption generation during the training process.

Inventive Principle:
Principle #10Preliminary action

Data Source

PatentUS20250336123A1Performance-aware image generation based on text
Publication Date: 2025.10.30 GOOGLE LLC
  • US20250336123A1 patent drawing
  • US20250336123A1 patent drawing
  • US20250336123A1 patent drawing

AI summary

A method of generating high-performance images includes generating, by one or more processors, a first plurality of captions each corresponding to a different one of a first plurality of images. Generating the first plurality of captions includes inputting the first plurality of images into a first generative artificial intelligence (AI) model. The method also includes training or finetuning, by the one or more processors, a second generative AI model using the first plurality of images and the first plurality of captions, and generating, by the one or more processors, a second plurality of images. Generating the second plurality of images includes inputting a plurality of text prompts into the trained or finetuned second generative AI model.