Two-Stage Image Recaptioning for More Accurate Image Generation

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing image generation models face challenges due to low-quality captions in training datasets, leading to incoherent, low-quality, or inaccurate image generation, and improving these datasets with human input is costly and time-consuming.

Innovation Solution

Implement an image captioner model trained in multiple stages to generate high-quality, detailed captions for images, enhancing the training datasets to improve the performance of image generation models.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Loss of time

If traditional automation is used to improve dataset quality, then time and cost are reduced, but computational power and memory usage increase significantly

Engineering Contradiction:
ImprovetimeVSAvoidcomputational power
Core Design Contradiction:
Loss of timeVSUse of energy by moving object

Solution Approach 1:

The patent changes the parameters of the automation system by using a two-stage captioning approach with different model configurations. The first stage uses a smaller, faster model for initial captioning, while the second stage uses a larger, more accurate model for refinement. This parameter change allows the system to achieve high dataset quality with optimized computational resource usage, balancing time efficiency and computational power requirements.

Inventive Principle:
Principle #35Parameter changes

2Measurement precision

If human input is used to improve dataset quality, then caption accuracy improves, but cost and time consumption increase

Engineering Contradiction:
Improvecaption accuracyVSAvoidcost and time efficiency
Core Design Contradiction:
Measurement precisionVSProductivity

Solution Approach 1:

The patent implements self-service by designing an automated two-stage captioning system that improves dataset quality without human intervention. The system uses the first stage to generate initial captions and the second stage to automatically refine and enhance them, achieving high caption accuracy while maintaining cost and time efficiency. This eliminates the need for manual dataset curation while preserving data quality.

Inventive Principle:
Principle #25Self-service

3Ease of operation

If simple user prompts are used for image generation, then user convenience improves, but image quality and detail may suffer

Engineering Contradiction:
Improveuser convenienceVSAvoidimage quality
Core Design Contradiction:
Ease of operationVSManufacturing precision

Solution Approach 1:

The patent applies preliminary action by pre-processing user prompts through the two-stage captioning system before they reach the image generation model. The first stage creates an initial interpretation of the prompt, and the second stage refines it with enhanced detail and accuracy. This preliminary enhancement of the prompt ensures that even simple user inputs are transformed into high-quality, detailed image generation instructions, maintaining both user convenience and image quality.

Inventive Principle:
Principle #10Preliminary action

Data Source

PatentUS20250259423A1Model image generation using recaptioned images
Publication Date: 2025.08.14 OPENAI OPCO LLC
  • US20250259423A1 patent drawing
  • US20250259423A1 patent drawing
  • US20250259423A1 patent drawing

AI summary

Disclosed herein are methods, systems, and computer-readable media for generating image captions for training a machine learning model. Current image generation models are hindered by the prevalence of improper or inaccurate captions, which leads to suboptimal training data. This results in less effective image generation models. Disclosed systems and methods involve obtaining a text-to-image dataset including one or more digital image-caption pairs. Systems and methods involve generating a recaptioned dataset by applying an image captioner model to images in the text-to-image dataset. An image captioner model can be trained with an improved image dataset, a first tuning stage, and a second tuning stage, for improved performance.