Iterative Text-to-Image Prompt Refinement Through Image Feedback

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing image generation systems from text prompts often require extensive user expertise and computational resources, as they lack efficient mechanisms for iteratively improving text prompts based on generated images, leading to suboptimal results and unnecessary resource consumption.

Innovation Solution

A system that iteratively generates images from text by receiving user input, analyzing generated images to provide additional descriptions and suggestions for improving the text prompts, allowing users to refine their inputs intelligently and efficiently.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Extent of automation

If automated text-to-image generation is used, then image generation capability is improved, but user expertise requirement increases

Engineering Contradiction:
Improveautomated text-to-image generationVSAvoiduser expertise requirement
Core Design Contradiction:
Extent of automationVSEase of operation

Solution Approach 1:

The system analyzes the generated first image and automatically generates a text description with additional words that provide feedback on what visual features were created. This feedback loop enables users to iteratively refine their text prompts based on actual image outcomes, reducing the expertise needed to achieve desired results

Inventive Principle:
Principle #23Feedback

Solution Approach 2:

The system performs self-analysis of generated images and automatically generates suggestions for prompt improvement without requiring user intervention. The automated image analysis and suggestion generation mechanisms allow the system to serve itself in improving the text-to-image generation process

Inventive Principle:
Principle #25Self-service

2Manufacturing precision

If iterative prompt refinement is implemented, then image quality is improved, but time consumption increases

Engineering Contradiction:
Improveimage qualityVSAvoidtime consumption
Core Design Contradiction:
Manufacturing precisionVSLoss of time

Solution Approach 1:

The system performs preliminary analysis of the generated image and pre-generates a list of relevant words and suggestions before the user needs to refine their prompt. This preliminary action reduces the time needed for iterative refinement by having improvement suggestions ready in advance

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The system replaces manual trial-and-error prompt refinement with an automated image analysis mechanism that uses AI to suggest improvements. This substitution of mechanical user effort with intelligent automation reduces time consumption while maintaining or improving image quality

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

3Measurement precision

If comprehensive image analysis is performed, then prompt improvement accuracy is improved, but computational resources increase

Engineering Contradiction:
Improveprompt improvement accuracyVSAvoidcomputational resources
Core Design Contradiction:
Measurement precisionVSUse of energy by moving object

Solution Approach 1:

The system extracts only the essential visual features from the generated image that are relevant to prompt improvement, rather than performing exhaustive analysis of all image attributes. This selective extraction maintains accuracy while reducing computational overhead

Inventive Principle:
Principle #2Taking out (Extraction)

Solution Approach 2:

The system applies different levels of analysis to different aspects of the image, focusing computational resources on visual features that most impact prompt quality. By prioritizing analysis of critical features over exhaustive analysis of all features, the system achieves high accuracy with reduced computational cost

Inventive Principle:
Principle #3Local quality

Data Source

PatentUS12462441B2Iterative image generation from text
Publication Date: 2025.11.04 SONY INTERACTIVE ENTERTAINMENT LLC
  • US12462441B2 patent drawing
  • US12462441B2 patent drawing
  • US12462441B2 patent drawing

AI summary

Methods and systems are presented for automatically identifying additional descriptors of an image generated by a text-to-image generator from an initial prompt. The additional descriptors are either incorporated into the initial prompt or made into a new prompt in order to produce another image from the text-to-image generator. The initial prompt and additional descriptors can describe visual features represented in images including content, artistic styles, visual perspectives, and other visible attributes of images. The additional descriptors can be incorporated into the initial prompt by replacing or supplementing existing descriptors. Subsequent images generated by the text-to-image generator can be used to iteratively produce additional descriptors.