Prompted Text-to-Image Generation via Iterative AI Refinement
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current text-to-image generation technologies lack an intuitive user interface for refining image characteristics, leading to suboptimal image outputs that do not accurately represent user intentions.
Innovation Solution
A method utilizing a graphical user interface that allows users to input image characteristics, receive prompts for additional details, and generate images using a combination of generative artificial intelligence language and text-to-image models, enabling iterative refinement of image outputs.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Device complexity
If a simple text input interface is used for image generation, then the device complexity is reduced, but the image quality and accuracy of representing user intentions deteriorates
Solution Approach 1:
The interface is segmented into multiple interaction stages: initial text input, followed by structured question-response pairs that break down image characteristics into discrete controllable parameters. This allows complex image specification to be achieved through simple, sequential interactions rather than requiring a complex simultaneous input interface.
Solution Approach 2:
The system performs preliminary action by proactively generating and presenting targeted questions about image characteristics before the image generation is finalized. This allows users to refine their intentions through guided questioning, improving image accuracy without requiring users to anticipate all parameters in advance.
2Manufacturing precision
If detailed image characteristics are requested through multiple questions, then the image quality improves, but the time required for image generation increases
Solution Approach 1:
The image generation process uses periodic action by implementing iterative refinement cycles. The system generates an initial image, then presents targeted questions for refinement, repeating this cycle until the user is satisfied. This breaks down the time-consuming detailed specification into manageable periodic interactions rather than requiring all details upfront.
Solution Approach 2:
The system applies partial action by only asking questions about image characteristics that are relevant to the current generation stage and user needs. Not all possible parameters are queried exhaustively, but only those that will meaningfully improve the current image output, balancing detail with efficiency.
3Ease of operation
If a generative AI language model is used to output questions and options, then the ease of operation improves, but the device complexity increases
Solution Approach 1:
The generative AI language model serves as an intermediary between the user's simple text input and the complex image generation parameters. It translates natural language into structured questions and options, shielding users from system complexity while enabling sophisticated image specification through conversational interaction.
Solution Approach 2:
The system uses self-service by employing the generative AI model to automatically generate context-relevant questions and options based on the user's input and the current image generation state. This eliminates the need for pre-programmed fixed question sets, allowing the interface to adapt autonomously to user needs without increasing operational complexity.
Data Source
AI summary
Methods, non-transitory computer-readable storage media and computer or computer systems are described which include or relate to inputting or receiving information on one or more image characteristics from a graphical user interface, outputting one or more questions or options for additional details of the one or more image characteristics on a graphical user interface by way of a generative artificial intelligence language model performed on one or more processor, inputting or receiving the additional details from the graphical user interface, and outputting one or more images by way of a generative artificial intelligence text-to-image model performed on one or more processor based on the one or more image characteristics and the additional details.


