AI Image Generation Using Multi-Step Prompting for Brand Consistency

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing image generation solutions struggle to produce high-fidelity images of branded products and contextually relevant images for brand-agnostic recommendations, lacking customization and precision in capturing product details and specific contexts.

Innovation Solution

A method involving a multi-step process using machine-learned language models and image generation models, where a fine-tuned image generation model is used to generate high-quality, contextually relevant, and branded images by inputting images of branded items and using a knowledge graph to inform prompts for image generation.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Manufacturing precision

If existing image generation solutions are used, then image generation can be performed, but the images lack high fidelity of branded products and contextual relevance

Engineering Contradiction:
Improveimage fidelityVSAvoidbrand consistency
Core Design Contradiction:
Manufacturing precisionVSReliability

Solution Approach 1:

The image generation process is divided into multiple sequential steps: first generating a theme, then creating a detailed prompt based on the theme, and finally generating images from the prompt. This segmentation allows each step to focus on specific aspects (contextual relevance, then visual details), improving both image fidelity and brand consistency without requiring a single complex model to handle everything at once.

Inventive Principle:
Principle #1Segmentation

2Manufacturing precision

If traditional image generation methods are used, then images can be produced, but they lack customization and precision in capturing product details

Engineering Contradiction:
Improveproduct detail precisionVSAvoidsystem complexity
Core Design Contradiction:
Manufacturing precisionVSDevice complexity

Solution Approach 1:

The system performs preliminary actions by first generating a theme and then creating a detailed prompt before actual image generation. This preliminary structuring of information (context → theme → prompt → image) allows the final image generation step to focus specifically on visual fidelity and product details, rather than having to infer everything from a single undifferentiated input.

Inventive Principle:
Principle #10Preliminary action

3Reliability

If high-quality branded images are generated, then brand consistency is maintained, but the process becomes more resource-intensive

Engineering Contradiction:
Improvebrand visual identity consistencyVSAvoidcomputational resource consumption
Core Design Contradiction:
ReliabilityVSUse of energy by moving object

Solution Approach 1:

The system introduces intermediate textual representations (theme and prompt) that act as mediators between the input context and the final image generation. These intermediaries encode brand identity and contextual information in a structured format that guides the image generation model, reducing the computational burden compared to trying to infer all constraints directly from raw input data.

Inventive Principle:
Principle #24Intermediary (Mediator)

Data Source

PatentUS20250037323A1Generating artificial intelligence (AI)-based images using large language machine-learned models
Publication Date: 2025.01.30 MAPLEBEAR INC
  • US20250037323A1 patent drawing
  • US20250037323A1 patent drawing
  • US20250037323A1 patent drawing

AI summary

An online system performs a task in conjunction with the model serving system or the interface system. The system generates a first prompt for input to a machine-learned language model, which specifies contextual information and a first request to generate a theme. The system provides the first prompt to a model serving system for execution by the machine-learned language model, receives a first response, and generates a second prompt. The second prompt specifies the theme and a second request to generate a third prompt for input to an image generation model that includes a third request to generate one or more images of one or more items associated with the theme. The system receives the third prompt by executing the model on the second prompt, provides the third prompt to the image generation model, and receives one or more images for presentation.