Text-to-Image Model Training With Human Feedback Ranking

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing text-to-image generative models often fail to produce images with satisfactory visual appeal and fidelity due to limitations in semantic understanding, training data quality, and alignment with human expectations.

Innovation Solution

A system that pre-trains a text-to-image diffusion model using a combination of pre-trained text encoders, base and high-resolution diffusion models, and human feedback mechanisms to improve image generation quality, incorporating a web intelligence engine for data retrieval, visual aesthetics and watermark detection, and content moderation models to ensure high-quality and appropriate image output.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Manufacturing precision

If existing text-to-image generative models are used, then image generation speed is maintained, but visual appeal and fidelity are insufficient

Engineering Contradiction:
Improveimage qualityVSAvoidalignment with human expectations
Core Design Contradiction:
Manufacturing precisionVSReliability

Solution Approach 1:

The system pre-trains the diffusion model with high-quality image-text pairs from curated datasets before deployment, preparing the model in advance to better understand semantic relationships and human preferences, thereby improving image quality and alignment without affecting generation speed during actual use

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

Human feedback mechanisms are integrated into the training process, where human annotators provide preferences and corrections on generated images. This feedback loop allows the model to learn from human expectations and continuously improve its output quality and alignment with human preferences

Inventive Principle:
Principle #23Feedback

2Manufacturing precision

If more training data is used to improve semantic understanding, then image quality improves, but training time and computational resources increase

Engineering Contradiction:
Improvesemantic understandingVSAvoidtraining time
Core Design Contradiction:
Manufacturing precisionVSLoss of time

Solution Approach 1:

The system pre-processes and curates training data in advance, creating high-quality image-text pairs from diverse sources before the actual training process. This preliminary data preparation ensures that the model learns from optimized datasets, improving semantic understanding efficiency while reducing overall training time

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The training process is divided into multiple stages: pre-training on large-scale data for general semantic understanding, fine-tuning on curated high-quality data for specific improvements, and reinforcement learning with human feedback for alignment optimization. This segmented approach allows efficient use of computational resources at each stage

Inventive Principle:
Principle #1Segmentation

3Manufacturing precision

If content moderation and aesthetics detection are added, then image appropriateness and quality are improved, but system complexity increases

Engineering Contradiction:
Improveimage quality and appropriatenessVSAvoidsystem architecture
Core Design Contradiction:
Manufacturing precisionVSDevice complexity

Solution Approach 1:

The content moderation model and visual aesthetics model are integrated into the same system architecture as the diffusion model, sharing computational resources and infrastructure. These models work together in a unified framework where the diffusion model generates images while the moderation and aesthetics models evaluate and guide the output, improving quality and appropriateness without requiring completely separate systems

Inventive Principle:
Principle #5Merging (Combining)

Data Source

PatentUS12499519B1Training and deployment of image generation models
Publication Date: 2025.12.16 CASTLE GLOBAL INC
  • US12499519B1 patent drawing
  • US12499519B1 patent drawing
  • US12499519B1 patent drawing

AI summary

In some embodiments, a method receives a text prompt. A text encoder is executed on the text prompt to generate a representation. The method generates a set of images based on the representation and a set of parameters of an image generation model. The set of images is ranked using reward values that are generated by a reward model. The reward model is trained using human input that provided feedback on a quality of generated images using the image generation model. The method outputs one or more images based on the ranking in response to the text prompt.