Knowledge Graph-Enhanced Training for Image Generation Models

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing image generation solutions struggle to produce realistic and semantically accurate images from descriptive text, particularly in cross-modality scenarios such as abstract painting and landscape generation.

Innovation Solution

A method involving knowledge graph-enhanced training of an image generation model, where image and text samples are enhanced using entity information and knowledge data, followed by aggregation and training with specific loss functions to improve semantic consistency.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Manufacturing precision

If existing image generation solutions are used, then image generation can be performed, but the generated images do not align well with commonsense and factual description text

Engineering Contradiction:
Improveimage generation accuracyVSAvoidsemantic consistency
Core Design Contradiction:
Manufacturing precisionVSReliability

Solution Approach 1:

The patent applies preliminary action by pre-training the image generation model with knowledge graph data before actual image generation. The knowledge graph is constructed in advance with entity relationships, and the model is trained on this structured knowledge to establish semantic understanding before encountering real generation tasks, thereby improving both accuracy and semantic consistency

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The patent introduces a knowledge graph as an intermediary between the input text and the image generation process. The knowledge graph acts as a mediator that structures and validates the semantic relationships in the input text, guiding the generation model to produce images that are semantically consistent with the description while maintaining factual accuracy

Inventive Principle:
Principle #24Intermediary (Mediator)

2Reliability

If knowledge graph enhancement is applied to improve semantic consistency, then image-text alignment improves, but training complexity and computational resources increase

Engineering Contradiction:
Improvesemantic consistencyVSAvoidtraining system complexity
Core Design Contradiction:
ReliabilityVSDevice complexity

Solution Approach 1:

The patent segments the training process into distinct phases: knowledge graph construction, model pre-training on knowledge graph data, and fine-tuning on image-text pairs. This segmentation allows each component to be optimized independently and simplifies the overall training workflow, reducing system complexity while maintaining semantic consistency

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent employs parameter changes by adjusting the knowledge graph depth, entity selection criteria, and training data composition to balance semantic consistency with training complexity. By dynamically modifying these parameters based on computational resources and performance requirements, the system achieves optimal trade-off without excessive complexity

Inventive Principle:
Principle #35Parameter changes

Data Source

PatentUS12406472B2Method of training image generation model, and method of generating image
Publication Date: 2025.09.02 BEIJING BAIDU NETCOM SCI & TECH CO LTD
  • US12406472B2 patent drawing
  • US12406472B2 patent drawing
  • US12406472B2 patent drawing

AI summary

A method of training an image generation model, and a method of generating an image. A specific implementation solution includes: acquiring a first image sample and a first text sample matched with the first image sample; performing an enhancement on at least one type of sample in the first image sample and the first text sample according to a predetermined knowledge graph, so as to obtain at least one type of sample in a second image sample obtained by the enhancement and a second text sample obtained by the enhancement; and training the image generation model according to a training set selected from a first training set, a second training set or a third training set, until the image generation model converges.