Text Compositing in AI Image Generation for Accurate Character Display

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing image generation models struggle to accurately incorporate specific characters in generated images, leading to unsatisfactory character display effects.

Innovation Solution

An image generating method that includes obtaining target text with background image and text description, determining the text to be displayed, generating a first image, and compositing the text onto the image to ensure accurate character inclusion.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Manufacturing precision

If existing image generation models are used to generate images with specific characters, then the image generation process is simple, but the character display accuracy is poor

Engineering Contradiction:
Improvecharacter display accuracyVSAvoidimage generation process complexity
Core Design Contradiction:
Manufacturing precisionVSDevice complexity

Solution Approach 1:

The patent divides the image generation process into two independent stages: first generating the background image without text, then separately adding the required text elements. This segmentation allows each stage to optimize for its specific function - image generation for aesthetics and text placement for accuracy - thereby improving character display accuracy without requiring the entire system to become significantly more complex.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent introduces an intermediary step of creating a mask image that defines the precise location and shape of text elements before compositing them onto the background image. This mask serves as a mediator between the text requirements and the final image, enabling accurate character display by providing precise spatial information that guides the text compositing process.

Inventive Principle:
Principle #24Intermediary (Mediator)

2Reliability

If text is directly integrated into image generation, then the process is efficient, but character accuracy and user intent matching are insufficient

Engineering Contradiction:
Improvecharacter accuracyVSAvoidimage generation efficiency
Core Design Contradiction:
ReliabilityVSProductivity

Solution Approach 1:

The patent performs preliminary actions by first generating the background image and creating a mask image that pre-defines where text elements will be placed. This preparation step ensures that when text is added in the second stage, it can be positioned with high accuracy according to user intent, while the overall process remains efficient through this structured approach rather than attempting to integrate all operations simultaneously.

Inventive Principle:
Principle #10Preliminary action

3Manufacturing precision

If existing models generate images with text requirements, then the workflow is simple, but the character display effect is unsatisfactory

Engineering Contradiction:
Improvecharacter display effectVSAvoidworkflow simplicity
Core Design Contradiction:
Manufacturing precisionVSEase of operation

Solution Approach 1:

The patent segments the workflow into clear, distinct steps: background image generation, mask image creation, and text compositing. Although this increases the number of operations, each step is simple and well-defined, making the overall process easier to control and adjust for specific character display requirements compared to attempting to handle all requirements in a single complex operation.

Inventive Principle:
Principle #1Segmentation

Data Source

PatentUS20260057585A1Image generating method, apparatus, electronic device and storage medium
Publication Date: 2026.02.26 BEIJING ZITIAO NETWORK TECH CO LTD
  • US20260057585A1 patent drawing
  • US20260057585A1 patent drawing
  • US20260057585A1 patent drawing

AI summary

The present disclosure relates to an image generating method and apparatus, an electronic device, and an storage medium, the method includes: obtaining a target text, the target text including description information of background image and text to be displayed; determining the text to be displayed based on the target text; generating a first image based on the target text; compositing the text to be displayed with the first image to obtain a target image, and the target image includes the text to be displayed.