A hierarchical layout driven arbitrary shape scene text image generation method, system, device and medium

By constructing a scene text image generation method driven by hierarchical layout, the problem of insufficient adaptability in generating arbitrary-shaped text in existing technologies is solved, and the generation of diverse and visually appealing scene text images is realized automatically.

CN120070666BActive Publication Date: 2025-11-21XIDIAN UNIV +1
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
CN202510126918.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-01-27
Publication Date
2025-11-21
Estimated Expiration
2045-01-27

AI Technical Summary

Technical Problem

Existing methods for generating scene text images are not adaptable enough when generating text of arbitrary shapes, making it difficult to meet the needs of expressing complex text forms. The generated results lack visual appeal and overall scene coordination, and cannot meet the application requirements of high quality and high fidelity.

Method used

By constructing a hierarchical layout-driven method for generating arbitrary-shaped scene text images, including a background image generation module, a hierarchical layout generation module, and a scene text image generation module, an image feature extraction network, a region hierarchical layout generation network, a sentence hierarchical layout generation network, and a character hierarchical layout generation network are used. Combined with Bézier curves to represent the character hierarchical layout, diverse scene text images are generated.

Benefits of technology

It achieves automated generation without relying on user-input layout information, and can generate scene text images with diverse layouts such as curvature and tilt, improving the flexibility and visual effect of generation, and meeting the expression needs of complex text forms.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120070666B_ABST
    Figure CN120070666B_ABST
Patent Text Reader

Abstract

A hierarchical layout driven arbitrary shape scene text image generation method, system, device and medium, the method comprising: preprocessing a scene text image training set to obtain a preprocessed scene text layout generation training set, a scene text image generation training set and a scene text image generation test set; constructing a background image generation module, a hierarchical layout generation module and a scene text image generation module; constructing a complete arbitrary shape scene text image generation model; training the hierarchical layout generation module and the scene text image generation module respectively to obtain the weight file of the trained hierarchical layout generation module and the scene text image generation module; model inference to obtain the final scene text image; the system, device and medium are used to realize the method; the present application can automatically generate arbitrary shape scene text images without relying on user input layout.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the interdisciplinary field of computer vision and computer graphics, specifically relating to a method, system, device, and medium for generating text images of arbitrary shapes in a hierarchical layout-driven environment. Background Technology

[0002] Scene text-to-image generation technology has broad application prospects in advertising design, game development, and data synthesis. Its research significance lies in expanding the technological boundaries of language and vision integration and solving key bottlenecks in text-to-image generation in complex scenes. This technology aims to generate semantically coherent and visually appealing images, involving deep collaboration between text semantic expression, visual layout, font style, and background integration. In advertising design, this technology can achieve creative text layouts that match the scene, significantly improving design efficiency; in game development, it can generate background text that meets the needs of virtual scenes, shortening the development cycle; in data synthesis, it can generate complex text images on a large scale, providing high-quality training data for text detection and recognition models. However, its implementation faces many challenges, such as balancing global semantics with local detail expression, adapting to complex and irregular text layouts, and ensuring the unity and coordination of various visual features such as fonts, colors, and backgrounds.

[0003] Existing methods for generating scene text images have made some progress in generating text with regular layouts, but significant technical bottlenecks remain for generating text of arbitrary shapes. Most current mainstream methods employ simple bounding box layout modeling, which is insufficient to meet the demands of practical applications for arbitrary text layouts, especially when generating irregular arrangements such as arcs, curves, or italics, demonstrating a clear lack of adaptability. This modeling approach not only limits the flexibility of scene text image generation but also results in a lack of visual appeal and overall scene harmony in the generated results, making it difficult to meet the requirements for expressing complex text forms.

[0004] Patent application CN116935169B discloses a training and application method for a text-based image model, belonging to the field of artificial intelligence technology. This method significantly enhances the ability of the text-based image model to generate diverse images by introducing differentiable diversity constraints during training. The diversity constraints play a crucial role in the model training and parameter tuning stages, enabling the model to generate diverse images. However, the optimization objective of this method mainly focuses on improving image diversity, without specifically optimizing for generating high-quality, high-fidelity scene text images. This results in low-quality generated scene text images that fail to meet the application requirements for high-precision expression of text content.

[0005] Patent application CN116977774A discloses an image generation method, apparatus, device, and storage medium, relating to the field of artificial intelligence technology. This method takes scene description text as input, extracts semantic features from the text to predict its corresponding scene layout information, and uses this layout information to characterize the relative positional relationships between objects in the scene. By fusing layout information into the scene description text to generate target description text, and combining it with an initial noisy image for noise reduction, a high-precision image satisfying the target scene layout is generated. This method effectively improves the accuracy of scene content in the generated image. However, the generation process only optimizes the layout of scene objects, without incorporating layout information from the scene text content, thus failing to highlight the semantic features of the scene text. Furthermore, the generated scene text content lacks fidelity, making it difficult to meet the practical needs of generating complex scene text images. Summary of the Invention

[0006] To overcome the shortcomings of the prior art, the present invention aims to provide a method, system, device and medium for generating arbitrary-shaped scene text images driven by hierarchical layout. By generating a background image, the invention perceives content information of different scales in the background image and generates three different hierarchical layouts: regional hierarchical layout, sentence hierarchical layout and character hierarchical layout. Based on the character hierarchical layout, arbitrary-shaped scene text images can be generated. The present invention does not rely on user input layout and generates text with diversity and realism.

[0007] To achieve the above objectives, the technical solution adopted by the present invention is as follows:

[0008] A method for generating text images of arbitrary-shaped scenes driven by hierarchical layout includes the following steps:

[0009] Step 1: Preprocess the scene text image training set. The scene text image training set includes a scene text layout generation training set, a scene text image generation training set, and a scene text image generation test set. Finally, the preprocessed scene text layout generation training set, scene text image generation training set, and scene text image generation test set are obtained.

[0010] Step 2: Construct a background image generation module, a hierarchical layout generation module, and a scene text image generation module;

[0011] Step 3: Based on the background image generation module, hierarchical layout generation module, and scene text image generation module constructed in Step 2, construct a complete arbitrary shape scene text image generation model;

[0012] Step 4: Use the preprocessed scene text layout and scene text image generation training sets from Step 1 to train the hierarchical layout generation module and scene text image generation module constructed in Step 2, respectively, to obtain the weight files of the trained hierarchical layout generation module and scene text image generation module.

[0013] Step 5: Based on the scene text image preprocessed in Step 1, generate a test set. Using the weight files of the hierarchical layout generation module and the scene text image generation module trained in Step 4, perform model inference through the complete arbitrary shape scene text image generation model constructed in Step 3 to obtain the final scene text image.

[0014] The specific method for step 1 is as follows:

[0015] Step 1.1: The scene text layout generation training set includes a textless background image and content containing region, sentence, and character hierarchical layouts. The textless background image and content containing region, sentence, and character hierarchical layouts in the scene text layout generation training set are used to form training background image-hierarchical layout pairs. The region and sentence hierarchical layouts consist of 4 values, with a total of 2 control points, which respectively represent the horizontal and vertical coordinates of the top left and bottom right corners of the layout. The character hierarchical layout consists of 16 values, which respectively represent the horizontal and vertical coordinates of the 4 control points of the top Bézier curve and the 4 control points of the bottom Bézier curve of the layout.

[0016] Step 1.2 involves preprocessing the training background image-hierarchical layout pair from Step 1.1. The textless background image in the training background image-hierarchical layout pair is read, resized to 3×512×512, and converted to tensor format. Simultaneously, the mean and variance are used to normalize the three color channels of the textless background image. The input textless background image color values ​​range from integers [0~255] to floating-point values ​​[-1~1]. The json library is used to read the three hierarchical layouts corresponding to the textless background image in Step 1.1: region, sentence, and character. The horizontal and vertical coordinate values ​​range from floating-point values ​​[0~511] to floating-point values ​​[0~1], resulting in the preprocessed training background image-hierarchical layout pair.

[0017] Step 1.3: The scene text image generation training set includes text prompts, images without text backgrounds, character hierarchy layout content, and scene text images. Preprocessing of the scene text image generation training set specifically involves: processing a single image without text background using the same process as in Step 1.2; and reading the text prompts and character hierarchy layout content corresponding to a single image without text background using a json library to obtain the preprocessed scene text image generation training set.

[0018] Step 1.4: The scene text image generation test set includes text prompts and scene text images; the scene text image generation test set is preprocessed, specifically by reading the text prompts using a json library to obtain the preprocessed scene text image generation test set.

[0019] The specific method for step 2 is as follows:

[0020] Step 2.1: Construct a background image generation module, including a text-generated image model and a scene text detection model, used to generate a textless background image based on the text prompts read in Step 1.3;

[0021] Step 2.1.1: Remove the text content to be rendered from the text prompts read in Step 1.3, and generate a background image using the text-generated image model;

[0022] Step 2.1.2: Use the scene text detection model to detect the background image generated by the text-generated image model in step 2.1.1. If text is detected, repeat steps 2.1.1-2.1.2 until no text is detected, and output a text-free background image.

[0023] Step 2.2: Construct a hierarchical layout generation module. The hierarchical layout generation module is a network model based on transformers, including an image feature extraction network, a region hierarchical layout generation network, a sentence hierarchical layout generation network, and a character hierarchical layout generation network. It is used to extract the features of the textless background image generated in Step 2.1, gradually generate the content of the region, sentence and character hierarchical layout, and finally output the character hierarchical layout.

[0024] Step 2.2.1: The image feature extraction network is an encoder based on ResNet50 fine-tuning;

[0025] The image features are extracted by using a ResNet50-based encoder fine-tuned on the textless background image generated in step 2.1;

[0026] Step 2.2.2: The region-level layout generation network is a network model based on the Transformer architecture. The region-level layout generation network includes two decoder layers. Each decoder layer includes an Embedding layer, a SelfAttention layer, two Layer Norm layers, a Cross Attention layer, a Feed Forward layer, an MLP layer, and a Linear layer. The output of the Embedding layer is connected to the input of the Self Attention layer. The output of the SelfAttention layer is connected to the input of the first Layer Norm layer. The output of the first Layer Norm layer is connected to the input of the Cross Attention layer. The output of the Cross Attention layer is connected to the input of the second Layer Norm layer. The output of the second Layer Norm layer is connected to the input of the Feed Forward layer. The output of the Feed Forward layer is connected to the input of the MLP layer and the Linear layer, respectively. The MLP layer outputs the region-level layout, and the Linear layer outputs its confidence score. The image features output in Step 2.2.1 are input into the region-level layout generation network. The region-level layout generation network predicts the region-level layout and its confidence score by judging the content of different regions of the image features, and outputs the region-level image features and the region-level layout.

[0027] Step 2.2.3: Construct a sentence hierarchy layout generation network. The sentence hierarchy layout generation network adds the region-level image features output by the region hierarchy layout generation network in Step 2.2.2 to the image features output in Step 2.2.1 as input features for encoding. The architecture of the sentence hierarchy layout generation network is the same as that of the region hierarchy layout generation network. By judging the relationship between the content of different feature regions, it predicts the layout of the sentence hierarchy and its confidence, and outputs the sentence hierarchy image features and sentence hierarchy layout.

[0028] Step 2.2.4: The character hierarchy layout generation network adds the sentence hierarchy image features output by the sentence hierarchy layout generation network in step 2.2.3 to the image features output in step 2.2.1 as input features for encoding. The architecture of the character hierarchy layout generation network is the same as that of the region hierarchy layout generation network. By judging the relationship between the content of different feature regions, it predicts the character hierarchy layout and its confidence, and finally outputs the character hierarchy layout.

[0029] Step 2.3: Construct a scene text image generation module, including a controllable graph generation model. Based on the text prompts read in Step 1.3, the character hierarchy layout generated in Step 2.2, and the textless background image generated in Step 2.1, the controllable graph generation model is used to generate scene text images that accurately match the description of the text prompts.

[0030] The controllable graph generation model consists of a VAE-based image encoder, a Transformer-based text encoder, and a Diffusion-based conditional diffusion model.

[0031] Step 2.3.1: Based on the characters and their number in the text prompts read in Step 1.3, further refine the character hierarchy layout generated in Step 2.2 to obtain an independent layout for each character and the characters of the text to be rendered;

[0032] Step 2.3.2: Place the characters of the text to be rendered obtained in Step 2.3.1 into the independent layout position of each character specified in Step 2.3.1. The font file used for placement is Arial Unicode, and the text placement image is obtained.

[0033] Step 2.3.3: Input the text prompts read in Step 1.4 into a Transformer-based text encoder to obtain text features. Combine the textless background image generated in Step 2.1.2 and the text placement image obtained in Step 2.3.2 and input them into a VAE-based image encoder for feature extraction to obtain text placement background image features, and add random Gaussian noise. Input the obtained text features and the text placement background image features with added random Gaussian noise into a Diffusion-based conditional diffusion model to generate a scene text image.

[0034] The specific method for step 3 is as follows:

[0035] Step 3.1: Connect the text prompt words read in step 1.3 with the background image generation module constructed in step 2.1. The text-generated image model and the scene text detection model in the background image generation module are both loaded with pre-trained model parameters. The output image size of the text-generated image model is the same as the input image size of the scene text detection model. The output result is a textless background image.

[0036] Step 3.2: Input the textless background image output from Step 3.1 into the hierarchical layout generation module constructed in Step 2.2. The image feature extraction network extracts image features by continuously convolving the image upwards. The image features are then input into the region hierarchical layout generation network. Based on the feature values ​​of the input image, the region hierarchical layout generation network continuously predicts the next region hierarchical layout, selects the best layout based on the confidence level, and obtains the region hierarchical image features and the region hierarchical layout. The image features are then added to the region hierarchical image features and input into the sentence hierarchical layout generation network. The same process is repeated to obtain the sentence hierarchical image features and the sentence hierarchical layout. This process continues until the character hierarchical layout is finally obtained, and the final character hierarchical layout is selected based on the best confidence level.

[0037] Step 3.3: The character hierarchy layout obtained in Step 3.2 is segmented and filled according to the characters of the text to be rendered in the text prompts read in Step 1.3 to obtain the text placement image. The text prompts read in Step 1.3, the textless background image generated in Step 2.1, and the text placement image are connected with the scene text image generation module constructed in Step 2.3. After combining the text placement image with the textless background image, the text placement background image is obtained. Feature extraction is performed using the VAE-based image encoder in Step 2.3 to obtain the text placement background image features, and random Gaussian noise is added. Conditional expansion based on Diffusion is then performed. The diffusion model takes the text placement background image features with added random Gaussian noise and the current time step as input, predicts the noise added between the current time step and the previous time step, subtracts the predicted noise from the input image, and obtains the predicted image corresponding to the previous time step. The diffusion model is then used again to take the predicted image, the text placement background image features with added random Gaussian noise, and the current time step as input, predicts the noise added between the current time step and the previous time step, and obtains the predicted image of the previous time step again. This process is repeated until the initial time step is reached, resulting in a complete arbitrary shape scene text image generation model.

[0038] The specific method for step 4 is as follows:

[0039] Step 4.1: Take the textless background image from the training background image-hierarchical layout pair after preprocessing in Step 1.2 as the input to the hierarchical layout generation module constructed in Step 2, and take the corresponding hierarchical layout as the real hierarchical layout.

[0040] Step 4.2: Take the text prompts, textless background images, and character hierarchy layout content from the training set of the scene text image generation in Step 1.3 as input to the scene text image generation module constructed in Step 2, and take the corresponding scene text images as real scene text images.

[0041] Step 4.3: During training, the loss function is used to evaluate the network model performance and optimize the network model parameters. The optimization of the hierarchical layout generation module adopts a hierarchical loss mechanism. The loss functions defined for the region hierarchical layout generation network, the sentence hierarchical layout generation network, and the character hierarchical layout generation network are as follows:

[0042] Step 4.3.1, for training the region-level layout generation network and the sentence-level layout generation network, the following three loss constraints are adopted:

[0043]

[0044] Among them, L L1 This represents the L1 loss between the region-level layout output in step 2.2.2 and the statement-level layout output in step 2.2.3, and the ground truth bounding box, measuring the difference between the predicted bounding box B and the ground truth bounding box. The coordinate difference between them; L GIoU This represents the calculation of the generalized IoU loss between bounding boxes, which extends the IoU calculation to the minimum bounding rectangle C, constraining the generated bounding boxes to be closer in extent to the true bounding boxes; L ol This involves calculating the same overlap loss, L. L1 and L GIoU Divide all by N to distribute the effect evenly across all bounding boxes, while L ol That is, averaged to Among several pairwise combinations of possible bounding boxes;

[0045] The total bounding box loss is:

[0046] L bbox =c1L L1 +c2L GIoU +c3L ol ,

[0047] Where c1 = 5, c2 = 2, c3 = 1;

[0048] Step 4.3.2: For training the character hierarchy layout generation network, calculate the Bézier curve loss using the character hierarchy layout generated in Step 2.2 and the real layout, constrained by L1 loss.

[0049]

[0050] Step 4.3.3: For training the region-level layout generation network, sentence-level layout generation network, and character-level layout generation network, add confidence loss and use standard cross-entropy loss to evaluate the confidence loss for each bounding box and Bézier curve:

[0051]

[0052] Where q is the confidence level of the model prediction. The confidence level of the target label (with a value of 1);

[0053] Step 4.3.4: For each decoder layer of the region-level layout generation network, sentence-level layout generation network, and character-level layout generation network, the corresponding loss is calculated, and the matching loss between the output and the actual layout is calculated using the Hungarian algorithm to obtain the final loss. Combining the total bounding box loss obtained in Step 4.3.1, the Bézier curve loss obtained in Step 4.3.2, and the confidence loss obtained in Step 4.3.3, the final total loss function is:

[0054] L=λ1L bbox +λ2L Bezier +λ3L conf ,

[0055] Where λ1=1, λ2=5, λ3=1;

[0056] Step 4.4, for the scene text image generation module, to quantify the differences between the original image and the predicted image across all text regions, a text-aware loss is used:

[0057]

[0058] Where h,w = 512, 512 represents the height and width of the image, and here it represents the feature map before the fully connected layer. The text writing information at position p in the original and predicted images is represented. Since the time step t is directly related to the text quality of the predicted image x'0, an adjustment function φ(t) is used to dynamically adjust the weights of the loss function, where... These are coefficients in the diffusion process;

[0059] Finally, the weight files for the trained hierarchical layout generation module and scene text image generation module are obtained.

[0060] The specific method for step 5 is as follows:

[0061] Step 5.1 Load the complete arbitrary-shape scene text image generation model constructed in Step 3, read the weight files of the hierarchical layout generation module and the scene text image generation module trained in Step 4, load the weight files into the structure of the hierarchical layout generation module and the scene text image generation module, and set the complete arbitrary-shape scene text image generation model to inference mode and fix the model parameters.

[0062] Step 5.2: Input the text prompts from the preprocessed scene text image generation test set in Step 1.4 into the complete arbitrary shape scene text image generation model constructed in Step 3. Use the background image generation module to modify the text prompts and perform text detection on the generated background image to obtain a textless background image. Use the hierarchical layout generation module to encode the textless background image, extract features, and generate a hierarchical character layout. Use the scene text image generation module to segment and fill the character layout, and combine it with the textless background image to obtain a text placement background image. After encoding the text placement background image, obtain the text placement background image features and add random Gaussian noise. Encode the text prompts to obtain text features. Use the text placement background image features with added random Gaussian noise and the text features as input to the conditional diffusion model to gradually diffuse and generate the final scene text image.

[0063] The present invention also provides a hierarchical layout-driven arbitrary shape scene text image generation system, comprising:

[0064] The scene text image training set preprocessing module is used to preprocess the scene text image training set, which includes a scene text layout generation training set, a scene text image generation training set, and a scene text image generation test set. Finally, the preprocessed scene text layout generation training set, scene text image generation training set, and scene text image generation test set are obtained.

[0065] The first module is used to build a background image generation module, a hierarchical layout generation module, and a scene text image generation module.

[0066] The arbitrary-shape scene text image generation model building module is used to build a complete arbitrary-shape scene text image generation model based on the background image generation module, the hierarchical layout generation module, and the scene text image generation module.

[0067] The first module is used to generate training sets for the preprocessed scene text layout and scene text image, and to train the constructed hierarchical layout generation module and scene text image generation module respectively, so as to obtain the weight files of the trained hierarchical layout generation module and scene text image generation module.

[0068] The model inference module is used to generate a test set based on scene text images. It uses the weight files of the trained hierarchical layout generation module and scene text image generation module to construct a complete arbitrary-shape scene text image generation model and perform model inference to obtain the final scene text image.

[0069] The present invention also provides a hierarchical layout-driven arbitrary shape scene text image generation device, comprising:

[0070] Memory: A computer program that stores the above-described hierarchical layout-driven method for generating text images of arbitrary-shaped scenes, and is a computer-readable device;

[0071] Processor: Used to implement the hierarchical layout-driven arbitrary shape scene text image generation method when executing the computer program.

[0072] The present invention also provides a computer-readable storage medium storing a computer program that, when executed by a processor, enables the implementation of the hierarchical layout-driven arbitrary shape scene text image generation method.

[0073] Compared with the prior art, the beneficial effects of the present invention are as follows:

[0074] 1. This invention constructs a complete arbitrary-shape scene text image generation model in step 3. The background image generation module reads and processes the preprocessed scene text image obtained in step 1.4 to generate text prompts in the test set, generating a textless background image. The hierarchical layout generation module reads and processes the textless background image to generate the final character hierarchy layout. The scene text image generation module reads and processes the text prompts, the textless background image, and the final character hierarchy layout to gradually generate the final scene text image. This invention can automatically generate the final scene text image from the input text prompts with one click without relying on user input layout information.

[0075] 2. The hierarchical layout generation module constructed in step 2.2 of this invention includes an image feature extraction network, a region hierarchical layout generation network, a sentence hierarchical layout generation network, and a character hierarchical layout generation network. It is used to extract the features of the textless background image generated in step 2.1, gradually generate the content of the region, sentence, and character hierarchical layouts, and finally output the character hierarchical layout. The final character hierarchical layout can include diverse layout forms such as curved and tilted shapes, generating diverse scene text images.

[0076] 3. The scene text image generation module constructed in step 2.3 of this invention includes a controllable graph generation model. The text prompt words read in step 1.3 are used as input, and the character hierarchy layout generated in step 2.2 and the textless background image generated in step 2.1 are used as control conditions. The controllable graph generation model can generate scene text images that accurately match the description of the text prompt words.

[0077] 4. During the training process of this invention, the optimization of the hierarchical layout generation module adopts a layered loss mechanism. The region hierarchical layout generation network and the sentence hierarchical layout generation network are constrained by the total bounding box loss designed in step 4.3.1; the character hierarchical layout generation network is constrained by the Bézier curve loss designed in step 4.3.2; and the region hierarchical layout generation network, the sentence hierarchical layout generation network, and the character hierarchical layout generation network are constrained by the confidence loss designed in step 4.3.3. The total loss function, constructed by weighting the above losses according to certain proportions, is used to optimize the hierarchical layout generation module. The scene text image generation module is optimized using the text perception loss designed in step 4.4. The final trained weights of the hierarchical layout generation module can accurately generate refined character hierarchical layouts, and the weights of the scene text image generation module can accurately generate scene text images described by text prompts.

[0078] In summary, this invention constructs a complete arbitrary shape scene text image generation model and its hierarchical layout generation module, which can automatically generate single-line or multi-line scene text images with various conditions such as bending and tilting without relying on user input layout.

[0079] This invention proposes a hierarchical layout representation method and a text-based image model that supports the generation of text images in arbitrary-shaped scenes. Through hierarchical layout generation technology, the generated text blends naturally with the background, resulting in a more realistic visual effect. This method not only significantly improves layout flexibility by introducing a character hierarchy layout represented by Bézier curves, supporting diverse text style generation, but also simplifies user operation through a complete arbitrary-shaped scene text image generation model. Users only need text prompts to generate scene text images that meet their needs, improving the ease of use and practicality of the technology and overcoming the shortcomings of existing scene text image generation algorithms. Attached Figure Description

[0080] Figure 1 This is a flowchart of the training process for the hierarchical layout generation module of the present invention.

[0081] Figure 2 This is a flowchart of the training process for the scene text image generation module of the present invention.

[0082] Figure 3 This is a complete reasoning flowchart of the present invention.

[0083] Figure 4 This is the network structure diagram for generating the regional hierarchical layout of the present invention.

[0084] Figure 5 This is a diagram of the statement / character hierarchical layout generation network structure of the present invention.

[0085] Figure 6These are comparison images of the generation effects of this invention. Detailed Implementation

[0086] To make the implementation process and features of the method clearer and easier to understand, the present invention will be further described in detail below with reference to the accompanying drawings and specific embodiments.

[0087] Existing methods often require predefined text layouts or text position prediction using simple rectangles, making it difficult to accurately represent complex text structures. To address the shortcomings of existing scene text image generation technologies, this invention proposes a hierarchical layout-driven method for generating arbitrary-shaped scene text images, filling the technological gap in flexibly generating arbitrary-shaped scene text images. Based on a fully constructed arbitrary-shaped scene text image generation model, this invention divides the generation of the overall scene text image into three steps: background image generation, hierarchical layout generation, and scene text image generation. This significantly improves the ability to generate complex text layouts and achieves automated generation from simple text prompts to high-quality scene text images, effectively solving the bottlenecks of traditional methods in terms of flexibility, adaptability, and generation quality.

[0088] The core content of this invention includes: a hierarchical layout generation module that generates text layout structures at the region, sentence, and character levels through a layer-by-layer refinement approach, particularly by introducing Bézier curves to represent character-level layouts, enabling flexible expression of complex text forms such as curves and italics, thus solving the problem of inaccurate representation by traditional rectangular boxes; a directional control mechanism that utilizes the multi-layer Transformer architecture of the hierarchical layout generation module to automatically generate character hierarchical layouts in conjunction with image content, providing more refined geometric and directional control, resulting in a high degree of integration between the generated text and the background visual content; and an innovative model for generating text images of arbitrary shapes that innovatively divides the generation of complete scene text images into three steps: background image generation, hierarchical layout generation, and scene text image generation, achieving automated generation from simple text prompts to high-quality scene text images. These technologies collectively solve the problems of poor layout flexibility and insufficient detail control in existing text-to-image models, significantly improving the effectiveness and practicality of text generation.

[0089] A method for generating text images of arbitrary-shaped scenes driven by hierarchical layout includes the following steps:

[0090] Step 1: Preprocess the scene text image training set. The scene text image training set includes a scene text layout generation training set, a scene text image generation training set, and a scene text image generation test set. Finally, the preprocessed scene text layout generation training set, scene text image generation training set, and scene text image generation test set are obtained.

[0091] The specific method for step 1 is as follows:

[0092] Step 1.1: The scene text layout generation training set includes a textless background image and content containing region, sentence, and character hierarchical layouts. The textless background image and content containing region, sentence, and character hierarchical layouts in the scene text layout generation training set are used to form training background image-hierarchical layout pairs. The region and sentence hierarchical layouts consist of 4 values, with a total of 2 control points, which respectively represent the horizontal and vertical coordinates of the top left and bottom right corners of the layout. The character hierarchical layout consists of 16 values, which respectively represent the horizontal and vertical coordinates of the 4 control points of the top Bézier curve and the 4 control points of the bottom Bézier curve of the layout.

[0093] Step 1.2: Preprocess the training background image-hierarchical layout pair from Step 1.1. Using the Pillow library, read the textless background image from the training background image-hierarchical layout pair, adjust the textless background image to 3×512×512 and convert it to tensor format to meet the input requirements of the arbitrary shape hierarchical text layout generation model. Simultaneously, normalize the three color channels of the textless background image using mean values ​​of 0.485, 0.456, and 0.406 and variance values ​​of 0.229, 0.224, and 0.225 respectively. The color value range of the input textless background image is mapped from integer values ​​of [0~255] to floating-point values ​​of [-1~1]. Using the json library, read the three hierarchical layouts corresponding to the textless background image in Step 1.1: region, sentence, and character, and map the horizontal and vertical coordinate values ​​from floating-point values ​​of [0~511] to floating-point values ​​of [0~1] to obtain the preprocessed training background image-hierarchical layout pair.

[0094] Step 1.3: The scene text image generation training set includes text prompts, images without text backgrounds, character hierarchy layout content, and scene text images. Preprocessing of the scene text image generation training set specifically involves: processing a single image without text background using the same process as in Step 1.2; and reading the text prompts and character hierarchy layout content corresponding to a single image without text background using a json library to obtain the preprocessed scene text image generation training set.

[0095] Step 1.4: The scene text image generation test set includes text prompts and scene text images; the scene text image generation test set is preprocessed, specifically by reading the text prompts using a json library to obtain the preprocessed scene text image generation test set.

[0096] Step 2: Construct a background image generation module, a hierarchical layout generation module, and a scene text image generation module;

[0097] The specific method for step 2 is as follows:

[0098] Step 2.1: Construct a background image generation module, including a text-generated image model and a scene text detection model, used to generate a textless background image based on the text prompts read in Step 1.3;

[0099] Step 2.1.1: Remove the text content to be rendered from the text prompts read in Step 1.3, and generate a background image using the text-generated image model. The text-generated image model uses a model based on the Diffusion principle for generation.

[0100] Step 2.1.2: Use the scene text detection model to detect the background image generated by the text-generated image model in step 2.1.1. If text is detected, repeat steps 2.1.1-2.1.2 until no text is detected, and output a text-free background image. The scene text detection model uses the basic PP-OCRv3 model for detection.

[0101] Step 2.2: Construct a hierarchical layout generation module. The hierarchical layout generation module is a network model based on transformers, including an image feature extraction network, a region hierarchical layout generation network, a sentence hierarchical layout generation network, and a character hierarchical layout generation network. It is used to extract the features of the textless background image generated in Step 2.1, gradually generate the content of the region, sentence and character hierarchical layout, and finally output the character hierarchical layout.

[0102] Step 2.2.1: The image feature extraction network is an encoder based on ResNet50 fine-tuning;

[0103] The image features are extracted by using a ResNet50-based encoder fine-tuned on the textless background image generated in step 2.1;

[0104] Step 2.2.2, as follows Figure 4As shown, the region-level layout generation network is a network model based on the Transformer architecture. The region-level layout generation network includes two decoder layers. Each decoder layer includes an Embedding layer, a SelfAttention layer, two Layer Norm layers, a Cross Attention layer, a Feed Forward layer, an MLP layer, and a Linear layer. The output of the Embedding layer is connected to the input of the SelfAttention layer. The output of the SelfAttention layer is connected to the input of the first Layer Norm layer. The output of the first Layer Norm layer is connected to the input of the Cross Attention layer. The output of the Cross Attention layer is connected to the input of the second Layer Norm layer. The output of the second Layer Norm layer is connected to the input of the Feed Forward layer. The output of the Feed Forward layer is connected to the input of the MLP layer and the Linear layer, respectively. The MLP layer outputs the region-level layout, and the Linear layer outputs its confidence. The image features output in step 2.2.1 are input into the region-level layout generation network. The region-level layout generation network predicts the layout of the region hierarchy and its confidence by judging the content of different regions of the image features, and outputs the region-level image features and the region-level layout.

[0105] Step 2.2.3, as follows Figure 5 As shown, a sentence hierarchy layout generation network is constructed. The sentence hierarchy layout generation network adds the region-level image features output by the region hierarchy layout generation network in step 2.2.2 to the image features output in step 2.2.1 as input features for encoding. The architecture of the sentence hierarchy layout generation network is the same as that of the region hierarchy layout generation network. By judging the relationship between the content of different feature regions, the layout of the sentence hierarchy and its confidence are predicted, and the sentence hierarchy image features and sentence hierarchy layout are output.

[0106] Step 2.2.4, as follows Figure 5 As shown, the character hierarchy layout generation network adds the sentence hierarchy image features output by the sentence hierarchy layout generation network in step 2.2.3 to the image features output in step 2.2.1 as input features for encoding. The architecture of the character hierarchy layout generation network is the same as that of the region hierarchy layout generation network. By judging the relationship between the content of different feature regions, it predicts the character hierarchy layout and its confidence, and finally outputs the character hierarchy layout.

[0107] Step 2.3: Construct a scene text image generation module, including a controllable graph generation model. Based on the text prompts read in Step 1.3, the character hierarchy layout generated in Step 2.2, and the textless background image generated in Step 2.1, the controllable graph generation model is used to generate a high-quality scene text image that accurately matches the description of the text prompts.

[0108] The controllable graph generation model consists of a VAE-based image encoder, a Transformer-based text encoder, and a Diffusion-based conditional diffusion model.

[0109] Step 2.3.1: Based on the characters and their number in the text prompts read in Step 1.3, further refine the character hierarchy layout generated in Step 2.2 to obtain an independent layout for each character and the characters of the text to be rendered;

[0110] Step 2.3.2: Using the Pillow library, the characters of the text to be rendered obtained in Step 2.3.1 are sequentially placed into the independent layout positions of each character specified in Step 2.3.1. The font file used for placement is Arial Unicode, resulting in a text placement image.

[0111] Step 2.3.3: Input the text prompts read in Step 1.3 into a Transformer-based text encoder to obtain text features. Combine the textless background image generated in Step 2.1.2 and the text placement image obtained in Step 2.3.2 and input them into a VAE-based image encoder for feature extraction to obtain text placement background image features, and add random Gaussian noise. Input the obtained text features and the text placement background image features with added random Gaussian noise into a Diffusion-based conditional diffusion model to generate a scene text image.

[0112] Step 3: Based on the background image generation module, hierarchical layout generation module, and scene text image generation module constructed in Step 2, construct a complete arbitrary shape scene text image generation model;

[0113] The specific method for step 3 is as follows:

[0114] Step 3.1: Connect the text prompt words read in step 1.3 with the background image generation module constructed in step 2.1. The text-generated image model and the scene text detection model in the background image generation module are both loaded with pre-trained model parameters. The output image size of the text-generated image model is the same as the input image size of the scene text detection model. The output result is a textless background image.

[0115] Step 3.2: Input the textless background image output from Step 3.1 into the hierarchical layout generation module constructed in Step 2.2. The image feature extraction network extracts image features by continuously convolving the image upwards. The image features are then input into the region hierarchical layout generation network. Based on the feature values ​​of the input image, the region hierarchical layout generation network continuously predicts the next region hierarchical layout, selects the best layout based on the confidence level, and obtains the region hierarchical image features and the region hierarchical layout. The image features are then added to the region hierarchical image features and input into the sentence hierarchical layout generation network. The same process is repeated to obtain the sentence hierarchical image features and the sentence hierarchical layout. This process continues until the character hierarchical layout is finally obtained, and the final character hierarchical layout is selected based on the best confidence level.

[0116] Step 3.3: The character hierarchy layout obtained in Step 3.2 is segmented and filled according to the characters of the text to be rendered in the text prompts read in Step 1.3 to obtain the text placement image. The text prompts read in Step 1.3, the textless background image generated in Step 2.1, and the text placement image are connected with the scene text image generation module constructed in Step 2.3. After combining the text placement image with the textless background image, the text placement background image is obtained. Feature extraction is performed using the VAE-based image encoder in Step 2.3 to obtain the text placement background image features, and random Gaussian noise is added. Conditional expansion based on Diffusion is then performed. The diffusion model takes the text placement background image features with added random Gaussian noise and the current time step as input, predicts the noise added between the current time step and the previous time step, subtracts the predicted noise from the input image, and obtains the predicted image corresponding to the previous time step. The diffusion model is then used again to take the predicted image, the text placement background image features with added random Gaussian noise, and the current time step as input, predicts the noise added between the current time step and the previous time step, and obtains the predicted image of the previous time step again. This process is repeated until the initial time step is reached, resulting in a complete arbitrary shape scene text image generation model.

[0117] Step 4: Use the preprocessed scene text layout and scene text image generation training sets from Step 1 to train the hierarchical layout generation module and scene text image generation module constructed in Step 2, respectively, to obtain the weight files of the trained hierarchical layout generation module and scene text image generation module.

[0118] The specific method for step 4 is as follows:

[0119] In the network structure corresponding to this invention, only the hierarchical layout generation module and the scene text image generation module need to be trained.

[0120] like Figure 1As shown, in step 4.1, the textless background image in the training background image-hierarchical layout pair after preprocessing in step 1.2 is used as the input of the hierarchical layout generation module constructed in step 2, and the corresponding hierarchical layout is used as the real hierarchical layout.

[0121] The hierarchical layout generation module uses AdamW as the parameter optimizer, with an initial learning rate of 0.0001 and a weight decay of 0.00001 every 200 epochs. The ResNet50 fine-tuned encoder in the image feature extraction network uses a trainable residual network with a learning rate of 0.00001 and an input image size of 512x512 pixels.

[0122] like Figure 2 As shown, in step 4.2, the text prompts, background images without text, and character hierarchy layout content in the training set of the scene text image generation training set after preprocessing in step 1.3 are used as inputs to the scene text image generation module constructed in step 2, and the corresponding scene text images are used as real scene text images.

[0123] The scene text image generation module also selects AdamW as the parameter optimizer, with an initial learning rate of 0.00002 and Betas parameters of (0.5, 0.999).

[0124] Step 4.3: During training, the loss function is used to evaluate the network model performance and optimize the network model parameters. The optimization of the hierarchical layout generation module adopts a hierarchical loss mechanism. The loss functions defined for the region hierarchical layout generation network, the sentence hierarchical layout generation network, and the character hierarchical layout generation network are as follows:

[0125] Step 4.3.1, for training the region-level layout generation network and the sentence-level layout generation network, the following three loss constraints are adopted:

[0126]

[0127]

[0128] Among them, L L1 This represents the L1 loss between the region hierarchy layout output in step 2.2.2 and the statement hierarchy layout generated in step 2.2.3, and the ground truth bounding box, measuring the difference between the predicted bounding box B and the ground truth bounding box. The coordinate difference between them; L GIoU This represents the calculation of the generalized IoU loss between bounding boxes, which extends the IoU calculation to the minimum bounding rectangle C, constraining the generated bounding boxes to be closer in extent to the true bounding boxes; L ol This involves calculating the same overlap loss to avoid multiple bounding boxes overlapping each other, L L1 and LGIoU Both need to be divided by N to distribute their effects evenly across all bounding boxes, while L ol This requires averaging. Among several pairwise combinations of possible bounding boxes;

[0129] The total bounding box loss is:

[0130] L bbox =c1L L1 +c2L GIoU +c3L ol ,

[0131] Where c1 = 5, c2 = 2, c3 = 1;

[0132] Step 4.3.2: For training the character hierarchy layout generation network, calculate the Bézier curve loss using the character hierarchy layout generated in step 2.2.4 and the real layout, constrained by a simple L1 loss:

[0133]

[0134] Step 4.3.3: For the training of the region-level layout generation network, sentence-level layout generation network, and character-level layout generation network, add confidence loss and use standard cross-entropy loss to evaluate the confidence loss of the region-level layout generated in step 2.2.2, the sentence-level layout generated in step 2.2.3, and the character-level layout generated in step 2.2.4.

[0135]

[0136] Where q is the confidence level of the model prediction. The confidence level of the target label (with a value of 1);

[0137] Step 4.3.4: For each decoder layer of the region-level layout generation network, sentence-level layout generation network, and character-level layout generation network, the corresponding loss is calculated, and the Hungarian algorithm is used to calculate the matching loss between the output and the actual layout to obtain the final loss. Combining the total bounding box loss obtained in Step 4.3.1, the Bézier curve loss obtained in Step 4.3.2, and the confidence loss obtained in Step 4.3.3, the final total loss function is:

[0138] L=λ1L bbox +λ2L Bezier +λ3L conf ,

[0139] Where λ1=1, λ2=5, λ3=1;

[0140] Step 4.4, for the scene text image generation module, to quantify the differences between the original image and the predicted image across all text regions, a text-aware loss is used:

[0141]

[0142] Text-aware loss improves the quality of text content in the generated image by minimizing the difference between the predicted image and the original image in the text region through the introduction of a mean squared error constraint. Here, h,w = 512, where 512 represents the height and width of the image, and is the feature map before the fully connected layer. The text writing information at position p in the original and predicted images is represented. Since the time step t is directly related to the text quality of the predicted image x'0, an adjustment function φ(t) is used to dynamically adjust the weights of the loss function, where... These are coefficients in the diffusion process;

[0143] Finally, the weight files for the trained hierarchical layout generation module and scene text image generation module are obtained.

[0144] like Figure 3 As shown, in step 5, a test set is generated based on the scene text image preprocessed in step 1. Using the weight files of the hierarchical layout generation module and the scene text image generation module trained in step 4, the final scene text image is obtained by using the complete arbitrary shape scene text image generation model constructed in step 3 and performing model inference.

[0145] The specific method for step 5 is as follows:

[0146] Step 5.1 Load the complete arbitrary-shape scene text image generation model constructed in Step 3, read the weight files of the hierarchical layout generation module and the scene text image generation module trained in Step 4, load the weight files into the structure of the hierarchical layout generation module and the scene text image generation module, and set the complete arbitrary-shape scene text image generation model to inference mode and fix the model parameters.

[0147] Step 5.2 Input the text prompts from the preprocessed scene text image generation test set in Step 1.4 into the complete arbitrary shape scene text image generation model constructed in Step 3. Use the background image generation module to modify the text prompts and perform text detection on the generated background image to obtain a textless background image. Use the hierarchical layout generation module to encode the textless background image, extract features, and generate a hierarchical character layout. Use the scene text image generation module to segment and fill the character layout, and combine it with the textless background image to obtain a text placement background image. After encoding the text placement background image, obtain the text placement background image features and add random Gaussian noise. Encode the text prompts to obtain text features. Use the text placement background image features with added random Gaussian noise and the text features as input to the conditional diffusion model to gradually diffuse and generate the final scene text image.

[0148] The effects of the present invention will be further illustrated below with simulation experiments.

[0149] Simulation conditions

[0150] The hardware platform for the simulation experiment of this invention is: Intel i9-10900X processor and NVIDIA GeForce RTX 4090 GPU.

[0151] The software platform for the simulation experiment of this invention is: Ubuntu 20.04 operating system, Python 3.9.0 programming language, PyTorch 2.2.1 deep learning framework, and Visual Studio Code code editor.

[0152] Simulation content

[0153] The training background image-hierarchical layout pair preprocessed in step 1.2 and the scene text image generation training set preprocessed in step 1.4 are used to train the hierarchical layout generation module and scene text image generation module constructed in step 2, respectively, to obtain the weight files of the trained hierarchical layout generation module and scene text image generation module. The complete arbitrary shape scene text image generation model constructed in step 3 is loaded, and the weight files of the trained hierarchical layout generation module and scene text image generation module are read and loaded into the structure of the hierarchical layout generation module and scene text image generation module. The complete arbitrary shape scene text image generation model is set to inference mode and the model parameters are fixed. The specified prompt word is input into the complete arbitrary shape scene text image generation model for processing, and the text layout and scene text image are generated to obtain the output layout and output image.

[0154] Simulation results

[0155] Two existing methods compared to the results obtained in this invention are GlyphControl and TextDiffuser. For GlyphControl, TextDiffuser, and the method of this invention, the same prompt word is input, with GlyphControl requiring additional glyph input. The output layouts of TextDiffuser and the method of this invention, as well as the output images of GlyphControl, TextDiffuser, and the method of this invention, are compared. Specific results are as follows: Figure 6 As shown, when inputting a prompt word, the GlyphControl method in the figure needs to rely on the input glyph to assist in the output image; the TextDiffuser method in the figure can output a layout and assist in the output image based on the output layout, but it can only generate an output image with a horizontal straight text layout; while the present invention can generate output images with various text layouts such as horizontal and curved while outputting the layout, and the result obtained is superior to the other two methods in terms of layout diversity and scene text image rendering quality.

[0156] The present invention also provides a hierarchical layout-driven arbitrary shape scene text image generation system, comprising:

[0157] The scene text image training set preprocessing module is used to preprocess the scene text image training set in step 1. The scene text image training set includes a scene text layout generation training set, a scene text image generation training set, and a scene text image generation test set. Finally, the preprocessed scene text layout generation training set, scene text image generation training set, and scene text image generation test set are obtained.

[0158] The first module is used to implement the background image generation module, hierarchical layout generation module and scene text image generation module in step 2;

[0159] The arbitrary-shape scene text image generation model construction module is used to implement the background image generation module, hierarchical layout generation module and scene text image generation module built in step 3 based on step 2, and to build a complete arbitrary-shape scene text image generation model.

[0160] The first module is used to generate training sets for scene text layout and scene text image in step 4 using the preprocessed scene text layout in step 1, and to train the hierarchical layout generation module and scene text image generation module constructed in step 2 respectively, so as to obtain the weight files of the trained hierarchical layout generation module and scene text image generation module.

[0161] The model inference module is used to generate a test set based on the scene text image preprocessed in step 1 in step 5. It uses the weight files of the hierarchical layout generation module and the scene text image generation module trained in step 4 to perform model inference through the complete arbitrary shape scene text image generation model constructed in step 3, and obtains the final scene text image.

[0162] The present invention also provides a hierarchical layout-driven arbitrary shape scene text image generation device, comprising:

[0163] Memory: A computer program that stores the above-described hierarchical layout-driven method for generating text images of arbitrary-shaped scenes, and is a computer-readable device;

[0164] Processor: Used to implement the hierarchical layout-driven arbitrary shape scene text image generation method when executing the computer program.

[0165] The present invention also provides a computer-readable storage medium storing a computer program that, when executed by a processor, enables the implementation of the hierarchical layout-driven arbitrary shape scene text image generation method.

Claims

1. A method for generating text images of arbitrary-shaped scenes driven by hierarchical layout, characterized in that, Includes the following steps: Step 1: Preprocess the scene text image training set. The scene text image training set includes a scene text layout generation training set, a scene text image generation training set, and a scene text image generation test set. Finally, the preprocessed scene text layout generation training set, scene text image generation training set, and scene text image generation test set are obtained. The scene text image generation training set includes text prompts, images without text backgrounds, character hierarchical layout content, and scene text images. Step 1.1: The scene text layout generation training set includes a textless background image and content containing region, sentence, and character hierarchical layouts. The textless background image and content containing region, sentence, and character hierarchical layouts in the scene text layout generation training set are used to form training background image-hierarchical layout pairs. The region and sentence hierarchical layouts consist of 4 values, with a total of 2 control points, which respectively represent the horizontal and vertical coordinates of the top left and bottom right corners of the layout. The character hierarchical layout consists of 16 values, which respectively represent the horizontal and vertical coordinates of the 4 control points of the top Bézier curve and the 4 control points of the bottom Bézier curve of the layout. Step 2: Construct a background image generation module, a hierarchical layout generation module, and a scene text image generation module; Step 3: Based on the background image generation module, hierarchical layout generation module, and scene text image generation module constructed in Step 2, construct a complete arbitrary shape scene text image generation model; Step 4: Use the preprocessed scene text layout generation training set and scene text image generation training set from Step 1 to train the hierarchical layout generation module and scene text image generation module constructed in Step 2, respectively, to obtain the weight files of the trained hierarchical layout generation module and scene text image generation module. Step 5: Based on the scene text image preprocessed in Step 1, generate a test set. Using the weight files of the hierarchical layout generation module and the scene text image generation module trained in Step 4, perform model inference through the complete arbitrary shape scene text image generation model constructed in Step 3 to obtain the final scene text image.

2. The method for generating arbitrary-shaped scene text images driven by hierarchical layout according to claim 1, characterized in that, Step 1.1 is followed by: Step 1.2 involves preprocessing the training background image-hierarchical layout pair from Step 1.

1. The textless background image in the training background image-hierarchical layout pair is read, resized to 3×512×512, and converted to tensor format. Simultaneously, the mean and variance are used to normalize the three color channels of the textless background image. The input textless background image color values ​​range from integers [0~255] to floating-point values ​​[-1~1]. The json library is used to read the three hierarchical layouts corresponding to the textless background image in Step 1.1: region, sentence, and character. The horizontal and vertical coordinate values ​​range from floating-point values ​​[0~511] to floating-point values ​​[0~1], resulting in the preprocessed training background image-hierarchical layout pair. Step 1.3: The scene text image generation training set includes text prompts, images without text backgrounds, character hierarchy layout content, and scene text images. Preprocessing of the scene text image generation training set specifically involves: processing a single image without text background using the same process as in Step 1.2; and reading the text prompts and character hierarchy layout content corresponding to a single image without text background using a json library to obtain the preprocessed scene text image generation training set. Step 1.4: The scene text image generation test set includes text prompts and scene text images; the scene text image generation test set is preprocessed, specifically by reading the text prompts using a json library to obtain the preprocessed scene text image generation test set.

3. The method for generating arbitrary-shaped scene text images driven by hierarchical layout according to claim 1, characterized in that, The specific method for step 2 is as follows: Step 2.1: Construct a background image generation module, including a text-generated image model and a scene text detection model, used to generate a textless background image based on the text prompts read in Step 1.3; Step 2.1.1: Remove the text content to be rendered from the text prompts read in Step 1.3, and generate a background image using the text-generated image model; Step 2.1.2: Use the scene text detection model to detect the background image generated by the text-generated image model in step 2.1.

1. If text is detected, repeat steps 2.1.1-2.1.2 until no text is detected, and output a text-free background image. Step 2.2: Construct a hierarchical layout generation module. The hierarchical layout generation module is a network model based on transformers, including an image feature extraction network, a region hierarchical layout generation network, a sentence hierarchical layout generation network, and a character hierarchical layout generation network. It is used to extract the features of the textless background image generated in Step 2.1, gradually generate the content of the region, sentence and character hierarchical layout, and finally output the character hierarchical layout. Step 2.2.1: The image feature extraction network is an encoder based on ResNet50 fine-tuning; The image features are extracted by using a ResNet50-based encoder fine-tuned on the textless background image generated in step 2.1; Step 2.2.2: The region-level layout generation network is a network model based on the Transformer architecture. The region-level layout generation network includes two decoder layers. Each decoder layer includes an Embedding layer, a SelfAttention layer, two Layer Norm layers, a Cross Attention layer, a Feed Forward layer, an MLP layer, and a Linear layer. The output of the Embedding layer is connected to the input of the Self Attention layer. The output of the SelfAttention layer is connected to the input of the first Layer Norm layer. The output of the first Layer Norm layer is connected to the input of the Cross Attention layer. The output of the Cross Attention layer is connected to the input of the second Layer Norm layer. The output of the second Layer Norm layer is connected to the input of the Feed Forward layer. The output of the Feed Forward layer is connected to the input of the MLP layer and the Linear layer, respectively. The MLP layer outputs the region-level layout, and the Linear layer outputs its confidence score. The image features output in Step 2.2.1 are input into the region-level layout generation network. The region-level layout generation network predicts the region-level layout and its confidence score by judging the content of different regions of the image features, and outputs the region-level image features and the region-level layout. Step 2.2.3: Construct a sentence hierarchy layout generation network. The sentence hierarchy layout generation network adds the region-level image features output by the region hierarchy layout generation network in Step 2.2.2 to the image features output in Step 2.2.1 as input features for encoding. The architecture of the sentence hierarchy layout generation network is the same as that of the region hierarchy layout generation network. By judging the relationship between the content of different feature regions, it predicts the layout of the sentence hierarchy and its confidence, and outputs the sentence hierarchy image features and sentence hierarchy layout. Step 2.2.4: The character hierarchy layout generation network adds the sentence hierarchy image features output by the sentence hierarchy layout generation network in step 2.2.3 to the image features output in step 2.2.1 as input features for encoding. The architecture of the character hierarchy layout generation network is the same as that of the region hierarchy layout generation network. By judging the relationship between the content of different feature regions, it predicts the character hierarchy layout and its confidence, and finally outputs the character hierarchy layout. Step 2.3: Construct a scene text image generation module, including a controllable graph generation model. Based on the text prompts read in Step 1.3, the character hierarchy layout generated in Step 2.2, and the textless background image generated in Step 2.1, the controllable graph generation model is used to generate scene text images that accurately match the description of the text prompts. The controllable graph generation model consists of a VAE-based image encoder, a Transformer-based text encoder, and a Diffusion-based conditional diffusion model. Step 2.3.1: Based on the characters and their number in the text prompts read in Step 1.3, further refine the character hierarchy layout generated in Step 2.2 to obtain an independent layout for each character and the characters of the text to be rendered; Step 2.3.2: Place the characters of the text to be rendered obtained in Step 2.3.1 into the independent layout position of each character specified in Step 2.3.

1. The font file used for placement is Arial Unicode, and the text placement image is obtained. Step 2.3.3: Input the text prompts read in Step 1.3 into a Transformer-based text encoder to obtain text features. Combine the textless background image generated in Step 2.1.2 and the text placement image obtained in Step 2.3.2 and input them into a VAE-based image encoder for feature extraction to obtain text placement background image features, and add random Gaussian noise. Input the obtained text features and the text placement background image features with added random Gaussian noise into a Diffusion-based conditional diffusion model to generate a scene text image.

4. The method for generating arbitrary-shaped scene text images driven by hierarchical layout according to claim 1, characterized in that, The specific method for step 3 is as follows: Step 3.1: Connect the text prompt words read in step 1.3 with the background image generation module constructed in step 2.

1. The text-generated image model and the scene text detection model in the background image generation module are both loaded with pre-trained model parameters. The output image size of the text-generated image model is the same as the input image size of the scene text detection model. The output result is a textless background image. Step 3.2: Input the textless background image output from Step 3.1 into the hierarchical layout generation module constructed in Step 2.

2. The image feature extraction network extracts image features by continuously convolving the image upwards. The image features are then input into the region hierarchical layout generation network. Based on the feature values ​​of the input image, the region hierarchical layout generation network continuously predicts the next region hierarchical layout, selects the best layout based on the confidence level, and obtains the region hierarchical image features and the region hierarchical layout. The image features are then added to the region hierarchical image features and input into the sentence hierarchical layout generation network. The same process is repeated to obtain the sentence hierarchical image features and the sentence hierarchical layout. This process continues until the character hierarchical layout is finally obtained, and the final character hierarchical layout is selected based on the best confidence level. Step 3.3: The character hierarchy layout obtained in Step 3.2 is segmented and filled according to the characters of the text to be rendered in the text prompts read in Step 1.3 to obtain the text placement image. The text prompts read in Step 1.3, the textless background image generated in Step 2.1, and the text placement image are connected with the scene text image generation module constructed in Step 2.

3. After combining the text placement image with the textless background image, the text placement background image is obtained. Feature extraction is performed using the VAE-based image encoder in Step 2.3 to obtain the text placement background image features, and random Gaussian noise is added. Conditional expansion based on Diffusion is then performed. The diffusion model takes the text placement background image features with added random Gaussian noise and the current time step as input, predicts the noise added between the current time step and the previous time step, subtracts the predicted noise from the input image, and obtains the predicted image corresponding to the previous time step. The diffusion model is then used again to take the predicted image, the text placement background image features with added random Gaussian noise, and the current time step as input, predicts the noise added between the current time step and the previous time step, and obtains the predicted image of the previous time step again. This process is repeated until the initial time step is reached, resulting in a complete arbitrary shape scene text image generation model.

5. The method for generating arbitrary-shaped scene text images driven by hierarchical layout according to claim 1, characterized in that, The specific method for step 4 is as follows: Step 4.1: Take the textless background image from the training background image-hierarchical layout pair after preprocessing in Step 1.2 as the input to the hierarchical layout generation module constructed in Step 2, and take the corresponding hierarchical layout as the real hierarchical layout. Step 4.2: Take the text prompts, textless background images, and character hierarchy layout content from the training set of the scene text image generation training set after preprocessing in Step 1.3 as input to the scene text image generation module constructed in Step 2, and take the corresponding scene text images as real scene text images. Step 4.3: During training, the loss function is used to evaluate the network model performance and optimize the network model parameters. The optimization of the hierarchical layout generation module adopts a hierarchical loss mechanism. The loss functions defined for the region hierarchical layout generation network, the sentence hierarchical layout generation network, and the character hierarchical layout generation network are as follows: Step 4.3.1, for training the region-level layout generation network and the sentence-level layout generation network, the following three loss constraints are adopted: Among them, L L1 This represents the L1 loss between the region-level layout output in step 2.2.2 and the statement-level layout output in step 2.2.3, and the ground truth bounding box, measuring the difference between the predicted bounding box B and the ground truth bounding box. The coordinate difference between them; L GIoU This represents the calculation of the generalized IoU loss between bounding boxes, which extends the IoU calculation to the minimum bounding rectangle C, constraining the generated bounding boxes to be closer in extent to the true bounding boxes; L ol This involves calculating the same overlap loss, L. L1 and L GIoU Divide all by N to distribute the effect evenly across all bounding boxes, while L ol That is, averaged to Among several pairwise combinations of possible bounding boxes; The total bounding box loss is: L bbox =c1L L1 +c2L GIoU +c3L ol , Where c1 = 5, c2 = 2, c3 = 1; Step 4.3.2: For training the character hierarchy layout generation network, calculate the Bézier curve loss using the character hierarchy layout generated in Step 2.2 and the real layout, constrained by L1 loss. Step 4.3.3: For training the region-level layout generation network, sentence-level layout generation network, and character-level layout generation network, add confidence loss and use standard cross-entropy loss to evaluate the confidence loss for each bounding box and Bézier curve: Where q is the confidence level of the model prediction. The confidence level of the target label (with a value of 1); Step 4.3.4: For each decoder layer of the region-level layout generation network, sentence-level layout generation network, and character-level layout generation network, the corresponding loss is calculated, and the matching loss between the output and the actual layout is calculated using the Hungarian algorithm to obtain the final loss. Combining the total bounding box loss obtained in Step 4.3.1, the Bézier curve loss obtained in Step 4.3.2, and the confidence loss obtained in Step 4.3.3, the final total loss function is: L=λ1L bbox +λ2L Bezier +λ3L conf , Where λ1=1, λ2=5, λ3=1; Step 4.4, for the scene text image generation module, to quantify the differences between the original image and the predicted image across all text regions, a text-aware loss is used: Where h,w = 512, 512 represents the height and width of the image, and here it represents the feature map before the fully connected layer. The text writing information at position p in the original and predicted images is represented. Since the time step t is directly related to the text quality of the predicted image x'0, an adjustment function φ(t) is used to dynamically adjust the weights of the loss function, where... These are coefficients in the diffusion process; Finally, the weight files for the trained hierarchical layout generation module and scene text image generation module are obtained.

6. The method for generating arbitrary-shaped scene text images driven by hierarchical layout according to claim 1, characterized in that, The specific method for step 5 is as follows: Step 5.1 Load the complete arbitrary-shape scene text image generation model constructed in Step 3, read the weight files of the hierarchical layout generation module and the scene text image generation module trained in Step 4, load the weight files into the structure of the hierarchical layout generation module and the scene text image generation module, and set the complete arbitrary-shape scene text image generation model to inference mode and fix the model parameters. Step 5.2: Input the text prompts from the preprocessed scene text image generation test set in Step 1.4 into the complete arbitrary shape scene text image generation model constructed in Step 3. Use the background image generation module to modify the text prompts and perform text detection on the generated background image to obtain a textless background image. Use the hierarchical layout generation module to encode the textless background image, extract features, and generate a hierarchical character layout. Use the scene text image generation module to segment and fill the character layout, and combine it with the textless background image to obtain a text placement background image. After encoding the text placement background image, obtain the text placement background image features and add random Gaussian noise. Encode the text prompts to obtain text features. Use the text placement background image features with added random Gaussian noise and the text features as input to the conditional diffusion model to gradually diffuse and generate the final scene text image.

7. A hierarchical layout-driven arbitrary shape scene text image generation system based on the method of any one of claims 1 to 6, characterized in that, include: The scene text image training set preprocessing module is used to preprocess the scene text image training set, which includes a scene text layout generation training set, a scene text image generation training set, and a scene text image generation test set. Finally, the preprocessed scene text layout generation training set, scene text image generation training set, and scene text image generation test set are obtained. The first module is used to build a background image generation module, a hierarchical layout generation module, and a scene text image generation module. The arbitrary-shape scene text image generation model building module is used to build a complete arbitrary-shape scene text image generation model based on the background image generation module, the hierarchical layout generation module, and the scene text image generation module. The first module is used to generate training sets for the preprocessed scene text layout and scene text image, and to train the constructed hierarchical layout generation module and scene text image generation module respectively, so as to obtain the weight files of the trained hierarchical layout generation module and scene text image generation module. The model inference module is used to generate a test set based on scene text images. It uses the weight files of the trained hierarchical layout generation module and scene text image generation module to construct a complete arbitrary-shape scene text image generation model and perform model inference to obtain the final scene text image.

8. A hierarchical layout-driven arbitrary shape scene text image generation device, characterized in that, include: Memory: A computer program for generating text images of arbitrary-shaped scenes driven by hierarchical layout as described in any one of claims 1-6, which is a computer-readable device; Processor: configured to implement, when executing the computer program, a hierarchical layout-driven arbitrary shape scene text image generation method as described in any one of claims 1-6.

9. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program that, when executed by a processor, enables the generation of a hierarchical layout-driven arbitrary-shape scene text image as described in any one of claims 1-6.

Citation Information

Patent Citations

  • Method for training a text-graph model and text-graph method

    CN116935169B

  • Image generation method and device, equipment and medium

    CN116977774A

  • Target Chinese poster generation method and device, computer equipment and storage medium

    CN118397148A

  • End-to-end recognition method for scene text in any shape

    WO2019192397A1