Hierarchical layout driven random shape scene text image generation method, system and device and medium

By building a hierarchical layout-driven text image generation model for arbitrary shape scenes, the problem of difficulty in generating high-quality text images in arbitrary shape is solved in the existing technology, and the generation of layouts that do not rely on user input is realized. The generated image visual effects are more realistic and support diverse text styles.

CN120070666AActive Publication Date: 2025-05-30XIDIAN UNIV +1
View PDF 9 Cites 0 Cited by

Patent Information

Application Number
CN202510126918.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-27
Publication Date
2025-05-30
Estimated Expiration
2045-01-27

AI Technical Summary

Technical Problem

Existing scene text image generation methods are difficult to generate high-quality text images of any shape, especially when irregular arrangements such as arcs, curves or italics, which show insufficient adaptability, resulting in the lack of visual aesthetics and overall coordination between the generated results.

Method used

By constructing a hierarchical layout-driven arbitrary-shaped scene text image generation model, generating background images, perceive content information of different scales in the background image, gradually generating layouts of area levels, statement levels and character levels, and finally generating arbitrary-shaped scene text images. This model does not rely on user input layout, and achieves more refined geometric and directional control through multi-layer Transformer architecture and Bezier curve characterization.

Benefits of technology

It realizes the generation of text images in any shape that does not depend on user input layout. The generated text is naturally integrated with the background, and the visual effect is more realistic. It supports diverse text style generation, which improves the ease of use and practicality of the technology.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120070666A_ABST
    Figure CN120070666A_ABST
Patent Text Reader

Abstract

The invention discloses a hierarchical layout-driven random-shape scene text image generation method, system and device and a medium, and the method comprises the steps: carrying out the preprocessing of a scene text image training set, and obtaining a preprocessed scene text layout generation training set, a scene text image generation training set and a scene text image generation test set; constructing a background image generation module, a hierarchical layout generation module and a scene text image generation module; constructing a complete random-shape scene text image generation model; respectively training the hierarchical layout generation module and the scene text image generation module to obtain weight files of the trained hierarchical layout generation module and the trained scene text image generation module; performing model reasoning to obtain a final scene text image; the system, the equipment and the medium are used for implementing the method. According to the method, scene text images in any shapes can be automatically generated without depending on user input layout.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the cross - technical field of computer vision and computer graphics, and particularly relates to a method, system, device and medium for generating arbitrary - shaped scene text images driven by hierarchical layout. Background Technique

[0002] The technology of generating scene text images has broad application prospects in fields such as advertising design, game development, and data synthesis. Its research significance lies in expanding the technical boundaries of the combination of language and vision and solving the key bottlenecks in text - to - image generation in complex scenes. This technology aims to generate images with semantic coherence and visual beauty, involving in - depth coordination of text semantic expression, visual layout, font style, and background integration. In advertising design, this technology can achieve creative text layouts that match the scene, significantly improving design efficiency; in game development, it can generate background text that meets the requirements of virtual scenes, shortening the development cycle; in data synthesis, it can generate text images with complex layouts on a large scale, providing high - quality training data for text detection and recognition models. However, there are many challenges in its implementation process, such as the need to balance global semantics and local detail expression, adapt to complex and irregular text arrangements, and ensure the unity and coordination of various visual features such as fonts, colors, and backgrounds.

[0003] Existing methods for generating scene text images have made some progress in generating text with regular layouts, but there are still significant technical bottlenecks for the task of generating text with arbitrary shapes. Most current mainstream methods use simple bounding - box layout modeling, which is difficult to meet the requirements of arbitrary - shaped text layouts in practical applications. Especially when generating non - regular arrangements such as arcs, curves, or italics, it shows obvious insufficient adaptability. This modeling method not only limits the flexibility of generating scene text images but also results in the generated results lacking visual beauty and overall coordination with the scene, making it difficult to meet the requirements of expressing complex text forms.

[0004] The patent application with the publication number CN116935169B discloses a training method and an application method for a text - to - image model, belonging to the field of artificial intelligence technology. By introducing differentiable diversity constraints during the training process, this method significantly enhances the ability of the text - to - image model to generate diverse images. The diversity constraints play an important role in the model training and parameter tuning stages, enabling the model to have the ability to generate diverse images. However, the optimization objective of this method mainly focuses on improving image diversity and does not conduct targeted optimization for generating high - quality and high - fidelity scene text images, resulting in low - quality generated scene text images and making it difficult to meet the application requirements for high - precision expression of text content.

[0005] The patent application with the publication number CN116977774A discloses an image generation method, apparatus, device, and storage medium, which relates to the field of artificial intelligence technology. This method takes a scene description text as input, predicts the corresponding scene layout information by extracting the semantic features of the text, and uses this layout information to represent the relative position relationship between various objects in the scene. By fusing the layout information into the scene description text to generate a target description text, and combining it with an initial noise image for noise reduction processing, a high-precision image that meets the target scene layout is generated. This method effectively improves the accuracy of the scene content in the generated image. However, during the generation process, only the layout of the scene objects is optimized, and the layout information of the scene text content is not introduced, so the semantic features of the scene text cannot be highlighted. In addition, the generated scene text content lacks fidelity and is difficult to meet the actual needs of complex scene text image generation. Summary of the Invention

[0006] In order to overcome the above-mentioned deficiencies of the prior art, the purpose of the present invention is to provide a method, system, device, and medium for generating arbitrary-shaped scene text images driven by hierarchical layout. By generating a background image, perceiving different-scale content information in the background image, generating three different hierarchical layouts: regional hierarchical layout, sentence hierarchical layout, and character hierarchical layout, and finally being able to generate arbitrary-shaped scene text images according to the character hierarchical layout. The present invention does not rely on user input layout, and the generated text has diversity and authenticity.

[0007] In order to achieve the above purpose, the technical solution adopted by the present invention is:

[0008] A method for generating arbitrary-shaped scene text images driven by hierarchical layout includes the following steps:

[0009] Step 1, preprocess the scene text image training set, which includes a scene text layout generation training set, a scene text image generation training set, and a scene text image generation test set, and finally obtain the preprocessed scene text layout generation training set, scene text image generation training set, and scene text image generation test set;

[0010] Step 2, construct a background image generation module, a hierarchical layout generation module, and a scene text image generation module;

[0011] Step 3, based on the background image generation module, hierarchical layout generation module, and scene text image generation module constructed in Step 2, construct a complete arbitrary-shaped scene text image generation model;

[0012] Step 4: Use the preprocessed scene text layout in Step 1 to generate a training set and a training set for scene text image generation. Respectively train the hierarchical layout generation module and the scene text image generation module constructed in Step 2 to obtain the weight files of the trained hierarchical layout generation module and the scene text image generation module.

[0013] Step 5: Based on the scene text images preprocessed in Step 1, generate a test set. Use the weight files of the trained hierarchical layout generation module and the scene text image generation module in Step 4, and pass through the complete arbitrary-shaped scene text image generation model constructed in Step 3 for model inference to obtain the final scene text images.

[0014] The specific method of Step 1 is as follows:

[0015] Step 1.1: The training set for scene text layout generation includes textless background images and the content of the hierarchical layout of regions, sentences, and characters. The textless background images and the content of the hierarchical layout of regions, sentences, and characters in the training set for scene text layout generation are used to form training background image - hierarchical layout pairs. Among them, the hierarchical layout of regions and sentences consists of 4 values, with a total of 2 control points, which respectively represent the horizontal and vertical coordinates of the upper left corner and the lower right corner of the layout. The hierarchical layout of characters consists of 16 values, which respectively represent the horizontal and vertical coordinates of the 4 control points of the top Bezier curve and the horizontal and vertical coordinates of the 4 control points of the bottom Bezier curve of the layout.

[0016] Step 1.2: Preprocess the training background image - hierarchical layout pairs in Step 1.1. Read the textless background images in the training background image - hierarchical layout pairs, adjust the textless background images to 3×512×512 and convert them to the tensor format. At the same time, normalize the three color channels of the textless background images using the mean and variance respectively, and map the input color values of the textless background images in the range of integer values from [0~255] to floating-point values in the range of [-1~1]. Read the hierarchical layouts of regions, sentences, and characters corresponding to the textless background images in the training background image - hierarchical layout pairs in Step 1.1 through the json library, and map the horizontal and vertical coordinate values from floating-point values in the range of [0~511] to floating-point values in the range of [0~1] to obtain the preprocessed training background image - hierarchical layout pairs.

[0017] Step 1.3: The training set for scene text image generation includes text prompts, textless background images, the content of the hierarchical layout of characters, and scene text images. The preprocessing of the training set for scene text image generation is specifically as follows: Perform the same process as in Step 1.2 on a single textless background image. Read the text prompts and the content of the hierarchical layout of characters corresponding to a single textless background image through the json library to obtain the preprocessed training set for scene text image generation.

[0018] Step 1.4, the scene text image generation test set includes text prompts and scene text images; preprocess the scene text image generation test set, specifically: read the text prompts through the json library to obtain the preprocessed scene text image generation test set.

[0019] The specific method of Step 2 is as follows:

[0020] Step 2.1, construct a background image generation module, including a text-to-image model and a scene text detection model, for generating a text-free background image through the text prompts read in Step 1.3;

[0021] Step 2.1.1, remove the text content to be rendered in the text prompts read in Step 1.3, so that it generates a background image through the text-to-image model;

[0022] Step 2.1.2, use the scene text detection model to detect the background image generated by the text-to-image model in Step 2.1.1. If text is detected, repeat Steps 2.1.1 - 2.1.2 until no text is detected, and output a text-free background image;

[0023] Step 2.2, construct a hierarchical layout generation module. The hierarchical layout generation module is a network model based on the transformer, including an image feature extraction network, a regional hierarchical layout generation network, a sentence hierarchical layout generation network, and a character hierarchical layout generation network, for extracting the features of the text-free background image generated in Step 2.1, gradually generating the content of the regional, sentence, and character hierarchical layouts, and finally outputting the character hierarchical layout;

[0024] Step 2.2.1, the image feature extraction network is an encoder fine-tuned based on ResNet50;

[0025] Extract the features of the text-free background image generated in Step 2.1 by using the encoder fine-tuned based on ResNet50, and output the image features;

[0026] Step 2.2.2, The region-level layout generation network is a network model based on the Transformer architecture. The region-level layout generation network includes two decoder layers. Each decoder layer includes an Embedding layer, a SelfAttention layer, two Layer Norm layers, a Cross Attention layer, a FeedForward layer, an MLP layer, and a Linear layer. The output of the Embedding layer is connected to the input of the SelfAttention layer. The output of the SelfAttention layer is connected to the input of the first Layer Norm layer. The output of the first LayerNorm layer is connected to the input of the Cross Attention layer. The output of the Cross Attention layer is connected to the input of the second LayerNorm layer. The output of the second Layer Norm layer is connected to the input of the Feed Forward layer. The output of the Feed Forward layer is respectively connected to the inputs of the MLP layer and the Linear layer. Among them, the MLP layer outputs the region-level layout, and the Linear layer outputs its confidence. The image features output in Step 2.2.1 are input into the region-level layout generation network. The region-level layout generation network predicts the layout of the region level and its confidence by judging the content of different regions of the image features, and outputs the region-level image features and the layout of the region level;

[0027] Step 2.2.3, Construct a sentence-level layout generation network. The sentence-level layout generation network takes the sum of the region-level image features output by the region-level layout generation network in Step 2.2.2 and the image features output in Step 2.2.1 as input features for encoding. The architecture of the sentence-level layout generation network is the same as that of the region-level layout generation network. By judging the connection between the content of different regions of the features, it predicts the layout of the sentence level and its confidence, and outputs the sentence-level image features and the sentence-level layout;

[0028] Step 2.2.4, The character-level layout generation network takes the sum of the sentence-level image features output by the sentence-level layout generation network in Step 2.2.3 and the image features output in Step 2.2.1 as input features for encoding. The architecture of the character-level layout generation network is the same as that of the region-level layout generation network. By judging the connection between the content of different regions of the features, it predicts the layout of the character level and its confidence, and finally outputs the character-level layout;

[0029] Step 2.3, construct a scene text image generation module, including a controllable image-to-image generation model. According to the text prompt read in Step 1.3, the character-level layout generated in Step 2.2, and the textless background image generated in Step 2.1, use the controllable image-to-image generation model to generate a scene text image that exactly conforms to the description of the text prompt;

[0030] The controllable image-to-image generation model consists of a VAE-based image encoder, a Transformer-based text encoder, and a Diffusion-based conditional diffusion model;

[0031] Step 2.3.1, according to the characters and their quantities of the text to be rendered in the text prompt read in Step 1.3, further refine and segment the character-level layout generated in Step 2.2 to obtain the independent layout of each character and the characters of the text to be rendered;

[0032] Step 2.3.2, sequentially place the characters of the text to be rendered obtained in Step 2.3.1 at the independent layout positions of each character specified in Step 2.3.1. The font file used for placement is Arial Unicode to obtain a text-placement image;

[0033] Step 2.3.3, input the text prompt read in Step 1.4 into the Transformer-based text encoder to obtain text features. Combine the textless background image generated in Step 2.1.2 and the text-placement image obtained in Step 2.3.2, input them into the VAE-based image encoder for feature extraction to obtain text-placement background image features, and add random Gaussian noise; input the obtained text features and the text-placement background image features after adding random Gaussian noise into the Diffusion-based conditional diffusion model to generate a scene text image.

[0034] The specific method of Step 3 is as follows:

[0035] Step 3.1, connect the text prompt read in Step 1.3 with the background image generation module constructed in Step 2. The text-to-image generation model and the scene text detection model in the background image generation module are both loaded with pre-trained model parameters. The output image size of the text-to-image generation model is the same as the input image size of the scene text detection model, and the output result is a textless background image;

[0036] Step 3.2, input the textless background image output in Step 3.1 into the hierarchical layout generation module constructed in Step 2.2. The image feature extraction network continuously performs upward convolution on the image to extract image features; input the image features into the regional hierarchical layout generation network. The regional hierarchical layout generation network continuously predicts the next regional hierarchical layout according to the feature values of the input image, selects the best layout based on the confidence, and obtains the regional hierarchical image features and the layout at the regional level; after adding the image features to the regional hierarchical image features, input them into the statement-level layout generation network. In the same process, obtain the statement-level image features and the statement-level layout, and so on, finally obtain the character-level layout, and select the final character-level layout according to the best confidence;

[0037] Step 3.3, segment and fill the character-level layout obtained in Step 3.2 according to the characters of the text to be rendered in the text prompt read in Step 1.3 to obtain the text placement image. Connect the text prompt read in Step 1.3, the textless background image generated in Step 2.1, and the text placement image with the scene text image generation module constructed in Step 2.3. After combining the text placement image and the textless background image, obtain the text placement background image. Perform feature extraction through the VAE-based image encoder in Step 2.3 to obtain the text placement background image features, and add random Gaussian noise; the conditional diffusion model based on Diffusion takes the text placement background image features added with random Gaussian noise and the current time step as inputs, predicts the noise added between the current time step and the previous time step, subtracts the predicted noise from the input image to obtain the predicted image corresponding to the previous time step. Again, use this conditional diffusion model based on Diffusion to take the predicted image, the text placement background image features added with random Gaussian noise, and the current time step as inputs, predict the noise added between the current time step and the previous time step, and obtain the predicted image of the previous time step again, and so on, iterate to the initial time step to obtain a complete arbitrary-shaped scene text image generation model.

[0038] The specific method of Step 4 is as follows:

[0039] Step 4.1, use the textless background image in the training background image-hierarchical layout pair preprocessed in Step 1.2 as the input of the hierarchical layout generation module constructed in Step 2, and the corresponding hierarchical layout as the true hierarchical layout;

[0040] Step 4.2, use the text prompt, the textless background image, and the character-level layout content in the scene text image generation training set preprocessed in Step 1.3 as the inputs of the scene text image generation module constructed in Step 2, and the corresponding scene text image as the true scene text image;

[0041] Step 4.3. During the training process, the loss function is used to evaluate the performance of the network model and optimize the network model parameters. The optimization of the hierarchical layout generation module adopts a hierarchical loss mechanism. For the region-level layout generation network, the sentence-level layout generation network, and the character-level layout generation network, the defined loss functions are as follows:

[0042] Step 4.3.1. For the training of the region-level layout generation network and the sentence-level layout generation network, the following three-part loss constraints are adopted:

[0043]

[0044] Among them, L L1 represents calculating the L1 loss between the region-level layout output in Step 2.2.2, the sentence-level layout output in Step 2.2.3, and the true bounding box, measuring the coordinate gap between the predicted bounding box B and the true bounding box ; L GIoU represents calculating the generalized IoU loss between the bounding boxes, which extends the calculation of IoU to the minimum bounding rectangle C, constraining the generated bounding boxes to be closer to the true bounding box in terms of range; L ol is to calculate the same overlap loss. Both L L1 and L GIoU are divided by N to act on each bounding box on average, while L ol is averaged over several pairwise combinations of possible generated bounding boxes;

[0045] The total bounding box loss is:

[0046] L bbox = c 1 L L1 + c 2 L GIoU + c 3 L ol ,

[0047] where c 1 = 5, c 2 = 2, c 3 = 1;

[0048] Step 4.3.2. For the training of the character-level layout generation network, the Bezier curve loss is calculated using the character-level layout generated in Step 2.2 and the true layout, and is constrained by the L1 loss:

[0049]

[0050] Step 4.3.3, for the training of the region-level layout generation network, the sentence-level layout generation network, and the character-level layout generation network, add the confidence loss, and use the standard cross-entropy loss to evaluate the confidence loss of each bounding box and Bezier curve:

[0051] L conf = -p·log(q),

[0052] where q is the confidence predicted by the model, and p is the confidence of the target label (taking the value of 1);

[0053] Step 4.3.4, for each decoder layer of the region-level layout generation network, the sentence-level layout generation network, and the character-level layout generation network, calculate the corresponding loss, and use the Hungarian algorithm to calculate the matching loss between the output and the true layout to obtain the final loss. Combining the total bounding box loss obtained in Step 4.3.1, the Bezier curve loss obtained in Step 4.3.2, and the confidence loss obtained in Step 4.3.3, the final total loss function is:

[0054] L = λ 1 L bbox + λ 2 L Bezier + λ 3 L conf ,

[0055] where λ 1 = 1, λ 2 = 5, λ 3 = 1;

[0056] Step 4.4, for the scene text image generation module, to quantify the difference between the original image and the predicted image in all text regions, use the text-aware loss:

[0057]

[0058] where h, w = 512, 512 represent the height and width of the image. Here, the feature map before the fully connected layer represents the text writing information of the original image and the predicted image at position p. Since the time step t is directly related to the text quality of the predicted image x' 0 use the adjustment function φ(t) to dynamically adjust the weight of the loss, where is the coefficient in the diffusion process;

[0059] Finally, obtain the weight files of the trained hierarchical layout generation module and the scene text image generation module.

[0060] The specific method of the said Step 5 is:

[0061] Step 5.1 Load the complete arbitrary-shaped scene text image generation model constructed in Step 3, read the weight files of the hierarchical layout generation module and the scene text image generation module after training in Step 4, load the weight files into the structures of the hierarchical layout generation module and the scene text image generation module, and set the complete arbitrary-shaped scene text image generation model to the inference mode to fix the model parameters;

[0062] Step 5.2, input the text prompts in the scene text image generation test set preprocessed in Step 1.4 into the complete arbitrary-shaped scene text image generation model constructed in Step 3. Use the background image generation module to modify the text prompts and perform text detection on the generated background image to obtain a text-free background image; use the hierarchical layout generation module to encode the text-free background image, extract features and hierarchically generate a character-level layout; use the scene text image generation module to segment and fill the character-level layout, and combine it with the text-free background image to obtain a text-placed background image. After encoding the text-placed background image, obtain the text-placed background image features and add random Gaussian noise. Encode the text prompts to obtain text features. Use the text-placed background image features with added random Gaussian noise and the text features as the input of the conditional diffusion model, and gradually diffuse to generate the final scene text image.

[0063] The present invention also provides a hierarchical layout-driven arbitrary-shaped scene text image generation system, including:

[0064] A scene text image training set preprocessing module, which is used to preprocess the scene text image training set. The scene text image training set includes a scene text layout generation training set, a scene text image generation training set, and a scene text image generation test set, and finally obtains the preprocessed scene text layout generation training set, scene text image generation training set, and scene text image generation test set;

[0065] A first module, which is used to construct a background image generation module, a hierarchical layout generation module, and a scene text image generation module;

[0066] An arbitrary-shaped scene text image generation model construction module, which is used to implement the construction of a complete arbitrary-shaped scene text image generation model based on the background image generation module, the hierarchical layout generation module, and the scene text image generation module;

[0067] A first module, which is used to use the preprocessed scene text layout generation training set and scene text image generation training set to train the constructed hierarchical layout generation module and scene text image generation module respectively, and obtain the weight files of the trained hierarchical layout generation module and scene text image generation module;

[0068] The model inference module is used to generate a test set based on the scene text image, utilize the weight files of the trained hierarchical layout generation module and the scene text image generation module, and perform model inference through the constructed complete arbitrary-shaped scene text image generation model to obtain the final scene text image.

[0069] The present invention also provides an arbitrary-shaped scene text image generation device driven by a hierarchical layout, including:

[0070] A memory: storing a computer program for the above-mentioned arbitrary-shaped scene text image generation method driven by a hierarchical layout, which is a computer-readable device;

[0071] A processor: used to implement the above-mentioned arbitrary-shaped scene text image generation method driven by a hierarchical layout when executing the computer program.

[0072] The present invention also provides a computer-readable storage medium, which stores a computer program. When the computer program is executed by a processor, it can implement the above-mentioned arbitrary-shaped scene text image generation method driven by a hierarchical layout.

[0073] Compared with the prior art, the beneficial effects of the present invention are:

[0074] 1. Through the complete arbitrary-shaped scene text image generation model constructed in step 3 of the present invention, the background image generation module reads and processes the text prompts in the preprocessed scene text image generation test set obtained in step 1.4 to generate a textless background image; the hierarchical layout generation module reads and processes the textless background image to hierarchically generate the final character hierarchical layout; the scene text image generation module reads and processes the text prompts, the textless background image, and the final character hierarchical layout to gradually generate the final scene text image; the present invention can automatically generate the final scene text image from the input text prompts with one key without relying on user input layout information.

[0075] 2. Through the hierarchical layout generation module constructed in step 2.2 of the present invention, which includes an image feature extraction network, a regional hierarchical layout generation network, a statement hierarchical layout generation network, and a character hierarchical layout generation network, it is used to extract the features of the textless background image generated in step 2.1, gradually generate the content of the regional, statement, and character hierarchical layouts, and finally output the character hierarchical layout; the final character hierarchical layout can include diverse layout forms such as bending and tilting, generating diverse scene text images.

[0076] 3. The scene text image generation module constructed by the present invention through step 2.3 includes a controllable image-to-image model, which uses the text prompt read in step 1.3 as input, and the character-level layout generated in step 2.2 and the textless background image generated in step 2.1 as control conditions. Using the controllable image-to-image model, a scene text image that precisely conforms to the description of the text prompt can be generated.

[0077] 4. During the training process of the present invention, the hierarchical layout generation module is optimized using a hierarchical loss mechanism. The total bounding box loss designed in step 4.3.1 is used to constrain the region-level layout generation network and the sentence-level layout generation network. The Bezier curve loss designed in step 4.3.2 is used to constrain the character-level layout generation network. The confidence loss designed in step 4.3.3 is used to constrain the region-level layout generation network, the sentence-level layout generation network, and the character-level layout generation network. The total loss function constructed by the above losses according to a certain weight ratio is used to optimize the hierarchical layout generation module; the scene text image generation module is optimized through the text perception loss designed in step 4.4. The weights of the finally trained hierarchical layout generation module can accurately generate refined character-level layouts, and the weights of the scene text image generation module can accurately generate the scene text images described by the text prompt.

[0078] In summary, the present invention constructs a complete arbitrary-shaped scene text image generation model and the hierarchical layout generation module therein, which can automatically generate single-line or multi-line scene text images with various situations such as bending and tilting without relying on user input layout;

[0079] The present invention proposes a hierarchical layout representation method; and proposes a text-to-image model that supports the generation of arbitrary-shaped scene text images based on the hierarchical layout. Through the hierarchical layout generation technology, the generated text and background are naturally integrated, and the visual effect is more realistic. This method not only significantly improves the layout flexibility by introducing the character-level layout represented by Bezier curves, supporting the generation of diverse text styles; at the same time, through the complete arbitrary-shaped scene text image generation model, it simplifies the user operation, and only requires a text prompt to generate the scene text image that meets the requirements, improving the usability and practicality of the technology and making up for the deficiencies of the existing scene text image generation algorithms. BRIEF DESCRIPTION OF THE DRAWINGS

[0080] Figure 1 is the training flowchart of the hierarchical layout generation module of the present invention.

[0081] Figure 2 is the training flowchart of the scene text image generation module of the present invention.

[0082] Figure 3 is the complete inference flowchart of the present invention.

[0083] Figure 4 is the network structure diagram of the regional hierarchical layout generation of the present invention.

[0084] Figure 5 is the network structure diagram of the sentence / character hierarchical layout generation of the present invention.

[0085] Figure 6 is the comparison diagram of the generation effect of the present invention. Detailed implementation manners

[0086] To make the implementation process and features of the method clearer and easier to understand, the present invention will be further described in detail below in conjunction with the accompanying drawings and specific embodiments.

[0087] Most existing methods need to pre-define the text layout or predict the text position through simple rectangular boxes, making it difficult to accurately present complex text structures. Aiming at the deficiencies of the existing scene text image generation technology, the present invention proposes a method for generating arbitrary-shaped scene text images driven by hierarchical layout, filling the technical gap in flexibly generating arbitrary-shaped scene text images; based on the constructed complete model for generating arbitrary-shaped scene text images, the generation of the overall scene text image is divided into three steps: background image generation, hierarchical layout generation, and scene text image generation, significantly improving the generation ability of complex text layouts, realizing the automated generation from simple text prompts to high-quality scene text images, and effectively solving the bottlenecks in flexibility, adaptability, and generation quality of traditional methods.

[0088] The core content of the present invention includes: the hierarchical layout generation module generates the layout structures at the regional, sentence, and character levels of the text in a layer-by-layer refinement manner. In particular, Bezier curves are introduced to represent the character-level layout, which can flexibly express complex text forms such as curves and italics, solving the problem of inaccurate representation by traditional rectangular boxes; the directional control mechanism uses the multi-layer Transformer architecture of the hierarchical layout generation module to automatically generate the character-level layout in combination with the image content, providing more refined geometric and directional control, making the generated text highly integrated with the visual content of the background; based on the complete model for generating arbitrary-shaped scene text images, the generation of the complete scene text image is innovatively divided into three steps: background image generation, hierarchical layout generation, and scene text image generation, realizing the automated generation from simple text prompts to high-quality scene text images. These technologies together solve the problems of poor layout flexibility and insufficient detail control in existing text-to-image models, significantly improving the generation effect and practicality of text generation.

[0089] A method for generating arbitrary-shaped scene text images driven by hierarchical layout includes the following steps:

[0090] Step 1: Preprocess the scene text image training set, which includes the scene text layout generation training set, the scene text image generation training set, and the scene text image generation test set, and finally obtain the preprocessed scene text layout generation training set, scene text image generation training set, and scene text image generation test set;

[0091] The specific method of Step 1 is as follows:

[0092] Step 1.1: The scene text layout generation training set includes textless background images and content with layout at the region, sentence, and character levels. The textless background images and the content with layout at the region, sentence, and character levels in the scene text layout generation training set are used to form training background image - layout pairs. Among them, the region and sentence level layouts are composed of 4 values, with a total of 2 control points, respectively representing the horizontal and vertical coordinates of the upper left and lower right corners of the layout. The character level layout is composed of 16 values, respectively representing the horizontal and vertical coordinates of the 4 control points of the top Bezier curve and the horizontal and vertical coordinates of the 4 control points of the bottom Bezier curve of the layout;

[0093] Step 1.2: Preprocess the training background image - layout pairs in Step 1.1. Read the textless background images in the training background image - layout pairs through the Pillow library, adjust the textless background images to 3×512×512 and convert them to the tensor format to meet the input requirements of the arbitrary - shaped hierarchical text layout generation model. At the same time, normalize the three color channels of the textless background images respectively using the means of 0.485, 0.456, 0.406 and the variances of 0.229, 0.224, 0.225. Map the color value range of the input textless background images from integer values in [0 - 255] to floating - point values in [-1 - 1]. Read the region, sentence, and character level layouts corresponding to the textless background images in the training background image - layout pairs in Step 1.1 through the json library, and map the horizontal and vertical coordinate values from floating - point values in [0 - 511] to floating - point values in [0 - 1] to obtain the preprocessed training background image - layout pairs;

[0094] Step 1.3: The scene text image generation training set includes text prompts, textless background images, character level layout content, and scene text images. The preprocessing of the scene text image generation training set is as follows: Process a single textless background image in the same way as in Step 1.2. Read the text prompts and character level layout content corresponding to a single textless background image through the json library to obtain the preprocessed scene text image generation training set;

[0095] Step 1.4, the scene text image generation test set includes text prompts and scene text images; preprocess the scene text image generation test set, specifically: read the text prompts through the json library to obtain the preprocessed scene text image generation test set.

[0096] Step 2, construct a background image generation module, a hierarchical layout generation module, and a scene text image generation module;

[0097] The specific method of step 2 is as follows:

[0098] Step 2.1, construct a background image generation module, including a text-to-image model and a scene text detection model, for generating a text-free background image through the text prompts read in step 1.3;

[0099] Step 2.1.1, remove the text content to be rendered in the text prompts read in step 1.3, so that it generates a background image through the text-to-image model, and the text-to-image model uses a model based on the Diffusion principle for generation.

[0100] Step 2.1.2, use the scene text detection model to detect the background image generated by the text-to-image model in step 2.1.1. If text is detected, repeat steps 2.1.1 - 2.1.2 until no text is detected, and output a text-free background image; the scene text detection model uses the basic PP-OCRv3 model for detection.

[0101] Step 2.2, construct a hierarchical layout generation module. The hierarchical layout generation module is a network model based on transformer, including an image feature extraction network, a regional hierarchical layout generation network, a sentence hierarchical layout generation network, and a character hierarchical layout generation network, for extracting the features of the text-free background image generated in step 2.1, gradually generating the content of the regional, sentence, and character hierarchical layouts, and finally outputting the character hierarchical layout;

[0102] Step 2.2.1, the image feature extraction network is an encoder fine-tuned based on ResNet50;

[0103] Extract the features of the text-free background image generated in step 2.1 through the encoder fine-tuned based on ResNet50, and output the image features;

[0104] Step 2.2.2, as Figure 4As shown, the regional hierarchical layout generation network is a network model based on the Transformer architecture. The regional hierarchical layout generation network includes two decoder layers. Each decoder layer includes an Embedding layer, a SelfAttention layer, two Layer Norm layers, a CrossAttention layer, a Feed Forward layer, an MLP layer, and a Linear layer. The output of the Embedding layer is connected to the input of the SelfAttention layer. The output of the SelfAttention layer is connected to the input of the first LayerNorm layer. The output of the first Layer Norm layer is connected to the input of the Cross Attention layer. The output of the CrossAttention layer is connected to the input of the second Layer Norm layer. The output of the second Layer Norm layer is connected to the input of the Feed Forward layer. The output of the Feed Forward layer is respectively connected to the inputs of the MLP layer and the Linear layer. Among them, the MLP layer outputs the regional hierarchical layout, and the Linear layer outputs its confidence. The image features output in step 2.2.1 are input into the regional hierarchical layout generation network. The regional hierarchical layout generation network predicts the layout of the regional hierarchy and its confidence by judging the content of different regions of the image features, and outputs the regional hierarchical image features and the layout of the regional hierarchy;

[0105] Step 2.2.3, as Figure 5 shown, construct the sentence hierarchical layout generation network. The sentence hierarchical layout generation network adds the regional hierarchical image features output by the regional hierarchical layout generation network in step 2.2.2 to the image features output in step 2.2.1 as input features for encoding. The architecture of the sentence hierarchical layout generation network is the same as that of the regional hierarchical layout generation network. By judging the connection between the content of different regions of the features, it predicts the layout of the sentence hierarchy and its confidence, and outputs the sentence hierarchical image features and the sentence hierarchical layout;

[0106] Step 2.2.4, as Figure 5 shown, the character hierarchical layout generation network adds the sentence hierarchical image features output by the sentence hierarchical layout generation network in step 2.2.3 to the image features output in step 2.2.1 as input features for encoding. The architecture of the character hierarchical layout generation network is the same as that of the regional hierarchical layout generation network. By judging the connection between the content of different regions of the features, it predicts the layout of the character hierarchy and its confidence, and finally outputs the character hierarchical layout.

[0107] Step 2.3, construct a scene text image generation module, including a controllable image-to-image generation model. According to the text prompt read in Step 1.3, the character-level layout generated in Step 2.2, and the textless background image generated in Step 2.1, use the controllable image-to-image generation model to generate a high-quality scene text image that precisely conforms to the description of the text prompt;

[0108] The controllable image-to-image generation model consists of a VAE-based image encoder, a Transformer-based text encoder, and a Diffusion-based conditional diffusion model;

[0109] Step 2.3.1, according to the characters and their quantities of the text to be rendered in the text prompt read in Step 1.3, further refine and segment the character-level layout generated in Step 2.2 to obtain the independent layout of each character and the characters of the text to be rendered;

[0110] Step 2.3.2, use the Pillow library to sequentially place the characters of the text to be rendered obtained in Step 2.3.1 at the independent layout positions of each character specified in Step 2.3.1. The font file used for placement is ArialUnicode to obtain a text-placement image;

[0111] Step 2.3.3, input the text prompt read in Step 1.3 into the Transformer-based text encoder to obtain text features. Combine the textless background image generated in Step 2.1.2 and the text-placement image obtained in Step 2.3.2, and input them into the VAE-based image encoder for feature extraction to obtain text-placement background image features, and add random Gaussian noise; input the obtained text features and the text-placement background image features after adding random Gaussian noise into the Diffusion-based conditional diffusion model to generate a scene text image.

[0112] Step 3, based on the background image generation module, hierarchical layout generation module, and scene text image generation module constructed in Step 2, construct a complete arbitrary-shaped scene text image generation model;

[0113] The specific method of Step 3 is as follows:

[0114] Step 3.1, connect the text prompt read in Step 1.3 with the background image generation module constructed in Step 2. The text-to-image generation model and the scene text detection model in the background image generation module are both loaded with pre-trained model parameters. The output image size of the text-to-image generation model is the same as the input image size of the scene text detection model, and the output result is a textless background image;

[0115] Step 3.2, input the text-free background image output in Step 3.1 into the hierarchical layout generation module constructed in Step 2.2. The image feature extraction network continuously performs upsampling convolution on the image to extract image features. The image features are input into the regional hierarchical layout generation network. The regional hierarchical layout generation network continuously predicts the next regional hierarchical layout based on the feature values of the input image, selects the best layout according to the confidence, and obtains the regional hierarchical image features and the layout at the regional level. After adding the image features to the regional hierarchical image features, they are input into the sentence-level layout generation network. Through the same process, sentence-level image features and sentence-level layouts are obtained, and so on. Finally, the character-level layout is obtained, and the final character-level layout is selected according to the best confidence.

[0116] Step 3.3, segment and fill the character-level layout obtained in Step 3.2 according to the characters of the text to be rendered in the text prompt read in Step 1.3 to obtain the text placement image. Connect the text prompt read in Step 1.3, the text-free background image generated in Step 2.1, and the text placement image to the scene text image generation module constructed in Step 2.3. After combining the text placement image and the text-free background image, the text placement background image is obtained. Feature extraction is performed through the VAE-based image encoder in Step 2.3 to obtain the text placement background image features, and random Gaussian noise is added. The conditional diffusion model based on Diffusion takes the text placement background image features added with random Gaussian noise and the current time step as inputs, predicts the noise added between the current time step and the previous time step, subtracts the predicted noise from the input image to obtain the predicted image corresponding to the previous time step. Again, use the conditional diffusion model based on Diffusion to take the predicted image, the text placement background image features added with random Gaussian noise, and the current time step as inputs, predict the noise added between the current time step and the previous time step, and obtain the predicted image of the previous time step again. And so on, iterate to the initial time step to obtain a complete arbitrary-shaped scene text image generation model.

[0117] Step 4, use the preprocessed scene text layout generated in Step 1 to generate a training set and a scene text image generation training set, and train the hierarchical layout generation module and the scene text image generation module constructed in Step 2 respectively to obtain the weight files of the trained hierarchical layout generation module and the scene text image generation module.

[0118] The specific method of Step 4 is as follows:

[0119] In the network structure corresponding to the present invention, only the hierarchical layout generation module and the scene text image generation module need to be trained.

[0120] Such as Figure 1As shown in the figure, in step 4.1, the text-free background image in the preprocessed training background image-hierarchical layout pair in step 1.2 is used as the input of the hierarchical layout generation module constructed in step 2, and the corresponding hierarchical layout is used as the true hierarchical layout;

[0121] The hierarchical layout generation module uses AdamW as the parameter optimizer, with an initial learning rate set to 0.0001 and a weight decay of 0.00001 every 200 epochs. The encoder of the ResNet50 fine-tuned in the image feature extraction network uses a trainable residual network, with a learning rate set to 0.00001 and an input image size of 512x512 pixels.

[0122] As Figure 2 shown in the figure, in step 4.2, the text prompt words, text-free background images, and character hierarchical layout content generated from the scene text images preprocessed in step 1.3 are used as the input of the scene text image generation module constructed in step 2, and the corresponding scene text images are used as the true scene text images;

[0123] The scene text image generation module also selects AdamW as the parameter optimizer, with an initial learning rate of 0.00002 and betas parameters of (0.5, 0.999).

[0124] In step 4.3, during the training process, the loss function is used to evaluate the performance of the network model and optimize the network model parameters. The optimization of the hierarchical layout generation module adopts a hierarchical loss mechanism. For the region hierarchical layout generation network, sentence hierarchical layout generation network, and character hierarchical layout generation network, the defined loss function is as follows:

[0125] In step 4.3.1, for the training of the region hierarchical layout generation network and the sentence hierarchical layout generation network, the following three parts of loss constraints are adopted:

[0126]

[0127]

[0128] Among them, L L1 represents calculating the L1 loss between the region hierarchical layout output in step 2.2.2 and the sentence hierarchical layout generated in step 2.2.3 and the true bounding box, measuring the coordinate gap between the predicted bounding box B and the true bounding box ; L GIoU represents calculating the generalized IoU loss between the bounding boxes, which extends the calculation of IoU to the minimum bounding rectangle C, constraining the generated bounding boxes to be closer to the true bounding box in terms of range; L ol is to calculate the same overlap loss to avoid the generated bounding boxes covering each other. L L1 and LGIoU All need to be divided by N to evenly apply to each bounding box, while L ol needs to be averaged over several pairwise combinations of the possible generated bounding boxes;

[0129] The total bounding box loss is:

[0130] L bbox = c 1 L L1 + c 2 L GIoU + c 3 L ol ,

[0131] where c 1 = 5, c 2 = 2, c 3 = 1;

[0132] Step 4.3.2, for the training of the character-level layout generation network, calculate the Bezier curve loss using the character-level layout and the real layout generated in Step 2.2.4, and constrain it through a simple L1 loss:

[0133]

[0134] Step 4.3.3, for the training of the region-level layout generation network, the sentence-level layout generation network, and the character-level layout generation network, add a confidence loss, and use the standard cross-entropy loss to evaluate the confidence loss of the region-level layout generated in Step 2.2.2, the sentence-level layout generated in Step 2.2.3, and the character-level layout generated in Step 2.2.4:

[0135] L conf = -p·log(q),

[0136] where q is the confidence predicted by the model, and p is the confidence of the target label (taking the value of 1).

[0137] Step 4.3.4, for each decoder layer of the region-level layout generation network, the sentence-level layout generation network, and the character-level layout generation network, calculate the corresponding loss, and use the Hungarian algorithm to calculate the matching loss between the output and the real layout to obtain the final loss. Combining the total bounding box loss obtained in Step 4.3.1, the Bezier curve loss obtained in Step 4.3.2, and the confidence loss obtained in Step 4.3.3, the final total loss function is:

[0138] L = λ 1 L bbox + λ 2 L Bezier + λ 3L conf ,

[0139] where λ 1 = 1, λ 2 = 5, λ 3 = 1;

[0140] Step 4.4. For the scene text image generation module, to quantify the difference between the original image and the predicted image in all text regions, use the text-aware loss:

[0141]

[0142] The text-aware loss introduces a mean squared error constraint to minimize the difference between the predicted image and the original image in the text regions, thereby improving the quality of the text content in the generated image. Among them, h, w = 512, 512 represent the height and width of the image. Here, the feature map before the fully connected layer represents the text writing information of the original image and the predicted image at position p. Since the time step t is directly related to the text quality of the predicted image x' 0 use the adjustment function φ(t) to dynamically adjust the weight of the loss, where is the coefficient in the diffusion process;

[0143] Finally, obtain the weight files of the trained hierarchical layout generation module and the scene text image generation module.

[0144] As Figure 3 shown, in Step 5, based on the scene text image generated in Step 1 for preprocessing, generate a test set, use the weight files of the hierarchical layout generation module and the scene text image generation module trained in Step 4, and through the complete arbitrary-shaped scene text image generation model constructed in Step 3 and perform model inference to obtain the final scene text image.

[0145] The specific method of the said Step 5 is as follows:

[0146] Step 5.1. Load the complete arbitrary-shaped scene text image generation model constructed in Step 3, read the weight files of the hierarchical layout generation module and the scene text image generation module trained in Step 4, load the weight files into the structures of the hierarchical layout generation module and the scene text image generation module, and set the complete arbitrary-shaped scene text image generation model to the inference mode and fix the model parameters;

[0147] Step 5.2 Input the scene text image preprocessed in Step 1.4 as the text prompt in the test set into the complete arbitrary-shaped scene text image generation model constructed in Step 3. Use the background image generation module to modify the text prompt and perform text detection on the generated background image to obtain a text-free background image; use the hierarchical layout generation module to encode the text-free background image, extract features, and hierarchically generate a character-level layout; use the scene text image generation module to segment and fill the character-level layout, and combine it with the text-free background image to obtain a text-placement background image. After encoding the text-placement background image, obtain the text-placement background image features and add random Gaussian noise. Encode the text prompt to obtain text features. Use the text-placement background image features with added random Gaussian noise and the text features as the input of the conditional diffusion model, and gradually diffuse to generate the final scene text image.

[0148] The effects of the present invention will be further described below in conjunction with simulation experiments.

[0149] Simulation conditions

[0150] The hardware platform for the simulation experiment of the present invention is: the processor is Intel i9-10900X, and the GPU is NVIDIA GeForce RTX 4090.

[0151] The software platform for the simulation experiment of the present invention is: Ubuntu 20.04 operating system, Python 3.9.0 programming language, Pytorch 2.2.1 deep learning framework, and Visual Studio Code code editor.

[0152] Simulation content

[0153] Use the training background image-hierarchical layout pairs preprocessed in Step 1.2 and the scene text image generated in Step 1.4 to preprocess the training set to train the hierarchical layout generation module and the scene text image generation module constructed in Step 2 respectively, and obtain the weight files of the trained hierarchical layout generation module and the scene text image generation module; load the complete arbitrary-shaped scene text image generation model constructed in Step 3, read the weight files of the trained hierarchical layout generation module and the scene text image generation module and load them into the hierarchical layout generation module and the scene text image generation module structure, and set the complete arbitrary-shaped scene text image generation model to the inference mode to fix the model parameters; input the specified prompt into the complete arbitrary-shaped scene text image generation model for processing, and generate a text layout and a scene text image to obtain the output layout and the output image.

[0154] Simulation results

[0155] The two existing methods for comparison with the results obtained by the present invention are GlyphControl and TextDiffuser respectively. For GlyphControl, TextDiffuser and the method of the present invention, the same prompt words are input. Among them, GlyphControl requires additional input of glyphs. The output layouts of TextDiffuser and the method of the present invention are compared, as well as the output images of GlyphControl, TextDiffuser and the method of the present invention. The specific effects are as Figure 6 shown. When the prompt words are input, the GlyphControl method in the figure needs to rely on the input glyphs to assist in outputting the image; the TextDiffuser method in the figure can output the layout and assist in outputting the image according to the output layout, but can only generate the output image of the horizontal straight text layout; while the present invention can generate the output images of various text layouts such as horizontal and curved while outputting the layout, and the results obtained are superior to the other two methods in terms of layout diversity and scene text image rendering quality.

[0156] The present invention also provides a system for generating arbitrary-shaped scene text images driven by a hierarchical layout, including:

[0157] A preprocessing module for the scene text image training set, which is used to preprocess the scene text image training set in step 1. The scene text image training set includes a scene text layout generation training set, a scene text image generation training set, and a scene text image generation test set, and finally obtains the preprocessed scene text layout generation training set, scene text image generation training set, and scene text image generation test set;

[0158] The first module is used to implement the construction of the background image generation module, the hierarchical layout generation module, and the scene text image generation module in step 2;

[0159] A module for constructing an arbitrary-shaped scene text image generation model, which is used to construct a complete arbitrary-shaped scene text image generation model based on the background image generation module, the hierarchical layout generation module, and the scene text image generation module constructed in step 2 in step 3;

[0160] The first module is used to implement the training of the hierarchical layout generation module and the scene text image generation module constructed in step 2 respectively using the preprocessed scene text layout generation training set and scene text image generation training set in step 1 in step 4, and obtain the weight files of the trained hierarchical layout generation module and scene text image generation module;

[0161] A model inference module, which is used to implement the generation of a test set from the scene text image preprocessed in step 1 in step 5, utilize the weight files of the hierarchical layout generation module and the scene text image generation module trained in step 4, and perform model inference through the complete arbitrary-shaped scene text image generation model constructed in step 3 to obtain the final scene text image.

[0162] The present invention also provides an arbitrary-shaped scene text image generation device driven by a hierarchical layout, including:

[0163] A memory: storing a computer program of the above-mentioned arbitrary-shaped scene text image generation method driven by a hierarchical layout, which is a computer-readable device;

[0164] A processor: used to implement the above-mentioned arbitrary-shaped scene text image generation method driven by a hierarchical layout when executing the computer program.

[0165] The present invention also provides a computer-readable storage medium, which stores a computer program, and when the computer program is executed by a processor, it can implement the above-mentioned arbitrary-shaped scene text image generation method driven by a hierarchical layout.

Claims

1. A hierarchical layout driven method for generating arbitrary shape scene text images, characterized in that: The following steps are involved: Step 1, preprocessing the scene text image training set, the scene text image training set includes a scene text layout generation training set, a scene text image generation training set and a scene text image generation test set, and finally obtaining the preprocessed scene text layout generation training set, scene text image generation training set and scene text image generation test set; Step 2, constructing a background image generation module, a hierarchical layout generation module and a scene text image generation module; Step 3, based on the background image generation module, hierarchical layout generation module and scene text image generation module constructed in step 2, a complete arbitrary shape scene text image generation model is constructed; Step 4, using the scene text layout generation training set and the scene text image generation training set preprocessed in step 1, respectively train the hierarchical layout generation module and the scene text image generation module constructed in step 2 to obtain weight files of the trained hierarchical layout generation module and the scene text image generation module; Step 5, based on the scene text image preprocessed in step 1, a test set is generated, and the weight files of the hierarchical layout generation module and the scene text image generation module trained in step 4 are used to generate the complete arbitrary shape scene text image model constructed in step 3 and perform model inference to obtain the final scene text image.

2. The method for generating a hierarchical layout-driven arbitrary-shaped scene text image according to claim 1, characterized in that: The specific method of step 1 is: Step 1.1, the scene text layout generation training set includes a text-free background image and content containing region, sentence and character hierarchical layout; the text-free background image and content containing region, sentence and character hierarchical layout in the scene text layout generation training set are used to form a training background image-hierarchical layout pair, wherein the region and sentence hierarchical layout are composed of 4 values, with a total of 2 control points, which respectively express the horizontal and vertical coordinates of the upper left corner and the lower right corner of the layout; the character hierarchical layout is composed of 16 values, which respectively express the horizontal and vertical coordinates of the four control points of the top Bezier curve and the four control points of the bottom Bezier curve of the layout; Step 1.2, preprocess the training background image-hierarchical layout pair in step 1.1, read the text-free background image in the training background image-hierarchical layout pair, adjust the text-free background image to 3×512×512 and convert it into tensor format, and use the mean and variance to normalize the three color channels of the text-free background image respectively, and the input text-free background image color value range is mapped from an integer value of [0 to 255] to a floating point value of [-1 to 1]; read the three hierarchical layouts of region, sentence and character corresponding to the text-free background image in the training background image-hierarchical layout pair in step 1.1 through the json library, and map the horizontal and vertical coordinate values ​​from floating point values ​​of [0 to 511] to floating point values ​​of [0 to 1] to obtain the preprocessed training background image-hierarchical layout pair; Step 1.3, the scene text image generation training set includes text prompt words, text-free background images, character level layout content and scene text images; the scene text image generation training set is preprocessed, specifically: a single text-free background image is processed in the same process as step 1.2; the text prompt words and character level layout content corresponding to the single text-free background image are read through the json library to obtain the preprocessed scene text image generation training set; Step 1.4, the scene text image generation test set includes text prompt words and scene text images; the scene text image generation test set is preprocessed, specifically: the text prompt words are read through the json library to obtain the preprocessed scene text image generation test set.

3. The method for generating a hierarchical layout-driven arbitrary-shaped scene text image according to claim 1, characterized in that: The specific method of step 2 is: Step 2.1, constructing a background image generation module, including a text image model and a scene text detection model, for generating a text-free background image using the text prompt words read in step 1.3; Step 2.1.1, removing the text content to be rendered from the text prompt word read in step 1.3, so that the text prompt word generates a background image through the text-generated graph model; Step 2.1.2, use the scene text detection model to detect the background image generated by the text image model in step 2.1.

1. If text is detected, repeat steps 2.1.1-2.1.2 until no text is detected, and output a background image without text; Step 2.2, construct a hierarchical layout generation module. The hierarchical layout generation module is a network model based on transformer, including an image feature extraction network, a region hierarchical layout generation network, a sentence hierarchical layout generation network, and a character hierarchical layout generation network. It is used to extract the features of the text-free background image generated in step 2.1, gradually generate the content of the region, sentence and character hierarchical layout, and finally output the character hierarchical layout; Step 2.2.1, the image feature extraction network is an encoder fine-tuned based on ResNet50; Extract features from the text-free background image generated in step 2.1 using an encoder fine-tuned on ResNet50 and output image features; Step 2.2.2, the regional hierarchical layout generation network is a network model based on the Transformer architecture. The regional hierarchical layout generation network includes two decoder layers, each decoder layer includes an Embedding layer, a SelfAttention layer, two Layer Norm layers, a Cross Attention layer, a Feed Forward layer, an MLP layer and a Linear layer. The output of the Embedding layer is connected to the input of the Self Attention layer, the output of the SelfAttention layer is connected to the input of the first Layer Norm layer, the output of the first LayerNorm layer is connected to the input of the Cross Attention layer, the output of the Cross Attention layer is connected to the input of the second LayerNorm layer, the output of the second Layer Norm layer is connected to the input of the Feed Forward layer, and the output of the Feed Forward layer is connected to the input of the MLP layer and the Linear layer respectively, wherein the MLP layer outputs the regional hierarchical layout, and the Linear layer outputs its confidence. The image features output in step 2.2.1 are input into the regional hierarchical layout generation network. The regional hierarchical layout generation network predicts the regional hierarchical layout and its confidence by judging the content of different regions of the image features, and outputs the regional hierarchical image features and the regional hierarchical layout; Step 2.2.3, construct a sentence level layout generation network. The sentence level layout generation network adds the regional level image features output by the regional level layout generation network in step 2.2.2 and the image features output by step 2.2.1 as input features for encoding. The architecture of the sentence level layout generation network is the same as that of the regional level layout generation network. By judging the connection between the contents of different feature regions, the sentence level layout and its confidence are predicted, and the sentence level image features and sentence level layout are output. Step 2.2.4, the character level layout generation network adds the sentence level image features output by the sentence level layout generation network in step 2.2.3 and the image features output by step 2.2.1 as input features for encoding. The architecture of the character level layout generation network is the same as that of the region level layout generation network. By judging the connection between the contents of different feature regions, the character level layout and its confidence are predicted, and finally the character level layout is output; Step 2.3, constructing a scene text image generation module, including a controllable image generation model, and using the controllable image generation model to generate a scene text image that accurately matches the description of the text prompt word according to the text prompt word read in step 1.3, the character hierarchy layout generated in step 2.2, and the text-free background image generated in step 2.1; The controllable graph generation model consists of a VAE-based image encoder, a Transformer-based text encoder, and a Diffusion-based conditional diffusion model; Step 2.3.1, according to the characters and the number of the text to be rendered in the text prompt words read in step 1.3, further refine the character hierarchy layout generated in step 2.2 to obtain an independent layout for each character and the characters of the text to be rendered; Step 2.3.2, placing the characters of the text to be rendered obtained in step 2.3.1 in turn into the independent layout position of each character specified in step 2.3.1, using Arial Unicode as the font file used for placement, to obtain a text placement image; In step 2.3.3, the text prompt words read in step 1.3 are input into the text encoder based on Transformer to obtain text features. The text-free background image generated in step 2.1.2 and the text placement image obtained in step 2.3.2 are combined and input into the image encoder based on VAE for feature extraction to obtain text placement background image features, and random Gaussian noise is added. The obtained text features and the text placement background image features after adding random Gaussian noise are uniformly input into the conditional diffusion model based on Diffusion to generate a scene text image.

4. The hierarchical layout driven arbitrary shape scene text image generation method according to claim 1, characterized in that: The specific method of step 3 is: Step 3.1, connecting the text prompt word read in step 1.3 and the background image generation module constructed in step 2.1, wherein the text graph model and the scene text detection model in the background image generation module are both loaded with pre-trained model parameters, the output image size of the text graph model is the same as the input image size of the scene text detection model, and the output result is a background image without text; Step 3.2, input the text-free background image output from step 3.1 into the hierarchical layout generation module constructed in step 2.2, and the image feature extraction network extracts image features by continuously convolving the image upwards; the image features are input into the region-level layout generation network, and the region-level layout generation network continuously predicts the next region-level layout according to the feature value of the input image, selects the best layout according to the confidence, and obtains the region-level image features and the region-level layout; the image features are added to the region-level image features and then input into the sentence-level layout generation network, and the same process is followed to obtain the sentence-level image features and the sentence-level layout, and so on, to finally obtain the character-level layout, and select the final character-level layout according to the best confidence; Step 3.3, segment and fill the character hierarchy layout obtained in step 3.2 according to the characters of the text to be rendered in the text prompt words read in step 1.3 to obtain a text placement image, connect the text prompt words read in step 1.3, the text-free background image generated in step 2.1, and the text placement image with the scene text image generation module constructed in step 2.3, combine the text placement image with the text-free background image to obtain a text placement background image, extract features through the VAE-based image encoder in step 2.3, obtain text placement background image features, and add random Gaussian noise; conditional diffusion based on Diffusion The scattered model takes the text placement background image features after adding random Gaussian noise and the current time step as input, predicts the noise added between the current time step and the previous time step, subtracts the predicted noise from the input image, and obtains the predicted image corresponding to the previous time step. The Diffusion-based conditional diffusion model uses the predicted image, the text placement background image features after adding random Gaussian noise, and the current time step as input, predicts the noise added between the current time step and the previous time step, and obtains the predicted image of the previous time step again. And so on, iterates to the initial time step to obtain a complete arbitrary shape scene text image generation model.

5. The hierarchical layout driven arbitrary shape scene text image generation method according to claim 1, characterized in that: The specific method of step 4 is: Step 4.1, using the text-free background image in the training background image-hierarchical layout pair preprocessed in step 1.2 as the input of the hierarchical layout generation module constructed in step 2, and the corresponding hierarchical layout as the real hierarchical layout; Step 4.2, the text prompt words, text-free background images, and character level layout contents in the scene text image generation training set preprocessed in step 1.3 are used as inputs of the scene text image generation module constructed in step 2, and the corresponding scene text images are used as real scene text images; Step 4.3, during the training process, the loss function is used to evaluate the performance of the network model and optimize the parameters of the network model. The optimization of the hierarchical layout generation module adopts a hierarchical loss mechanism. For the region hierarchical layout generation network, the sentence hierarchical layout generation network, and the character hierarchical layout generation network, the defined loss functions are as follows: Step 4.3.1, for the training of the region-level layout generation network and the sentence-level layout generation network, the following three loss constraints are adopted: Among them, L L1 It represents the L1 loss between the regional level layout output by step 2.2.2 and the sentence level layout output by step 2.2.3 and the real bounding box, which measures the difference between the predicted bounding box B and the real bounding box. The coordinate difference between GIoU represents the generalized IoU loss between bounding boxes, which extends the IoU calculation to the minimum bounding rectangle C, constraining the generated bounding box to be closer to the real bounding box in range; L ol is to calculate the same overlap loss, L L1 and L GIoU Divide by N to make it evenly applied to each bounding box, and L ol The average Among several possible pairwise combinations of generated bounding boxes; The total bounding box loss is: L bbox =c1L L1 +c2L GIoU +c3L ol , Where c1=5, c2=2, c3=1; Step 4.3.2, for the training of the character hierarchy layout generation network, use the character hierarchy layout generated in step 2.2 and the real layout to calculate the Bezier curve loss, through the L1 loss constraint: Step 4.3.3, for the training of the region-level layout generation network, the sentence-level layout generation network, and the character-level layout generation network, add confidence loss, and use the standard cross entropy loss to evaluate the confidence loss of each bounding box and Bezier curve: L conf =-p·log(q), Among them, q is the confidence of the model prediction, and p is the confidence of the target label (the value is 1); In step 4.3.4, for each decoder layer of the region-level layout generation network, the sentence-level layout generation network, and the character-level layout generation network, the corresponding loss is calculated, and the Hungarian algorithm is used to calculate the matching loss between the output and the true layout to obtain the final loss. The total bounding box loss obtained in step 4.3.1, the Bezier curve loss obtained in step 4.3.2, and the confidence loss obtained in step 4.3.3 are combined. The final total loss function is: L=λ1L bbox +λ2L Bezier +λ3L conf , Among them, λ1=1,λ2=5,λ3=1; Step 4.4, for the scene text image generation module, to quantify the difference between the original image and the predicted image in all text regions, use the text-aware loss: Among them, h,w = 512, 512 represents the height and width of the image, and the feature map before the fully connected layer here Characterize the text writing information of the original image and the predicted image at position p. Since the time step t is directly related to the text quality of the predicted image x'0, the adjustment function φ(t) is used to dynamically adjust the weight of the loss, where is the coefficient in the diffusion process; Finally, the weight files of the trained hierarchical layout generation module and scene text image generation module are obtained.

6. The hierarchical layout driven arbitrary shape scene text image generation method according to claim 1, characterized in that: The specific method of step 5 is: Step 5.1 loads the complete arbitrary shape scene text image generation model constructed in step 3, reads the weight files of the hierarchical layout generation module and the scene text image generation module trained in step 4, loads the weight files into the hierarchical layout generation module and the scene text image generation module structure, and sets the complete arbitrary shape scene text image generation model to inference mode, fixing the model parameters; Step 5.2, input the text prompt words in the scene text image generation test set preprocessed in step 1.4 into the complete arbitrary shape scene text image generation model constructed in step 3, use the background image generation module to modify the text prompt words and perform text detection on the generated background image to obtain a text-free background image; use the hierarchical layout generation module to encode the text-free background image, extract features and hierarchically generate a character hierarchy layout; use the scene text image generation module to segment and fill the character hierarchy layout, and combine it with the text-free background image to obtain a text placement background image, after encoding the text placement background image, obtain the text placement background image features and add random Gaussian noise, encode the text prompt words to obtain text features, use the text placement background image features after adding random Gaussian noise and the text features as the input of the conditional diffusion model, and gradually diffuse to generate the final scene text image.

7. A hierarchical layout driven arbitrary shape scene text image generation system based on the method according to any one of claims 1 to 6, characterized in that: include: A scene text image training set preprocessing module is used to preprocess the scene text image training set, which includes a scene text layout generation training set, a scene text image generation training set and a scene text image generation test set, and finally obtain the preprocessed scene text layout generation training set, scene text image generation training set and scene text image generation test set; The first module is used to construct a background image generation module, a hierarchical layout generation module and a scene text image generation module; An arbitrary shape scene text image generation model construction module is used to realize the construction of a complete arbitrary shape scene text image generation model based on the background image generation module, the hierarchical layout generation module and the scene text image generation module; The first module is used to use the preprocessed scene text layout generation training set and the scene text image generation training set to train the constructed hierarchical layout generation module and the scene text image generation module respectively, and obtain the weight files of the trained hierarchical layout generation module and the scene text image generation module; The model reasoning module is used to realize the generation of a test set based on scene text images. It uses the weight files of the trained hierarchical layout generation module and the scene text image generation module to construct a complete arbitrary shape scene text image generation model and perform model reasoning to obtain the final scene text image.

8. A hierarchical layout driven arbitrary shape scene text image generation device, characterized in that: include: Memory: a computer program storing the method for generating a hierarchical layout-driven arbitrary-shaped scene text image according to any one of claims 1 to 6, which is a computer-readable device; Processor: used to implement the hierarchical layout driven arbitrary shape scene text image generation method described in any one of claims 1-6 when executing the computer program.

9. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, it can implement the hierarchical layout driven arbitrary shape scene text image generation method described in any one of claims 1-6.

Citation Information

Patent Citations

  • Method for training a text-graph model and text-graph method

    CN116935169B

  • Image generation method and device, equipment and medium

    CN116977774A

  • Text generation image model training method and system and text generation image method and system

    CN117058673A

  • Poster generation method and device and storage medium

    CN117876534A

  • Target Chinese poster generation method and device, computer equipment and storage medium

    CN118397148A