Controllable Southeast Asia text image generation method and device based on conditional diffusion model

By using a conditional diffusion model and an image generation method based on Southeast Asian language features, the problem of low quality in Southeast Asian language text image generation was solved. The generated images are closer to the real scene, thus improving the performance of the recognition model.

CN120894599APending Publication Date: 2025-11-04KUNMING UNIV OF SCI & TECH
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510841949.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-23
Publication Date
2025-11-04

AI Technical Summary

Technical Problem

Existing methods for generating text images in Southeast Asian language scenes have a large gap between the generated quality and the real scene, resulting in poor performance of the trained recognition models in practical applications and difficulty in capturing the morphological features of Southeast Asian language text.

Method used

By employing a conditional diffusion model-based approach, Southeast Asian language text images that more closely resemble real-world scenarios are generated through text sketches in Southeast Asian languages, local attention constraints, and image-level cues. Techniques such as rendering, adaptive thresholding Canny operator, region growing algorithm, pre-trained embedding model, and attention mechanism are used to ensure the accuracy and readability of the generated images.

Benefits of technology

It improved the visual quality and structural consistency of text images in Southeast Asian language scenes, enhanced the recognition performance of the recognition model in real scenes, and achieved an FID score of 20.42 for the generated Burmese text images, with a sequence recognition accuracy of over 93%.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120894599A_ABST
    Figure CN120894599A_ABST
Patent Text Reader

Abstract

The invention relates to a controllable Southeast Asia text image generation method and device based on a conditional diffusion model, and belongs to the technical field of natural language processing. The Southeast Asia language belongs to a low-resource language, and aims to solve the problem that the performance of a trained scene text recognition model in practical application is reduced due to the fact that the quality of synthetic data is low and the difference between the synthetic data and a real scene is large in Southeast Asia language scene text image generation. The invention provides a controllable Southeast Asia text image generation method based on a conditional diffusion model. The controllable Southeast Asia text image generation method mainly comprises three parts of Southeast Asia language text sketch image construction, text coding fusing scene style information and text image control generation based on an attention mechanism. The controllable Southeast Asia text image generation device based on the conditional diffusion model is modularly developed according to the three functions, so that the visual quality and the structural consistency of the generated Southeast Asia language scene text image are effectively improved, and the recognition performance of a recognition model in a real scene is favorably improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to a controllable Southeast Asian text image generation method and apparatus based on a conditional diffusion model, belonging to the field of natural language processing technology. Background Technology

[0002] In the field of scene text image generation, methods for generating scene text image data have achieved certain results in common languages ​​such as Chinese and English, but there are still many challenges in generating scene text images in Southeast Asian languages. The current mainstream generation methods are mainly divided into: (1) Algorithm-based methods: By setting specific rules, text is converted into images. Although this method is simple, its universality is poor and it is difficult to deal with complex Southeast Asian language texts; (2) Template-based methods: A template is set in advance, and text information is filled into the template to generate text images. This method has a certain degree of flexibility, but template design and matching still have limitations; (3) Background overlay methods: Deep learning models such as convolutional neural networks (CNN) are used to convert text into images and overlay them with a preset background. This method largely depends on the quality of training data and the adjustment of model parameters. It has a certain degree of universality, but the data synthesized using the above methods is significantly different from the real scene, specifically as follows: Figure 4 As shown, there are several main problems: 1. The text image data generated by the algorithm is too simple and neat, which is quite different from the image data in the real scene; 2. The synthesized image may have text rendering effects similar to the background layer, resulting in poor quality of the synthesized image; 3. This method cannot enable the model to learn the font texture style and background style of the text in the image, resulting in poor fusion effect of the synthesized image.

[0003] Most existing generation methods follow fixed rules, which results in a relatively homogeneous distribution of the generated data in real Southeast Asian language datasets. As a result, Southeast Asian language scene text recognition models trained using images from current synthesis methods will suffer from poor performance in real-world applications.

[0004] In recent years, with the advent of diffusion models, significant progress has been made in generating realistic and instantly aligned images. Large-scale text-guided diffusion models such as DALLE-2, Imagen, and StableDiffusion have revolutionized the way high-quality images are generated by leveraging the semantics of text input. These models can follow user instructions to generate customized high-fidelity images with specified content, styles, and attributes. However, despite these significant advancements in generation quality, the character characteristics of Southeast Asian languages ​​differ significantly from mainstream languages. The complex arrangement of characters within words, along with numerous nested combinations (such as consonant nesting and vowel superscripts and subscripts), makes it difficult for current mainstream generative models to accurately capture the morphological features of Southeast Asian language texts.

[0005] To solve the above problems, the application provides a controllable Southeast Asian text image generation method and device based on a conditional diffusion model, which generates sketch and contour information of an image through a text label of a Southeast Asian language to control the generation position of the text, and integrates prompt information to constrain the content and style of the generated scene text image, thereby generating a Southeast Asian language text image closer to the real scene. SUMMARY

[0006] The application solves the technical problem that the application provides a controllable Southeast Asian text image generation method and device based on a conditional diffusion model to solve the problem that the quality of synthesized data is low and the gap with the real scene is large in the generation of Southeast Asian language scene text images, which leads to the performance decline of the trained scene text recognition model in actual application. The application particularly introduces image contour sketch, local attention constraint method and image-level prompt to guide the generation process of the text image. The local attention constraint method enables the model to pay more attention to the details and local features of the text when generating the image, thereby improving the accuracy and readability of the generated image. The image-level prompt provides overall image style information to help the model generate an image that meets the expected visual effect. The application effectively improves the visual quality and structural consistency of the generated Southeast Asian language scene text image, which helps to improve the recognition performance of the recognition model in the real scene.

[0007] The technical solution of the application is a controllable Southeast Asian text image generation method based on a conditional diffusion model, which comprises:

[0008] Step 1, constructing a Southeast Asian language text sketch image:

[0009] When using rendering technology, ensure the visual integrity of the combined characters;

[0010] Combine the adaptive threshold Canny operator to alleviate the edge breakage problem of Southeast Asian language curved characters;

[0011] Fill the contour gap of Southeast Asian language nested characters through a region growing algorithm to generate a binary text sketch image;

[0012] Step 2, text encoding with integrated scene style information:

[0013] An automatic detection and conversion module is built in to standardize the input prompt text;

[0014] A pre-trained Southeast Asian language Embedding model is used to capture the context dependency relationship of Southeast Asian languages and generate semantic representations suitable for processing by the generation model;

[0015] Step3、Text-to-Image Generation Based on Attention Mechanism:

[0016] Transforming images into latent representations through Variational Autoencoder (VAE);

[0017] Designing a multi-scale latent space to encode global scene layout and local character details separately;

[0018] Placing scene text in the scene reasonably through conditional UNet and attention constraint mechanism;

[0019] Using the semantic representation converted from the prompt text as a text-level condition and image-level condition to guide the UNet to predict noise and gradually denoise through the reverse diffusion process, ultimately reconstructing high-quality images.

[0020] Further, the Step1 includes:

[0021] Step1.1、Performing preliminary processing on the input Southeast Asian language text I text and rendering it into a text sketch image I s . When using rendering techniques, ensure the visual integrity of combined characters; the model randomly selects a font and draws the text content in different shapes in black on a white background;

[0022] Step1.2、Using Canny operator to extract structured edge features of Southeast Asian language text to alleviate the edge breakage problem of Southeast Asian language curved characters, and filling the contour gaps of Southeast Asian language nested characters through region growing algorithm to obtain binary text sketch image containing key contour information; the extracted structured edge features are integrated into the conditional input of the control branch as important visual prior knowledge.

[0023] Further, the Step2 includes:

[0024] Step2.1、First, pre-process the input prompt text I prompt , including character segmentation, encoding conversion, and semantic embedding steps, to ensure that the text information can be accurately mapped into the latent space of the generation model;

[0025] To address the multiple encoding conflicts of Southeast Asian languages, an encoding automatic detection and conversion module is built in to ensure that the input prompt text I prompt is unified into a standardized Unicode format prompt text I unicode ;

[0026] Step2.2, a pre-trained Southeast Asian language Embedding model is used to capture the context-dependent relationship of Southeast Asian languages, map the Unicode encoded character sequence to a high-dimensional semantic vector, and map each character to a fixed-length vector; finally, these fixed-length vectors will prompt the text I unicode encoding to be converted into a word vector Emb input with practical significance, providing a high-quality semantic basis for subsequent image generation.

[0027] Further, in Step3, the transformation of the image into a latent representation by the variational autoencoder VAE includes:

[0028] A multi-scale latent space is designed, and then the global scene layout and local character details are encoded respectively. For an input image X0∈R H×W×3 , the variational autoencoder VAE transforms it into a latent representation Z0∈R h×w×c , where f = H / h = W / w is the down-sampling factor, and c is the dimension of the latent feature.

[0029] Further, in Step3, the reasonable placement of scene text in the scene by the conditional UNet and attention constraint mechanism includes:

[0030] A diffusion process is performed on the latent space, where a conditional UNet denoiser ε θ is used to predict noise with the current time step t, latent noise Z t , and generated condition information C; the generated condition information C is input into each cross-attention block i of the UNet model, and the calculation process is as follows:

[0031]

[0032] where d represents the output dimension of the key K and query Q features, φ i (Z t ) is the intermediate representation of the latent noise Z t implemented by the UNet, is a learnable matrix;

[0033] By introducing the attention constraint mechanism, in each forward propagation process at each time step, the cross-attention map of all layers of the diffusion model is traversed , where H and W are the latent space resolution, and d t represents the maximum length of the token; in this framework, the position i s of the text is user-specified or randomly generated by the sketch generation module; based on this, a mask image of the text region is generated using the bounding box, and it is defined as m box ∈R H×Wand down-sampled to the resolution of h x w, and the attention map is reconstructed; the specific formula is as follows:

[0034]

[0035] wherein, lambda represents a hyperparameter, represents the attention map corresponding to different text prompts, and N represents a set of label indexes corresponding to the words in the prompt that may contain text.

[0036] Further, in Step 3, the semantic representation converted from the provided input prompt text is used as a text-level condition together with an image-level condition to guide the UNet to predict noise and gradually denoise through a reverse diffusion process, and finally reconstruct a high-quality image, including:

[0037] The modified attention map M t_map participate in the prediction of noise Z t-1 at time t-1; at the same time, the provided input prompt text I prompt is processed by the encoder and used as a text-level condition together with an image-level condition to guide the UNet to predict noise epsilon θ ; the UNet branch predicts noise Z t at time t and uses the predicted noise Z t-1 at time t-1 to reconstruct the output image from the Gaussian noise, and the specific formula is as follows:

[0038] Z t =Z t-1 +UNet(Z t-1 );

[0039]

[0040] wherein, UNet(Z t-1 ) is the prediction output of the neural network (UNet) on the previous noise, x t is the image sample at time t, x t-1 is the image sample at time t-1, q(x t |x t-1 ) is the conditional probability of the forward diffusion process, N() is the Gaussian distribution, t is in [1, T], T is the total time step, beta t is in [0, 1], which is a preset noise scheduling parameter, and through the reverse diffusion process, the high-quality image is gradually denoised and finally reconstructed.

[0041] The application provides a multi-modal aspect sentiment analysis system based on prompt-based double-layer cross-modal distillation learning, which comprises a module for executing the controllable Southeast Asian text image generation method based on the conditional diffusion model.

[0042] The application provides an electronic device, comprising a memory, a processor and a computer program stored on the memory and executable on the processor, wherein the processor implements the controllable Southeast Asian text image generation method based on a conditional diffusion model when executing the program.

[0043] The application provides a non-transitory computer-readable storage medium, which stores a computer program, wherein the computer program is executed by a processor to implement the controllable Southeast Asian text image generation method based on a conditional diffusion model.

[0044] The application provides a computer program product, comprising a computer program, wherein the computer program is executed by a processor to implement the controllable Southeast Asian text image generation method based on a conditional diffusion model.

[0045] The application has the following beneficial effects:

[0046] 1. The application introduces an image contour sketch and local attention constraint method, so that the model can pay more attention to the details and local features of the text when generating an image, thereby improving the accuracy and readability of the generated image.

[0047] 2. The application uses image-level prompts to provide overall image style information, helping the model generate images that meet the expected visual effect.

[0048] 3. The application effectively improves the visual quality and structural consistency of the generated Southeast Asian language scene text image, which helps to improve the recognition performance of the recognition model in real scenes.

[0049] 4. Experimental results on a manually collected dataset show that the Myanmar text image generated by the application is closer to the image in a real scene, the FID score of the generated image is 20.42, and the sequence recognition accuracy is over 93%, effectively improving the generation effect of Southeast Asian language text images. BRIEF DESCRIPTION OF DRAWINGS

[0050] Figure 1 The figure is a method flowchart in the application.

[0051] Figure 2 The figure is a controllable Southeast Asian text image generation model framework based on a conditional diffusion model in the application.

[0052] Figure 3 The figure is an image contour information generation schematic diagram in the application.

[0053] Figure 4 The figure is an image sample generated by the current different generation methods. DETAILED DESCRIPTION

[0054] Example 1: as Figures 1-3As shown, the controllable Southeast Asian language text image generation method based on the conditional diffusion model comprises:

[0055] Step1, construct a Southeast Asian language text sketch image:

[0056] When using rendering technology, ensure the visual integrity of the combined characters;

[0057] Combine the adaptive threshold Canny operator to alleviate the edge breakage problem of Southeast Asian language curved characters;

[0058] Fill the contour gap of Southeast Asian language nested characters by region growing algorithm to generate binary text sketch image;

[0059] Step2, text encoding with fusion scene style information:

[0060] For the multi-coding conflict problem of Southeast Asian languages, an encoding automatic detection and conversion module is built in to standardize the input prompt text; ensure that the prompt text with fusion scene style information is unified into standardized Unicode format. For example, convert the Zawgyi encoding to Unicode Avoid semantic loss caused by garbled characters;

[0061] Use a pre-trained Southeast Asian language Embedding model to capture the context dependency of Southeast Asian languages and generate semantic representations suitable for the generation model;

[0062] Step3, text image control generation based on attention mechanism:

[0063] Transform the image into latent representation through variational autoencoder VAE;

[0064] Design a multi-scale latent space to encode global scene layout and local character details respectively;

[0065] Place the scene text in the scene reasonably through conditional UNet and attention constraint mechanism;

[0066] Use the semantic representation converted from the prompt text as the text-level condition and image-level condition to guide the UNet to predict noise and gradually denoise through the reverse diffusion process, and finally reconstruct high-quality images.

[0067] Further, the Step1 comprises:

[0068] Step1.1, the input Southeast Asian language text I text is processed and rendered into a text sketch image I s, when using rendering techniques, ensure the visual integrity of combined characters; for example, Burmese characters need to be placed precisely above the consonants to avoid semantic ambiguity caused by positional shifts; the model randomly selects a font and draws the text content in different shapes such as slanting and bending on a whiteboard background in black; this randomness not only increases the diversity of generated images, but also simulates various shapes that real-world text may take on;

[0069] Step1.2, use Canny operator to extract structured edge features of Southeast Asian language text, to alleviate the edge breakage problem of Southeast Asian language curved characters (such as Khmer characters ); and fill the contour gap of Southeast Asian language nested characters through region growing algorithm to obtain binary text sketch image containing key contour information; the extracted structured edge features are integrated into the conditional input of the control branch as important visual prior knowledge. In this way, the model can fully utilize edge information to guide the generation process, ensuring that the output image not only maintains the accuracy of the text semantics, but also accurately presents the expected visual style and structural features.

[0070] Further, the Step2 includes:

[0071] Step2.1, first preprocess the input prompt text I prompt , including character segmentation, encoding conversion and semantic embedding, etc., to ensure that the text information can be accurately mapped to the latent space of the generation model;

[0072] For the problem of encoding conflict of Southeast Asian languages (such as Zawgyi and Unicode of Burmese), an encoding automatic detection and conversion module is built in to ensure that the input prompt text I prompt is unified into a standardized Unicode format prompt text I unicode ; for example, convert Zawgyi encoded to Unicode to avoid semantic inconsistency caused by different encoding formats;

[0073] Step2.2, use a pre-trained Southeast Asian language Embedding model to capture the context dependency of Southeast Asian languages, map the Unicode encoded character sequence to a high-dimensional semantic vector, and map each character to a fixed-length vector; this process not only preserves the form and semantic information of the characters, but also further enhances the richness of semantic representation by capturing the context relationship between characters; finally, these fixed-length vectors will prompt the text I unicodeThe encoding is transformed into a meaningful word vector (Emb). input This provides a high-quality semantic foundation for subsequent image generation.

[0074] Furthermore, in Step 3, transforming the image into a latent representation using a variational autoencoder (VAE) includes:

[0075] Design multi-scale potential spaces, and then encode global scene layouts (such as billboard shapes) and local character details (such as Cambodian language) separately. (a ring structure), for the input image X0∈R H×W×3 The variational autoencoder (VAE) transforms it into a latent representation Z0∈R. h×w×c , where f=H / h=W / w is the downsampling factor, and c is the dimension of the latent feature.

[0076] Furthermore, in Step 3, placing the scene text appropriately within the scene using conditional UNet and attention constraint mechanisms includes:

[0077] A diffusion process is performed in the latent space, where a conditional UNet denoiser ε is used. θ To predict with current time step t and potential noise Z t The noise that generates the conditional information C is input into each cross-interest block i of the UNet model, and the calculation process is as follows:

[0078]

[0079] Where d represents the output dimension of the key K and query Q features, φ i (Z t The latent noise Z is implemented through UNet. t The middle representation, It is a learnable matrix;

[0080] By introducing an attention constraint mechanism, the cross-attention graphs of all layers of the diffusion model are traversed during each forward propagation at each time step. Where H and W are the potential spatial resolutions, and d t Indicates the maximum length of the token; within this frame, the position i of the text. s It is either user-specified or randomly generated by the sketch generation module; based on this, a mask image of the text region is generated using a bounding box, and defined as m. box ∈R H×W The attention map is then downsampled and adjusted to an h×w resolution to reconstruct the attention map; the specific formula is as follows:

[0081]

[0082] where λ denotes a hyper-parameter, denotes the attention map corresponding to different text prompts, and N denotes a set of label indexes corresponding to word pairs in the text that may be contained in the prompt.

[0083] Further, in Step 3, the semantic representation converted from the provided input prompt text is used as a text-level condition together with an image-level condition to guide the UNet to predict noise and gradually denoise through a reverse diffusion process, and finally reconstruct a high-quality image, including:

[0084] The modified attention map M t_map participates in the calculation of the predicted noise Z t-1 at time t-1; at the same time, the provided input prompt text I prompt is processed by the encoder and used as a text-level condition together with an image-level condition to guide the UNet to predict noise ε θ ; the UNet branch predicts the noise Z t at time t and uses the predicted noise Z t-1 at time t-1 to reconstruct the output image from the Gaussian noise, and the specific formula is as follows:

[0085] Z t =Z t-1 +UNet(Z t-1 );

[0086]

[0087] where UNet(Z t-1 ) is the predicted output of the neural network (UNet) on the previous noise, x t is the image sample at time t, x t-1 is the image sample at time t-1, q(x t |x t-1 ) is the conditional probability of the forward diffusion process, N() is the Gaussian distribution, t∈[1,T] is the total time step, β t ∈[0,1] is a preset noise scheduling parameter, and the image is gradually denoised through the reverse diffusion process and finally reconstructed to a high-quality image.

[0088] The application provides a multi-modal aspect sentiment analysis system based on prompt-based double-layer cross-modal distillation learning, and the system comprises:

[0089] A Southeast Asian language text sketch image module is constructed to realize the following functions:

[0090] When the rendering technology is adopted, the visual integrity of the combined characters is ensured.

[0091] The adaptive threshold Canny operator is combined to alleviate the edge fracture problem of Southeast Asian language curved characters.

[0092] Fill the contour gap of Southeast Asian language nested characters by a region growing algorithm to generate a binary text sketch image;

[0093] The text encoding module fuses scene style information, and is used for realizing the following functions:

[0094] The built-in encoding automatic detection and conversion module is used for standardizing the input prompt text;

[0095] The pre-trained Southeast Asian language Embedding model is adopted to capture the context dependency relationship of Southeast Asian languages, and semantic representation suitable for processing of a generation model is generated;

[0096] The text image control generation module based on an attention mechanism is used for realizing the following functions:

[0097] The image is transformed into a latent representation through a variational autoencoder VAE;

[0098] A multi-scale latent space is designed to encode global scene layout and local character details respectively;

[0099] The scene text is reasonably placed in the scene through a conditional UNet and an attention constraint mechanism;

[0100] The semantic representation converted from the prompt text is used as a text-level condition and an image-level condition to jointly guide the UNet to predict noise and gradually denoise through a reverse diffusion process, and finally reconstruct a high-quality image.

[0101] The application provides an electronic device, including a memory, a processor and a computer program stored on the memory and executable on the processor, wherein the processor implements the controllable Southeast Asian text image generation method based on a conditional diffusion model when executing the program.

[0102] The application provides a non-transitory computer readable storage medium, which stores a computer program, and the computer program is executed by a processor to implement the controllable Southeast Asian text image generation method based on a conditional diffusion model.

[0103] The application provides a computer program product, which includes a computer program, and the computer program is executed by a processor to implement the controllable Southeast Asian text image generation method based on a conditional diffusion model.

[0104] To verify the effect of the controllable Southeast Asian text image generation method based on the conditional diffusion model proposed in the present application, in view of the scarcity of Southeast Asian language image data, 10,000 Myanmar language text image data in real environment were collected through photographing, printing and other ways, and the obtained Myanmar language text image data were manually annotated. In order to obtain high-quality embedded visual text images, a data preprocessing process was implemented, aiming to finely screen the data and extract key image and position information.

[0105] Experimental setup: The model architecture of the present application integrates the core components of stable-diffusion-V15 and sd-controlnet-canny. In terms of optimization strategy, the Adam algorithm is used for parameter update, wherein the initial value of learning rate is set to 1×10-5, and the CosineAnnealing scheduling strategy is used to realize adaptive adjustment of learning rate. During the training process, the batch size is set to 128, and the total iteration number is set to 1,000,000 steps. To objectively evaluate the model performance, the image set similarity (Frechet Inception Distance, FID) of the test set is selected as the main evaluation index, and the model weight with the best performance during the training process is reserved as the final output.

[0106] Evaluation index: The image set similarity evaluation index is used in the experiment, as shown below:

[0107]

[0108] wherein μ r and μ g respectively represent the mean vectors of the real samples and the generated samples in the feature space, Σ r and Σ g are the covariance matrices of the two, used to describe the second-order statistical characteristics of the feature distribution.

[0109] Experiment 1: main experiment and result analysis

[0110] Table 2 is the main experimental results

[0111]

[0112] From the experimental results in Table 2, it can be concluded that:

[0113] The scores of the rule and template based methods are 348.06 and 298.39 respectively, indicating that there is a significant difference between the text images generated by these traditional methods and the text images in the real scene. The GAN based methods (DC-GAN and WGAN, scores are 114.64 and 128.13 respectively) show strong fitting ability and can generate samples close to the real data distribution, but there are problems of instability and mode collapse. The methods based on VAE and Flow-based Model (scores are 162.46 and 89.05 respectively) have advantages and disadvantages: VAE can better handle high-dimensional data but the generated quality is low; Flow-based Model is stable in training but has low computational efficiency and high requirements for data. The score of the method proposed in the application is 20.42, which proves the effectiveness of the method in generating high-quality text images.

[0114] Experiment two: comparative analysis of training effects of different generated data sets

[0115] In order to verify the training effect of the same recognition model on different generated data sets, two kinds of currently more mainstream recognition model frameworks were used for experiments, and the experimental results were expressed by sequence accuracy (SA) of the test set. As follows:

[0116]

[0117] Among them, SA, SL and LN respectively represent the string accuracy of Myanmar language text image recognition, the length of correctly recognized text image characters and the length of text.

[0118] Table 3 is the effect of the recognition model on different data sets

[0119]

[0120]

[0121] As shown in Table 3, the present experiment evaluates the influence of the quality of generated data on the performance of Myanmar language text image recognition by comparing the training effects of different generated data sets on two mainstream recognition models (VGG16+BiLSTM+CTC and Resnet50+Transformer). The experimental results show that the selection of the generation method has a significant impact on the performance of the model. The methods based on algorithm generation and template generation perform the worst, with sequence accuracy (SA) of 48.3% / 55.3% and 57.6% / 63.2%, respectively, mainly due to the lack of diversity and authenticity of the generated images. In contrast, the methods based on deep generation models significantly improve the performance of the model, with SA of DC-GAN, VAE and Flow-based Model reaching 83.2% / 85.6%, 79.5% / 82.4% and 80.7% / 83.1%, respectively, indicating that these methods can generate higher quality image data. In addition, WGAN, DiffusionModel and StableDiffusion perform better, indicating that diffusion models have a significant advantage in generating high-quality and diverse images. Furthermore, the method proposed in the present invention performs best among all the generated methods, with SA reaching 92.4% / 93.8%, proving that the data generated by the method proposed in the present invention is of better quality and closer to the Myanmar language text images in real scenarios.

[0122] The specific embodiments of the present application are described in detail above with reference to the accompanying drawings, but the present application is not limited to the above-described embodiments, and various changes can be made within the knowledge of those skilled in the art without departing from the spirit of the present application.

Claims

1. A method for controllable Southeast Asian text image generation based on a conditional diffusion model, characterized in that: The method includes: Step 1: Create a sketch image of the Southeast Asian language text: When using rendering techniques, ensure the visual integrity of the combined characters; By combining the adaptive threshold Canny operator, the edge breakage problem of curved characters in Southeast Asian languages ​​can be alleviated; A binary text sketch image is generated by filling the outline gaps of nested characters in Southeast Asian languages ​​using a region growing algorithm. Step 2: Perform text encoding to integrate scene style information: A built-in automatic encoding detection and conversion module is used to standardize the input prompt text; We employ a pre-trained Southeast Asian language embedding model to capture the contextual dependencies of Southeast Asian languages ​​and generate semantic representations suitable for generative models. Step 3: Text-image control generation based on attention mechanism: The image is transformed into a latent representation using a variational autoencoder (VAE). Design a multi-scale potential space to encode the global scene layout and local character details respectively; The scene text is placed appropriately within the scene using conditional UNet and attention constraint mechanisms; The semantic representation converted from the prompt text is used as a text-level condition and an image-level condition to guide UNet in predicting noise and performing a reverse diffusion process to gradually denoise the image, ultimately reconstructing a high-quality image.

2. The method of claim 1, wherein the method is based on a conditional diffusion model. Step 1 includes: Step 1.1, the input Southeast Asian language text I text is processed and rendered into a text sketch image I s . When using rendering technology, ensure the visual integrity of the combined characters; the model randomly selects a font and draws the text content in different shapes on a whiteboard background in black. Step 1.2: The Canny operator is used to extract structured edge features from Southeast Asian language texts to alleviate the edge breakage problem of curved characters in Southeast Asian languages. The region growing algorithm is used to fill the contour gaps of nested characters in Southeast Asian languages ​​to obtain a binary text sketch image containing key contour information. The extracted structured edge features are integrated into the conditional input of the control branch as important visual prior knowledge.

3. The method of claim 1, wherein the method further comprises: determining a condition diffusion model based on a plurality of Southeast Asian text images; and generating a plurality of Southeast Asian text images based on the condition diffusion model. Step 2 includes: Step 2.1, first pre-process the input prompt text I prompt including character segmentation, encoding conversion and semantic embedding steps, to ensure that the text information can be accurately mapped into the latent space of the generation model; For the problem of multi-coding conflict of Southeast Asian languages, an automatic detection and conversion module of coding is built in to ensure that the input prompt text I prompt is unified into the standardized Unicode format prompt text I unicode ; Step 2.2: Employ a pre-trained Southeast Asian language embedding model to capture the contextual dependencies of Southeast Asian languages. Map Unicode-encoded character sequences to high-dimensional semantic vectors, and map each character to a fixed-length vector. Ultimately, these fixed-length vectors will prompt the text I... unicode The encoding is transformed into a meaningful word vector (Emb). input This provides a high-quality semantic foundation for subsequent image generation.

4. The method of claim 1, wherein the method further comprises: In Step 3, transforming the image into a latent representation using a variational autoencoder (VAE) includes: A multi-scale latent space is designed, and then the global scene layout and local character details are encoded respectively. For the input image X0∈R H×W×3 , the variational autoencoder VAE transforms it into a latent representation Z0∈R h×w×c , where f=H / h=W / w is the down-sampling factor, and c is the dimension of the latent feature.

5. The method of claim 1, wherein the method further comprises: In Step 3, placing the scene text appropriately within the scene using conditional UNet and attention constraint mechanisms includes: A diffusion process is performed on the latent space, where a conditional UNet denoiser ε θ is used to predict a noise with current time step t, latent noise Z t and generated conditional information C; the generated conditional information C is input into each cross-attention block i of the UNet model, the calculation process is as follows: where d denotes the output dimension of the key K and query Q features, φ i (Z t ) is a latent noise Z t implemented by a UNet, is a learnable matrix; By introducing an attention constraint mechanism, in the process of each forward propagation, cross-attention maps of all layers of the diffusion model are traversed where H, W are latent space resolutions, d t denotes the maximum length of tokens; in this framework, the position i s is user-specified or randomly generated by the sketch generation module; based on this, a mask image of the text region is generated using the bounding box, and it is defined as m box ∈R H×W , and is adjusted to the resolution of h x w by downsampling, and the attention map is reconstructed; the specific formula is as follows: Where λ represents the hyperparameter. This represents the attention map corresponding to different text prompts, where N represents the set of tag indices corresponding to the words that may be contained in the text within the prompt.

6. The controllable Southeast Asian text image generation method based on the conditional diffusion model according to claim 1, characterized in that: In Step 3, the semantic representation of the provided input prompt text is used as both text-level and image-level conditions to guide UNet in predicting noise and performing a reverse diffusion process to gradually denoise the image, ultimately reconstructing a high-quality image, including: Corrected attention map The noise Z predicted at time t-1 t-1 The calculation; at the same time, the provided input prompt text I prompt After being processed by the encoder, these text-level conditions, together with the image-level conditions, guide UNet in predicting the noise ε. θ Noise Z at UNet branch prediction time t t And use time t-1 to predict the noise Z t-1 The specific formula for reconstructing the output image from Gaussian noise is as follows: Z t = Z t-1 + UNet(Z t-1 ); where UNet(Z t-1 ) is the predicted output of the neural network (UNet) on the previous noise, x t is the image sample at time t, x t-1 is the image sample at time t-1, q(x t | x t-1 ) is the conditional probability of the forward diffusion process, N() is a Gaussian distribution, t ∈ [1, T] is the total time step, and β t ∈ [0, 1] is a preset noise scheduling parameter. Through the inverse diffusion process, the high-quality image is finally reconstructed.

7. A multimodal aspect sentiment analysis system based on cue-driven two-layer cross-modal distillation learning, characterized in that, The system includes a module for performing the controllable Southeast Asian text image generation method based on the conditional diffusion model as described in any one of claims 1 to 6.

8. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, When the processor executes the program, it implements the controllable Southeast Asian text image generation method based on the conditional diffusion model as described in any one of claims 1 to 6.

9. A non-transitory computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, it implements the controllable Southeast Asian text image generation method based on the conditional diffusion model as described in any one of claims 1 to 6.

10. A computer program product, comprising a computer program, characterized in that, When the computer program is executed by the processor, it implements the controllable Southeast Asian text image generation method based on the conditional diffusion model as described in any one of claims 1 to 6.