Text image generation method based on structural semantic prompt constraint

Through the method of structural semantic prompt constraints, combined with multimodal feature coding and difficult-to-negative sample mining losses, the problem of insufficient generation quality and efficiency in text image generation is solved, and efficient and semantically consistent text image generation is achieved.

CN120298522APending Publication Date: 2025-07-11HARBIN INST OF TECH +1
View PDF 0 Cites 4 Cited by

Patent Information

Application Number
CN202510376583.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-28
Publication Date
2025-07-11

AI Technical Summary

Technical Problem

The existing text image generation technology has shortcomings in the generation quality and efficiency, especially the target control accuracy of the diffusion model and the slow iteration inference speed, the semantic consistency between the generated image and the text description, the visual performance of the generative adversarial networks, and the existing combination methods have failed to effectively improve the fine-grained consistency.

Method used

Using a method based on structural semantic cue constraints, a fine-grained visual concept is extracted from the text through a structured semantic cue generator, combined with a multimodal feature encoder and a relation-perceptual attention mechanism, visual feature adjustment is layered, and the training process is optimized using difficult-negative sample mining to match perceptual loss.

Benefits of technology

It significantly improves the semantic consistency and visual quality of the generated images, reduces the training resource requirements, and can quickly generate high-quality text images that meet the description requirements in an environment of limited computing resources.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120298522A_ABST
    Figure CN120298522A_ABST
Patent Text Reader

Abstract

The invention discloses a structure semantic prompt constraint-based text image generation method, which comprises the following steps of: taking a generative adversarial network as a generative model of a main body, and respectively learning and updating the generative model and a discrimination model by utilizing an existing sufficient training data set so as to finish model updating of the model and a training model; and on the basis of the generation model which completes model updating, a new image is generated by inputting the description text, so that the purpose of text image generation is achieved. According to the method, the consistency of the text and the image is remarkably enhanced while the image synthesis quality is improved. In order to reduce the dependence of the model on the training batch size and the training round, improve the training efficiency and optimize the resource demand, the difficult-to-load sample mining is introduced into the image generation task for the first time, and the difficult-to-load sample mining matching perception loss is proposed, so that the model can concentrate on the most challenging sample, and the image generation efficiency is improved. In practical application, good performance can be obtained even in an environment where computing resources are limited.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to a method for generating text images, and more particularly to a method for generating text images based on a pre-trained model with structural semantic prompt constraints. Background Art

[0002] Text-to-image generation technology aims to generate high-quality images that are semantically consistent with the described content according to conditional text. How to quickly synthesize text-related images has always been an important challenge. Although, in recent years, diffusion models have made significant progress in generation quality relying on large-scale data pre-training and model characteristics, they are also limited by problems such as insufficient target control accuracy and slow iterative inference speed, and cannot meet the requirements for efficiency and precise control in practical applications. Different from diffusion models, generative adversarial networks (GANs) have efficient generation capabilities, flexible controllability, and lower data requirements, but are lacking in visual performance.

[0003] To solve these problems, some methods attempt to combine generative adversarial networks with pre-trained models to achieve the purpose of quickly generating high-quality images, such as GigaGAN, GALIP, etc. However, these methods mainly focus on improving the visual quality of the generated images, and pay less attention to the fine-grained consistency between the generated images and text descriptions at the semantic level, resulting in the generated images being prone to missing details or semantic confusion. Summary of the Invention

[0004] To solve the above problems existing in the background art, the present invention provides a method for generating text images based on structural semantic prompt constraints, which captures fine-grained visual concepts from text and constructs structured semantic prompts to hierarchically guide the adjustment of visual features in the latent space. While improving the quality of image synthesis, this method significantly enhances the consistency between text and images. At the same time, in order to reduce the model's dependence on the training batch size and the number of training rounds, improve training efficiency, and optimize resource requirements, the present invention first introduces hard negative sample mining into the image generation task and proposes a hard negative sample mining matching perception loss, enabling the model to focus on the most challenging samples and achieving good performance even in environments with limited computing resources in practical applications.

[0005] The object of the present invention is achieved by the following technical solutions:

[0006] A method for generating text images based on structural semantic prompt constraints, comprising the following steps:

[0007] Step 1: Select a set of samples from a preset training sample set, each set of sample pairs contains a number of samples, and each sample contains a piece of text and a corresponding image;

[0008] Step 2: The text in the sample is fed into the structured semantic prompt generator, which uses a pre-trained grammar parser to extract and separate the text semantic concepts from the description content: entities, attributes, and relationships, and constructs a semantic scene graph;

[0009] Step 3: Combine the extracted text semantic concepts with the prompt words, and use a multi-modal feature encoder pre-trained on a large-scale data to encode, map the features of the text modality to the visual modality features, and replace the text features in the semantic scene graph;

[0010] Step 4: Use the relation-aware attention mechanism to update the semantic scene graph, construct a fine-grained local semantic prompt embedding, and use the global text feature encoding as the global semantic prompt embedding;

[0011] Step 5: Use the visual encoder to predict the coarse-grained image features from the condition-enhanced text encoding features;

[0012] Step 6: Concatenate the local semantic prompt embedding and the global semantic prompt embedding into a structural semantic prompt, and use it as the category embedding to respectively hierarchically prompt and guide the feature transformation adapter to hierarchically update the predicted coarse-grained image features, and the participation degree of each concept prompt is controlled by a self-attention gating structure;

[0013] Step 7: Feed the modulated image features into the image generator to synthesize the target image;

[0014] Step 8: Use the discriminator to calculate the generative adversarial loss between the synthesized image, the real image, and the conditional text in each sample, use the maximum matching perception loss of all sample pairs in a group of samples to replace the average matching perception loss, and pay special attention to the hard negative samples;

[0015] Step 9: Use the optimizer to update the parameters of the generative adversarial model according to the generative adversarial loss in Step 8;

[0016] Step 10: Repeat Step 1 to Step 9 until the model converges to a relative value, and use the trained generator to input the text to generate the required image.

[0017] Compared with the prior art, the present invention has the following advantages:

[0018] 1. The method using the structural semantic prompt constraint of the present invention fully considers all the semantic information in the description text. In particular, by separately extracting visual concepts and hierarchically adjusting, it can further focus on fine-grained attributes and greatly restore all visual elements in the text description.

[0019] 2. The present invention significantly improves the semantic consistency and visual quality between the synthesized image and the conditional text, and can quickly generate text images that meet the description requirements on the specified data.

[0020] 3. The hard negative sample-aware loss proposed by the present invention optimizes the training method of traditional image adversarial generation, enabling the rapid acquisition of a generation model on a given dataset with fewer training batches and training time, reducing the demand for large-scale computing resources, and enabling model training to be achieved in an environment with limited computing resources. BRIEF DESCRIPTION OF THE DRAWINGS

[0021] Figure 1 It is a schematic diagram of a training framework for text-image adversarial generation based on structural semantic prompt constraints;

[0022] Figure 2 It is a schematic diagram of the structure of the relationship-aware attention module;

[0023] Figure 3 It is a structural block diagram of a hierarchical prompt-guided feature transformation adapter;

[0024] Figure 4 It is a visual structural comparison between the method of the present invention and advanced text-image generation methods;

[0025] Figure 5 It is a training comparison of the method of the present invention on the COCO dataset with or without using hard negative sample matching-aware loss under different batches. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0026] The technical solutions of the present invention will be further described below with reference to the accompanying drawings, but are not limited thereto. Any modification or equivalent replacement of the technical solutions of the present invention without departing from the spirit and scope of the technical solutions of the present invention shall be covered by the protection scope of the present invention.

[0027] The present invention provides a text-image generation method based on structural semantic prompt constraints. The method uses a generative adversarial network as the main generative model, utilizes an existing sufficient training dataset, respectively learns and updates the models of the generative model and the discriminative model to complete the model update of the model and the training model, and then generates a new image by inputting a description text based on the generative model with the completed model update, so as to achieve the purpose of text-image generation. As Figure 1 shown, it specifically includes the following steps:

[0028] Step 1: Select a group of samples from a preset training sample set. Each group of sample pairs contains several samples, and each sample contains a piece of text and a corresponding image.

[0029] Step 2: The text in the sample is sent into a structured semantic prompt generator. The structured semantic prompt generator uses a pre-trained semantic parser to separate text semantic concepts: entities, attributes, and relationships from the description content, and constructs a semantic scene graph.

[0030] Step 3: Combine the semantic concepts of the extracted text with the prompt words, and use a multi-modal feature encoder pre-trained on a large-scale dataset to encode them, mapping the features of the text modality to the visual modality features and replacing the text features in the semantic scene graph.

[0031] Step 4: Use a relation-aware attention mechanism to update the semantic scene graph, construct a fine-grained local semantic prompt embedding, and use the global text feature encoding as the global semantic prompt embedding.

[0032] Step 5: Use an image feature predictor to predict the coarse-grained image features from the conditionally enhanced text encoding features.

[0033] Step 6: Concatenate the local semantic prompt embedding and the global semantic prompt embedding into a structural semantic prompt, and feed them into a hierarchical prompt-guided feature transformation adapter as class embeddings respectively to hierarchically update the predicted coarse-grained image features. The participation degree of each concept prompt is controlled by a self-attention gating structure (i.e., a gated attention affine layer).

[0034] Step 7: Feed the modulated image features into an image generator to synthesize the target image.

[0035] Step 8: Use a discriminator to calculate the generative adversarial loss between the synthesized image, the real image, and the conditional text in each sample. Replace the average matching perception loss with the maximum matching perception loss of all sample pairs in a set of samples, and pay special attention to hard negative samples.

[0036] Step 9: Use an optimizer to update the parameters of the generative adversarial model according to the generative adversarial loss in Step 8.

[0037] Step 10: Repeat Steps 1 to 9 until the model converges to a relative value. Use the trained generator to input the text to generate the required image.

[0038] In the present invention, the structured semantic prompt generator consists of a semantic parser for extracting entity concepts and generating a scene graph, a multi-modal feature encoder pre-trained on a large-scale text-image dataset for feature mapping, and a relation-aware attention unit for aggregating semantic information. Specifically: Given a sample text T, use the semantic parser to extract a set of fine-grained text semantic concepts (entity E, attribute A, relation R) and a semantic scene graph G (represented in the form of relation tuples) from the sample text T. These extracted text semantic concepts are concatenated with a padding prompt (such as "a photo of") and encoded into latent visual features f{f∈f E ,f A ,f R} by a pre-trained multi-modal feature encoder (CLIP-VIT), where f E represents the entity feature, fA Represents the attribute feature, f R Represents the entity relationship feature. These visual features contain all the detailed information required for the target image, and the structured information is maintained through the relational tuple. The relation-aware attention unit is used to effectively aggregate this information contained in the visual features to obtain complete conceptual cues, ensuring that the structured information of entities, attributes, and their mutual relationships is retained and accurately reflected in the synthesized image. The attribute feature f A Maps the affine parameters θ and γ through the multi-layer perceptron MLPs, and projects and adjusts the entity feature f E to obtain the adjusted entity feature f E ' with attribute information:

[0039] γ, θ = MLPs(f A )

[0040]

[0041] where, Represents element-wise addition, Represents element-wise multiplication.

[0042] To be able to participate in the calculation, the relational tuple in the semantic scene graph G is replaced by the form of a mask matrix to represent the relationship between entities, that is, the relationship mask matrix E mask . Multiply the relationship mask matrix E mask and the relationship feature f R to aggregate the relationship information between different entities and use it as the query of the cross-attention layer. Integrate the modulated entity relationship feature f R , and then through the self-attention layer and the semantic cue predictor, predict and generate the local semantic concept cue feature f p :

[0043] f h = Attention(Q(E Mask *f R ), K(f E '), V(f E '))

[0044] f p = Prerdictor(Attention(Q(f h ), K(f h ), V(f h )))

[0045] where, f hDenote the intermediate features through the cross-attention layer, where Q, K, and V represent query, key, and value respectively, Attention represents the attention layer, and Predictor represents the semantic prompt predictor. The feature prompt predictor is a set of fully connected mapping layers used to modulate the feature space of the aggregated semantic prompts. The global semantic prompt f g is generated by inputting the text encoding concatenated with the random noise z into a multi-layer perceptron:

[0046] f g = MLP(T, z)

[0047] In the present invention, the hierarchical prompt-guided feature transformation adapter includes a multi-layer vision Transformer encoding layer and a gated attention affine layer stacked in sequence. The structural semantic prompts are fed into the feature sequences of each layer of the Transformer as category Token embeddings according to the hierarchical order to adaptively adjust the visual image features:

[0048]

[0049] where f o represents the original image feature patch predicted by the generator from the noise text features, represents the structural semantic prompt, represents the image features through the nth layer of the Transformer, and the initial structure of the transformer module is inherited from the vision transformer of CLIP-VIT. The gated attention affine layer uses a gated attention mechanism to adjust the degree of participation of the structural semantic prompts in the adapter, control the influence of different concept information in the image generation process, and improve the consistency and accuracy between the text and the image:

[0050]

[0051] where represents the structural semantic prompt, and Gate represents the gated attention affine layer.

[0052] In the present invention, the hard negative sample mining perception loss is used to replace the average matching perception loss as the optimization objective for the generation task training, so as to better guide the generator and discriminator to move faster towards the Nash equilibrium direction during the training process. The formula is as follows:

[0053]

[0054] where z is a noise vector sampled from the N(0, 1) distribution, e is the sentence vector, where represents the mismatched sentence vector in the training batch, k and p are two hyperparameters used to balance the effectiveness of the gradient penalty, D represents the discriminator, and G represents the generator. Represents gradient derivation.

[0055] The present invention captures fine-grained visual concepts from text and constructs structured semantic cues to hierarchically guide the adjustment of visual features in the latent space. While improving the quality of image synthesis, the method significantly enhances the consistency between text and image. To reduce the model's dependence on the training batch size and the number of training epochs, the present invention introduces hard negative sample mining into the image generation task and proposes a hard negative sample mining matching-aware loss, enabling the model to focus on the most challenging samples and achieve good performance even under computational resource constraints. Compared with existing text-to-image generation methods, the present invention can quickly generate semantically consistent text images and achieve good training results with less training time and fewer training batches under low computing power resources.

[0056] Embodiment:

[0057] This embodiment provides a text-to-image generation method based on structural semantic cue constraints. The method uses a generative adversarial network as the main generative model, and uses an existing sufficient training dataset to learn and update the model of the generative model and the discriminative model respectively to complete the model update of the model and the training model. Then, based on the generative model that has completed the model update, new images are generated by inputting the description text, so as to achieve the purpose of text-to-image generation. As Figure 1 shown, it specifically includes the following steps:

[0058] Step S1: Preparation of the pre-trained model.

[0059] The pre-trained model adopts the structure of Vision Transformer (VIT) and is pre-trained using image-text pairs on a large dataset. The encoding of the image uses the method of image patches plus position encoding embedding. This embodiment is described by taking CLIP-VIT-B / 32 as an example.

[0060] Step S2: Construction of the structural semantic cue generator.

[0061] The structural semantic cue generator includes a semantic parser for extracting entity concepts and generating a scene graph, a multi-modal feature encoder pre-trained on a large-scale image-text dataset for feature mapping, and a relation-aware attention unit for aggregating semantic information.

[0062] In this embodiment, a number of image-text sample pairs are randomly sampled from the training dataset as a sample group with a training batch size of n, and fine-grained semantic concepts (such as entities E, attributes A, relations R) and semantic scene graphs are extracted from the sample text T using natural language processing tools. These extracted concepts are concatenated with padding cues (such as "a photo of") and the latent visual features f {f ∈ f are extracted through CLIP-VITE , f A , f R}. These visual features contain all the detailed information required for the target image and maintain the structured information through relational tuples. As Figure 2 shown, this embodiment uses a relation-aware attention mechanism to effectively aggregate this information to obtain conceptual cues, ensuring that the structured information of entities, attributes, and their interrelationships is retained and accurately reflected in the synthesized image.

[0063] γ, θ = MLPs(f A )

[0064]

[0065] f h = Attetion(Q(E Mask * f R ), K(f E '), V(f E '))

[0066] f p = Prerdictor(Attetion(Q(f h ), K(f h ), V(f h )))

[0067] Among them, represents element-wise addition, represents element-wise multiplication, and E Mask is a relational mask matrix generated from the relational tuples parsed from the semantic scene graph. The feature predictor is a set of fully connected mapping layers for modulating the feature space. In addition, the noise text encoding used to predict the initial image features is input into a multi-layer perceptron (MLP) to form the global semantic cue f g :

[0068] f g = MLP(T, z)

[0069] This structural design significantly improves the detail performance and semantic consistency of the generated image while retaining the global semantic information by combining local (fine-grained) and global (coarse-grained) semantic information.

[0070] Step S3: Hierarchical embedding of semantic cues guides image feature prediction.

[0071] To make full use of the fine-grained information aggregated by these structured semantic cues, this embodiment treats them as class tokens in a Vision Transformer and hierarchically embeds them into each module of the hierarchical cue-guided feature transformation adapter. These cues enable the model to precisely fuse text information when processing image features by adjusting the patch tokens of the predicted image, enhancing the correlation between the generated image and the text, and thus generating an image that is highly consistent with the text description.

[0072] The specific structure of the hierarchical cue-guided feature transformation adapter is as Figure 3 shown. The structural semantic cues are embedded as class tokens into the feature sequences of each layer of the Transformer in hierarchical order to adaptively adjust the visual image features:

[0073]

[0074] where f o represents the original image feature patches predicted by the generator from the noise text features, and the initial structure of the transformer module inherits from the visual transformer of CLIP-VIT. Considering that the participation degrees of the structured visual concepts aggregated from different text descriptions may vary, this embodiment adopts a gated attention mechanism to regulate the participation degrees of these aggregated concepts in the hierarchical cue-guided feature transformation adapter. In this way, the model can flexibly control the influence of different text information in the image generation process, thereby further improving the consistency and accuracy between the text and the image.

[0075]

[0076] Step S4: Adversarial generation training based on hard negative sample-aware loss.

[0077] To accelerate the training speed of the generative adversarial network and mitigate the accuracy loss caused by reducing the training data batch size, this embodiment introduces the hard negative sample mining loss into the generation task and redefines the training optimization objective formula as follows:

[0078]

[0079] where z is a noise vector sampled from the N(0,1) distribution, and e is the sentence vector where, Denote the unmatched sentence vectors in the training batch. k and p are two hyperparameters used to balance the effectiveness of gradient penalty. In a training with a batch size of n, replace the original average loss with the matching-aware loss of hard negative samples to focus on the most challenging samples in each batch, thus better guiding the generator and discriminator to develop towards the Nash equilibrium during training and improving the learning efficiency of the model.

[0080] In this embodiment, two datasets, the multi-object dataset MS-COCO and the single-object dataset CUB_BIRD, are used for training and testing the model. The pre-trained model is implemented using Pytorch. The model is trained using 4×Nvidia RTX 3090 GPUs and tested using 1×Nvidia RTX 3090 GPU. The default training batch is 32, and the ADAM optimizer is used for training with an initial learning rate of 1e -5 。

[0081] In the performance test stage, two main evaluation metrics are used for comparison, including the Frechet Inception Distance (FID) for measuring the generation quality and the CLIP Score (CS) for measuring the text-image semantic consistency. Table 1 shows the performance comparison between the present invention and other text-image generation methods on the CUB and COCO datasets.

[0082] Table 1 Performance comparison with other generative models on the CUB and COCO test sets

[0083]

[0084] The experimental results in Table 1 show that the model of the present invention exhibits significant improvements both on the single-object dataset CUB and the multi-object dataset COCO, and achieves state-of-the-art performance in terms of both image quality and text-image consistency. Especially in terms of text-image consistency, compared with other GAN- and CLIP-based methods, the method of the present invention has a significant improvement, improving by 1.26 and 1.69 respectively compared to the second place on the two datasets, reaching 33.36 and 35.07.

[0085] Figure 4Shows a visual comparison between the method of the present invention and state-of-the-art GAN models, GANs combined with pre-trained models, and well-received diffusion models. Apparently, compared with other methods that combine GANs with pre-trained models, the method of the present invention not only improves the visual quality of the generated images, but also significantly enhances the semantic consistency between the images and the text descriptions. The model of the present invention can capture the fine-grained visual concepts mentioned in the text and accurately reflect these concepts in the synthesized images, effectively alleviating common problems in other methods, such as concept confusion, concept loss, and attribute misalignment (for example, in the first row, "black spotted primaries" is not the same as "black primaries", and in the fourth row, "orange juice" is also different from "orange and juice"). Compared with diffusion models, although the model of the present invention is slightly lacking in visual quality, it performs better in capturing the visual concepts in the text descriptions, and the generated images are more similar to the original images, which is more crucial for certain tasks. In addition, the method of the present invention demonstrates a faster synthesis speed (SD: 3.84s, Ours: 0.05s) and requires fewer computing resources, as shown in Table 2.

[0086] Table 2 Comparison of training resource requirements and inference performance with other text-image generation models

[0087]

[0088] Figure 5 Shows the impact of the hard sample matching-aware loss in reducing training resource requirements and accelerating the training process of the generation model. Compared with the training without using the hard sample matching-aware loss, the model using this loss converges approximately 400 training epochs earlier when the batch size is 32. In addition, for the model with a batch size of 16, using the hard sample matching-aware loss significantly reduces the performance gap with the model with a batch size of 32. This indicates that the proposed hard sample matching-aware loss of the present invention can not only significantly accelerate the convergence speed of the model, but also achieve comparable or equivalent performance under limited training resources, especially in the training of smaller batch data. This makes the training process more efficient and enables good performance to be obtained in a shorter time.

Claims

1. A text image generation method based on structural semantic prompt constraints, characterized in that The method includes the following steps: Step 1: Select a set of samples from a preset training sample set. Each set of sample pairs contains several samples, and each sample contains a piece of text and a corresponding image; Step 2: The text in the sample is fed into a structured semantic prompt generator. The structured semantic prompt generator uses a pre-trained syntax parser to extract and separate text semantic concepts from the description content: entities, attributes, and relationships, and constructs a semantic scene graph; Step 3: Combine the extracted text semantic concepts with prompt words, and use a multi-modal feature encoder pre-trained on a large-scale data set to encode, map the features of the text modality to the visual modality features, and replace the text features in the semantic scene graph; Step 4: Use a relationship-aware attention mechanism to update the semantic scene graph, construct a fine-grained local semantic prompt embedding, and use the global text feature encoding as the global semantic prompt embedding; Step 5: Use a visual encoder to predict coarse-grained image features from the conditionally enhanced text encoding features; Step 6: The local semantic prompt embedding and the global semantic prompt embedding are concatenated into a structured semantic prompt, which is fed into a hierarchical prompt-guided feature transformation adapter as category embeddings respectively to hierarchically update the predicted coarse-grained image features. The participation degree of each concept prompt is controlled by a self-attention gating structure; Step 7: Feed the modulated image features into an image generator to synthesize the target image; Step 8: Use a discriminator to calculate the generative adversarial loss between the synthesized image, the real image, and the conditional text in each sample. Use the maximum matching perception loss of all sample pairs in a set of samples to replace the average matching perception loss, and pay special attention to hard negative samples; Step 9: Use an optimizer to update the parameters of the generative adversarial model according to the generative adversarial loss in Step 8; Step 10: Repeat Step 1 to Step 9 until the model converges to a relative value. Use the trained generator to input text to generate the required image.

2. The text image generation method based on structural semantic prompt constraints according to claim 1, wherein In the said Step 2, the structured semantic prompt generator consists of a semantic parser for extracting entity concepts and generating a scene graph, a multi-modal feature encoder pre-trained on a large-scale text-image data set for feature mapping, and a relationship-aware attention unit for aggregating semantic information; Given a sample text T, a set of fine-grained text semantic concepts are extracted from the sample text T using a semantic parser: entity E, attribute A, relation R, and semantic scene graph G. The extracted text semantic concepts are concatenated with the padding prompt and encoded into latent visual features f{f∈f E ,f A ,f R}, where f E represents entity features, f A represents attribute features, f R represents entity relation features. The relation-aware attention unit is used to aggregate the information contained in the visual features to obtain a complete conceptual prompt, ensuring that the structured information of entities, attributes, and their mutual relations is retained and accurately reflected in the synthesized image.

3. The method for generating a text image based on structural semantic hint constraints according to claim 2, wherein In the relation-aware attention unit, the attribute feature f A maps the affine parameters θ and γ through multi-layer perceptrons (MLPs), and projects and adjusts the entity feature f E to obtain the adjusted entity feature f E ' with attribute information: γ,θ = MLPs(f A ) Among them, represents element-wise addition, represents element-wise multiplication. Using the relation mask matrix E mask Replace the relation tuples in the semantic scene graph G that represent the relationships between entities with the relation mask matrix E mask and the relation feature f R Multiply them to aggregate the relationship information between different entities, and use it as the query of the cross-attention layer. For the modulated entity relation feature f R Integrate it, and then through the self-attention layer and the semantic prompt predictor, predict and generate the local semantic concept prompt feature f p : f h = Attention(Q(E Mask * f R ), K(f E '), V(f E ')) It should be noted that the original text seems to have some errors or be an incomplete or unclear formula. The translation is done strictly according to the rules while maintaining the original text's integrity. f p = Prerdictor(Attetion(Q(f h ),K(f h ),V(f h ))) Among them, f h represents the intermediate feature passed through the cross-attention layer, Q, K, and V respectively represent the query, key, and value, Attention represents the attention layer, and Predictor represents the semantic prompt predictor; Global semantic hint f g Generated by inputting the text encoding concatenated with random noise z into a multi-layer perceptron: f g = MLP(T, z).

4. The text image generation method based on structural semantic hint constraints according to claim 1, characterized in that In the said Step 6, the hierarchical prompt-guided feature transformation adapter includes a stack of multiple visual Transformer encoding layers and gated attention affine layers in sequence. The structured semantic prompt is fed into the feature sequence of each layer of Transformer as a category Token embedding according to the hierarchical order to adaptively adjust the visual image features; Among them, f o represents the original image feature patch predicted by the generator from the noise text feature, represents the structural semantic hint, represents the image feature passing through the n-th layer of the Transformer; the gated attention affine layer adopts a gated attention mechanism to adjust the degree of participation of the structural semantic hint in the adapter, control the influence of different concept information in the image generation process, and improve the consistency and accuracy between the text and the image: Among them, represents structural semantic cues, and Gate represents the gated attention affine layer.

5. The text image generation method based on structural semantic prompt constraints according to claim 1, characterized in that In the said Step 9, use the hard negative sample mining perception loss as the optimization objective for generative task training, and its formula is as follows: where z is a noise vector sampled from an N(0,1) distribution, e is the sentence vector, denotes the mismatched sentence vectors in the training batch, k and p are two hyperparameters used to balance the effectiveness of the gradient penalty, D denotes the discriminator, and G denotes the generator, denotes the gradient derivative.

Citation Information

Cited By

  • Text image generation method based on wavelet representation and domain adaptive strategy

    CN120876643A

  • Text-to-image generation method and system based on spatial perception decoupling

    CN120894456A

  • Text-to-image generation method and system based on spatial perception decoupling

    CN120894456B

  • Depth supervision type polyp segmentation improvement method, system and equipment based on prompt guidance and medium

    CN121437550A