Text-to-image generation method and device, equipment and medium

Through the text-to-image generation method, the scaling and adjustment module of the pre-training module is used to solve the problem of training instability of the pre-training model, and achieve more efficient and stable image generation.

CN119991837APending Publication Date: 2025-05-13BEIJING KNOWLEDGE ATLAS TECHNOLOGY CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202311428637.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2023-10-31
Publication Date
2025-05-13

AI Technical Summary

Technical Problem

The pre-trained model has instability during the training process, and the existing solutions are complex, time-consuming and computational resources are expensive.

Method used

A text-to-image generation method is adopted to obtain the images and description text in the dataset, perform word metamorphism processing, and input vectorized text lexicons and image lexicons equally into the pre-training module for training to obtain an image generation model. The model sets a scaling module in the normalization layer to avoid the output value explosion, and sets a adjustment module in the attention mechanism to prevent the output value from overflowing.

Benefits of technology

Improves the stability of the training process, reduces the consumption of computing resources, and generates the image that best matches the text through post-processing processes such as adjusting super resolution and reordering.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119991837A_ABST
    Figure CN119991837A_ABST
Patent Text Reader

Abstract

The invention provides a text-to-image generation method and device, equipment and a medium. The text-to-image generation method comprises the following steps: acquiring an image in a data set and a corresponding description text; performing lexical element processing on the description text and the image to obtain vectorized text lexical elements and vectorized image lexical elements; equally inputting the vectorized text lexical units and the image lexical units into a pre-training module for training to obtain an image generation model; and inputting a test text into the image generation model, and generating an image corresponding to the test text. According to the method, large-scale generation combination and training of texts and images can be realized, and the problem of instability in a large-scale text-to-image generation pre-training process is solved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of computer vision, and more particularly to a method, device, equipment and medium for generating text into an image. Background Art

[0002] In recent years, with the development of autoregressive generative models, large-scale pre-trained language models represented by generative pre-trained models (GPT) have demonstrated amazing strength in various natural language processing tasks, and have also opened up new research directions in the field of vision. However, due to the huge number of pixels in visual images, it is impossible to build a fine-grained sequence model. The current largest model, ImageGPT, is only trained at an image resolution of 96×96.

[0003] Traditional pre-trained models may have overflow and underflow during training, which may lead to unstable training. The existing solution proposes a mixed precision architecture for each residual loss scale, and stores all gains, biases, embedded and non-embedded vectors with 32-bit precision instead of the usual 16-bit precision. However, this method has the problems of complexity, time-consuming and high consumption of computing resources. Therefore, the present invention provides a method, device, equipment and medium for text-to-image generation. Summary of the invention

[0004] The present invention is proposed based on the above-mentioned needs of the prior art. The technical problem to be solved by the present invention is to provide a method, device, equipment and medium for text-to-image generation in view of the instability of the pre-trained model during the training process and the problems that the existing solutions are complex, time-consuming and consume a lot of computing resources.

[0005] In order to solve the above problems, the present invention is implemented by adopting the following technical solutions:

[0006] A first aspect of the present invention provides a method for generating text from an image, the method comprising:

[0007] Get the images and corresponding description texts in the dataset;

[0008] Performing tokenization processing on the description text and the image to obtain vectorized text tokens and image tokens;

[0009] The vectorized text word unit and the image word unit are equally input into a pre-training module for training to obtain an image generation model; the normalization layer in the pre-training module includes a scaling module, and the scaling module is used to scale the input of the normalization layer;

[0010] The test text is input into the image generation model to generate an image corresponding to the test text.

[0011] Furthermore, the scaling module is used to scale the input of the normalization layer, including:

[0012] The maximum value of the original input of the normalization layer is obtained, and the input is scaled by the maximum value, which is expressed as: f(x)=x / max(x), where x represents the original input of the normalization layer, max(x) represents the maximum value of the original input of the normalization layer, and f(x) represents the output of the scaling module.

[0013] Furthermore, the attention mechanism in the pre-training module includes an adjustment module; the adjustment module is used to adjust the output of the attention mechanism, including:

[0014] The output of the attention mechanism is adjusted according to the softmax function and the calculated query vector Q and the queried vector K, as shown below:

[0015]

[0016] In the formula, represents the output of the attention mechanism, d represents the dimension of vectors Q and K, and α represents a value greater than 32. 5 The constant.

[0017] Furthermore, the lemma processing of the description text and the image includes:

[0018] Inputting the description text into a text encoder to obtain the text word unit;

[0019] The image is input into a discrete image encoder to obtain the image word unit.

[0020] Furthermore, the step of inputting the test text into the image generation model and generating an image corresponding to the test text further includes a post-processing process; the post-processing process includes adjusting super-resolution and re-ordering.

[0021] Further, adjusting the super-resolution includes:

[0022] The low-resolution image is enhanced to a high-resolution image according to the super-resolution adjustment strategy.

[0023] Further, the reordering includes:

[0024] The correlation between the image and the text is evaluated based on the description text score, and the image that best matches the text is selected.

[0025] A second aspect of the present invention provides a text-to-image generation device, comprising:

[0026] The text image acquisition module is used to obtain the images and corresponding description texts in the data set;

[0027] A text image processing module, used for performing word-meta processing on the description text and image to obtain vectorized text word-meta and image word-meta;

[0028] A model training module, which equally inputs the vectorized text word unit and the image word unit into a pre-training module for training to obtain an image generation model; the normalization layer in the pre-training module includes a scaling module, and the scaling module is used to scale the input of the normalization layer;

[0029] The text-to-image generation module inputs the test text into the image generation model to generate an image corresponding to the test text.

[0030] According to a third aspect of the present invention, there is provided an electronic device, comprising a memory and a processor, wherein the memory stores a computer program, and when the computer program is executed by the processor, the method for generating text to image as described in any one of the embodiments is implemented.

[0031] A fourth aspect of the present invention provides a computer-readable storage medium having a computer program stored thereon, wherein the computer program, when executed by a processor, implements the method for generating text to image as described in any one of the embodiments.

[0032] Compared with the prior art, the present invention can achieve at least one of the following beneficial effects:

[0033] (1) By using a discrete image encoder to tokenize the image, the image is compressed into a low-dimensional discrete latent space, thereby achieving large-scale generative joint pre-training of description text and image;

[0034] (2) By setting a scaling module in the normalization layer to scale the input of the normalization layer, the output value of the normalization layer is prevented from exploding, thereby improving the stability of the training process;

[0035] (3) By setting an adjustment module in the second attention mechanism, the output of the second attention mechanism is appropriately adjusted to prevent the output value of the second attention mechanism from overflowing, thereby further improving the stability of the training process;

[0036] (4) By performing post-processing operations such as adjusting super-resolution and re-ranking before generating the image, the correlation between the image and the text is evaluated based on the description text score, so that the image that best matches the text is selected as the generated image.

[0037] In the present invention, the above-mentioned technical solutions can also be combined with each other to achieve more preferred combination solutions. Other features and advantages of the present invention will be described in the subsequent description, and some advantages can become obvious from the description, or can be understood by practicing the present invention. The purpose and other advantages of the present invention can be realized and obtained through the contents particularly pointed out in the description and the drawings. BRIEF DESCRIPTION OF THE DRAWINGS

[0038] In order to more clearly illustrate the embodiments of this specification or the technical solutions in the prior art, the drawings required for use in the embodiments or the description of the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments recorded in the embodiments of this specification. For ordinary technicians in this field, other drawings can also be obtained based on these drawings.

[0039] Figure 1 A flowchart of a method for generating text to image is provided for the first embodiment of the present invention;

[0040] Figure 2 A schematic diagram of the structure of the Transformer model provided in the first embodiment of the present invention;

[0041] Figure 3 A schematic diagram of the pre-normalization strategy structure provided in the first embodiment of the present invention;

[0042] Figure 4 A schematic diagram of a device for generating text to image provided in Embodiment 2 of the present invention;

[0043] Figure 5 This is a schematic diagram of the electronic device architecture provided in Embodiment 3 of the present invention. DETAILED DESCRIPTION

[0044] In order to make the purpose, technical solutions and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are part of the embodiments of the present invention, not all of the embodiments. It should be noted that, in the absence of conflict, the embodiments and features in the embodiments disclosed in this disclosure can be combined, separated, interchanged and / or rearranged with each other. Based on the embodiments in the present invention, all other embodiments obtained by ordinary technicians in this field without making creative work are within the scope of protection of the present invention.

[0045] Embodiment 1

[0046] The following is an introduction to the text-to-image generation method provided by the present invention through a specific embodiment. Figure 1 The method for generating text to image provided by the first aspect embodiment includes:

[0047] S1. Get the images and corresponding description texts in the dataset.

[0048] In some embodiments, a dataset of images and corresponding texts is obtained by crawling or collecting public datasets, and the dataset is manually screened to remove unclear text and image data, and to modify text that inaccurately describes the image.

[0049] S2. Performing tokenization processing on the description text and image to obtain vectorized text tokens and image tokens.

[0050] Furthermore, S2 includes:

[0051] S2.1, input the description text into the text encoder to obtain text words;

[0052] S2.2. Input the image into a discrete image encoder to obtain image word units.

[0053] Optionally, SentencePiece, BPE, WordPiece, or subword can be used as a text encoder to tokenize the description text.

[0054] Preferably, SentencePiece is used as a text encoder to tokenize the description text. SentencePiece contains four main components: normalizer, trainer, encoder, and decoder. The normalizer is used to normalize semantically equivalent Unicode characters into a canonical form; the trainer trains a subword segmentation model from a standardized corpus; the encoder executes the normalizer internally to standardize the input text and tokenizes it into a subword sequence using the subword model trained by the trainer; the decoder converts the subword sequence into standardized text.

[0055] Compared with the other text encoders mentioned above, SentencePiece has the advantages of being language independent, supporting multiple algorithms, being fast and lightweight.

[0056] Optionally, the image encoder can be a variational autoencoder (VAE), a convolutional autoencoder (CAE) or a denoising autoencoder (DAE), etc. The image encoder is used to encode the image into different attributes one by one, and then comprehensively decode these attributes to obtain a reconstructed image.

[0057] Preferably, a vector quantized variational autoencoder (VQ-VAE) is used as the image encoder to tokenize the image.

[0058] Specifically, VQ-VAE compresses the image into a low-dimensional discrete latent space by training the encoder, and then reconstructs the image from the latent space by training the decoder. More specifically, the encoder maps the image of shape 256×256×3 to an intermediate code of shape 32×32×256, then maps the intermediate code to a quantized vector in the codebook through nearest neighbor search, and finally reconstructs the quantized vector into an image through the decoder.

[0059] The method of using VQ-VAE to lemmatize images provided in this embodiment loses less fidelity than direct downsampling, while maintaining the spatial correlation of pixels and being able to process images with more pixels.

[0060] S3. The vectorized text word unit and the image word unit are equally input into a pre-training module for training to obtain an image generation model.

[0061] Optionally, the vectorized text words and image words are input into a pre-training module in a first order for training.

[0062] Optionally, this embodiment may also perform training by inputting the vectorized text words and image words into the pre-training module in a second order that reverses the first order.

[0063] The training method provided in this embodiment, in which vectorized text words and image words are input into a pre-training module in a second order that reverses the first order for training, can be compared with the image generation model obtained by training in the first order, and further image words and text words with lower correlation can be separated, thereby improving the accuracy of the image generation model.

[0064] Specifically, before training, four delimiters representing the beginning of text, the end of text, the beginning of image, and the end of image are used to indicate the boundaries of text and image.

[0065] Optionally, the pre-training module can be a model such as a convolutional neural network, a recurrent convolutional neural network, and a generative adversarial network.

[0066] Preferably, a unidirectional Transformer model is used as a pre-training module for training vectorized text words and image words.

[0067] Specifically, Figure 2 As shown in Figure 1, the above Transformer model consists of multiple encoder modules and decoder modules.

[0068] Specifically, the encoder module includes a first attention mechanism, a first feedforward neural network and a first normalization layer; the first attention mechanism is composed of multiple self-attention mechanisms; the first feedforward neural network is composed of a two-layer fully connected layer; the first normalization layer is composed of a residual module and a normalization module.

[0069] Specifically, the decoder module includes a second attention mechanism, a third attention mechanism, a second feedforward neural network, an adjustment module and a second normalization layer; the second attention mechanism is composed of multiple self-attention mechanisms after masking operations; the third attention mechanism is composed of multiple self-attention mechanisms and an adjustment module; the second feedforward neural network is composed of the same as the first feedforward neural network; the second normalization layer is composed of a residual module, a scaling module and a normalization module.

[0070] Specifically, a scaling module is added to the normalization layer to appropriately scale the input of the normalization layer, thereby appropriately adjusting the output of the normalization layer; the scaling module can be configured to any one of the above-mentioned normalization layers or multiple normalization layers, and the scaling module can be configured as a constant or a function that changes with the original input.

[0071] Preferably, the scaling module is configured in the second normalization layer and configured as the maximum value of the original input of the second normalization layer, for scaling the input of the second normalization layer, including:

[0072] Get the maximum value of the original input of the second normalization layer, and scale the input of the second normalization layer by using the maximum value, which is expressed as: f(x)=x / max(x), where x represents the original input of the second normalization layer, max(x) represents the maximum value of the original input of the second normalization layer, and f(x) represents the output of the scaling module, i.e., the input of the second normalization layer after scaling.

[0073] This embodiment provides that by configuring the scaling module in the second normalization layer and configuring it as the maximum value of the original input of the second normalization layer, the input of the second normalization layer can be effectively scaled, and then the output of the second normalization layer can be scaled, which can prevent the output value of the second normalization layer from suddenly increasing and causing overflow, thereby maintaining the stability of the training process.

[0074] Specifically, the output of the attention mechanism is appropriately adjusted by adding an adjustment module in the attention mechanism; the adjustment module can be configured into any one of the above-mentioned attention mechanisms or multiple attention mechanisms, and the adjustment module can be configured as a constant or a function that varies with the original input.

[0075] Preferably, the adjustment module is configured in the third attention mechanism to adjust the output of the third attention mechanism, including:

[0076] The query vector Q and the queried vector K are calculated based on the input matrix and the linear transformation matrix, and then the output of the third attention mechanism is adjusted to:

[0077]

[0078] Where d represents the dimension of the query vector Q and the queried vector K, and α represents a value greater than 32. 5 The constant.

[0079] This embodiment provides that by configuring the adjustment module into the third attention mechanism, the output of the third attention mechanism can be effectively adjusted, thereby preventing the output value of the third attention mechanism from overflowing, thereby further maintaining the stability of the training process.

[0080] Optionally, the pre-training module may also be configured to suppress the inputs of the attention mechanism and the feedforward neural network in the encoder module and the decoder module using a pre-normalization strategy.

[0081] Preferably, a Sandwich LayerNorm strategy is used as the pre-normalization strategy.

[0082] Specifically, Figure 3 As shown, a) is the original post-normalization structure, b) is the mainstream pre-normalization structure, and c) is the sandwich hierarchical normalization structure provided in this embodiment.

[0083] The sandwich hierarchical normalization structure provided in this embodiment adds a normalization layer at the end of the attention mechanism and the feedforward neural network to suppress the input of the attention mechanism and the feedforward neural network in the encoder module and the decoder module within a reasonable range, which is conducive to better convergence of the model and maintaining the stability of the training process.

[0084] S4. Input the test text into the image generation model to generate an image corresponding to the test text.

[0085] Optionally, before inputting the test text into the image generation model and generating an image corresponding to the test text, a post-processing process is also included; the post-processing process includes adjusting super-resolution and re-ranking. The post-processing process may also include operations such as style transfer and image annotation.

[0086] Specifically, the adjusting super-resolution includes enhancing the low-resolution image into a high-resolution image according to a super-resolution adjustment strategy. The adjustment strategy may be an adjustment strategy based on interpolation or an adjustment strategy based on deep learning.

[0087] Preferably, the super-resolution of the image is adjusted using a central sliding window strategy, including:

[0088] Set a window of appropriate size and step length, slide the preset window in the center of the image according to the preset step length, and enhance the resolution of the image block by block.

[0089] The super-resolution of an image is adjusted through a central sliding window strategy provided in this embodiment, so that a low-resolution image can be enhanced into a high-resolution image, and the image can be made clearer without losing image information.

[0090] Specifically, the re-ranking includes evaluating the relevance between the image and the text according to the description text score, and selecting the image that best matches the text.

[0091] Optionally, the CapLoss function is used to describe the degree of association between the image and the text, as shown below:

[0092]

[0093] In the formula, x represents the image word sequence, t is the text word sequence, and p(t i |x,t 0:i-1 ) represents the conditional distribution of text words based on image words.

[0094] Embodiment 2

[0095] The present invention provides a device for generating text to image, such as Figure 4 As shown, including:

[0096] The text image acquisition module is used to obtain the images and corresponding description texts in the data set;

[0097] A text image processing module, used for performing word-meta processing on the description text and image to obtain vectorized text word-meta and image word-meta;

[0098] A model training module, used for inputting the vectorized text word unit and the image word unit equally into a pre-training module for training to obtain an image generation model; wherein the pre-training module includes a plurality of encoder modules and a decoder module; the normalization layer in the decoder module includes a scaling module; the scaling module is used for scaling the input of the normalization layer;

[0099] The text-to-image generation module inputs the test text into the image generation model to generate an image corresponding to the test text.

[0100] Embodiment 3

[0101] The present invention provides an electronic device, such as Figure 5As shown, it includes a memory and a processor, the memory stores a computer program, and when the computer program is executed by the processor, the training method of the packaging design language model as described in any of the above embodiments is implemented.

[0102] Embodiment 4

[0103] The present invention provides a computer-readable storage medium having a computer program stored thereon. When the computer program is executed by a processor, the training method of the packaging design language model as described in any of the above embodiments is implemented.

[0104] Computer readable storage media include permanent and non-permanent, removable and non-removable media that can be implemented by any method or technology to store information. Information can be computer readable instructions, data structures, program modules or other data. Examples of computer storage media include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technology, compact disk read-only memory (CD-ROM), digital versatile disk (DVD) or other optical storage, magnetic cassettes, magnetic tape magnetic disk storage or other magnetic storage devices or any other non-transmission media that can be used to store information that can be accessed by a computing device. As defined herein, computer readable media does not include temporary computer readable media (transitory media), such as modulated data signals and carrier waves.

[0105] The professionals should further realize that the units and algorithm steps of each example described in conjunction with the embodiments disclosed herein can be implemented by electronic hardware, computer software, or a combination of the two. In order to clearly illustrate the interchangeability of hardware and software, the composition and steps of each example have been generally described in the above description according to the function. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Professionals and technicians can use different methods to implement the described functions for each specific application, but such implementation should not be considered to be beyond the scope of the present invention.

[0106] The steps of the method or algorithm described in conjunction with the embodiments disclosed herein may be implemented using hardware, a software module executed by a processor, or a combination of the two. The software module may be placed in a random access memory (RAM), a memory, a read-only memory (ROM), an electrically programmable ROM, an electrically erasable programmable ROM, a register, a hard disk, a removable disk, a CD-ROM, or any other form of storage medium known in the art.

[0107] The specific implementation methods described above further illustrate the objectives, technical solutions and beneficial effects of the present invention in detail. It should be understood that the above description is only a specific implementation method of the present invention and is not intended to limit the scope of protection of the present invention. Any modifications, equivalent substitutions, improvements, etc. made within the spirit and principles of the present invention should be included in the scope of protection of the present invention.

Claims

1. A method for generating text from an image, characterized in that: The method comprises: Get the images and corresponding description texts in the dataset; Performing tokenization processing on the description text and the image to obtain vectorized text tokens and image tokens; The vectorized text word unit and the image word unit are equally input into a pre-training module for training to obtain an image generation model; the normalization layer in the pre-training module includes a scaling module, and the scaling module is used to scale the input of the normalization layer; The test text is input into the image generation model to generate an image corresponding to the test text.

2. The method for generating text to image according to claim 1, characterized in that: The scaling module is used to scale the input of the normalization layer, including: The maximum value of the original input of the normalization layer is obtained, and the input is scaled by the maximum value, which is expressed as: f(x)=x / max(x), where x represents the original input of the normalization layer, max(x) represents the maximum value of the original input of the normalization layer, and f(x) represents the output of the scaling module.

3. The method for generating text to image according to claim 1 or 2, characterized in that: The attention mechanism in the pre-training module includes an adjustment module; the adjustment module is used to adjust the output of the attention mechanism, including: The output of the attention mechanism is adjusted according to the softmax function and the calculated query vector Q and the queried vector K, as shown below: In the formula, represents the output of the attention mechanism, d represents the dimension of vectors Q and K, and α represents a value greater than 32. 5 The constant.

4. The method of text-to-image generation according to claim 1, characterized in that: The lemma processing of the description text and the image includes: Input the description text into a text encoder to obtain the text word unit; The image is input into a discrete image encoder to obtain the image word unit.

5. The method of text-to-image generation according to claim 1, characterized in that: The method further includes a post-processing process before inputting the test text into the image generation model and generating an image corresponding to the test text; the post-processing process includes adjusting super-resolution and re-ranking.

6. The method of text-to-image generation according to claim 5, characterized in that: The adjusting super-resolution comprises: The low-resolution image is enhanced to a high-resolution image according to the super-resolution adjustment strategy.

7. The method of text-to-image generation according to claim 5, characterized in that: The reordering includes: The correlation between the image and the text is evaluated based on the description text score, and the image that best matches the text is selected.

8. A text-to-image generation device, characterized in that: The device comprises: The text image acquisition module is used to obtain the images and corresponding description texts in the data set; A text image processing module, used for performing word-meta processing on the description text and image to obtain vectorized text word-meta and image word-meta; A model training module, which equally inputs the vectorized text word unit and the image word unit into a pre-training module for training to obtain an image generation model; the normalization layer in the pre-training module includes a scaling module, and the scaling module is used to scale the input of the normalization layer; The text-to-image generation module inputs the test text into the image generation model to generate an image corresponding to the test text.

9. An electronic device, characterized in that: The invention comprises a memory and a processor, wherein the memory stores a computer program, and when the computer program is executed by the processor, the text-to-image generation method according to any one of claims 1 to 7 is implemented.

10. A storage medium, characterized in that: A computer program is stored thereon, and when the program is executed by a processor, the text-to-image generation method described in any one of claims 1-7 is implemented.