Image generation model and training and image generation method thereof
By introducing a feature extraction module into the image generation model, using the transformer module to encode spatial information, and optimizing image feature processing through the compression and decompression modules, the problem that image generation models in the prior art are difficult to achieve high quality and low resource consumption at the same time, and efficient image generation is achieved.
Patent Information
- Application Number
- CN202411909126.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-23
- Publication Date
- 2025-05-30
AI Technical Summary
Existing image generation models are difficult to achieve high-quality image generation and low resource consumption at the same time.
An image generation model including a text encoder, an image encoder, a mask module, a feature extraction module and an image decoder are adopted. The feature extraction module encodes spatial information through the transformer module, and optimizes image feature processing through the compression and decompression modules to reduce resource consumption.
On the basis of ensuring the quality of image generation, it effectively reduces resource consumption, improves the performance of image generation model, and enables it to run on ordinary consumer hardware.
Smart Images

Figure CN120071040A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the technical field of image generation, and in particular, to an image generation model, its training, and an image generation method. Background Art
[0002] With the development of artificial intelligence technology, using an image generation model to generate images from text has become a popular research field. Currently, image generation models mainly include diffusion models, autoregressive-based image generation models, mask-based image generation models, etc. Among them, the images generated by diffusion models have relatively high quality, but the resource consumption during operation is relatively large; the resource consumption during operation of autoregressive-based image generation models is relatively large, and the image generation quality is relatively lower than that of diffusion models; the resource consumption during operation of mask-based image generation models is relatively smaller than the previous two models, but the quality of the images generated by them is also relatively lower than that of diffusion models. In summary, current image generation models are difficult to achieve both relatively high image generation quality and relatively low resource consumption. Summary of the Invention
[0003] In a first aspect, an embodiment of this application provides an image generation model, which includes: a text encoder for encoding the text in the input image-text pair into text features; an image encoder for encoding the image in the image-text pair into image features; a mask module for randomly masking the image features output by the image encoder; a feature extraction module including a compression module, a decompression module, and a transformer module, where the compression module is used to perform feature compression on the image features output by the mask module, the transformer module is used to perform spatial information encoding on the text features and the compressed image features, and fuse and output the encoded text features and image features, and the decompression module is used to perform feature decompression on the features output by the transformer module; and an image decoder for generating an output image based on the features output by the decompression module.
[0004] In the embodiment of this application, the feature extraction module is applied to a mask-based image generation model. The feature extraction module includes a transformer module, a compression module, and a decompression module. The transformer module can perform spatial information encoding, enabling the image generation model to effectively maintain the details in high-resolution images and improving the quality of the generated images; the compression module and the decompression module can compress and decompress images, and can effectively reduce resource consumption when processing high-resolution images. In summary, the image generation model of this application can effectively reduce resource consumption on the basis of ensuring image generation quality.
[0005] Second aspect, an embodiment of the present application provides a training method for an image generation model, which is used to train the image generation model described in the first aspect. The method includes: masking the compression module and the decompression module, and training the image generation model based on a first sample set; canceling the masking of the compression module and the decompression module, and training the trained image generation model based on a second sample set; wherein, both the first sample set and the second sample set include image-text pairs, and the resolution of the images in the second sample set is higher than the resolution of the images in the first sample set.
[0006] In the embodiment of the present application, the image generation model is trained in a two-stage training manner. In the first stage, a sample set with a lower resolution is used to train the image generation model. Since less resources are occupied when processing low-resolution images, the compression module and the decompression module can be masked during this stage of training to improve the training efficiency. In the second stage, a sample set with a higher resolution is used to train the image generation model to improve the quality of the images generated by the model. Since more resources are occupied when processing high-resolution images, the masking of the compression module and the decompression module is canceled during this stage of training to reduce the resource occupation during the training process. In summary, the training method of the present application can train an image generation model with high performance with relatively low resource consumption.
[0007] Third aspect, an embodiment of the present application provides an image generation method based on the image generation model described in the first aspect. The method includes: encoding the text in the input image-text pair into text features through a text encoder; encoding the image in the image-text pair into image features through an image encoder; randomly masking the image features output by the image encoder through a masking module; compressing the image features output by the masking module through a compression module; performing spatial information encoding on the text features and the compressed image features through a transformer module, and fusing the encoded text features and image features; decompressing the fused features through a decompression module; and generating an output image based on the features output by the transformer module through an image decoder.
[0008] In the embodiment of the present application, a feature extraction module is applied to the mask-based image generation model. The feature extraction module includes a transformer module, a compression module, and a decompression module. The transformer module can perform spatial information encoding, enabling the image generation model to effectively retain the details in high-resolution images and improving the quality of the generated images. The compression module and the decompression module can compress and decompress image features, and can effectively reduce resource consumption when processing high-resolution images. In summary, the image generation model of the present application can effectively reduce resource consumption while ensuring the quality of image generation.
[0009] In a fourth aspect, an embodiment of the present application provides a computer-readable storage medium, on which the image generation model described in the first aspect, or a computer program is stored. When the program is executed by a processor, the method described in the second aspect or the third aspect is implemented.
[0010] In a fifth aspect, an embodiment of the present application provides a computer device, including a memory, a processor, and a memory; a computer program that can run on the processor is stored on the memory. When the processor executes the program, the method described in the second aspect or the third aspect is implemented; or, the image generation model described in the first aspect, and a computer program that can run on the processor are stored on the memory. When the processor executes the computer program, the method described in the second aspect or the third aspect is implemented.
[0011] In a sixth aspect, an embodiment of the present application provides a computer program product, including the image generation model described in the first aspect, or including a computer program. When the computer program is executed by a processor, the method described in the second aspect or the third aspect is implemented.
[0012] It should be understood that the above general description and the following detailed description are only exemplary and explanatory, and do not limit the present application. Description of the Drawings
[0013] The drawings here are incorporated into the specification and constitute a part of the present application. These drawings show embodiments that conform to the present application and are used together with the specification to explain the technical solutions of the present application.
[0014] Figure 1 It is a schematic structural diagram of the image generation model according to an embodiment of the present application.
[0015] Figure 2 It is a schematic structural diagram of the transformer module according to an embodiment of the present application.
[0016] Figure 3 It is a schematic structural diagram of the transformer module according to another embodiment of the present application.
[0017] Figure 4A and Figure 4B It is a schematic structural diagram of the transformer sub-module according to an embodiment of the present application.
[0018] Figure 5 It is a flowchart of the training method of the image generation model according to an embodiment of the present application.
[0019] Figure 6 It is a flowchart of the image generation method according to an embodiment of the present application.
[0020] Figure 7It is a schematic diagram of the computer device according to an embodiment of the present application. Detailed implementation manners
[0021] Here, the exemplary embodiments will be described in detail, and the examples are shown in the drawings. When the following description refers to the drawings, unless otherwise indicated, the same numbers in different drawings represent the same or similar elements. The implementation manners described in the following exemplary embodiments do not represent all implementation manners consistent with the present application. On the contrary, they are merely examples of the apparatus and methods consistent with some aspects of the present application as detailed in the appended claims.
[0022] The terms used in the present application are for the purpose of describing specific embodiments only and are not intended to limit the present application. The singular forms "a", "the" and "said" used in the present application and the appended claims are also intended to include the plural forms unless the context clearly indicates otherwise. It should also be understood that the term "and / or" used herein refers to and includes any or all possible combinations of one or more of the associated listed items. In addition, the term "at least one" herein means any one of a plurality or any combination of at least two of a plurality.
[0023] It should be understood that although the terms first, second, third, etc. may be used in the present application to describe various information, such information should not be limited to these terms. These terms are only used to distinguish the same type of information from each other. For example, without departing from the scope of the present application, the first information may also be referred to as the second information, and similarly, the second information may also be referred to as the first information. Depending on the context, the word "if" as used herein may be interpreted as "when" or "while" or "in response to determining".
[0024] In order to enable those skilled in the art of the present technology to better understand the technical solutions in the embodiments of the present application and make the above objects, features, and advantages of the embodiments of the present application more obvious and understandable, the technical solutions in the embodiments of the present application will be further described in detail below with reference to the drawings.
[0025] Currently, image generation models mainly include diffusion models, autoregressive-based image generation models, mask-based image generation models, etc. Diffusion models have high image quality and generation ability, but the models are large, require high computing resources, and their operation paradigm is quite different from autoregressive language models. Some autoregressive-based image generation models attempt to use discrete VQ-VAE tokens for autoregressive image generation. However, due to the large number of image tokens, the time overhead of the image generation process is large, and the image generation quality of autoregressive-based image generation models is relatively low compared to diffusion models. Mask-based image generation models cannot effectively handle high-resolution image generation, and the image generation quality is also low. In summary, current image generation models are difficult to achieve both high image generation quality and low resource consumption.
[0026] Based on this, the present application provides an image generation model 10, see Figure 1 , the image generation model 10 includes:
[0027] A text encoder 102 for encoding the text in the input image-text pair into text features;
[0028] An image encoder 104 for encoding the image in the image-text pair into image features;
[0029] A mask module 106 for randomly masking the image features output by the image encoder 104;
[0030] A feature extraction module 108, including a compression module 1082, a decompression module 1084, and a transformer module 1086. The compression module 1082 is used to compress the features of the image features output by the mask module 106. The transformer module 1086 is used to perform spatial information encoding on the text features and the compressed image features, and fuse the encoded text features and image features and then output. The decompression module 1084 is used to decompress the features output by the transformer module 1086; and
[0031] An image decoder 110 for generating an output image based on the features output by the decompression module 1084.
[0032] In the embodiments of the present application, the transformer module 1086 can perform spatial information encoding, enabling the image generation model 10 to effectively preserve details in high-resolution images and improving the quality of the generated images; the compression module 1082 and the decompression module 1084 can compress and decompress image features, effectively reducing resource consumption when processing high-resolution images. In summary, the image generation model 10 of the present application can effectively reduce resource consumption while ensuring the quality of image generation. The following will illustrate the specific implementation details of the present application by examples.
[0033] The image generation model 10 of the present application can generate images based on image-text pairs and is an end-side text-to-image model. The image-text pairs include the original image and the text information used to describe the original image. The image generation model 10 can use the original image as a basis and modify and edit the original image based on the text information in the image-text pairs to obtain the output image. For example, if the original image includes a cloudless sky and the text information is "Add white clouds to the image", the image generation model 10 can output an image of a sky with white clouds. In particular, the image in the image-text pairs can be empty. In this case, the image generation model 10 generates an image only based on the input text information. For ease of distinction, the task of generating images based on both images and text information (i.e., the image in the image-text pairs is not empty) will be referred to as an image editing task hereinafter, and the task of generating images only based on text information (i.e., the image in the image-text pairs is empty) will be referred to as an image generation task.
[0034] The image-text pairs can be input into the image generation model 10. After the image generation model 10 obtains the image-text pairs, it can first perform text encoding on the text therein through the text encoder 102 to obtain text features, and perform image encoding on the image therein through the image encoder 104 to obtain image features.
[0035] In some embodiments, the text encoder 102 can be a large language model (such as T5 or LLaMa). In some other embodiments, the text encoder 102 can also be a single text encoder based on the CLIP (Contrastive Language-Image Pretraining) model. Compared with the large language model, the single text encoder based on the CLIP model can significantly reduce the GPU memory requirements and computational costs without affecting the visual quality.
[0036] The image encoder 104 can be a vector quantization variational autoencoder (VQ-VAE), which is capable of converting image pixels into discrete semantic tokens, which are the image features. The masking module 106 can randomly mask the image features output by the image encoder 104. In some embodiments, a preset sampling rate parameter can be obtained, and based on this sampling rate parameter, the masking ratio r (r is a random number between 0 and 1) of the masking module 106 for the image features is determined, and the image features output by the image encoder 104 are randomly masked based on a specific masking strategy and the masking ratio r. For example, the masking strategy can be a cosine scheduling strategy. The masking ratio r can be input as an input parameter into the inverse cosine (arccos) distribution, the output value of this inverse cosine distribution is obtained, and this output value is used as the probability distribution for masking the image features, and the image features are masked according to this probability distribution. This random masking method enables the image generation model 10 to learn the conditional distribution of any subset of tokens, providing flexibility for the parallel sampling strategy.
[0037] In some embodiments, when the image in the image-text pair is empty, the image features output by the image encoder 104 are also empty. At this time, the masking ratio is 1. In other embodiments, when the image in the image-text pair is non-empty, the image features output by the image encoder 104 are also non-empty. At this time, the masking ratio is less than 1. That is to say, when the data input to the image generation model 10 only includes text information and does not include an image (i.e., the image in the image-text pair is empty), all the image features will be masked, and the features that are all masked are propagated forward. When the data input to the image generation model 10 includes both text information and an image (i.e., the image in the image-text pair is non-empty), only some of the image features are masked. The determination method of the masking ratio can refer to the foregoing embodiments and will not be elaborated here. In this way, the present application can achieve zero shot, that is, the same network architecture can implement two functions of image generation and image editing. When performing image editing, as long as the image area to be edited is masked and then restored through image decoding.
[0038] The image features and text features can be fused by the feature extraction module 108. The feature extraction module 108 may include a compression module 1082, a decompression module 1084, and a transformer module 1086. Among them, the text features can be directly input into the transformer module 1086, and the image features can be first input into the compression module 1082 for compression, then input into the transformer module 1086, and then decompressed by the decompression module 1084. Assume that the size of the image features input into the compression module 1082 is H*W, and the downsampling rate of the compression module 1082 is f. Then the size of the image features output by the compression module 1082 is (H / f)*(W / f). For example, the downsampling rate f can be 16. Assume that the size of the image features input into the compression module 1082 is 1024×1024. Then the size of the image features output by the compression module 1082 is 64×64.
[0039] In some embodiments, referring to Figure 2 , the transformer module 1086 includes a plurality of cascaded multimodal transformer modules 1086a. The first multimodal transformer module 1086a among the plurality of multimodal transformer modules 1086a is respectively connected to the output ends of the text encoder 102 and the compression module 1082. The last multimodal transformer module 1086a among the plurality of multimodal transformer modules 1086a is connected to the input end of the decompression module 1084. The compression module 1082 is used to perform feature compression on the image features output by the mask module 106. Each multimodal transformer module 1086a among the plurality of multimodal transformer modules 1086a is used to respectively perform spatial information encoding on the text features and image features output by the previous module connected to this module, fuse the encoded text features and encoded image features to obtain fused features, and output the text features and image features based on the fused features to the next module connected to this module.
[0040] In this embodiment, multiple cascaded multi-modal Transformer modules 1086a are used to extract and fuse text features and image features. Among them, the text features input to the first multi-modal Transformer module 1086a are the text features output by the text encoder 102, and the image features input to the first multi-modal Transformer module 1086a are the image features output by the compression module 1082. The text features and image features output by the first multi-modal Transformer module 1086a can be input to the second multi-modal Transformer module 1086a, and the text features and image features output by the second multi-modal Transformer module 1086a can be input to the third multi-modal Transformer module 1086a, and so on. The image features output by the last multi-modal Transformer module 1086a can be input to the decompression module 1084 for decompression. The multi-modal Transformer module 1086a can effectively capture cross-modal interactions, thereby extracting the correlation between text features and image features.
[0041] Further, referring to Figure 3 , the Transformer module 1086 may further include multiple cascaded single-modal Transformer modules 1086b. The first single-modal Transformer module 1086b among the multiple single-modal Transformer modules 1086b is connected to the last multi-modal Transformer module 1086a, and the last single-modal Transformer module 1086b among the multiple single-modal Transformer modules 1086b is connected to the input end of the decompression module 1084. Each single-modal Transformer module 1086b among the multiple single-modal Transformer modules 1086b is used to perform spatial information encoding on the image features output by the previous module connected to this module and then output.
[0042] In this embodiment, a plurality of cascaded single-modal Transformer modules 1086b are used to extract features from image features. Among them, the image features input to the first single-modal Transformer module 1086b are the image features output by the last multi-modal Transformer module 1086a. The image features output by the first single-modal Transformer module 1086b can be input to the second single-modal Transformer module 1086b, and the image features output by the second single-modal Transformer module 1086b can be input to the third single-modal Transformer module 1086b, and so on. The image features output by the last single-modal Transformer module 1086b can be input to the decompression module 1084 for decompression. The single-modal Transformer module 1086b can refine the visual representation, thereby improving the detail richness of the output image so as to obtain a high-quality output image. In addition, since the single-modal Transformer module 1086b only needs to process image features and does not need to process text features, the number of parameters of the single-modal Transformer module 1086b is much less than the number of parameters of the multi-modal Transformer module 1086a. Combining the multi-modal Transformer module 1086a and the single-modal Transformer module 1086b can effectively reduce the scale and resource consumption of the image generation model 10.
[0043] In some embodiments, the quantity ratio of the multi-modal Transformer module 1086a and the single-modal Transformer module 1086b can be pre-configured. For example, the above quantity ratio can be 1:2, that is, the number of single-modal Transformer modules 1086b is twice the number of multi-modal Transformer modules 1086a. The applicant has found that when the quantity ratio of the multi-modal Transformer module 1086a and the single-modal Transformer module 1086b is 1:2, the optimal performance can be obtained. Of course, in other examples, other quantity ratios can also be adopted, and the present application does not limit this.
[0044] In some embodiments, the transformer module 1086 is used to perform spatial information encoding on the text features and the compressed image features based on preset micro-conditions. Micro-Conditions refer to the subtle conditions or environmental factors that affect the system, model, or behavior in certain specific situations or contexts. By adding micro-conditions, the performance of the image generation model 10 can be enhanced. These micro-conditions can be converted into sine embeddings and connected as additional channels to the final pooling hidden state of the text encoder 102 to enhance the stability and performance of the image generation model 10. Among them, the micro-conditions include but are not limited to at least any one of the following:
[0045] (1) Sampling rate parameter, which is used to determine the masking ratio of the masking module for masking the image features. By setting different sampling rate parameters, the masking module 106 can mask different proportions of image features.
[0046] (2) Cropping coordinates, which are used to determine the effective image region of the image in the image-text pair. In some embodiments, the cropping coordinates can be randomly generated. For example, assuming that the size of the effective image region (i.e., the image region processed by the image generation model 10) of the image is 1024*1024, and the size of the image in the image-text pair is 1920*1080, then random cropping coordinates (for example, the coordinates of the upper left corner of the effective image region) can be added to the micro-conditions to crop out a 1024*1024 effective image region from the image with the size of 1920*1080, so that the size of the image meets the requirements of the image generation model 10.
[0047] (3) Score of the image-text pair, which is used to determine the quality of the image-text pair. The score of the image-text pair can be evaluated by a pre-trained scoring model, and this score is used to determine the quality of the image-text pair. The quality of the image-text pair can be evaluated from two dimensions. One dimension is the matching degree between the image and the text information in the image-text pair, and the other is whether the image in the image-text pair conforms to human aesthetic preferences. In some embodiments, the above score is the Human Preference Score. By using the score of the image-text pair as a micro-condition, the generated image can be more matched with the text description information and the generated image can be more in line with human aesthetic preferences.
[0048] (4) Resolution of the image in the image-text pair.
[0049] (5) Text features output by the text encoder. The text features contain the semantic information of the text. Using the text features as micro-conditions enables the generated image to more accurately reflect the content of the text.
[0050] In some embodiments, the transformer module 1086 includes a plurality of cascaded transformer sub-modules 108y. The first transformer sub-module 108y among the plurality of transformer sub-modules 108y is connected to the output end of the compression module 1082, and the last transformer sub-module 108y among the plurality of transformer sub-modules 108y is connected to the input end of the decompression module 110. See Figure 4A and Figure 4B , each transformer sub-module 108y includes:
[0051] A first attention layer 108y-1, configured to perform cross-attention processing on the text features (denoted as c in the figure) of the input sub-module based on the micro-condition (denoted as y in the figure) to obtain a query Q corresponding to the text feature c c , a key K corresponding to the text feature c c and a value V corresponding to the text feature c c , and perform cross-attention processing on the image features (denoted as x in the figure) of the input sub-module based on the micro-condition y to obtain a query Q corresponding to the image feature x x , a key K corresponding to the image feature x x and a value V corresponding to the image feature x x ;
[0052] A spatial information encoding layer 108y-2, configured to perform spatial information encoding on the query Q corresponding to the text feature c c and the key K corresponding to the text feature c c as well as the query Q corresponding to the image feature x x and the key K x respectively;
[0053] A second attention layer 108y-3, configured to obtain a target query Q, a target key K, and a target value V, and perform cross-attention processing on the target query Q, the target key K, and the target value V to obtain an output feature F; the target query Q is obtained by concatenating the query Q corresponding to the text feature c c and the query Q corresponding to the image feature x x , the target key K is obtained by concatenating the key K corresponding to the text feature c c and the key K corresponding to the image feature x x , and the target value V is obtained by concatenating the value V corresponding to the text feature c c and the value V corresponding to the image feature x x ;
[0054] The third attention layer 108y-4 is used to obtain the text feature c split from the output feature F of the second attention module 108y-3 ’ and the image feature x ’ , and perform cross-attention processing on the split text feature c based on the micro-condition y ’ and output it, and perform cross-attention processing on the split image feature x based on the micro-condition y ’ and output it.
[0055] In the above embodiment, each transformer sub-module 108y can be the multi-modal transformer module 1086a or the single-modal transformer module 1086b in the foregoing embodiment.
[0056] When a certain transformer sub-module 108y is the multi-modal transformer module 1086a, the first attention layer 108y-1, the spatial information encoding layer 108y-2, the second attention layer 108y-3, and the third attention layer 108y-4 in this transformer sub-module 108y all include two parts, one part corresponding to the text feature c (as shown in the upper half of Figure 4A ), and the other part corresponding to the image feature x (as shown in the lower half of Figure 4A ). On this basis, both the input text feature c and the input image feature x to this transformer sub-module 108y are non-empty.
[0057] When a certain transformer sub-module 108y is the single-modal transformer module 1086b, the first attention layer 108y-1, the spatial information encoding layer 108y-2, the second attention layer 108y-3, and the third attention layer 108y-4 in this transformer sub-module 108y only include the part corresponding to the image feature x (as shown in Figure 4B ). On this basis, the input text feature c to this transformer sub-module 108y is empty, and the input image feature x to this transformer sub-module 108y is non-empty. At this time, the target query Q is the query Q corresponding to the image feature x x , the target key is the key K corresponding to the image feature x x , the target value is the value V corresponding to the image feature x x , and the output feature F of the second attention layer 108y-3 is the image feature x ’ .
[0058] Further, each transformer sub-module 108y may further include a normalization layer 108y-5 for normalizing the queries and keys corresponding to the text features obtained by the spatial information encoding layer 108y-2 and the queries and keys corresponding to the image features obtained by the spatial information encoding layer 108y-2.
[0059] Further, each transformer sub-module 108y may further include at least one of the following network layers:
[0060] A normalization layer 108y-6 for normalizing the text feature c and the image feature x input to this sub-module;
[0061] A linear processing layer 108y-7 for linearly processing the text feature and the image feature output by the first attention layer 108y-1;
[0062] A normalization layer 108y-8 for normalizing the text feature c ’ and the image feature x ’ split from the output feature F of the second attention module 108y-3; and
[0063] A linear processing layer 108y-9 for linearly processing the text feature and the image feature output by the third attention layer 108y-4.
[0064] Further, before inputting the micro-condition y to the first attention layer 108y-1, the micro-condition y may also be linearly processed.
[0065] The image feature output by the last transformer sub-module 108y may be input to the image decoder 110. The image decoder 110 may reconstruct an image by predicting the masked image feature based on the obtained image feature.
[0066] In some embodiments, the feature output by the feature extraction module 108 may also be returned to the input end of the masking module, and the functions of the masking module 106 and the feature extraction module 108 may be repeatedly executed. After several cycles, the feature output by the feature extraction module 108 is input to the image decoder 110 for decoding. In this way, the feature extraction module 108 can extract features more fully, thereby further improving the quality of the generated image. The specific implementation manner in each cycle may refer to the foregoing embodiments and will not be elaborated herein.
[0067] Referring to Figure 5 , this application also provides a training method for an image generation model 10, and the method includes:
[0068] Step S12: Mask the compression module 1082 and the decompression module 1084, and train the image generation model 10 based on the first sample set;
[0069] Step S14: Unmask the compression module 1082 and the decompression module 1084, and train the trained image generation model 10 based on the second sample set; wherein, both the first sample set and the second sample set include image-text pairs, and the resolution of the images in the second sample set is higher than the resolution of the images in the first sample set.
[0070] During the training process, the specific structure and principle of the image generation model 10 can be referred to the foregoing embodiments, and will not be elaborated here. In the embodiments of the present application, the image generation model 10 is first trained with the first sample set having a lower resolution. Since the low-resolution image processing occupies less resources, the compression module and the decompression module can be masked during this stage of training to improve the training efficiency; then the image generation model 10 is trained with the second sample set having a higher resolution to improve the quality of the images generated by the model. Since the high-resolution image processing occupies more resources, the masking of the compression module 1082 and the decompression module 1084 is cancelled during this stage of training to reduce the resource occupation during the training process.
[0071] In some embodiments, the first sample set includes a first sample subset and a second sample subset. The number of image-text pairs in the first sample subset is greater than the number of image-text pairs in the second sample subset, and the quality of the image-text pairs in the second sample subset is higher than the quality of the image-text pairs in the first sample subset. When training the image generation model based on the first sample set, the image generation model can be first trained based on the first sample subset, and then the image generation model trained based on the first sample subset can be trained based on the second sample subset.
[0072] That is to say, the process of training the image generation model with the first sample set in the present application is further divided into two stages. In the first stage, the image generation model 10 is trained with the first sample subset. The first sample subset can be a training set including a large number of image-text pairs, but the resolution of the images therein can be relatively low (for example, the resolution of the images in the first sample set is 256×256), and the quality of the image-text pairs therein can also be relatively low (for example, the matching degree between the images and the text information in the image-text pairs is relatively low). On the one hand, the acquisition difficulty of the low-quality image-text pairs is relatively low. Therefore, a large amount of sample data can be obtained at a relatively low data acquisition cost during this stage; on the other hand, by learning from a large number of image-text pairs, the image generation model 10 can be initially equipped with the ability to generate images.
[0073] In the second stage, the second subset of samples is used to train the image generation model 10. The number of image-text pairs in the second subset of samples is less than that in the first subset of samples, but the resolution is higher, and the matching degree between the image and the text information in the image-text pair is also higher. For example, the resolution of the images in the second sample set is 512×512. Through the training of this stage, the image generation model 10 can be enabled to have the ability to understand text information, so that the images generated by the image generation model 10 match the input text information better. In some embodiments, the second subset of samples may include multiple groups of image-text pairs, and the number and quality of different groups of image-text pairs may be different. For example, the second subset of samples may include three groups of image-text pairs. The number of the first group of image-text pairs (also referred to as low-quality image-text pairs) is the largest among the three groups of image-text pairs, but the quality of the first group of image-text pairs is the lowest among the three groups of image-text pairs; the number and quality of the second group of image-text pairs (also referred to as medium-quality image-text pairs) are both medium among the three groups of image-text pairs, and the number of the third group of image-text pairs (also referred to as high-quality image-text pairs) is the smallest among the three groups of image-text pairs, but the quality of the third group of image-text pairs is the highest among the three groups of image-text pairs. By using sample sets with various characteristics to train the image generation model 10, the generalization ability and image generation performance of the model can be improved.
[0074] In some embodiments, when training the trained image generation model after adding the image compression module and the image decompression module based on the second sample set, the parameters of the text encoder can be fixed first, and the model parameters of the image encoder, the mask module, the feature extraction module, and the image decoder can be fine-tuned based on the second sample set and the first learning rate. Then, the model parameters of the text encoder, the image encoder, the mask module, the feature extraction module, and the image decoder can be fine-tuned based on the second sample set and the second learning rate. Among them, the first learning rate is greater than the second learning rate.
[0075] In some embodiments, the resolution of the images in the second sample set may be 1024×1024. When training the image generation model using the second sample set, it can also be divided into two stages. In the first stage, a larger learning rate is used for training, which can improve the training efficiency. In the second stage, a smaller learning rate is used for training, which can further improve the performance of the trained image generation model.
[0076] In some embodiments, gradient clipping can also be performed during training, that is, if any of the trained model parameters is greater than the preset parameter upper limit, the model parameter is truncated based on the parameter upper limit; similarly, if any of the trained model parameters is less than the preset parameter lower limit, the model parameter is truncated based on the parameter lower limit.
[0077] In some embodiments, a checkpoint reloading strategy can also be adopted during training. The model parameters obtained during training can be recorded at regular intervals. If the value of the loss function obtained based on the model parameters from the most recent training is an invalid value (this phenomenon is called NaN Loss), the model parameters are rolled back to the model parameters obtained from the previous training.
[0078] See Figure 6 , this application also provides an image generation method based on the image generation model 10, and the method includes:
[0079] Step S22: Encode the text in the input image-text pair into text features through the text encoder 102;
[0080] Step S24: Encode the image in the image-text pair into image features through the image encoder 104;
[0081] Step S26: Randomly mask the image features output by the image encoder through the mask module 106;
[0082] Step S28: Compress the image features output by the mask module through the compression module 1082;
[0083] Step S30: Perform spatial information encoding on the text features and the compressed image features through the transformer module 1086, and fuse the encoded text features and image features;
[0084] Step S32: Decompress the fused features through the decompression module 1084; and
[0085] Step S34: Generate an output image based on the features output by the transformer module through the image decoder 110.
[0086] For the specific implementation details of this embodiment, reference can be made to the embodiments of the aforementioned image generation model 10, which will not be elaborated here.
[0087] The image generation model and its training and image generation method of this application have the following technical effects:
[0088] This application improves the performance of masked image modeling technology in high-resolution image generation, making it comparable to advanced diffusion models and even surpassing them in some aspects;
[0089] The architecture of this application adopts an image generation method based on the masking technique. During the image generation process, a model is used to predict the masked image features. This method is similar to the operation paradigms and optimization objectives of language models such as text encoders (such as large language models, CLIP-based single text encoders), which facilitates unifying with the text encoder. Thus, the principle of the text encoder can be borrowed to make the image generation model more intelligent;
[0090] This application optimizes the model architecture and training strategy, reduces the demand for computing resources, enables the model to run on ordinary consumer-grade hardware (such as a GPU with 8GB VRAM), and improves the accessibility and practicality of the model;
[0091] Adopting a masking strategy with a variable masking ratio and cosine scheduling enables the image generation model 10 to learn the conditional distribution of any subset of tokens, provides flexibility for the parallel sampling strategy, and supports the zero-shot image editing function, that is, image generation and image editing can be implemented using the same architecture without changing the image generation architecture to implement the image editing function;
[0092] This application can quickly generate high-quality and high-resolution images, reduce the time and cost of manual image creation, and improve the production efficiency in industries such as creative design and advertising and marketing;
[0093] For example, in the field of creative design, a designer's creative copy (such as "Future City" or "Ocean-style office space") and a reference image can be input into the image generation model. The image generation model can generate corresponding creative images based on the input creative copy and combined with features such as color matching and element combination in the reference image. The designer can further adjust the generated image through tools, such as modifying colors, proportions, details, etc., and even adding specific design elements (such as logos, icons, etc.). The generated image can be used as a reference for design inspiration or directly as the basis for a design draft to enter a more detailed design process. Through the above method, a tool for quickly generating creative images can be provided for designers to assist in inspiring design inspiration and visualizing concepts.
[0094] Another example is in the advertising and marketing industry. The original advertising and promotional pictures, as well as advertising copy, product descriptions, or themes, can be input into the image generation model. For example, inputting "High-end smartwatch" or "Summer promotion event", the image generation model can quickly generate relevant advertising images or renderings based on the input data and combined with information such as marketing needs, target audiences, and brand tones. The image may include product displays, scene backgrounds, advertising slogans, etc. Marketers can modify the generated image through simple editing tools, such as adjusting the color of the product, the style of the background, the position of the copy, etc. The generated advertising images can be directly used for advertising releases, social media promotions, product displays, etc.
[0095] For another example, in the content creation and entertainment industries, the chapter content of a novel, the plot summary of a comic, or a description of a specific scene (such as "a mysterious figure in the forest") can be input into an image generation model. The image generation model understands the content through natural language processing technology and combines information such as the theme, character features, and scene settings of the literary work to generate corresponding illustrations, comics, or animation frames, thereby further enriching the content form in the work. The creator can modify based on the generated image, adjust the character's expression, actions, clothing, background, etc., and even select different painting styles (such as realistic, cartoon, etc.). The generated illustrations or comic frames can be directly used as illustrations in novels, comics, or animations, enriching the visual performance and enhancing the attractiveness of the content.
[0096] This application provides a simple and efficient image creation tool for users, meeting the users' needs for personalized and diverse images and enhancing the users' experience in content creation, entertainment, etc.
[0097] Due to the high efficiency of the image generation model of this application and its low demand for computing resources, it can be applied in resource-constrained environments such as mobile devices, expanding the application scope of the text-to-image synthesis technology. The image generation model can be deployed on a mobile device. Users can import the images on the local mobile device or the images taken by the mobile device into the image generation model and perform image generation through the image generation model to obtain the output image. In this way, the users' need for image creation anytime and anywhere can be met.
[0098] The embodiment of this application also provides a computer device, which at least includes a memory, a processor, and a memory; a computer program that can run on the processor is stored on the memory, and when the processor executes the program, it implements the method described in any embodiment of this application; or, the image generation model described in any embodiment of this application and a computer program that can run on the processor are stored on the memory, and when the processor executes the computer program, it implements the method described in any embodiment of this application.
[0099] Figure 7 FIG. shows a more specific schematic diagram of the hardware structure of a computer device provided by the embodiment of this application. The device may include: a processor 42, a memory 44, an input / output interface 46, a communication interface 48, and a bus 50. Among them, the processor 42, the memory 44, the input / output interface 46, and the communication interface 48 are communicatively connected to each other inside the device through the bus 50.
[0100] The processor 42 can be implemented in the form of a general-purpose central processing unit (CPU), a microprocessor, an application specific integrated circuit (ASIC), or one or more integrated circuits, etc., and is used to execute relevant programs to implement the technical solutions provided in the embodiments of the present application. The processor 42 may further include a graphics card, and the graphics card may be an Nvidia titan X graphics card or a 1080Ti graphics card, etc.
[0101] The memory 44 can be implemented in the form of a read-only memory (ROM), a random access memory (RAM), a static storage device, a dynamic storage device, etc. The memory 44 can store an operating system and other application programs. When implementing the technical solutions provided in the embodiments of the present application through tools or firmware, the relevant program codes are stored in the memory 44 and called and executed by the processor 42.
[0102] The input / output interface 46 is used to connect to the input / output module to implement information input and output. The input / output module can be configured as a component in the device (not shown in the figure) or externally connected to the device to provide corresponding functions. Among them, the input device may include a keyboard, a mouse, a touch screen, a microphone, various sensors, etc., and the output device may include a display, a speaker, a vibrator, an indicator light, etc.
[0103] The communication interface 48 is used to connect to a communication module (not shown in the figure) to implement communication interaction between this device and other devices. Among them, the communication module can implement communication in a wired manner (such as USB, network cable, etc.) or in a wireless manner (such as mobile network, WIFI, Bluetooth, etc.).
[0104] The bus 50 includes a path for transmitting information between various components of the device (such as the processor 42, the memory 44, the input / output interface 46, and the communication interface 48).
[0105] It should be noted that although the above device only shows the processor 42, the memory 44, the input / output interface 46, the communication interface 48, and the bus 50, in the specific implementation process, the device may further include other components necessary for normal operation. In addition, those skilled in the art can understand that the above device may also only include the components necessary to implement the solution of the embodiments of the present application, and does not necessarily include all the components shown in the figure.
[0106] An embodiment of the present application provides a computer program product, including a computer program which, when executed by a processor, implements the method described in any embodiment of the present application.
[0107] An embodiment of the present application further provides a computer-readable storage medium, on which an image generation model described in any embodiment of the present application is stored, or a computer program is stored, which, when executed by a processor, implements the method described in any of the foregoing embodiments.
[0108] Computer-readable media includes both permanent and non-permanent, removable and non-removable media and can be implemented by any method or technology for storing information. The information can be computer-readable instructions, data structures, program modules, or other data. Examples of computer storage media include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technologies, compact disc read-only memory (CD-ROM), digital versatile discs (DVDs) or other optical storage, magnetic cassettes, magnetic tape magnetic disk storage or other magnetic storage devices, or any other non-transitory medium that can be used to store information that can be accessed by a computing device. As defined herein, computer-readable media does not include transitory computer-readable media such as modulated data signals and carrier waves.
[0109] The present application also provides a computer program product, including an image generation model described in any embodiment of the present application, or including a computer program which, when executed by a processor, implements the method described in any embodiment of the present application.
[0110] Each embodiment in the present application is described in a progressive manner. For the same or similar parts among the embodiments, reference can be made to each other. Each embodiment focuses on the differences from other embodiments. In particular, for the device embodiments, since they are basically similar to the method embodiments, they are described relatively simply. For the relevant parts, reference can be made to the description of the method embodiments. The device embodiments described above are merely illustrative. The modules described as separate components may or may not be physically separated. When implementing the solutions of the embodiments of the present application, the functions of the modules can be implemented in the same or multiple tools and / or hardware. It is also possible to select some or all of the modules according to actual needs to achieve the purpose of the solutions of this embodiment. Those of ordinary skill in the art can understand and implement it without creative efforts.
[0111] The above description is only the specific implementation manners of the embodiments of the present application. It should be noted that for those of ordinary skill in the art, without departing from the principle of the embodiments of the present application, several improvements and refinements can be made, and these improvements and refinements should also be regarded as the protection scope of the embodiments of the present application.
Claims
1. An image generation model, the image generation model comprising: A text encoder, used to encode the text in the input image-text pair into text features; An image encoder, configured to encode an image in the image-text pair into image features; A mask module, used for randomly masking the image features output by the image encoder; A feature extraction module, comprising a compression module, a decompression module and a transformer module, wherein the compression module is used to perform feature compression on the image features output by the mask module, the transformer module is used to perform spatial information encoding on the text features and the compressed image features, and fuse the encoded text features and image features and output them, and the decompression module is used to perform feature decompression on the features output by the transformer module; as well as An image decoder is used to generate an output image based on the features output by the decompression module.
2. The image generation model according to claim 1, wherein the transformer module comprises a plurality of cascaded multimodal transformer modules, wherein the first multimodal transformer module in the plurality of multimodal transformer modules is connected to the output ends of the text encoder and the compression module respectively, and the last multimodal transformer module in the plurality of multimodal transformer modules is connected to the input end of the decompression module; The compression module is used to perform feature compression on the image features output by the mask module; Each of the multiple multimodal transformer modules is used to encode spatial information of text features and image features output by a previous module connected to the present module, fuse the encoded text features and the encoded image features to obtain fused features, and output text features and image features to a next module connected to the present module based on the fused features.
3. The image generation model according to claim 2, wherein the transformer module further comprises a plurality of cascaded unimodal transformer modules, the first unimodal transformer module among the plurality of unimodal transformer modules is connected to the last multimodal transformer module, and the last unimodal transformer module among the plurality of unimodal transformer modules is connected to the input end of the decompression module; Each of the multiple single-modal transformer modules is used to encode the spatial information of the image features output by the previous module connected to the module and then output it.
4. According to the image generation model according to claim 3, the ratio of the number of the multimodal transformer modules to the number of the unimodal transformer modules is 1:
2.
5. According to the image generation model of claim 1, the transformer module is used to encode the spatial information of the text features and the compressed image features based on the preset micro-conditions; wherein, The micro-conditions include at least any of the following: A sampling rate parameter, used to determine the mask ratio of the mask module to mask the image features; cropping coordinates for determining a valid image region of an image in the image-text pair; a score of the image-text pair, used to determine the quality of the image-text pair; the resolution of the image in the image-text pair; and The text encoder outputs text features.
6. The image generation model according to claim 5, wherein the transformer module comprises a plurality of cascaded transformer submodules, wherein the first transformer submodule among the plurality of transformer submodules is connected to the output end of the compression module, and the last transformer submodule among the plurality of transformer submodules is connected to the input end of the decompression module; and each transformer submodule comprises: A first attention layer is used to perform cross-attention processing on the text features of the input book submodule based on the micro-conditions to obtain queries, keys, and values corresponding to the text features, and to perform cross-attention processing on the image features of the input book submodule based on the micro-conditions to obtain queries, keys, and values corresponding to the image features; A spatial information encoding layer, for encoding spatial information of the query and key corresponding to the text feature and the query and key corresponding to the image feature respectively; The second attention layer is used to obtain a target query, a target key, and a target value, and perform cross-attention processing on the target query, the target key, and the target value to obtain output features; The target query is obtained by concatenating the query corresponding to the text feature and the query corresponding to the image feature, the target key is obtained by concatenating the key corresponding to the text feature and the key corresponding to the image feature, and the target value is obtained by concatenating the value corresponding to the text feature and the value corresponding to the image feature; The third attention layer is used to obtain text features and image features separated from the output features of the second attention module, perform cross-attention processing on the separated text features based on the micro-conditions and output them, and perform cross-attention processing on the separated image features based on the micro-conditions and output them.
7. The image generation model according to claim 6, wherein each transformer submodule further comprises: The normalization layer is used to normalize the query and key corresponding to the text feature obtained by the spatial information coding layer and the query and key corresponding to the image feature obtained by the spatial information coding layer.
8. According to the image generation model according to claim 1, the text encoder is a single text encoder based on the CLIP model.
9. The image generation model according to claim 1, wherein the mask module is used to randomly mask the image features output by the image encoder based on a predetermined mask ratio; If the image feature output by the image encoder is empty, the mask ratio is 1; If the image feature output by the image encoder is not empty, the mask ratio is less than 1; in, When the image in the image-text pair is empty, the image feature output by the image encoder is empty; when the image in the image-text pair is not empty, the image feature output by the image encoder is not empty.
10. A method for training an image generation model, used for training the image generation model according to any one of claims 1 to 9, the method comprising: Masking the compression module and the decompression module, and training the image generation model based on the first sample set; Unmasking the compression module and the decompression module, and training the trained image generation model based on the second sample set; The first sample set and the second sample set both include image-text pairs, and the resolution of the images in the second sample set is higher than the resolution of the images in the first sample set.
11. The method according to claim 10, wherein the first sample set comprises a first sample subset and a second sample subset, the number of image-text pairs in the first sample subset is greater than the number of image-text pairs in the second sample subset, and the quality of the image-text pairs in the second sample subset is higher than the quality of the image-text pairs in the first sample subset; The step of training the image generation model based on the first sample set includes: Training the image generation model based on the first sample subset; The image generation model trained based on the first sample subset is trained based on the second sample subset.
12. The method according to claim 10, wherein the training of the trained image generation model after adding the image compression module and the image decompression module based on the second sample set comprises: Fixing the parameters of the text encoder, and fine-tuning the model parameters of the image encoder, the mask module, the feature extraction module, and the image decoder based on the second sample set and the first learning rate; Fine-tune the model parameters of the text encoder, image encoder, mask module, feature extraction module and image decoder based on the second sample set and the second learning rate; The first learning rate is greater than the second learning rate.
13. An image generation method based on the image generation model according to any one of claims 1 to 9, the method comprising: Encode the text in the input image-text pair into text features through a text encoder; encoding the image in the image-text pair into image features by an image encoder; Randomly masking the image features output by the image encoder through a mask module; Performing feature compression on the image features output by the mask module through a compression module; Encoding the spatial information of the text features and the compressed image features through a transformer module, and fusing the encoded text features and image features; Decompress the fused features through the decompression module; as well as The output image is generated by the image decoder based on the features output by the transformer module.
14. A computer-readable storage medium storing thereon the image generation model according to any one of claims 1 to 9, or a computer program, which implements the method according to any one of claims 10 to 13 when executed by a processor.
15. A computer device comprising a memory, a processor and a memory; The memory stores a computer program that can be run on a processor, and when the processor executes the computer program, the method according to any one of claims 10 to 13 is implemented; or The memory stores the image generation model described in any one of claims 1 to 9 and a computer program that can be run on a processor, and the processor implements the method described in any one of claims 10 to 13 when executing the computer program.
16. A computer program product, comprising the image generation model according to any one of claims 1 to 9, or comprising a computer program, which implements the method according to any one of claims 10 to 13 when executed by a processor.