Text-to-image generation via masked generation TRANSFORMER
By utilizing the combination of pre-trained language model and visual word prediction model in discrete word space, the problem of high computing resources and low accuracy of existing text-to-image generation models is solved, and efficient and accurate image generation and editing is achieved.
Patent Information
- Application Number
- CN202380089848.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2022-12-16
- Filing Date
- 2023-12-15
- Publication Date
- 2025-07-29
AI Technical Summary
Existing text-to-image generation models have high demand for computing resources and many sampling iterations, making it difficult to achieve real-time or high-throughput applications, and it is difficult to maintain fine-grained language understanding, resulting in insufficient accuracy and fidelity of image generation.
The text embedding is extracted using a pre-trained large language model, using masking modeling tasks in discrete word space, combining the basic visual word prediction model and super-resolution visual word prediction model, and generating high-fidelity images through parallel decoding.
It significantly improves the computing efficiency of the model and the accuracy of image generation, and can understand complex visual concepts such as objects, spatial relationships and poses, which are suitable for real-time image editing applications.
Smart Images

Figure CN120390939A_ABST
Abstract
Description
[0001] Related Applications
[0002] This application claims the priority and benefit of U.S. Provisional Patent Application No. 63 / 433,311, filed on December 16, 2022. U.S. Provisional Patent Application No. 63 / 433,311 is hereby incorporated by reference in its entirety. Technical Field
[0003] The present disclosure generally relates to machine learning. More specifically, the present disclosure relates to text-to-image generation via masked generative transformers. Background Art
[0004] In the fields of computer vision and natural language processing, the text-to-image generation task refers to generating high-quality images from text descriptions. Despite recent advances in deep learning and artificial intelligence, the task of creating realistic images that accurately represent the semantic content of a given text input remains a complex and computationally intensive problem.
[0005] Certain prior arts for text-to-image generation, such as diffusion or autoregressive models, effectively generate images from text prompts. However, these methods typically require a large amount of computational resources and a large number of sampling iterations, which results in significant inefficiencies and delays in the image generation process. This hinders their practical applicability, especially in real-time or high-throughput applications.
[0006] Furthermore, existing models typically operate in the pixel space, which further complicates the computational process and increases resource requirements. These models also struggle to maintain a fine-grained understanding of language, which often results in lower accuracy and fidelity in image generation. This inability to accurately represent complex visual concepts such as objects, spatial relationships, poses, cardinality, etc. based on text descriptions (a concept that can be referred to as "text and image alignment") is a shortcoming of these conventional text-to-image generation models.
[0007] In addition, when it comes to image editing applications such as inpainting, image extension, and maskless editing, current models typically require additional fine-tuning or inversion, which adds another layer of complexity and computational requirements. This makes them less desirable for real-time or on-demand image editing tasks. Summary of the Invention
[0008] Aspects and advantages of embodiments of the present disclosure will be set forth in part in the following description, or may be learned from the description, or may be learned through practice of the embodiments.
[0009] A system of one or more computers can be configured to perform particular operations or actions by installing software, firmware, hardware, or a combination thereof on the system, which, in operation, cause the system to perform these actions. One or more computer programs can be configured to perform particular operations or actions by means of instructions that, when executed by a data processing device, cause the device to perform the actions.
[0010] One general aspect includes a computing system configured to perform text-to-image generation with improved computational efficiency. The computing system further includes one or more processors. The system further includes one or more non-transitory computer-readable media that collectively store: a machine-learned text-to-image generation model, which can include a base visual token prediction model, a super-resolution visual token prediction model, and a visual token decoder model. The one or more non-transitory computer-readable media also store instructions that, when executed by the computing system, cause the computing system to perform operations. The operations include obtaining a text embedding associated with a text prompt that textually describes image content. The operations further include processing the text embedding with the base visual token prediction model to predict one or more of a first set of latent visual tokens associated with a first resolution. The operations further include processing the first set of latent visual tokens with the super-resolution visual token prediction model to generate one or more of a second set of latent visual tokens associated with a second resolution, the second resolution being greater than the first resolution. The operations further include processing the second set of latent visual tokens with the visual token decoder model to generate a synthetic image, wherein the synthetic image depicts the image content textually described by the text prompt. Other embodiments of this aspect include corresponding computer systems, devices, and computer programs recorded on one or more computer storage devices, each configured to perform the actions of the method.
[0011] One general aspect includes a computer-implemented method for training a text-to-image generation model. The computer-implemented method includes obtaining, by a computing system that may include one or more computing devices, training tuples that may include text prompts and training images depicting the content described by the text prompts. The method further includes processing, by the computing system, a lower-resolution version of the training images using a first tokenizer model to generate a first set of conditional tokens. The method further includes masking, by the computing system, one or more of the first set of conditional tokens to generate a first set of masked tokens. The method further includes processing, by the computing system, the first set of masked tokens using the base vision token prediction model conditioned on the text embedding to predict one or more of a first set of latent vision tokens associated with a first resolution. The method further includes training, by the computing system, the base vision token prediction model using a first loss function that compares one or more of the masked first set of conditional tokens with one or more of the first set of latent vision tokens predicted by the base vision token prediction model. The method further includes processing, by the computing system, a higher-resolution version of the training images using a second tokenizer model to generate a second set of conditional tokens. The method further includes masking, by the computing system, one or more of the second set of conditional tokens to generate a second set of masked tokens. The method further includes processing, by the computing system, the second set of masked tokens and the first set of latent vision tokens using the super-resolution vision token prediction model to generate one or more of a second set of latent vision tokens associated with a second resolution that is greater than the first resolution. The method further includes training, by the computing system, the super-resolution vision token prediction model using a second loss function that compares one or more of the masked second set of conditional tokens with one or more of the second set of latent vision tokens generated by the super-resolution vision token prediction model. Other embodiments of this aspect include corresponding computer systems, devices, and computer programs recorded on one or more computer storage devices, each configured to perform the actions of the method.
[0012] One general aspect includes one or more non-transitory computer-readable media that collectively store a text-to-image generation model. The text-to-image generation model includes a base visual token prediction model and a super-resolution visual token prediction model, and the text-to-image generation model has been trained by performing training operations. The training operations include obtaining, by a computing system that may include one or more computing devices, training tuples that may include text prompts and training images depicting the content described by the text prompts. The training operations also include processing, by the computing system, a lower-resolution version of the training images using a first token analyzer model to generate a first set of conditional tokens. The training operations also include masking, by the computing system, one or more of the first set of conditional tokens to generate a first set of masked tokens. The training operations also include processing, by the computing system, the first set of masked tokens conditioned on the text embeddings using the base visual token prediction model to predict one or more of a first set of latent visual tokens associated with a first resolution. The training operations also include training, by the computing system, the base visual token prediction model using a first loss function that compares one or more of the masked first set of conditional tokens with one or more of the first set of latent visual tokens predicted by the base visual token prediction model. The training operations also include processing, by the computing system, a higher-resolution version of the training images using a second token analyzer model to generate a second set of conditional tokens. The training operations also include masking, by the computing system, one or more of the second set of conditional tokens to generate a second set of masked tokens. The training operations also include processing, by the computing system, the second set of masked tokens and the first set of latent visual tokens using the super-resolution visual token prediction model to generate one or more of a second set of latent visual tokens associated with a second resolution that is greater than the first resolution. The training operations also include training, by the computing system, the super-resolution visual token prediction model using a second loss function that compares one or more of the masked second set of conditional tokens with one or more of the second set of latent visual tokens generated by the super-resolution visual token prediction model. Other embodiments of this aspect include corresponding computer systems, devices, and computer programs recorded on one or more computer storage devices, each configured to perform the actions of the method.
[0013] Other aspects of the present disclosure relate to various systems, devices, non-transitory computer-readable media, user interfaces, and electronic devices.
[0014] These and other features, aspects, and advantages of the various embodiments of the present disclosure will be better understood with reference to the following description and the appended claims. The accompanying drawings, which are incorporated in and constitute a part of this specification, illustrate example embodiments of the present disclosure and, together with the description, serve to explain the relevant principles. Description of the Drawings
[0015] With reference to the accompanying drawings, a detailed discussion of the embodiments for those of ordinary skill in the art is set forth in this specification. In the drawings:
[0016] Figure 1 A graphical diagram depicting an example framework for training an example model in accordance with an example embodiment of the present disclosure is shown.
[0017] Figure 2 A graphical diagram depicting an example framework for performing image generation using an example model in accordance with an example embodiment of the present disclosure is shown.
[0018] Figure 3 A graphical diagram depicting an example architecture of an example super-resolution model in accordance with an example embodiment of the present disclosure is shown.
[0019] Figure 4A A block diagram depicting an example computing system in accordance with an example embodiment of the present disclosure is shown.
[0020] Figure 4B A block diagram depicting an example computing device in accordance with an example embodiment of the present disclosure is shown.
[0021] Figure 4C A block diagram depicting an example computing device in accordance with an example embodiment of the present disclosure is shown.
[0022] Reference numerals repeated across multiple drawings are intended to identify the same features in various implementations. Detailed Description
[0023] Overview
[0024] Example aspects of the present disclosure relate to a text-to-image generation model that provides significant improvements in image generation performance and computational efficiency. Prior art such as diffusion or autoregressive models, while effective, typically require a large amount of computational resources and sampling iterations. The systems and methods described herein address these challenges by implementing masked modeling tasks in a discrete token space, which significantly improves the efficiency of the model.
[0025] Specifically, example implementations of the present disclosure utilize a pre-trained large language model (LLM) to extract text embeddings, and then use the text embeddings to predict randomly masked image tokens. This approach not only improves the efficiency of the model, but also enables fine-grained language understanding, which translates to high-fidelity image generation and understanding of visual concepts such as objects and their spatial relationships, poses, cardinality, etc.
[0026] The proposed image generation model can include a series of models. These can include a pair of "tokenizer" models (e.g., VQGAN, VQVAE, etc.), namely, a basic visual token prediction model and a super-resolution visual token prediction model. These components work together to create a powerful and efficient system for generating high-quality images from text prompts.
[0027] Furthermore, due to the use of quantized image tokens and parallel decoding, the proposed model is significantly faster than comparable models. This improves computational efficiency without sacrificing the quality of the generated images or the semantic understanding of the input text prompts. The proposed model can be applied to perform a series of image editing applications, including inpainting, outpainting, and maskless editing, without fine-tuning or inverting the model.
[0028] More specifically, an example aspect of the present disclosure relates to an efficient text-to-image generation model that can be used for various image editing applications. This model - an example implementation of which can be referred to as 'Muse' - is designed to generate high-fidelity images from text prompts. The model can be trained on masked modeling tasks in a discrete token space, where the model is trained to predict randomly masked image tokens given text embeddings extracted from a pre-trained large language model (e.g., LLM). Using discrete tokens and fewer sampling iterations makes this model significantly more efficient than pixel-space diffusion models such as Imagen and DALL-E 2. Thus, in some cases, the proposed approach can be referred to as "discrete diffusion".
[0029] An example aspect involves using a pre-trained text encoder (e.g., LLM) to extract text embeddings. This approach enables the model to have a fine-grained understanding of language, which translates to high-fidelity image generation. The model can understand visual concepts such as objects, their spatial relationships, poses, cardinality, etc. For example, given a text prompt describing "red apple on a table", the model can generate an image accurately depicting this scene.
[0030] Another example aspect involves the use of parallel decoding, which makes the proposed approach more efficient than autoregressive models. For example, compared to some existing autoregressive models, example implementations of the present disclosure can generate images faster due to their ability to predict and / or decode multiple image tokens in parallel. This significantly reduces the time required to generate an image from a given text prompt. The proposed approach is also more efficient than a simple "vanilla" diffusion approach. Specifically, via parallel diffusion, the example models proposed herein can predict multiple final token values in a single forward pass of the model. Standard text-to-image diffusion can predict multiple pixels in a single forward pass, but those predictions are not final - diffusion typically requires feeding those predictions back as input into a series of forward passes to generate the final pixels.
[0031] Another example aspect of the present disclosure relates to a new framework for learning to perform text-to-image synthesis using masked image modeling approaches. Specifically, example implementations of the present disclosure can apply a cascaded visual token prediction approach that includes both a base visual token prediction model and a super-resolution visual token prediction model. The base model can take a series of partially masked low-resolution tokens and predict the marginal distribution of each masked token conditioned on the unmasked tokens and the text embedding. On the other hand, the super-resolution visual token prediction model can again condition on the text embedding to transform the unmasked low-resolution tokens into high-resolution tokens.
[0032] In terms of image editing capabilities, the proposed model can be used to perform many zero-shot image editing capabilities. This includes inpainting, outpainting, and maskless editing. These capabilities are enabled by a mask-based training approach that allows the model to learn to predict masked tokens given a set of conditional tokens generated by a vector-quantized tokenizer model.
[0033] Accordingly, the present disclosure describes an efficient text-to-image generation model that leverages pre-trained large language models, masked image modeling, variable masking rates, and / or classifier-free guidance to generate high-quality images from text prompts. The efficiency and image editing capabilities of the model make it suitable for a variety of use cases in the fields of computer vision and natural language processing.
[0034] The systems and methods of the present disclosure provide many technical solutions to technical problems in the fields of computer vision and natural language processing. One example technical challenge is generating high-quality images from text descriptions. Due to the complexity and computational intensity of creating realistic images that accurately map the semantic content of a given text input, this technical problem poses significant challenges. Conventional techniques such as diffusion or autoregressive models, while effective, have been hindered by high computational resource requirements and many sampling iterations. This results in a significant inefficiency and latency in the image generation process, thus limiting their practical applicability in real-time or high-throughput applications.
[0035] The techniques proposed herein provide a technical solution to this technical problem. By implementing masked modeling tasks in a discrete token space, the efficiency of the model is significantly improved. This approach greatly reduces the need for large amounts of computational resources and sampling iterations, thus addressing a key limitation of existing methods. This technical solution does not require sacrificing the quality of the generated images or the semantic understanding of the input text prompt, thus demonstrating a balance between computational efficiency and performance effectiveness.
[0036] Furthermore, the proposed techniques not only improve the efficiency of image generation, but also improve its accuracy and fidelity. By leveraging pre-trained large language models to extract text embeddings and predict randomly masked image tokens, the model achieves fine-grained language understanding. This translates into high-fidelity image generation and a comprehensive understanding of visual concepts such as objects and their spatial relationships, poses, cardinality, etc.
[0037] In addition, the proposed techniques also provide a solution to the technical problem of image editing. Current models typically require additional fine-tuning or inversion, which increases computational requirements. However, the proposed model can perform a range of image editing applications, including inpainting, image extension, and maskless editing, without the need for model fine-tuning or inversion.
[0038] The techniques proposed herein have many practical applications as it can be used to generate high-quality images from text prompts, which are useful in a wide range of scenarios. For example, it can be used to generate images to accompany various forms of text content, thus eliminating the need for manual acquisition or creation of images. As another example, the technique can be used in an educational environment, where it can generate images based on text descriptions to aid learning and understanding. It can also be used in entertainment and gaming applications, where it generates images for character design, scenes, and other visual elements based on text descriptions. As yet another example, the techniques proposed herein have great potential in image editing applications. For example, it can be used for inpainting, image extension, and maskless editing, thus making it very useful for photo retouching, restoration, and other image editing tasks.
[0039] Figure 1illustrates an example text-to-image generation model and an example framework for training the model. As Figure 1 depicted, an initial step of the training process involves obtaining training tuples by a computing system, which may include one or more computing devices. This training tuple includes a text prompt (labeled 12 in the figure) and a corresponding training image that visually represents the content described by the text prompt. The text prompt can be in any language or script and can contain various types of descriptions ranging from simple object descriptions to complex scene narratives.
[0040] Subsequently, the computing system processes the text prompt using a text encoder, which is labeled 14 in Figure 1 . This text encoder can be a pre-existing language model, such as Google's T5 model, which is trained on a large amount of text data. The text encoder takes the text prompt and transforms it into a text embedding, which is labeled 16 in Figure 1 . This text embedding can be a numerical representation of the text prompt in the latent space, capturing its semantic and syntactic information.
[0041] Once the text embedding is generated, the computing system processes a lower-resolution version of the training image, which is depicted as 18 in Figure 1 . This processing is completed using a first tokenizer model labeled 20 in the figure. As an example, this tokenizer model can be an image encoder, such as a VQGAN (Vector Quantized Generative Adversarial Network) model, which is trained to convert an image into a sequence of discrete tokens. The tokenizer model generates a first set of conditional tokens, which are used as an intermediate representation of the image. The tokenizer can produce tokens in the quantization space. Having a (quantized) list of possible token values instead of a range of possible values allows the application of cross-entropy loss to the encoder output, as further discussed elsewhere in this document.
[0042] The VQGAN model can include an encoder, a decoder, and a quantization layer that converts the input image into a sequence of tokens from a learned codebook. In some implementations, the encoder and decoder can be constructed entirely with convolutional layers to support encoded images at different resolutions. Two VQGAN models can be trained, one with a downsampling rate of 16 and the other with a downsampling rate of 8, to obtain tokens for the base model and the super-resolution model, respectively. These tokens - being discrete in nature - allow the use of cross-entropy loss at the output to predict masked tokens in subsequent stages.
[0043] Still referring to Figure 1 , in Figure 1In the next step depicted as 22, the computing system randomly masks one or more of the first set of conditional tokens, thereby generating a first set of masked tokens labeled 24. This masking process is to create an incomplete representation of the image, which the model will then attempt to complete in subsequent steps.
[0044] Then, in Figure 1 the basic visual token prediction model labeled 26 in processes the first set of masked tokens. This model is conditioned on the previously generated text embeddings. The basic visual token prediction model is trained to predict the masked tokens, thereby generating a first set of latent visual tokens ( Figure 1 labeled 28 in) associated with a lower-resolution version of the image. This prediction takes into account both the masked visual tokens and the text embeddings, thus essentially bridging the gap between the text description and the visual representation.
[0045] In some implementations, the basic prediction model 26 can be a masked transformer whose inputs are the projected text embeddings and the image tokens. The image tokens can be randomly masked and replaced with special [MASK] tokens. The basic model can employ several transformer layers including self-attention blocks, cross-attention blocks, and MLP blocks to extract features. An MLP can be used at the output layer to convert each masked image embedding into a set of logits corresponding to the VQGAN codebook size.
[0046] The training of the basic visual token prediction model 26 can be completed using the first loss function denoted as 29 in Figure 1 This loss function compares the masked tokens from the first set of conditional tokens with the predicted tokens from the first set of latent visual tokens. This comparison helps to refine the prediction model, thus ensuring that the predicted tokens accurately represent the masked tokens. As an example, cross-entropy loss can be applied with the ground truth token labels as the target. The basic model 26 can be trained to predict all the masked tokens at each step during training. However, for inference, mask prediction can be performed iteratively, thereby significantly improving the quality.
[0047] The computing system can also process the higher-resolution version of the training image, which is depicted as 30 in Figure 1 This processing can be completed using the second token analyzer model labeled 32 in the figure. Very similar to the first token analyzer model, this second token analyzer model also generates a set of conditional tokens, but these tokens are associated with the higher-resolution version of the image.
[0048] Similar to the previous masking step, the computing system then masks one or more of the second set of conditional tokens, thereby generating a second set of masked tokens, which is in Figure 1Labeled as 36 in. These masked tokens, along with the first set of potential visual tokens, are then processed by a super-resolution visual token prediction model labeled as 38 in the figure. This model generates a second set of potential visual tokens ([ Figure 1 labeled as 40 in). This super-resolution model essentially transforms or maps the lower-resolution tokens generated by the basic visual token prediction model to higher-resolution tokens.
[0049] Specifically, in some implementations, the proposed framework can incorporate a VQGAN model, where one model has a 16x16 latent resolution and a 256x256 spatial resolution, and the second model has a 64x64 latent resolution and a 512x512 spatial resolution. Since in some implementations the basic model 26 outputs tokens corresponding to a 16x16 latent map, the super-resolution model 38 can learn to "transform" the lower-resolution latent map into a higher-resolution latent map. This super-resolution model 38 can also utilize text conditioning and cross-attention to be trained in a similar manner to the basic model.
[0050] Can be used in Figure 1 labeled as 42 in to train the super-resolution visual token prediction model 38. This loss function compares the masked tokens from the second set of conditional tokens with the generated tokens from the second set of potential visual tokens. This comparison further refines the super-resolution model, ensuring that the high-resolution tokens accurately represent the masked tokens.
[0051] In some implementations, a variable masking rate based on cosine scheduling can also be used to train the model. For example, for each training example, the masking rate can be sampled from a truncated distribution with a density function of . This has an expected masking rate of 0.64 and is strongly biased towards higher masking rates. Biasing towards higher masking rates makes the prediction problem more difficult. Compared to the autoregressive approach of learning the conditional distribution of some fixed order of tokens , the random masking with variable masking rates allows the proposed model to learn for any subset of tokens .
[0052] In some implementations, classifier-free guidance (CFG) is used to improve the model's generation quality and text-image alignment. For example, during training, the text conditioning can be removed from randomly selected 10% of the samples. During inference, some example implementations calculate the conditional log odds and the unconditional log odds for each masked token. Then, the model can be shifted by a certain amount from the unconditional log odds(Guidance Scale) to form the final logits :
[0053]
[0054] Intuitively, the CFG trades diversity for fidelity. Different from previous approaches, some example implementations linearly increase the guidance scale through a sampling process to reduce the impact on diversity. This allows early tokens to be sampled more freely with little or no guidance, but increases the influence of the conditional prompt on later tokens.
[0055] Some example implementations also utilize this mechanism to replace the unconditional logits with logits conditioned on a "negative prompt" to implement negative prompting. This encourages the resulting image to have features associated with the positive prompt and removes features associated with the negative prompt associated. Thus, Figure 1 the framework in can be used to train a text-to-image generation model employing a two-stage token prediction process, involving a basic visual token prediction model and a super-resolution visual token prediction model. This process ensures the generation of high-quality images from text prompts, while also allowing for efficient and scalable training of the model.
[0056] Figure 2 illustrates the use of a trained text-to-image generation model to generate an output image 244 based on a text prompt 212. Figure 2 illustrates the structure of the model, which includes a basic visual token prediction model 226, a super-resolution visual token prediction model 238, and a visual token decoder model 242. These three components work in sequence to generate the output image 244 from the text prompt 212.
[0057] The process begins by receiving the text prompt 212, which is a text description of the content of the image to be generated. This prompt is then processed by a pre-trained language model 214 to obtain a text embedding 216. The text embedding 216, which is a representation of the text prompt in the latent space, carries rich information about the content to be depicted in the image. This text embedding 216 is used as the input to the basic visual token prediction model 226.
[0058] As Figure 2 shown, the basic visual token prediction model 226 processes the text embedding 216 to predict a set of latent visual tokens 228. These tokens 228, represented at a first lower resolution, capture the semantic content of the image at a coarse scale. The prediction of these tokens 228 is conditioned on the text embedding 216, thus ensuring that the generated tokens are consistent with the content described by the text prompt 212.
[0059] Then, the first set of latent visual tokens 228 is used as the input to the super-resolution visual token prediction model 238. This model is designed to transform the lower-resolution tokens 228 into a second set of latent visual tokens 240 represented at a higher resolution. This process is analogous to enhancing the level of detail in an image representation, transforming the coarse semantic content captured by the first set of tokens into a more refined and detailed representation in the second set of tokens 240.
[0060] In some implementations, the efficiency of the model's operation during inference can be attributed to using parallel decoding (e.g., token prediction) to predict multiple output tokens in a single forward pass. Decoding (e.g., token prediction) can be performed at the base model 226 and / or the super-resolution model 238 based on a cosine schedule that selects a fixed fraction of the most confident masked tokens to be predicted at that step. For example, this allows the model to perform inference on 256 tokens using only 24 decoding steps in the base model and on 4096 tokens using 8 decoding steps in the super-resolution model.
[0061] The final stage of the process involves the visual token decoder model 242. This model 242 takes the second set of latent visual tokens 240 as input and decodes them into the synthetic image 244. The decoder model 244 ensures that the generated image 244 is aligned with the content described by the text prompt 212, effectively decoding the high-dimensional token representation into a standard image format.
[0062] In some implementations, the decoder model 242 can be the decoder part of a token analyzer model (e.g., VQGAN) used during the training of the predictor model 238. In some implementations, to enhance the model's ability to generate fine details, the capacity of the decoder 242 can be increased by adding more residual layers and channels while maintaining the encoder capacity. The newly added decoder layers can then be fine-tuned while keeping the encoder weights, codebook, and the base model and super-resolution model unchanged. This allows for an improvement in visual quality without the need to retrain any other model components.
[0063] Figure 3 An example architecture of the super-resolution model is shown, providing a visual representation of the process by which the model enhances the resolution of the image tokens generated by the base visual token prediction model. The architecture of the model includes a series of self-attention Transformer layers that process the low-resolution tokens. These low-resolution tokens generated by the base visual token prediction model represent the image content at a coarser scale. After being processed by the self-attention Transformer layers, output embeddings are generated. These embeddings represent a higher-level understanding of the image content, capturing more detailed and specific features of the image.
[0064] After generating output embeddings, the model concatenates these embeddings with text embeddings. These text embeddings, extracted from the conditional text prompts, carry semantic information about what is being depicted in the image. The concatenation of the output embeddings with the text embeddings produces a combined representation that bridges the gap between textual descriptions and visual representations.
[0065] like Figure 3 The next stage of the super-resolution model depicted in [1] involves the application of crisscross attention. This process applies attention from the concatenated embeddings to masked high-resolution tokens. Masked high-resolution tokens represent the image content at a higher resolution, but are incomplete, with some tokens being masked. The crisscross attention process conditions the predictions for these masked tokens on the concatenated embeddings, ensuring that the predicted tokens align with both the low-resolution tokens and the text embedding.
[0066] The super-resolution model learns to predict these masked tokens using a loss function. This loss function compares the predicted tokens to the actual tokens, refining the model's predictions over time. Through this process, the super-resolution model learns to transform the low-resolution tokens generated by the basic visual token prediction model into high-resolution tokens that capture the image content at a finer scale.
[0067] It should be noted that, as mentioned above and in Figure 3 The architecture, operation, and performance of the super-resolution model depicted in
[15] are exemplary. Changes and modifications to the model may be made without departing from the scope of this disclosure. For example, the number and type of self-attention Transformer layers, the method for generating and processing text embeddings, the details of the cross-attention process, and the nature of the loss function may vary depending on the specific requirements and constraints of the application.
[0068] Figure 4A A block diagram of an example computing system 100 is depicted, according to an example embodiment of the present disclosure. System 100 includes a user computing device 102, a server computing system 130, and a training computing system 150 communicatively coupled via a network 180.
[0069] The user computing device 102 may be any type of computing device, such as, for example, a personal computing device (e.g., a laptop or desktop computer), a mobile computing device (e.g., a smartphone or tablet computer), a gaming console or controller, a wearable computing device, an embedded computing device, or any other type of computing device.
[0070] The user computing device 102 includes one or more processors 112 and a memory 114. The one or more processors 112 can be any suitable processing device (e.g., a processor core, a microprocessor, an ASIC, an FPGA, a controller, a microcontroller, etc.) and can be a single processor or multiple processors operatively connected. The memory 114 can include one or more non-transitory computer-readable storage media, such as RAM, ROM, EEPROM, EPROM, flash memory devices, disks, etc., and combinations thereof. The memory 114 can store data 116 and instructions 118 that are executed by the processor 112 to cause the user computing device 102 to perform operations.
[0071] In some implementations, the user computing device 102 can store or include one or more machine learning models 120. For example, the machine learning model 120 can be or otherwise include various machine learning models, such as neural networks (e.g., deep neural networks) or other types of machine learning models, including non-linear models and / or linear models. The neural network can include a feedforward neural network, a recurrent neural network (e.g., a long short-term memory recurrent neural network), a convolutional neural network, or other forms of neural networks. Some example machine learning models can utilize an attention mechanism, such as self-attention. For example, some example machine learning models can include a multi-head self-attention model (e.g., a transformer model). Refer to Figures 1 to 2 discussion of example machine learning models 120.
[0072] In some implementations, one or more machine learning models 120 can be received via the network 180 from the server computing system 130, stored in the user computing device memory 114, and then used or otherwise implemented by the one or more processors 112. In some implementations, the user computing device 102 can implement multiple parallel instances of a single machine learning model 120 (e.g., to perform parallel image generation across multiple instances of the input).
[0073] Additionally or alternatively, one or more machine learning models 140 can be included in or otherwise stored and implemented by the server computing system 130, which communicates with the user computing device 102 according to a client-server relationship. For example, the machine learning model 140 can be implemented by the server computing system 140 as part of a web service (e.g., an image generation service). Thus, one or more models 120 can be stored and implemented at the user computing device 102, and / or one or more models 140 can be stored and implemented at the server computing system 130.
[0074] The user computing device 102 may also include one or more user input components 122 that receive user input. For example, the user input component 122 may be a touch-sensitive component (e.g., a touch-sensitive display screen or a touchpad) that is sensitive to a user input object (e.g., a finger or a stylus). The touch-sensitive component may be used to implement a virtual keyboard. Other example user input components include microphones, traditional keyboards, or other devices by which a user may provide user input.
[0075] The server computing system 130 includes one or more processors 132 and a memory 134. The one or more processors 132 may be any suitable processing device (e.g., a processor core, a microprocessor, an ASIC, an FPGA, a controller, a microcontroller, etc.) and may be a single processor or multiple processors operatively connected. The memory 134 may include one or more non-transitory computer-readable storage media, such as RAM, ROM, EEPROM, EPROM, flash memory devices, magnetic disks, etc., and combinations thereof. The memory 134 may store data 136 and instructions 138 that are executed by the processor 132 to cause the server computing system 130 to perform operations.
[0076] In some implementations, the server computing system 130 includes or is otherwise implemented by one or more server computing devices. In cases where the server computing system 130 includes multiple server computing devices, such server computing devices may operate according to a sequential computing architecture, a parallel computing architecture, or some combination thereof.
[0077] As described above, the server computing system 130 may store or otherwise include one or more machine learning models 140. For example, the model 140 may be or may otherwise include various machine learning models. Example machine learning models include neural networks or other multi-layer non-linear models. Example neural networks include feedforward neural networks, deep neural networks, recurrent neural networks, and convolutional neural networks. Some example machine learning models may utilize an attention mechanism, such as self-attention. For example, some example machine learning models may include a multi-head self-attention model (e.g., a transformer model). Refer to Figures 1 to 2 Discuss example model 140.
[0078] The user computing device 102 and / or the server computing system 130 may train the model 120 and / or 140 via interaction with a training computing system 150 communicatively coupled via a network 180. The training computing system 150 may be separate from the server computing system 130 or may be a part of the server computing system 130.
[0079] The training computing system 150 includes one or more processors 152 and a memory 154. The one or more processors 152 can be any suitable processing device (e.g., a processor core, a microprocessor, an ASIC, an FPGA, a controller, a microcontroller, etc.) and can be a single processor or multiple processors operatively connected. The memory 154 can include one or more non-transitory computer-readable storage media, such as RAM, ROM, EEPROM, EPROM, flash memory devices, magnetic disks, etc. and combinations thereof. The memory 154 can store data 156 and instructions 158 that are executed by the processor 152 to cause the training computing system 150 to perform operations. In some implementations, the training computing system 150 includes or is otherwise implemented by one or more server computing devices.
[0080] The training computing system 150 can include a model trainer 160 that trains machine learning models 120 and / or 140 stored at the user computing device 102 and / or the server computing system 130 using various training or learning techniques (such as, for example, error backpropagation). For example, a loss function can be backpropagated through the model to update one or more parameters of the model (e.g., based on the gradient of the loss function). Various loss functions can be used, such as mean squared error, likelihood loss, cross-entropy loss, hinge loss, and / or various other loss functions. Gradient descent techniques can be used to iteratively update the parameters over multiple training iterations.
[0081] In some implementations, performing error backpropagation can include performing truncated backpropagation through time. The model trainer 160 can perform various generalization techniques (e.g., weight decay, dropout, etc.) to improve the generalization ability of the model being trained.
[0082] Specifically, the model trainer 160 can train the machine learning models 120 and / or 140 based on a training data set 162. In some implementations, if the user has provided consent, the training examples can be provided by the user computing device 102. Thus, in such implementations, the model 120 provided to the user computing device 102 can be trained by the training computing system 150 based on user-specific data received from the user computing device 102. In some cases, this process can be referred to as personalizing the model.
[0083] The model trainer 160 includes computer logic for providing desired functionality. The model trainer 160 can be implemented in hardware, firmware, and / or software that controls a general-purpose processor. For example, in some implementations, the model trainer 160 includes program files stored on a storage device, loaded into memory, and executed by one or more processors. In other implementations, the model trainer 160 includes one or more sets of computer-executable instructions stored on a tangible computer-readable storage medium (such as RAM, a hard disk, or optical or magnetic media).
[0084] The network 180 can be any type of communication network, such as a local area network (e.g., an intranet), a wide area network (e.g., the Internet), or some combination thereof, and can include any number of wired or wireless links. Generally, communication over the network 180 can be carried out via any type of wired and / or wireless connection, using various communication protocols (e.g., TCP / IP, HTTP, SMTP, FTP), encoding or formats (e.g., HTML, XML), and / or protection schemes (e.g., VPN, secure HTTP, SSL).
[0085] Figure 4A An example computing system that can be used to implement the present disclosure is illustrated. Other computing systems can also be used. For example, in some implementations, the user computing device 102 can include the model trainer 160 and the training data set 162. In such implementations, the model 120 can be trained and used locally on the user computing device 102. In some of such implementations, the user computing device 102 can implement the model trainer 160 to personalize the model 120 based on user-specific data.
[0086] Figure 4B A block diagram of an example computing device 10 performing in accordance with an example embodiment of the present disclosure is depicted. The computing device 10 can be a user computing device or a server computing device.
[0087] The computing device 10 includes a plurality of applications (e.g., Application 1 to Application N). Each application contains its own machine learning library and machine learning model. For example, each application can include a machine learning model. Example applications include a text messaging application, an email application, a dictation application, a virtual keyboard application, a browser application, and the like.
[0088] As Figure 4BAs shown, each application can communicate with multiple other components of the computing device, such as one or more sensors, a context manager, a device status component, and / or additional components. In some implementations, each application can communicate with each device component using an API (e.g., a common API). In some implementations, the API used by each application is specific to that application.
[0089] Figure 4C FIG. shows a block diagram of an example computing device 50 executing in accordance with an example embodiment of the present disclosure. The computing device 50 can be a user computing device or a server computing device.
[0090] The computing device 50 includes multiple applications (e.g., Application 1 to Application N). Each application communicates with a central intelligence layer. Example applications include a text messaging application, an email application, a dictation application, a virtual keyboard application, a browser application, etc. In some implementations, each application can use an API (e.g., a common API across all applications) to communicate with the central intelligence layer (and the models stored therein).
[0091] The central intelligence layer includes multiple machine learning models. For example, as Figure 4C shown, a corresponding machine learning model can be provided for each application, and the corresponding machine learning model can be managed by the central intelligence layer. In other implementations, two or more applications can share a single machine learning model. For example, in some implementations, the central intelligence layer can provide a single model for all applications. In some implementations, the central intelligence layer is included within or otherwise implemented by the operating system of the computing device 50.
[0092] The central intelligence layer can communicate with a central device data layer. The central device data layer can be a centralized data repository of the computing device 50. As Figure 4C shown, the central device data layer can communicate with multiple other components of the computing device, such as one or more sensors, a context manager, a device status component, and / or additional components. In some implementations, the central device data layer can communicate with each device component using an API (e.g., a private API).
[0093] The technology discussed in this document relates to servers, databases, software applications, and other computer-based systems, as well as the actions taken and the information sent to and from such systems. The inherent flexibility of computer-based systems allows for a variety of possible configurations, combinations, and divisions of tasks and functions among and within components. For example, the processes discussed in this document can be implemented using a single device or component or multiple devices or components working in combination. Databases and applications can be implemented on a single system or distributed across multiple systems. Distributed components can operate sequentially or in parallel.
[0094] Although the subject matter has been described in detail with respect to various specific example embodiments thereof, each example is provided by way of illustration and not limitation of the disclosure. Those skilled in the art will readily conceive of changes, variations, and equivalents to such embodiments upon understanding the foregoing. Accordingly, the disclosure does not exclude including such modifications, variations, and / or additions to the subject matter that would be readily apparent to one of ordinary skill in the art. For example, features shown or described as part of one embodiment can be used with another embodiment to yield yet a further embodiment. Accordingly, the disclosure is intended to cover such changes, variations, and equivalents.
Claims
1. A computing system configured to perform text-to-image generation with improved computational efficiency, the computer system comprising: One or more processors; And One or more non-transitory computer-readable media that collectively store: A machine-learned text-to-image generation model, the machine-learned text-to-image generation model including a base visual token prediction model, a super-resolution visual token prediction model, and a visual token decoder model; And Instructions that, when executed by the computing system, cause the computing system to perform operations, the operations including: Obtaining a text embedding associated with a text prompt that textually describes image content; Processing the text embedding with the base visual token prediction model to predict one or more of a first set of latent visual tokens associated with a first resolution; Processing the first set of latent visual tokens with the super-resolution visual token prediction model to generate one or more of a second set of latent visual tokens associated with a second resolution, the second resolution being greater than the first resolution; and Processing the second set of latent visual tokens with the visual token decoder model to generate a synthetic image, wherein the synthetic image depicts the image content textually described by the text prompt.
2. The computer system of claim 1, wherein obtaining the text embedding associated with the text prompt includes processing the text prompt with a pre-trained language model.
3. The computer system of claim 1, wherein processing the text embedding with the base visual token prediction model to predict the one or more of the first set of latent visual tokens includes processing one or more first masked image tokens conditioned on the text embedding with the base visual token prediction model to predict the one or more of the first set of latent visual tokens.
4. The computer system of claim 1, wherein processing the text embedding with the base visual token prediction model to predict the one or more of the first set of latent visual tokens includes processing the text embedding with the base visual token prediction model to predict all of the latent visual tokens in the first set of latent visual tokens.
5. The computer system of claim 1, wherein processing the text embedding with the base visual token prediction model to predict the one or more of the first set of latent visual tokens includes processing the text embedding with the base visual token prediction model to predict a plurality of latent visual tokens in the first set of latent visual tokens, wherein the base visual token prediction model predicts at least two of the plurality of latent visual tokens in the first set of latent visual tokens in parallel.
6. The computer system according to claim 1, wherein processing the text embedding using the basic visual token prediction model to predict the one or more of the first set of potential visual tokens includes processing one or more masked image tokens and one or more conditional visual tokens using the basic visual token prediction model conditioned on the text embedding to predict the one or more of the first set of potential visual tokens, wherein the first set of potential visual tokens includes the one or more conditional visual tokens and the one or more visual tokens predicted by the basic visual token prediction model.
7. The computer system according to claim 1, wherein the basic visual token prediction model and the super-resolution visual token prediction model include transformer models.
8. The computer system according to claim 1, wherein the basic visual token prediction model and the super-resolution visual token prediction model have been jointly trained on a masked image modeling task.
9. The computer system according to claim 8, wherein the masked image modeling task includes predicting masked tokens given a set of conditional tokens generated by a vector-quantized tokenizer model.
10. A computer-implemented method for training a text-to-image generation model, the text-to-image generation model including a basic visual token prediction model and a super-resolution visual token prediction model, the method comprising: obtaining, by a computing system including one or more computing devices, a training tuple including a text prompt and a training image depicting the content described by the text prompt; processing, by the computing system, a lower-resolution version of the training image using a first token analyzer model to generate a first set of conditional tokens; masking, by the computing system, one or more of the first set of conditional tokens to generate a first set of masked tokens; processing, by the computing system, the first set of masked tokens using the basic visual token prediction model conditioned on the text embedding to predict one or more of a first set of potential visual tokens associated with a first resolution; training, by the computing system, the basic visual token prediction model using a first loss function that compares the one or more of the masked first set of conditional tokens with the one or more of the first set of potential visual tokens predicted by the basic visual token prediction model; processing, by the computing system, a higher-resolution version of the training image using a second token analyzer model to generate a second set of conditional tokens; masking, by the computing system, one or more of the second set of conditional tokens to generate a second set of masked tokens; processing, by the computing system, the second set of masked tokens and the first set of potential visual tokens using the super-resolution visual token prediction model to generate one or more of a second set of potential visual tokens associated with a second resolution greater than the first resolution; and The computing system trains the super-resolution visual token prediction model using a second loss function that compares one or more of the masked second set of conditional tokens to one or more of the second set of latent visual tokens generated by the super-resolution visual token prediction model.
11. The computer-implemented method of claim 10, further comprising processing the text prompt using a pre-trained language model to generate the text embedding.
12. The computer-implemented method of claim 10, wherein the super-resolution visual token prediction model is further conditioned on the text embedding.
13. The computer-implemented method of claim 10, wherein masking one or more of the first set of conditional tokens by the computing system to generate the first set of masked tokens comprises masking two or more of the first set of conditional tokens by the computing system to generate the first set of masked tokens, and wherein the base visual token prediction model is configured to predict all of the masked tokens among the two or more masked tokens in a single step.
14. The computer-implemented method of claim 10, wherein the base visual token prediction model and the super-resolution visual token prediction model comprise transformer models.
15. The computer-implemented method of claim 10, wherein the first token analyzer model and the second token analyzer model comprise vector-quantized generative adversarial networks.
16. The computer-implemented method of claim 10, further comprising fine-tuning a decoder network of the second token analyzer model to decode the second set of latent visual tokens into a synthetic image.
17. The computer-implemented method of claim 10, wherein masking one or more of the first set of conditional tokens by the computing system to generate the first set of masked tokens comprises masking one or more of the first set of conditional tokens by the computing system according to a variable masking rate.
18. The computer-implemented method according to claim 10, wherein the method comprises: First, the computing system trains the base visual token prediction model using the first loss function; subsequently, the computing system trains the super-resolution visual token prediction model using the second loss function.
19. The computer-implemented method of claim 10, wherein the first loss function and the second loss function comprise cross-entropy loss functions.
20. One or more non-transitory computer-readable media that collectively store a text-to-image generation model, the text-to-image generation model comprising a base visual token prediction model and a super-resolution visual token prediction model, the text-to-image generation model having been trained by performing training operations that include: obtaining, by a computing system comprising one or more computing devices, training tuples that include text prompts and training images depicting the content described by the text prompts; The computing system processes the lower-resolution version of the training image using the first token analyzer model to generate a first set of conditional tokens; The computing system masks one or more of the first set of conditional tokens to generate a first set of masked tokens; The computing system processes the first set of masked tokens using the basic visual token prediction model conditioned on the text embedding to predict one or more of a first set of latent visual tokens associated with a first resolution; The computing system trains the basic visual token prediction model using a first loss function that compares one or more of the masked first set of conditional tokens with one or more of the first set of latent visual tokens predicted by the basic visual token prediction model; The computing system processes the higher-resolution version of the training image using the second token analyzer model to generate a second set of conditional tokens; The computing system masks one or more of the second set of conditional tokens to generate a second set of masked tokens; The computing system processes the second set of masked tokens and the first set of latent visual tokens using the super-resolution visual token prediction model to generate one or more of a second set of latent visual tokens associated with a second resolution, the second resolution being greater than the first resolution; and The computing system trains the super-resolution visual token prediction model using a second loss function that compares one or more of the masked second set of conditional tokens with one or more of the second set of latent visual tokens generated by the super-resolution visual token prediction model.