Contrastive image captioning neural networks with vector quantization

The system trains an image processing neural network to generate discrete representations using vector quantization, addressing the challenge of capturing semantic properties of images, and enhancing performance in various downstream tasks.

WO2025128886A1PCT designated stage expired Publication Date: 2025-06-19GOOGLE LLC
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
PCT/US2024/059876
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2023-12-12
Filing Date
2024-12-12
Publication Date
2025-06-19

AI Technical Summary

Technical Problem

Existing image processing neural networks struggle to generate discrete representations that capture semantic properties of images, especially across different spatial regions, which limits their effectiveness in downstream tasks such as image generation, compression, and visual understanding.

Method used

A system that trains an image processing neural network to generate discrete representations by using a vector quantization encoder neural network, which processes encoded representations of images to produce discrete embeddings from a codebook, allowing for better semantic capture and utilization in various tasks.

Benefits of technology

The proposed system effectively generates discrete representations that capture semantic properties of images, enabling improved performance in downstream tasks such as image generation, compression, and visual understanding, by facilitating easier modeling and processing by other neural networks.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure US2024059876_19062025_PF_FP_ABST
    Figure US2024059876_19062025_PF_FP_ABST
Patent Text Reader

Abstract

Methods, systems, and apparatus, including computer programs encoded on computer storage media, for training an image processing neural network to generate discrete representations of input images. The discrete representations can then be used for any of a variety of downstream tasks.
Need to check novelty before this filing date? Find Prior Art

Description

[0001]Attorney Docket No.56113-0532WO1 CONTRASTIVE IMAGE CAPTIONING NEURAL NETWORKS WITH VECTOR QUANTIZATION CROSS-REFERENCE TO RELATED APPLICATIONS This application claims the benefit of priority to U.S. Provisional Application No. 63 / 609,299 filed on December 12, 2023, the contents of which are hereby incorporated by reference. BACKGROUND This specification relates to processing images using machine learning models. As one example, neural networks are machine learning models that employ one or more layers of nonlinear units to predict an output for a received input. Some neural networks include one or more hidden layers in addition to an output layer. The output of each hidden layer is used as input to another layer in the network, e.g., the next hidden layer or the output layer. Each layer of the network generates an output from a received input in accordance with current values of a respective set of weights. SUMMARY This specification describes a system implemented as computer programs on one or more computers that trains an image processing neural network that generates discrete representations of images. Particular embodiments of the subject matter described in this specification can be implemented so as to realize one or more of the following advantages. By training the image processing neural network as described in this specification, the image processing neural network, the discrete representations generated by the image processing neural network, or both, can effectively be used in a variety of downstream tasks. In particular, by training the image processing neural network as described in this specification, the representations generated by the neural network capture semantic properties of the contents of the image and, more specifically, can capture different semantic properties of the contents of different spatial regions of the image. Because the representations are discrete, rather than continuous, the representations can more easily be modeled by other neural networks when performing processing that is required to perform the downstream task. For example, the discrete representations generated by the image processing neural network can be used to train an image generation neural network that generates Attorney Docket No.56113-0532WO1 images conditioned on discrete representations. In particular, the image generation neural network can generate a sequence of tokens from the vocabulary, e.g., conditioned on another sequence of tokens from the vocabulary or on another sequence representing a different type of conditioning input or both, and then an “inverter” or “decoder” neural network can decode the sequence of tokens to generate an image. As another example, the discrete representations can be used for image compression, e.g., so that the discrete representations are used to later reconstruct an input image by an image reconstruction neural network. As yet another example, the discrete representations can be used as a representation of the image in visual understanding tasks, e.g., image-text retrieval tasks, image classification tasks, image captioning tasks, and visual question answering tasks. As a particular example, the discrete representations can be provided as input to a multi-modal language model, i.e., that operates on tokens of multiple different modalities, in order to allow the multi-modal language model to effectively perform visual understanding tasks. The details of one or more embodiments of the subject matter of this specification are set forth in the accompanying drawings and the description below. Other features, aspects, and advantages of the subject matter will become apparent from the description, the drawings, and the claims. BRIEF DESCRIPTION OF THE DRAWINGS FIG.1 shows an example neural network system. FIG.2 shows an example architecture of the neural network system. FIG.3 is a flow diagram of an example process for training the neural network. FIG.4 shows an example of performing an inversion process to generate an output image from discrete tokens. Like reference numbers and designations in the various drawings indicate like elements. DETAILED DESCRIPTION FIG.1 shows an example neural network system 100. The neural network system 100 is an example of a system implemented as computer programs on one or more computers in one or more locations, in which the systems, components, and techniques described below can be implemented. Attorney Docket No.56113-0532WO1 The system 100 trains an image processing neural network 110 that generates discrete representations 120 of images 102. The representations 120 are referred to as “discrete” because the representation of a given image 102 represents the image 102 as a collection of embedding vectors selected from a discrete codebook 135 (also referred as a “set” or a “vocabulary”) of embedding vectors, i.e., a codebook 135 that includes a fixed number of embedding vectors. For example, the discrete representation 120 can identify, for each of multiple spatial regions in the input image 102, a respective embedding vector from the discrete codebook 135. This is in contrast to a “continuous” representation of an image, where the only constraint on the embedding vectors in the representation is that each numeric value in each embedding vector be representable in the numerical format used by the system when processing inputs through the neural network. In particular, the image processing neural network 110 includes a visual encoder neural network 130 that is configured to process an input image 102 to generate an encoded representation 132 of the input image 102. The input image 102 can have a plurality of pixel values, e.g., respective intensity values for each of the pixels of the image. The image processing neural network 110 also includes a vector quantization encoder neural network 140 that is configured to process the encoded representation 132 of the input image 102 to generate a set of encoded vectors 142 representing the input image. The image processing neural network 110 is then configured to generate the discrete representation 120 of the input image 102 by, for each encoded vector, selecting a corresponding embedding vector from the codebook of vectors in the codebook 135. For example, the neural network 110 can select, for each encoded vector, the closest embedding vector from the codebook, e.g., based on Euclidean distance or cosine similarity. The visual encoder neural network 130 can be any appropriate neural network that processes an image to generate an encoded representation that includes a respective vector (“initial embedding vector”) for each of multiple spatial regions within the image. For example, the visual encoder neural network can be a vision Transformer (ViT) neural network or a convolutional neural network, e.g., a ResNet. Attorney Docket No.56113-0532WO1 As will be described in more detail below, in some cases, prior to the training of the image processing neural network 110, the visual encoder neural network 130 may have been pre-trained as part of a different image processing neural network on a different image representation learning task. The system 100 can train the image processing neural network 110 jointly with a language model neural network 150. For example, the language model neural network 150 can include a set of initial uni-modal neural network layers 160 and a set of subsequent neural network layers 170 that include both cross-modal layers and uni-modal layers. That is, when processing a multi-modal input that includes both an image and a text sequence, the representation generated by the initial layers 160 within the language model neural network 150 depends only on the text sequence while the representation generated by the subsequent layers 170 depends on both the representation of the text sequence generated by the initial layers and the image, i.e., on the discrete representation 120. Generally, the language model neural network 150 is configured to process a current text sequence to generate an output defining a new token 128 to be appended to the current text sequence. The output defining a new token to be appended to the current text sequence generally includes a respective score for each token in a vocabulary of tokens. The vocabulary of tokens can include any of: characters, subwords, words, punctuation marks, sign tokens (e.g., the #, $, and other signs), mathematical symbols, and so on. The vocabulary of tokens can also include one or more special tokens that are appended to input text sequences that processed by the neural network, e.g., a start of sequence token, an end of sequence token, a designated “class” token, and so on. During training, the language model neural network 150 can generate a respective output for each of multiple tokens in an input sequence in a single forward pass, i.e., in parallel. The language model neural network 150 can have any appropriate architecture that allows the language model neural network 150 to map the tokens in the text sequence to a respective uni-modal representation 124 for each of the tokens and then map the uni- modal representations 124 to an output defining the next token 128. In a particular example, the language model neural network 150 can have an attention-based architecture, e.g., the architecture of a decoder-only Transformer neural network. Attorney Docket No.56113-0532WO1 In this example, the set of initial uni-modal neural network layers 160 can include a sequence of initial attention layers, where each initial attention layer is configured to receive as input a respective current representation of each of the text tokens in the current text sequence and to process the respective current representations to generate as output a respective updated representation of each of the text tokens in the current text sequence. For example, each initial attention layer can apply a causally masked self- attention mechanism over the respective current representations to generate the respective updated representations. A self-attention mechanism over the respective current representations refers to an attention mechanism that computes queries, keys, and values from the respective current representations. A causally masked self-attention mechanism over the respective current representations refers to an attention mechanism in which any given position in the current text sequence does not attend over, i.e., does not have a non-zero attention weight for, any positions after the given position in the current text sequence. Each attention layer can optionally apply other operations to the representations as part of updating the representations, e.g., by making use of a position-wise feed-forward neural network, by applying layer normalization, by making use of residual connections, and so on. In this example, the respective current representations that are received as input by the first initial attention layer in the sequence of initial attention layers are respective embeddings of each of the text tokens in the current text sequence, e.g., as generated by an embedding layer of the language model neural network 150 and the respective current representations that are received as input by each subsequent initial attention layer, i.e., each initial attention layer after the first initial attention layer in the sequence of initial attention, layers are respective updated representations of the text tokens in the current text sequence that are generated as output by a preceding initial attention layer in the sequence of initial attention layers. Thus, the respective uni-modal representations of the text tokens in the current text sequence are the respective updated representations of the text tokens in the current text sequence that are generated as output by the last initial attention layer in the sequence of initial attention layers. More specifically, when the language model neural network 150 has an attention- based architecture, the initial attention layers include multiple self-attention layers but do Attorney Docket No.56113-0532WO1 not include any cross-attention layers to ensure that the respective updated representations of the text tokens in the current text sequence are uni-modal representations that depend only on the current text sequence and not on the image. Additionally, the set of subsequent neural network layers 170 includes a sequence of subsequent attention layers, with each subsequent attention layer being configured to receive as input a respective current representation of each of the text tokens in the current text sequence and to process the respective current representations to generate as output a respective updated representation of each of the text tokens in the current text sequence. Thus, the respective current representations that are received as input by the first subsequent attention layer in the sequence of subsequent attention layers are the respective uni-modal representations of each of the text tokens in the current text sequence (generated by the initial attention layers) and the respective current representations that are received as input by each subsequent attention layer after the first subsequent attention layer in the sequence of subsequent attention layers are respective updated representations of the text tokens in the current text sequence that are generated as output by the preceding subsequent attention layer in the sequence of subsequent attention layers. In this example, like the initial attention layers, the sequence of subsequent attention layers also includes one or more self-attention layers. That is, for one or more of the subsequent neural network layers, processing the respective current representations to generate as output a respective updated representation of each of the text tokens in the current text sequence includes applying a causally masked self-attention mechanism. Each of these attention layers can optionally apply other operations to the representations as part of updating the representations, e.g., by making use of a position-wise feed- forward neural network, by applying layer normalization, by making use of residual connections, and so on. Unlike the initial layers, the sequence of subsequent attention layers also includes one or more cross-modal layers. Each cross-modal layer processes the respective current representations to generate as output a respective updated representation of each of the text tokens in the current text sequence by applying a cross-attention mechanism between an input derived from (generated from) the discrete representation 120 of the image and the respective current representations of the text tokens in the current text sequence received as input by the cross-modal layer. Attorney Docket No.56113-0532WO1 “Cross-attention” between the input derived from the encoded representation of the image and the respective current representations of the text tokens in the current text sequence received as input by the cross-modal layer refers to an attention mechanism that uses queries derived from the respective current representations of the text tokens in the current text sequence and keys and values derived from the input generated from the encoded representation of the image. Specific examples of self-attention, cross-attention, and causally masked self- attention mechanisms that can be employed by the system are described in Vaswani et al. “Attention is all you need”, 31st Conference on Neural Information Processing Systems (NIPS 2017), Long Beach, CA, USA; Colin Raffel, Noam Shazeer, Adam Roberts, Katherine Lee, Sharan Narang, Michael Matena, Yanqi Zhou, Wei Li, and Peter J Liu. Exploring the limits of transfer learning with a unified text-to-text transformer. arXiv preprint arXiv:1910.10683, 2019; Daniel Adiwardana, Minh-Thang Luong, David R. So, Jamie Hall, Noah Fiedel, Romal Thoppilan, Zi Yang, Apoorv Kulshreshtha, Gaurav Nemade, Yifeng Lu, and Quoc V. Le. Towards a human-like open-domain chatbot. CoRR, abs / 2001.09977, 2020; and Tom B Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, et al. Language models are few-shot learners. arXiv preprint arXiv:2005.14165, 2020, Hua, et al, Transformer Quality in Linear Time, arXiv preprint arXiv:2202.10447, 2022. Generally, the system 100 can jointly train the image processing neural network 110 and the language model neural network 150 on (i) a contrastive loss, (ii) a captioning loss, or (iii) both. The contrastive loss can depend on the discrete representation 120 generated by the image processing neural network 110 and the representations generated by the initial neural network layers 160 while the captioning loss can depend on the discrete representation 120 generated by the image processing neural network 110 and the representations generated by the subsequent neural network layers 170. After the image processing neural network 110 has been pre-trained, the image processing neural network 110 can be used, e.g., by the system 100 or by a different inference system, to perform a downstream task. For example, the discrete representations generated by the image processing neural network can be used to train an image generation neural network that generates images conditioned on discrete representations. In particular, the image generation neural Attorney Docket No.56113-0532WO1 network can generate a sequence of tokens from the vocabulary, e.g., conditioned on another sequence of tokens from the vocabulary or on another sequence representing a different type of conditioning input or both, and then an “inverter” or “decoder” neural network can decode the sequence of tokens to generate an image. An example of such an approach is described below with reference to FIG.4. As another example, the discrete representations can be used for image compression, e.g., so that the discrete representations are used to later reconstruct an input image by an image reconstruction neural network. As the discrete representation includes selections from a codebook, these can be stored as codebook indices, thereby providing a much smaller, compressed representation of the input image. As yet another example, the discrete representations can be used as a representation of the image in visual understanding tasks. One example of a visual understanding task is an image-text retrieval task, where the input includes an image or text or both and the output is an image that is received from an image datastore. Another example of a visual understanding task is an image classification task, where the input is an image and the output is an identification of objects depicted in the image). Another example of a visual understanding task is an image captioning task, where the input is an image and the output describes in natural language the objects depicted in the image. Another example of a visual understanding task is a visual question answering task, where the input is an image and a query about the image and the output is a response to the query. As another example, the downstream task can be a video processing task that requires processing respective discrete representations generated by the image processing neural network of each video frame in an input video. For example, the task can be a video question answering task, a video classification task, an action recognition task, and so on. For example, to perform the above tasks, the discrete representations of a query image can be provided as input to a language model neural network or multi-modal language model neural network along with a prompt that specifies the task to be performed. The neural network can then process the input to generate the output for the task. Attorney Docket No.56113-0532WO1 FIG.2 shows an example 200 of the architecture of the neural network system 100. As shown in the example 200, the system 100 includes the image processing neural network 110 and the language model neural network 150. The language model neural network 150 includes the initial, unimodal layers 160 and the subsequent, multi-modal layers 170. For example, as described above, the initial layers 160 can include self-attention layers while the subsequent layers 170 can include self-attention layers and cross-attention layers. In the example 200, the cross-attention layers cross-attend into a contrastive representation 250 of a training image that is generated by applying attention pooling 220 to the final representation of the training image. Applying attentional pooling is described in more detail below. The image processing neural network 110 includes the visual encoder neural network 130 that is configured to process an input image to generate an encoded representation of the input image, the vector quantization encoder neural network 140 that is configured to process the encoded representation of the input image to generate a set of encoded vectors representing the input image, and the codebook 135 of embedding vectors. In the example 200, the visual encoder neural network 130 generates an encoded representation that includes a respective initial encoded vector corresponding to each of a plurality of regions of the image. In particular, the visual encoder neural network generates a respective initial encoded vector for each patch in a 16 x 16 grid of patches from the input image. Similarly, the set of encoded vectors includes a respective encoded vector for each of the plurality of regions of the image. For example, the vector quantization encoder neural network 140 can be a Transformer neural network that is configured to update each of the initial encoded vectors to generate the set of encoded vectors. That is, the vector quantization encoder neural network 140 maintains the correspondence between encoded vectors and regions of the input image. In the example 200, the system 100 also includes a vector quantization decoder neural network 210. The vector quantization decoder neural network 210 processes the discrete representation to generate the final representation. For example, the vector quantization decoder neural network 210 can process the embedding vectors that are identified by the discrete representation to update each of the embedding vectors, so that the final Attorney Docket No.56113-0532WO1 representation includes the updated embedding vectors generated by the vector quantization decoder neural network 210. For example, the vector quantization decoder neural network 210 can be a Transformer neural network. The system can then train the image processing neural network 110 using the final representation, as will be described in more detail below. For example, the system can train the image processing neural network 110 using a contrastive loss. In the example of FIG.2, the contrastive loss is computed using a contrastive representation 250 of the image that is computed from the final representation by applying attentional pooling 220 and a uni-modal representation 240 of a “CLS” token in the text sequence generated by the initial, uni-modal layers 160. As another example, the system can train the image processing neural network 110 using a captioning loss that is computed based on the probability distributions generated by the subsequent, multi-modal layers 170. After training, the system can use the image processing neural network 110 to perform downstream tasks, as described above. In these cases, the system can discard the other components of the system that may not be necessary for the downstream tasks, e.g., the vector quantization decoder neural network 210, the attentional pooling 220, or the language model neural network 150. FIG.3 is a flow diagram of an example process 300 for training the image processing neural network. For convenience, the process 300 will be described as being performed by a system of one or more computers located in one or more locations. For example, a neural network system, e.g., the neural network system 100 of FIG.1, appropriately programmed, can perform the process 300. The system can repeatedly perform the process 300 on different batches of image – text sequence pairs in order to train the image processing neural network. The system obtains a plurality of image – text sequence pairs (step 302). For example, the system can sample the pairs from a larger training data set of training pairs. That is, each pair in the training data set includes a respective image and a respective input text sequence. In particular, for each pair, the input text sequence in the pair has been determined by the system or an external source to describe the contents of the image or otherwise be relevant to the image in the pair. In other words, the image and the input text sequence have been determined to be semantically similar, where, for example, an image and a text Attorney Docket No.56113-0532WO1 sequence can be similar if the text sequences describes or otherwise relates to the objects depicted in the image. For example, within a given training pair, the text sequence can be a text annotation of the image from a set of manually or automatically generated image annotations or can be alt text associated with the image in a set of alt-text data. Alt text is text that is displayed in place of an image on a web page, e.g., if the image cannot be rendered properly or otherwise fails to load. For example, the system can obtain the alt- text data from data maintained by an Internet search engine or other software that automatically crawls web pages on the Internet. The system then performs steps 304-312 for each of the training pairs. The system processes the training image in the training pair using the visual encoder neural network to generate an encoded representation of the training image (step 304). As described above, the visual encoder neural network can generally have any appropriate architecture that maps an image to an encoded representation that includes a set of feature vectors. For example, the visual encoder can be a convolutional neural network, e.g., a ResNet, or a Transformer neural network, e.g., a Vision Transformer. The system processes the encoded representation of the training image using the vector quantization encoder neural network to generate a set of encoded vectors representing the training image (step 306). As described above, the vector quantization encoder neural network generally updates each of the vectors in the encoded representation to generate the set of encoded vectors. For example, the vector quantization encoder neural network can be a Transformer neural network or a multi- layer perceptron (MLP) or a linear layer that processes each vector in the encoded representation independently. The system generates a discrete representation of the training image by, for each encoded vector, selecting a corresponding embedding vector from the codebook of embedding vectors (step 308). For example, the system can select, for each encoded vector, the closest embedding vector from the codebook. The system generates a final representation of the training image from the discrete representation (step 310). For example, the final representation can be the embedding vectors identified by the discrete representation. Attorney Docket No.56113-0532WO1 As another example, the system can process the discrete representation using the vector quantization decoder neural network as described above to generate the final representation. The system processes the training text sequence in the training pair using a language model neural network to generate a text representation of the text sequence (step 312). For example, as described above, the language model neural network can have any appropriate architecture that allows the language model neural network to map the tokens in the text sequence to a respective uni-modal (text) representation for each of the tokens and then map the uni-modal representations to an output defining the next token. That is, the uni-modal representation can be the text representation of the text sequence. For example, the language model neural network can include a set of initial uni- modal neural network layers and a set of subsequent neural network layers that include both cross-modal layers and uni-modal layers. That is, when processing a multi-modal input that includes both an image and a text sequence, the representation generated by the initial layers within the language model neural network depends only on the text sequence while the representation generated by the subsequent layers depends on both the representation of the text sequence generated by the initial layers and the image, i.e., on the discrete representation. For example, as described above, the subsequent layers can include one or more layers that each receive an input derived from the discrete representation. For example, this input can be the final representation of the image described above. As another example, this input can be a captioning representation of the image that is generated by processing the final representation of the image. For example, the system can use learned attentional pooling to generate the contrastive representation. To perform learned attentional pooling, the system can incorporate an attentional pooling layer. The attentional pooling layer applies attention over the embeddings in the final representation and a set of learned query tokens to generate a respective updated query token for each of the learned query tokens in the second set. That is, the system uses the learned query tokens to generate queries for the attention mechanism applied by the attentional pooling layer and the embeddings in the final representation to generate keys Attorney Docket No.56113-0532WO1 and values for the attention mechanism. The output of the attentional pooling layer is therefore an updated query token for each of the set of query tokens. The set of learned query tokens and the parameters of the attention mechanism are learned jointly with the parameters of the language model neural network and the image processing neural network during the pre-training. To generate the captioning representation, the system can use a set of learned query tokens that has multiple tokens, and can use the updated query tokens as the captioning representation. The system trains the image processing neural network using, for each training pair, the final representation of the training image in the training pair and the text representation of the text sequence in the training pair (step 314). When the system includes the vector quantization decoder neural network, the system also trains the vector quantization decoder neural network using, for each training pair, the final representation of the training image in the training pair and the text representation of the text sequence in the training pair. In particular, the system can train the image processing neural network to minimize a loss function that is computed based on, for each training pair, the final representation of the training image in the training pair and the text representation of the text sequence in the training pair. The system also updates the codebook of embedding vectors. For example, the system can update the codebook of embedding vectors using a VQ-VAE loss, e.g., using a VQ-VAE moving averages update or using a VQ-VAE gradient-based update. The system can also train the language model neural network on the same loss function. By repeatedly performing the process 300 on different batches of training pairs, the system trains the image generation neural network to generate discrete representations that effectively represent properties of training images, e.g., that accurately represent semantic properties of the training images despite the “loss” of information that occurs due to discretizing the representations. For example, the loss function can include a contrastive loss term. The contrastive loss term measures a contrastive loss that is computed using, for each training pair, a contrastive image final representation of the training image in the training pair and a contrastive text final representation of the training text in the training pair. Attorney Docket No.56113-0532WO1 The contrastive image final representation of the training image is generated from the final representation of the training image. The system can generate the contrastive image final representation of the training image from the final representation of the training image in any of a variety of ways. As one example, the system can use pooling to generate the contrastive representation. As one example, the system can apply a pooling operation, e.g., global average pooling (GAP), on the embeddings in the final representation and use the resulting pooled embedding as the contrastive representation. As another example, the system can use learned attentional pooling to generate the contrastive representation. To perform learned attentional pooling, the system can incorporate an attentional pooling layer. The attentional pooling layer applies attention over the embeddings in the final representation, and a set of learned query tokens to generate a respective updated query token for each of the learned query tokens in the set. That is, the system uses the learned query tokens to generate queries for the attention mechanism applied by the attentional pooling layer and the embeddings in the final representation to generate keys and values for the attention mechanism. The output of the attentional pooling layer is therefore an updated query token for each of the set of query tokens. The set of learned query tokens and the parameters of the attention mechanism are learned jointly with the parameters of the language model neural network and the image processing neural network during the pre-training. To generate the contrastive representation, the system can use a set of learned query tokens that has only a single query token, and can use the updated query token for the single query token as the contrastive representation. The system can generate the contrastive text representation in any of a variety of ways. As a particular example, each text sequence in the batch can include the same designated token located at the same position within each text sequence. For example, the system or another system can augment each text sequence with a designated token, e.g., a “CLS” token placed at the end of every text sequence. The system can then use the uni-modal representation of the designated token to compute the contrastive loss. The goal of the contrastive loss is to train the image processing neural network and the language model so that they can embed image and text inputs into the Attorney Docket No.56113-0532WO1 representation space, i.e., the space of the image and text embeddings, in such a way that inputs with similar semantics are mapped to nearby points regardless of their modalities. Thus, the system can train the image processing neural network and the language modeling neural network on a contrastive loss that encourages, for all training pairs in the batch that include a training image xi and a text sequence yi, the text embedding of xi and the contrastive representation of yito be closer together while being farther from all other embeddings of all other images and text segments in the batch. A particular example of a contrastive loss will be described next. Based on the embeddings for the images and the text segments in the pairs in the mini-batch, an N x N similarity matrix A is computed, where Ai;jis a value that represents how similar the embedding of xi is to the embedding of yj. For example, Ai;j can be the dot product between the embedding of xiand the embedding of yj. The system can then train the language model neural network and the visual encoder neural network using gradients of a contrastive loss computed using the matrix A. For example, the contrastive loss can be the cross-entropy loss on the rows and columns of A, where the diagonal entries are treated as correct classes while other entries are treated as incorrect classes. A specific example of such a loss is: ಲ^,^ಲೕ,ೕ^^ ൌ െ^^∑ே lo ^^ ∑ே^ ^ ^^^ g ^ ^ ^ log ^^^ , where ^^ is the to steepen or dampen the softmax distributions in the rows and columns of A, and N is the total number of training pairs in the batch. In some cases, prior to computing the matrix A, the system normalizes the contrastive representations and the uni-modal representations of the images and text sequences in the batch. As this loss is minimized, for all pairs in the batch, the embeddings of xi and yi become closer together while becoming farther from all other embeddings of all other training images and text segments in the batch, thereby achieving the goal of the contrastive learning. As another example, the system can instead use a focal contrastive loss. The focal contrastive loss can be expressed as: ^^௩^^ೕ ^^^^ ^^ ൌെ^ ேே^ ^ ^^^^ ൌே^∑ ∑^ୀ^ ^1 െ ^^^^ఊ, ^^ Attorney Docket No.56113-0532WO1 As another example, the overall loss can also include a captioning loss. Generally, the captioning loss term is based on, for each training pair, the respective score distributions (generated by the language model neural network) for the plurality of text tokens in the respective text sequence. As a particular example, the captioning loss for a given training pair may be given by: ^^^^^ ൌ െ∑்௧ୀ^ log ^^ఏ^^^௧|^^ழ௧,^^^ ,with the overall captioning loss being the average of the captioning losses for the training pairs in the batch, T being the total number of positions in the training text sequence in the training pair, and ^^ఏ^^^௧|^^ழ௧,^^^ being the score assigned, in the score distribution that was generated conditioned on the tokens preceding the token at position t in the training text sequence and the image x in the training pair, to the token ^^௧at the position t in the training text sequence. As another example, the overall loss can also include one or more reconstruction losses. The reconstruction losses can include any of a variety of losses that attempt to reconstruct certain ones of the values that are processed or generated during the processing of the image processing neural network. As one example of this, the overall loss can also include an embedding reconstruction loss that is based on the sets of encoded vectors generated for each of the training images. For example, the loss can measure a reconstruction error of an auxiliary neural network in processing the final representation or the embedding vectors identified in the discrete representation to generate a predicted reconstruction of the encoded vectors. As another example of this, the overall loss can also include a pixel reconstruction loss that measures a reconstruction error of an auxiliary neural network in processing the final representation or the embedding vectors identified in the discrete representation to generate a predicted reconstruction of the training image. In some cases, the system initializes the parameters of some or all of the components of the system using a pre-trained neural network. For example, the system can initialize the parameters of the visual encoder neural network to values determined by pre-training the visual encoder neural network as part of a different system. In this example, the system can then keep the visual encoder neural network fixed during the Attorney Docket No.56113-0532WO1 training of the image processing neural network. As a particular example, the visual encoder neural network can have been pre-trained as part of system that uses both contrastive and captioning losses as described above, but uses continuous rather than discrete representations, i.e., does not include the codebook or the vector quantization encoder and decoder neural networks. FIG.4 shows an example 400 of performing an inversion process to generate an output image from discrete tokens. In particular, as shown on the left-hand side of FIG.4, the system can be trained to generate a discrete representation 404 of an input image 402 that represents a given image as a set of embedding vectors from the codebook 135 using the image processing neural network 110. The system can then train a downstream neural network to perform “inversion” on these generated discrete representations 404, i.e., to map a discrete representation 404 to an output image 410 that accurately reconstructs the original input image 402. Thus, by modifying a given discrete representation 404, the system can use the downstream neural network to generate a different output image. Similarly, although not shown in FIG.4, by processing a discrete representation 404 and another conditioning input, e.g., a text input, the system can generate an output image that is similar to the image that is represented by the discrete representation but that differs in certain properties that are specified by the conditioning input. Specifically, two example approaches for performing the inversion task are to use a generative adversarial network (GAN) approach to train the downstream neural network, i.e., to perform GAN inversion 406, or to implement the downstream neural network as a diffusion neural network which uses the discrete representation as a conditioning input, i.e., to perform diffusion inversion 408. An example of using a diffusion neural network to perform diffusion inversion 408 starting from a discrete representation 404 is shown on the right-hand side of FIG.4. In particular, the system generates an image by performing a reverse diffusion process to generate an output image 410 from a given conditioning input that includes the discrete representation 404. To perform the reverse diffusion process, the system initializes a representation x of the output image. For example, the system can sample each value in each representation from a noise distribution, e.g., a Gaussian distribution. The system then updates the representation at each of a plurality of reverse Attorney Docket No.56113-0532WO1 diffusion steps (also referred to as “iterations” or “updating iterations”) using the diffusion neural network, which, in the example of FIG.4 is a convolutional neural network, e.g., a U-Net, that has one or more cross-attention layers that are conditioned on the discrete representation. Each reverse diffusion step is associated with a noise level for the iteration. Generally, each updating iteration has a corresponding time step t and the noise level for the iteration depends on the time step. For example, the noise level can be a decreasing function of the time step t. Examples of such functions include a linear function, a cosine function, and a sigmoid function. Thus, early iterations are associated with higher noise levels and later iterations are associated with lower noise levels, resulting in the diffusion neural network gradually “denoising” the representation to generate the final representation. As part of the updating at any given step, the system generates a denoising output for the reverse diffusion step by processing the representation of the output image (and the conditioning input) using the diffusion neural network. The system then updates the representation of the output image using the denoising output for the reverse diffusion step. For example, the system can map the denoising output to an initial updated representation and then apply a diffusion sampler, e.g., the DDPM (Denoising Diffusion Probabilistic Model) sampler, the DDIM (Denoising Diffusion Implicit Model) sampler or another appropriate sampler, to the initial updated representation to generate an updated representation. Details of the DDPM sampler and the DDIM sampler can be found in, for example, Ho et al. “Denoising Diffusion Probabilistic Models” arXiv: 2006.11239 and Song et al. “Denoising Diffusion Implicit Models” arXiv: 2010.02502.Optionally, after the last reverse diffusion iteration, the system can refrain from using the diffusion sampler and can instead use the initial updated representation as the final representation. That is, the system can keep the initial updated representation from the last reverse diffusion step as the final updated representation and not apply a diffusion sampler to the representation. After the last reverse diffusion step, the system uses the final representation to generate the final output image. For example, the system can use the final representation as the final output image, can apply an upsampling neural network to the final output image, or, when the reverse diffusion steps are performed in latent space rather than in pixel space, apply a decoder neural network to map the final representation to pixel space. Attorney Docket No.56113-0532WO1 This specification uses the term “configured” in connection with systems and computer program components. For a system of one or more computers to be configured to perform particular operations or actions means that the system has installed on it software, firmware, hardware, or a combination of them that in operation cause the system to perform the operations or actions. For one or more computer programs to be configured to perform particular operations or actions means that the one or more programs include instructions that, when executed by data processing apparatus, cause the apparatus to perform the operations or actions. Embodiments of the subject matter and the functional operations described in this specification can be implemented in digital electronic circuitry, in tangibly-embodied computer software or firmware, in computer hardware, including the structures disclosed in this specification and their structural equivalents, or in combinations of one or more of them. Embodiments of the subject matter described in this specification can be implemented as one or more computer programs, i.e., one or more modules of computer program instructions encoded on a tangible non transitory storage medium for execution by, or to control the operation of, data processing apparatus. The computer storage medium can be a machine-readable storage device, a machine-readable storage substrate, a random or serial access memory device, or a combination of one or more of them. Alternatively or in addition, the program instructions can be encoded on an artificially generated propagated signal, e.g., a machine-generated electrical, optical, or electromagnetic signal, that is generated to encode information for transmission to suitable receiver apparatus for execution by a data processing apparatus. The term “data processing apparatus” refers to data processing hardware and encompasses all kinds of apparatus, devices, and machines for processing data, including by way of example a programmable processor, a computer, or multiple processors or computers. The apparatus can also be, or further include, special purpose logic circuitry, e.g., an FPGA (field programmable gate array) or an ASIC (application specific integrated circuit). The apparatus can optionally include, in addition to hardware, code that creates an execution environment for computer programs, e.g., code that constitutes processor firmware, a protocol stack, a database management system, an operating system, or a combination of one or more of them. A computer program, which may also be referred to or described as a program, software, a software application, an app, a module, a software module, a script, or code, can be written in any form of programming language, including compiled or interpreted Attorney Docket No.56113-0532WO1 languages, or declarative or procedural languages; and it can be deployed in any form, including as a stand alone program or as a module, component, subroutine, or other unit suitable for use in a computing environment. A program may, but need not, correspond to a file in a file system. A program can be stored in a portion of a file that holds other programs or data, e.g., one or more scripts stored in a markup language document, in a single file dedicated to the program in question, or in multiple coordinated files, e.g., files that store one or more modules, sub programs, or portions of code. A computer program can be deployed to be executed on one computer or on multiple computers that are located at one site or distributed across multiple sites and interconnected by a data communication network. In this specification, the term “database” is used broadly to refer to any collection of data: the data does not need to be structured in any particular way, or structured at all, and it can be stored on storage devices in one or more locations. Thus, for example, the index database can include multiple collections of data, each of which may be organized and accessed differently. Similarly, in this specification the term “engine” is used broadly to refer to a software-based system, subsystem, or process that is programmed to perform one or more specific functions. Generally, an engine will be implemented as one or more software modules or components, installed on one or more computers in one or more locations. In some cases, one or more computers will be dedicated to a particular engine; in other cases, multiple engines can be installed and running on the same computer or computers. The processes and logic flows described in this specification can be performed by one or more programmable computers executing one or more computer programs to perform functions by operating on input data and generating output. The processes and logic flows can also be performed by special purpose logic circuitry, e.g., an FPGA or an ASIC, or by a combination of special purpose logic circuitry and one or more programmed computers. Computers suitable for the execution of a computer program can be based on general or special purpose microprocessors or both, or any other kind of central processing unit. Generally, a central processing unit will receive instructions and data from a read only memory or a random access memory or both. The essential elements of a computer are a central processing unit for performing or executing instructions and one or more memory devices for storing instructions and data. The central processing unit and the memory can be supplemented by, or incorporated in, special purpose logic Attorney Docket No.56113-0532WO1 circuitry. Generally, a computer will also include, or be operatively coupled to receive data from or transfer data to, or both, one or more mass storage devices for storing data, e.g., magnetic, magneto optical disks, or optical disks. However, a computer need not have such devices. Moreover, a computer can be embedded in another device, e.g., a mobile telephone, a personal digital assistant (PDA), a mobile audio or video player, a game console, a Global Positioning System (GPS) receiver, or a portable storage device, e.g., a universal serial bus (USB) flash drive, to name just a few. Computer readable media suitable for storing computer program instructions and data include all forms of non volatile memory, media and memory devices, including by way of example semiconductor memory devices, e.g., EPROM, EEPROM, and flash memory devices; magnetic disks, e.g., internal hard disks or removable disks; magneto optical disks; and CD ROM and DVD-ROM disks. To provide for interaction with a user, embodiments of the subject matter described in this specification can be implemented on a computer having a display device, e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor, for displaying information to the user and a keyboard and a pointing device, e.g., a mouse or a trackball, by which the user can provide input to the computer. Other kinds of devices can be used to provide for interaction with a user as well; for example, feedback provided to the user can be any form of sensory feedback, e.g., visual feedback, auditory feedback, or tactile feedback; and input from the user can be received in any form, including acoustic, speech, or tactile input. In addition, a computer can interact with a user by sending documents to and receiving documents from a device that is used by the user; for example, by sending web pages to a web browser on a user’s device in response to requests received from the web browser. Also, a computer can interact with a user by sending text messages or other forms of message to a personal device, e.g., a smartphone that is running a messaging application, and receiving responsive messages from the user in return. Data processing apparatus for implementing machine learning models can also include, for example, special-purpose hardware accelerator units for processing common and compute-intensive parts of machine learning training or production, i.e., inference, workloads. Machine learning models can be implemented and deployed using a machine learning framework, e.g., a TensorFlow framework or a Jax framework. Attorney Docket No.56113-0532WO1 Embodiments of the subject matter described in this specification can be implemented in a computing system that includes a back end component, e.g., as a data server, or that includes a middleware component, e.g., an application server, or that includes a front end component, e.g., a client computer having a graphical user interface, a web browser, or an app through which a user can interact with an implementation of the subject matter described in this specification, or any combination of one or more such back end, middleware, or front end components. The components of the system can be interconnected by any form or medium of digital data communication, e.g., a communication network. Examples of communication networks include a local area network (LAN) and a wide area network (WAN), e.g., the Internet. The computing system can include clients and servers. A client and server are generally remote from each other and typically interact through a communication network. The relationship of client and server arises by virtue of computer programs running on the respective computers and having a client-server relationship to each other. In some embodiments, a server transmits data, e.g., an HTML page, to a user device, e.g., for purposes of displaying data to and receiving user input from a user interacting with the device, which acts as a client. Data generated at the user device, e.g., a result of the user interaction, can be received at the server from the device. While this specification contains many specific implementation details, these should not be construed as limitations on the scope of any invention or on the scope of what may be claimed, but rather as descriptions of features that may be specific to particular embodiments of particular inventions. Certain features that are described in this specification in the context of separate embodiments can also be implemented in combination in a single embodiment. Conversely, various features that are described in the context of a single embodiment can also be implemented in multiple embodiments separately or in any suitable subcombination. Moreover, although features may be described above as acting in certain combinations and even initially be claimed as such, one or more features from a claimed combination can in some cases be excised from the combination, and the claimed combination may be directed to a subcombination or variation of a subcombination. Similarly, while operations are depicted in the drawings and recited in the claims in a particular order, this should not be understood as requiring that such operations be performed in the particular order shown or in sequential order, or that all illustrated operations be performed, to achieve desirable results. In certain circumstances, Attorney Docket No.56113-0532WO1 multitasking and parallel processing may be advantageous. Moreover, the separation of various system modules and components in the embodiments described above should not be understood as requiring such separation in all embodiments, and it should be understood that the described program components and systems can generally be integrated together in a single software product or packaged into multiple software products. Particular embodiments of the subject matter have been described. Other embodiments are within the scope of the following claims. For example, the actions recited in the claims can be performed in a different order and still achieve desirable results. As one example, the processes depicted in the accompanying figures do not necessarily require the particular order shown, or sequential order, to achieve desirable results. In some cases, multitasking and parallel processing may be advantageous. What is claimed is:

Claims

Attorney Docket No.56113-0532WO1 CLAIMS 1. A method of training an image processing neural network that is configured to process an input image to generate a discrete representation of the image, wherein the image processing neural network comprises: a visual encoder neural network that is configured to process an input image to generate an encoded representation of the input image, and a vector quantization encoder neural network that is configured to process the encoded representation of the input image to generate a set of encoded vectors representing the input image, and wherein the image processing neural network is configured to generate the discrete representation of the input image by, for each encoded vector, selecting a corresponding embedding vector from the codebook of embedding vectors, the method comprising: obtaining a set of one or more training pairs that each include a respective training image and a respective training text sequence; for each training pair: processing the training image in the training pair using the visual encoder neural network to generate an encoded representation of the training image; processing the encoded representation of the training image using the vector quantization encoder neural network to generate a set of encoded vectors representing the training image; generating a discrete representation of the training image by, for each encoded vector, selecting a corresponding embedding vector from the codebook of embedding vectors; generating a final representation of the training image from the discrete representation; processing the training text sequence in the training pair using a language model neural network to generate a text representation of the text sequence; and training the image processing neural network using, for each training pair, the final representation of the training image in the training pair and the text representation of the text sequence in the training pair.

2. The method of claim 1, further comprising updating the codebook of embedding vectors using the discrete representations of the training images in the pair.Attorney Docket No.56113-0532WO1 3. The method of claim 1 or claim 2, wherein training the image processing neural network comprises holding the visual encoder neural network fixed while training the vector quantization encoder neural network.

4. The method of any preceding claim, wherein the encoded representation comprises a respective initial encoded vector corresponding to each of a plurality of regions of the image, and wherein the set of encoded vectors comprises a respective encoded vector for each of the plurality of regions of the image.

5. The method of claim 4, wherein the vector quantization encoder neural network is a Transformer neural network that is configured to update each of the initial encoded vectors to generate the set of encoded vectors.

6. The method of any preceding claim, wherein generating a final representation of the training image from the discrete representation comprises: processing the discrete representation using a vector quantization decoder neural network to generate the final representation.

7. The method of claim 6, wherein the vector quantization decoder neural network is a Transformer neural network.

8. The method of claim 6 or 7, further comprising: training the vector quantization decoder neural network using, for each training pair, the final representation of the training image in the training pair and the text representation of the text sequence in the training pair.

9. The method of any preceding claim, wherein training the image processing neural network using, for each training pair, the final representation of the training image in the training pair and the text representation of the text sequence in the training pair comprises training the image processing neural network to minimize a loss function.

10. The method of claim 9, wherein the loss function comprises a contrastive loss term that measures a contrastive loss that is computed using, for each training pair, a contrastive image final representation of the training image in the training pair generated from the final representation of the training image and a contrastive text final representation of the training text in the training pair.Attorney Docket No.56113-0532WO1 11. The method of claim 9 or claim 10, wherein the language model neural network comprises: a set of initial neural network layers that are configured to process the text sequence to generate the text representation of the text sequence, wherein the text representation is a uni-modal representation of the text sequence that is independent of the training image; and a set of subsequent neural network layers that are configured to process an input comprising the text representation of the text sequence to generate a respective score distribution for each text token in the text sequence, wherein the subsequent neural network layers comprise one or more cross-modal layers that are conditioned on the final representation of the training image.

12. The method of claim 11, wherein the loss function comprises a captioning loss term that is based on, for each training pair, the respective score distributions for the plurality of text tokens in the respective text sequence.

13. The method of any one of claims 9-12, wherein the loss function comprises an embedding reconstruction loss that is based on the sets of encoded vectors generated for each of the training images.

14. The method of any one of claims 9-13, when dependent on claim 10, wherein the contrastive image final representation of the training image in the training pair is an output of a second attentional pooling layer that is trained jointly with the image processing neural network and that processes as input the final representation of the training image.

15. The method of any one of claims 9-14, when dependent on claim 11, wherein each cross-modal layer receives as input a generative image final representation of the training image in the training pair that is an output of a first attentional pooling layer that is trained jointly with the image processing neural network and that processes as input the final representation of the training image.

16. The method of any preceding claim, further comprising: after the training, using the image processing neural network to perform a downstream task.Attorney Docket No.56113-0532WO1 17. The method of claim 16, wherein the downstream task is an image generation task.

18. The method of claim 16, wherein the downstream task is an image understanding task.

19. The method of claim 16, wherein the downstream task is an image reconstruction task.

20. The method of claim 16, wherein the downstream task is an image classification task.

21. The method of claim 16, wherein the downstream task is an image compression task.

22. The method of claim 16, wherein the downstream task is a video processing task that requires processing respective discrete representations of each video frame in an input video.

23. A system comprising: one or more computers; and one or more storage devices storing instructions that, when executed by the one or more computers, cause the one or more computers to perform the respective operations of the method of any one of claims 1-22.

24. One or more computer-readable storage media storing instructions that when executed by one or more computers cause the one or more computers to perform the respective operations of the method of any one of claims 1-22.

Citation Information

Cited By

  • Visual content generation method and training method and device of visual content generation model

    CN121792813A