Control Caption Neural Network

The Contrastive Caption neural network addresses the inefficiencies in processing multimodal inputs by decoupling the language model and jointly pre-training with contrastive and captioning objectives, resulting in reduced computational and environmental impact while achieving state-of-the-art performance.

JP2025517085AActive Publication Date: 2025-06-03GOOGLE LLC
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
JP2024563401
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2022-05-03
Filing Date
2023-04-28
Publication Date
2025-06-03
Estimated Expiration
2043-04-28

AI Technical Summary

Technical Problem

Existing machine learning models struggle to efficiently process multimodal inputs such as images and text, often requiring multiple stages and data sources, which increases computational overhead and environmental impact.

Method used

The development of a Contrastive Caption (CoCa) neural network that decouples the language model into unimodal and multimodal components, allowing for joint pre-training with contrastive and captioning objectives, thereby reducing computational requirements and environmental footprint.

Benefits of technology

CoCa achieves improved pre-training efficiency with fewer FLOPs and training iterations, leading to reduced carbon emissions and electricity consumption while maintaining state-of-the-art performance on various downstream tasks.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2025517085000001_ABST
    Figure 2025517085000001_ABST
Patent Text Reader

Abstract

A method, system, and apparatus including a computer program encoded on a computer storage medium for processing multimodal input using a contrast caption neural network.
Need to check novelty before this filing date? Find Prior Art

Description

Background Art

[0001] This specification relates to processing an input using a machine learning model.

[0002] As an example, a neural network is a machine learning model that employs one or more layers of non-linear units to predict an output for a received input. Some neural networks include one or more hidden layers in addition to an output layer. The output of each hidden layer is used as input to other layers in the network, such as the next hidden layer or the output layer. Each layer of the network generates an output from the received input according to the current values of each set of weights.

Summary of the Invention

[0003] This specification describes a system implemented as a computer program on one or more computers that processes multimodal input including both visual input, i.e., images or multiple video frames from a video, and text, using a contrastive caption neural network. As described below, since the neural network can be pre-trained using both a contrastive learning loss and a caption loss jointly, the neural network is referred to as a "contrastive caption" neural network.

[0004] Certain embodiments of the subject matter described in this specification may be implemented to realize one or more of the following advantages.

[0005] This specification describes a Contrastive Captioning (CoCa) neural network having an architecture that enables the neural network to be pre-trained jointly with contrastive objectives and caption loss. Unlike a standard encoder-decoder transformer where all decoder layers attend to the encoder output, CoCa omits cross-attention in the first set of decoder layers such that the first set of decoder layers encode unimodal text representations. In other words, CoCa has a decoder (language model neural network) with a plurality of initial self-attention layers that do not have any cross-attention layers. Next, CoCa cascades the remaining decoder layers that cross-attend to the visual encoder to produce multimodal image-text representations. Thus, CoCa effectively decouples the language model neural network into a unimodal text decoder followed by a multimodal text decoder.

[0006] The contrastive loss is applied between the unimodal visual embedding and the text embedding, along with the caption loss in the multimodal decoder output predicting text tokens.

[0007] By sharing the same computational graph, i.e., sharing the same neural network architecture, the two training objectives are efficiently computed with minimal computational overhead. That is, the quantities required for computing both losses are obtained by a single forward pass through the CoCa network.

[0008] This enables the neural network to be pre-trained from scratch in a single step in a unified format of image-text pairs, including one or more of, for example, web-scale alternative text data or annotated images, and natural language supervision for representation learning is seamlessly integrated.

[0009] In other words, for each training pair, the system applies both a contrastive objective between the output of the visual encoder and the output of the unimodal text decoder, and a captioning objective in the output of the multimodal decoder.

[0010] Furthermore, CoCa can be trained on both image annotation data and noisy image-text data by simply treating all labels as text. Thus, the generation loss on image annotation text provides a fine-grained training signal similar to the cross-entropy loss approach of a single encoder, effectively encompassing all three pre-training paradigms in a single integrated way.

[0011] Moreover, as a result of CoCa's decoupled decoder (language model) design, both training losses can be efficiently accounted for. Since the autoregressive language model is trained using causal masking on complete sentences, the decoder can efficiently generate outputs for both the contrastive loss and the generation loss in a single forward pass (compared to the two passes of the bidirectional approach).

[0012] Therefore, most of the computation is shared between the two losses, and CoCa induces minimal overhead compared to standard encoder-decoder models. In contrast, while many existing methods train model components in multiple stages on various data sources and / or modalities, CoCa is trained end-to-end directly from scratch using various data sources (e.g., using both annotated images and noisy alternative text images) by treating all labels as text for both contrastive and generative purposes.

[0013] Therefore, the described technique can achieve improved pre-training efficiency, i.e., achieve performance equivalent to or better than the prior art using fewer FLOPs and fewer training iterations.

[0014] Pre-training a large-scale multimodal model that can be used for real-world tasks generally results in a large amount of carbon dioxide (CO 2 ) emissions and a large amount of electricity consumption, because, for example, the dataset on which pre-training is performed is very large and the model has a fairly large number of parameters. For the reasons described above, by reducing the number of FLOPs that need to be performed and performing fewer training iterations, the technology described significantly reduces the CO 2 footprint of the pre-training process while also significantly reducing the amount of electricity consumed by the pre-training process.

[0015] Additionally, with this pre-training scheme, the neural network can achieve state-of-the-art performance on a wide range of downstream tasks, either through zero-shot transfer or minimal task-specific adaptation. Specific examples of downstream tasks will be described in more detail below.

[0016] Details of one or more embodiments of the subject matter of this specification are set forth in the accompanying drawings and the description below. Other features, aspects, and advantages of the subject matter will become apparent from the description, the drawings, and the claims.

Brief Description of the Drawings

[0017]

Figure 1

Figure 2

Figure 3

Figure 4

Modes for Carrying Out the Invention

[0018] Like reference numerals and designations in the various drawings refer to like elements.

[0019] FIG. 1 shows an example of a neural network system 100. The neural network system 100 is an example of a system implemented as a computer program on one or more computers in one or more locations, and the systems, components, and techniques described below may be implemented.

[0020] System 100 is a system that processes a visual input 102, i.e., a multimodal input including both a plurality of video frames from an image or video and text, using a contrastive caption neural network 110.

[0021] As will be described below, since the neural network 110 can be pre-trained jointly using both a contrastive learning loss and a caption loss, the neural network 110 is referred to as a “contrastive caption” neural network.

[0022] The contrastive caption neural network 110 includes a visual encoder neural network 112 configured to process a visual input 102 including one or more images to generate an encoded representation 114 of the visual input 102, and (ii) a set of initial unimodal neural network layers 122 and a set of subsequent neural network layers 126 including both cross-modal layers and unimodal layers, a language model neural network 120.

[0023] That is, when processing multimodal input that includes both visual input 102 and a text sequence, the representation generated by the initial layer 122 within the language model neural network 120 is a unimodal representation 124 that depends only on the text sequence, while the representation generated by the subsequent layer 126 is a multimodal representation that depends on both the text sequence representation 124 of the text sequence generated by the initial layer 122 and the visual input 102.

[0024] Generally, the language model neural network 120 is configured to process the current text sequence 104 in order to generate an output that defines a new token 128 that is to be added to the current text sequence 104.

[0025] The output that defines the new token 128 that is to be added to the current text sequence 104 generally includes respective scores for each token within the token vocabulary. The token vocabulary can include any of characters, subwords, words, punctuation marks, sign tokens (e.g., #, $, and other signs), mathematical symbols, etc. The token vocabulary can also include one or more special tokens that are added to the input text sequence processed by the neural network, such as a sequence start token, a sequence end token, a designated "class" token, etc.

[0026] During training, the language model neural network 120 can generate respective outputs for each of the multiple tokens of the input sequence in a single forward pass, i.e., in parallel, by processing a single "current sequence" 104 that represents the entire input text sequence.

[0027] After training, at each time step, the current text sequence 104 is processed at that time step, and then the current text sequence 104 is updated by selecting a token from the vocabulary using the output for the current text sequence, and then the selected token is appended to the end of the current text sequence 104, so that the language model neural network 120 can be used to autoregressively generate the text sequence.

[0028] The visual encoder neural network 112 has parameters (referred to as "visual encoder neural network parameters" or "visual encoder parameters") to generate an encoded representation 114 of the visual input 102, receives the visual input 102, and is a neural network that processes the visual input 102 according to the parameters.

[0029] Generally, the encoded representation 114 includes an embedding (also referred to as an "updated token") for each of a plurality of patches of the visual input 102, for example, for each of a plurality of spatial patches (regions) of each image of the visual input 102, or in some cases where the visual input 102 includes a plurality of images, for each of a plurality of spatio-temporal patches (regions) of the visual input 102.

[0030] As used herein, an "embedding" is a numerical value having a predetermined dimensionality, such as a floating-point value or a vector of other values. The space of possible vectors having a predetermined dimensionality is referred to as an "embedding space".

[0031] The visual encoder neural network 112 can have any suitable architecture that enables the neural network 112 to map the input visual input 102 to an encoded representation 114. For example, the visual encoder neural network 112 can be a convolutional neural network. As another example, the visual encoder neural network 112 can be a neural network of a vision transformer having one or more self-attention layers. As yet another example, the visual encoder neural network 112 can be a neural network having a mix of both convolutional layers and self-attention layers.

[0032] The language model neural network 120 can have any suitable architecture that enables the language model neural network 120 to map the tokens of a text sequence to respective unimodal representations 124 for each of the tokens and then map the unimodal representations 124 to an output that defines the next token 128.

[0033] In a particular example, the language model neural network 120 can have an attention-based architecture, such as the architecture of a decoder-only transformer neural network.

[0034] In this example, a set of initial unimodal neural network layers 122 can include a sequence of initial attention layers, where each initial attention layer is configured to receive, as input, the respective current representations of each text token of the current text sequence and process the respective current representations to generate, as output, the respective updated representations of each text token of the current text sequence. For example, each initial attention layer can apply a causally masked self-attention mechanism to the respective current representations to generate the respective updated representations.

[0035] The self-attention mechanism for each current representation refers to an attention mechanism that calculates a query, a key, and a value from each current representation.

[0036] The causally masked self-attention mechanism for each current representation refers to an attention mechanism in which a given position in the current text sequence does not attend to any position after the given position in the current text sequence, i.e., has no non-zero attention weights.

[0037] Each attention layer can optionally apply other operations to the representation as part of updating the representation, such as by using a position-wise feed-forward neural network, by applying layer normalization, by using residual connections, etc.

[0038] In this example, each current representation received as input by the first initial attention layer in the sequence of initial attention layers is the respective embedding of each text token of the current text sequence (e.g., generated by the embedding layer of the language model neural network 120), and each subsequent current representation received as input by each initial attention layer, i.e., each initial attention layer after the first initial attention layer in the sequence of initial attention layers, is the respective updated representation of the text tokens of the current text sequence generated as output by the preceding initial attention layer in the sequence of initial attention layers.

[0039] Therefore, each unimodal representation of the text tokens of the current text sequence is the respective updated representation of the text tokens of the current text sequence generated as output by the last initial attention layer in the sequence of initial attention layers.

[0040] More specifically, when the language model neural network 120 has an attention-based architecture, to ensure that the updated representation of each text token in the current text sequence is a unimodal representation that depends only on the current text sequence and not on the visual input, the initial attention layer includes a plurality of self-attention layers but does not include any cross-attention layers.

[0041] Furthermore, the set of subsequent neural network layers 126 includes a sequence of subsequent attention layers, and each subsequent attention layer is configured to receive, as input, the current representation of each text token in the current text sequence and to process each current representation to generate, as output, an updated representation of each text token in the current text sequence.

[0042] Accordingly, each current representation received as input by the first subsequent attention layer in the sequence of subsequent attention layers is a unimodal representation of each text token in the current text sequence (generated by the initial attention layer), and each current representation received as input by each subsequent attention layer after the first subsequent attention layer in the sequence of subsequent attention layers is an updated representation of each text token in the current text sequence generated as output by the preceding subsequent attention layer in the sequence of subsequent attention layers.

[0043] In this example, similar to the initial attention layer, the sequence of subsequent attention layers also includes one or more self-attention layers. That is, for one or more of the subsequent neural network layers, processing each current representation in order to generate, as output, an updated representation for each text token of the current text sequence involves applying a causally masked self-attention mechanism. Optionally, each of these attention layers can apply other operations to the representation as part of updating the representation, such as by utilizing a position-wise feed-forward neural network, by applying layer normalization, by utilizing residual connections, and so on.

[0044] Unlike the initial layer, the sequence of subsequent attention layers also includes one or more cross-modal layers. Each cross-modal layer processes each current representation in order to generate, as output, an updated representation for each text token of the current text sequence by applying a cross-attention mechanism between an input derived from the encoded representation of the visual input (generated from the encoded representation) and the current representation of each text token of the current text sequence received as input by the cross-modal layer.

[0045] "Cross-attention" between an input derived from the encoded representation of the visual input and the current representation of each text token of the current text sequence received as input by the cross-modal layer refers to an attention mechanism that uses a query derived from the current representation of each text token of the current text sequence and keys and values derived from the input generated from the encoded representation of the visual input.

[0046] Specific examples of self-attention, cross-attention, and causally masked self-attention mechanisms that can be adopted by the system are described below. Vaswani et al. “Attention is all you need”, 31st Conference on Neural Information Processing Systems (NIPS 2017), Long Beach, CA, USA; Colin Raffel, Noam Shazeer, Adam Roberts, Katherine Lee, Sharan Narang, Michael Matena, Yanqi Zhou, Wei Li, Peter J Liu “Exploring the limits of transfer learning with a unified text-to-text transformer”, arXiv preprint arXiv:1910.10683, 2019; Daniel Adiwardana, Minh-Thang Luong, David R. So, Jamie Hall, Noah Fiedel, Romal Thoppilan, Zi Yang, Apoorv Kulshreshtha, Gaurav Nemade, Yifeng Lu, Quoc V. Le “Towards a human-like open-domain chatbot”, CoRR, abs / 2001.09977, 2020; Tom B Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, et al. “Language models are few-shot learners”, arXiv preprint arXiv:2005.14165, 2020; Hua, et al. “Transformer Quality in Linear Time”, arXiv preprint arXiv:2202.10447, 2022。

[0047] In some embodiments, the input derived from the encoded representation (hereinafter also referred to as the caption representation) is an embedding of the encoded representation 114.

[0048] In some other embodiments, the neural network 110 applies one or more transformations to the encoded representation 114 to generate the input provided to the cross-modal layer. An example of these transformations will be described in more detail below with reference to FIG. 3.

[0049] Thus, the updated representation generated by a given cross-modal layer is a multimodal representation that depends on the visual input and the text tokens of the current text sequence. Each of these cross-modal layers can optionally apply other operations to the representation as part of updating the representation, such as by using a position-wise feed-forward neural network, by applying layer normalization, by using residual connections, etc.

[0050] As an example, the sequence of subsequent attention layers can alternate between self-attention layers and cross-modal layers. As another example, the sequence of subsequent attention layers can include a cross-modal layer after every two, three, or four self-attention layers.

[0051] Thus, due to the presence of the cross-modal layer, the representation generated by the last subsequent attention layer of the sequence is a multimodal representation as described above.

[0052] To generate the score distribution, the set of subsequent neural network layers 126 can also include an output layer block.

[0053] The output layer block is a set of one or more neural network layers, e.g., one or more fully connected layers followed by a softmax layer, and receives one or more of the updated representations of each of the text tokens of the current text sequence that are generated as output by the last subsequent attention layer of a sequence of subsequent attention layers, and processes one or more of each of the updated representations to generate an output that defines a new token to be added to the current text sequence, i.e., to generate a score distribution over the tokens of the vocabulary.

[0054] For example, during training, when the current output sequence is the entire training sequence, the output layer block can generate, for each text token, a respective score distribution over the text tokens by processing the updated representations of the tokens immediately preceding the text tokens of the training sequence, in parallel for each of the text tokens. In this example, the system can extend the training sequence to have a specified sequence start of tokens before processing the training sequence using a language model neural network.

[0055] After training, when the system 100 is operating autoregressively, the output layer block can generate a single score distribution over the current output sequence by processing the updated representation for the last token of the current output sequence. Next, the system 100 can select the next token to be added to the current output sequence using the score distribution generated by the output layer block. For example, the system 100 can select the token with the highest score within the score distribution or can sample a token from the score distribution.

[0056] Generally, system 100 or other training systems can pre-train the contrast caption neural network 110 with both a contrast loss and a caption loss.

[0057] The contrast loss may depend on the encoded representation generated by the visual encoder 112 and the representation 124 generated by the initial neural network layer 122, while the caption loss may depend on the encoded representation 124 generated by the visual encoder and the representation generated by the subsequent neural network layer 126.

[0058] That is, due to the architecture of the neural network 110, the contrast caption neural network 110 can be effectively pre-trained by jointly using both the contrast loss and the caption loss without increasing the number of forward passes that need to be performed through the language model neural network 120 and the visual encoder neural network 112.

[0059] Pre-training the neural network 110 will be described in more detail below with reference to FIGS. 2 and 3.

[0060] After the contrast caption neural network 110 has been pre-trained, the visual encoder 112, the initial layer 122, the subsequent layer 126, or some combination of the foregoing can be used for the downstream task.

[0061] In some embodiments, the downstream task can be performed in a zero-shot manner, i.e., without further training any of the components of the contrast caption neural network 110.

[0062] In some other embodiments, the downstream task can be performed after fine-tuning one or more of the components of the contrast caption neural network 110 with labeled training data for the downstream task.

[0063] For example, while fixing any part of the visual encoder 112 and the language model neural network 120 used for the downstream task, the system 100 can learn a customized attention pooling layer and, optionally, one or more additional output layers specialized for the downstream task that receive the output of the attention pooling layer, the output of one of the layers of the language model neural network 120, or both.

[0064] As another example, the system 100 can also fine-tune any part of the visual encoder 112 and the language model neural network 120 used for the downstream task.

[0065] In some examples, the downstream task may be an image or video processing task.

[0066] In some examples, the downstream task is a visual classification task that requires classifying visual input into one of a set of categories, each corresponding to a different object type.

[0067] In some other examples, the downstream task is a visual action recognition task that requires classifying video input into one of a set of action categories.

[0068] In some examples, the downstream task is a cross-modal retrieval task that requires (i) retrieving one or more text sequences that are most similar to the visual input, or (ii) retrieving one or more visual inputs that are most similar to the text sequence.

[0069] In some examples, the downstream task is a multimodal understanding task. For example, the task can be a visual question answering task (VQA) that requires generating an answer to a question posed with respect to the visual input.

[0070] In some examples, the downstream task is an image captioning task that requires generating a text caption for visual input.

[0071] The downstream task is described in more detail below with reference to FIG. 4.

[0072] FIG. 2 is a flowchart of an exemplary process 200 for training a caption neural network. For convenience, process 200 is described as being performed by a system of one or more computers located in one or more locations. For example, a neural network system, such as neural network system 100 of FIG. 1 appropriately programmed, can perform process 200.

[0073] The system can repeat iterations of process 200 with different batches of training examples to update the parameters of the visual encoder neural network, the language model neural network, or both.

[0074] That is, in each iteration of process 200, the system obtains a batch of training pairs, for example, by sampling a batch from a larger set of training data, and uses the batch of one or more training pairs to update the parameters of the visual encoder neural network and the language model neural network.

[0075] The system can continue to perform iterations of process 200 until an end criterion for training the neural network is met, for example, until the parameters converge, until a threshold amount of wall clock time has elapsed, or until a threshold number of iterations of process 200 have been performed.

[0076] The system obtains a batch of one or more training pairs (step 202).

[0077] Each training pair includes a visual input and an input text sequence.

[0078] Specifically, the input text sequence is determined by the system or an external source to describe the content of the visual input or be related to the visual input. In other words, the visual input and the input text sequence are determined to be semantically similar.

[0079] For example, within a given training pair, the text sequence can be the text annotation of the visual input from a set of manually or automatically generated image annotations, or the alternative text associated with the visual input of a set of alternative text data. Alternative text is, for example, the text that is displayed on a web page instead of an image when the image cannot be rendered properly or cannot be loaded. For example, the system can obtain alternative text data from data maintained by an Internet search engine or other software that automatically crawls web pages on the Internet.

[0080] For each training pair, the system processes each visual input and each text sequence of the training pair using a contrastive caption neural network (step 204).

[0081] Specifically, for each training pair, the system processes the visual input of the training pair using a visual encoder neural network to generate an encoded representation of the visual input.

[0082] The system processes the text sequence of the training pair using a set of initial neural network layers to generate a respective unimodal representation for each text token of the text sequence. Since the representation does not depend on the visual input and depends only on the text tokens of the text sequence, the representation is called "unimodal".

[0083] For each of the plurality of text tokens from each text sequence, the system processes the respective unimodal representation of the text tokens of the text sequence using a set of subsequent neural network layers to generate a respective score distribution for the vocabulary of the text tokens. As described above, the system can generate these score distributions in parallel for each of the plurality of text tokens.

[0084] For each training pair, processing the text sequence of the training pair using a set of initial neural network layers and processing the respective unimodal representations using a set of subsequent neural network layers are performed in a single forward pass through the language modeling neural network. That is, the system only needs to perform a single forward pass through the language model neural network to generate both the unimodal representation (used to calculate the contrastive loss) and the score distribution (used to calculate the caption loss).

[0085] The system trains the neural network (step 206) to minimize a loss function that includes (i) a contrastive learning loss term based on the similarity between a contrastive representation derived from the encoded representation of the visual input and one or more unimodal representations of the text tokens from each of the text sequences of the training pairs, and (ii) for each training pair, a caption loss term based on the respective score distributions for the plurality of text tokens of the respective text sequences.

[0086] That is, the system can calculate the gradient of the loss function with respect to the parameters of the visual encoder neural network and the language model neural network, for example, through backpropagation of errors, and then apply an optimizer to the gradient to update the parameters of the visual encoder neural network and the language model neural network.

[0087] As described above, the system only needs to perform a single forward pass through the language model neural network to generate both the unimodal representation (used to calculate the contrastive loss) and the score distribution (used to calculate the caption loss). Therefore, the system can use different outputs of the same forward pass to calculate the amounts required for each of the two losses.

[0088] Thus, even if the neural network is trained on both the contrastive loss and the caption loss, only a single forward pass through the visual encoder neural network and the language model neural network is required to evaluate both losses.

[0089] More specifically, the contrastive loss is based on the "contrastive representation" of each of the visual inputs in the batch and one or more of the unimodal representations for one or more of the text sequences in the batch, for each text sequence in the batch.

[0090] As a specific example, each text sequence in the batch can include the same specified token placed at the same position within each text sequence. For example, the system or another system can extend each text sequence with a specified token, such as a "CLS" token placed at the end of all text sequences. Next, the system can calculate the contrastive loss using the unimodal representation of that specified token.

[0091] Calculating the contrastive representation will be described in more detail below with reference to Figure 3.

[0092] The goal of the contrastive loss is to train the visual encoder 112 and the language model 120 such that inputs with similar semantics are mapped to nearby points regardless of their modality, so that the visual encoder 112 and the language model 120 can embed the image and text inputs into a representation space, i.e., the space of contrastive representations and unimodal representations.

[0093] Accordingly, the system can train the neural network 112 and the neural network 120 to push the embeddings of x i and the text sequence y i close to each other for all training pairs in a batch that include x i and y i while pushing them away from all other embeddings of all other visual inputs and text segments in the batch.

[0094] Next, a specific example of the contrastive loss 130 is described.

[0095] Based on the embeddings of the image and text segments of the mini-batch pairs, an N×N similarity matrix A is calculated, where A i;j represents how similar the embedding of x i is to the embedding of y j . For example, A i;j can be the dot product between the embedding of x i and the embedding of y j .

[0096] Next, the system can train the language model neural network and the visual encoder neural network using the gradients of the contrastive loss calculated using the matrix A. For example, the contrastive loss can be the cross-entropy loss over the rows and columns of A, where the diagonal entries are treated as the correct classes while the other entries are treated as incorrect classes. A specific example of such a loss is as follows. [Number] Where σ is the softmax temperature that scales the logits, for example, it functions to sharpen or flatten the softmax distribution of the rows and columns of A, and N is the total number of training pairs in the batch. In some cases, before calculating the matrix A, the system normalizes the contrastive and unimodal representations of the visual inputs and text sequences in the batch.

[0097] As this loss is minimized, for all pairs in the batch, the embeddings of x i and y i become closer to each other, while moving away from all other embeddings of all other visual inputs and text segments in the batch, thereby achieving the goal of contrastive learning.

[0098] The caption loss term measures the quality of each score distribution for a token, compared to the corresponding token of the text sequence, for each training pair and for each of the multiple tokens from each text sequence. As described above, since the score distribution is generated using the output of the subsequent attention layer of the language model neural network, the score distribution depends on both the visual input and the text sequence.

[0099] As a specific example, the caption loss for a given training pair may be given by [Number] and the overall caption loss is the average of the caption losses for the training pairs in the batch, where T is the total number of positions in the training text sequence of the training pair, and P θ (y t |y <t, x) is the score distribution generated conditional on the token preceding the token at position t of the visual input x of the training text sequence and the training pair, for the token at position t of the training text sequence y t is the score assigned to.

[0100] The overall loss for pre-training can be, for example, a weighted sum of the caption loss and the contrastive loss.

[0101] FIG. 3 is a diagram showing an example of training the neural network 110 with the contrastive loss and the caption loss.

[0102] Specifically, FIG. 3 shows an example of how the neural network 110 is trained with a training pair including the image 310 and the text sequence 320 "two dogs running in a field".

[0103] As shown in FIG. 3, the system processes the image 310 using the visual encoder neural network 112 to generate an encoded representation of the image 310.

[0104] Next, the system generates two representations from the encoded representation: a contrastive representation used for the contrastive loss as described above, and a caption representation (i.e., this is an input derived from the encoded representation) used to condition the cross-modal layer within the language model neural network 120 as described above.

[0105] The system can generate the contrastive representation in any of a variety of ways.

[0106] As an example, the system can use the embedding of a specified token in the encoded representation as a contrastive representation. That is, when splitting a given visual input into patches, the neural network 112 can add placeholder patches that do not correspond to any of the patches in the visual input. Next, the system can use the embedding of this placeholder patch as a contrastive representation.

[0107] As another example, the system can use pooling to generate a contrastive representation. As an example, the system can apply a pooling operation, e.g., global average pooling (GAP), to the embedding of the encoded representation and use the resulting pooled embedding as a contrastive representation.

[0108] As yet another example, the system can use learned attention pooling to generate a contrastive representation. To perform learned attention pooling, the system can incorporate an attention pooling layer within the contrastive caption neural network 110.

[0109] For each of the learned query tokens in the second set, the attention pooling layer applies attention to the updated tokens, i.e., the embeddings of the encoded representation, and the set of learned query tokens to generate respective updated query tokens. That is, the system generates a query for the attention mechanism applied by the attention pooling layer using the learned query tokens and generates keys and values for the attention mechanism using the embeddings of the encoded representation. Thus, the output of the attention pooling layer is the updated query tokens for each of the set of query tokens.

[0110] The set of learned query tokens and the parameters of the attention mechanism are co-learned with the parameters of the language model neural network and the visual encoder neural network during pre-training.

[0111] To generate contrastive representations, the system can use a set of learned query tokens that have only a single query token, and can use the query tokens updated for the single query token as contrastive representations.

[0112] The system can generate caption representations in any of a variety of ways.

[0113] As an example, the system can directly use the encoded representation as the caption representation.

[0114] As another example, the system incorporates another attention pooling layer where the set of learned query vectors has multiple query vectors, and then can use the updated query tokens generated by the attention pooling layer as the encoded representation.

[0115] When used, the attention pooling layer can function as a "task adapter", thereby (i) ensuring that the caption task receives a more fine-grained input that separately represents different regions within the visual input, while the contrast task receives a global representation that represents the entire visual input, and (ii) enabling the output of the visual encoder to be adapted in different learned ways for each of the tasks, which can improve the quality of pre-training in many situations.

[0116] Next, as described above, the system processes the text sequence 320 using the language model neural network 120. Specifically, the language model neural network 120 first generates a unimodal representation by processing the text sequence 320 using an initial layer 122 (“unimodal text decoder”), and then, to generate an output that defines the tokens of the text sequence 320, processes these unimodal representations conditioned by cross-attention on the caption representation using a subsequent layer 126 (“multimodal text decoder”).

[0117] In the example of FIG. 3, the system then calculates a contrastive loss using the unimodal representation for the “[CLS]” token and the contrastive representation generated using attention pooling while calculating a caption loss using the output for the tokens of the text sequence 320 as described above.

[0118] FIG. 4 shows how the contrastive caption neural network 110 can be used for various downstream tasks.

[0119] As shown in FIG. 4, the system first performs pre-training 402 of the “CoCa” neural network 110 as described above.

[0120] Next, the system can perform zero-shot, frozen feature, or fine-tuning downstream adaptation 404 to use at least a portion of the CoCa neural network 110 for downstream tasks.

[0121] That is, in some embodiments, the neural network 110 can be adapted to downstream tasks in a zero-shot manner, i.e., without further training any of the components of the contrastive caption neural network 110 or any additional components.

[0122] In some other embodiments, the downstream task can be performed after fine-tuning one or more of the components of the contrast caption neural network 110 with labeled training data for the downstream task, e.g., via supervised learning.

[0123] For example, in the case of frozen feature adaptation, the system can learn a task-customized attention pooling layer while fixing any part of the visual encoder 112 and the language model neural network 120 used for the downstream task, and, optionally, one or more additional output layers specialized for the downstream task that receive the output of the attention pooling layer, the output of one of the layers of the language model neural network 120, or both.

[0124] As another example, in the case of fine-tuning adaptation, the system can also fine-tune any part of the visual encoder 112 and the language model neural network 120 used for the downstream task.

[0125] In some examples, the downstream task is a visual classification task 406 that requires classifying the visual input into one of a set of categories corresponding to different object types. In this example, the system can use at least the visual encoder 112 and then process the encoded representation generated by the visual encoder 112 to generate the classification.

[0126] For example, the system can perform zero-shot visual classification by using the language model neural network to process the text labels for the set of categories to generate unimodal representations for each category, and using the visual encoder 112 to process the visual input to generate a contrastive representation for the visual input. Next, the system can select the category with the unimodal representation that is most similar to the contrastive representation as the classification of the visual input.

[0127] In some other examples, the downstream task is a visual action recognition task that requires classifying video input into one of a set of action categories.

[0128] In these examples, the system can use only the visual encoder 112 and can learn an attention pooling layer and one or more output layers customized for the visual action recognition task. As a specific example, the system can obtain multiple frames of a video and feed each frame individually to the shared visual encoder. In the case of frozen feature evaluation or fine-tuning, the system can use the cross-entropy loss of softmax to learn additional poolers in addition to the spatial and temporal feature tokens. Note that since the pooler has a single query token, the calculation of pooling for all spatial and temporal tokens is not expensive.

[0129] In some examples, the downstream task is a cross-modal alignment task 408 that requires (i) retrieving one or more text sequences that are most similar to the visual input, or (ii) retrieving one or more visual inputs that are most similar to the text sequence. In these examples, the system can use the visual encoder neural network 112 and the initial layer of the language model neural network (unimodal text decoder). For example, in the case of zero-shot video-text retrieval, the system can use a simple approach where the system calculates the average embedding of a set of frames of the video (the frames are evenly sampled from the video) and uses the average embedding as the representation of the video.

[0130] In some examples, the downstream task is a multimodal understanding task 410. In these examples, the system can use the entire language model neural network 120 (both the unimodal decoder and the multimodal decoder) and the visual encoder 112.

[0131] For example, the task may be a visual question answering task (VQA) that requires generating an answer to a question posed with respect to a visual input.

[0132] As another example, the downstream task is an image captioning task that requires generating a text caption for a visual input.

[0133] Thus, after performing downstream adaptation 404, the system can receive a new input for the downstream task, process the new input using the neural network of the downstream task, and generate a task output for the downstream task. Depending on the downstream task, the neural network of the downstream task can include one or more of (i) a visual encoder, (ii) an initial set of neural network layers, or (iii) a subsequent set of neural network layers, and generate an output task for the downstream task.

[0134] Table 1 shows examples of the performance of CoCa in two downstream tasks - image classification (left) and video action recognition (right) - compared to existing techniques. Table 1 shows the performance of downstream adaptation using a frozen technique where the components of CoCa are not further trained, and a fine-tuning technique where the components of CoCa are fine-tuned.

Table 1

[0135] As can be seen from Table 1, the described technique can compete with the existing technique even without fine-tuning, and outperforms the existing technique in both tasks by performing fine-tuning.

[0136] This specification uses the term "configured" in connection with components of a system and computer programs. Being configured to perform a particular operation or action by one or more computer systems means that software, firmware, hardware, or combinations thereof that cause the operation or action to be performed are installed on the systems during operation. Being configured to perform a particular operation or action by one or more computer programs means that the one or more programs include instructions that cause the operation or action to be performed when executed by a data processing apparatus.

[0137] Embodiments of the subject matter and the functional operations described in this specification can be implemented in digital electronic circuitry, in tangibly embodied computer software or firmware, in computer hardware, or in one or more combinations thereof, including the structures disclosed in this specification and their structural equivalents. Embodiments of the subject matter described in this specification can be implemented as one or more modules of computer program instructions, i.e., as one or more computer programs encoded on a tangible non-transitory storage medium for execution by, or to control the operation of, a data processing apparatus. The computer storage medium may be a machine-readable storage device, a machine-readable storage substrate, a random or serial access memory device, or one or more combinations thereof. Alternatively or additionally, the program instructions may be encoded in an artificially generated propagated signal, e.g., a mechanically generated electrical, optical, or electromagnetic signal, generated to encode information for transmission to a receiver device suitable for execution by a data processing apparatus.

[0138] The term "data processing apparatus" refers to data processing hardware and includes, by way of example, all kinds of devices, apparatuses, and machines for processing data, including programmable processors, computers, or multiple processors or computers. The apparatus can also be, or further include, special-purpose logic circuits, such as FPGAs (Field Programmable Gate Arrays) or ASICs (Application Specific Integrated Circuits). The apparatus can optionally include, in addition to the hardware, code for creating an execution environment for a computer program, such as processor firmware, protocol stacks, database management systems, operating systems, or code constituting one or more combinations thereof.

[0139] A computer program (also referred to or described as a program, software, software application, app, module, software module, script, or code) can be written in any form of programming language, including compiled or interpreted languages, or declarative or procedural languages, and it can be deployed in any form, including as a stand-alone program or as a module, including components, subroutines, or other units suitable for use in a computing environment. The program can correspond to a file in a file system, but does not necessarily have to. The program can be stored in a part of a file that holds one or more scripts stored in a document of a markup language, in a single file dedicated to the program of interest, or in multiple associated files, such as files that store one or more modules, subprograms, or portions of code. A computer program can be deployed to be executed on one computer or on multiple computers located in one place or distributed across multiple places and interconnected by a data communication network.

[0140] As used herein, the term "database" is used broadly to refer to any collection of data. The data need not be structured in any particular way, or indeed structured at all, and can be stored on storage devices in one or more locations. Thus, for example, an index database can include multiple collections of data, each of which can be organized and accessed in a different way.

[0141] Similarly, as used herein, the term "engine" is used broadly to refer to a software-based system, subsystem, or process programmed to perform one or more specific functions. Generally, an engine is implemented as one or more software modules or components and installed on one or more computers in one or more locations. In some cases, one or more computers are dedicated to a particular engine, while in other cases, multiple engines can be installed and run on the same one or more computers.

[0142] The processes and logical flows described herein can be performed by one or more programmable computers that execute one or more computer programs to perform functions by operating on input data to generate output. Also, the processes and logical flows can be performed by special-purpose logic circuits, such as FPGAs or ASICs, or by a combination of special-purpose logic circuits and one or more programmed computers.

[0143] A computer suitable for executing a computer program can be based on a general-purpose or special-purpose microprocessor, or both, or other types of central processing units. Generally, the central processing unit receives instructions and data from read-only memory, random access memory, or both. The basic elements of a computer are a central processing unit for executing and running instructions and one or more memory devices for storing instructions and data. The central processing unit and memory can be complemented or incorporated by dedicated logic circuits. Generally, a computer also includes, or is operatively coupled to receive data from, or transmit data to, one or more mass storage devices for storing data, such as magnetic disks, magneto-optical disks, or optical disks, for example. However, a computer need not have such devices. Further, by way of example, a computer can be embedded in other devices, such as cellular telephones, personal digital assistants (PDAs), mobile audio or video players, game consoles, global positioning system (GPS) receivers, or portable storage devices, such as universal serial bus (USB) flash drives.

[0144] A computer-readable medium suitable for storing computer program instructions and data includes any form of non-volatile memory, media, and memory devices, including, by way of example, semiconductor memory devices, such as EPROM, EEPROM, and flash memory devices, magnetic disks, such as internal hard disks or removable disks, magneto-optical disks, and CD ROM and DVD-ROM disks.

[0145] To provide interaction with a user, embodiments of the subject matter described herein may be implemented on a computer having a display device for displaying information to the user, such as a CRT (cathode ray tube) or LCD (liquid crystal display) monitor, and a keyboard and a pointing device, such as a mouse or trackball, by which the user can provide input to the computer. Other types of devices may also be used to provide interaction with the user. For example, the feedback provided to the user may be any form of sensory feedback, such as visual feedback, auditory feedback, or tactile feedback, and the input received from the user may be received in any form, including acoustic input, voice input, or tactile input. Further, the computer can interact with the user by sending and receiving documents to and from the devices used by the user, such as by sending a web page to a web browser on the user's device in response to a request received from the web browser, or by sending a text message or other form of message to a personal device, such as a smartphone on which a messaging application is running, and receiving a response message from the user in return.

[0146] A data processing apparatus for implementing a machine learning model may also include, for example, a dedicated hardware accelerator unit for processing the general and numerical computing portions of a machine learning training or machine learning production, i.e., inference, workload.

[0147] The machine learning model may be implemented and deployed using a machine learning framework, such as the TensorFlow framework or the Jax framework.

[0148] Embodiments of the subject matter described herein may be implemented in a computing system that includes back-end components, such as a data server, or middleware components, such as an application server, or front-end components, such as a graphical user interface, a web browser, or an app with which a user can interact with an embodiment of the subject matter described herein, or any combination of one or more such back-end, middleware, or front-end components. The components of the system may be interconnected by any form or medium of digital data communication, such as a communication network. Examples of communication networks include local area networks (LANs), wide area networks (WANs), such as the Internet.

[0149] A computing system may include clients and servers. Clients and servers are generally remote from each other and typically interact through a communication network. The relationship of client and server arises by virtue of computer programs running on respective computers and having a client-server relationship to each other. In some embodiments, a server may send data, such as an HTML page, to a user device for the purpose of displaying data to a user who interacts with a device acting as a client, for example, and receiving user input from the user. Data generated on the user device, such as the results of user interactions, may be received at the server from the device.

[0150] This specification includes details of many specific embodiments, which should not be construed as limiting the scope of any invention or the scope of what can be claimed, but rather as descriptions of features that may be specific to particular embodiments of a particular invention. Specific features described in the context of individual embodiments herein may also be implemented in combination in a single embodiment. Conversely, the various features of the invention described in the context of a single embodiment can also be implemented separately or in any suitable sub-combination in multiple embodiments. Further, even if a feature is described above as functioning in a particular combination and was initially claimed as such, one or more features from the claimed combination may in some cases be deleted from the combination, and the claimed combination may be directed to a sub-combination or a variation of a sub-combination.

[0151] Similarly, even if a plurality of operations are shown in the drawings and recited in the claims in a particular order, that should not be construed as requiring that such operations be performed in the particular order or sequence shown, nor should it be construed as requiring that all of the operations shown be performed in order to obtain a desirable result. In certain circumstances, multitasking and parallel processing may be advantageous. Further, the separation of the various system modules and components in the embodiments described above should not be understood as requiring such separation in all embodiments, and the described program components and systems can generally be integrated into a single software product or packaged into multiple software products.

[0152] Particular embodiments of the subject matter have been described. Other embodiments are within the scope of the following claims. For example, even if the actions recited in the claims are performed in a different order, desirable results can still be obtained. As an example, the processes illustrated in the accompanying figures do not necessarily require to be in the particular order shown, or in sequential order, to obtain desirable results. In some cases, multitasking and parallel processing may be advantageous.

Claims

1. A system comprising one or more computers and one or more storage devices storing instructions that, when executed by the one or more computers, cause the one or more computers to implement a contrastive caption neural network, wherein the neural network comprises: A visual encoder neural network configured to process a visual input comprising one or more images to generate an encoded representation of the visual input; A language model neural network configured to process the current text sequence to generate an output that defines a new token to be added to the current text sequence, wherein the current text sequence comprises respective text tokens at each of one or more input positions, and the language model neural network comprises: A set of initial neural network layers configured to process an input comprising each text token of the current text sequence to generate respective unimodal representations of each text token of the current text sequence that are independent of the visual input; A set of subsequent neural network layers configured to process an input comprising the respective unimodal representations of the text tokens of the current text sequence to generate the output that defines the new token to be added to the current text sequence, wherein the subsequent neural network layers comprise one or more cross-modal layers conditioned on the encoded representation of the visual input; a language model neural network; A system comprising the same.

2. The set of initial unimodal neural network layers comprises a sequence of initial attention layers, each initial attention layer being configured to receive as input the respective current representations of each text token of the current text sequence and to process the respective current representations to generate as output the respective updated representations of each text token of the current text sequence. Each of the current representations received as input by the first initial attention layer of the sequence of the initial attention layers is an embedding of each text token of the current text sequence, Each of the current representations received as input by each initial attention layer after the first initial attention layer of the sequence of the initial attention layers is an updated representation of each text token of the current text sequence generated as output by a preceding initial attention layer of the sequence of the initial attention layers, The system according to claim 1.

3. The system according to claim 2, wherein each of the unimodal representations of the text tokens of the current text sequence is an updated representation of each text token of the current text sequence generated as output by the last initial attention layer of the sequence of the initial attention layers.

4. Processing each of the current representations to generate as output an updated representation of each text token of the current text sequence includes applying a causally masked self-attention mechanism, according to the system of claim 2 or claim 3.

5. The output defining the new token to be added to the current text sequence includes a score distribution that assigns a respective score to each text token within the vocabulary of text tokens, according to the system of any one of claims 1 to 4.

6. The set of subsequent neural network layers includes a sequence of subsequent attention layers, Each subsequent attention layer is configured to receive as input each of the current representations of each text token of the current text sequence and to process each of the current representations to generate as output an updated representation of each text token of the current text sequence, Each of the current representations received as input by the first subsequent attention layer of the sequence of the subsequent attention layers is each of the unimodal representations of each text token of the current text sequence, Each current representation received as input by each subsequent attention layer after the first subsequent attention layer of the sequence of the subsequent attention layers is an updated representation of each text token of the current text sequence generated as output by a preceding subsequent attention layer of the sequence of the subsequent attention layers. The system according to any one of claims 1 to 5. **Claim 7** The set of subsequent neural network layers receives one or more of the respective updated representations of the text tokens of the current text sequence generated as output by the last initial subsequent attention layer of the sequence of the subsequent attention layers, and processes the one or more respective updated representations to generate the output that defines the new token to be added to the current text sequence, an output layer block configured to perform The system according to claim 6, further comprising **Claim 8** For one or more of the subsequent neural network layers, processing the respective current representations to generate, as output, an updated representation of each text token of the current text sequence includes applying a causally masked self-attention mechanism. The system according to claim 6 or claim 7. **Claim 9** Each of the one or more cross-modal layers is one of the respective subsequent attention layers of the sequence of the subsequent attention layers, and for each cross-modal layer, processing the respective current representations to generate, as output, an updated representation of each text token of the current text sequence includes applying a cross-attention mechanism between an input derived from the encoded representation of the visual input and the respective current representations of the text tokens of the current text sequence received as input by the cross-modal layer. The system according to any one of claims 6 to 8. **Claim 10** The encoded representation includes respective updated tokens for each of a plurality of patches of the visual input. The system according to any one of claims 1 to 9. **Claim 11** The control caption neural network, a first attention pooling layer that applies attention to the updated tokens and a first set of the learned query tokens to generate respective updated query tokens for each of the learned query tokens, wherein each cross-modal layer further includes the first attention pooling layer that receives the respective updated query tokens as inputs, The system according to claim 10. **Claim 12** The system according to claim 11 when dependent on claim 9, wherein the input derived from the encoded representation is the respective updated query token. **Claim 13** The system according to any one of claims 10 to 12, wherein the visual encoder neural network is a neural network of a vision transformer. **Claim 14** The control caption neural network, further comprising a second attention pooling layer that applies attention to the updated tokens and the second set of the learned query tokens to generate respective updated query tokens for each of the learned query tokens in the second set, The system according to any one of claims 10 to 13. **Claim 15** The system according to claim 14, wherein the second set of the learned query tokens includes only a single learned query token. **Claim 16** A method of training a control caption neural network according to any one of claims 1 to 15, comprising: obtaining a set of one or more training pairs each including a respective visual input and a respective text sequence; for each training pair, processing the respective visual input and the respective text sequence of the training pair using the control caption neural network, including: processing the visual input of the training pair using the visual encoder neural network to generate an encoded representation of the visual input; Processing the text sequence of the training pair using the set of initial neural network layers to generate a unimodal representation for each text token of the text sequence; Processing the respective unimodal representations of the text tokens of the text sequence using the set of subsequent neural network layers to generate a respective score distribution for each of the plurality of text tokens from the respective text sequences; Training the neural network to minimize a loss function, the loss function including (i) a contrastive learning loss term based on the similarity between a contrastive representation derived from the encoded representation of the visual input and one or more unimodal representations of text tokens from each of the text sequences of the training pair, and (ii) for each training pair, a caption loss term based on the respective score distributions for the plurality of text tokens of the respective text sequences; A method comprising the above. **Claim 17** The method according to claim 16, wherein, for each training pair, processing the text sequence of the training pair using the set of initial neural network layers and processing the respective unimodal representations using the set of subsequent neural network layers are performed in a single forward pass through the language model neural network. **Claim 18** The method according to claim 16 or claim 17, wherein the caption loss term measures the quality of the respective score distribution for the token compared to the corresponding token of the text sequence for each training pair and for each of the plurality of tokens from the respective text sequences. **Claim 19** Each training text sequence in each training pair contains the same specified tokens, and the contrastive learning loss term is based on the similarity between the contrastive representation derived from the encoded representation for the visual input of the training pair and the respective unimodal representations for the specified tokens of the training text sequence of the training pair. The method according to any one of claims 16 to 18.

20. The method according to any one of claims 16 to 19 when dependent on claim 15, wherein the contrastive representation for each visual input is the updated query token for the single query token within the second set.

21. The method according to any one of claims 16 to 19 when dependent on claim 10, wherein the contrastive representation for each visual input is generated by pooling the respective updated tokens for each of the plurality of patches of the visual input.

22. The method according to any one of claims 16 to 21, further comprising using one or more of (i) the visual encoder, (ii) an initial set of the neural network layers, or (iii) a subsequent set of the neural network layers after the training to perform a downstream task.

23. The method according to claim 22, further comprising fine-tuning one or more components of the contrastive caption neural network with labeled training data for the downstream task before using one or more of (i) the visual encoder, (ii) an initial set of the neural network layers, or (iii) a subsequent set of the neural network layers after the training and for performing a downstream task.

24. The method according to claim 22, further comprising fine-tuning a downstream neural network including the one or more components of the contrastive caption neural network with labeled training data for the downstream task before using one or more of (i) the visual encoder, (ii) an initial set of the neural network layers, or (iii) a subsequent set of the neural network layers after the training and for performing a downstream task.

25. Fine-tuning the downstream neural network comprises training one or more additional components of the downstream neural network while fixing and holding one or more of (i) the visual encoder, (ii) an initial set of the neural network layers, or (iii) a subsequent set of the neural network layers, the method of claim 24.

26. A method performed by one or more computers, comprising: receiving new input for a downstream task; and processing the new input using a neural network of the downstream task that includes one or more of (i) a visual encoder, (ii) an initial set of neural network layers, or (iii) a subsequent set of neural network layers to generate a task output for the downstream task, wherein the one or more of (i) a visual encoder, (ii) an initial set of neural network layers, or (iii) a subsequent set of neural network layers for generating the task output for the downstream task are trained by performing the respective operations of any one of claims 16-25. Method.

27. One or more computer-readable storage media storing instructions that, when executed by one or more computers, cause the one or more computers to implement the contrast caption neural network of any one of claims 1-15.

28. A method comprising the respective operations performed by the contrast caption neural network of any one of claims 1-15.

29. A system comprising one or more computers and one or more storage devices storing instructions that, when executed by the one or more computers, cause the one or more computers to perform the respective operations of the method of any one of claims 16-26.

30. One or more computer-readable storage media storing instructions that, when executed by one or more computers, cause the one or more computers to perform each operation of the method according to any one of claims 16 to 26.

Citation Information

Patent Citations

  • Spatial attention model for image captioning

    JP2020123372A

  • Image captioning augmented with understanding of the surrounding text

    US20200175063A1

  • Bidirectional attention-based image-text cross-modal retrieval method

    US20210012150A1

  • Systems and methods for contrastive learning of visual representations

    US20210319266A1

  • Contrastive pre-training for language tasks

    WO2021061555A1