Systems and Methods for Visual and Verbal Expression Learning

By employing an intermediate image-text contrast loss and momentum distillation, the VLP systems address the inefficiencies and noise-related issues in conventional VLP frameworks, achieving enhanced performance and efficiency in visual and language tasks.

JP7684005B2Active Publication Date: 2025-05-27SALESFORCE INC
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
JP2023572887
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Priority Date
2021-07-08
Filing Date
2022-01-26
Publication Date
2025-05-27
Estimated Expiration
2042-01-26

AI Technical Summary

Technical Problem

Conventional Visual and Learning Pretraining (VLP) frameworks face limitations such as insufficient modeling efficiency, high annotation and computing overhead, and overfitting to noise, which hinder their performance in visual and language tasks.

Method used

The proposed VLP systems and methods utilize an intermediate image-text contrast (ITC) loss to align image and text features, and momentum distillation (MoD) to generate pseudo-targets for improved learning under noisy supervision, enabling more efficient and effective cross-modal representation learning.

Benefits of technology

The approach enhances the alignment of image and text features, improves the understanding of semantic meanings, and allows for more effective image-text matching, leading to improved performance in downstream visual and language tasks while reducing computational overhead.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007684005000037
    Figure 0007684005000037
  • Figure 0007684005000038
    Figure 0007684005000038
  • Figure 0007684005000039
    Figure 0007684005000039
Patent Text Reader

Abstract

[0005] Embodiments described herein provide a visual and linguistic (V+L) system and method for learning visual and linguistic representations. Specifically, the method may include receiving a training dataset including a plurality of image samples and a plurality of text samples, encoding the plurality of image samples into a plurality of encoded image samples and encoding the plurality of text samples into a plurality of encoded text samples, computing a first loss target based on the plurality of encoded image samples and the plurality of encoded text samples, encoding a first subset of the plurality of encoded image samples and a second subset of the plurality of encoded text samples into a plurality of encoded image-text samples, computing a second loss target based on the plurality of encoded image-text samples, and updating a V+L model based at least in part on the first loss target and the second loss target.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application claims priority to U.S. Non-Provisional Application No. 17 / 370,524, filed Jul. 8, 2021, and U.S. Provisional Application No. 63 / 193,286, filed May 26, 2021, which are hereby expressly incorporated by reference in their entireties.

[0002] The present disclosure generally relates to machine learning models and neural networks, and more specifically, to visual and language representation learning.

Background Art

[0003] Visual and Learning Pretraining (VLP) aims to learn multimodal representations from large-scale image-text pairs that can improve downstream visual and language (V+L) tasks such as image-text retrieval, image-text relationships, visual question answering, or natural language prediction for visual inference. VLP approaches have made some progress with respect to visual and language tasks, but conventional VLP frameworks have suffered from several limitations including insufficient modeling efficiency, high annotation and / or computing overhead, and / or overfitting to noise.

[0004] Accordingly, there is a need for improved VLP systems and methods that avoid the drawbacks associated with conventional VLP methods.

Brief Description of the Drawings

[0005]

Figure 1

[0006]

Figure 2

[0007]

Figure 3

[0008]

Figure 4A

Figure 4B

[0009] In the figure, elements having the same reference numerals have the same or similar functions.

DETAILED DESCRIPTION OF THE INVENTION

[0010] Machine learning methods have been applied to vision and language (V+L) tasks. Such machine learning methods often use vision and language pre-training (VLP) aimed at learning multimodal representations from large-scale image-text pairs. This conventional VLP framework may suffer from several limitations. First, since image features and word token embeddings exist in very different spaces, it can be difficult for a multimodal encoder to learn to model the interaction between features and embeddings. Second, conventional VLP frameworks require bounding box annotations for pre-training and / or high-resolution images, resulting in high annotation and / or computing overhead. Third, the image-text datasets used to train conventional VLP methods may be noisy, resulting in overfitting to the noise and accompanying performance degradation.

[0011] In view of the need for improved VLP systems and methods that avoid the drawbacks associated with conventional VLP methods, the embodiments described herein provide VLP systems and methods such as a method for pre-training a V+L model that utilizes an intermediate image-text contrast (ITC) loss. For example, training inputs are fed into unimodal image and text encoders and transformed into unimodal outputs, and the ITC loss is computed by calculating the loss between the predicted similarity of the unimodal outputs from the image-text pair and the ground truth similarity. The ITC loss is computed based at least in part on the representations output by the unimodal image and text encoders, and these encoders can be updated based at least in part on the ITC loss. In this way, image features and text features are aligned through a training process based on the ITC loss, facilitating the multimodal encoder to perform cross-modal learning. Additionally, the ability of the unimodal encoders to understand the semantic meaning of images and text may be improved through training based on the ITC loss. A common embedding space for images and text may also be learned based on the ITC loss, enabling the purpose of image-text matching to find more beneficial samples.

[0012] In one embodiment, the VLP systems and methods described herein use momentum distillation (MoD) to generate pseudo-targets for capturing visual concepts that may not be described by the ground truth text. MoD utilizes a momentum model to generate pseudo-targets as additional teachers during training and feeds these pseudo-targets to train the image encoder, text encoder, and multimodal encoder, enabling improved learning under noisy supervision and the use of larger uncurated training datasets.

[0013] As used herein, the term "network" may include any hardware or software-based framework that includes any artificial intelligence network or system, neural network or system, and / or any training or learning model implemented in or with it.

[0014] As used herein, the term "module" may include a hardware or software-based framework that executes one or more functions. In some embodiments, the module may be implemented on one or more neural networks. VLP System and Method

[0015] FIG. 1 is a simplified diagram of a computing device for implementing a VLP system for training a vision-and-learning (V+L) model according to some embodiments. As shown in FIG. 1, the computing device 100 includes a processor 110 coupled to a memory 110. The operation of the computing device 100 is controlled by the processor 110. Also, although the computing device 100 is shown as having only one processor 110, it is understood that the processor 110 may represent one or more central processing units, multi-core processors, microprocessors, microcontrollers, digital signal processors, field programmable gate arrays (FPGAs), application specific integrated circuits, graphics processing units (GPUs), etc. within the computing device 100. The computing device 100 may be implemented as a stand-alone subsystem, as a board added to a computing device, and / or as a virtual machine.

[0016] Memory 120 may be used to store software executed by computing device 100 and / or one or more data structures used during the operation of computing device 100. Memory 120 may include one or more types of machine-readable media. Some common forms of machine-readable media include, for example, floppy disks, flexible disks, hard disks, magnetic tapes, any other magnetic media, CD-ROMs, any other optical media, punch cards, paper tapes, any other physical media with patterns of holes, RAM, PROM, EPROM, FLASH-EPROM, any other memory chip or cartridge, and / or any other media adapted to be read by a processor or computer.

[0017] Processor 110 and / or memory 120 may be arranged in any suitable physical configuration. In some embodiments, processor 110 and / or memory 120 may be implemented on the same board, in the same package (e.g., system-in-package), on the same chip (e.g., system-on-chip), etc. In some embodiments, processor 110 and / or memory 120 may include distributed, virtualized, and / or containerized computing resources. Matching such embodiments, processor 110 and / or memory 120 may be located in one or more data centers and / or cloud computing facilities.

[0018] In some examples, memory 120 may include a non-transitory tangible machine-readable medium containing executable code that, when operated on by one or more processors (e.g., processor 110), can cause the one or more processors to execute methods described in more detail herein. For example, as shown, memory 120 may contain instructions for VLP module 130 that can be used to implement and / or emulate systems and models and / or implement any of the methods further described herein. In some examples, VLP module 130 may receive several inputs, such as image input 142 and text input 144, from a user via data interface 115. Data interface 115 may be either a user interface that receives image and text inputs from a user or a communication interface that receives or retrieves image and text inputs from a database. VLP module 130 may generate outputs 150, such as one or more output image-text pairs.

[0019] In some embodiments, VLP module 130 includes an image encoder module 131 and a text encoder module 132. Specifically, the image encoder module is configured to form an encoding of image input 142. The text encoder module is configured to form an encoding of text input 144. In some embodiments, VLP module 130 includes a multimodal encoder 133. The multimodal encoder is configured to receive the encoding of the image input and the encoding of the text input. The multimodal encoder is configured to fuse the encoding of the image input and the encoding of the text input. In some embodiments, VLP module 130 includes a momentum module 134. During training, the momentum module is configured to receive the output from the multimodal encoder and perform momentum distillation (MoD) that generates a pseudo-target of the output, such as an exponentially weighted moving average version of the output.

[0020] Some examples of computing devices, such as computing device 100, may include a non-transitory tangible machine-readable medium containing executable code. Some common forms of machine-readable media are, for example, floppy disks, flexible disks, hard disks, magnetic tapes, any other magnetic media, CD-ROMs, any other optical media, punch cards, paper tapes, any other physical media with patterns of holes, RAM, PROM, EPROM, FLASH-EPROM, any other memory chip or cartridge, and / or any other media adapted to be read by a processor or computer.

[0021] FIG. 2 is a simplified diagram of a process flow for training a V+L model using one or more loss targets, according to some embodiments. As shown in FIG. 2, an image input 210 is passed to a feed-forward image encoder 212 to generate an embedding 214. The input image I is encoded into a sequence of embeddings 214 such as {v cls , v 1 , … v N}, where v cls is the embedding of the [CLS] token. A text input 220 is passed to a feed-forward text encoder 222 to generate an embedding 224. For example, the text encoder transforms the input text T into a sequence of embeddings 224 such as {w cls , w 1 , … w N}.

[0022] The V+L model 200 can include an image encoder 212, a text encoder 222, and a multimodal encoder 240. The image-text contrastive loss 230 can be generated to align the unimodal representations of the image-text pair before fusion in the multimodal encoder 240. The image-text matching loss 242 (using hard negatives 250 mined by contrast similarity) and the masked language modeling loss 244 are applied to learn the multimodal interaction between the image and the text. To improve learning with noisy data, a momentum model 260 (e.g., a moving average version of the base model) can be used as additional supervision during training of the V+L model 200 to generate pseudo targets.

[0023] The image encoder 212 and the text encoder 222 can include one or more feed-forward layers and one or more self-attention layers. The multimodal encoder 240 can include one or more feed-forward layers, one or more cross-attention layers, and one or more self-attention layers. For example, a 12-layer Transformer can be used for the image encoder 212, and a 6-layer Transformer can be used for both the text encoder 222 and the multimodal encoder 240. The text encoder 222 is initialized using the first 6 layers of the BERT model, and the multimodal encoder is initialized using the last 6 layers of the BERT model. The image features can be fused with the text features through cross-attention in each layer of the multimodal encoder 240.

[0024] The encoding 214 from the image encoder 212 and the encoding 224 from the text encoder 222 are used to generate a first loss objective including an image - text contrastive learning (ITC) loss function 230, and can align and compare the encoding 214 from the image encoder 212 and the encoding 224 from the text encoder 222. Image - text contrastive learning (ITC) aims to learn better unimodal representations before the fusion of the encoding 214 from the image encoder 212 and the encoding 224 from the text encoder 222.

[0025] To generate the image - text contrastive learning (ITC) loss for each image and text, the similarity between each image and each text in a plurality of image - text pairs, and the similarity between unpaired images and texts can be generated. For example, a similarity function

Number

[0026] The image - text contrastive learning (ITC) loss can further incorporate the latest M image - text representations of the encoded image samples and encoded text samples from the momentum unimodal encoder 260 into two queues. The normalized features of the encoding 214 and encoding 224 from the momentum unimodal encoder 260 are

Number

Number

Number

Number

Number

[0027] The one-hot similarity of the ground truth may be shown as

Number

Number

[0028] The image-text contrastive learning (ITC) loss function is computed as the average expected sum of the cross-entropy between the computed and softmax-normalized similarity from image to text and the similarity from image to text of the labeled ground truth, and the cross-entropy between the computed and softmax-normalized similarity from text to image and the similarity from text to image of the labeled ground truth.

[0029] For example, the image-text contrast (ITC) learning loss can be defined as the cross-entropy H between the predicted similarity p between the encoded image sample and the encoded text sample, and the ground truth one-hot similarity y. For example,

Number

[0030] In one embodiment, the encodings from the image encoder 212 and the text encoder 222 are further passed to a feed-forward multimodal encoder 240 to generate an encoded image-text sample. The multimodal encoder 240 is configured to generate a second loss objective that includes an image-text matching (ITM) loss 242 and a masked language modeling (MLM) loss 244. The ITM loss 242 is computed based on the expected cross-entropy between the predicted image-text matching of the encoded image-text sample and the corresponding ground truth image-text matching of the encoded image-text sample. The ITM loss 242 can be generated using hard negatives 250 mined through the contrast similarity of the encoded image-text sample.

[0031] The image-text matching (ITM) 242 predicts the two-class likelihood of the encoded image-text sample, e.g., whether the pair of image and text in the encoded image-text sample is positive (matching) or negative (not matching). The output embedding of the [CLS] token of the multimodal encoder 240 can be used as the combined representation of the pair of image and text within the encoded image-text sample, a fully connected (FC) layer is added, and then, by the softmax function, the two-class likelihood p of the image-text pair itmIt is possible to predict (i.e., whether the image-text pair is positive or negative). The ITM loss can be the cross-entropy H between the predicted two-class probabilities of the image-text pair and the one-hot two-class probabilities of the ground truth, for example, [Number] where, y itm is a two-dimensional one-hot vector representing the ground truth label.

[0032] The multimodal encoder 240 is also configured to generate a masked language modeling (MLM) loss 244 to learn the multimodal interaction between the image input 210 and the text input 220. The MLM loss 244 can be defined as a loss function between the predicted probabilities of one or more masked tokens in the encoded image-text sample and the ground truth identities of one or more masked tokens of the encoded image-text sample.

[0033] Masked language modeling (MLM) uses both the image and the context text from the encoded image-text sample to predict the masked words in the encoded image-text sample. The input tokens are randomly masked with a predetermined probability such as 15% and replaced with a special token [MASK]. For example, the replacement is 10% random tokens, 10% no change, and 80% [MASK].

[0034] The MLM learning loss 244 can be the cross-entropy H between the predicted probabilities for the masked tokens within the encoded image-text sample and the one-hot vocabulary distribution of the ground truth, for example, [Number] where, [Number] can be used to show masked text,

Number

[0035] A subset of the encoded images and text samples can be selected based at least in part on negative mining before being encoded into the image-text samples encoded by the multimodal encoder. Hard negatives can be sampled for the ITM task with zero computing overhead. Negative image-text pairs are hard if they share similar semantics and differ in fine-grained details. The contrast similarity from Equation (1) can be used to find hard negatives. For each image in the mini-batch, one negative text can be sampled from the same batch according to the contrast similarity distribution, and texts more similar to the image are more likely to be sampled. Similarly, one hard negative image can be sampled for each text.

[0036] In some embodiments, the vision-language learning (V+L) model is updated based on a combination of a first loss objective and a second objective, e.g., a combination of a first loss objective and a second loss objective such as an ITC loss, an MLM loss, and an ITM loss. For example,

Number

[0037] In one embodiment, the final loss objective may be a weighted sum of the ITC loss, the MLM loss, and the ITM loss, and the weighting coefficients are empirical or predefined.

[0038] In one embodiment, to improve learning when there is noisy input data for training the model, for example, pseudo-targets are generated using momentum distillation (MoD) as an alternative to the original noisy data for training the model. For all of the encoders (e.g., the image encoder 212, the text encoder 222, and the multimodal encoder 240), the pseudo-targets are generated by the momentum model 260. The momentum model is a continuously evolving teacher model and includes an exponentially weighted moving average version of all encoders including unimodal and multimodal encoders.

[0039] During training, the vision and language-based models can be trained such that their predictions match the predictions from the momentum model. Specifically, to modify the ITC, the image-text similarity can be adjusted with the pseudo-targets generated by the momentum model, e.g.,

Number

Number

Number

[0040] Similarly, to modify the MLM, the predicted probability of the momentum model for the masked tokens is, e.g.,

Number

[0041] In some embodiments, the vision and language learning (V+L) model is at least partially updated with a combination of a first loss target and a second target, e.g., a first loss target and a second loss target modified by a pseudo-target generated by a momentum model.

[0042] FIG. 3 is a simplified logic flow diagram illustrating a method 300 for visual and language representation learning implementing sub-modules 131-134 of FIG. 1 according to some embodiments. One or more of processes 310-360 of method 300 may be implemented in the form of executable code stored in a non-transitory tangible machine-readable medium that can cause one or more processors to execute one or more of processes 310-360 when executed by the one or more processors. In some embodiments, method 300 may correspond to the method used by module 130.

[0043] In process 310, a training data set including a plurality of image samples and a plurality of text samples may be received, e.g., via data interface 115 of FIG. 1. In some embodiments, at least one image sample of the plurality of image samples corresponds to at least one text sample of the plurality of text samples.

[0044] In process 320, the image encoder may encode a plurality of image samples into a plurality of encoded image samples. In process 320, the text encoder may encode a plurality of text samples into a plurality of encoded text samples. The encoding by the image encoder or the text encoder may be performed simultaneously or at different times. For example, the encoding by the image encoder may be performed before the encoding by the text encoder. For example, the encoding by the image encoder may be performed after the encoding by the text encoder. In some embodiments, the image encoder is a transformer. In further embodiments, the text encoder is a transformer.

[0045] In process 330, a first loss target may be computed based on a plurality of encoded image samples and a plurality of encoded text samples. The first loss target may include an image-text contrastive (ITC) loss target that refers to a loss function between a predicted similarity between an encoded image sample and an encoded text sample and a corresponding ground truth similarity.

[0046] In additional and alternative embodiments, method 300 or process 330 may further include using momentum distillation (MoD) to form a momentum model, using the momentum model to generate a plurality of modeled image samples and a plurality of modeled text samples, including the plurality of modeled image samples in the plurality of image samples, including the plurality of modeled text samples in the plurality of text samples, and using the modeled image samples and the modeled image samples to generate a first target such as an ITC loss target.

[0047] In process 340, the multimodal encoder may encode a first subset of a plurality of encoded image samples and a second subset of a plurality of encoded text samples into a plurality of encoded image-text samples. In some embodiments, the multimodal encoder is a transformer. The first subset and the second subset may be selected based at least in part on negative mining or mining of negative image-text pairs that share similar semantics but differ in fine-grained details. The negative image-text pairs can be selected based at least on the contrast similarity distribution from equation (1).

[0048] In process 350, a second loss target is computed based on a plurality of encoded image-text samples and includes an image-text matching (ITM) loss target and a masked language modeling (MLM) loss target. The ITM loss can be a loss function between the predicted image-text matching of the encoded image-text sample and the corresponding ground truth image-text matching of the encoded image-text sample. The MLM loss can be a loss function between what is predicted for the masked tokens in the encoded image-text sample and the ground truth vocabulary distribution of the encoded image-text sample.

[0049] In additional alternative embodiments, method 300 or process 350 may further include generating a second target, such as an MLM loss target, using modeled image samples and modeled image samples from a momentum model.

[0050] In process 360, the V+L model may be updated at least partially based on a first loss target and a second loss target. For example, updating the V+L model may include updating an image encoder, a text encoder, and a multimodal encoder based on a combination of the first loss target and the second loss target. In another example, the step of updating the V+L model includes updating the image encoder and the text encoder at least partially based on the first loss target, and updating the multimodal encoder at least partially based on the second loss target.

[0051] In a further embodiment, method 300 may further include fine-tuning the V+L model for a task selected from the group consisting of an image-text retrieval task, an image-to-text retrieval (TR) task, a text-to-image retrieval (IR) task, a visual entailment (VE) task, a visual question answering (VQA) task, and a natural language for visual reasoning for real (NLVR) task.

[0052] In one embodiment, the lower bound of the mutual information (MI) between different "views" of an image-text pair can be maximized.

[0053] Formally, given two random variables a and b, the mutual information (MI) measures their dependence and is

Equation

[0054] To maximize the lower bound of the mutual information, a self-supervised learning method known as InfoNCE has been proposed. That is,

Equation

Number

Number

Number

[0055]

Number

[0056] MLM can be interpreted as maximizing the MI between the masked word token and its masked context (i.e., image + masked text). Specifically, an alternative version of the MLM loss using one-hot labels (a variation of eqn(3)) is

Number

[0057] where

Number

Number

Number

[0058] Both ITC and MLM generate views by taking partial information from the image-text pair. Momentum distillation improves ITC and MLM and generates different views from the proposed distribution as a whole. In ITC, alternative views of the image-text pair can be generated by finding semantically similar images and texts within the training dataset. In MLM, alternative views of the masked word can be generated from the entire vocabulary set. Therefore, MoD can be considered as performing data augmentation on the original view. MoD generates a set of diverse views that do not exist in the original image-text pair, which can improve the generalization performance of the model.

[0059] Exemplary System Architecture and Performance Exemplary experiments were conducted to evaluate the performance of a VLP system (e.g., a pre-trained vision and learning model or V+L model) in downstream tasks. In some embodiments, the pre-trained V+L model is fine-tuned and can be applied to one or more downstream tasks including image-text retrieval, visual entailment, visual question answering, and natural language for realistic visual reasoning.

[0060] The V+L model consists of BERT with 123.7M parameters and ViT-B / 16 with 85.8M parameters. This model was pre-trained for 30 epochs using 8 NVIDIA A100 GPUs with a batch size of 512. The AdamW optimizer was used with a weight decay of 0.02. Further details of the AdamW optimizer are provided in oshchilov, Decoupled Weight Decay Regularization, arXiv preprint arXiv:1711.05101, 2017, which is incorporated herein by reference in its entirety. The learning rate was warmed up to 1e -4 until the first 1,000 iterations and decayed to 1e -5 according to the cosine schedule.

[0061] For example, the pre-training data was generated using two web datasets (Conceptual Captions and SBU Captions) and two in-domain datasets (COCO and Visual Genome). The total number of unique images is 4.0M and the number of image-text pairs is 5.1M. To show that the V+L model is scalable with large-scale web data, the noisier Conceptual 12M dataset can also be included, increasing the total number of images to 14.1M2.

[0062] During pre-training, random image crops of resolution 256×256 were taken as input and RandAugment was also applied. Further details of RandAugment are provided in Cubuk et al., RandAugment: Practical automated data augmentation with a reduced search space, CVPR Workshops, pages 702-03, 2020, which is incorporated herein by reference in its entirety. Since the text often contains pornographic information, color changes were removed from RandAugment.

[0063] During fine-tuning, the image resolution was increased to 384×384, and the positional encoding of the image patches was interpolated. The momentum parameter for updating the momentum model was set to 0.995, and the size of the queue used for image-text contrastive learning was set to 65536. The distillation weight α was ramped up linearly within the first epoch.

[0064] Image-text retrieval involves two subtasks, namely, text retrieval from images (TR) and image retrieval from text (IR). The V+L model was evaluated on the benchmarks of Flickr30K and COCO after being fine-tuned using the training samples from each of the Flickr30K and COCO datasets. For zero-shot retrieval on Flickr30K, the V+L model fine-tuned on COCO was evaluated.

[0065] During fine-tuning, both the ITC loss (Equation (2)) and the ITM loss (Equation (4)) were optimized. ITC learns an image-text scoring function based on the similarity of unimodal features, while ITM models the fine-grained interaction between images and texts to predict a matching score. Since the downstream dataset contains multiple texts for each image, the ground-truth labels for ITC were changed to consider multiple positives in the queue, and each positive has a ground-truth probability of 1 / #positives.

[0066] During inference, the feature similarity score s itc was first computed for all image-text pairs. Then, the top-k candidates were selected and used to compute the ITM score s itm for ranking. The inference speed of the V+L model is much faster than the method that requires computing the ITM score for all image-text pairs.

[0067] Visual Entailment (SNLI-VE) is a fine-grained visual reasoning task for predicting whether the relationship between an image and text is entailment, neutral, or contradiction. Visual Entailment can be considered a three-way classification problem. Class probabilities can be predicted using a multi-layer perceptron (MLP) on the multimodal encoder representation of the [CLS] token.

[0068] Visual Question Answering (``VQA'') requires that an image and a question be given and that the model predict an answer. Different from existing research that formulates VQA as a multiple-choice classification problem, VQA can be formulated as an answer generation problem. Specifically, a six-layer Transformer decoder can be used to generate an answer.

[0069] Figures 4A-4B are simplified diagrams of model architectures for using a VLP system according to some embodiments described herein. As shown in Figure 4A, substantially the same model as in Figure 2 is used for visual question answering, except that an autoregressive decoder 450 is added to generate an answer when an image-question embedding is given. The image encoder 420 encodes the image input 410 into an image embedding, and the text encoder 422 encodes the question input 412 into a question embedding. The image embedding is passed to the multimodal encoder 430 via the cross-attention input 440 to generate a multimodal image-question embedding using the question embedding from the text encoder 422. The autoregressive answer decoder 450 receives the multimodal image-question embedding via the cross-attention input 440, and the sequence start token ([CLS]) 460 is used as the initial input token of the decoder. Similarly, a sequence end token ([SEP]) is added at the end of the decoder output to indicate the completion of generation. The answer decoder 450 is initialized using the pre-trained weights from the multimodal encoder 430 and fine-tuned with a language modeling loss. For a fair comparison with existing methods, the answer decoder 450 is constrained to generate only from 3,192 candidate answers during inference.

[0070] As shown in FIG. 4B, for realistic visual reasoning, a natural language uses a model to predict whether the text accurately describes a pair of images. A natural extension can be made to a multimodal encoder 470 that enables reasoning for two images 490 and 492. The two images 490 and 492 can be fed to two image encoders 494 and 496 that share all parameters to generate embeddings and then fed to the multimodal encoder 470. The text input 475 can also be fed to a text encoder 485 to generate an embedding that enters the multimodal encoder 470. Each layer of the multimodal encoder 470 is replicated to have two consecutive transformer blocks 480, and each block includes a self-attention layer, a cross-attention layer, and a feed-forward layer (see FIG. 2). The multimodal blocks 480 can also share a cross-attention layer. The two multimodal blocks 480 within each layer are initialized using the same pre-trained weights, and the two cross-attention layers share the same linear projection weights for keys and values.

[0071] During training, the two multimodal blocks 480 receive two different sets of image embeddings for the image pair 490 and 492. The MLP classifier can learn with the multimodal encoder representation of the [CLS] token to predict "true" or "false".

[0072] To prepare a new multimodal encoder for image pair input, additional pre-training steps can be performed. The text assignment (TA) task can be designed such that when given an image and text pair, the model needs to assign the text to either the first image, the second image, or not assign it to either. This can be considered a three-way classification problem, and the FC layer can be used to predict the assignment class on the [CLS] representation. This model was pre-trained with text alignment (TA) for only 1 epoch using 4M images.

[0073] The V+L model was evaluated in downstream tasks (including image-text contrastive learning, contrastive hard negative mining, and momentum distillation) as shown in Table 1. Table 1 shows the performance of downstream tasks using various variations of the V+L model. Compared to the baseline pre-training tasks (MLM+ITM), adding ITC significantly improved the performance of the pre-trained model across all tasks. The proposed hard negative mining improved ITM by finding more beneficial training samples. Furthermore, adding momentum distillation improved both the learning of ITC, MLM, and all downstream tasks (image-to-text retrieval (or TR), text-to-image retrieval (or IR), visual entailment (or VE), visual question answering (or VQA), and natural language for visual reasoning for reality (or NLVR)). The V+L model can effectively utilize more noisy web data to improve the performance of pre-training, such as pre-training on 14M pre-trained images.

Table 1

[0074] In Table 1, the averages of R@1, R@5, and R@10 were reported for text retrieval (TR) and image retrieval (IR). Also, in Table 1, ITC refers to image-text contrastive learning, MLM refers to masked language modeling, and ITMhard refers to image-text matching using contrastive hard negative mining.

[0075] MoD: Momentum distillation Tables 2 and 3 report the results of fine-tuning and zero-shot image-text retrieval, respectively. The V+L model achieves state-of-the-art performance and outperforms other methods trained on significantly larger datasets. Considering the substantial improvement of the V+L model when the number of training images increased from 4M to 14M, the V+L model can be trained on larger web image-text pairs.

Table 2

Table 3

[0076] Table 4 reports a comparison with existing methods for other V+L understanding tasks. With 4M pre-trained images, the V+L model achieved state-of-the-art performance. With 14M pre-trained images, the V+L model was significantly superior to existing methods including those that require additional object tags or adversarial data augmentation. Compared with VILLA, the V+L model achieved absolute improvements of 2.47% on the VQA test-std, 3.84% on the NLVR2 test-P, and 1.88% on the SNLI-VE test. The V+L model does not require a detector and requires low-resolution images, and thus also enjoys a much faster inference speed (more than 10 times faster than UNITER or VILLA) compared to existing methods.

Table 4

[0077] Visual grounding aims to identify regions within an image that correspond to a specific text description. The V+L model has been shown to achieve visual grounding by seeking its attention without being trained on bounding box annotations. The experiments were conducted on the widely used RefCOCO+ dataset. The pre-trained model was fine-tuned on the training set of RefCOCO+ using only image-text monitoring. A similar fine-tuning strategy was followed for image-text retrieval. Table 5 reports these results.

Table 5

[0078] In Table 6, the impact of various design choices on image-text retrieval was studied. The contrast similarity score s itcSince it is used to filter the top-k candidates during inference, k can be varied to report its effect. Generally, s itm The final ranking results obtained by s itc are not sensitive to changes in k. The reason is that by using only s

Table 6

[0079] In Table 7, the effects of text assignment (TA) pre-training and parameter sharing were studied for NLVR2. Three sharing strategies were considered. That is, (1) all parameters are shared among two consecutive multimodal blocks, (2) only the cross-attention (CA) layers are shared, and (3) nothing is shared. Without TA, sharing the entire block improves performance. By pre-training the model on image pairs input with TA, sharing the cross-attention layers brings the best performance.

Table 7

[0080] This description and the accompanying drawings, which illustrate embodiments, implementations, or uses of the invention, should not be construed as limiting. Various mechanical, compositional, structural, electrical, and operational changes may be made without departing from the spirit and scope of this description and the claims. In some instances, well-known circuits, structures, or techniques have not been shown or described in detail so as not to obscure the embodiments of the present disclosure. Similar numerals in two or more figures represent the same or similar elements.

[0081] In this description, specific details are set forth that describe some embodiments that do not conflict with the present disclosure. To provide a complete understanding of the embodiments, numerous details are set forth. It will be apparent to those skilled in the art that some embodiments may be practiced without some or all of these specific details. The specific embodiments disclosed herein are illustrative but not limiting. Those skilled in the art may recognize other elements that are within the scope and spirit of the present disclosure but are not specifically described herein. Additionally, to avoid unnecessary repetition, one or more features shown and described in connection with one embodiment may be incorporated into other embodiments unless specifically described otherwise in connection with those other embodiments or the one or more features render one embodiment non-functional.

[0082] Exemplary embodiments have been shown and described, but a wide range of modifications, changes, and substitutions are contemplated in the foregoing disclosure, and in some instances, some features of the embodiments may be employed without the corresponding use of other features. Those skilled in the art will recognize many variations, alternatives, and modifications. Accordingly, the scope of the invention should be limited only by the following claims, and it is appropriate that the claims be broadly construed in a manner consistent with the scope of the embodiments disclosed herein.

Claims

A method for training a visual and language learning (V+L) model including an image encoder, a text encoder, and a multimodal encoder, which is executed by a computer, comprising: Receiving, via a data interface, a training data set including a plurality of image samples and a plurality of text samples, wherein at least one image sample among the plurality of image samples corresponds to at least one text sample among the plurality of text samples; Encoding the plurality of image samples into a plurality of encoded image samples by an image encoder, and encoding the plurality of text samples into a plurality of encoded text samples by a text encoder; Computing a first loss target based on the plurality of encoded image samples and the plurality of encoded text samples; Encoding a first subset of the plurality of encoded image samples and a second subset of the plurality of encoded text samples into a plurality of encoded image-text samples by a multimodal encoder; Computing a second loss target based on the plurality of encoded image-text samples; Updating the V+L model at least partially based on the first loss target and the second loss target, The method further includes selecting the first subset and the second subset at least partially based on mining negative image-text pairs through the contrast similarity of the encoded image-text samples. Claim 2 The first loss target includes an image-text contrast (ITC) loss target that is the average expected sum of the cross-entropy between the similarity from the computed and softmax-normalized image to text and the similarity from the labeled ground truth image to text, and the cross-entropy between the similarity from the computed and softmax-normalized text to image and the similarity from the labeled ground truth text to image. The method according to claim 1. Claim 3 The second loss target includes an image-text matching (ITM) loss target computed as the cross-entropy between the predicted two-class probabilities of the image-text pair and the one-hot two-class probabilities of the ground truth, and an MLM loss target computed as the cross-entropy between the predicted probabilities of one or more masked tokens in the encoded image-text sample and the identities of the ground truth of the one or more masked tokens in the encoded image-text sample. The method according to claim 1.

4. Updating the V+L model comprises updating the image encoder and the text encoder at least partially based on the first loss target; and updating the multimodal encoder at least partially based on the second loss target. The method according to claim 1.

5. forming a momentum model using momentum distillation (MoD); using the momentum model to generate a plurality of modeled image samples and a plurality of modeled text samples; including the plurality of modeled image samples in the plurality of image samples; and including the plurality of modeled text samples in the plurality of text samples. The method according to claim 1 further comprises.

6. The image encoder, the text encoder, and the multimodal encoder each include a transformer. The method according to claim 1.

7. further fine-tuning the V+L model for a task selected from the group consisting of an image-text retrieval task, an image-to-text retrieval (TR) task, a text-to-image retrieval (IR) task, a visual entailment (VE) task, a visual question answering (VQA) task, and a natural language for visual reasoning for reality (NLVR) task. The method according to claim 1.

8. A system for training a V+L model, comprising a non-transitory memory; and one or more processors coupled to the non-transitory memory and configured to read instructions from the non-transitory memory to cause the system to perform operations, the operations comprising Receiving, via a data interface, a training data set including a plurality of image samples and a plurality of text samples, wherein at least one image sample among the plurality of image samples corresponds to at least one text sample among the plurality of text samples; Encoding, by an image encoder, the plurality of image samples into a plurality of encoded image samples, and encoding, by a text encoder, the plurality of text samples into a plurality of encoded text samples; Computing a first loss target based on the plurality of encoded image samples and the plurality of encoded text samples; Encoding, by a multimodal encoder, a first subset of the plurality of encoded image samples and a second subset of the plurality of encoded text samples into a plurality of encoded image-text samples; Computing a second loss target based on the plurality of encoded image-text samples; Updating the V+L model of the image encoder, the text encoder, and the multimodal encoder, at least partially based on the first loss target and the second loss target; The system further includes selecting the first subset and the second subset, at least partially based on mining negative image-text pairs through a contrast similarity of the encoded image-text samples. Claim 9 Updating the V+L model includes updating the image encoder and the text encoder, at least partially based on the first loss target, and updating the multimodal encoder, at least partially based on the second loss target. The system according to claim 8. Claim 10 The operations include Using momentum distillation (MoD) to form a momentum model; Generating, using the momentum model, a plurality of modeled image samples and a plurality of modeled text samples; Including the plurality of modeled image samples in the plurality of image samples; The system according to claim 8, further comprising including the plurality of modeled text samples in the plurality of text samples.

11. The system according to claim 8, wherein the image encoder, the text encoder, and the multimodal encoder each include a transformer.

12. The operation further includes fine-tuning the V+L model for a task selected from the group consisting of an image-text extraction task, a text extraction from image (TR) task, an image extraction from text (IR) task, a visual entailment (VE) task, a visual question answering (VQA) task, and a natural language for visual inference for reality (NLVR) task, the system according to claim 8.

13. A non-transitory machine-readable medium storing machine-readable instructions executable to cause a system to perform an operation, the operation comprising: Receiving, via a data interface, a training data set including a plurality of image samples and a plurality of text samples, wherein at least one image sample of the plurality of image samples corresponds to at least one text sample of the plurality of text samples; Encoding, by an image encoder, the plurality of image samples into a plurality of encoded image samples, and encoding, by a text encoder, the plurality of text samples into a plurality of encoded text samples; Computing a first loss target based on the plurality of encoded image samples and the plurality of encoded text samples; Encoding, by a multimodal encoder, a first subset of the plurality of encoded image samples and a second subset of the plurality of encoded text samples into a plurality of encoded image-text samples; Computing a second loss target based on the plurality of encoded image-text samples; Updating the image encoder, the text encoder, and the multimodal encoder based at least in part on the first loss target and the second loss target. The non-transitory machine-readable medium, wherein the operation further includes selecting the first subset and the second subset based at least in part on mining negative image-text pairs through a contrast similarity of the encoded image-text samples.

Citation Information

Patent Citations

  • Multi-modal push-text named entity identification method based on text-picture relationship pre-training

    CN112257445A

  • System and method for learning interactive language

    JP2019185748A

  • System and method for a dialogue response generation system

    WO2021049199A1