Video-text modeling with zero sample migration from comparative commentary

By freezing or fine-tuning the attention pooling layer of the pretrained image-text processing model, it is directly applied to the video comprehension task, solving the problem of high computing resource consumption in the prior art and achieving efficient video comprehension capabilities.

CN120283271APending Publication Date: 2025-07-08GOOGLE LLC
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202380084725.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2022-12-08
Filing Date
2023-12-08
Publication Date
2025-07-08

AI Technical Summary

Technical Problem

Existing video understanding models require a large amount of computing resources when training and adapting to new tasks, especially when the model involves a large number of parameters, and additional training of additional parameters when facing new types of video data, resulting in high computational costs.

Method used

Using pre-trained image-text processing models, especially CoCa models, is applied directly to video comprehension tasks by freezing or fine-tuning their attention pooling layer, reducing or eliminating the need for additional training, and adapting it with contrast and generative loss functions.

Benefits of technology

Significantly reduces the computing resources required for video comprehension tasks, enables zero-sample or low-sample video comprehension capabilities, reduces training costs and maintains high performance.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120283271A_ABST
    Figure CN120283271A_ABST
Patent Text Reader

Abstract

An efficient approach is provided for establishing a base video-text model for tasks including open vocabulary video classification, text-to-video retrieval, video commentary addition, and video questions and answers. Some example implementations include a model that may be referred to as Video CoCa. Example implementations reuse a pre-trained image-text comparative commentary adder (CoCa) model and adapt it to video-text tasks with little or minimal additional training. While previous work employs an image-text model with various cross-frame fusion modules (e.g., a cross-frame attention layer or perceptron resampler), and fine tuning of the modified architecture on the video-text data, the image-text model may be used as an image-text model. However, aspects of the present disclosure utilize the discovery that generative and comparative attention pooling layers in an image-text CoCa design may be immediately adapted to "flattened frame embedding", resulting in strong zero sample migration baselines for many video-text tasks.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] Related Applications

[0002] This application claims the priority and benefit of U.S. Provisional Patent Application No. 63 / 431,224, filed on Dec. 8, 2022. U.S. Provisional Patent Application No. 63 / 431,224 is hereby incorporated by reference in its entirety. Technical Field

[0003] The present disclosure generally relates to machine learning models. More specifically, the present disclosure relates to applying a pre-trained image-text processing model to video understanding tasks. Background Art

[0004] In recent years, due to the development of innovative computing models, the field of video understanding, including tasks such as video classification, video question answering, video retrieval, and video captioning, has made significant progress.

[0005] However, a significant challenge in this field is the computational resources required for the initial training of the model and the subsequent task-specific fine-tuning of the model. Each time the model is trained or fine-tuned, a large amount of computational resources are consumed. This is especially the case when the model involves a large number of parameters.

[0006] Additionally, when the model is applied to new types of tasks or data, such as video data, additional parameters are typically added to the model, and then these additional parameters are trained from scratch. This further increases the computational resources required to train the model. Summary of the Invention

[0007] Aspects and advantages of embodiments of the present disclosure will be set forth in part in the following description, or may be learned from the description, or may be learned through practice of the embodiments.

[0008] An example aspect of the present disclosure relates to a computer-implemented method for performing video understanding tasks with improved computational efficiency. The method includes a computing system including one or more computing devices accessing a pre-trained image-text processing model, where the pre-trained image-text processing model includes one or more pre-trained attention pooling layers having a plurality of parameters, and where the pre-trained image-text processing model has been pre-trained on a joint contrastive and generative image captioning loss function. The method includes the computing system obtaining an input video including a plurality of image frames. The method includes the computing system using the pre-trained image-text processing model having one or more pre-trained attention pooling layers with the same number of parameters to process the input video to generate a prediction for the video understanding task as an output of the pre-trained image-text processing model. The method includes the computing system providing the prediction for the video understanding task as an output.

[0009] Another example aspect of the present disclosure relates to one or more non-transitory computer-readable media that collectively store: a pre-trained image-text processing model, wherein the pre-trained image-text processing model includes one or more pre-trained attention pooling layers having a plurality of parameters, and wherein the pre-trained image-text processing model has been pre-trained on a joint contrastive and generative image captioning loss function; and computer-executable instructions for performing an operation that includes processing an input video including a plurality of image frames by a computing system using the pre-trained image-text processing model having one or more pre-trained attention pooling layers with the same number of parameters to generate a prediction for a video understanding task as an output of the pre-trained image-text processing model.

[0010] Other aspects of the present disclosure relate to various systems, devices, non-transitory computer-readable media, user interfaces, and electronic devices.

[0011] These and other features, aspects, and advantages of the various embodiments of the present disclosure will be better understood with reference to the following description and the appended claims. The accompanying drawings, which are incorporated in and constitute a part of this specification, illustrate example embodiments of the present disclosure and, together with the description, serve to explain the relevant principles. BRIEF DESCRIPTION OF THE DRAWINGS

[0012] Reference is made to the accompanying drawings, in which a detailed discussion of embodiments is presented to those of ordinary skill in the art, in which:

[0013] Figure 1 A graphical diagram of an example video-text model in accordance with an example embodiment of the present disclosure is depicted.

[0014] Figure 2 A graphical diagram of an example attention pooler and flattened frame token embeddings in accordance with an example embodiment of the present disclosure is depicted.

[0015] Figure 3 A flowchart of an example method for performing a video understanding task using a pre-trained image-text processing model in accordance with an example embodiment of the present disclosure is depicted.

[0016] Figure 4A A block diagram of an example computing system in accordance with an example embodiment of the present disclosure is depicted.

[0017] Figure 4B A block diagram of an example computing device in accordance with an example embodiment of the present disclosure is depicted.

[0018] Figure 4C A block diagram of an example computing device in accordance with an example embodiment of the present disclosure is depicted.

[0019] Reference numerals repeated across multiple figures are intended to identify like features in various implementations. DETAILED DESCRIPTION

[0020] Aspects of the present disclosure provide efficient methods for building a foundation video-text model for various video understanding tasks such as open-vocabulary video classification, text-to-video retrieval, video captioning, and video question answering. In particular, some example implementations can reuse a pre-trained image-text model, such as the Contrastive Captioner (CoCa) model, and adapt it to video-text tasks with zero or minimal additional training. An example implementation of the proposed method of repurposing the Contrastive Captioner (CoCa) model for video tasks may be referred to as VideoCoCa.

[0021] In particular, some prior work has sought to adapt image-text models by modifying the image-text model to include various cross-frame fusion modules (e.g., cross-frame attention layers or perceptual resamplers) or other novel layer or architectural aspects. After modifying the model architecture, these prior works then train the newly added parameters on video-text data. Retraining a new set of parameters in this way increases the computational resources required to adapt the model for video tasks.

[0022] In contrast, example implementations of the present disclosure instead directly adapt a pre-trained image-text model to video tasks. Specifically, as an example, the frozen image encoder of the pre-trained image-text CoCa can be used to separately process each video frame of an input video in order to generate per-frame token embeddings for each video frame. Then, some example implementations can flatten the token embeddings into a long sequence of the frozen video representation and apply the generative and contrastive attention pooling layers of CoCa to this representation to generate predictions for video tasks.

[0023] In some example implementations, all model weights including the attention pooling layer can be directly loaded from the pre-trained image-text CoCa model while still achieving state-of-the-art performance on video understanding tasks. Additional example implementations can perform various forms of lightweight fine-tuning on top of the image-text model to provide further performance gains.

[0024] More specifically, an example aspect of the present disclosure relates to systems and methods for performing video understanding tasks with improved computational efficiency. The example methods utilize a pre-trained image-text processing model that has been pre-trained on a joint contrastive and generative image captioning loss function. An example of such a model is the Contrastive Captioner (CoCa) model described in Yu et al., CoCa: Contrastive Captioners are Image-Text Foundation Models, arXiv:2205.01917. The image-text processing model can include one or more pre-trained attention pooling layers with multiple parameters. The proposed method can include using the pre-trained model to process an input video including multiple image frames to generate predictions for the video understanding task.

[0025] Specifically, the proposed techniques can utilize a pre-trained image-text processing model that combines contrastive pre-training methods with generative pre-training methods. The image-text processing model can be designed to facilitate image, text, and image-text representation learning. The model can include a cascaded decoder design, where the lower half unimodal decoder encodes text context using self-attention with a causal mask, and the upper half multimodal decoder uses cross-attention to align images and text. The model can be trained using a joint contrastive loss and a captioning loss.

[0026] Thus, in some implementations, the pre-trained image-text processing model includes a pre-trained unimodal image encoder. The encoder processes the input image to generate one or more frame embeddings. The pre-trained attention pooling layers of the image-text processing model then process these frame embeddings to generate one or more contrastive embeddings and one or more generative embeddings. When using the model to process an input video, the method can include separately processing each of the image frames in the input video using the pre-trained unimodal image encoder to generate multiple frame embeddings.

[0027] According to one aspect of the present disclosure, in some implementations, the frame embeddings generated from processing the image frames of the input video are combined to form a set of combined frame embeddings. This can be achieved in several ways. For example, the frame embeddings can be concatenated along the temporal dimension to generate a set of flattened frame embeddings. Alternatively, the frame embeddings can be reshaped into a joint spatio-temporal representation. The combined frame embeddings are then processed by the attention layers of the model to generate one or more generative embeddings and one or more contrastive embeddings.

[0028] In some implementations, after pre-training an image-text processing model, the parameters of the pre-trained attention pooling layer of the pre-trained image-text processing model can be kept fixed. This allows the model to be directly applied to video understanding tasks, including zero-shot video understanding tasks, without further training.

[0029] However, other implementations of the present disclosure also provide for further fine-tuning of various parts of the image-text model, such as the parameters of the pre-trained attention pooling layer of the pre-trained image-text processing model. This fine-tuning can be performed using a joint contrastive and generative image captioning addition loss function applied to video data. The fine-tuning can involve unfreezing all the parameters of the pre-trained image-text processing model, or the fine-tuning can involve freezing the parameters of the encoder and decoder and only tuning the parameters of the generative and contrastive poolers.

[0030] Some implementations of the present disclosure also allow an additional encoder model to be added to the attention pooling layer before processing the input video. This additional encoder model can be a Transformer encoder that models the interaction between tokens from different frames. Then, the output tokens from this encoder can be used for the captioning addition loss, and their globally averaged pooled embeddings in the temporal dimension can be used for the contrastive loss.

[0031] The proposed method can be applied to various video understanding tasks. For example, it can be used for video classification tasks, where the goal is to classify an input video into one of several predefined categories. The method can also be used for video question answering tasks, where the goal is to generate answers to questions about the content of the input video. Additionally, the method can be used for video captioning addition tasks, where the goal is to generate a text description of the content of the input video.

[0032] The systems and methods of the present disclosure provide several technical effects and benefits. As an example, the proposed technique solves the technical problem of computational resource consumption in video understanding tasks. In particular, the proposed technique that utilizes the pre-trained image-text processing model significantly reduces the computational resources required for such tasks, which is achieved by reusing the parameters of the pre-trained image-text processing model without the need to add new parameters or extensive retraining.

[0033] In particular, some example implementations allow the parameters of the pre-trained attention pooling layer to remain fixed after pre-training of the image-text processing model, enabling the model to be directly applied to various video understanding tasks, including zero-shot video understanding tasks. This approach eliminates the need for further training, thus significantly reducing computational resources. In other implementations, the proposed techniques enable fine-tuning of the parameters of the pre-trained attention pooling layer of a pre-trained image-text processing model, which can be performed using a joint contrastive and generative image captioning loss function applied to video data. Relative to training a brand-new model, this fine-tuning process still represents a reduced computational expense, as the pre-trained parameters represent a strong starting point from which the model can be fine-tuned.

[0034] An additional technical advantage of the present disclosure is the ability to utilize a video understanding model that does not require training on large amounts of computationally complex video data. Traditional methods of video understanding typically require training on extensive video datasets, which is a computationally costly and time-consuming process due to the high dimensionality and complexity of video data. However, the proposed method overcomes this obstacle by training only on image data while retaining the ability to perform video understanding tasks.

[0035] The techniques described herein can be used to perform various video understanding tasks. One such task is video classification. In this task, the specific input data can be an unclassified video from a media library, and the output data can be class labels that accurately describe the video content, thus enabling organized storage and retrieval.

[0036] Another task to which the technique is applicable is video question answering. In this scenario, the input data can be a video and a specific question about the video content, such as "what is the main action in the video?". The output data can be an accurate answer to the question asked, thus enhancing, for example, the interactive learning experience in an educational environment.

[0037] Additionally, the technique can also be used for the video captioning task. For this task, the specific input data can be a video without any text description. The output data generated by the technique can be a detailed text description of the video content, thus providing, for example, accessibility benefits such as assisting hearing-impaired individuals by providing captions.

[0038] In the field of robotics, the technique can interpret video input to understand and navigate a robot's environment. Here, the input data can be a real-time video stream capturing the robot's surrounding environment, and the output data can be a set of instructions for the robot to navigate its environment safely and efficiently.

[0039] As another example, the technology can also be used to perform video retrieval based on a text query. In this case, the input data can be a specific text query, and the output data can be a list of videos that match the query, thereby improving the search efficiency in information retrieval tools.

[0040] Example video processing model

[0041] Example image-text processing model

[0042] Thus, the sub-section describes an example image-text processing model. Other image-text processing models exist and can be used according to the techniques described herein.

[0043] The Contrastive Captioner (CoCa) is an encoder-decoder architecture that combines contrastive pre-training and generative pre-training methods. It is designed to facilitate image, text, and image-text representation learning. CoCa employs a cascaded decoder design, where the lower half unimodal decoder encodes text context using self-attention with a causal mask, and the upper half multimodal decoder uses cross-attention to align images and text. The model is trained using a joint contrastive loss and a captioning loss. For text representation, the [CLS] token from the unimodal decoder is used as the global text representation for the contrastive loss, and the captioning loss is applied per text token to learn fine-grained visual-text information.

[0044] CoCa uses two attention pooling layers (simply called pooling layers) to extract image representations, where the generative attention pooling layer is used to generate embeddings for the captioning loss, and the contrastive attention pooling layer and its output together serve as the contrastive image embedding. The benefits of such an architecture design are twofold. First, regardless of the input image resolution, the pooling layers produce a fixed number of tokens (e.g., 256 tokens as the generative embedding and 1 token as the contrastive embedding), making the image encoder and text decoder more modular and adaptable to other modalities. Second, the pooling layers serve as lightweight adapters, keeping the pre-trained ViT as the backbone frozen for many downstream tasks. For example, it is shown that by only fine-tuning the pooling layers, the pre-trained CoCa can already achieve 90.6% top-1 ImageNet accuracy, in which case the frozen ViT has not seen any ImageNet data. The example implementation of this disclosure adopts this design and adapts it to the video-text domain.

[0045] Example techniques for migrating a pre-trained model to video-text tasks

[0046] This sub - section describes how to quickly transform the example image CoCa model into a VideoCoCa model by tuning a small fraction of the parameters. A small batch of input videos can be represented as , where is the number of frames evenly sampled from the video. Some example implementations extract tokens from images by dividing the image into non - overlapping patches and linearly projecting them. Then, all the tokens can be concatenated together to form a sequence, resulting in a mini - batch of sequence tokens of shape , where . Position embeddings can be added to this sequence to obtain . Some example implementations also use a text decoder to extract text representations. The frame - level representation can be obtained by forwarding to an image encoder, where is the number of encoder layers.

[0047] There are various ways to adapt CoCa to videos. Some example methods are described in the following paragraphs.

[0048] Attention Pooling.

[0049] To obtain a video representation, some example implementations concatenate all the spatial tokens along the time dimension as , and then feed it into a generative pooler and a contrastive pooler. See Figure 2 for a detailed illustration. This model corresponds to a late fusion of temporal information, similar to a factorized encoder. Compared to alternative methods where an additional new pooler is added on top of the frame - level representation to learn the video representation, the model described in this paragraph does not add any novel learnable layers, allowing all parameters from the pre - trained CoCa model to be reused with minimal additional computation, enabling zero - shot transfer from an image - text model to a video - text task.

[0050] Factorized Encoder.

[0051] This adaptation additionally adds a Transformer encoder on top of the contrastive pooler. The frame - level representation is first fed into a generative pooler and a contrastive pooler to obtain a spatial embedding . Then, the spatial embedding is reshaped to and fed into a Transformer encoder consisting of ​A transformer encoder consisting of multiple layers to model the interactions between tokens from different frames. The output tokens are used for the caption addition loss, and their globally averaged pooled embeddings in the time dimension are used for the contrastive loss. Some example implementations can use . Note that, unlike alternative methods in which the spatial embeddings are summarized by a learnable class token prefixed, some example implementations of the present disclosure can use the output of the contrastive pooler as the representation.

[0052] Joint spatio-temporal encoder.

[0053] This adaptation can include the use of a spatio-temporal attention model or a joint spatio-temporal model. Specifically, some example implementations can reshape the sequence tokens into , and then add position embeddings to obtain which can be initialized by temporarily repeating the position embeddings from a pre-trained image model. This allows the CoCa image encoder to encode the pairwise interactions between all spatio-temporal tokens from the first layer. The spatio-temporal representation is then fed into a generative pooler and a contrastive pooler to obtain the final task-specific representation. This model adaptation corresponds to the early fusion of temporal information and does not add any new learnable layers. However, due to the linearly increasing number of tokens, this makes the self-attention calculations in the encoder heavier.

[0054] Average pooling.

[0055] In this adaptation, after the attention pooler, the frame-level representations are simply individually average pooled in the time dimension. It ignores the temporal information.

[0056] Example model visualization

[0057] Figure 1 FIG. shows an example framework for fine-tuning a pre-trained image-text model 12 to perform video understanding tasks. A video can include a sequence of images or a series of images (e.g., digital images) that, when displayed in a fast and continuous manner, create an illusion of motion. These images, also referred to as frames, can be captured or created in various ways including by using digital cameras, computer graphics, or animation techniques. Each frame in the video can include pixel data that collectively represents the visual information in the frame. The pixel data can include information about the color, brightness, and other visual attributes of each individual pixel in the frame. The video can also include audio data synchronized with the image data.

[0058] As Figure 1As shown, the input video 14 is processed by a pre-trained image-text model 12 that is configured to perform video understanding tasks. The input video 14 contains multiple image frames, and the image-text model 12 is pre-trained on an image-text dataset.

[0059] The pre-trained image-text model 12 includes several components. One of the important components is the pre-trained unimodal image encoder 16. The unimodal image encoder 16 can individually process each of the image frames of the input video 14 to generate a plurality of frame embeddings 18. These frame embeddings 18 can represent the visual content of each frame in the latent space, thereby capturing the important visual features of the video.

[0060] The generated frame embeddings 18 are then processed by attention pooling layers 20. These attention pooling layers 20 further refine the frame embeddings 18 by focusing on the most prominent features and discarding less relevant information. As a result, the attention pooling layers 20 can generate two sets of embeddings: contrastive embeddings 22 and generative embeddings 24. The contrastive embeddings 22 can correspond to an overall representation of the discriminative features of the input video, while the generative embeddings 24 can be used to generate a text description or prediction about the video content.

[0061] In some implementations, the frame embeddings 18 can be combined to form a set of combined frame embeddings. This can be achieved by concatenating the frame embeddings 18 along the temporal dimension. As a result, a set of flattened frame embeddings representing the temporal sequence of frames in the video is generated.

[0062] Considering the optional text aspect of the model, the image-text processing model 12 also includes a unimodal text decoder 26. This component processes a set of input text tokens 28 and generates text embeddings. These text embeddings include global text tokens 30 that serve as an integrated representation of the input text.

[0063] Furthermore, the image-text processing model 12 incorporates a multimodal decoder 32. The multimodal decoder 32 processes the generative embeddings 24 obtained from the attention pooling layers 20 and the text embeddings obtained from the unimodal text decoder 26. The output of the multimodal decoder 32 is a set of text 34 representing the output of the video understanding task.

[0064] Two types of loss function terms, namely, a generative loss term 36 and a contrastive loss term 38, can be applied to the output of the image-text processing model 12. The generative loss term 36 is applied to a set of output texts 34 generated by the multimodal decoder 32, e.g., aiming to minimize the difference between the generated text and the ground-truth text. The contrastive loss term 38 is applied between the contrastive embedding 22 and the global text tokens 30. For example, the objective of the contrastive loss term 38 can be to bring the contrastive embedding 22 and the global text tokens 30 closer in the latent space if the inputs are positive inputs associated with each other; and to push the contrastive embedding 22 and the global text tokens 30 apart if the inputs are negative inputs not associated with each other.

[0065] In Figure 2 is depicted the processing of the input video 214. As shown in this figure, the input video 214 is a sequence of frames containing important visual information. These frames are processed by a pre-trained unimodal image encoder ( Figure 2 not explicitly depicted in Figure 1 ), which is a component of the pre-trained image-text model mentioned in the previous description of . The unimodal image encoder processes each of the frames in the input video 214 to generate a set of frame token embeddings 218.

[0066] The frame token embeddings 218 represent the visual content of each frame in the latent space. These embeddings 218 capture the visual features of the video, which can be further processed to extract more complex representations. In this context, the frame token embeddings 218 are building blocks for understanding video content as they serve as the main input for subsequent processing stages.

[0067] Then, the frame token embeddings 218 are flattened into a set of N×T flattened tokens 219. This flattening process allows the model to represent the time series of frames in the video as a single unified data structure. The flattened tokens 219 maintain the temporal order of the frames and thus preserve the temporal information contained in the video. The dimension N×T of the flattened tokens 219 reflects the number of frames (T) in the video and the number of tokens (N) generated for each frame.

[0068] The flattened tokens 219 are then input into two attention pooling layers: a generative pooling layer 220 and a contrastive pooling layer 221. These attention pooling layers 220, 221 further process the flattened tokens 219, thereby focusing on the most important features and discarding less relevant information.

[0069] The generative pooling layer 220 processes the flattened tokens 219 to generate a set of generative embeddings 224. These generative embeddings 224 can be used to generate a text description or prediction about the video content. Figure 2The generation process is not explicitly depicted in [the figure], but the generation process can be performed by a separate generative model (not shown), which can be part of the entire video understanding system.

[0070] The contrastive pooling layer 221 processes the flattened tokens 219 to generate a contrastive embedding 222. The contrastive embedding 222 is an overall representation of the discriminative features of the input video. It captures the unique aspects of the video that distinguish it from other videos. The contrastive embedding 222 can be used to perform various video understanding tasks, such as video classification or video retrieval.

[0071] In some embodiments, the contrastive embedding 222 and / or the generative embedding 224 can be directly used as the model output. In other embodiments, these embeddings 222, 224 can be further processed (e.g., using additional feed-forward layers and / or softmax layers) to generate model predictions for video understanding tasks. The prediction can be a class label (for video classification tasks), a text description (for video captioning tasks), an answer to a question (for video question answering tasks), or any other suitable form of prediction or output.

[0072] Example techniques for fine-tuning video-text data

[0073] In addition to the direct zero-shot transfer from image-text CoCa to video-text tasks, some example implementations can further push the limits of VideoCoCa by continuously pre-training on large-scale video-text paired data. This sub-section explores four different learning options.

[0074] Fine-tuning (FT).

[0075] In this setting, the training system can unfreeze all the parameters of the pre-trained CoCa model during continuous video-text pre-training, including the parameters of the encoder and decoder, as well as the parameters of the generative pooler and contrastive pooler. These parameters are fine-tuned together with the newly added learnable layers.

[0076] Frozen encoder-decoder tuning (frozen).

[0077] In this method, the training system can freeze the parameters of the encoder and decoder and only tune the parameters of the generative pooler and contrastive pooler. This allows most of the parameters of the pre-trained CoCa model to be reused.

[0078] Frozen tuning then fine-tuning (frozen + FT).

[0079] With a small amount of parameters of the given pooler, the frozen encoder-decoder tuning may converge very fast (see Figure 4). Therefore, in this method, the training system can first perform frozen feature tuning and then fine-tuning. In this way, the parameters of the pooler can be trained quickly, thus making the fine-tuning more stable. Note that this is a two-step tuning method.

[0080] Frozen encoder tuning (LiT).

[0081] In this example method, only the parameters of the pre-trained CoCa image encoder are frozen, and the parameters of the pooler and the decoder are tuned. Since the computation of image representations is much heavier than that of text representations, this not only allows the training system to pre-compute the frame-level embeddings once to save the TPU memory and computation for development, but also provides a sufficient amount of learnable parameters for task adaptation.

[0082] Example method

[0083] Figure 3 A flowchart illustrating an example method for performing video understanding tasks with improved computational efficiency.

[0084] In block 302, the method begins with a computing system accessing a pre-trained image-text processing model. The image-text processing model may be stored in a memory device or storage medium of the computing system. The image-text processing model may include one or more pre-trained attention pooling layers with multiple parameters. The image-text processing model may have been pre-trained on a joint contrastive and generative image captioning loss function. For example, the image-text processing model may be a contrastive caption adder (CoCa) model that has been pre-trained on a dataset of pairs including images and image captions.

[0085] In block 304, the computing system obtains an input video containing a plurality of image frames. The input video may be obtained from various sources such as a user device, a network server, a storage device, a camera, or a video streaming service. The input video may include a sequence of image frames depicting a dynamic scene or event. Each image frame may be a two-dimensional array of pixel values, and the sequence of image frames may represent the temporal evolution of the scene or event.

[0086] In block 306, the computing system utilizes a pre-trained image-text processing model to process the input video to generate predictions for video understanding tasks. The processing can involve applying the pre-trained image-text processing model to each image frame of the input video to generate a set of frame embeddings. The frame embeddings can be processed by a pre-trained attention pooling layer of the image-text processing model to generate a set of generative embeddings and a set of contrastive embeddings. The generative embeddings and contrastive embeddings can be used to generate predictions for video understanding tasks, such as classification of the video, responses to questions about the video, captions for the video, or text descriptions of the video.

[0087] In block 308, the computing system provides the predictions for the video understanding tasks as output. The output can be provided to a user, a client device, a server, a database, a display device, or another component of the computing system. The output can be provided in various forms, such as a text string, a data file, a database entry, a network message, a display signal, or an audio signal.

[0088] Figure 3 The flowchart provides a high-level overview of an example method for performing video understanding tasks with improved computational efficiency. The method can be implemented by a computing system including one or more computing devices. The method utilizes a pre-trained image-text processing model to process the input video and generate predictions for video understanding tasks. The pre-trained image-text processing model can include one or more pre-trained attention pooling layers with multiple parameters. The method can significantly reduce the computational resources required to perform video understanding tasks by reusing the parameters of the pre-trained image-text processing model without the need to add new parameters or extensive retraining.

[0089] The method can be applied to various types of video understanding tasks, including video classification, video question answering, video retrieval, and video captioning. The method is particularly beneficial for applications that require real-time or near-real-time processing of video data, such as video streaming services, video surveillance systems, autonomous driving systems, and interactive gaming systems.

[0090] The method can be implemented in various types of computing systems, including personal computers, server computers, mobile devices, cloud computing systems, machine learning platforms, and distributed computing systems. The method can be implemented in various forms, including software programs, hardware devices, firmware modules, machine learning models, cloud services, or combinations thereof.

[0091] Example apparatus and system

[0092] Figure 4AFIG. 0 depicts a block diagram of an example computing system 100 that can implement the video-text model described herein, according to an example embodiment of the present disclosure. The system 100 includes a user computing device 102, a server computing system 130, and a training computing system 150 communicatively coupled via a network 180.

[0093] The user computing device 102 can be any type of computing device, such as, for example, a personal computing device (e.g., a laptop computer or a desktop computer), a mobile computing device (e.g., a smart phone or a tablet computer), a gaming console or controller, a wearable computing device, an embedded computing device, or any other type of computing device.

[0094] The user computing device 102 includes one or more processors 112 and a memory 114. The one or more processors 112 can be any suitable processing device (e.g., a processor core, a microprocessor, an ASIC, an FPGA, a controller, a microcontroller, etc.) and can be one processor or multiple processors operatively connected. The memory 114 can include one or more non-transitory computer-readable storage media, such as RAM, ROM, EEPROM, EPROM, flash memory devices, disks, etc., and combinations thereof. The memory 114 can store data 116 and instructions 118 that are executed by the processor 112 to cause the user computing device 102 to perform operations.

[0095] In some implementations, the user computing device 102 can store or include one or more machine learning models 120. For example, the machine learning model 120 can be or otherwise include various machine learning models, such as neural networks (e.g., deep neural networks) or other types of machine learning models, including non-linear models and / or linear models. Neural networks can include feed-forward neural networks, recurrent neural networks (e.g., long short-term memory recurrent neural networks), convolutional neural networks, or other forms of neural networks. Some example machine learning models can utilize an attention mechanism, such as self-attention. For example, some example machine learning models can include a multi-head self-attention model (e.g., a transformer model).

[0096] In some implementations, one or more machine learning models 120 can be received from the server computing system 130 via the network 180, stored in the user computing device memory 114, and then used or otherwise implemented by the one or more processors 112. In some implementations, the user computing device 102 can implement multiple parallel instances of a single machine learning model 120 (e.g., to perform parallel video-text processing across multiple instances of video-text input).

[0097] Additionally or alternatively, one or more machine learning models 140 may be included in or otherwise stored and implemented by the server computing system 130, which communicates with the user computing device 102 according to a client-server relationship. For example, the machine learning model 140 may be implemented by the server computing system 140 as part of a web service (e.g., a video processing service). Thus, one or more models 120 may be stored and implemented at the user computing device 102, and / or one or more models 140 may be stored and implemented at the server computing system 130.

[0098] The user computing device 102 may also include one or more user input components 122 that receive user input. For example, the user input component 122 may be a touch-sensitive component (e.g., a touch-sensitive display screen or a touchpad) that is sensitive to a user input object (e.g., a finger or a stylus). The touch-sensitive component may be used to implement a virtual keyboard. Other example user input components include microphones, traditional keyboards, or other devices by which a user may provide user input.

[0099] The server computing system 130 includes one or more processors 132 and a memory 134. The one or more processors 132 may be any suitable processing device (e.g., a processor core, a microprocessor, an ASIC, an FPGA, a controller, a microcontroller, etc.) and may be one processor or multiple processors operatively connected. The memory 134 may include one or more non-transitory computer-readable storage media, such as RAM, ROM, EEPROM, EPROM, flash memory devices, disks, etc. and combinations thereof. The memory 134 may store data 136 and instructions 138, which are executed by the processor 132 to cause the server computing system 130 to perform operations.

[0100] In some implementations, the server computing system 130 includes one or more server computing devices or is otherwise implemented by the one or more server computing devices. In the case where the server computing system 130 includes multiple server computing devices, such server computing devices may operate according to a sequential computing architecture, a parallel computing architecture, or some combination thereof.

[0101] As described above, the server computing system 130 may store or otherwise include one or more machine learning models 140. For example, the model 140 may be or otherwise include various machine learning models. Example machine learning models include neural networks or other multi-layer non-linear models. Example neural networks include feedforward neural networks, deep neural networks, recurrent neural networks, and convolutional neural networks. Some example machine learning models may utilize attention mechanisms, such as self-attention. For example, some example machine learning models may include multi-head self-attention models (e.g., Transformer models).

[0102] The user computing device 102 and / or the server computing system 130 may train the models 120 and / or 140 via interaction with a training computing system 150 communicatively coupled via a network 180. The training computing system 150 may be separate from the server computing system 130 or may be part of the server computing system 130.

[0103] The training computing system 150 includes one or more processors 152 and a memory 154. The one or more processors 152 may be any suitable processing device (e.g., a processor core, a microprocessor, an ASIC, an FPGA, a controller, a microcontroller, etc.) and may be a single processor or multiple processors operatively connected. The memory 154 may include one or more non-transitory computer-readable storage media, such as RAM, ROM, EEPROM, EPROM, flash memory devices, disks, etc. and combinations thereof. The memory 154 may store data 156 and instructions 158, which are executed by the processor 152 to cause the training computing system 150 to perform operations. In some implementations, the training computing system 150 includes one or more server computing devices or is otherwise implemented by the one or more server computing devices.

[0104] The training computing system 150 may include a model trainer 160 that trains the machine learning models 120 and / or 140 stored at the user computing device 102 and / or the server computing system 130 using various training or learning techniques, such as, for example, error backpropagation. For example, a loss function may be backpropagated through the model to update one or more parameters of the model (e.g., based on the gradient of the loss function). Various loss functions may be used, such as mean squared error, likelihood loss, cross-entropy loss, hinge loss, and / or various other loss functions. Gradient descent techniques may be used to iteratively update the parameters over multiple training iterations.

[0105] In some implementations, performing backpropagation of errors can include performing truncated backpropagation through time. The model trainer 160 can perform various generalization techniques (e.g., weight decay, dropout, etc.) to improve the generalization ability of the model being trained.

[0106] Specifically, the model trainer 160 can train the machine learning models 120 and / or 140 based on the training data set 162. In some implementations, if the user has provided consent, the training examples can be provided by the user computing device 102. Thus, in such implementations, the model 120 provided to the user computing device 102 can be trained by the training computing system 150 based on user-specific data received from the user computing device 102. In some cases, this process can be referred to as personalizing the model.

[0107] The model trainer 160 includes computer logic for providing the desired functionality. The model trainer 160 can be implemented in hardware, firmware, and / or software that controls a general-purpose processor. For example, in some implementations, the model trainer 160 includes program files stored on a storage device, loaded into memory, and executed by one or more processors. In other implementations, the model trainer 160 includes one or more sets of computer-executable instructions stored on a tangible computer-readable storage medium (such as RAM, a hard disk, or optical or magnetic media).

[0108] The network 180 can be any type of communication network, such as a local area network (e.g., an intranet), a wide area network (e.g., the Internet), or some combination thereof, and can include any number of wired or wireless links. Generally, communication on the network 180 can be performed via any type of wired and / or wireless connection, using various communication protocols (e.g., TCP / IP, HTTP, SMTP, FTP), encodings or formats (e.g., HTML, XML), and / or security schemes (e.g., VPN, secure HTTP, SSL).

[0109] Figure 4A An example computing system that can be used to implement the present disclosure is illustrated. Other computing systems can also be used. For example, in some implementations, the user computing device 102 can include the model trainer 160 and the training data set 162. In such implementations, the model 120 can be trained and used locally on the user computing device 102. In some of such implementations, the user computing device 102 can implement the model trainer 160 to personalize the model 120 based on user-specific data.

[0110] Figure 4BDepicts a block diagram of an example computing device 10 performing in accordance with an example embodiment of the present disclosure. The computing device 10 can be a user computing device or a server computing device.

[0111] The computing device 10 includes a plurality of applications (e.g., Application 1 to Application N). Each application contains its own machine learning library and machine learning model. For example, each application can include a machine learning model. Example applications include a text messaging application, an email application, a dictation application, a virtual keyboard application, a browser application, and the like.

[0112] As Figure 4B shown, each application can communicate with a plurality of other components of the computing device, such as one or more sensors, a context manager, a device status component, and / or additional components. In some implementations, each application can communicate with each device component using an API (e.g., a common API). In some implementations, the API used by each application is specific to that application.

[0113] Figure 4C Depicts a block diagram of an example computing device 50 performing in accordance with an example embodiment of the present disclosure. The computing device 50 can be a user computing device or a server computing device.

[0114] The computing device 50 includes a plurality of applications (e.g., Application 1 to Application N). Each application communicates with a central intelligence layer. Example applications include a text messaging application, an email application, a dictation application, a virtual keyboard application, a browser application, and the like. In some implementations, each application can use an API (e.g., a common API across all applications) to communicate with the central intelligence layer (and the models stored therein).

[0115] The central intelligence layer includes a plurality of machine learning models. For example, as Figure 4C shown, a corresponding machine learning model can be provided for each application, and the corresponding machine learning model can be managed by the central intelligence layer. In other implementations, two or more applications can share a single machine learning model. For example, in some implementations, the central intelligence layer can provide a single model for all applications. In some implementations, the central intelligence layer is included within or otherwise implemented by the operating system of the computing device 50.

[0116] The central intelligence layer can communicate with a central device data layer. The central device data layer can be a centralized data repository of the computing device 50. As Figure 4CAs shown, the central device data layer can communicate with multiple other components of the computing device, such as one or more sensors, a context manager, a device status component, and / or additional components. In some implementations, the central device data layer can communicate with each device component using an API (e.g., a private API).

[0117] Additional disclosure

[0118] The techniques discussed herein relate to servers, databases, software applications, and other computer-based systems, as well as the actions taken and the information sent to and from such systems. The inherent flexibility of computer-based systems allows for many possible configurations, combinations, and divisions of tasks and functions between and among components. For example, the processes discussed herein can be implemented using a single device or component or multiple devices or components working in combination. Databases and applications can be implemented on a single system or distributed across multiple systems. Distributed components can operate sequentially or in parallel.

[0119] Although the subject matter has been described in detail with respect to various specific example embodiments thereof, each example is provided by way of illustration and not limitation of the disclosure. Those skilled in the art will readily generate alterations, variations, and equivalents to such embodiments upon understanding the foregoing. Accordingly, the disclosure does not exclude including such modifications, variations, and / or additions to the subject matter that would be readily understood by those of ordinary skill in the art. For example, features shown or described as part of one embodiment can be used with another embodiment to yield yet a further embodiment. Accordingly, the disclosure is intended to cover such alterations, variations, and equivalents.

Claims

1. A computer-implemented method for performing video understanding tasks with improved computational efficiency, the method comprising: accessing, by a computing system including one or more computing devices, a pre-trained image-text processing model, wherein the pre-trained image-text processing model includes one or more pre-trained attention pooling layers having a plurality of parameters, and wherein the pre-trained image-text processing model has been pre-trained on a joint contrastive and generative image captioning loss function; obtaining, by the computing system, an input video including a plurality of image frames; processing, by the computing system, the input video using the pre-trained image-text processing model having the one or more pre-trained attention pooling layers with the same number of parameters to generate a prediction for the video understanding task as an output of the pre-trained image-text processing model; and providing, by the computing system, the prediction for the video understanding task as an output.

2. The computer-implemented method of claim 1, wherein: the pre-trained image-text processing model includes a pre-trained unimodal image encoder configured to process an input image to generate one or more frame embeddings; the one or more pre-trained attention pooling layers are configured to process the one or more frame embeddings to generate one or more contrastive embeddings and one or more generative embeddings; and and processing, by the computing system, the input video using the pre-trained image-text processing model includes: processing each of the plurality of image frames separately using the pre-trained unimodal image encoder to generate a plurality of frame embeddings respectively for the plurality of image frames; combining the plurality of frame embeddings to form a set of combined frame embeddings; and processing the set of combined frame embeddings using the one or more attention layers to generate one or more generative embeddings and one or more contrastive embeddings.

3. The computer-implemented method according to claim 2, wherein, Combining the plurality of frame embeddings to form a set of combined frame embeddings includes concatenating the plurality of frame embeddings along a temporal dimension to generate a set of flattened frame embeddings.

4. The computer-implemented method according to claim 2, wherein, Combining the plurality of frame embeddings to form a set of combined frame embeddings includes reshaping the plurality of frame embeddings into a joint spatio-temporal representation.

5. The computer-implemented method according to claim 1, wherein, After the pre-training of the pre-trained image-text processing model on the joint contrastive and generative image captioning loss function, the parameters of the one or more pre-trained attention pooling layers of the pre-trained image-text processing model have remained fixed.

6. The computer-implemented method according to claim 5, wherein, The video understanding task includes a zero-shot video understanding task.

7. The computer-implemented method according to claim 5, wherein, The pre-trained image-text processing model has been trained only on training data including only still images.

8. The computer-implemented method according to claim 1, wherein, After the pre-training of the pre-trained image-text processing model on the joint contrastive and generative image captioning loss function, all the parameters of the pre-trained image-text processing model have remained fixed.

9. The computer-implemented method according to claim 1, wherein, After the pre-training of the pre-trained image-text processing model on the joint contrastive and generative image captioning loss function, the parameters of the one or more pre-trained attention pooling layers of the pre-trained image-text processing model have been further fine-tuned.

10. The computer-implemented method according to claim 9, wherein, The parameters of the one or more pre-trained attention pooling layers of the pre-trained image-text processing model have been further fine-tuned using the joint contrastive and generative image captioning loss function applied to video data.

11. The computer-implemented method according to claim 1, wherein, After the pre-training of the pre-trained image-text processing model on the joint contrastive and generative image captioning loss function, all the parameters of the pre-trained image-text processing model have been further fine-tuned.

12. The computer-implemented method according to claim 1, wherein: The pre-trained image-text processing model includes a pre-trained unimodal image encoder configured to process an input image to generate one or more frame embeddings; The one or more pre-trained attention pooling layers are configured to process the one or more frame embeddings to generate one or more contrastive embeddings and one or more generative embeddings; The pre-trained image-text processing model includes a pre-trained multimodal decoder configured to process at least the one or more generative embeddings to generate a generative output; And The parameters of the pre-trained unimodal image encoder have been kept fixed, while the parameters of the pre-trained attention pooling layers and the pre-trained multimodal decoder have been further fine-tuned using the joint contrastive and generative image captioning loss function applied to video data.

13. The computer-implemented method according to claim 1, wherein, The one or more pre-trained attention pooling layers include a generative pooling layer configured to generate one or more generative embeddings and a contrastive pooling layer configured to generate one or more contrastive embeddings.

14. The computer-implemented method according to claim 1, wherein, The method further includes, before processing the input video, attaching an additional encoder model by the computing system to at least one of the one or more attention pooling layers.

15. The computer-implemented method according to claim 1, wherein, The pre-trained image-text processing model includes a decoder configured to process the embeddings generated by the one or more attention layers to generate a text output.

16. The computer-implemented method according to claim 1, wherein, The pre-trained image-text processing model further includes a unimodal text decoder, and wherein processing the input video includes using the unimodal text decoder to process a set of input texts associated with the input video.

17. The computer-implemented method according to claim 1, wherein, The video understanding task includes a video classification task.

18. The computer-implemented method according to claim 1, wherein, The video understanding task includes a video question answering task.

19. The computer-implemented method according to claim 1, wherein, The video understanding task includes a video captioning task.

20. One or more non-transitory computer-readable media that together store: Pre-trained image-text processing model, wherein, The pre-trained image-text processing model includes one or more pre-trained attention pooling layers having a plurality of parameters, and wherein the pre-trained image-text processing model has been pre-trained on a joint contrastive and generative image captioning loss function; and Computer-executable instructions for performing operations, the operations including processing, by the computing system, an input video including a plurality of image frames using the pre-trained image-text processing model having the one or more pre-trained attention pooling layers with the same number of parameters to generate a prediction of a video understanding task as an output of the pre-trained image-text processing model.