Retrieval-augmented video processing
The neural network system addresses the inflexibility of existing video processing systems by dividing videos into segments and using shared neural networks to generate coherent outputs for videos of varying lengths, achieving efficient and accurate processing.
Patent Information
- Application Number
- PCT/EP2024/082752
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-10-01
- Filing Date
- 2024-11-18
- Publication Date
- 2025-05-22
AI Technical Summary
Existing video processing systems are inflexible and require tedious reconfiguration to process videos of varying lengths, limiting their ability to efficiently generate outputs for long or continuously streaming videos.
A neural network system that divides videos into smaller segments and processes them using shared video encoder and decoder neural networks, along with an autoregressive transformer network, to generate outputs that characterize each segment, allowing for flexible and efficient processing of videos of any length.
The system can generate high-quality, spatial-temporally coherent outputs for long videos by parallelizing the processing of video segments, improving efficiency and scalability while maintaining contextual accuracy.
Smart Images

Figure EP2024082752_22052025_PF_FP_ABST
Abstract
Description
[0001] RETRIEVAL-AUGMENTED VIDEO PROCESSING
[0002] CROSS-REFERENCE TO RELATED APPLICATIONS
[0003] This application claims priority to U.S. Provisional Application No. 63 / 600,621, filed on November 17, 2023 and to U.S. Provisional Application No. 63 / 702,082, filed on October 1, 2024. The disclosure of the prior applications is considered part of and is incorporated by reference in the disclosure of this application.
[0004] BACKGROUND
[0005] This specification relates to processing videos using neural networks.
[0006] Neural networks are machine learning models that employ one or more layers of nonlinear units to predict an output for a received input. Some neural networks include one or more hidden layers in addition to an output layer. The output of each hidden layer is used as input to the next layer in the network, i.e., the next hidden layer or the output layer. Each layer of the network generates an output from a received input in accordance with current values of a respective set of parameters.
[0007] SUMMARY
[0008] This specification describes a neural network system implemented as computer programs on one or more computers in one or more locations that processes a video using neural networks and generates outputs that characterize video segments of the video. The video includes a respective video frame at each of a plurality time steps.
[0009] According to an aspect, there is provided a method performed by one or more computers, the method comprising: obtaining a temporal sequence of video frames that span a plurality of time steps; partitioning the temporal sequence of video frames into a plurality of video segments, wherein each video segment comprises a respective video frame at each time step in a respective subset of the plurality time steps; and for each video segment in the plurality of video segments: generating an encoded representation corresponding to the video segment using one or more neural networks to process respective video frames included in the video segment and in any video segments that precede the video segment in the plurality of video segments; and generating, using a text decoder neural network and from the encoded representation corresponding to the video segment, a text output that characterizes the video segment in parallel with generating text outputs that respectively characterize other video segments in the plurality of video segments using the text decoder neural network and from encoded representations corresponding respectively to the other video segments.
[0010] The one or more neural networks may comprise a video encoder neural network that processes the respective video frames included in the video segment to generate a first initial encoded representation of the video segment.
[0011] Processing the respective video frames included in the video segment to generate the first initial encoded representation of the video segment may comprise processing the respective video frames included in the video segment to generate the first initial encoded representation of the video segment in parallel with processing the respective video frames included in the other video segments to generate the first initial encoded representations of the other video segments using the video decoder neural network.
[0012] The one or more neural networks may comprise a video memory transformer neural network that generates a second initial encoded representation of the video segment based on the first initial encoded representation of the video segment.
[0013] The generating may comprise: obtaining a video memory transformer network input that comprise (i) a set of memory vectors and (ii) a set of feature vectors selected from the first initial encoded representation; and processing the video memory transformer network input using the video memory transformer neural network to generate the second initial encoded representation of the video segment based on applying an attention mechanism over (i) the set of memory vectors and (ii) the set of feature vectors to determine (i) a set of updated memory vectors and (ii) a set of updated feature vectors.
[0014] The one or more neural networks may comprise an autoregressive transformer neural network that generates the encoded representation corresponding to the video segment based on the second initial encoded representation of the video segment.
[0015] The generating may comprise, for each video segment in the plurality of video segments: obtaining an autoregressive transformer network input that comprises (i) the second initial encoded representation of the video segment and (ii) the second initial encoded representations of any video segments that precede the video segment in the plurality of video segments; and processing the autoregressive transformer network input using the autoregressive transformer neural network to generate the encoded representation of the video segment based on applying an attention mechanism over (i) the second initial encoded representation of the video segment and (ii) the second initial encoded representations of any video segments that precede the video segment in the plurality of video segments. Generating the text output for each video segment in parallel with generating text outputs that respectively characterize the other video segments may comprise generating a plurality of instances of the text decoder neural network that correspond respectively to the plurality of video segments and that have identical parameter values with each other.
[0016] Generating the text output for each video segment in parallel with generating text outputs that respectively characterize the other video segments may comprise processing, by each of the plurality of instances of the text decoder neural network, a different encoded representation in accordance with the identical parameter values to generate the text output for each video segment.
[0017] The text decoder neural network may be an autoregressive neural network that generates the text output by, at each of multiple output steps, generating a text token conditioned on any text tokens that have already been generated in preceding output steps.
[0018] The text output may comprise one of: a text caption output for the video segment, an action recognition output for the video segment, or an action localization output for the video segment.
[0019] The method may further comprise training the one or more neural networks and the text decoder neural network on a plurality of training video segments that are each associated with a ground truth text output, wherein for each training video segment, the ground truth text output identifies a beginning time step, an end time step, or both of the training video segment, and describes one or more actions depicted in the training video segment.
[0020] During training, the text decoder neural network may generate training text outputs that respectively characterize the plurality of training video segments in a sequential order by using a masked attention mechanism.
[0021] According to another aspect, there is provided a method performed by one or more computers, the method comprising: obtaining a video segment that comprises a respective video frame at each of a plurality of time steps; processing, using a video encoder neural network, a video encoder input that includes the respective video frames included in the video segment to generate an initial encoded video segment representation that corresponds to the video segment; performing a search in a retrieval dataset that stores a plurality of embeddings for one or more most similar embeddings to the initial encoded video segment representation according to a similarity measure, wherein each embedding corresponds to a respective text sequence; and generating, based on (i) the initial encoded video segment representation and (ii) the one or more most similar embeddings, an output that characterizes the video segment. Generating, based on (i) the initial encoded video segment representation and (ii) the one or more most similar embeddings, an output that characterizes the video segment may comprise processing, using a dimensionality reduction neural network, a first dimensionality reduction input that includes the initial encoded video segment representation to generate a first reduced encoded video segment representation; and processing, using a decoder neural network, a first decoder input that includes (i) data derived from the first reduced encoded video segment representation and (ii) the respective text sequences that correspond to the one or more most similar embeddings or (iii) the one or more most similar embeddings or both (ii) and (iii) to generate the output that characterizes the video segment.
[0022] Generating, based on (i) the initial encoded video segment representation and (ii) the one or more most similar embeddings, an output that characterizes the video segment may comprise: processing, using the dimensionality reduction neural network, a second dimensionality reduction input that includes the initial encoded video segment representation to generate a second reduced encoded video segment representation; generating a second decoder input based on (i) the second reduced encoded video segment representation and (ii) the one or more most similar embeddings; and processing, using the decoder neural network, the second decoder input to generate the output that characterizes the video segment.
[0023] Generating the second decoder input based on (i) the second reduced encoded video segment representation and (ii) the one or more most similar embeddings may comprise: generating, from the one or more most similar embeddings, a first pooled embedding; generating a first updated encoded video segment representation that corresponds to the video segment based on processing at least (i) the second reduced encoded video segment representation and (ii) the first pooled embedding using the autoregressive transformer neural network; and using the first updated encoded video segment representation as the second decoder input.
[0024] Generating, based on (i) the initial encoded video segment representation and (ii) the one or more most similar embeddings, an output that characterizes the video segment may comprise: processing, using the dimensionality reduction neural network, a third dimensionality reduction input that includes the initial encoded video segment representation to generate a third reduced encoded video segment representation; generating a second updated encoded video segment representation that corresponds to the video segment based on (i) the third reduced encoded video segment representation and (ii) the one or more most similar embeddings; and processing, using the decoder neural network, a third decoder input that includes (i) the second updated encoded video segment representation and (ii) the respective text sequences that correspond to the one or more most similar embeddings or (iii) the one or more most similar embeddings or both (ii) and (iii) to generate the output that characterizes the video segment.
[0025] Generating the second updated encoded video segment representation may comprise: generating, from the one or more most similar embeddings, a second pooled embedding; generating the second updated encoded video segment representation that corresponds to the video segment based on processing at least (i) the third reduced encoded video segment representation and (ii) the second pooled embedding using the autoregressive transformer neural network.
[0026] The respective text sequences that correspond to the plurality of embeddings stored in the retrieval dataset may be generated by, for each embedding: obtaining the respective text sequence, wherein the respective text sequence characterizes one or more respective prestored video frames; and processing, using an embedding neural network, the respective text sequence to generate the embedding.
[0027] Obtaining the respective text sequence may comprise obtaining, from a video caption dataset, an initial text sequence that characterizes the one or more respective prestored video frames; and processing, using a summarization neural network, the initial text sequence to generate the respective sequence.
[0028] The output may comprise one of: a text caption output for the video segment, an action recognition output for the video segment, or an action localization output for the video segment.
[0029] The video segment follows one or more preceding video segments that each may comprise a respective video frame at each of a plurality of preceding time steps.
[0030] Generating the output that characterizes the video segment may comprise generating the output based on one or more initial encoded video segment representations that correspond to the one or more preceding video segments.
[0031] The video encoder input may be a composite image that is generated based on applying frame resizing to the respective video frames included in the video segment.
[0032] The method may further comprise training the video encoder neural network and the decoder neural network on a plurality of training video segments that are each associated with a ground truth output.
[0033] For each training video segment, the ground truth output may identify a beginning time step, an end time step, or both of the training video segment, and may describe one or more actions depicted in the training video segment. The training may comprise: processing, using the video encoder neural network and the decoder neural network, (i) a first training video segment and (ii) one or more most similar embeddings obtained from the retrieval dataset to generate a first training output; processing, using the video encoder neural network and the decoder neural network, a second training video segment to generate a second training output; and determining an update to values of parameters of the video encoder neural network and the decoder neural network based on a difference between the first training output and a first ground truth output associated with the first training video segment, and on a difference between the second training output and a second ground truth output associated with the second training video segment.
[0034] According to another aspect, there is provided one or more computer-readable storage media encoded with instructions that, when executed by one or more computers, cause the one or more computers to perform the operations of the above method aspect.
[0035] According to yet another aspect, there is provided a system comprising one or more computers and one or more storage devices storing instructions that when executed by one or more computers cause the one or more computers to perform the respective operations of the above method aspect.
[0036] It will be appreciated that features described in the context of one aspect may be combined with features described in the context of another aspect.
[0037] The subject matter described in this specification can be implemented in particular embodiments so as to realize one or more of the following advantages.
[0038] Many existing video processing systems are configured to operate on videos that have a fixed, predetermined length. Repeated reconfiguration of the architecture of the neural networks (e.g., to change the input / output dimension of the input / output neural network layers and possibly other components) included in those systems, which can be tedious and time consuming, are often required in order to use those systems to process videos that have varying lengths.
[0039] Using some techniques described in the specification, a neural network system can generate outputs based on a video in a manner that is not constrained in the length of video, i.e., that is not constrained to processing a video having at most a certain number of frames. The neural network system can thus generate outputs that characterize respective segments of a very long, potentially continuously streaming, video that includes a sequence of many video frames. The neural network system may generate the outputs by dividing a video into smaller video segments and then processing these video segments using a shared video encoder neural network, a shared decoder neural network, and an autoregressive transformer neural network to generate the outputs that correspond respectively to these video segments. Because of this, the neural network system may be able to flexibly and computationally efficiently generate outputs for videos of varying lengths, e.g., by generating additional instances of the shared neural networks that can operate in parallel with each other, without having to restructure the neural networks to modify their output dimensions.
[0040] Parallelized processing can decrease the overall time needed for generating the outputs that characterize the video segments. This parallel processing scheme also makes the neural network system more suitable for deployment on modern parallel computing hardware, including tensor processing units (TPUs), graphics processing units (GPUs), and other hardware accelerator devices that perform parallel computing using dedicated circuitries. Such parallel computing hardware may include at least one integrated circuit including multiple processing cores and / or multiple integrated circuits within a single housing and / or multiple integrated circuits within different respective housings (e.g. having different respective electrical supplies and using different respective clock signals). Thus, different instances of the neural networks can be implemented in parallel using different corresponding ones of the cores and / or different respective ones of the integrated circuits.
[0041] As a result, the neural network system can be scaled up substantially while maintaining long range coherence between the content in the text outputs. As a particular example, this scalability allows the neural network system to generate text captions that are spatial-temporally coherent and are highly detailed for many frames in a long video.
[0042] Some techniques described in the specification can further augment the output generation process performed by the neural network system with relevant external information retrieved from a retrieval dataset by performing an online data retrieval process that can be repeated for each video segment of the video.
[0043] By augmenting the output generation process performed by using the neural network system with the repeated online retrieval process, the disclosed techniques further improve the performance of the neural network system to enable the system to generate higher quality outputs that are more accurate, contextually appropriate, and temporally aligned.
[0044] It will be appreciated that the retrieval -based augmentation techniques and the parallelization techniques can be applied either jointly or independently of one another, such that the advantages of the parallelization techniques can be achieved with or without the retrieval-based augmentation techniques (and vice versa).
[0045] The details of one or more embodiments of the subject matter described in this specification are set forth in the accompanying drawings and the description below. Other features, aspects, and advantages of the subject matter will become apparent from the description, the drawings, and the claims.
[0046] BRIEF DESCRIPTION OF THE DRAWINGS
[0047] FIG. 1 shows an example neural network system.
[0048] FIG. 2 is an example illustration of operations performed by a neural network system.
[0049] FIG. 3 is a flow diagram of an example process for processing a video.
[0050] FIG. 4 shows another example neural network system.
[0051] FIG. 5 is another example illustration of operations performed by a neural network system.
[0052] FIG. 6 is a flow diagram of an example process for processing a video.
[0053] FIG. 7 is a flow diagram of sub-steps of an implementation of one of the steps of the process of FIG. 6.
[0054] FIG. 8 is a flow diagram of sub-steps of another implementation of one of the steps of the process of FIG. 6.
[0055] FIG. 9 is a flow diagram of sub-steps of another implementation of one of the steps of the process of FIG. 6.
[0056] FIG. 10 is a flow diagram of sub-steps of another implementation of one of the steps of the process of FIG. 6.
[0057] FIG. 11 shows a quantitative example of the performance gains that can be achieved by the neural network system 100 of FIG. 1 on a video captioning task compared to existing video processing systems.
[0058] FIG. 12 shows a quantitative example of the performance gains that can be achieved by the neural network system 400 of FIG. 4 on a video captioning task compared to existing video processing systems.
[0059] Like reference numbers and designations in the various drawings indicate like elements. DETAILED DESCRIPTION
[0060] FIG. 1 shows an example neural network system 100. The neural network system 100 is an example of a system implemented as computer programs on one or more computers in one or more locations that processes one or more video segments of a video and generates outputs that characterize the one or more video segments of the video.
[0061] For example, FIG. 1 illustrates that the neural network system 100 receives a video segment 102 that includes multiple video frames, and processes the video segment 102 to generate an output 104 that characterizes the video segment 102.
[0062] The video includes a respective video frame at each of multiple time steps. Each video frame is an image that can be represented as a two-dimensional (2D) array of pixels, where each pixel is represented as a vector of one or more values. Thus, processing the video includes processing the values of the pixels of the video frames in the video.
[0063] Each video segment includes a respective video frame at each time step in a respective subset of the multiple time steps of the video. FIG. 1 illustrates that the video segment 102 includes a total of three video frames at three time steps of the video. However, in other examples, the video segment 102 can include fewer or more video frames of the video.
[0064] For example, if the video frames are black-and-white images, each pixel can be represented as an integer or floating point number (i.e., a vector with one component) representing the brightness of the pixel. As another example, if the video frames are red- green-blue (RGB) images, each pixel can be represented as a vector with three integer or floating point components, which respectively represent the intensity of the red, green, and blue color of the pixel. In some other examples YUV or HSV color channels may be employed. The pixels of a video frame can be indexed by (x, y) coordinates, where x and y are integer values.
[0065] Some examples of outputs that the neural network system 100 can generate from the video are as follows.
[0066] For example, the output for each video segment is a video captioning output. The video captioning output can include a natural language output sequence, e.g., a sequence of words, that is descriptive of the video segment.
[0067] As another example, the output for each video segment is an action recognition output. An action recognition output for a video segment can recognize an action that spans multiple video frames in the video segment. The action can for example be an action, activity, and / or other temporally varying occurrence which involves a human actor and / or a non-human actor, such as an animal, a robot, an inanimate object, or portions thereof.
[0068] When the action recognition output of the video is expressed as text, the action recognition output of the video may include at least one verb that is descriptive of the action, activity, and / or other temporally varying occurrence. Action recognition outputs may be used to facilitate video retrieval, video captioning, and / or visual question-and-answer, among other tasks.
[0069] As another example, the output for each video segment is an action localization output. The action localization output can identify an action spatially, e.g., by defining the coordinates of bounding boxes that enclose respective actions depicted in the video frames. Additionally or alternatively, the action localization output can identify an action temporally, e.g., by identifying one or more time step within the video segment during which the action is depicted in the corresponding video frames.
[0070] As another example, the output for each video segment is an object detection output. The object detection output can identify regions within each respective video frame in the video segment that are predicted to include objects. For example, the object detection output can include data defining a plurality of bounding boxes in a video frame; and, optionally, for each of the plurality of bounding boxes, a respective confidence score that represents a likelihood that an object belonging to an object category from a predetermined set of one or more object categories is present in the region of the video frame shown in the bounding box.
[0071] As another example, the neural network system 100 can receive data representing a question that is posed about the video segment 102, e.g., either together with or separate from the video segment 102, and the output for each video segment can be an answer to the question that is posed about the video segment 102.
[0072] In some implementations, the output for each video segment can be represented by text tokens selected from a vocabulary of text tokens that includes, e.g., one or more of characters, sub-words, words, punctuation marks, numbers, or other symbols that appear in a natural language or a computer programming language.
[0073] As illustrated in FIG. 1, the neural network system 100 includes a plurality of neural networks. The plurality of neural networks include a video encoder neural network 110, a decoder neural network 140, and, optionally, in some implementations, a dimensionality reduction neural network 120 and an autoregressive transformer neural network 130.
[0074] The video encoder neural network 110 is configured to receive a video encoder input that includes a video segment 102 in a video and process the video encoder input to generate an initial encoded representation (also referred to as “an initial encoded video segment representation” or “a first initial encoded representation”) that corresponds to the video segment. An “encoded representation” or “embedding” as used in this specification is a sequence of one or more vectors of numeric values, e.g., floating point values or other values, each vector having a pre-determined dimensionality.
[0075] The video encoder neural network 110 can have any appropriate neural network architecture that allows the video encoder neural network to map the video segment 102 that includes multiple video frames to an initial encoded video segment representation that corresponds to the video segment.
[0076] For example, the video encoder neural network 110 can include any appropriate types of neural network layers (e.g., embedding layers, convolution layers, attention layers, and so forth) in any appropriate number (e.g., 2 layers, or 5 layers, or 10 layers) and connected in any appropriate configuration (e.g., as a directed graph of layers).
[0077] In some implementations, the video encoder neural network 110 can have a vision transformer (ViT) architecture. For example, the video encoder neural network 110 can have one of the ViT architectures described in Piergiovanni, AJ, et al. Rethinking video vits: Sparse video tubes for joint image and video learning. CVPR, 2023; and Radford, Alec, et al. Learning transferable visual models from natural language supervision. International conference on machine learning. PMLR, 2021.
[0078] The dimensionality reduction neural network 120, when included, is configured to receive a dimensionality reduction input that includes the initial encoded video segment representation and process the dimensionality reduction input to generate a reduced encoded representation (also referred to as a “reduced encoded video segment representation” or “a second initial encoded representation”) that corresponds to the video segment 102.
[0079] The reduced encoded video segment representation has a reduced dimensionality (a reduced number of independent numerical values, e.g. such that it can be represented using a reduced number of bits) compared to the initial encoded video segment representation. Hence the reduced encoded video segment representation is more data efficient, e.g., more compact, than the initial encoded video segment representation.
[0080] For example, the dimensionality reduction neural network 120 can include any appropriate types of neural network layers (e.g., fully connected layers, pooling layers, attention layers, and so forth) in any appropriate number (e.g., 2 layers, or 5 layers, or 10 layers) and connected in any appropriate configuration (e.g., as a directed graph of layers). In some implementations, the dimensionality reduction neural network 120 can have a transformer architecture that includes one or more attention layers. For example, the dimensionality reduction neural network 120 can have one of the transformer architectures described in Ryoo, Michael S., et al. Tokenlearner: Adaptive space-time tokenization for videos. In Adv. Neural Inform. Process. Syst., 2021; and Jaegle, Andrew, et al. Perceiver: General perception with iterative attention, 2021.
[0081] In some implementations, the dimensionality reduction neural network 120 can be configured as a video memory transformer neural network. To generate the reduced encoded representation, the dimensionality reduction neural network 120 obtains a video dimensionality reduction input that includes the initial encoded representation, generates (i) a set of memory vectors and (ii) a set of feature vectors selected from the initial encoded representation, and then generates the reduced encoded representation that corresponds to the video segment based on applying an attention mechanism over (i) the set of memory vectors and (ii) the set of feature vectors to determine (i) a set of updated memory vectors and (ii) a set of updated feature vectors.
[0082] In these implementations, the dimensionality reduction neural network 120 can maintain a respective set of memory vectors for each initial encoded representation. A few examples of how such a set of memory vectors can be generated from an initial encoded representation are described in Burtsev, Mikhail S., et al. Memory transformer. arXiv preprint arXiv:2006.11527 (2020); Bulatov, Aydar, et al. Recurrent memory transformer. Advances in Neural Information Processing Systems, 35: 11079-11091, 2022; and Ryoo, Michael S., et al. Token Turing Machines. In CVPR, 2023.
[0083] The autoregressive transformer neural network 130, when included, is configured to receive a transformer input that includes (i) the initial encoded representation (in implementations where the dimensionality reduction neural network 120 is not included) or the reduced encoded representation (in implementations where the dimensionality reduction neural network 120 is included) that corresponds to the video segment 102 and, in cases where the video segment 102 is not the first video segment in the video, (ii) the initial encoded representation or the reduced encoded representation that corresponds to each of one or more preceding video segments that temporally precede the video segment 102 in the video, and to process the transformer input to generate an updated encoded representation (also referred to as an “updated encoded video segment representation” or “an encoded representation”) that corresponds to the video segment 102, by updating the initial encoded representation or the reduced encoded representation that corresponds to the video segment 102 based on applying an attention mechanism using the initial encoded representations or the reduced encoded representations that correspond to the one or more preceding video segments.
[0084] The updated encoded representation encodes information about a longer temporal relationship between different video frames at different time steps in the video that is extracted by the autoregressive transformer neural network 130 based on attending over the initial or reduced encoded representations that correspond to the video segment 102 and one or more preceding video segments that temporally precede the video segment 102 in the video.
[0085] The autoregressive transformer neural network 130 can have any of a variety of transformer-based neural network architectures. Examples of such architectures include those described in Raffel, Colin, et al., Exploring the limits of transfer learning with a unified text-to-text transformer. arXiv preprint arXiv: 1910.10683, 2019; Adiwardana, Daniel, et al., Towards a human-like open-domain chatbot. CoRR, abs / 2001.09977, 2020; Brown, Tom B, et al. Language models are few-shot learners. arXiv preprint arXiv:2005.14165, 2020; and Vaswani, Ashish, et al, Attention is all you need. Advances in Neural Information Processing Systems (2017).
[0086] In some implementations, the autoregressive transformer neural network 130 has one or more self-attention layers that can apply a masked causal self-attention mechanism along the video segment axis (the temporal axis) when processing the transformer input to generate the updated encoded video segment representation.
[0087] When processing the video segment 102 in the video to generate the updated encoded representation that corresponds to the video segment 102, each of one or more self-attention layers within the autoregressive transformer neural network 130 applies a masked causal selfattention mechanism over the preceding video segments in the video that precede the video segment 102, so that the video segment 102 does not attend over, i.e., the self-attention layer does not generate a non-zero attention weight for, any video segment that temporally follows the video segment 102.
[0088] Put another way, the self-attention layer generates non-zero weights only to the video segment 102 and the video segments that temporally precede the video segment 102 in the video. Masked causal self-attention incorporates information from both current and preceding video segments, allowing the decoder neural network 140 to generate evolving and contextualized outputs as video segments are processed incrementally. The decoder neural network 140 is configured to receive a decoder input and process the decoder input generate an output 104 that characterizes the video segment 102. The output 104 can be any one of the example outputs mentioned above.
[0089] In implementations where both the dimensionality reduction neural network 120 and the autoregressive transformer neural network 130 are not included, the decoder input can include the initial encoded video segment representation that is generated by the video encoder neural network 110.
[0090] Alternatively, in implementations where the dimensionality reduction neural network 120 is included but the autoregressive transformer neural network 130 is not, the decoder input can include the reduced encoded video segment representation that is generated by the dimensionality reduction neural network 120.
[0091] Further alternatively, in implementations where the autoregressive transformer neural network 130 is included, the decoder input can include the updated encoded video segment representation that is generated by the autoregressive transformer neural network 130.
[0092] For example, the decoder neural network 140 can include any appropriate types of neural network layers (e.g., fully connected layers, attention layers, activation layers, and so forth) in any appropriate number (e.g., 2 layers, or 5 layers, or 10 layers) and connected in any appropriate configuration (e.g., as a directed graph of layers).
[0093] In some implementations, the decoder neural network 140 can have a text decoder neural network architecture. For example, the decoder neural network 140 can have one of the text decoder architectures described in Radford, Alec, et al. Learning transferable visual models from natural language supervision. International conference on machine learning. PMLR, 2021, Alayrac, Jean-Baptiste, et al. Flamingo: a visual language model for few-shot learning, Advances in neural information processing systems 35 (2022): 23716-23736, and Yang, Antoine, et al., Vid2seq: Large-scale pretraining of a visual language model for dense video captioning. Proceedings of the IEEE / CVF Conference on Computer Vision and Pattern Recognition. 2023.
[0094] In some implementations, the decoder neural network 140 can operate autoregressively to generate the output 104 as an output sequence of text tokens across multiple output time steps. More specifically, at each output time step, the autoregressively generated output is created by generating the particular text token in the output sequence that corresponds to the output time step conditioned on a current input sequence that includes (i) any text tokens that precede the particular text token in the output sequence, i.e., the tokens that have already been generated in preceding output time steps for any previous positions in the output sequence that precede the particular position of the particular token and (ii) the decoder input.
[0095] FIG. 2 is an example illustration of operations performed by the neural network system 100 of FIG. 1. At a high level, FIG. 2 illustrates that the neural network system 100 receives multiple video segments 102A-E of a video, generates multiple instances 110A-E of the video encoder neural network 110 and multiple instances 140A-E of the decoder neural network 140, and then uses the multiple instances of video encoder neural network 110A-E, the autoregressive transformer neural network 130, and the multiple instances of the decoder neural network 140A-E to perform at least some of the operations in parallel with each other to generate the multiple outputs 104A-E that each characterize a respective video segment of the multiple video segment 102A-E.
[0096] In some implementations, the neural network system 100 can generate as many video encoder (and decoder) neural network instances as needed, e.g., one instance of the encoder neural network for each video segment that has been received. In some other implementations, the neural network system 100 can generate at most a fixed number of video encoder (and decoder) neural networks, e.g., based on the amount of computational resources that is available to the system.
[0097] Each one of the multiple instances of video encoder neural network 110A-E has an identical architecture and identical parameter values (“shared weights”) as another one of the multiple instances of video encoder neural network 110A-E. Likewise, each one of the multiple instances of decoder neural network 140A-E has an identical architecture and identical parameter values (“shared weights”) as another one of the multiple instances of decoder neural network 140A-E.
[0098] The system processes each of the multiple video segments 102A-E using a corresponding one of the multiple instances of video encoder neural network 110A-E to generate multiple initial encoded representations 112A-E that correspond respectively to the multiple video segments 102A-E.
[0099] The multiple instances of video encoder neural network 110A-E are configured to operate in parallel, e.g., such that the generation process of an initial encoded representation 112A that corresponds to a first video segment 102A occurs substantially in parallel with (e.g., at least partly overlap) the generation process of an initial encoded representation 112B that corresponds to a second video segment 102B. Parallelized processing can decrease the overall time needed for encoding the multiple initial video segments 102A-E. Although not illustrated in FIG. 2, in implementations where the dimensionality reduction neural network 120 is included, the neural network system 100 can similarly use multiple instances of the dimensionality reduction neural network 120 to process the multiple initial encoded representations 112A-E in parallel to generate multiple reduced encoded representations that correspond respectively to the multiple video segments 102A-E.
[0100] For each video segment, the neural network system 100 processes a transformer input that includes (i) the initial encoded representation that corresponds to the video segment and (ii) the initial encoded representation that corresponds to each of one or more preceding video segments that temporally precede the video segment in the video, and processes the transformer input using the autoregressive transformer neural network 130 to generate an updated encoded representation that corresponds to the video segment.
[0101] That is, the neural network system 100 uses the autoregressive transformer neural network 130 to update the multiple initial encoded representations 112A-E to generate the updated encoded representations 122A-E that correspond respectively to the multiple video segments 102A-E.
[0102] For example, for the first video segment 102A, because there is no preceding video segment, the transformer input includes the encoded representation 112A that corresponds to the first video segment 102A but not any other initial encoded representations. Thus the updated encoded representations 122A is generated based on the initial encoded representation 112A alone.
[0103] As another example, for the second video segment 102B, the transformer input includes the initial encoded representation 112B that corresponds to the second video segment 102B, and the initial encoded representation 112A that corresponds to the first video segment 102 A, which temporally precede the second video segment 102B. Thus the updated encoded representation 122B is generated based on not only the initial encoded representation 112B, but also the initial encoded representation 112A.
[0104] The neural network system 100 processes each of the multiple updated encoded representations 122A-E using a corresponding one of the multiple instances of decoder neural network 140A-E to generate multiple outputs 104A-E that respectively characterize the multiple video segments 102A-E.
[0105] The multiple instances of decoder neural network 140A-E are configured to operate in parallel, e.g., such that the generation process of an output 104A that characterizes the first video segment 102A occurs substantially in parallel with (e.g., at least partly overlap) the generation process of an output 104B that characterizes the second video segment 102B. Parallelized processing can decrease the overall time needed for generating the outputs 104A- E that respectively characterize the multiple video segments 102A-E.
[0106] In the example of FIG. 2, even though the multiple outputs 104A-E are generated in parallel, because of the way the updated encoded representations 122A-E are generated by the autoregressive transformer neural network 130, each updated encoded representation encodes information about a longer temporal relationship between different video frames at different time steps, and hence the multiple outputs 104A-E will generally include coherent content. For example, in FIG. 2, the first output 104A “A woman is getting a harness put on” is different and yet, related to the second output 104B “She rides up in a yellow cart”.
[0107] FIG. 3 is a flow diagram of an example process 300 for processing a video. For convenience, the process 300 will be described as being performed by a system of one or more computers located in one or more locations. For example, a neural network system, e.g., the neural network system 100 of FIG. 1, appropriately programmed, can perform the process 300.
[0108] The system obtains a video (step 302). The video includes a temporal sequence of video frames that span multiple time steps. Each of the video frames is an image that can be represented as a two-dimensional (2D) array of pixels, where each pixel is represented as a vector of one or more values. For example, the system can receive the video as an upload from a client computing device. As another example, the system can receive the video from a server. As another example, the system can retrieve the video from a remote source.
[0109] The system partitions the temporal sequence of video frames into multiple video segments (step 304). Each video segment includes a respective video frame at each time step in a respective subset of the multiple steps. For example, the video can be divided into video segments based at least in part on a fixed segment length.
[0110] For each video segment in the multiple video segments, the system generates an encoded representation corresponding to the video segment using one or more neural networks (step 306). The one or more neural networks include a video encoder neural network that processes a video encoder input that includes the respective video frames included in the video segment to generate an initial encoded representation that corresponds to the video segment.
[0111] In some implementations, the system includes multiple instances of the video encoder neural network that can operate in parallel, and the system processes each of the multiple video segments using a corresponding one of the multiple instances of video encoder neural network to generate multiple initial encoded representations that correspond respectively to the multiple video segments with at least some degree of parallelization.
[0112] In other words, for each video segment in the multiple video segments, the system processes the respective video frames included in the video segment using an instance of the video decoder neural network to generate the initial encoded representation of the video segment in parallel with processing the respective video frames included in the other video segments using other instances of the video decoder neural network to generate the initial encoded representations of the other video segments.
[0113] Optionally, in some implementations, the one or more neural networks include a dimensionality reduction neural network that processes a reduction input that includes the initial encoded video segment representation to generate a reduced encoded representation that corresponds to the video segment.
[0114] In some of these implementations, the system includes multiple instances of the dimensionality reduction neural network that can operate in parallel, and the system processes each of the multiple initial encoded video segment representations using a corresponding one of the multiple instances of dimensionality reduction neural network to generate multiple reduced encoded representations that correspond respectively to the multiple video segments with at least some degree of parallelization.
[0115] Further optionally, in some implementations, the one or more neural networks include an autoregressive transformer neural network that processes a transformer input that includes (i) the initial encoded representation (in implementations where the dimensionality reduction neural network is not included) or the reduced encoded representation (in implementations where the dimensionality reduction neural network is included) that corresponds to the video segment and, in cases where the video segment is not the first video segment in the video, (ii) the initial encoded representation or the reduced encoded representation that corresponds to each of one or more preceding video segments that temporally precede the video segment in the video, to generate an updated encoded representation that corresponds to the video segment.
[0116] For each video segment in the plurality of video segments, the system processes a decoder input using a decoder neural network to generate an output that characterizes the video segment (step 308).
[0117] In implementations where both the dimensionality reduction neural network and the autoregressive transformer neural network are not included, the decoder input can include the initial encoded video segment representation that is generated by the video encoder neural network.
[0118] Alternatively, in implementations where the dimensionality reduction neural network is included but the autoregressive transformer neural network is not, the decoder input can include the reduced encoded video segment representation that is generated by the dimensionality reduction neural network.
[0119] Alternatively, in implementations where the autoregressive transformer neural network is included, the decoder input can include the updated encoded video segment representation that is generated by the autoregressive transformer neural network.
[0120] In some implementations, the system includes multiple instances of the decoder neural network that can operate in parallel, and the system processes each of the multiple decoder inputs using a corresponding one of the multiple instances of decoder neural network to generate multiple outputs that respectively characterize the multiple video segments with at least some degree of parallelization.
[0121] In other words, for each video segment in the multiple video segments, the system processes the decoder input that corresponds to the video segment using an instance of the decoder neural network to generate the output that characterizes the video segment in parallel with processing the decoder inputs that correspond respectively to other video segments using other instances of the decoder neural network to generate the outputs that respective characterize the other video segments.
[0122] Generally, in implementations where the output for each video segment can include text tokens selected from a vocabulary of text tokens, the decoder neural network can autoregressively generate such an output by generating one text token after another. The output can be any one of the example outputs mentioned above.
[0123] As an example for illustration and not limitation, when the output is an action recognition output, the output can be a sequence of text tokens that has the format of “<start_of_segment> <start_time> <end_time> <action> <end_of_segment>”. In this example, “< start of segment >” and “< end of segment >” are optional, predetermined tokens that represent the beginning and the end of an output for a video segment and that surrounds the output. “<start_time>” and “<end_time>” are optional tokens that identifies a beginning time step and an end time step, respectively, of a recognized action. For example, “<start_time>” can be “lm52s” while “<end_time>” can be “Im 55s”. “<action>” are tokens that define the recognized action. For example, “<action>” can be “A person eats”, “A person walks”, or “A person looks around”. The system can repeat the process 300 to process the videos for which the desired outputs, i.e., the output that should be generated by the system for the video segments included in those videos, are not known.
[0124] The system or another training system can also repeat the process 300 on video segments in a set of training video segments, i.e., a plurality of training video segments for which the ground truth outputs that should be generated by the system are known, in order to train the video encoder neural network and the decoder neural network, and, when included, the dimensionality reduction neural network and the autoregressive transformer neural network, i.e., to determine the trained values for the parameters of these neural networks.
[0125] For example, for each training video segment, the ground truth output that is associated with the training video segment can describe one or more actions depicted in the training video segment; the ground truth output can also identify a beginning time step, an end time step, or both of the training video segment.
[0126] The process 300 can be performed repeatedly on training video segments selected from the plurality of training video segments as part of a gradient-based training technique to train the neural networks based on optimizing a supervised loss function. The supervised loss function includes one or more terms that measure, for each training video segment, the quality of a training output for the training video segment generated by the neural networks, relative to a respective ground truth text output associated with the training video segment. Some common examples of the loss terms include cross entropy loss terms, mean squared error loss terms, negative log likelihood loss terms, to name just a few.
[0127] During training, because the training video segment as well as the outputs that should be generated are known in advance, the computations performed by the neural networks can be accelerated to reduce the amount of time and computing resources (e.g., memory usage) necessary to generate the training outputs by processing the training video segments and, therefore, to decrease the time required for training, to improve the performance of the trained neural network, or both.
[0128] For example, when the decoder neural network includes one or more attention layers, the system can use the decoder neural network to generate the training outputs that respectively characterize a plurality of training video segments in a sequential order while each of the one or more attention layers applies a masked attention mechanism.
[0129] That is, instead of running the decoder neural network T times, i.e., once per video segment, to generate training outputs that respectively characterize T training video segments, the system runs the decoder neural network once to generate a full sequence that includes T training outputs arranged in a sequential order.
[0130] During the one-time processing of the T training video segments by using the decoder neural network, the system masks the attention mechanism used by each of the one or more attention layers such that, for each training video segment, the decoder neural network can only access the decoder input that corresponds to the training video segments, so that the training video segment does not attend over, i.e., the attention layer does not generate a nonzero attention weight for any data that is not included in the decoder input that corresponds to the training video segment.
[0131] FIG. 4 shows another example neural network system 400. The neural network system 400 is an example of a system implemented as computer programs on one or more computers in one or more locations that processes one or more segments of a video and generates outputs that characterize the one or more segments of the video.
[0132] For example, FIG. 4 illustrates that the neural network system 400 receives a video segment 402 that includes multiple video frames, and processes the video segment 402 to generate an output 404 that characterizes the video segment 402. The video segment 402 can be one of multiple video segments included in a video. The output can be any one of the example outputs mentioned above.
[0133] The neural network system 400 includes a plurality of neural networks. The plurality of neural networks include a video encoder neural network 410, a decoder neural network 440, and, optionally, in some implementations, a dimensionality reduction neural network 420 and an autoregressive transformer neural network 430.
[0134] In some implementations, the video encoder neural network 410 can have the same or similar architecture as the video encoder neural network 110 described with reference to FIG. 1. In some implementations, the dimensionality reduction neural network 420, when included, can have the same or similar architecture as the dimensionality reduction neural network 120 described with reference to FIG. 1. In some implementations, the autoregressive transformer neural network 430, when included, can have the same or similar architecture as the autoregressive transformer neural network 130 described with reference to FIG. 1. In some implementations, the decoder neural network 440 can have the same or similar architecture as the decoder neural network 140 described with reference to FIG. 1.
[0135] The neural network system 400 has access to a retrieval dataset 450. The retrieval dataset 450 stores a plurality of embeddings. Each embedding corresponds (or maps) to a respective text sequence. In some implementations, the neural network system 400 can access different retrieval datasets when generating different types of outputs for the same video. Each different retrieval dataset stores a different plurality of embeddings that are generated based on different text sequences. For example, the neural network system 400 can receive, e.g., from a user, an identifier and then use the received identifier to identify the retrieval dataset 450 from a plurality of retrieval datasets that are associated with different identifiers. As another example, the neural network system 400 can automatically determine which retrieval dataset from the plurality of retrieval datasets should be used based on the types of outputs that it is used to generate, or on the types of video segments that it is used to process.
[0136] In some implementations, each respective text sequence includes text in a natural language that characterizes the content (e.g., one or more actions or events) that occur across a temporal sequence of one or more respective prestored video frames. In some examples, each respective text sequence can be one of: a text caption for a prestored video frame, an action recognition output for a prestored video frame, or another text sequence that characterizes one or more aspects of a sequence of one or more prestored video frames.
[0137] As a particular example, when the output 404 that characterizes the video segment 402 is an action recognition output, the retrieval dataset 450 can store a plurality of embeddings that correspond to respective text sequences that are each an action-object phrase. For example, each respective text sequence can be in the format of “<action verb> <target object>”, e.g., “twirling discus”, “seeing replay”, or “throwing disc”.
[0138] In some implementations, the respective text sequences are stored together and in association with the plurality of embeddings in the retrieval dataset 450, whereas, in other implementations, the respective text sequences need not be stored in the retrieval dataset 450. For example, the retrieval dataset 450 stores only the plurality of embeddings, or alternatively stores an identifier together with each of the plurality of embeddings, where the identifier can be used to retrieve the respective text sequence that corresponds to the embedding from a remote location.
[0139] The plurality of embeddings can be generated by the neural network system 400 or another system in many different ways based on a plurality of text sequences.
[0140] In some implementations, the neural network system 400 or another system can generate the embedding for each respective text sequence by processing the respective text sequence using an embedding neural network that includes one or more neural network layers of any appropriate type, or some other machine learning model. The embedding neural network can have any appropriate neural network architecture that allows the embedding neural network to map a text sequence to an embedding. For example, the embedding neural network can include any appropriate types of neural network layers (e.g., embedding layers, fully connected layers, attention layers, and so forth) in any appropriate number (e.g., 2 layers, or 5 layers, or 10 layers) and connected in any appropriate configuration (e.g., as a directed graph of layers).
[0141] Such an embedding neural network can be trained on unlabeled training data based on optimizing a self-supervised or unsupervised loss function to generate vectors in an embedding space that have a fixed dimensionality. In some implementations, the embedding neural network can be trained as part of another neural network (that e.g. has a larger architecture) on tasks that involve generating embedding space representations, e.g., text classification or semantic analysis tasks.
[0142] For example, the neural network system 400 or another system can obtain a plurality of text sequences and then process, using the embedding neural network, each of the plurality of text sequences to generate an embedding that corresponds to the text sequence. For example, the system can receive the plurality of text sequences as an upload from a client computing device. As another example, the system can receive the plurality of text sequences from a server. As another example, the system can retrieve the plurality of text sequences from a remote source.
[0143] In some implementations, in addition to the embedding neural network, the neural network system 400 or another system also makes use of a summarization neural network that includes one or more neural network layers of any appropriate type, or some other machine learning model.
[0144] The summarization neural network can have any appropriate neural network architecture that allows the embedding neural network to map an input text sequence to an output text sequence that is shorter than the input text sequence. For example, the summarization neural network can include any appropriate types of neural network layers (e.g., fully connected layers, attention layers, and so forth) in any appropriate number (e.g., 2 layers, or 5 layers, or 10 layers) and connected in any appropriate configuration (e.g., as a directed graph of layers). Such a summarization neural network can be trained on labeled training data based on optimizing a supervised loss function to perform a summarization task, e.g., an extractive summarization task or an abstractive summarization task, where the input to the summarization neural network is an input natural language text sequence and the output of the summarization neural network is a summary natural language sequence that is shorter than the input sequence but summarizes the input sequence, i.e., represents the most important or relevant information within the input sequence.
[0145] For example, the neural network system 400 or another system can access a video caption dataset that stores a plurality of video segment-text pairs. Each video segment-text pair includes (i) a video segment that includes one or more video frames and (ii) an initial text sequence that is paired with the video segment, e.g., that describes or otherwise characterizes one or more events that occur during the video segment.
[0146] The neural network system 400 or another system can then process each initial text sequence included in the video caption dataset using the summarization neural network to generate a summarized text sequence — and the summarized text sequence will then be processed by using the embedding neural network generate an embedding that corresponds to the initial text sequence.
[0147] It will be appreciated that in other examples, the neural network system 400 or another system can access a different dataset that stores a plurality of image-text pairs, where the text with each pair generally includes a long text sequence that includes many text tokens.
[0148] When processing the video frame 402 in the video, the video encoder neural network 410 is configured to receive a video encoder input that includes the video segment 402 and process the video encoder input to generate an initial encoded representation 452 (also referred to as “an initial encoded video segment representation 452”) that corresponds to the video segment.
[0149] Having generated the initial encoded video segment representation 452 using the video encoder neural network 410, the neural network system 400 performs a search in the retrieval dataset 450 for K most similar embeddings 454 to the initial encoded video segment representation 452 generated by the video encoder neural network according to some similarity measure.
[0150] For some similarity measures, e.g., Euclidean distance or Manhattan distance or other distance measures, the most similar embeddings 454 are those that are closest to the query vector (have the smallest similarity measure with the initial encoded video segment representation 452). For some other similarity measures, e.g., inner product or cosine similarity, the most similar embeddings 454 are those that have the largest similarity measure with the initial encoded video segment representation 452.
[0151] K can generally be any positive integer, i.e., any integer greater than or equal to one, but is generally much smaller than the total number N of embeddings stored in the retrieval dataset 450.
[0152] Then, the neural network system 400 generates the output 404 that characterizes the video segment 402 based on (i) the initial encoded video segment representation 452 and (ii) the one or more most similar embeddings 454 retrieved from the retrieval dataset 450, by using the decoder neural network 440, and, optionally, the dimensionality reduction neural network 420 and the autoregressive transformer neural network 430.
[0153] The relevant external information (in the form of the K most similar embeddings 454) retrieved from the retrieval dataset 450 is used to augment the generation process of the output 404 that characterizes the video segment 402, thereby improving the performance of the neural network system 400 to enable the system to generate higher quality outputs that are more accurate, contextually appropriate, and temporally aligned.
[0154] There are many ways in which the retrieval-augmented output generation process can be performed, and the way retrieval-augmented output generation process is performed may vary from one implementation of the neural network system to another. Generally, however, because of the relevant external information (in the form of the K most similar embeddings 454) retrieved from the retrieval dataset 450, the quality of the outputs generated by these systems can be improved.
[0155] A few examples of performing the retrieval-augmented output generation process will be discussed below.
[0156] FIG. 5 is an example illustration of operations performed by the neural network system 400 of FIG. 4. At a high level, FIG. 5 illustrates that the neural network system 400 receives a video segment 402 of a video, retrieves K most similar embeddings from a retrieval dataset 450, and then generates an output 404 that characterizes the video segment 402 based on the K most similar embeddings.
[0157] The neural network system 400 processes the video segment 402 using the video encoder neural network 410 to generate an initial encoded representation that corresponds to the video segment 402. In the example of FIG. 5, the video segment 402 includes a total of L=3 video frames. Each of the video frames is an image that can be represented as a two- dimensional (2D) array of pixels in the dimension of H * W, where each pixel is represented as a vector of 3 intensity values. The initial encoded representation has the dimension of M * D, where M is the number of vectors and D is the dimension of each vector.
[0158] The neural network system 400 performs a search in the retrieval dataset 450 for K most similar embeddings to the initial encoded video segment representation. For example, the system can use a brute-force search technique, a k nearest neighbors search (a kNN search) technique, an approximate K nearest neighbors search (an approximate kNN search) technique, or another search technique to identify the K most similar embeddings from a larger number of embeddings stored in the retrieval dataset 450.
[0159] In the example of FIG. 5, each embedding stored in the retrieval dataset 450 has the same dimensionality as the initial encoded video segment representation. That is, each embedding stored in the retrieval dataset 450 can be a vector having the dimension of D. However, this is not required. In other examples, each embedding stored in the retrieval dataset 450 may have different dimensions.
[0160] Further, in the example of FIG. 5, in addition to retrieving the K most similar embeddings, the neural network system 400 also retrieves the respective text sequences that corresponds to the K most similar embeddings. However, this is also not required. In other examples, only the K most similar embeddings may need to be retrieved. Nor need the respective text sequences be stored together with the plurality of embeddings in the retrieval dataset 450.
[0161] The neural network system 400 processes the initial encoded video segment representation using the dimensionality reduction neural network 420 to generate a reduced encoded representation that corresponds to the video segment 402. The reduced encoded video segment representation has the dimension of N x D, where N « M. Hence the reduced encoded video segment representation has a reduced dimensionality compared to the initial encoded video segment representation.
[0162] The neural network system 400 generates a retrieval-augmented representation based on the reduced encoded video segment representation and the K most similar embeddings. The retrieval-augmented representation can include K the most similar embeddings, a combined embedding generated based on the K most similar embeddings, or both.
[0163] In the example of FIG. 5, to do this, the neural network system 400 applies a combination function to the K most similar embeddings to generate a combined embedding. In the example of FIG. 5, the combination function is an average pooling function, and the combined embedding is a vector having the dimension of D. However, in other examples, the combination function can be another pooling function (e.g., a global pooling function) or another projection function, and the combined embedding can have a same or different dimension.
[0164] Then, the neural network system 400 generates a combination, e.g., a concatenation or a summation, of the combined embedding and the K most similar embeddings. In the example of FIG. 5, the neural network system 400 concatenates the combined embedding and the K most similar embeddings along the vertical dimension to generate, as the combination, a concatenation having the dimension of (N+l) x D. The combination can then be used as the retrieval -augmented representation.
[0165] The neural network system 400 processes the transformer input that includes (i) the retrieval-augmented representation that corresponds to the video segment 402 and (ii) the retrieval-augmented representation that corresponds to each of one or more preceding video segments that temporally precede the video segment 402 in the video, and processes the transformer input using the autoregressive transformer neural network 430 to generate an updated encoded representation that corresponds to the video segment 402. The updated encoded representation has the same dimensionality as retrieval-augmented representation. That is, updated encoded representation has the dimension of (N+l) x D.
[0166] The neural network system 400 processes a decoder input that includes at least the updated encoded representation using the decoder neural network 440 to generate the output 404 that characterizes the video segment 402.
[0167] In the example of FIG. 5, the decoder input also includes, for each of the K most similar embeddings, the respective text sequence that corresponds to the most similar embedding that has been retrieved from the retrieval dataset 450. For example, FIG. 5 illustrates that the text sequence “baking ham” is provided as part of the decoder input to the decoder neural network 440 for processing. Hence the decoder input can be an enriched, multi-modal input that includes both the text sequences and the updated encoded representation.
[0168] In the example of FIG. 5, the decoder output is a video captioning output that includes a natural language output sequence “[52s-68s] the person places the ham on the pan and bake it.” that is descriptive of the video segment, where “[52s-68s]” identifies the beginning time step and the ending time step during which the described event(s) “the person places the ham on the pan and bake it” occur. However, this is also not required. In other examples, the decoder output can be a different output, e.g., another one of the outputs mentioned above. Nor need the decoder output include data that identifies the beginning and / or ending time steps. FIG. 6 is a flow diagram of an example process 600 for processing a video. For convenience, the process 600 will be described as being performed by a system of one or more computers located in one or more locations. For example, a neural network system, e.g., the neural network system 400 of FIG. 4, appropriately programmed, can perform the process 600.
[0169] The system obtains a video segment that includes a respective video frame at each of multiple time steps (step 602). Each of the video frames is an image that can be represented as a two-dimensional (2D) array of pixels, where each pixel is represented as a vector of one or more values.
[0170] In some implementations, the video segment can one of multiple video segments of a prestored video that includes a longer temporal sequence of video frames that span many time steps. In some implementations, the video segment can be part of a live streaming video that is being received as a stream of video frames.
[0171] For example, the system can receive the video segment, e.g., as part of a prestored video or a live streaming video that includes multiple video segments, as an upload from a client computing device. As another example, the system can receive the video segment from a server. As another example, the system can retrieve the video segment from a remote source, e.g., from a sensor system, which can include a still or video camera.
[0172] The system processes, using a video encoder neural network, a video encoder input that includes the respective video frames included in the video segment to generate an initial encoded representation that corresponds to the video segment (step 604). In some implementations, the video encoder input is in the form of a composite image that is generated based on the respective video frames included in the video segment. Optionally the system can apply resizing to the respective video frames and then generate the composite image based on the resized video frames. For example, each video segment includes 9 video frames, where each video frame has 176 * 176 pixels, the 9 video frames can then be combined into a 3 x 3 grid, resulting in a video encoder input that is a 528 x 528 composite image for each video segment.
[0173] The system performs a search in a retrieval dataset for one or more most similar embeddings to the initial encoded representation according to a similarity measure (step 606). The retrieval dataset stores a plurality of embeddings, where each embedding corresponds to a respective text sequence. For example, the similarity measure can be determined based on one of: a Euclidean distance, a Manhattan distance, an inner product, or a cosine similarity.
[0174] The system generates, based on (i) the initial encoded representation that is generated by the video encoder neural network and (ii) the one or more most similar embeddings obtained from the retrieval dataset, an output that characterizes the video segment (step 608). There are many ways in which step 608 can be performed. A few examples are described in FIGS. 7-10.
[0175] Notably, in some implementations where the video segment is part of a live streaming video that is being received as a stream of video frames, the system can generate the output in an online manner. That is, the system can receive a video in real-time and generate an output that characterize a video segment as new video frames that make up the video segment arrive. Thus, up-to-date outputs that characterize the latest video segments can be continuously generated.
[0176] FIG. 7 is a flow diagram of sub-steps 702-706 of an implementation of step 608 of process 600. In the example of FIG. 7, the system includes the video encoder neural network, an autoregressive transformer neural network, and a decoder neural network. Optionally, the system includes a dimensionality reduction neural network. The system uses a decoder prefixing approach to generate the output that characterizes the video segment.
[0177] In implementations where the dimensionality reduction neural network is included, the system processes, using the dimensionality reduction neural network, a dimensionality reduction input that includes the initial encoded representation to generate a reduced encoded representation (step 702). The reduced encoded representation has a reduced dimensionality compared to the initial encoded representation.
[0178] The system processes, using the autoregressive transformer neural network, a transformer input that includes (i) the initial encoded representation (in implementations where the dimensionality reduction neural network is not included) or the reduced encoded representation (in implementations where the dimensionality reduction neural network is included) that corresponds to the video segment and, in cases where the video segment is not the first video segment in the video, (ii) the initial encoded representation or the reduced encoded representation that corresponds to each of one or more preceding video segments that temporally precede the video segment in the video, to generate an updated encoded representation that corresponds to the video segment (step 704).
[0179] The system processes, using the decoder neural network, a decoder input that includes (i) the updated encoded representation, and (ii) the respective text sequences that correspond to the one or more most similar embeddings, or (iii) the one or more most similar embeddings, or both (ii) and (iii), to generate the output that characterizes the video segment (step 706). That is, the decoder neural network generates the output conditioned not only on (i) the updated encoded representation, but also on (ii) the respective text sequences that correspond to the one or more most similar embeddings, (iii) the one or more most similar embeddings, or both (ii) and (iii).
[0180] In this example, the decoder prefixing approach improves the accuracy and relevance (with respect to the video segment) of the output by providing the neural networks with information about the content (e.g., the key actions, objects, or events that occur) within one or more prestored video frames that are similar to the video frames included in the video segment. Such information can be represented either in the form of natural language text, embeddings, or a combination of natural language text and embeddings.
[0181] FIG. 8 is a flow diagram of sub-steps 802-808 of another implementation of step 608 of process 600. In the example of FIG. 8, the system includes the video encoder neural network, an autoregressive transformer neural network, and a decoder neural network. Optionally, the system includes a dimensionality reduction neural network. The system uses an embedding fusion approach to generate the output that characterizes the video segment.
[0182] In implementations where the dimensionality reduction neural network is included, the system processes, using the dimensionality reduction neural network, a dimensionality reduction input that includes the initial encoded representation to generate a reduced encoded representation (step 802). The reduced encoded representation has a reduced dimensionality compared to the initial encoded representation.
[0183] The system generates a transformer input that corresponds to the video segment based on (i) the initial encoded representation (in implementations where the dimensionality reduction neural network is not included) or the reduced encoded representation (in implementations where the dimensionality reduction neural network is included) and (ii) the one or more most similar embeddings (step 804).
[0184] For example, the system can combine, e.g., concatenate or add, the initial encoded representation or the reduced encoded representation and the one or more most similar embeddings, and then use the combination as the transformer input.
[0185] As another example, the system can generate a pooled embedding from the one or more most similar embeddings by applying a pooling function to the one or more most similar embeddings, and then concatenate the initial encoded representation or the reduced encoded representation to the pooled embedding. The concatenate can then be used as the transformer input. As a similar example, the system can generate a combined embedding from the one or more most similar embeddings by applying a combination function, e.g., a projection function, to the one or more most similar embeddings, and then concatenate the initial encoded representation or the reduced encoded representation to the combined embedding. The concatenate can then be similarly used as the transformer input.
[0186] The system processes, using the autoregressive transformer neural network, the transformer input that corresponds to the video segment and, in cases where the video segment is not the first video segment in the video, (ii) the transformer input that corresponds to each of one or more preceding video segments that temporally precede the video segment in the video, to generate an updated encoded representation that corresponds to the video segment (step 806).
[0187] The system processes, using the decoder neural network, a decoder input that includes the updated encoded representation to generate the output that characterizes the video segment (step 808). That is, the updated encoded representation is used as the decoder input, and the decoder neural network generates the output conditioned on the updated encoded representation.
[0188] In this example, the embedding fusion approach improves the quality of the output by generating a comprehensive representation of each video segment. The comprehensive representation includes the initial or reduced encoded representation and the one or more most similar embeddings; such a comprehensive representation is then processed by the autoregressive Transformer neural network to generate an updated encoded representation, enabling the updated encoded representation to incorporate both visual and retrieved text information for the video segment, along with such information extracted from preceding video segments.
[0189] Because the comprehensive representation is a temporally and causally aware representation of the video segment, the embedding fusion approach can be advantageous in cases where the output is a video captioning output, where understanding the temporal span and boundaries of events is critical.
[0190] FIG. 9 is a flow diagram of sub-steps 902-908 of another implementation of step 608 of process 600. In the example of FIG. 9, the system includes the video encoder neural network, an autoregressive transformer neural network, and a decoder neural network. Optionally, the system includes a dimensionality reduction neural network. The system uses a combined approach that combines decoder prefixing with embedding fusion to generate the output that characterizes the video segment. In implementations where the dimensionality reduction neural network is included, the system processes, using the dimensionality reduction neural network, a dimensionality reduction input that includes the initial encoded representation to generate a reduced encoded representation (step 902). The reduced encoded representation has a reduced dimensionality compared to the initial encoded representation.
[0191] The system generates a transformer input that corresponds to the video segment based on (i) the initial encoded representation (in implementations where the dimensionality reduction neural network is not included) or the reduced encoded representation (in implementations where the dimensionality reduction neural network is included) and (ii) the one or more most similar embeddings (step 904). A few example ways of how the transformer input can be generated based on (i) and (ii) are described above with reference to step 804 of FIG. 8.
[0192] The system processes, using the autoregressive transformer neural network, the transformer input that corresponds to the video segment and, in cases where the video segment is not the first video segment in the video, (ii) the transformer input that corresponds to each of one or more preceding video segments that temporally precede the video segment in the video, to generate an updated encoded representation that corresponds to the video segment (step 906).
[0193] The system processes, using the decoder neural network, a decoder input that includes (i) the updated encoded representation and (ii) the respective text sequences that correspond to the one or more most similar embeddings or (iii) the one or more most similar embeddings or both (ii) and (iii) to generate the output that characterizes the video segment (step 908).
[0194] That is, the decoder neural network generates the output conditioned not only on (i) the updated encoded representation, but also on (ii) the respective text sequences that correspond to the one or more most similar embeddings, (iii) the one or more most similar embeddings, or both (ii) and (iii). Note that the updated encoded representation, in turn, is generated based at least on the one or more most similar embeddings.
[0195] The combined approach has the benefits of both decoder prefixing and embedding fusion approaches. In cases where the video is a continuously streaming video, the combined approach enhances the neural networks’ capability to generate accurate and temporally aligned outputs by maintaining awareness of immediate actions (or events) in the video and the evolving context throughout the video that is being continuously streamed.
[0196] FIG. 10 is a flow diagram of sub-steps 1002-1004 of another implementation of step 608 of process 600. In the example of FIG. 10, the system includes the video encoder neural network and a decoder neural network; the system uses a decoder prefixing approach to generate the output that characterizes the video segment.
[0197] The system generates a decoder input that includes (i) the initial encoded representation and (ii) the respective text sequences that correspond to the one or more most similar embeddings (step 1002).
[0198] The system processes, using the decoder neural network, the decoder input to generate the output that characterizes the video segment (step 1004). In this example, the retrieved text sequences can be added as a natural language prefix to the decoder neural network, when it is processing the initial encoded representation to generate the output.
[0199] In any of these examples described with reference to FIGS. 7-10, in implementations where the output for each video segment can include text tokens selected from a vocabulary of text tokens, the decoder neural network can autoregressively generate such an output by generating one text token after another. The output can be any one of the example outputs mentioned above.
[0200] The system can repeat the process 600 to process videos for which the desired outputs, i.e., the output that should be generated by the system for the video segments included in those videos, are not known.
[0201] The system or another training system can also repeat the process 600 on video segments in a set of training video segments, i.e., a plurality of training video segments for which the ground truth outputs that should be generated by the system are known, in order to train the video encoder neural network and the decoder neural network, and, when included, the dimensionality reduction neural network and the autoregressive transformer neural network, i.e., to determine the trained values for the parameters of these neural networks.
[0202] For example, for each training video segment, the ground truth output that is associated with the training video segment can describe one or more actions depicted in the training video segment; the training video segment can also identify a beginning time step, an end time step, or both of the training video segment.
[0203] The process 600 can be performed repeatedly on training video segments selected from the plurality of training video segments as part of a gradient-based training technique to train the neural networks based on optimizing a supervised loss function. The supervised loss function includes one or more terms that measure, for each training video segment, the quality of a training output for the training video segment generated by the neural networks, relative to a respective ground truth text output associated with the training video segment. Some common examples of the loss terms include cross entropy loss terms, mean squared error loss terms, negative log likelihood loss terms, to name just a few.
[0204] During training, the system can incorporate any number of techniques to improve the speed, the effectiveness, or both of the training process. For example, the system can use a mixed training strategy that alternates between retrieval-augmented and non-retrieval- augmented training to enhance the neural networks’ adaptability to inference settings without retrieval augmentation, e.g., when a retrieval dataset is not available to the system. The mixed training strategy improves the non-augmented inference performance of the neural networks while maintaining strong inference performance when augmentation is available.
[0205] To apply the mixed training strategy, the system samples a first training video segment and a second training video segment from the plurality of training video segments. The system processes, using the video encoder neural network and the decoder neural network, (i) the first training video segment and (ii) one or more most similar embeddings obtained from the retrieval dataset to generate a first training output. In addition, the system processes, using the video encoder neural network and the decoder neural network, the second training video segment to generate a second training output, without processing any embeddings obtained from the retrieval dataset. Then, the system determines an update to values of parameters of the video encoder neural network and the decoder neural network and, when included, the dimensionality reduction neural network and the autoregressive transformer neural network, based on a difference between the first training output and a first ground truth output associated with the first training video segment, and on a difference between the second training output and a second ground truth output associated with the second training video segment.
[0206] FIG. 11 shows a quantitative example of the performance gains that can be achieved by the neural network system 100 of FIG. 1 on a video captioning task (evaluated on the ActivityNet dense captioning dataset) compared to existing video processing systems, including the PDVC system (described in Teng Wang, et al., End-to-end dense video captioning with parallel decoding. In ICCV, 2021), the Vid2Seq system (described in Antoine Yang, et al., Vid2seq: Large-scale pretraining of a visual language model for dense video captioning. CVPR, 2023), and the MT system (described in Luowei Zhou, et al., End- to-end dense video captioning with masked transformer. CVPR, 2018).
[0207] In the table shown in FIG. 11, the “pretr.” column specifies whether the system has been pre-trained and if so, on which dataset the system has been pre-trained, the “backbone” column specifies the architecture of a backbone neural network included in the system, the “modalities” column specifies the modalities of the data included in the input processed by the system, the “S” column specifies the SODA score (described in Soichiro Fujita, et al., Soda: Story oriented dense video captioning evaluation framework. In Computer Vision- ECCV 2020: 16th European Conference, Glasgow, UK, August 23-28, 2020, Proceedings, Part VI 16, pages 517-531. Springer, 2020) of the video captioning outputs generated by the system, the “C” column specifies the CiDER score (described in Ramakrishna Vedantam, et al., Cider: Consensus-based image description evaluation. In Proceedings of the IEEE conference on computer vision and pattern recognition, pages 4566-4575, 2015) of the video captioning outputs generated by the system, the “M” column specifies the METEOR score (described in Satanjeev Baneijee, at al., Meteor: An automatic metric for mt evaluation with improved correlation with human judgments. In Proceedings of the acl workshop on intrinsic and extrinsic evaluation measures for machine translation and / or summarization, pages 65- 72, 2005) of the video captioning outputs generated by the system. Higher scores indicate better performance on the video captioning task. It will be appreciated that the neural network system 100 of FIG. 1 has better performance compared to these existing video processing systems, even without pre-training.
[0208] FIG. 12 shows a quantitative example of the performance gains that can be achieved by the neural network system 400 of FIG. 4 on a video captioning task (evaluated on the ViTT dataset, the YouCook2 dataset, and ActivityNet dense captioning dataset) compared to existing video processing systems, including the E2ESG system (described in Wanrong Zhu, et al., End-to-end dense video captioning as sequence generation. In COLING, 2022), the MT system (described in Luowei Zhou, et al. End-to-end dense video captioning with masked transformer. CVPR, 2018), the PDVC system (described in Teng Wang, et al. End-to-end dense video captioning with parallel decoding. In ICCV, 2021), the GIT system (Jianfeng Wang, et al., Git: A generative image-to-text transformer for vision and language. arXiv preprint arXiv:2205.14100, 2022), the OmniViD system (described in Junke Wang, et al., Omnivid: A generative framework for universal video understanding. In Proceedings of the IEEE / CVF Conference on Computer Vision and Pattern Recognition, pp. 18209-18220, 2024), the TimeChat system (described in Shuhuai Ren, et al., Timechat: A time-sensitive multimodal large language model for long video understanding. In Proceedings of the IEEE / CVF Conference on Computer Vision and Pattern Recognition, pp. 14313-14323, 2024), the Vid2Seq system (described in Antoine Yang, et al. Vid2seq: Large-scale pretraining of a visual language model for dense video captioning. CVPR, 2023), the DoYou system (described in Minkuk Kim, et al., Do you remember? dense video captioning with cross-modal memory retrieval. In Proceedings of the IEEE / CVF Conference on Computer Vision and Pattern Recognition, pp. 13894-13904, 2024), the DIBS system (described in Hao Wu, et al., Dibs: Enhancing dense video captioning with unlabeled videos via pseudo boundary enrichment and online refinement. In Proceedings of the IEEE / CVF Conference on Computer Vision and Pattern Recognition, pp. 18699-18708, 2024), and the streaming system (described in Xingyi Zhou, et al., Streaming dense video captioning. In Proceedings of the IEEE / CVF Conference on Computer Vision and Pattern Recognition, pp. 18243- 18252, 2024).
[0209] In the table shown in FIG. 12, the “backbone” column specifies the architecture of a backbone neural network included in the system. Higher scores indicate better performance on the video captioning task. It will be appreciated that the neural network system 400 of FIG. 4 has better performance compared to these existing video processing systems.
[0210] This specification uses the term “configured” in connection with systems and computer program components. For a system of one or more computers to be configured to perform particular operations or actions means that the system has installed on it software, firmware, hardware, or a combination of them that in operation cause the system to perform the operations or actions. For one or more computer programs to be configured to perform particular operations or actions means that the one or more programs include instructions that, when executed by data processing apparatus, cause the apparatus to perform the operations or actions.
[0211] Embodiments of the subject matter and the functional operations described in this specification can be implemented in digital electronic circuitry, in tangibly-embodied computer software or firmware, in computer hardware, including the structures disclosed in this specification and their structural equivalents, or in combinations of one or more of them. Embodiments of the subject matter described in this specification can be implemented as one or more computer programs, i.e., one or more modules of computer program instructions encoded on a tangible non transitory storage medium for execution by, or to control the operation of, data processing apparatus. The computer storage medium can be a machine- readable storage device, a machine-readable storage substrate, a random or serial access memory device, or a combination of one or more of them. Alternatively or in addition, the program instructions can be encoded on an artificially generated propagated signal, e.g., a machine-generated electrical, optical, or electromagnetic signal, that is generated to encode information for transmission to suitable receiver apparatus for execution by a data processing apparatus. The term “data processing apparatus” refers to data processing hardware and encompasses all kinds of apparatus, devices, and machines for processing data, including by way of example a programmable processor, a computer, or multiple processors or computers. The apparatus can also be, or further include, special purpose logic circuitry, e.g., an FPGA (field programmable gate array) or an ASIC (application specific integrated circuit). The apparatus can optionally include, in addition to hardware, code that creates an execution environment for computer programs, e.g., code that constitutes processor firmware, a protocol stack, a database management system, an operating system, or a combination of one or more of them.
[0212] A computer program, which may also be referred to or described as a program, software, a software application, an app, a module, a software module, a script, or code, can be written in any form of programming language, including compiled or interpreted languages, or declarative or procedural languages; and it can be deployed in any form, including as a stand alone program or as a module, component, subroutine, or other unit suitable for use in a computing environment. A program may, but need not, correspond to a file in a file system. A program can be stored in a portion of a file that holds other programs or data, e.g., one or more scripts stored in a markup language document, in a single file dedicated to the program in question, or in multiple coordinated files, e.g., files that store one or more modules, sub programs, or portions of code. A computer program can be deployed to be executed on one computer or on multiple computers that are located at one site or distributed across multiple sites and interconnected by a data communication network.
[0213] In this specification, the term “database” is used broadly to refer to any collection of data: the data does not need to be structured in any particular way, or structured at all, and it can be stored on storage devices in one or more locations. Thus, for example, the index database can include multiple collections of data, each of which may be organized and accessed differently.
[0214] Similarly, in this specification the term “engine” is used broadly to refer to a software-based system, subsystem, or process that is programmed to perform one or more specific functions. Generally, an engine will be implemented as one or more software modules or components, installed on one or more computers in one or more locations. In some cases, one or more computers will be dedicated to a particular engine; in other cases, multiple engines can be installed and running on the same computer or computers.
[0215] The processes and logic flows described in this specification can be performed by one or more programmable computers executing one or more computer programs to perform functions by operating on input data and generating output. The processes and logic flows can also be performed by special purpose logic circuitry, e.g., an FPGA or an ASIC, or by a combination of special purpose logic circuitry and one or more programmed computers.
[0216] Computers suitable for the execution of a computer program can be based on general or special purpose microprocessors or both, or any other kind of central processing unit. Generally, a central processing unit will receive instructions and data from a read only memory or a random access memory or both. The essential elements of a computer are a central processing unit for performing or executing instructions and one or more memory devices for storing instructions and data. The central processing unit and the memory can be supplemented by, or incorporated in, special purpose logic circuitry. Generally, a computer will also include, or be operatively coupled to receive data from or transfer data to, or both, one or more mass storage devices for storing data, e.g., magnetic, magneto optical disks, or optical disks. However, a computer need not have such devices. Moreover, a computer can be embedded in another device, e.g., a mobile telephone, a personal digital assistant (PDA), a mobile audio or video player, a game console, a Global Positioning System (GPS) receiver, or a portable storage device, e.g., a universal serial bus (USB) flash drive, to name just a few.
[0217] Computer readable media suitable for storing computer program instructions and data include all forms of non volatile memory, media and memory devices, including by way of example semiconductor memory devices, e.g., EPROM, EEPROM, and flash memory devices; magnetic disks, e.g., internal hard disks or removable disks; magneto optical disks; and CD ROM and DVD-ROM disks.
[0218] To provide for interaction with a user, embodiments of the subject matter described in this specification can be implemented on a computer having a display device, e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor, for displaying information to the user and a key vectorboard and a pointing device, e.g., a mouse or a trackball, by which the user can provide input to the computer. Other kinds of devices can be used to provide for interaction with a user as well; for example, feedback provided to the user can be any form of sensory feedback, e.g., visual feedback, auditory feedback, or tactile feedback; and input from the user can be received in any form, including acoustic, speech, or tactile input. In addition, a computer can interact with a user by sending documents to and receiving documents from a device that is used by the user; for example, by sending web pages to a web browser on a user’s device in response to requests received from the web browser. Also, a computer can interact with a user by sending text messages or other forms of message to a personal device, e.g., a smartphone that is running a messaging application, and receiving responsive messages from the user in return.
[0219] Data processing apparatus for implementing machine learning models can also include, for example, special-purpose hardware accelerator units for processing common and compute-intensive parts of machine learning training or production, i.e., inference, workloads.
[0220] Machine learning models can be implemented and deployed using a machine learning framework, e.g., a TensorFlow framework or a JAX framework.
[0221] Embodiments of the subject matter described in this specification can be implemented in a computing system that includes a back end component, e.g., as a data server, or that includes a middleware component, e.g., an application server, or that includes a front end component, e.g., a client computer having a graphical user interface, a web browser, or an app through which a user can interact with an implementation of the subject matter described in this specification, or any combination of one or more such back end, middleware, or front end components. The components of the system can be interconnected by any form or medium of digital data communication, e.g., a communication network. Examples of communication networks include a local area network (LAN) and a wide area network (WAN), e.g., the Internet.
[0222] The computing system can include clients and servers. A client and server are generally remote from each other and typically interact through a communication network. The relationship of client and server arises by virtue of computer programs running on the respective computers and having a client-server relationship to each other. In some embodiments, a server transmits data, e.g., an HTML page, to a user device, e.g., for purposes of displaying data to and receiving user input from a user interacting with the device, which acts as a client. Data generated at the user device, e.g., a result of the user interaction, can be received at the server from the device.
[0223] While this specification contains many specific implementation details, these should not be construed as limitations on the scope of any invention or on the scope of what may be claimed, but rather as descriptions of features that may be specific to particular embodiments of particular inventions. Certain features that are described in this specification in the context of separate embodiments can also be implemented in combination in a single embodiment. Conversely, various features that are described in the context of a single embodiment can also be implemented in multiple embodiments separately or in any suitable subcombination. Moreover, although features may be described above as acting in certain combinations and even initially be claimed as such, one or more features from a claimed combination can in some cases be excised from the combination, and the claimed combination may be directed to a subcombination or variation of a subcombination.
[0224] Similarly, while operations are depicted in the drawings and recited in the claims in a particular order, this should not be understood as requiring that such operations be performed in the particular order shown or in sequential order, or that all illustrated operations be performed, to achieve desirable results. In certain circumstances, multitasking and parallel processing may be advantageous. Moreover, the separation of various system modules and components in the embodiments described above should not be understood as requiring such separation in all embodiments, and it should be understood that the described program components and systems can generally be integrated together in a single software product or packaged into multiple software products.
[0225] Particular embodiments of the subject matter have been described. Other embodiments are within the scope of the following claims. For example, the actions recited in the claims can be performed in a different order and still achieve desirable results. As one example, the processes depicted in the accompanying figures do not necessarily require the particular order shown, or sequential order, to achieve desirable results. In some cases, multitasking and parallel processing may be advantageous.
Claims
CLAIMS1. A method performed by one or more computers, the method comprising: obtaining a temporal sequence of video frames that span a plurality of time steps; partitioning the temporal sequence of video frames into a plurality of video segments, wherein each video segment comprises a respective video frame at each time step in a respective subset of the plurality time steps; and for each video segment in the plurality of video segments: generating an encoded representation corresponding to the video segment using one or more neural networks to process respective video frames included in the video segment and in any video segments that precede the video segment in the plurality of video segments; and generating, using a text decoder neural network and from the encoded representation corresponding to the video segment, a text output that characterizes the video segment in parallel with generating text outputs that respectively characterize other video segments in the plurality of video segments using the text decoder neural network and from encoded representations corresponding respectively to the other video segments.
2. The method of claim 1, wherein the one or more neural networks comprise a video encoder neural network that processes the respective video frames included in the video segment to generate a first initial encoded representation of the video segment.
3. The method of claim 2, wherein processing the respective video frames included in the video segment to generate the first initial encoded representation of the video segment comprises: processing the respective video frames included in the video segment to generate the first initial encoded representation of the video segment in parallel with processing the respective video frames included in the other video segments to generate the first initial encoded representations of the other video segments using the video decoder neural network.
4. The method of any one of claims 1-3, wherein the one or more neural networks comprise a video memory transformer neural network that generates a second initial encoded representation of the video segment based on the first initial encoded representation of the video segment, wherein the generating comprises: obtaining a video memory transformer network input that comprises (i) a set of memory vectors and (ii) a set of feature vectors selected from the first initial encoded representation; and processing the video memory transformer network input using the video memory transformer neural network to generate the second initial encoded representation of the video segment based on applying an attention mechanism over (i) the set of memory vectors and (ii) the set of feature vectors to determine (i) a set of updated memory vectors and (ii) a set of updated feature vectors.
5. The method of any one of claims 1-4, wherein the one or more neural networks comprise an autoregressive transformer neural network that generates the encoded representation corresponding to the video segment based on the second initial encoded representation of the video segment, wherein the generating comprises, for each video segment in the plurality of video segments: obtaining an autoregressive transformer network input that comprises (i) the second initial encoded representation of the video segment and (ii) the second initial encoded representations of any video segments that precede the video segment in the plurality of video segments; and processing the autoregressive transformer network input using the autoregressive transformer neural network to generate the encoded representation of the video segment based on applying an attention mechanism over (i) the second initial encoded representation of the video segment and (ii) the second initial encoded representations of any video segments that precede the video segment in the plurality of video segments.
6. The method of any one of claims 1-5, wherein generating the text output for each video segment in parallel with generating text outputs that respectively characterize the other video segments comprises: generating a plurality of instances of the text decoder neural network that correspond respectively to the plurality of video segments and that have identical parameter values with each other.
7. The method of claim 6, wherein generating the text output for each video segment in parallel with generating text outputs that respectively characterize the other video segments comprises: processing, by each of the plurality of instances of the text decoder neural network, a different encoded representation in accordance with the identical parameter values to generate the text output for each video segment.
8. The method of any one of claims 1-7, wherein text decoder neural network is an autoregressive neural network that generates the text output by, at each of multiple output steps, generating a text token conditioned on any text tokens that have already been generated in preceding output steps.
9. The method of any one of claims 1-8, wherein the text output comprises one of: a text caption output for the video segment, an action recognition output for the video segment, or an action localization output for the video segment.
10. The method of any one of claims 1-9, further comprising training the one or more neural networks and the text decoder neural network on a plurality of training video segments that are each associated with a ground truth text output, wherein for each training video segment, the ground truth text output identifies a beginning time step, an end time step, or both of the training video segment, and describes one or more actions depicted in the training video segment.
11. The method of claim 10, wherein during training, the text decoder neural network generates training text outputs that respectively characterize the plurality of training video segments in a sequential order by using a masked attention mechanism.
12. A system comprising: one or more computers; and one or more storage devices storing instructions that, when executed by the one or more computers, cause the one or more computers to perform the respective operations of any one of claims 1-11.
13. One or more computer-readable storage media storing instructions that when executed by one or more computers cause the one or more computers to perform the respective operations of the method of any one of claims 1-11.
14. A method performed by one or more computers, the method comprising: obtaining a video segment that comprises a respective video frame at each of a plurality of time steps; processing, using a video encoder neural network, a video encoder input that includes the respective video frames included in the video segment to generate an initial encoded video segment representation that corresponds to the video segment; performing a search in a retrieval dataset that stores a plurality of embeddings for one or more most similar embeddings to the initial encoded video segment representation according to a similarity measure, wherein each embedding corresponds to a respective text sequence; and generating, based on (i) the initial encoded video segment representation and (ii) the one or more most similar embeddings, an output that characterizes the video segment.
15. The method of claim 14, wherein generating, based on (i) the initial encoded video segment representation and (ii) the one or more most similar embeddings, an output that characterizes the video segment comprises: processing, using a dimensionality reduction neural network, a first dimensionality reduction input that includes the initial encoded video segment representation to generate a first reduced encoded video segment representation; and processing, using a decoder neural network, a first decoder input that includes (i) data derived from the first reduced encoded video segment representation and (ii) the respective text sequences that correspond to the one or more most similar embeddings or (iii) the one or more most similar embeddings or both (ii) and (iii) to generate the output that characterizes the video segment.
16. The method of claim 14, wherein generating, based on (i) the initial encoded video segment representation and (ii) the one or more most similar embeddings, an output that characterizes the video segment comprises: processing, using a dimensionality reduction neural network, a second dimensionality reduction input that includes the initial encoded video segment representation to generate a second reduced encoded video segment representation; generating a second decoder input based on (i) the second reduced encoded video segment representation and (ii) the one or more most similar embeddings; and processing, using a decoder neural network, the second decoder input to generate the output that characterizes the video segment.
17. The method of claim 16, wherein generating the second decoder input based on (i) the second reduced encoded video segment representation and (ii) the one or more most similar embeddings comprises: generating, from the one or more most similar embeddings, a first pooled embedding; generating a first updated encoded video segment representation that corresponds to the video segment based on processing at least (i) the second reduced encoded video segment representation and (ii) the first pooled embedding using the autoregressive transformer neural network; and using the first updated encoded video segment representation as the second decoder input.
18. The method of claim 14, wherein generating, based on (i) the initial encoded video segment representation and (ii) the one or more most similar embeddings, an output that characterizes the video segment comprises: processing, using a dimensionality reduction neural network, a third dimensionality reduction input that includes the initial encoded video segment representation to generate a third reduced encoded video segment representation; generating a second updated encoded video segment representation that corresponds to the video segment based on (i) the third reduced encoded video segment representation and (ii) the one or more most similar embeddings; and processing, using a decoder neural network, a third decoder input that includes (i) the second updated encoded video segment representation and (ii) the respective text sequences that correspond to the one or more most similar embeddings or (iii) the one or more most similar embeddings or both (ii) and (iii) to generate the output that characterizes the video segment.
19. The method of claim 18, wherein generating the second updated encoded video segment representation comprises: generating, from the one or more most similar embeddings, a second pooled embedding; generating the second updated encoded video segment representation that corresponds to the video segment based on processing at least (i) the third reduced encoded video segment representation and (ii) the second pooled embedding using the autoregressive transformer neural network.
20. The method of any one of claims 14-19, wherein the respective text sequences that correspond to the plurality of embeddings stored in the retrieval dataset are generated by, for each embedding: obtaining the respective text sequence, wherein the respective text sequence characterizes one or more respective prestored video frames; and processing, using an embedding neural network, the respective text sequence to generate the embedding.
21. The method of claim 20, wherein obtaining the respective text sequence comprises: obtaining, from a video caption dataset, an initial text sequence that characterizes the one or more respective prestored video frames; and processing, using a summarization neural network, the initial text sequence to generate the respective sequence.
22. The method of any one of claims 14-21, wherein the output comprises one of: a text caption output for the video segment, an action recognition output for the video segment, or an action localization output for the video segment.
23. The method of any one of claims 14-22, wherein the video segment follows one or more preceding video segments that each comprise a respective video frame at each of a plurality of preceding time steps, and wherein generating the output that characterizes the video segment comprises generating the output based on one or more initial encoded video segment representations that correspond to the one or more preceding video segments.
24. The method of any one of claims 14-23, wherein the video encoder input is a composite image that is generated based on applying frame resizing to the respective video frames included in the video segment.
25. The method of any one of claims 15-24, further comprising training the video encoder neural network and the decoder neural network on a plurality of training video segments that are each associated with a ground truth output, wherein for each training video segment, the ground truth output identifies a beginning time step, an end time step, or both of the training video segment, and describes one or more actions depicted in the training video segment.
26. The method of claim 25, wherein the training comprises: processing, using the video encoder neural network and the decoder neural network, (i) a first training video segment and (ii) one or more most similar embeddings obtained from the retrieval dataset to generate a first training output; processing, using the video encoder neural network and the decoder neural network, a second training video segment to generate a second training output; and determining an update to values of parameters of the video encoder neural network and the decoder neural network based on a difference between the first training output and a first ground truth output associated with the first training video segment, and on a differencebetween the second training output and a second ground truth output associated with the second training video segment.
27. A system comprising: one or more computers; and one or more storage devices storing instructions that, when executed by the one or more computers, cause the one or more computers to perform the respective operations of any one of claims 14-26.
28. One or more computer-readable storage media storing instructions that when executed by one or more computers cause the one or more computers to perform the respective operations of the method of any one of claims 14-26.