Visual TRANSFORMER with sparse application of video cores
Through sparse video tube technology, the problem of low computing efficiency of transformer model in video understanding tasks is solved, seamless fusion of images and videos and optimization of computing resources is achieved, and the performance of video understanding is improved.
Patent Information
- Application Number
- CN202380080607.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2022-11-22
- Filing Date
- 2023-11-22
- Publication Date
- 2025-07-01
AI Technical Summary
Existing transformer models are incomputed in video comprehension tasks and are difficult to effectively process long videos. Existing methods usually treat images and videos as independent inputs, resulting in waste of computing resources and insufficient performance.
The sparse video tube technology is adopted to sample videos by sparsely applying 3D space-time tubes, generate video marks, and use visual transformers to process them to achieve seamless fusion of images and videos and improve computing efficiency.
It realizes the reduction of computing resources and performance improvements in image and video tasks, and can better understand the action and spatiotemporal information in video, and is suitable for various computer vision tasks.
Smart Images

Figure CN120239875A_ABST
Abstract
Description
[0001] Related Applications
[0002] This application claims the benefit and priority of U.S. Provisional Patent Application No. 63 / 427,238, filed Nov. 22, 2022. U.S. Provisional Patent Application No. 63 / 427,238 is hereby incorporated by reference in its entirety. Technical Field
[0003] The present disclosure generally relates to machine learning. More specifically, the present disclosure relates to using a vision transformer in combination with a sparse application of a video kernel for joint image and video learning. Background Art
[0004] Transformer models are a type of machine learning model that utilize self-attention mechanisms on sequences of tokens or embeddings at each of multiple layers. An example “vanilla” transformer architecture is described in the following document: Vaswani et al., Attention is all you need, in NeurIPS, 2017.
[0005] Although transformer models were initially applied to natural language settings, they have since been adapted and widely applied to single-image vision tasks such as image classification. Transformers applied to vision tasks may be referred to as visual transformers or vision transformers. An example transformer model configured to process a single image is the ViT model described in the following document: Dosovitskiy et al., An image is worth 16x16 words: Transformers for image recognition at scale, in ICLR, 2021. When applied to vision tasks, transformers have become a ubiquitous backbone for visual representation learning, leading to many advances in areas such as image understanding, multimodal tasks, and self-supervised learning.
[0006] Video understanding is a fundamental computer vision task. However, adapting transformer models to video is both challenging and computationally intensive. Thus, video versions of transformer models have been specifically designed to handle larger numbers of frames.
[0007] In particular, due to the quadratic cost and dense sampling of self-attention, the use of transformers for videos has required different elements, such as space-time factorized attention. However, these video transformers have not been truly tested on longer videos as most have been evaluated on short clips. The ability to handle a larger number of input frames and understand long-term actions and their relationships is crucial but computationally becomes prohibitive for current models.
[0008] Other work has investigated ways to reduce the number of tokens in video transformer models. However, all of this work still uses the initial dense sampling of the video and then uses some heuristics to reduce the number of inputs. Thus, there is a desire in the art for a transformer model architecture that enables the processing of videos with improved efficiency.
[0009] In addition, most prior work has treated images and videos as completely different inputs and thus provided separate methods for videos or images, as designing a model that can handle both is challenging. For example, some prior methods for co-training images and videos adapt the architecture to do so, where significant portions of the network are designed separately for each input. As another example, some work resamples the input and compresses it into a fixed number of features. However, such resampling can still be expensive for long videos, and some approaches in this vein treat the video as individual frames sampled at 1 FPS, which limits the temporal information. For datasets that rely on motion and temporal understanding or for identifying fast and short actions, such low FPS sampling and per-frame modeling are generally insufficient. On the other hand, using one of the above approaches for dense frames is computationally infeasible. SUMMARY OF THE INVENTION
[0010] Aspects and advantages of embodiments of the present disclosure will be set forth in part in the following description, or may be learned from the description, or may be learned through practice of the embodiments.
[0011] One example aspect of the present disclosure relates to a computer system for performing video processing tasks with improved computational efficiency, the computer system including: one or more processors and one or more non-transitory computer-readable media. The one or more non-transitory computer-readable media collectively store a machine learning model, the machine learning model including: a video kernel configured to be applied to multiple data samples from a set of video data to respectively generate multiple video tokens, where each data sample includes at least a portion of multiple image frames included in the set of video data; and a vision transformer configured to process the multiple video tokens to generate a model output. The one or more non-transitory computer-readable media collectively store instructions that, when executed by the one or more processors, cause the computing system to perform operations. The operations include: processing the set of video data using the machine learning model to generate the model output; where processing the set of video data using the machine learning model includes sparsely applying the video kernel to the set of video data.
[0012] Other aspects of the present disclosure relate to various systems, devices, non-transitory computer-readable media, user interfaces, and electronic devices.
[0013] These and other features, aspects, and advantages of the various embodiments of the present disclosure will be better understood with reference to the following description and the appended claims. The accompanying drawings, which are incorporated in and constitute a part of this specification, illustrate example embodiments of the present disclosure and, together with the description, serve to explain the relevant principles. BRIEF DESCRIPTION OF THE DRAWINGS
[0014] Referring to the accompanying drawings, a detailed discussion of embodiments for those of ordinary skill in the art is set forth in this specification, in which:
[0015] Figure 1 A graphical diagram of an example machine learning model in accordance with an example embodiment of the present disclosure is depicted.
[0016] Figure 2 A graphical diagram of an example machine learning model in accordance with an example embodiment of the present disclosure is depicted.
[0017] Figure 3 A graphical diagram of an example machine learning model in accordance with an example embodiment of the present disclosure is depicted.
[0018] Figures 4A to 4C A block diagram of an example computing system and apparatus in accordance with an example embodiment of the present disclosure is depicted.
[0019] Like reference numerals repeated across multiple figures are intended to identify the same features in various implementations. DETAILED DESCRIPTION
[0020] Overview
[0021] In general, the present disclosure relates to machine learning models for performing video processing with improved efficiency. In particular, a machine learning model can perform a sparse application of one or more video kernels to a set of video data to generate video tokens, which can be provided, for example, as input to a vision transformer. Thus, example implementations of the present disclosure relate to a way to transform a vision transformer (e.g., a ViT encoder) into an efficient video model. Additionally, the example implementations described herein can work seamlessly with both image inputs and video inputs. Specifically, by sparsely sampling the input, the model is able to be trained and inferred based on both types of inputs. The proposed model is easily scalable and can optionally be adapted to large-scale pre-trained vision transformers without full fine-tuning.
[0022] More specifically, to address the computational challenges associated with various existing approaches, the present disclosure proposes an efficient model, example implementations of which may be referred to as "TubeViT". The proposed model can seamlessly leverage existing vision transformer architectures (e.g., standard ViT models) for both images and videos. Specifically, the proposed model can implement Sparse Video Tubes, which are a lightweight approach for video learning (e.g., joint image and video learning). Example implementations of the proposed technique work by sparsely sampling one or more 3D spatio-temporal tubes of various sizes from a video to generate learnable tokens used by the vision transformer.
[0023] Using the sparse video tubes, the model is easily adaptable to image inputs and / or video inputs and can better utilize either or both data sources for training and fine-tuning. The sparse video tubes naturally handle the raw video signal and / or image signal, which helps in understanding actions and other spatio-temporal information in the video. As an example, Figure 1 An example model in accordance with aspects of the present disclosure is shown. Specifically, using the sparse video tubes, the vision transformer 12 can be used to process image input 14, video input 16, and / or both inputs 14 and 16, thus providing an efficient video backbone and more accurate performance. Additionally, due to the flexibility of the transformer model in accepting various types and / or lengths of tokens, the input to the transformer can also include tokens generated from other data modalities in addition to image and / or video tokens. For example, speech tokens, natural language tokens, and / or other tokens can be provided as input for other forms of multimodal processing tasks.
[0024] By using sparse video tubes, example implementations of the present disclosure can better share the weights learned for both images and videos. This is in contrast to prior work that makes kernels dilated or adds new time-specific layers. Further, due to sparse sampling, the number of labels remains low, which is also beneficial for both reducing computational expenditure (e.g., in terms of FLOP) and improving performance. Thus, the technical effects and benefits of the present disclosure include both a reduced consumption of computational resources for various computer vision tasks and an improved performance or functionality of a computer system. For example, video-based computer vision tasks can include video classification, video recognition, object detection or classification in video, embedding generation (e.g., for video representation), and / or other video processing or vision tasks.
[0025] Accordingly, the present disclosure provides sparse video tubes, which can be obtained by sparsely sampling videos by using 3D spatio-temporal tubes or video kernels of various sizes. Using these sparse video kernels, example implementations of the present disclosure can achieve some or all of the following: (1) a general vision backbone that easily adapts a vision transformer (e.g., ViT) to video; (2) seamless use of joint image and video understanding for any input; and / or (3) an easy-to-scale approach for video understanding that can also utilize a (large) pre-trained vision transformer model.
[0026] In addition, compared to prior work, the proposed technique enables computationally efficient application of a single transformer model to a large set of video data. Example differences that achieve this computational efficiency include: the tubes are sparsely applied to the raw input, the model is composed of tubes of different shapes that may overlap, and the model uses a single shared backbone network. This results in a model that is both more efficient and more accurate. Further, the model can be fully shared between the image modality and the video modality. This is a valuable distinction as it not only improves performance for both tasks but is also more generally applicable to vision tasks.
[0027] Another aspect of the present disclosure relates to a technique for efficiently scaling from a small video model to a large video model using a pre-trained image model. Specifically, video models are typically computationally expensive to train, and previous work has studied ways to leverage already trained models, such as using frozen models or adapting them to video. Example implementations of the present disclosure build on these ideas and utilize lightweight training, using sparse video tubes to adapt a much larger pre-trained vision transformer model to video. Thus, powerful large video models can be created while consuming fewer computational resources, such as processor usage, memory usage, network bandwidth, etc.
[0028] Referring now to the accompanying drawings, example embodiments of the present disclosure will be discussed in further detail.
[0029] Example models and techniques
[0030] Example dense sampling method
[0031] The standard ViT architecture takes an image and converts it into patch embeddings, for example, by using a 16×16 2D convolutional kernel with a 16×16 stride. This produces a sequence of patches as the image representation, e.g., 196 patches for a 224×224 input image.
[0032] Given a video Previous applications of vision transformers either used the same dense 2D patches independently for all frames (e.g., applying 2D patches densely on a per-frame basis) or used entirely dense 3D kernels (e.g., applying 3D patches densely to multiple frames with a time stride of 1). As an example, ViVit (Arnab et al., ViVit: A video vision transformer, in ICCV, 2021) applies 2×16×16 or 4×16×16 densely to all frames in a video. In both cases (dense application of 2D kernels or dense application of 3D kernels), this produces significantly more tokens, e.g., T * 196, where T is the number of frames.
[0033] In the previous approach, these tubes or patches are then linearly projected into the embedding space, This sequence of tokens is then processed by a transformer encoder using standard components, such as MSA - multi-head self-attention and MLP - standard transformer projection layers. For example, for a sequence of layers l ∈ [0, 1,... L], the transformer can compute representations i for all z tokens and the next token features (LN represents layer normalization):
[0034]
[0035]
[0036] To reduce computational costs, the system can factorize the attention mechanism to have spatial and temporal attention, or utilize smaller view-level transformers to use multiple views. However, factorizing the attention mechanism does not directly address the challenges associated with the number of tokens.
[0037] Example sparse video tube
[0038] In contrast to the dense sampling techniques described above, example implementations of the present disclosure implement a simple and straightforward method that seamlessly applies to both images and videos. For example, some example implementations of the present disclosure can follow the standard ViT tokenization for images: 2D convolution with a 16×16 kernel. However, the example implementations are based on the observation that sparsity is effective for videos. Therefore, instead of following previous work that densely tokenizes videos, some example implementations use the same 2D kernel but with a large time stride applied, for example, to every 16th frame. Thus, for a 32×224×224 input video clip, this results in only 392 tokens, as opposed to 6,000 in TimeSFormer or 1,000 to 2,000 in ViViT.
[0039] However, this sparse spatial sampling may lose information, especially for fast or short actions. Therefore, some example implementations can create sparse video tubes of different shapes (e.g., various 3D shapes). Specifically, video kernels with these shapes can be applied to video data sampled according to these shapes to generate video tokens. The kernel can also be referred to as a filter.
[0040] 3D shapes generally can refer to shapes having a temporal length (e.g., length in terms of the number of frames), a spatial height, and a spatial width. When considering the channel depth of the data, these shapes can also be considered 4D shapes. Some example implementations can also optionally add an offset to the starting position so that the patches do not always start at (0,0,0), and this allows for a reduction in the overlap between tubes. This is shown in Figure 2 .
[0041] Specifically, we can represent the tube as (T×H×W) for the kernel shape, represented as (T s , H s , W s) is used for the spatio-temporal stride applied to the kernel, and the offset represented as (x, y, z) as the starting point of the convolution. As an example, a 16-frame × 4-pixel × 4-pixel tube can be used to obtain information from many frames at a low spatial resolution. However, the tube can have any shape. Importantly, in some implementations, the tube also has a large stride, so as to sparsely sample the video in different views.
[0042] Tubes of various sizes are also used in the multi-view (MultiView) approach (Yan et al., Multiview transformers for video recognition, CVPR, 2022) for video classification. However, in the multi-view, 3D tubes are densely sampled and processed separately by multiple different view-specific transformers, resulting in a more computationally intensive approach. In addition, in contrast to previous work, some example implementations of the present disclosure also allow overlap between tubes.
[0043] Using the proposed design, the example implementations of the present disclosure achieve seamless fusion of image visual information and video visual information. Sparse spatial sampling allows sharing of image tokens and frame tokens, and sparse video tubes create a small number of video-specific tokens. This enables better sharing of the vision transformer model between images and videos.
[0044] As an example, Figure 2 An example application of model 52 to a set of video data 54 is shown. Specifically, the example machine learning model 52 can include a video kernel 56, which is configured to be applied to multiple data samples from the set of video data 54 to respectively generate multiple video tokens 58. Each data sample can include at least a portion of multiple image frames included in the set of video data 54. Although reference is made to Figure 2 invoking the kernel 56, additional example video kernels 60 and 62 are shown, which can also be sparsely applied to generate additional video tokens.
[0045] The model 52 can also include a vision transformer 64, which is configured to process multiple video tokens 58 (and any other tokens present) to generate a model output 66.
[0046] Therefore, the machine learning model 52 can process the set of video data 54 to generate a model output 66. Specifically, processing the set of video data 54 using the machine learning model 52 can include sparsely applying the video kernel 56 to the set of video data.
[0047] In some implementations, the video kernel 56 has a spatial dimension size, and sparsely applying the video kernel 56 to the set of video data 54 can include: applying the video kernel 56 with a spatial stride larger than the spatial dimension size of the video kernel 56 to achieve spatial sparsity.
[0048] In some implementations, the video kernel 56 has a temporal dimension size, and sparsely applying the video kernel 56 to the set of video data 54 can include: applying the video kernel 56 with a temporal stride larger than the temporal dimension size of the video kernel 56 to achieve temporal sparsity.
[0049] In some implementations, the video kernel 56 can be directly applied to the set of video data 54 (e.g., rather than first linearly projecting or transforming the video data 54). For example, the video kernel 56 can be directly applied to the pixel values included in the set of video data 54.
[0050] In addition, in some implementations, the machine learning model 52 can further include one or more image kernels (e.g., image kernel 68), which are configured to be applied to individual image frames of the set of video data (e.g., frame 70) to generate a plurality of image tokens (e.g., token 74) from the individual image frames.
[0051] In some implementations, the machine learning model 52 can include a single vision transformer (e.g., transformer 64), which is configured to jointly process both a plurality of video tokens (e.g., video token 58) and a plurality of image tokens (e.g., image token 74) to generate a model output 66. For example, this is in contrast to certain multi-view approaches that use a separate transformer for each "view" obtained from the video. Due to the combination of the sampling density and the quadratic cost of the attention mechanism, these multi-transformer approaches are required. By using a single transformer to process all tokens, the resulting quality can be improved because cross-attention can be performed across all available tokens.
[0052] In some implementations, sparsely applying the video kernel to the set of video data 54 can include: starting to apply the video kernel at a predefined offset point different from the origin of the set of video data 54. For example, starting to apply the video kernel 62 at a spatial location not equal to (0,0). In another example, the offset can be a temporal offset such that the first application of the kernel starts at a frame different from the first frame of the video data.
[0053] In some implementations, the data samples to which different kernels are applied can be overlapping. For example, in Figure 2In this case, video cores 56 and 60 are applied to overlapping samples in the frames shown on the left side of video data 54.
[0054] As an example, model output 66 can be or include a video classification output. However, other forms of output for other video understanding or computer vision tasks can also be output, including for example video identification, object identification, object detection, image classification, image identification, video representation, image representation, and / or others.
[0055] Example position embedding for sparse video tube
[0056] Another example aspect of the proposed approach is the implementation of positional embeddings. In language models, relative positional embeddings are a common and effective way. However, here, the relative position between two tokens has minimal significance and there is no true reference to where the patches / tubes come from in the original video or image.
[0057] Some existing vision models use learnable positional embeddings for patches. Here, such an approach can be difficult for the model because these learned embeddings do not necessarily reflect where the patches come from in the original video, especially in the case of overlapping patches.
[0058] Instead, some example implementations of the present disclosure can use fixed sine / cosine embeddings. When applying positional embeddings, the stride, kernel shape, and offset of each tube can be considered. This ensures that the positional embeddings of each patch and tube have the global spatio-temporal position of that tube.
[0059] As an example, the embeddings can be calculated as follows. Here, τ is a constant hyperparameter (e.g., 10,000). For j from 0 to d / / 6 (where d is the number of features), and for t, x, y from 0 to T, H, W,
[0060]
[0061] p j,t = sin(t * ω j ), cos(t * ω j ) (4)
[0062] p j,x = sin(x * ω j ), cos(x * ω j ) (5)
[0063] p j,y = sin(y * ω j ), cos(y * ω j ) (6)
[0064] zi [t, x, y, 6j:6(j + 1)] += [p j,t , p j,x , p j,y (7)
[0065] This adds each spatio - temporal position embedding to the feature dimension of the token z i This can be done for different wavelengths for each channel. When there are 6 elements (sine and cosine values for each x, y, t), d / / 6 can be used, which creates position values for each channel of the representation.
[0066] Importantly, here, z i [t, x, y] represents the center of the tube, thus taking into account any strides and offsets used in the tube construction (the channel dimension is not shown here).
[0067] After the tokenization step, some example implementations can concatenate all the tokens together and apply a standard transformer model. This simple structure allows the model to share most of the weights between all inputs, which is quite beneficial.
[0068] Example sparse tube construction
[0069] Any number of methods or shapes can be used to create any number of visual tubes. One example way can include 2 tubes: a 1×16×16×d tube for tokenizing images and another 8×8×8×d tube for videos, where d represents the channel depth. Both tubes can have a stride of 16×16×16. This basic tokenizer provides strong performance and any number of variations can be made from this example.
[0070] Multi - Tube. Some example implementations add multiple tubes of various sizes to the core approach. For example, some example implementations can include tubes that are long in time and small in space, such as 16×4×4 for learning long actions, or include more spatially concentrated tubes, such as 2×16×16 tubes. There are many possible variations in tube shapes and strides.
[0071] Space-to-Depth. Another extension of the core approach is to reduce the number of channels in the tube, e.g., divide by a factor of 2. Thus, the tube shape becomes T×H×W×d / 2. Next, the model can concatenate 2 tokens along the channel axis. Optionally, the stride of the tube (e.g., the temporal stride) can also be reduced. This results in the same number of tokens and dimensions as the original, but effectively increases the kernel size without changing the number of parameters. In other words, when reducing the stride on the temporal axis, the tokens now represent T·2×H×W positions, but only T*H*W parameters are used. Although a factor of 2 is given as an example, any factor proportional to the number of available channels can be used.
[0072] Interpolated Kernels. In some implementations, the model can learn a single 3D kernel of a certain shape (e.g., 8×8×8) instead of having a unique kernel for each tube. Then trilinear interpolation can be performed to reshape the kernel into any number of various different sizes depending on the tube configuration, e.g., 4×16×16 or 32×4×4, etc. Kernels of any size can be created from this single kernel. This method has several advantages. (1) It reduces the number of learned parameters used only on video streams. (2) It enables more flexible use of the kernel, e.g., it can become longer to handle longer videos, or larger spatially to find small objects.
[0073] Example image and video joint training
[0074] As described above, the proposed approach can seamlessly adapt to images, videos, or both inputs. Although image + video joint input is rare, given that many datasets with valuable annotations (e.g., ImageNet, Kinetics) come either from image sources or video sources, but not from both, the ability to use them together during training is beneficial.
[0075] It is easy to jointly train using the proposed approach - images are tokenized with 2D kernels, and videos are tokenized with both 2D patches (e.g., with a large temporal stride) and sparse tubes. Then both are passed to a standard ViT; positional embeddings can be provided in either case. The positional embedding approach helps to make the joint training effective.
[0076] Example image-to-video expansion of the model
[0077] Some example implementations also utilize more efficient ways to scale the model (e.g., by using in Figure 3(as shown in the example). Training large vision transformer models is computationally expensive, especially for videos. Since almost all components of the proposed model are shared between images and videos, an example approach is to leverage large models without heavy fine-tuning.
[0078] First, a smaller model can be jointly trained on images and videos. This gives a set of weights for the tubes used to generate video tokens (e.g., the values of the parameters of the video kernels). Then, the large pre-trained image ViT can be modified by further adding the learned tubes. These tubes can use the same kernel weights as the smaller model, and thus some example implementations can avoid further training them (but alternatively, they can optionally continue to be refined / learned if needed). Since larger ViTs typically use more channel dimensions than smaller ViTs, a spatial-to-depth transformation can optionally be used again here to create tokens with appropriate channel dimensions without new weights.
[0079] Next, the scaling approach can include selecting a point in the network and freezing all layers before that point, e.g., the 26th layer out of 32 layers in ViT-H (but alternatively, the entire model can continue to be refined / learned if needed). At that point, some example implementations can add gated connections to the network:
[0080] z s = MLP(LN(y s )) + y s + tanh(α)z 0 (8)
[0081] where s is the layer of the ViT model at which the network is frozen (e.g., 26), and z 0 is the original input token from the tube. α is the learned gating parameter, initialized to 0 (e.g., and as training continues, it is updated to a non-zero value).
[0082] In the first few steps of training, the gate has no effect on the representation, and thus the ViT remains unchanged. However, it can learn to incorporate the original tubes at that point and further refine the weights that follow.
[0083] Therefore, Figure 3Shows an example way of scaling for the TubeViT model. Given that building large-scale video models is expensive, the proposed scaling way can utilize a large pre-trained ViT to expand the model capacity of the video model. TubeViT can easily train small-scale models on both image data and video data. Then, sparse video tubes can be imported into a much larger ViT trained only on images, which can be mostly frozen and / or partially fine-tuned.
[0084] Example devices and systems
[0085] Figure 4A Depicts a block diagram of an example computing system 100 according to an example embodiment of the present disclosure. The system 100 includes a user computing device 102, a server computing system 130, and a training computing system 150 communicatively coupled via a network 180.
[0086] The user computing device 102 can be any type of computing device, such as, for example, a personal computing device (e.g., a laptop computer or a desktop computer), a mobile computing device (e.g., a smart phone or a tablet computer), a gaming console or controller, a wearable computing device, an embedded computing device, or any other type of computing device.
[0087] The user computing device 102 includes one or more processors 112 and a memory 114. The one or more processors 112 can be any suitable processing device (e.g., a processor core, a microprocessor, an ASIC, an FPGA, a controller, a microcontroller, etc.) and can be one processor or multiple processors operatively connected. The memory 114 can include one or more non-transitory computer-readable storage media, such as RAM, ROM, EEPROM, EPROM, flash memory devices, disks, etc. and combinations thereof. The memory 114 can store data 116 and instructions 118 executed by the processor 112 to cause the user computing device 102 to perform operations.
[0088] In some implementations, the user computing device 102 can store or include one or more machine learning models 120. For example, the machine learning model 120 can be or otherwise include various machine learning models, such as neural networks (e.g., deep neural networks) or other types of machine learning models, including non-linear models and / or linear models. Neural networks can include feedforward neural networks, recurrent neural networks (e.g., long short-term memory recurrent neural networks), convolutional neural networks, or other forms of neural networks. Some example machine learning models can utilize an attention mechanism, such as self-attention. For example, some example machine learning models can include a multi-head self-attention model (e.g., a transformer model). Reference Figures 1 to 3The example machine learning model 120 is discussed.
[0089] In some implementations, one or more machine learning models 120 can be received via network 180 from server computing system 130, stored in user computing device memory 114, and then used or otherwise implemented by one or more processors 112. In some implementations, user computing device 102 can implement multiple parallel instances of a single machine learning model 120.
[0090] Additionally or alternatively, one or more machine learning models 140 can be included in or otherwise stored and implemented by server computing system 130, which communicates with user computing device 102 according to a client-server relationship. For example, machine learning model 140 can be implemented by server computing system 140 as part of a web service. Thus, one or more models 120 can be stored and implemented at user computing device 102, and / or one or more models 140 can be stored and implemented at server computing system 130.
[0091] User computing device 102 can also include one or more user input components 122 that receive user input. For example, user input component 122 can be a touch-sensitive component (e.g., a touch-sensitive display screen or touchpad) that is sensitive to user input objects (e.g., a finger or a stylus). The touch-sensitive component can be used to implement a virtual keyboard. Other example user input components include a microphone, a traditional keyboard, or other devices by which a user can provide user input.
[0092] Server computing system 130 includes one or more processors 132 and memory 134. One or more processors 132 can be any suitable processing device (e.g., a processor core, a microprocessor, an ASIC, an FPGA, a controller, a microcontroller, etc.) and can be a single processor or multiple processors operatively connected. Memory 134 can include one or more non-transitory computer-readable storage media, such as RAM, ROM, EEPROM, EPROM, flash memory devices, disks, etc., and combinations thereof. Memory 134 can store data 136 and instructions 138 that are executed by processor 132 to cause server computing system 130 to perform operations.
[0093] In some implementations, server computing system 130 includes one or more server computing devices or is otherwise implemented by the one or more server computing devices. In the case where server computing system 130 includes multiple server computing devices, such server computing devices can operate according to a sequential computing architecture, a parallel computing architecture, or some combination thereof.
[0094] As described above, the server computing system 130 may store or otherwise include one or more machine learning models 140. For example, the model 140 may be or otherwise include various machine learning models. Example machine learning models include neural networks or other multi-layer non-linear models. Example neural networks include feedforward neural networks, deep neural networks, recurrent neural networks, and convolutional neural networks. Some example machine learning models may utilize an attention mechanism, such as self-attention. For example, some example machine learning models may include a multi-head self-attention model (e.g., a transformer model). Refer to Figures 1 to 3 Discuss example model 140.
[0095] The user computing device 102 and / or the server computing system 130 may train the model 120 and / or 140 via interaction with a training computing system 150 communicatively coupled via a network 180. The training computing system 150 may be separate from the server computing system 130 or may be part of the server computing system 130.
[0096] The training computing system 150 includes one or more processors 152 and a memory 154. The one or more processors 152 may be any suitable processing device (e.g., a processor core, a microprocessor, an ASIC, an FPGA, a controller, a microcontroller, etc.) and may be one processor or multiple processors operatively connected. The memory 154 may include one or more non-transitory computer-readable storage media, such as RAM, ROM, EEPROM, EPROM, flash memory devices, disks, etc. and combinations thereof. The memory 154 may store data 156 and instructions 158, which are executed by the processor 152 to cause the training computing system 150 to perform operations. In some implementations, the training computing system 150 includes one or more server computing devices or is otherwise implemented by the one or more server computing devices.
[0097] The training computing system 150 may include a model trainer 160 that uses various training or learning techniques (such as, for example, error backpropagation) to train the machine learning models 120 and / or 140 stored at the user computing device 102 and / or the server computing system 130. For example, a loss function may be backpropagated through the model to update one or more parameters of the model (e.g., based on the gradient of the loss function). Various loss functions may be used, such as mean squared error, likelihood loss, cross-entropy loss, hinge loss, and / or various other loss functions. Gradient descent techniques may be used to iteratively update the parameters over multiple training iterations.
[0098] In some implementations, performing backpropagation of error can include performing truncated backpropagation through time. The model trainer 160 can perform various generalization techniques (e.g., weight decay, dropout, etc.) to improve the generalization ability of the model being trained.
[0099] In particular, the model trainer 160 can train the machine learning models 120 and / or 140 based on a set of training data 162. The training data 162 can include, for example, video data with associated labels, image data with associated labels, and / or other forms of training data.
[0100] In some implementations, if the user has provided consent, the training examples can be provided by the user computing device 102. Thus, in such implementations, the model 120 provided to the user computing device 102 can be trained by the training computing system 150 based on user-specific data received from the user computing device 102. In some cases, this process can be referred to as personalizing the model.
[0101] The model trainer 160 includes computer logic for providing the desired functionality. The model trainer 160 can be implemented in hardware, firmware, and / or software that controls a general-purpose processor. For example, in some implementations, the model trainer 160 includes program files stored on a storage device, loaded into memory, and executed by one or more processors. In other implementations, the model trainer 160 includes one or more sets of computer-executable instructions stored on a tangible computer-readable storage medium (such as RAM, a hard disk, or an optical or magnetic medium).
[0102] The network 180 can be any type of communication network, such as a local area network (e.g., an intranet), a wide area network (e.g., the Internet), or some combination thereof, and can include any number of wired or wireless links. Generally, communication over the network 180 can be performed using a variety of communication protocols (e.g., TCP / IP, HTTP, SMTP, FTP), encodings or formats (e.g., HTML, XML), and / or security schemes (e.g., VPN, secure HTTP, SSL) via any type of wired and / or wireless connection.
[0103] The machine learning models described in this specification can be used for a variety of tasks, applications, and / or use cases.
[0104] In some implementations, the input to the machine learning model of the present disclosure can be image data. The machine learning model can process the image data to generate an output. As an example, the machine learning model can process the image data to generate an image recognition output (e.g., recognition of the image data, latent embedding of the image data, encoded representation of the image data, hash of the image data, etc.). As another example, the machine learning model can process the image data to generate an image segmentation output. As another example, the machine learning model can process the image data to generate an image classification output. As another example, the machine learning model can process the image data to generate an image data modification output (e.g., change of the image data, etc.). As another example, the machine learning model can process the image data to generate an encoded image data output (e.g., encoded representation and / or compressed representation of the image data, etc.). As another example, the machine learning model can process the image data to generate an enlarged image data output. As another example, the machine learning model can process the image data to generate a prediction output.
[0105] In some implementations, the input to the machine learning model of the present disclosure can be text or natural language data. The machine learning model can process the text or natural language data to generate an output. As an example, the machine learning model can process the natural language data to generate a language encoding output. As another example, the machine learning model can process the text or natural language data to generate a latent text embedding output. As another example, the machine learning model can process the text or natural language data to generate a translation output. As another example, the machine learning model can process the text or natural language data to generate a classification output. As another example, the machine learning model can process the text or natural language data to generate a text segmentation output. As another example, the machine learning model can process the text or natural language data to generate a semantic intent output. As another example, the machine learning model can process the text or natural language data to generate an enlarged text or natural language output (e.g., text or natural language data with higher quality than the input text or natural language). As another example, the machine learning model can process the text or natural language data to generate a prediction output.
[0106] In some implementations, the input to the machine learning model of the present disclosure can be speech data. The machine learning model can process the speech data to generate an output. As an example, the machine learning model can process the speech data to generate a speech recognition output. As another example, the machine learning model can process the speech data to generate a speech translation output. As another example, the machine learning model can process the speech data to generate a latent embedding output. As another example, the machine learning model can process the speech data to generate an encoded speech output (e.g., an encoded representation and / or a compressed representation of the speech data, etc.). As another example, the machine learning model can process the speech data to generate an amplified speech output (e.g., speech data with higher quality than the input speech data, etc.). As another example, the machine learning model can process the speech data to generate a text representation output (e.g., a text representation of the input speech data, etc.). As another example, the machine learning model can process the speech data to generate a prediction output.
[0107] In some implementations, the input to the machine learning model of the present disclosure can be latent encoded data (e.g., an input latent space representation, etc.). The machine learning model can process the latent encoded data to generate an output. As an example, the machine learning model can process the latent encoded data to generate a recognition output. As another example, the machine learning model can process the latent encoded data to generate a reconstruction output. As another example, the machine learning model can process the latent encoded data to generate a search output. As another example, the machine learning model can process the latent encoded data to generate a reclustering output. As another example, the machine learning model can process the latent encoded data to generate a prediction output.
[0108] In some implementations, the input to the machine learning model of the present disclosure can be statistical data. The statistical data can be, represent, or otherwise include data calculated and / or derived from some other data source. The machine learning model can process the statistical data to generate an output. As an example, the machine learning model can process the statistical data to generate a recognition output. As another example, the machine learning model can process the statistical data to generate a prediction output. As another example, the machine learning model can process the statistical data to generate a classification output. As another example, the machine learning model can process the statistical data to generate a segmentation output. As another example, the machine learning model can process the statistical data to generate a visualization output. As another example, the machine learning model can process the statistical data to generate a diagnostic output.
[0109] In some implementations, the input of the machine learning model of the present disclosure can be sensor data. The machine learning model can process the sensor data to generate an output. As an example, the machine learning model can process the sensor data to generate an identification output. As another example, the machine learning model can process the sensor data to generate a prediction output. As another example, the machine learning model can process the sensor data to generate a classification output. As another example, the machine learning model can process the sensor data to generate a segmentation output. As another example, the machine learning model can process the sensor data to generate a visualization output. As another example, the machine learning model can process the sensor data to generate a diagnostic output. As another example, the machine learning model can process the sensor data to generate a detection output.
[0110] In some cases, the machine learning model can be configured to perform a task that includes encoding the input data for reliable and / or efficient transmission or storage (and / or corresponding decoding). For example, the task can be an audio compression task. The input can include audio data, and the output can include compressed audio data. In another example, the input includes visual data (e.g., one or more images or videos), the output includes compressed visual data, and the task is a visual data compression task. In another example, the task can include generating an embedding for the input data (e.g., input audio or visual data).
[0111] In some cases, the input includes visual data, and the task is a computer vision task. In some cases, the input includes pixel data for one or more images, and the task is an image processing task. For example, the image processing task can be image classification, where the output is a set of scores, each score corresponding to a different object class and representing the likelihood that the one or more images depict an object belonging to the object class. The image processing task can be object detection, where the image processing output identifies one or more regions in the one or more images, and for each region, identifies the likelihood that the region depicts an object of interest. As another example, the image processing task can be image segmentation, where the image processing output defines, for each pixel in the one or more images, the corresponding likelihood of each category in a predetermined set of categories. For example, the set of categories can be foreground and background. As another example, the set of categories can be object classes. As another example, the image processing task can be depth estimation, where the image processing output defines, for each pixel in the one or more images, a corresponding depth value. As another example, the image processing task can be motion estimation, where the network input includes multiple images, and the image processing output defines, for each pixel in one of the input images, the motion of the scene depicted at that pixel between the images in the network input.
[0112] In some cases, the input includes audio data representing spoken words, and the task is a speech recognition task. The output may include a text output mapped to the spoken words. In some cases, the task includes encrypting or decrypting the input data. In some cases, the task includes microprocessor performance tasks such as branch prediction or memory address translation.
[0113] Figure 4A An example computing system that can be used to implement the present disclosure is shown. Other computing systems may also be used. For example, in some implementations, the user computing device 102 may include a model trainer 160 and a training data set 162. In such implementations, the model 120 can be both trained and used locally on the user computing device 102. In some of such implementations, the user computing device 102 can implement the model trainer 160 to personalize the model 120 based on user-specific data.
[0114] Figure 4B A block diagram of an example computing device 10 executing in accordance with an example embodiment of the present disclosure is depicted. The computing device 10 can be a user computing device or a server computing device.
[0115] The computing device 10 includes multiple applications (e.g., Application 1 to Application N). Each application contains its own machine learning library and machine learning model. For example, each application can include a machine learning model. Example applications include text messaging applications, email applications, dictation applications, virtual keyboard applications, browser applications, and the like.
[0116] As Figure 4B shown, each application can communicate with multiple other components of the computing device, such as one or more sensors, a context manager, a device status component, and / or additional components. In some implementations, each application can communicate with each device component using an API (e.g., a common API). In some implementations, the API used by each application is specific to that application.
[0117] Figure 4C A block diagram of an example computing device 50 executing in accordance with an example embodiment of the present disclosure is depicted. The computing device 50 can be a user computing device or a server computing device.
[0118] The computing device 50 includes multiple applications (e.g., Application 1 to Application N). Each application communicates with a central intelligence layer. Example applications include text messaging applications, email applications, dictation applications, virtual keyboard applications, browser applications, and the like. In some implementations, each application can use an API (e.g., a common API across all applications) to communicate with the central intelligence layer (and the models stored therein).
[0119] The central intelligence layer includes multiple machine learning models. For example, as Figure 4C shown, a corresponding machine learning model can be provided for each application, and the corresponding machine learning model can be managed by the central intelligence layer. In other implementations, two or more applications can share a single machine learning model. For example, in some implementations, the central intelligence layer can provide a single model for all applications. In some implementations, the central intelligence layer is included within or otherwise implemented by the operating system of computing device 50.
[0120] The central intelligence layer can communicate with the central device data layer. The central device data layer can be a centralized data repository of computing device 50. As Figure 4C shown, the central device data layer can communicate with multiple other components of the computing device, such as one or more sensors, a context manager, a device status component, and / or additional components. In some implementations, the central device data layer can communicate with each device component using an API (e.g., a private API).
[0121] Additional disclosure
[0122] The techniques discussed herein relate to servers, databases, software applications, and other computer-based systems, as well as the actions taken and the information sent to and from such systems. The inherent flexibility of computer-based systems allows for many possible configurations, combinations, and divisions of tasks and functions among and within components. For example, the processes discussed herein can be implemented using a single device or component or multiple devices or components working in combination. Databases and applications can be implemented on a single system or distributed across multiple systems. Distributed components can operate sequentially or in parallel.
[0123] Although the subject matter has been described in detail with respect to various specific example embodiments thereof, each example is provided by way of explanation and not limitation of the disclosure. Those skilled in the art will readily generate alterations, variations, and equivalents to such embodiments upon understanding the foregoing. Accordingly, the disclosure does not exclude including such modifications, variations, and / or additions to the subject matter that would be readily apparent to a person of ordinary skill in the art. For example, features shown or described as part of one embodiment can be used with another embodiment to yield yet another embodiment. Accordingly, the disclosure is intended to cover such alterations, variations, and equivalents.
Claims
1. A computer system for performing video processing tasks with improved computational efficiency, the computer system comprising: One or more processors; And One or more non-transitory computer-readable media that collectively store: A machine learning model, the machine learning model comprising: A video kernel configured to be applied to multiple data samples from a set of video data to respectively generate multiple video tokens, where each data sample includes at least a portion of multiple image frames included in the set of video data; and A vision transformer configured to process the multiple video tokens to generate a model output; and Instructions that, when executed by the one or more processors, cause the computer system to perform operations, the operations including: Processing the set of video data using the machine learning model to generate the model output; Wherein processing the set of video data using the machine learning model includes: sparsely applying the video kernel to the set of video data.
2. The computer system according to any one of the preceding claims, wherein: The video kernel has a spatial dimension size; and Sparsely applying the video kernel to the set of video data includes: applying the video kernel with a spatial stride larger than the spatial dimension size of the video kernel to achieve spatial sparsity.
3. The computer system according to any one of the preceding claims, wherein: The video kernel has a temporal dimension size; and Sparsely applying the video kernel to the set of video data includes: applying the video kernel with a temporal stride larger than the temporal dimension size of the video kernel to achieve temporal sparsity.
4. The computer system according to any one of the preceding claims, wherein sparsely applying the video core to the set of video data comprises: Directly applying the video kernel to the pixel values included in the set of video data.
5. The computer system according to any one of the preceding claims, wherein the machine learning model further includes one or more image kernels configured to be applied to individual image frames of the set of video data to generate multiple image tokens from the individual image frames.
6. The computer system according to claim 5, wherein the machine learning model includes a single vision transformer configured to jointly process both the multiple video tokens and the multiple image tokens to generate the model output.
7. The computer system according to any one of the preceding claims, wherein sparsely applying the video core to the set of video data comprises: Start applying the video kernel at a predefined offset point different from the origin of the set of video data.
8. The computer system according to any one of the preceding claims, wherein: The machine learning model further includes at least a second kernel configured to be applied to a second set of data samples from the set of video data; and Wherein at least one data sample in the second set of data samples overlaps with at least one data sample in the multiple data samples to which the video kernel is applied.
9. The computer system according to any one of the preceding claims, wherein processing the set of video data using the machine learning model further comprises: Generating a plurality of fixed sine / cosine positional embeddings for the plurality of video tokens respectively, wherein the fixed sine / cosine positional embedding of each token indicates the video kernel relative to the center of the set of video data.
10. The computer system according to any one of the preceding claims, wherein each data sample of the plurality of data samples includes data for a subgroup of a plurality of channels in only the channel dimension of the set of video data, and wherein at least one of the plurality of tokens is generated by concatenating two temporally shifted data samples along the channel dimension.
11. The computer system according to any one of the preceding claims, wherein the machine learning model includes a pre-trained visual encoder that has been fine-tuned using a set of video training data.
12. The computer system according to any one of the preceding claims, wherein the model output includes a video classification output.
13. A computer-implemented method, the method comprising: Obtaining, by a computing system including one or more computing devices, a set of video data and video labels; Processing, by the computing system, the set of video data using a machine learning model to generate the model output, wherein processing the set of video data using the machine learning model includes: Sparsely applying, by the computing system, a video kernel of the machine learning model to the set of video data to generate a plurality of video tokens, the video kernel having a temporal dimension size greater than 1; and Processing, by the computing system, the plurality of video tokens using a vision transformer of the machine learning model to generate the model output; Evaluating, by the computing system, a loss function that generates a loss value based on the model output and the video labels; and Modifying, by the computing system, one or more values of one or more parameters of the machine learning model based on the loss function.
14. The computer-implemented method according to claim 13, wherein modifying the one or more values of the one or more parameters of the machine learning model by the computing system based on the loss function comprises: Updating the parameter values of the video kernel based on the loss function.
15. The computer-implemented method according to claim 13, further comprising: Importing the video kernel into a larger pre-trained image transformer.
16. The computer-implemented method according to claim 13, wherein modifying the one or more values of the one or more parameters of the machine learning model by the computing system based on the loss function comprises: Fine-tuning one or more layers of the pre-trained image transformer while keeping one or more other layers of the pre-trained image transformer fixed.
17. The computer-implemented method according to any one of claims 13 to 16, wherein the machine learning model further includes one or more image kernels configured to be applied to individual image frames of the set of video data to generate a plurality of image tokens from the individual image frames, and wherein the machine learning model includes a single vision transformer configured to jointly process both the plurality of video tokens and the plurality of image tokens to generate the model output.
18. The computer-implemented method according to claim 17, wherein the method further comprises: An image loss function is evaluated by the computing system, and the image loss function generates an image loss value based on the model output and an image label associated with the individual image frame; and one or more values of one or more parameters of the machine learning model are modified by the computing system based on the image loss function.
19. One or more non-transitory computer-readable media that collectively store: A machine learning model, the machine learning model comprising: A video kernel configured to be applied to a plurality of data samples from a set of video data to respectively generate a plurality of video tokens, wherein each data sample includes at least a portion of a plurality of image frames included in the set of video data; and A vision transformer configured to process the plurality of video tokens to generate a model output; and Instructions that, when executed by the one or more processors, cause the computer system to perform operations, the operations including: Processing the set of video data using the machine learning model to generate the model output; wherein processing the set of video data using the machine learning model includes: sparsely applying the video kernel to the set of video data.
20. The one or more non-transitory computer-readable media according to claim 19, wherein: The video kernel has a spatial dimension size and a temporal dimension size; and Sparsely applying the video kernel to the set of video data includes: Applying the video kernel with a spatial stride greater than the spatial dimension size of the video kernel to achieve spatial sparsity; or Applying the video kernel with a temporal stride greater than the temporal dimension size of the video kernel to achieve temporal sparsity.