Sequential representation modeling with enforced future collapse via latent projection

WO2026177933A1PCT designated stage Publication Date: 2026-08-27GDM HOLDING LLC
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
PCT/US2026/014934
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2025-02-18
Filing Date
2026-02-11
Publication Date
2026-08-27

Smart Images

  • Figure US2026014934_27082026_PF_FP_ABST
    Figure US2026014934_27082026_PF_FP_ABST
Patent Text Reader

Abstract

Systems and methods described herein provide for: obtaining sequential source data including a plurality of data elements; generating a first representation of one or more first data elements of the plurality of data elements by a first element encoder; generating a second representation of a second data element of the plurality of data elements by a second element encoder; generating a latent projection based on the second representation; generating a predicted second representation based on the first representation and the latent projection by a representation predictor model; determining a loss between the second representation and the predicted second representation; and training at least the first element encoder based on the loss.
Need to check novelty before this filing date? Find Prior Art

Description

SEQUENTIAL REPRESENTATION MODELING WITH ENFORCED FUTURE COLLAPSE VIA LATENT PROJECTIONPRIORITY CLAIM

[0001] The present application is based on and claims priority to United States Provisional Application Number 63 / 759,793 having a filing date of February718, 2025. The present application claims priority to and the benefit of each of such applications and incorporates all such applications herein by reference in their entirety.FIELD

[0002] The present disclosure relates generally to machine-learning and artificial intelligence systems. More particularly, the present disclosure relates to sequential representation modeling with enforced future collapse via latent projection.BACKGROUND

[0003] Machine learning (ML) is a field of study in artificial intelligence (Al) that allows machines to leam and improve from data without being explicitly programmed. ML uses statistical algorithms to analyze large amounts of data, identify patterns, and make decisions.SUMMARY

[0004] Aspects and advantages of embodiments of the present disclosure will be set forth in part in the following description, or can be learned from the description, or can be learned through practice of the embodiments.

[0005] In an aspect, the present disclosure provides a computer-implemented method. The method includes obtaining sequential source data including a plurality of data elements. The method includes generating a first representation of one or more first data elements of the plurality of data elements by a first element encoder. The method includes generating a second representation of a second data element of the plurality of data elements by a second element encoder. The method includes generating a latent projection based on the second representation. The method includes generating a predicted second representation based on the first representation and the latent projection by a representation predictor model. The method includes determining a loss between the second representation and the predictedsecond representation. The method includes training at least the first element encoder based on the loss.

[0006] In some implementations, the sequential source data includes video data, and the plurality of data elements include a plurality of video patches, each of the video patches including at least a portion of one or more frames of the video data.

[0007] In some implementations, the first representation of the one or more first data elements includes a first embedding of the one or more first data elements, and the second representation of the second data element includes a second embedding of the second data element.

[0008] In some implementations, the one or more first data elements is ordered earlier in the sequential source data than the second data element.

[0009] In some implementations, the first element encoder and the second element encoder are initialized from a common encoder model.

[0010] In some implementations, the second element encoder is frozen during the training of the first element encoder.

[0011] In some implementations, at least one of the first element encoder or the second element encoder includes a masked autoencoder.

[0012] In some implementations, at least one of the first element encoder or the second element encoder includes smoothed parameters of another of the first element encoder or the second element encoder.

[0013] In some implementations, generating a latent projection based on the second representation includes generating, using a projection generation model, the latent projection based on the second representation.

[0014] In some implementations, the projection generation model includes a neural network.

[0015] In some implementations, the latent projection includes a latent vector.

[0016] In some implementations, the latent projection includes a latent description of the second element.

[0017] In some implementations, the latent description includes text data.

[0018] In some implementations, the representation predictor model includes a transformer model.

[0019] In some implementations, the transformer model includes a vision transformer (ViT) model.

[0020] In some implementations, the method further includes obtaining an inference data element. In some implementation, the method further includes, subsequent to training at least the first element encoder based on the loss, generating a first inference representation of the inference data element using the first element encoder.

[0021] In some implementations, the method further includes generating a second inference representation using the representation predictor model based on the first inference representation. In some implementations, the method further includes generating a second inference data element based on the second inference representation.

[0022] In some implementations, generating the second inference representation is based on a guidance input.

[0023] In some implementations, the guidance input includes one or more of: noise data or a description of the second inference data element.

[0024] In an aspect, the present disclosure provides a computer-implemented method. The method includes obtaining video data including a plurality of video data elements, each video data element respectively including at least a portion of one or more video frames. The method includes generating a first representation of a first video data element of the plurality of video data elements by a first element encoder. The method includes generating a second representation of a second video data element of the plurality of data elements by a second element encoder, the second data element including data ordered later in the video data than the one or more first data elements. The method includes generating a latent projection based on the second representation, the latent projection being descriptive of the second video data element. The method includes generating a predicted second representation based on the first representation and the latent projection by a representation predictor model, the predicted second representation approximating the second representation. The method includes determining a loss between the second representation and the predicted second representation. The method includes training at least the first element encoder based on the loss.

[0025] In an aspect, the present disclosure provides a computing system, including one or more processors and one or more non-transitory, computer-readable media storing instructions that, when implemented, cause the one or more processors to perform operations. The operations include obtaining sequential source data including a plurality of data elements. The operations include generating a first representation of one or more first data elements of the plurality of data elements by a first element encoder. The operations include generating a second representation of a second data element of the plurality’ of data elementsby a second element encoder. The operations include generating a latent projection based on the second representation. The operations include generating a predicted second representation based on the first representation and the latent projection by a representation predictor model. The operations include determining a loss between the second representation and the predicted second representation. The operations include training at least the first element encoder based on the loss.

[0026] In some implementations, the sequential source data includes video data, and wherein the plurality of data elements include a plurality of video patches, each of the video patches including at least a portion of one or more frames of the video data.

[0027] In some implementations, the first representation of the one or more first data elements includes a first embedding of the one or more first data elements, and wherein the second representation of the second data element includes a second embedding of the second data element.

[0028] In some implementations, the one or more first data elements is ordered earlier in the sequential source data than the second data element.

[0029] In some implementations, the first element encoder and the second element encoder are initialized from a common encoder model.

[0030] In some implementations, the second element encoder is frozen during the training of the first element encoder.

[0031] In some implementations, at least one of the first element encoder or the second element encoder includes a masked autoencoder.

[0032] In some implementations, at least one of the first element encoder or the second element encoder includes smoothed parameters of another of the first element encoder or the second element encoder.

[0033] In some implementations, generating a latent projection based on the second representation includes generating, using a projection generation model, the latent projection based on the second representation.

[0034] In some implementations, the projection generation model includes a neural network.

[0035] In some implementations, the latent projection includes a latent vector.

[0036] In some implementations, the latent projection includes a latent description of the second element.

[0037] In some implementations, the latent description includes text data.

[0038] In some implementations, the representation predictor model includes a transformer model.

[0039] In some implementations, the transformer model includes a vision transformer (ViT) model.

[0040] In some implementations, the operations further include obtaining an inference data element. In some implementations, the operations further include, subsequent to training at least the first element encoder based on the loss, generating a first inference representation of the inference data element using the first element encoder.

[0041] In some implementations, the operations further include generating a second inference representation using the representation predictor model based on the first inference representation. In some implementations, the operations further include generating a second inference data element based on the second inference representation.

[0042] In some implementations, generating the second inference representation is based on a guidance input.

[0043] In some implementations, the guidance input includes one or more of: noise data or a description of the second inference data element.

[0044] In an aspect, the present disclosure provides a computing system, including one or more processors and one or more non-transitory, computer-readable media storing instructions that, when implemented, cause the one or more processors to perform operations. The operations include obtaining video data including a plurality of video data elements, each video data element respectively including at least a portion of one or more video frames. The operations include generating a first representation of a first video data element of the plurality of video data elements by a first element encoder. The operations include generating a second representation of a second video data element of the plurality of data elements by a second element encoder, the second data element including data ordered later in the video data than the one or more first data elements. The operations include generating a latent projection based on the second representation, the latent projection being descriptive of the second video data element. The operations include generating a predicted second representation based on the first representation and the latent projection by a representation predictor model, the predicted second representation approximating the second representation. The operations include determining a loss between the second representation and the predicted second representation. The operations include training at least the first element encoder based on the loss.

[0045] In an aspect, the present disclosure provides one or more non-transitory, computer-readable media storing instructions that, when implemented, cause one or more processors to perform operations. The operations include obtaining sequential source data including a plurality of data elements. The operations include generating a first representation of one or more first data elements of the plurality of data elements by a first element encoder. The operations include generating a second representation of a second data element of the plurality of data elements by a second element encoder. The operations include generating a latent projection based on the second representation. The operations include generating a predicted second representation based on the first representation and the latent projection by a representation predictor model. The operations include determining a loss between the second representation and the predicted second representation. The operations include training at least the first element encoder based on the loss.

[0046] In some implementations, the sequential source data includes video data, and wherein the plurality of data elements include a plurality of video patches, each of the video patches including at least a portion of one or more frames of the video data.

[0047] In some implementations, the first representation of the one or more first data elements includes a first embedding of the one or more first data elements, and wherein the second representation of the second data element includes a second embedding of the second data element.

[0048] In some implementations, the one or more first data elements is ordered earlier in the sequential source data than the second data element.

[0049] In some implementations, the first element encoder and the second element encoder are initialized from a common encoder model.

[0050] In some implementations, the second element encoder is frozen during the training of the first element encoder.

[0051] In some implementations, at least one of the first element encoder or the second element encoder includes a masked autoencoder.

[0052] In some implementations, at least one of the first element encoder or the second element encoder includes smoothed parameters of another of the first element encoder or the second element encoder.

[0053] In some implementations, generating a latent projection based on the second representation includes generating, using a projection generation model, the latent projection based on the second representation.

[0054] In some implementations, the projection generation model includes a neural network.

[0055] In some implementations, the latent projection includes a latent vector.

[0056] In some implementations, the latent projection includes a latent description of the second element.

[0057] In some implementations, the latent description includes text data.

[0058] In some implementations, the representation predictor model includes a transformer model.

[0059] In some implementations, the transformer model includes a vision transformer (ViT) model.

[0060] In some implementations, the operations further include obtaining an inference data element. In some implementations, the operations further include, subsequent to training at least the first element encoder based on the loss, generating a first inference representation of the inference data element using the first element encoder.

[0061] In some implementations, the operations further include generating a second inference representation using the representation predictor model based on the first inference representation. In some implementations, the operations further include generating a second inference data element based on the second inference representation.

[0062] In some implementations, generating the second inference representation is based on a guidance input.

[0063] In some implementations, the guidance input includes one or more of: noise data or a description of the second inference data element.

[0064] In an aspect, the present disclosure provides one or more non-transitory, computer-readable media storing instructions that, when implemented, cause one or more processors to perform operations. The operations include obtaining video data including a plurality of video data elements, each video data element respectively including at least a portion of one or more video frames. The operations include generating a first representation of a first video data element of the plurality of video data elements by a first element encoder. The operations include generating a second representation of a second video data element of the plurality of data elements by a second element encoder, the second data element including data ordered later in the video data than the one or more first data elements. The operations include generating a latent projection based on the second representation, the latent projection being descriptive of the second video data element. The operations include generating a predicted second representation based on the first representation and the latent projection by arepresentation predictor model, the predicted second representation approximating the second representation. The operations include determining a loss between the second representation and the predicted second representation. The operations include training at least the first element encoder based on the loss.

[0065] Other example aspects of the present disclosure are directed to other systems, methods, apparatuses, tangible non-transitory computer-readable media, and devices for performing functions described herein. These and other features, aspects, and advantages of various implementations will become better understood with reference to the following description and appended claims. The accompanying drawings, which are incorporated in and constitute a part of this specification, illustrate implementations of the present disclosure and, together with the description, help explain the related principles.BRIEF DESCRIPTION OF THE DRAWINGS

[0066] FIGS. 1A - 1C are block diagrams of example networked computing systems according to example implementations of aspects of the present disclosure;

[0067] FIG. 2 is a block diagram illustrating an example computing system configured to implement an agent system according to example implementations of aspects of the present disclosure;

[0068] FIGS. 3A - 3B are example systems for sequential representation modeling with enforced future collapse via latent projection according to example implementations of aspects of the present disclosure;

[0069] FIG. 4 is a flow chart diagram illustrating an example method for sequential representation modeling with enforced future collapse via latent projection according to example implementations of aspects of the present disclosure;

[0070] FIG. 5 is a flow' chart diagram illustrating an example method for sequential representation modeling with enforced future collapse via latent projection according to example implementations of aspects of the present disclosure;

[0071] FIG. 6 is a flow chart diagram illustrating an example method for training a machine-learned model according to example implementations of aspects of the present disclosure;

[0072] FIG. 7 is a block diagram of an example processing How for using machine-learned model(s) to process input(s) to generate output(s) according to example implementations of aspects of the present disclosure;

[0073] FIG. 8 is a block diagram of an example sequence processing model according to example implementations of aspects of the present disclosure;

[0074] FIG. 9 is a block diagram of an example technique for populating an example input sequence for processing by a sequence processing model according to example implementations of aspects of the present disclosure;

[0075] FIG. 10 is a block diagram of an example model development platform according to example implementations of aspects of the present disclosure;

[0076] FIG. 11 is a block diagram of an example training workflow for training a machine-learned model according to example implementations of aspects of the present disclosure; and

[0077] FIG. 12 is a block diagram of an inference system for operating one or more machine-learned model(s) to perform inference according to example implementations of aspects of the present disclosure.DETAILED DESCRIPTION

[0078] The present disclosure provides for sequential representation modeling with enforced future collapse via latent projection. In particular, the present disclosure provides for improved prediction and understanding of elements in sequential data, such as video data or other ordered data with sequential dependency. One approach according to example aspects of the present disclosure is extracting meaningful features from video data for tasks such as prediction, planning, or understanding physical dynamics. Some existing approaches can produce blurred predictions that are averaged over multiple possible futures, due in part to limitations in handling multimodal future possibilities.

[0079] The present disclosure provides for enforcing “future collapse” in predicted sequential data. Future collapse refers to the ability of predictive models to predict a single, specific future series from given data (e.g., “past” data). According to example aspects of the present disclosure, during training of a predictive model, one or more first data elements (e.g., past data) and a second data element (e.g., future data) are encoded into respective first and second representations. The predictive model can predict a predicted second representation from the first representation (e.g., to approximate the second representation). The “true” second representation can be projected into a latent projection and provided as input to the predictive model. The latent projection can be, for example, the second representation projected into a lower-dimensional space. Intuitively, the latent projection provides for “leaking” some information about the future. Furthermore, the latent projectionprovides for the predictive model to focus on more general predictive capabilities (e.g., physics simulation, human motion) for a single future. In some implementations, the machine-learned models described herein can be incorporated into an artificial intelligence agent or “Al agent.”

[0080] Systems and methods according to example aspects of the present disclosure can provide for obtaining sequential source data comprising a plurality of data elements. As used herein, sequential source data can be or can include any sequential data. Sequential data refers to data (e.g., binary data) that is ordered according to an expected sequence. As examples, the sequential data can be or can include audiovisual data, video data, audio data, textual data, measured data, or other suitable sequential data. The ordering can be explicit (e.g., through pointers, metadata, element identifiers etc.) or implicit (e.g., by relative positions of stored data). A data element can refer to a portion of the sequential source data. The data element can be or can include at least a portion of one or more data items in the sequential source data. As one example, if the sequential source data is video data, the data element can be or can include at least a portion of one or more frames of the video data. For instance, the data element can include a single frame, a crop of a single frame, a lurality of frames, a crop over multiple frames, and / or any other suitable portion of the video data. As one example, in some implementations, the data element can be or can include a video patch. As used herein, a video patch can include at least a portion of one or more frames of the video data.

[0081] Additionally and / or alternatively, systems and methods according to example aspects of the present disclosure can provide for generating a first representation of one or more first data elements of the plurality of data elements by a first element encoder.Additionally and / or alternatively, systems and methods according to example aspects of the present disclosure can provide for generating a second representation of a second data element of the plurality of data elements by a second element encoder. The representations of a data element can be a reduced representation of the data element. For instance, the representations of the data element can require fewer computing resources (e.g., memory, bandwidth, compute cycles, etc.) to store, transmit, and / or process than the data element itself. As examples, the representations can be an embedding or an encoding. For instance, in some implementations, the first representation of the one or more first data elements can be a first embedding of the one or more first data elements. Additionally and / or alternatively the second representation of the second data element can be a second embedding of the second data element.

[0082] The one or more first data elements can be ordered earlier in the sequential source data than the second data element. For instance, the one or more first data elements can be a ‘'past”or"present" data element, whereas the second data element can be a “future’’ data element. For instance, if the sequential source data is an already-extant video, the one or more first data elements can represent an arbitrary frame or frames of the video, and the second data element can be some portion of the video that would not yet have been played if the video was paused at the frame(s) of the one or more first data elements. In generation tasks, for example, the second data element may have been generated after the one or more first data elements, and / or may have been generated in the same generation task that generated the one or more first data elements. When training on videos, however, the second data element can be used to guide training of the encoder models without encouraging the models to be overly dependent on this future information using the techniques described herein.

[0083] The first element encoder and / or the second element encoder can be any suitable encoder or encoding model. As examples, the encoder can be a video encoder, a video autoencoder, another type of autoencoder, or any other suitable autoencoder. For instance, in some implementations, the encoder can be or can include a neural network, a transformer model, or other suitable machine-learned model. Furthermore, in some implementations, the encoder can be a masked encoder or masked autoencoder.

[0084] In some implementations, the first element encoder and / or the second element encoder can be a same or similar model. For example, in some implementations, both the first element encoder and the second element encoder can be the same model and / or the same instance of model. As another example, in some implementations, the first element encoder may be initialized as two identical instances of a same model. For instance, the first element encoder and the second element encoder can be initialized from a common encoder model. However, in some implementations, the first element encoder and the second element encoder may diverge during training. For example, in some implementations, the second element encoder can be frozen during the training of the first element encoder. For instance, the parameters of the second element encoder may not be modified during some training iterations whereas the parameters of the first element encoder may be updated at some or all training iterations. Freezing the second element encoder can provide reduced occurrences of the training converging to a trivial solution due to the commonality between the element encoders.

[0085] Additionally and / or alternatively, systems and methods according to example aspects of the present disclosure can provide for generating a latent projection based on the second representation. The latent projection can be or can include the second representation projected into a lower-dimensional space. The latent projection can convey some limited information about the second representation, but less than the entirety of the information conveyed by the second representation. For example, the latent projection may capture high-level features of the second representation. The latent projection can be, for example, a latent vector. As another example, the latent projection can include other suitable description-level information about the second representation or second data element such as, for example, a latent description of the second element (e.g., a text data description).

[0086] The latent projection can be generated using a projection generation model (which may also be referred to as a ‘’projector’ or “projector model”). The projection generation model can be any suitable model or machine-learned model such as, for example, a neural -network or neural-network based model. The projection generation model may be configured, for example, to perform regression, classification, anomaly detection, clustering, or other projection techniques.

[0087] Additionally and / or alternatively, systems and methods according to example aspects of the present disclosure can provide for generating a predicted second representation based on the first representation and the latent projection by a representation predictor model. The representation predictor model can be any suitable model or machine-learned model, such as, for example, a transformer model, such as a vision transformer (ViT) model. The representation predictor model can be configured to predict future data in the time series data provided as input at the representation level (e.g., embedding level). Because of the wide variety of potential futures at any point in time-series data, the representation predictor model can utilize the latent projection to collapse these possible futures to a single possibility and in turn provide more focused predicted second representations. Generally, the predicted second representation can approximate the second representation due in part to the information available in the latent projection. Intuitively, the information available in the latent projection can provide for the representation predictor model to “select” a specific future by “leaking” limited information from the “true” future. However, the predicted second representation may not necessarily be identical to the second representation, as the entire second representation is not available to the representation prediction model.

[0088] Additionally and / or alternatively, systems and methods according to example aspects of the present disclosure can provide for determining a loss between the secondrepresentation and the predicted second representation. The loss can be any suitable loss representative of differences between the second representation and the predicted second representation with respect to training objectives. The loss can, for example, be based on a dot product between the second representation and the predicted second representation. Additionally and / or alternatively, systems and methods according to example aspects of the present disclosure can provide for training at least the first element encoder based on the loss. For instance, the loss can be used to compute one or more deltas or updates for one or more parameters of the first element encoder. The deltas or updates can be applied to update the parameters of the first element encoder. Other machine-learned components, such as the projection generation model, the latent predictor model, and / or the second element encoder can be trained using the loss in addition to or alternatively to the first element encoder.

[0089] In some implementations, training the first element encoder can be performed over multiple training iterations or batches. For instance, the second representation can be input into the representation predictor model (e.g., in place of the first representation generated by the first element encoder) and used to generate additional subsequent representations, such as a third representation. This process can be repeated to generate a plurality of predicted representations. A loss from each of these predictions can be evaluated and accumulated over a batch. The batch may be selected from a single sequence or parts of multiple sequences. Furthermore, in some implementations, multiple batches may be generated, and at each step of the training process, a total loss of a selected batch (e.g.. a randomly selected batch for stochastic gradient descent) can be reduced.

[0090] After training the model(s), the systems and methods described herein can be used for a variety of inference-time objectives. For instance, the present disclosure can find applications in a number of computing-technology’ related fields, such as, for example, data representation tasks (e.g., self-supervised video representations), data generation tasks (e.g., video data generation from a description, outline, or other suitable latent data) and / or timeseries data forecasting tasks. For instance, the present disclosure can provide improved video understanding, such as providing for understanding non-verbal concepts directly from video data. As another example, the present disclosure can provide improved multi-resolution planning and control in dynamic environments for robotics control. As another example, the present disclosure can provide for enhanced prediction of future events for vehicle navigation (e.g., autonomous driving). As another example, the present disclosure can provide for improved generation of realistic and coherent video sequences.

[0091] As one example, the first element encoder can be used to generate representations of new data. For instance, the systems and methods described herein can further provide for obtaining an inference data element and, subsequent to training at least the first element encoder based on the loss, generating a first inference representation of the inference data element using the first element encoder. This first inference representation itself can be a valuable output. For instance, the first inference representation represents an improvement over existing video compression tasks by improvedly capturing spatiotemporal aspects of the inference data element that provide for understanding of actor behavior in the inference data element (e.g., video data).

[0092] Furthermore, in some implementations, this first inference representation can be used in generation or forecasting tasks. For example, the systems and methods described herein can further provide for generating a second inference representation using the representation predictor model based on the first inference representation and / or generating a second inference data element based on the second inference representation. For instance, the second inference representation can be generated using the representation prediction model. Due in part to the use of a latent projection during training, the representation prediction model can more accurately generate predicted representations that converge to a single “selected” future rather than being blurred across a plurality of potential futures.

[0093] In some implementations, generating the second inference representation is based on a guidance input. The guidance input can be similar to the latent projection that was used to train the models described herein. However, rather than being a projection of “true” second representations, the guidance input can be other guiding or seeding data that encourages the second inference representation to converge to a single future. In some implementations, for example, the guidance input can be noise data. The noise data can provide for some randomness in generation while still encouraging convergence to a single future. As another example, in some implementations, the guidance data can be a description of the second inference data element. For instance, in some implementations such as those where the latent projection includes a latent description, the description can guide the models to generate second inference representations that adhere to the high-level description of what the predicted data should contain.

[0094] Example aspects of the present disclosure can provide a number of technical effects and benefits, including improvements to computing technology. For instance, using a latent projection to generate a predicted separate representation can improve the computingtechnology-related field of data forecasting by enabling a predictive model to predict a single,plausible future rather than averaging over multiple possibilities. This can in turn provide for improved generation of sharper and / or more realistic generated predictions and / or reduced ■‘blurring’’ of generated data. The improved plausibility of generated data can, for instance, be considered to be representative of a more accurate prediction.

[0095] As another example, using a latent projection to generate a predicted separate representation can improve the computing-technology-related field of data forecasting by separating future disambiguation from general predictive capabilities. This can improve functionality of computing systems by providing for a predictive model to focus on general predictive capabilities rather than future disambiguation, which can provide for the model to improvedly generalize to unseen data and / or more accurately depict aspects attributable to those general predictive capabilities, such as physics simulation, actor movement, consistency and plausibility of features in the data across multiple data items, and so on.

[0096] As another example, using a latent projection to generate a predicted separate representation can improve the computing-technolog -related field of machine-learned model training. For instance, the use of a latent projection can simplify the training process for the machine-learned models, enable the models to be trained end-to-end in the desired machinelearning task, and reduce training time necessary to reach plausible outputs.

[0097] Various example implementations are described herein with respect to the accompanying FIGS.Example Model Systems and Architectures

[0098] Figure 1 A depicts a block diagram of an example computing system 100 that performs tasks according to example embodiments of the present disclosure. The system 100 includes a user computing device 102. a server computing system 130, and a training computing system 150 that are communicatively coupled over a network 180.

[0099] The user computing device 102 can be any type of computing device, such as, for example, a personal computing device (e.g., laptop or desktop), a mobile computing device (e.g., smartphone or tablet), a gaming console or controller, a wearable computing device, an embedded computing device, or any other type of computing device.

[0100] The user computing device 102 includes one or more processors 112 and a memory' 114. The one or more processors 112 can be any suitable processing device (e.g., a processor core, a microprocessor, an ASIC, an FPGA, a controller, a microcontroller, etc.) and can be one processor or a plurality of processors that are operatively connected. The memory 114 can include one or more non-transitory computer-readable storage media, suchas RAM, ROM, EEPROM, EPROM, flash memory devices, magnetic disks, etc., and combinations thereof. The memory 114 can store data 116 and instructions 118 which are executed by the processor 112 to cause the user computing device 102 to perform operations.

[0101] In some implementations, the user computing device 102 can store or include one or more machine-learned models 120. For example, the machine-learned models 120 can be or can otherwise include various machine-learned models such as block generation models, as described herein. Additionally and / or alternatively, as described herein, the models 120 can be or can include neural networks (e.g., deep neural networks) or other types of machine-learned models, including non-linear models and / or linear models. Neural networks can include feed-forward neural networks, recurrent neural networks (e.g., long short-term memory recurrent neural networks), convolutional neural networks or other forms of neural networks. Some example machine-learned models can leverage an attention mechanism such as self-attention. For example, some example machine-learned models can include multi-headed self-attention models (e.g., transformer models).

[0102] In some implementations, the one or more machine-learned models 120 can be received from the server computing system 130 over network 180, stored in the user computing device memory7114, and then used or otherwise implemented by the one or more processors 112. In some implementations, the user computing device 102 can implement multiple parallel instances of a single machine-learned model 120 (e.g., to perform parallel tasks across multiple instances of machine-learned models).

[0103] Additionally or alternatively, one or more machine-learned models 140 can be included in or otherwise stored and implemented by the server computing system 130 that communicates with the user computing device 102 according to a client-server relationship. For example, the machine-learned models 140 can be implemented by the server computing system 140 as a portion of a web service. Thus, one or more models 120 can be stored and implemented at the user computing device 102 and / or one or more models 140 can be stored and implemented at the server computing system 130.

[0104] The user computing device 102 can also include one or more user input components 122 that receives user input. For example, the user input component 122 can be a touch-sensitive component (e g., a touch-sensitive display screen or a touch pad) that is sensitive to the touch of a user input object (e.g., a finger or a stylus). The touch-sensitive component can serve to implement a virtual keyboard. Other example user input components include a microphone, a traditional keyboard, or other means by which a user can provide user input.

[0105] The server computing system 130 includes one or more processors 132 and a memory 134. The one or more processors 132 can be any suitable processing device (e.g., a processor core, a microprocessor, an ASIC, an FPGA, a controller, a microcontroller, etc.) and can be one processor or a plural ity of processors that are operatively connected. The memory 134 can include one or more non-transitory computer-readable storage media, such as RAM, ROM, EEPROM, EPROM, flash memory devices, magnetic disks, etc., and combinations thereof. The memory 134 can store data 136 and instructions 138 which are executed by the processor 132 to cause the server computing system 130 to perform operations.

[0106] In some implementations, the server computing system 130 includes or is otherwise implemented by one or more server computing devices. In instances in which the server computing system 130 includes plural server computing devices, such server computing devices can operate according to sequential computing architectures, parallel computing architectures, or some combination thereof.

[0107] As described above, the server computing system 130 can store or otherwise include one or more machine-learned models 140. For example, the models 140 can be or can otherwise include various machine-learned models. Example machine-learned models include neural networks or other multi-layer non-linear models. Example neural networks include feed forward neural networks, deep neural networks, recurrent neural networks, and convolutional neural networks. Some example machine-learned models can leverage an attention mechanism such as self-attention. For example, some example machine-learned models can include multi-headed self-attention models (e.g., transformer models). Some example machine-learned models can include diffusion models.

[0108] The user computing device 102 and / or the server computing system 130 can train the models 120 and / or 140 via interaction with the training computing system 150 that is communicatively coupled over the network 180. The training computing system 150 can be separate from the server computing system 130 or can be a portion of the server computing system 130.

[0109] The training computing system 150 includes one or more processors 152 and a memory 154. The one or more processors 152 can be any suitable processing device (e.g., a processor core, a microprocessor, an ASIC, an FPGA, a controller, a microcontroller, etc.) and can be one processor or a plurality of processors that are operatively connected. The memory 154 can include one or more non-transitory computer-readable storage media, such as RAM, ROM, EEPROM, EPROM, flash memory devices, magnetic disks, etc., andcombinations thereof. The memory 154 can store data 156 and instructions 158 which are executed by the processor 152 to cause the training computing system 150 to perform operations. In some implementations, the training computing system 150 includes or is otherwise implemented by one or more server computing devices.

[0110] The training computing system 150 can include a model trainer 160 that trains the machine-learned models 120 and / or 140 stored at the user computing device 102 and / or the server computing system 130 using various training or learning techniques, such as, for example, backwards propagation of errors. For example, a loss function can be backpropagated through the model(s) to update one or more parameters of the model(s) (e.g., based on a gradient of the loss function). Various loss functions can be used such as mean squared error, likelihood loss, cross entropy loss, hinge loss, and / or various other loss functions. Gradient descent techniques can be used to iteratively update the parameters over a number of training iterations.

[0111] In some implementations, performing backwards propagation of errors can include performing truncated backpropagation through time. The model trainer 160 can perform a number of generalization techniques (e.g., weight decays, dropouts, etc.) to improve the generalization capability of the models being trained.

[0112] In particular, the model trainer 160 can train the machine-learned models 120 and / or 140 based on a set of training data 162. The training data 162 can include, for example, examples of data that generally correspond to data that would be consumed by a machine-learned model configured to perform a particular task. In some implementations, the examples may be labeled with an expected or desired output from the models.

[0113] In some implementations, if the user has provided consent, the training examples can be provided by the user computing device 102. Thus, in such implementations, the model 120 provided to the user computing device 102 can be trained by the training computing system 150 on user-specific data received from the user computing device 102. In some instances, this process can be referred to as personalizing the model.

[0114] The model trainer 160 includes computer logic utilized to provide desired functionality. The model trainer 160 can be implemented in hardware, firmware, and / or software controlling a general purpose processor. For example, in some implementations, the model trainer 160 includes program files stored on a storage device, loaded into a memory7and executed by one or more processors. In other implementations, the model trainer 160 includes one or more sets of computer-executable instructions that are stored in a tangible computer-readable storage medium such as RAM, hard disk, or optical or magnetic media.

[0115] The network 180 can be any type of communications network, such as a local area network (e.g., intranet), wide area network (e.g., Internet), or some combination thereof and can include any number of wired or wireless links. In general, communication over the network 180 can be carried via any type of wired and / or wireless connection, using a wide variety of communication protocols (e.g., TCP / IP, HTTP, SMTP, FTP), encodings or formats (e.g., HTML, XML), and / or protection schemes (e.g., VPN, secure HTTP, SSL).

[0116] The machine-learned models described in this specification may be used in a variety of tasks, applications, and / or use cases.

[0117] In some implementations, the input to the machine-learned model(s) of the present disclosure can be image data. The machine-learned model(s) can process the image data to generate an output. As an example, the machine-learned model(s) can process the image data to generate an image recognition output (e.g., a recognition of the image data, a latent embedding of the image data, an encoded representation of the image data, a hash of the image data, etc ). As another example, the machine-learned model(s) can process the image data to generate an image segmentation output. As another example, the machine-learned model(s) can process the image data to generate an image classification output. As another example, the machine-learned model(s) can process the image data to generate an image data modification output (e.g., an alteration of the image data, etc.). As another example, the machine-learned model(s) can process the image data to generate an encoded image data output (e.g., an encoded and / or compressed representation of the image data, etc.). As another example, the machine-learned model(s) can process the image data to generate an upscaled image data output. As another example, the machine-learned model(s) can process the image data to generate a prediction output.

[0118] In some implementations, the input to the machine-learned model(s) of the present disclosure can be text or natural language data. The machine-learned model(s) can process the text or natural language data to generate an output. As an example, the machine-learned model(s) can process the natural language data to generate a language encoding output. As another example, the machine-learned model(s) can process the text or natural language data to generate a latent text embedding output. As another example, the machine-learned model(s) can process the text or natural language data to generate a translation output. As another example, the machine-learned model(s) can process the text or natural language data to generate a classification output. As another example, the machine-learned model(s) can process the text or natural language data to generate a textual segmentation output. As another example, the machine-learned model(s) can process the text or naturallanguage data to generate a semantic intent output. As another example, the machine-learned model(s) can process the text or natural language data to generate an upscaled text or natural language output (e.g., text or natural language data that is higher quality than the input text or natural language, etc.). As another example, the machine-learned model(s) can process the text or natural language data to generate a prediction output.

[0119] In some implementations, the input to the machine-learned model(s) of the present disclosure can be speech data. The machine-learned model(s) can process the speech data to generate an output. As an example, the machine-learned model(s) can process the speech data to generate a speech recognition output. As another example, the machine-learned model(s) can process the speech data to generate a speech translation output. As another example, the machine-learned model(s) can process the speech data to generate a latent embedding output. As another example, the machine-learned model(s) can process the speech data to generate an encoded speech output (e.g., an encoded and / or compressed representation of the speech data, etc.). As another example, the machine-learned model(s) can process the speech data to generate an upscaled speech output (e.g., speech data that is higher quality than the input speech data, etc.). As another example, the machine-learned model(s) can process the speech data to generate a textual representation output (e.g., a textual representation of the input speech data, etc.). As another example, the machine-learned model(s) can process the speech data to generate a prediction output.

[0120] In some implementations, the input to the machine-learned model(s) of the present disclosure can be latent encoding data (e.g., a latent space representation of an input, etc.). The machine-learned model(s) can process the latent encoding data to generate an output. As an example, the machine-learned model(s) can process the latent encoding data to generate a recognition output. As another example, the machine-learned model(s) can process the latent encoding data to generate a reconstruction output. As another example, the machine-learned model(s) can process the latent encoding data to generate a search output. As another example, the machine-learned model(s) can process the latent encoding data to generate a reclustering output. As another example, the machine-learned model(s) can process the latent encoding data to generate a prediction output.

[0121] In some implementations, the input to the machine-learned model(s) of the present disclosure can be statistical data. Statistical data can be, represent, or otherwise include data computed and / or calculated from some other data source. The machine-learned model(s) can process the statistical data to generate an output. As an example, the machine-learned model(s) can process the statistical data to generate a recognition output. As anotherexample, the machine-learned model(s) can process the statistical data to generate a prediction output. As another example, the machine-learned model(s) can process the statistical data to generate a classification output. As another example, the machine-learned model(s) can process the statistical data to generate a segmentation output. As another example, the machine-learned model(s) can process the statistical data to generate a visualization output. As another example, the machine-learned model(s) can process the statistical data to generate a diagnostic output.

[0122] In some implementations, the input to the machine-learned model(s) of the present disclosure can be sensor data. The machine-learned model(s) can process the sensor data to generate an output. As an example, the machine-learned model(s) can process the sensor data to generate a recognition output. As another example, the machine-learned model(s) can process the sensor data to generate a prediction output. As another example, the machine-learned model(s) can process the sensor data to generate a classification output. As another example, the machine-learned model(s) can process the sensor data to generate a segmentation output. As another example, the machine-learned model(s) can process the sensor data to generate a visualization output. As another example, the machine-learned model(s) can process the sensor data to generate a diagnostic output. As another example, the machine-learned model(s) can process the sensor data to generate a detection output.

[0123] In some cases, the machine-learned model(s) can be configured to perform a task that includes encoding input data for reliable and / or efficient transmission or storage (and / or corresponding decoding). For example, the task may be an audio compression task. The input may include audio data and the output may include compressed audio data. In another example, the input includes visual data (e.g. one or more images or videos), the output includes compressed visual data, and the task is a visual data compression task. In another example, the task may include generating an embedding for input data (e.g. input audio or visual data).

[0124] In some cases, the input includes visual data and the task is a computer vision task. In some cases, the input includes pixel data for one or more images and the task is an image processing task. For example, the image processing task can be image classification, where the output is a set of scores, each score corresponding to a different object class and representing the likelihood that the one or more images depict an object belonging to the object class. The image processing task may be object detection, where the image processing output identifies one or more regions in the one or more images and, for each region, a likelihood that region depicts an object of interest. As another example, the image processingtask can be image segmentation, where the image processing output defines, for each pixel in the one or more images, a respective likelihood for each category in a predetermined set of categories. For example, the set of categories can be foreground and background. As another example, the set of categories can be object classes. As another example, the image processing task can be depth estimation, where the image processing output defines, for each pixel in the one or more images, a respective depth value. As another example, the image processing task can be motion estimation, where the network input includes multiple images, and the image processing output defines, for each pixel of one of the input images, a motion of the scene depicted at the pixel between the images in the network input.

[0125] In some cases, the input includes audio data representing a spoken utterance and the task is a speech recognition task. The output may include a text output which is mapped to the spoken utterance. In some cases, the task includes encrypting or decrypting input data. In some cases, the task includes a microprocessor performance task, such as branch prediction or memory address translation.

[0126] Figure 1 A illustrates one example computing system that can be used to implement the present disclosure. Other computing systems can be used as well. For example, in some implementations, the user computing device 102 can include the model trainer 160 and the training dataset 162. In such implementations, the models 120 can be both trained and used locally at the user computing device 102. In some of such implementations, the user computing device 102 can implement the model trainer 160 to personalize the models 120 based on user-specific data.

[0127] Figure IB depicts a block diagram of an example computing device 10 that performs according to example embodiments of the present disclosure. The computing device 10 can be a user computing device or a server computing device.

[0128] The computing device 10 includes a number of applications (e.g., applications 1 through N). Each application contains its own machine learning library and machine-learned model(s). For example, each application can include a machine-learned model. Example applications include a text messaging application, an email application, a dictation application, a virtual keyboard application, a browser application, etc.

[0129] As illustrated in Figure IB, each application can communicate with a number of other components of the computing device, such as, for example, one or more sensors, a context manager, a device state component, and / or additional components. In some implementations, each application can communicate with each device component using anAPI (e.g., a public API). In some implementations, the API used by each application is specific to that application.

[0130] Figure 1C depicts a block diagram of an example computing device 50 that performs according to example embodiments of the present disclosure. The computing device 50 can be a user computing device or a server computing device.

[0131] The computing device 50 includes a number of applications (e.g., applications I through N). Each application is in communication with a central intelligence layer.Example applications include a text messaging application, an email application, a dictation application, a virtual keyboard application, a browser application, etc. In some implementations, each application can communicate with the central intelligence layer (and model(s) stored therein) using an API (e.g., a common API across all applications).

[0132] The central intelligence layer includes a number of machine-learned models. For example, as illustrated in Figure 1C, a respective machine-learned model can be provided for each application and managed by the central intelligence layer. In other implementations, two or more applications can share a single machine-learned model. For example, in some implementations, the central intelligence layer can provide a single model for all of the applications. In some implementations, the central intelligence layer is included within or otherwise implemented by an operating system of the computing device 50.

[0133] The central intelligence layer can communicate with a central device data layer. The central device data layer can be a centralized repository of data for the computing device 50. As illustrated in Figure 1 C, the central device data layer can communicate with a number of other components of the computing device, such as, for example, one or more sensors, a context manager, a device state component, and / or additional components. In some implementations, the central device data layer can communicate with each device component using an API (e.g., a private API).

[0134] Referring now to FIG. 2, a block diagram illustrates an example computing system 200 configured to implement an agent system 202, according to example implementations of aspects of the present disclosure. The depicted computing system 200 is designed to receive multiple types of input data, process this data, and generate outputs that are responsive to the inputs in a contextually appropriate manner.

[0135] The agent system 202 within the computing system 200 is configured to receive visual data 204, audio data 206, and additional context data 208. Each type of data is processed by the agent system 202 using one or more machine-learned model(s) 205 to facilitate interaction within its operational environment. For example, visual data 204 caninclude live video streams from a camera or recorded video streams from a web resource, while audio data 206 can include spoken commands or ambient sounds captured by microphones.

[0136] Additional context data 208 can include sensor data, textual information, or other forms of digital data that provide further insights into the environment or the context of the interaction. As one example, the additional context data 208 can include sensor data that captures user inputs beyond speech inputs, such as touch-screen inputs, gestures, facial expressions, and / or other inputs. These user inputs can, in some implementations, be merged with other inputs such as visual data 204 to create combined inputs. In one example, a user can be provided with an interface that displays a real-time field of view of the agent system (e.g., which may correspond to visual data 204). The interface can enable the user to “draw” on or otherwise interact with the interface to mark up the real-time field of view. For example, the user could draw an arrow or make a circle to identify a particular object included within the scene displayed on the interface. The user’s graphical input can be added onto or merged with the visual data 204 to form a combined input. For example, the visual data 204 can be amended to include the arrow or circle, which can then be processed by the agent system 202. In such manner, interactive interfaces can provide the ability’ for the user to more granularly interact with or identify portions of the environment when query ing the agent system 202.

[0137] Furthermore, it should be appreciated that in some cases the user will be able to control the type, nature, content, or other characteristics of the visual data 204, audio data 206, and / or additional context data 208. As one example, the user can manipulate a field of view of a camera to alter the content of the visual data 204 that is provided to the agent system 202. Similarly, by speaking into a microphone, the user can provide additional audio data 206 as an input for the agent system 202. The agent system's ability to process and combine visual, auditory, and textual information allows it to generate more comprehensive and nuanced responses, carefully tailored to the user's multi-modal context.

[0138] The agent system 202 processes these diverse inputs to generate an agent action 210, w hich can include an output designed to respond to the processed inputs effectively. As examples, this action can range from textual responses, vocal responses, displaying information, controlling connected devices, or any other form of interaction output that is deemed appropriate based on the input data. Specifically, the agent system 202 can provide concise answers, generate detailed explanations, offer step-by-step instructions,display information through visual highlights or augmented reality overlays, control connected devices, and / or other forms of actions 210.

[0139] In some implementations, the agent system 202 can include and use specialized sequence processing models to integrate and analyze the input data. These models are configured to process complex patterns across different data modalities, enabling the agent system 202 to generate more accurate and contextually relevant responses. The sequence processing models may be specifically fine-tuned to handle various interaction dynamics, such as tum-based dialogues or more open-ended conversational formats, enhancing the flexibility and adaptability of the agent system. Additionally and / or alternatively, the agent system 202 can include block generation models.

[0140] Furthermore, the computing system 200 can be connected to a real-time communication framework that facilitates the immediate and efficient exchange of data, including the inputs and outputs to and from the agent system 202. This configuration reduces latency in data processing and response generation.

[0141] The agent system 202 can include or can have access to a user-specific memory layer 212. The user-specific memory layer 212 can provide for the agent system to access user-specific data, such as image data, video data, documents, and other data provided by the user to the agent system. As one example, the user-specific memory' layer 212 can access data at a designated local repository on a calling device belonging to the user and / or other memory within the computing system 200 or accessible by the computing system 200. For example, the user-specific memory layer can be a directory, folder, file repository, or other non-tangible, computer-readable media in which the user consents to store video data, pictures or image data, documents, files, music or audio data, or other computer-readable data that the user wishes for the agent system 202 to have access to. Additionally or alternatively, the user can ask the agent system 202 to store data in the user-specific memory layer 212, such as by asking the agent system 202 to record and store video data from a camera of the user device. For example, the user may instruct the agent system 202 to “remember where I parked.” in response to which the agent system 202 may capture image data and / or geopositional data of a vehicle of the user. As another example, the user may instruct the agent system 202 to “remember that for later,” in which case the agent system 202 may capture image data or video data of the environment at which the user is looking (e.g., through a camera on a wearable device, such as smart glasses).

[0142] As another example, the user can provide the agent system 202 with access instructions for user-specific data streams, such as video watch history, historical geodata,and other data that the user wishes for the agent system 202 to have access to, which can either be or can provide data to the user-specific memory layer 212. The user-specific memory layer 212 can be, in some implementations, a long-term memory layer that can, with the consent of the user, provide context relating to long-term memories of the user, such as birthdays, anniversaries, and so on.

[0143] In some implementations, in addition to the user-specific memory layer 212, the agent system 202 can include or have access to a model memory layer 214 or other memory system. The agent system 202 can store and retrieve various types of information to and from the model memoiy' layer 214. For example, the agent system 202 can store past interactions, observations, preferences, and / or information from the environment in the model memory layer 214. The agent system 202 can then recall this information for use in generating new predictions, outputs, or agent actions.

[0144] A number of different types of data can be stored in the model memory layer 214. One example of data stored within the model memory layer 214 can include object detections. This can include indexed records of objects that the agent system encounters during its operations, complete with metadata such as timestamps, location coordinates, and / or contextual tags. By archiving these detections, the agent system 202 can recognize and recall objects from a "history" of observed scenes. The agent system 202 can leverage this information to refine interactions and bolster situational awareness, potentially spanning different sessions of user interaction.

[0145] As another example data type, the model memory layer 214 can store embeddings of observed visual content, textual content, or other inputs. These embeddings can be low-dimensional numerical representations that encode the essential features of input data into a latent embedding space. The storage of embeddings associated with observed inputs allows the agent system 202 to conduct rapid comparisons and recognition tasks efficiently. In particular, these embeddings, which can be derived from various layer(s) of the agent system’s machine-learned models, can be used to perform similarity7searches to facilitate quick data retrieval.

[0146] As another example, intermediate model activations can be stored in the model memory layer 214. Capturing and preserving the state of model activations at various stages can enable the agent system 202 to efficiently resume or adjust its processing activities as needed. This feature can be used in scenarios involving long running or complex processing tasks that may be interrupted or require dynamic adjustments such as resetting the agent system to a prior state associated with a prior time.

[0147] As another example, the model memory layer 214 can store raw tokens generated by the agent system’s natural language processing, image processing, or other tokenization mechanisms. For example, a cache of tokens can be stored, with each being associated with a specific timestamp. This data allows for the reconstruction of the sequence of inputs and internal states over time, which can be used to retrieve and replay perceptual inputs associated with a particular timestamp or setting, or to otherwise provide the raw tokens as a contextual input for a later prediction.

[0148] By maintaining a repository of these data types, the agent system 202 can be equipped with a knowledge base that supports advanced functionalities such as context-aware computing, personalized interactions, and information retrieval from past observations. For example, upon retrieving stored information from the model memory layer 214, the agent system 202 can integrate the retrieved data into the current processing workflow. This integration can include aligning historical and current data to enhance the accuracy and relevance of the output.

[0149] In some implementations, the agent system 202 can include or have access to both short-term and long-term memory components. The short-term memory may be volatile, designed for the temporary storage of recent interactions and sensory inputs. In contrast, the long-term memory' may be non-volatile, storing valuable learned information, user preferences, historical interaction data, and significant environmental events for longer-term recall and usage. In addition, the design of the model memory’ layer 214 can accommodate both structured and unstructured data. As an example, for immediate processing needs, volatile memory’ such as Random Access Memory' (RAM) can be used. As another example, for the purpose of long-term data retention, non-volatile storage solutions such as Hard Disk Drives (HDDs) or Solid-State Drives (SSDs) can be used. Furthermore, the model memory’ layer 214 can include hybrid memory solutions that combine the rapid access capabilities of RAM with the extensive storage capacity of disk storage, thereby7optimizing the performance of the agent system 202 across various tasks.

[0150] FIG. 3A illustrates a system 300 for sequential representation modeling with enforced future collapse via latent projection according to example implementations of the present disclosure. The system 300 can obtain sequential source data 302 comprising a plurality’ of data elements. As used herein, sequential source data 302 can be or can include any sequential data. Sequential data refers to data (e.g., binary7data) that is ordered according to an expected sequence. As examples, the sequential source data 302 can be or can include audiovisual data, video data, audio data, textual data, measured data, or other suitablesequential data. The ordering can be explicit (e.g., through pointers, metadata, element identifiers etc.) or implicit (e.g., by relative positions of stored data). A data element can refer to a portion of the sequential source data 302. The data element can be or can include at least a portion of one or more data items in the sequential source data 302. As one example, if the sequential source data 302 is video data, the data element can be or can include at least a portion of one or more frames of the video data. For instance, the data element can include a single frame, a crop of a single frame, a plurality of frames, a crop over multiple frames, and / or any other suitable portion of the video data. As one example, in some implementations, the data element can be or can include a video patch. As used herein, a video patch can include at least a portion of one or more frames of the video data.

[0151] The system 300 can provide for generating a first representation 312 of one or more first data elements of the plurality of data elements by a first element encoder 310. Additionally and / or alternatively, the system 300 can provide for generating a second representation 314 of a second data element of the plurality of data elements by a second element encoder 311. The representations of a data element can be a reduced representation of the data element. For instance, the representations of the data element can require fewer computing resources (e.g., memory, bandwidth, compute cycles, etc.) to store, transmit, and / or process than the data element itself. As examples, the representations 312, 314 can be an embedding or an encoding. For instance, in some implementations, the first representation 312 of the one or more first data elements can be a first embedding of the one or more first data elements. Additionally and / or alternatively the second representation 314 of the second data element can be a second embedding of the second data element.

[0152] The one or more first data elements can be ordered earlier in the sequential source data 302 than the second data element. For instance, the one or more first data elements can be a ‘'past”or‘’Present’’ data element, whereas the second data element can be a “future” data element. For instance, if the sequential source data 302 is an already -extant video, the one or more first data elements can represent an arbitrary frame or frames of the video, and the second data element can be some portion of the video that would not yet have been played if the video was paused at the frame(s) of the one or more first data elements. In generation tasks, for example, the second data element may have been generated after the one or more first data elements, and / or may have been generated in the same generation task that generated the one or more first data elements. When training on videos, however, the second data element can be used to guide training of the encoder models without encouraging themodels to be overly dependent on this future information using the techniques described herein.

[0153] The first element encoder 310 and / or the second element encoder 311 can be any suitable encoder or encoding model. As examples, the encoder(s) 310, 311 can be a video encoder, a video autoencoder, another type of autoencoder, or any other suitable autoencoder. For instance, in some implementations, the encoder(s) 310, 311 can be or can include a neural network, a transformer model, or other suitable machine-learned model. Furthermore, in some implementations, the encoder can be a masked encoder or masked autoencoder.

[0154] In some implementations, the first element encoder 310 and / or the second element encoder 311 can be a same or similar model. For example, in some implementations, both the first element encoder 310 and the second element encoder 311 can be the same model and / or the same instance of model. As another example, in some implementations, the first element encoder 310 may be initialized as two identical instances of a same model. For instance, the first element encoder 310 and the second element encoder 311 can be initialized from a common encoder model. However, in some implementations, the first element encoder 310 and the second element encoder 311 may diverge dunng training. For example, in some implementations, the second element encoder 311 can be frozen during the training of the first element encoder 310. For instance, the parameters of the second element encoder 311 may not be modified during some training iterations whereas the parameters of the first element encoder 310 may be updated at some or all training iterations. Freezing the second element encoder 311 can provide reduced occurrences of the training converging to a trivial solution due to the commonality between the element encoders.

[0155] Additionally and / or alternatively, the system 300 can provide for generating a latent projection 322 based on the second representation 314. The latent projection 322 can be or can include the second representation 314 projected into a lower-dimensional space. The latent projection 322 can convey some limited information about the second representation 314, but less than the entirety of the information conveyed by the second representation 314. For example, the latent projection 322 may capture high-level features of the second representation 314. The latent projection 322 can be, for example, a latent vector. As another example, the latent projection 322 can include other suitable description-level information about the second representation 314 or second data element such as, for example, a latent description of the second element (e.g., a text data description).

[0156] The latent projection 322 can be generated using a projection generation model 320 (which may also be referred to as a “projector” or “projector model”). Theprojection generation model 320 can be any suitable model or machine-learned model such as, for example, a neural-network or neural-network based model. The projection generation model 320 may be configured, for example, to perform regression, classification, anomaly detection, clustering, or other projection techniques.

[0157] Additionally and / or alternatively, the system 300 can provide for generating a predicted second representation 324 based on the first representation 312 and the latent projection 322 by a representation predictor model 315. The representation predictor model 315 can be any suitable model or machine-learned model, such as, for example, a transformer model, such as a vision transformer (ViT) model. The representation predictor model 315 can be configured to predict future data in the time series data provided as input at the representation level (e.g., embedding level). Because of the wide variety of potential futures at any point in time-series data, the representation predictor model 315 can utilize the latent projection 322 to collapse these possible futures to a single possibility and in turn provide more focused predicted second representations 324. Generally, the predicted second representation 324 can approximate the second representation 314 due in part to the information available in the latent projection 322. Intuitively, the information available in the latent proj ection 322 can provide for the representation predictor model 315 to “select” a specific future by “leaking” limited information from the “true” future. However, the predicted second representation 324 may not necessarily be identical to the second representation 314. as the entire second representation 314 is not available to the representation predictor model 315.

[0158] Additionally and / or alternatively, the system 300 can provide for determining a loss 325 between the second representation 314 and the predicted second representation 324. The loss 325 can be any suitable loss representative of differences between the second representation 314 and the predicted second representation 324 with respect to training objectives. The loss 325 can, for example, be based on a dot product between the second representation 314 and the predicted second representation 324. Additionally and / or alternatively ,the system 300 can provide for training at least the first element encoder 310 based on the loss 325. For instance, the loss 325 can be used to compute one or more deltas or updates for one or more parameters of the first element encoder 310. The deltas or updates can be applied to update the parameters of the first element encoder 310. Other machine-learned components, such as the projection generation model 320, the representation predictor model 315. and / or the second element encoder 311 can be trained using the loss 325 in addition to or alternatively to the first element encoder 310.

[0159] After training the model(s), the systems and methods described herein can be used for a variety’ of inference-time objectives. As one example, the first element encoder 310 can be used to generate representations of new data. FIG. 3B illustrates an example system 350 for generating representations of new data using the first element encoder 310 (as trained first element encoder 360). The system 350 can provide for obtaining an inference data element 352. The inference data element 352 can be any suitable new data, such as a same type of data as the sequential source data 302 of FIG. 3 A. Subsequent to training the trained first element encoder 360 based on the loss 325, the system 350 can provide for generating a first inference representation 354 of the inference data element 352 using the trained first element encoder 360. This first inference representation 354 itself can be a valuable output. For instance, the first inference representation 354 represents an improvement over existing video compression tasks by improvedly capturing spatiotemporal aspects of the inference data element 352 that provide for understanding of actor behavior in the inference data element 352 (e g., video data).

[0160] Furthermore, in some implementations, this first inference representation 354 can be used in generation or forecasting tasks. For example, the system 350 can further provide for generating a second inference representation 356 using the representation predictor model 315 based on the first inference representation 354. From the second inference representation 356, the system 350 can further provide for generating a second inference data element 358 based on the second inference representation 356. The second inference data element 358 may, for example, be generated using a decoder model, such as a decoder model corresponding to the trained first elemental encoder 360. Due in part to the use of a latent projection 322 during training, the representation prediction model 315 can more accurately generate predicted representations 356 that converge to a single “selected” future rather than being blurred across a plurality of potential futures.

[0161] In some implementations, generating the second inference representation 356 is based on a guidance input 362. The guidance input 362 can be similar to the latent projection 322 that was used to train the models described herein. However, rather than being a projection of “true” second representations (e.g., 314), the guidance input 362 can be other guiding or seeding data that encourages the second inference representation 356 to converge to a single future. In some implementations, for example, the guidance input 362 can be noise data. The noise data can provide for some randomness in generation while still encouraging convergence to a single future. As another example, in some implementations, the guidance data 362 can be a description of the second inference data element 358. For instance, in someimplementations such as those where the latent projection 322 includes alatent description, the description can guide the models to generate second inference representations 356 that adhere to the high-level description of what the predicted data should contain.Example Methods

[0162] FIG. 4 depicts a flowchart of a method 400 for training an element encoder according to aspects of the present disclosure. For instance, an example agent system can include one or more machine-learned models and / or other systems configured to perform tasks in response to a query from a user.

[0163] One or more portion(s) of example method 400 can be implemented by a computing system that includes one or more computing devices such as, for example, computing systems described with reference to the other figures. Each respective portion of example method 400 can be performed by any (or any combination) of one or more computing devices. Moreover, one or more portion(s) of example method 400 can be implemented on the hardware components of the device(s) described herein, for example, to train one or more systems or models. FIG. 4 depicts elements performed in a particular order for purposes of illustration and discussion. Those of ordinary skill in the art, using the disclosures provided herein, will understand that the elements of any of the methods discussed herein can be adapted, rearranged, expanded, omitted, combined, or modified in various ways without deviating from the scope of the present disclosure. FIG. 4 is described with reference to elements / terms described with respect to other systems and figures for exemplary illustrated purposes and is not meant to be limiting. One or more portions of example method 400 can be performed additionally, or alternatively, by other systems.

[0164] At 402, the method 400 can include obtaining sequential source data comprising a plurality of data elements. As used herein, sequential source data can be or can include any sequential data. Sequential data refers to data (e.g., binary data) that is ordered according to an expected sequence. As examples, the sequential data can be or can include audiovisual data, video data, audio data, textual data, measured data, or other suitable sequential data. The ordering can be explicit (e.g., through pointers, metadata, element identifiers etc.) or implicit (e.g., by relative positions of stored data). A data element can refer to a portion of the sequential source data. The data element can be or can include at least a portion of one or more data items in the sequential source data. As one example, if the sequential source data is video data, the data element can be or can include at least a portion of one or more frames of the video data. For instance, the data element can include a singleframe, a crop of a single frame, a plurality of frames, a crop over multiple frames, and / or any other suitable portion of the video data. As one example, in some implementations, the data element can be or can include a video patch. As used herein, a video patch can include at least a portion of one or more frames of the video data. The sequential source data can have any suitable dimensionality'. For instance, the sequential source data can be ordered over one or more dimensions. As one example, the sequential source data can be a one-dimensional sequence (e.g., a sequence of image frames comprising video data). As another example, the sequential source data can be a multidimensional sequence, such as a set of video data patches ordered by x-coordinate, y-coordinate, and time.

[0165] At 404, the method 400 can include generating a first representation of one or more first data elements of the plurality of data elements by a first element encoder.Additionally and / or alternatively, at 406, the method 400 can include generating a second representation of a second data element of the plurality of data elements by a second element encoder. The representations of a data element can be a reduced representation of the data element. For instance, the representations of the data element can require fewer computing resources (e.g., memory, bandwidth, compute cycles, etc.) to store, transmit, and / or process than the data element itself. As examples, the representations can be an embedding or an encoding. For instance, in some implementations, the first representation of the one or more first data elements can be a first embedding of the one or more first data elements.Additionally and / or alternatively the second representation of the second data element can be a second embedding of the second data element.

[0166] The one or more first data elements can be ordered earlier in the sequential source data than the second data element. For instance, the one or more first data elements can be a "‘past” or "‘present” data element, whereas the second data element can be a “future” data element. For instance, if the sequential source data is an already-extant video, the one or more first data elements can represent an arbitrary frame or frames of the video, and the second data element can be some portion of the video that would not yet have been played if the video was paused at the frame(s) of the one or more first data elements. In generation tasks, for example, the second data element may have been generated after the one or more first data elements, and / or may have been generated in the same generation task that generated the one or more first data elements. When training on videos, however, the second data element can be used to guide training of the encoder models without encouraging the models to be overly dependent on this future information using the techniques described herein.

[0167] The first element encoder and / or the second element encoder can be any suitable encoder or encoding model. As examples, the encoder can be a video encoder, a video autoencoder, another type of autoencoder, or any other suitable autoencoder. For instance, in some implementations, the encoder can be or can include a neural network, a transformer model, or other suitable machine-learned model. Furthermore, in some implementations, the encoder can be a masked encoder or masked autoencoder.

[0168] In some implementations, the first element encoder and / or the second element encoder can be a same or similar model. For example, in some implementations, both the first element encoder and the second element encoder can be the same model and / or the same instance of model. As another example, in some implementations, the first element encoder may be initialized as two identical instances of a same model. For instance, the first element encoder and the second element encoder can be initialized from a common encoder model. However, in some implementations, the first element encoder and the second element encoder may diverge during training. For example, in some implementations, the second element encoder can be frozen during the training of the first element encoder. For instance, the parameters of the second element encoder may not be modified during some training iterations whereas the parameters of the first element encoder may be updated at some or all training iterations. Freezing the second element encoder can provide reduced occurrences of the training converging to a trivial solution due to the commonality between the element encoders.

[0169] Additionally and / or alternatively, in some implementations, one of the first element encoder or the second element encoder can include smoothed weights of the other of the first element encoder or the second element encoder over one or more previous training iteration. As one example, in some implementations, the second element encoder can utilize a smoothed average of weights of the first element encoder over previous training iterations. The smoothed average of weights can be determined by any suitable smoothing process, such as exponential moving average (EMA) smoothing, simple moving average (SMA) smoothing, or other suitable smoothing techniques. As one example, the second element encoder can be the exponential moving average of weights of the first element encoder over previous batches of data, such that the second element encoder ’‘lags” behind the first element encoder. In some cases, using an identical model as the first element encoder and the second element encoder can risk convergence to a trivial solution and introducing a minor variation through smoothing can mitigate this risk.

[0170] At 408, the method 400 can include generating a latent projection based on the second representation. The latent projection can be or can include the second representation projected into a lower-dimensional space. The latent projection can convey some limited information about the second representation, but less than the entirety of the information conveyed by the second representation. For example, the latent projection may capture high-level features of the second representation. The latent projection can be, for example, a latent vector. As another example, the latent projection can include other suitable description-level information about the second representation or second data element such as, for example, a latent description of the second element (e.g., a text data description).

[0171] The latent projection can be generated using a projection generation model (which may also be referred to as a "‘projector’ or “projector model”). The projection generation model can be any suitable model or machine-learned model such as, for example, a neural -network or neural-network based model. The projection generation model may be configured, for example, to perform regression, classification, anomaly detection, clustering, or other projection techniques.

[0172] At 410, the method 400 can include generating a predicted second representation based on the first representation and the latent projection by a representation predictor model. The representation predictor model can be any suitable model or machine-learned model, such as, for example, a transformer model, such as a vision transformer (ViT) model. The representation predictor model can be configured to predict future data in the time series data provided as input at the representation level (e g., embedding level). Because of the wide variety of potential futures at any point in time-series data, the representation predictor model can utilize the latent projection to collapse these possible futures to a single possibility and in turn provide more focused predicted second representations. Generally, the predicted second representation can approximate the second representation due in part to the information available in the latent projection. Intuitively, the information available in the latent projection can provide for the representation predictor model to “select” a specific future by “leaking” limited information from the “true” future. However, the predicted second representation may not necessarily be identical to the second representation, as the entire second representation is not available to the representation prediction model.

[0173] At 412, the method 400 can include determining a loss between the second representation and the predicted second representation. The loss can be any suitable loss representative of differences between the second representation and the predicted second representation with respect to training objectives. The loss can, for example, be based on adot product between the second representation and the predicted second representation. Additionally and / or alternatively, at 414, the method 400 can include training at least the first element encoder based on the loss. For instance, the loss can be used to compute one or more deltas or updates for one or more parameters of the first element encoder. The deltas or updates can be applied to update the parameters of the first element encoder. Other machine-learned components, such as the projection generation model, the latent predictor model, and / or the second element encoder can be trained using the loss in addition to or alternatively to the first element encoder.

[0174] In some implementations, training the first element encoder can be performed over multiple training iterations or batches. For instance, the second representation can be input into the representation predictor model (e.g., in place of the first representation generated by the first element encoder) and used to generate additional subsequent representations, such as a third representation. This process can be repeated to generate a plurality of predicted representations. A loss from each of these predictions can be evaluated and accumulated over a batch. The batch may be selected from a single sequence or parts of multiple sequences. Furthermore, in some implementations, multiple batches may be generated, and at each step of the training process, a total loss of a selected batch (e g., a randomly selected batch for stochastic gradient descent) can be reduced.

[0175] After training the model(s), the systems and methods described herein can be used for a variety7of inference-time objectives. For instance, the present disclosure can find applications in a number of computing-technology related fields, such as, for example, data representation tasks (e.g., self-supervised video representations), data generation tasks (e.g., video data generation from a description, outline, or other suitable latent data) and / or timeseries data forecasting tasks. For instance, the present disclosure can provide improved video understanding, such as providing for understanding non-verbal concepts directly from video data. As another example, the present disclosure can provide improved multi-resolution planning and control in dynamic environments for robotics control. As another example, the present disclosure can provide for enhanced prediction of future events for vehicle navigation (e.g., autonomous driving). As another example, the present disclosure can provide for improved generation of realistic and coherent video sequences.

[0176] As one example, the first element encoder can be used to generate representations of new data. For instance, the systems and methods described herein can further provide for obtaining an inference data element and. subsequent to training at least the first element encoder based on the loss, generating a first inference representation of theinference data element using the first element encoder. This first inference representation itself can be a valuable output. For instance, the first inference representation represents an improvement over existing video compression tasks by improvedly capturing spatiotemporal aspects of the inference data element that provide for understanding of actor behavior in the inference data element (e.g., video data).

[0177] Furthermore, in some implementations, this first inference representation can be used in generation or forecasting tasks. For example, the systems and methods described herein can further provide for generating a second inference representation using the representation predictor model based on the first inference representation and / or generating a second inference data element based on the second inference representation. For instance, the second inference representation can be generated using the representation prediction model. Due in part to the use of a latent projection during training, the representation prediction model can more accurately generate predicted representations that converge to a single “selected” future rather than being blurred across a plurality of potential futures.

[0178] In some implementations, generating the second inference representation is based on a guidance input. The guidance input can be similar to the latent projection that was used to train the models described herein. However, rather than being a projection of “true” second representations, the guidance input can be other guiding or seeding data that encourages the second inference representation to converge to a single future. In some implementations, for example, the guidance input can be noise data. The noise data can provide for some randomness in generation while still encouraging convergence to a single future. As another example, in some implementations, the guidance data can be a description of the second inference data element. For instance, in some implementations such as those where the latent projection includes a latent description, the description can guide the models to generate second inference representations that adhere to the high-level description of what the predicted data should contain.

[0179] FIG. 5 depicts a flowchart of a method 500 for training an element encoder according to aspects of the present disclosure. For instance, an example agent system can include one or more machine-learned models and / or other systems configured to perform tasks in response to a query from a user.

[0180] One or more portion(s) of example method 500 can be implemented by a computing system that includes one or more computing devices such as, for example, computing systems described with reference to the other figures. Each respective portion of example method 500 can be performed by any (or any combination) of one or morecomputing devices. Moreover, one or more portion(s) of example method 500 can be implemented on the hardware components of the device(s) described herein, for example, to train one or more systems or models. FIG. 5 depicts elements performed in a particular order for purposes of illustration and discussion. Those of ordinary skill in the art, using the disclosures provided herein, will understand that the elements of any of the methods discussed herein can be adapted, rearranged, expanded, omitted, combined, or modified in various ways without deviating from the scope of the present disclosure. FIG. 5 is described with reference to elements / terms described with respect to other systems and figures for exemplary illustrated purposes and is not meant to be limiting. One or more portions of example method 500 can be performed additionally, or alternatively, by other systems.

[0181] At 502, the method 500 can include obtaining video data comprising a plurality of video data elements. Each video data element can respectively include at least a portion of one or more video frames. At 504, the method 500 can include generating a first representation of a first video data element of the plurality of video data elements by a first element encoder. At 506, the method 500 can include generating a second representation of a second video data element of the plurality of data elements by a second element encoder. The second data element can include data ordered later in the video data than the one or more first data elements. At 508, the method 500 can include generating a latent projection based on the second representation. The latent projection can be descriptive of the second video data element. At 510. the method 500 can include generating a predicted second representation based on the first representation and the latent projection by a representation predictor model. The predicted second representation can approximate the second representation. At 512, the method 500 can include determining a loss between the second representation and the predicted second representation. At 514. the method 500 can include training at least the first element encoder based on the loss.

[0182] FIG. 6 depicts a flowchart of a method 1000 for training one or more machine-learned models according to aspects of the present disclosure. For instance, an example machine-learned model can include a sequence processing model.

[0183] One or more portion(s) of example method 1000 can be implemented by a computing system that includes one or more computing devices such as, for example, computing systems described with reference to the other figures. Each respective portion of example method 1000 can be performed by any (or any combination) of one or more computing devices. Moreover, one or more portion(s) of example method 1000 can be implemented on the hardware components of the device(s) described herein, for example, totrain one or more systems or models. FIG. 6 depicts elements performed in a particular order for purposes of illustration and discussion. Those of ordinary skill in the art, using the disclosures provided herein, will understand that the elements of any of the methods discussed herein can be adapted, rearranged, expanded, omitted, combined, or modified in various ways without deviating from the scope of the present disclosure. FIG. 6 is described with reference to elements / terms described with respect to other systems and figures for exemplary illustrated purposes and is not meant to be limiting. One or more portions of example method 1000 can be performed additionally, or alternatively, by other systems.

[0184] At 1002, example method 1000 can include obtaining a training instance. A set of training data can include a plurality of training instances divided between multiple datasets (e.g., a training dataset, a validation dataset, or testing dataset). A training instance can be labeled or unlabeled. Although referred to in example method 1000 as a “training’’ instance, it is to be understood that runtime inferences can form training instances when a model is trained using an evaluation of the model's performance on that runtime instance (e.g., online training / leaming). Example data types for the training instance and various tasks associated therewith are described throughout the present disclosure.

[0185] At 1004, example method 1000 can include processing, using one or more machine-learned models, the training instance to generate an output. The output can be directly obtained from the one or more machine-learned models or can be a downstream result of a chain of processing operations that includes an output of the one or more machine-learned models.

[0186] At 1006, example method 1000 can include receiving an evaluation signal associated with the output. The evaluation signal can be obtained using a loss function. Various determinations of loss can be used, such as mean squared error, likelihood loss, cross entropy loss, hinge loss, contrastive loss, or various other loss functions. The evaluation signal can be computed using known ground-truth labels (e.g., supervised learning), predicted or estimated labels (e.g., semi- or self-supervised learning), or without labels (e.g., unsupervised learning). The evaluation signal can be a reward (e.g., for reinforcement learning). The reward can be computed using a machine-learned reward model configured to generate rewards based on output(s) received. The reward can be computed using feedback data describing human feedback on the output(s).

[0187] At 1008, example method 1000 can include updating the machine-learned model using the evaluation signal. For example, values for parameters of the machine-learned model(s) can be learned, in some embodiments, using various training or learning techniques,such as, for example, backwards propagation. For example, the evaluation signal can be backpropagated from the output (or another source of the evaluation signal) through the machine-learned model(s) to update one or more parameters of the model(s) (e.g., based on a gradient of the evaluation signal with respect to the parameter value(s)). For example, system(s) containing one or more machine-learned models can be trained in an end-to-end manner. Gradient descent techniques can be used to iteratively update the parameters over a number of training iterations. In some implementations, performing backwards propagation of errors can include performing truncated backpropagation through time. Example method 1000 can include implementing a number of generalization techniques (e.g., weight decays, dropouts, etc.) to improve the generalization capability of the models being trained.

[0188] In some implementations, example method 1000 can be implemented for training a machine-learned model from an initialized state to a fully trained state (e.g., when the model exhibits a desired performance profile, such as based on accuracy, precision, recall, etc.).

[0189] In some implementations, example method 1000 can be implemented for particular stages of a training procedure. For instance, in some implementations, example method 1000 can be implemented for pre-training a machine-learned model. Pre-training can include, for instance, large-scale training over potentially noisy data to achieve a broad base of performance levels across a variety of tasks / data types. In some implementations, example method 1000 can be implemented for fine-tuning a machine-learned model. Fine-tuning can include, for instance, smaller-scale training on higher-quality (e.g., labeled, curated, etc.) data. Fine-tuning can affect all or a portion of the parameters of a machine-learned model. For example, various portions of the machine-learned model can be “frozen’" for certain training stages. For example, parameters associated with an embedding space can be “frozen” during fine-tuning (e.g., to retain information learned from a broader domain(s) than present in the fine-tuning dataset(s)). An example fine-tuning approach includes reinforcement learning. Reinforcement learning can be based on user feedback on model performance during use.Example Machine-learned Models

[0190] FIG. 7 is a block diagram of an example processing flow for using machine-learned model(s) 1 to process input(s) 2 to generate output(s) 3.

[0191] Machine-learned model(s) 1 can be or include one or multiple machine-learned models or model components. Example machine-learned models can include neuralnetworks (e.g., deep neural networks). Example machine-learned models can include nonlinear models or linear models. Example machine-learned models can use other architectures in lieu of or in addition to neural networks. Example machine-learned models can include decision tree based models, support vector machines, hidden Markov models, Bayesian networks, linear regression models, k-means clustering models, etc.

[0192] Example neural networks can include feed-forward neural networks, recurrent neural networks (RNNs), including long short-term memory (LSTM) based recurrent neural networks, convolutional neural networks (CNNs), diffusion models, generative-adversarial networks, or other forms of neural networks. Example neural networks can be deep neural networks. Some example machine-learned models can leverage an attention mechanism such as self-attention. For example, some example machine-learned models can include multiheaded self-attention models. For example, the machine-learned models can be or include transformer models.

[0193] Machine-learned model(s) 1 can include a single or multiple instances of the same model configured to operate on data from input(s) 2. Machine-learned model(s) 1 can include an ensemble of different models that can cooperatively interact to process data from input(s) 2. For example, machine-learned model(s) 1 can employ a mixture-of-experts structure. See, e.g., Zhou et al., Mixture-of-Experts with Expert Choice Routing, ARXIV:2202.09368V2 (Oct. 14, 2022).

[0194] Input(s) 2 can generally include or otherwise represent various types of data. Input(s) 2 can include one type or many different types of data. Output(s) 3 can be data of the same type(s) or of different types of data as compared to input(s) 2. Output(s) 3 can include one type or many different types of data.

[0195] Example data types for input(s) 2 or output(s) 3 include natural language text data, software code data (e.g., source code, object code, machine code, or any other form of computer-readable instructions or programming languages), machine code data (e.g., binary code, assembly code, or other forms of machine-readable instructions that can be executed directly by a computer’s central processing unit), assembly code data (e.g., low-level programming languages that use symbolic representations of machine code instructions to program a processing unit), genetic data or other chemical or biochemical data, image data, audio data, audiovisual data, haptic data, biometric data, medical data, financial data, statistical data, geographical data, astronomical data, historical data, sensor data generally (e.g., digital or analog values, such as voltage or other absolute or relative level measurementvalues from a real or artificial input, such as from an audio sensor, light sensor, displacement sensor, etc.), and the like. Data can be raw or processed and can be in any format or schema.

[0196] In multimodal inputs 2 or outputs 3, example combinations of data types include image data and audio data, image data and natural language data, natural language data and software code data, image data and biometric data, sensor data and medical data, etc. It is to be understood that any combination of data types in an input 2 or an output 3 can be present.

[0197] An example input 2 can include one or multiple data types, such as the example data types noted above. An example output 3 can include one or multiple data types, such as the example data types noted above. The data type(s) of input 2 can be the same as or different from the data type(s) of output 3. It is to be understood that the example data types noted above are provided for illustrative purposes only. Data types contemplated within the scope of the present disclosure are not limited to those examples noted above.Example Machine-Learned Sequence Processing Models

[0198] FIG. 8 is a block diagram of an example implementation of an example machine-learned model configured to process sequences of information. For instance, an example implementation of machine-learned model(s) 1 can include machine-learned sequence processing model(s) 4. An example system can pass input(s) 2 to sequence processing model(s) 4. Sequence processing model(s) 4 can include one or more machine-learned components. Sequence processing model(s) 4 can process the data from input(s) 2 to obtain an input sequence 5. Input sequence 5 can include one or more input elements 5-1, 5-2, . . . , 5-M, etc. obtained from input(s) 2. Sequence processing model 4 can process input sequence 5 using prediction layer(s) 6 to generate an output sequence 7. Output sequence 7 can include one or more output elements 7-1, 7-2, . . . , 7-A, etc. generated based on input sequence 5. The system can generate output(s) 3 based on output sequence 7.

[0199] Sequence processing model (s) 4 can include one or multiple machine-learned model components configured to ingest, generate, or otherwise reason over sequences of information. For example, some example sequence processing models in the text domain are referred to as “Large Language Models,” or LLMs. See, e.g., PaLM 2 Technical Report, GOOG E, https: / / ai.google / static / documents / palm2techreport.pdf (n.d.). Other example sequence processing models can operate in other domains, such as image domains, see. e.g., Dosovitskiy et al., An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale, ARXIV:2010.11929v2 (Jun. 3, 2021), audio domains, see, e.g., Agostinelli et al.,MusicLM: Generating Music From Text, ARXIV:2301.11325V! (Jan. 26, 2023), biochemical domains, see. e.g, Jumper et al.. Highly accurate protein structure prediction with AlphaFold, 596 Nature 583 (Aug. 26, 2021), by way of example. Sequence processing model(s) 4 can process one or multiple types of data simultaneously. Sequence processing model(s) 4 can include relatively large models (e.g., more parameters, computationally expensive, etc.), relatively small models (e.g., fewer parameters, computationally lightweight, etc ), or both.

[0200] In general, sequence processing model(s) 4 can obtain input sequence 5 using data from input(s) 2. For instance, input sequence 5 can include a representation of data from input(s) 2 in a format understood by sequence processing model(s) 4. One or more machine-learned components of sequence processing model(s) 4 can ingest the data from input(s) 2, parse the data into pieces compatible with the processing architectures of sequence processing model(s) 4 (e.g., via “tokenization”), and project the pieces into an input space associated with prediction layer(s) 6 (e.g., via “embedding”).

[0201] Sequence processing model(s) 4 can ingest the data from input(s) 2 and parse the data into a sequence of elements to obtain input sequence 5. For example, a portion of input data from input(s) 2 can be broken down into pieces that collectively represent the content of the portion of the input data. The pieces can provide the elements of the sequence.

[0202] Elements 5-1, 5-2, . . . , 5-M can represent, in some cases, building blocks for capturing or expressing meaningful information in a particular data domain. For instance, the elements can describe “atomic units” across one or more domains. For example, for textual input source(s), the elements can correspond to groups of one or more words or sub-word components, such as sets of one or more characters.

[0203] For example, elements 5-1, 5-2, . . . , 5-M can represent tokens obtained using a tokenizer. For instance, a tokenizer can process a given portion of an input source and output a series of tokens (e.g., corresponding to input elements 5-1, 5-2, . . . , 5-M) that represent the portion of the input source. Various approaches to tokenization can be used. For instance, textual input source(s) can be tokenized using a byte-pair encoding (BPE) technique. See, e.g., Kudo et al., SentencePiece: A simple and language independent subword tokenizer and detokenizer for Neural Text Processing, PROCEEDINGS OF THE 2018 CONFERENCE ON EMPIRICAL METHODS IN NATURAL LANGUAGE PROCESSING (System Demonstrations), pages 66-71 (October 31-November 4, 2018), https: / / aclanthology.org / D18-2012.pdf. Image-based input source(s) can be tokenized by extracting and serializing patches from an image. Other tokenization approaches can beperformed as well, including linear projections, non-linear transformations, and / or other data transformations.

[0204] In general, arbitrary data types can be serialized and processed into input sequence 5. It is to be understood that element(s) 5-1, 5-2, . . . , 5-M depicted in FIG. 8 can be the tokens or can be the embedded representations thereof.

[0205] Prediction layer(s) 6 can predict one or more output elements 7-1, 7-2, . . . , 7-N based on the input elements. Prediction layer(s) 6 can include one or more machine-learned model architectures, such as one or more layers of learned parameters that manipulate and transform the input(s) to extract higher-order meaning from, and relationships between, input element(s) 5-1, 5-2, . . . , 5-M. In this manner, for instance, example prediction layer(s) 6 can predict new output element(s) in view of the context provided by input sequence 5.

[0206] Prediction layer(s) 6 can evaluate associations between portions of input sequence 5 and a particular output element. These associations can inform a prediction of the likelihood that a particular output follows the input context. For example, consider the textual snippet, “The carpenter’s toolbox was small and heavy. It was full of .” Example prediction layer(s) 6 can identify that “It” refers back to “toolbox” by determining a relationship between the respective embeddings. Example prediction layer(s) 6 can also link “It” to the attributes of the toolbox, such as “small” and “heavy.” Based on these associations, prediction layer(s) 6 can, for instance, assign a higher probability to the word “nails” than to the word “sawdust.”

[0207] A transformer is an example architecture that can be used in prediction layer(s) 4. See, e.g., Vaswani et al., Attention Is All You Need, ARXlV:1706.03762v7 (Aug. 2, 2023). A transformer is an example of a machine-learned model architecture that uses an attention mechanism to compute associations between items within a context window. The context window can include a sequence that contains input sequence 5 and potentially one or more output element(s) 7-1, 7-2, . . . , 7-N. A transformer block can include one or more attention layer(s) and one or more post-attention layer(s) (e.g., feedforw ard layer(s), such as a multi-layer perceptron).

[0208] Prediction layer(s) 6 can include other machine-learned model architectures in addition to or in lieu of transformer-based architectures. For example, recurrent neural networks (RNNs) and long short-term memory (LSTM) models can also be used, as well as convolutional neural networks (CNNs). In general, prediction layer(s) 6 can leverage various kinds of artificial neural networks that can understand or generate sequences of information.

[0209] Output sequence 7 can include or otherwise represent the same or different data types as input sequence 5. For instance, input sequence 5 can represent textual data, and output sequence 7 can represent textual data. Input sequence 5 can represent image, audio, or audiovisual data, and output sequence 7 can represent textual data (e.g., describing the image, audio, or audiovisual data). It is to be understood that prediction layer(s) 6, and any other interstitial model components of sequence processing model(s) 4. can be configured to receive a variety of data types in input sequence(s) 5 and output a variety of data types in output sequence(s) 7.

[0210] Output sequence 7 can have various relationships to input sequence 5. Output sequence 7 can be a continuation of input sequence 5. Output sequence 7 can be complementary to input sequence 5. Output sequence 7 can translate, transform, augment, or otherwise modify input sequence 5. Output sequence 7 can answer, evaluate, confirm, or otherwise respond to input sequence 5. Output sequence 7 can implement (or describe instructions for implementing) an instruction provided via an input sequence 5.

[0211] Output sequence 7 can be generated autoregressively. For instance, for some applications, an output of one or more prediction layer(s) 6 can be passed through one or more output layers (e g., SofitMax layer) to obtain a probability distribution over an output vocabulary' (e.g., a textual or symbolic vocabulary ) conditioned on a set of input elements in a context window. In this manner, for instance, output sequence 7 can be autoregressively generated by sampling a likely next output element, adding that element to the context window, and re-generating the probability distribution based on the updated context window7, and sampling a likely next output element, and so forth.

[0212] Output sequence 7 can also be generated non-autoregressively. For instance, multiple output elements of output sequence 7 can be predicted together without explicit sequential conditioning on each other. See, e.g., Saharia et al., Non-Autoregressive Machine Translation with Latent Alignments, ARXlV:2004.07437v3 (NOV. 16, 2020).

[0213] Output sequence 7 can include one or multiple portions or elements. In an example content generation configuration, output sequence 7 can include multiple elements corresponding to multiple portions of a generated output sequence (e.g., a textual sentence, values of a discretized w aveform, computer code, etc.). In an example classification configuration, output sequence 7 can include a single element associated with a classification output. For instance, an output “vocabulary"’ can include a set of classes into which an input sequence is to be classified. For instance, a vision transformer block can pass latent stateinformation to a multilayer perceptron that outputs a likely class value associated with an input image.

[0214] FIG. 9 is a block diagram of an example technique for populating an example input sequence 8. Input sequence 8 can include various functional elements that form part of the model infrastructure, such as an element 8-0 obtained from a task indicator 9 that signals to any model(s) that process input sequence 8 that a particular task is being performed (e.g., to help adapt a performance of the model(s) to that particular task). Input sequence 8 can include various data elements from different data modalities. For instance, an input modality 10-1 can include one modality of data. A data-to-sequence model 11-1 can process data from input modality 10-1 to project the data into a format compatible with input sequence 8 (e.g., one or more vectors dimensioned according to the dimensions of input sequence 8) to obtain elements 8-1, 8-2, 8-3. Another input modality 10-2 can include a different modality of data. A data-to-sequence model 11-2 can project data from input modality 10-2 into a format compatible with input sequence 8 to obtain elements 8-4, 8-5, 8-6. Another input modality 10-3 can include yet another different modality of data. A data-to-sequence model 11-3 can project data from input modality 10-3 into a format compatible with input sequence 8 to obtain elements 8-7, 8-8, 8-9.

[0215] Input sequence 8 can be the same as or different from input sequence 5. Input sequence 8 can be a multimodal input sequence that contains elements that represent data from different modalities using a common dimensional representation. For instance, an embedding space can have P dimensions. Input sequence 8 can be configured to contain a plurality of elements that have P dimensions. In this manner, for instance, example implementations can facilitate information extraction and reasoning across diverse data modalities by projecting data into elements in the same embedding space for comparison, combination, or other computations therebetween.

[0216] For example, elements 8-0, . . . , 8-9 can indicate particular locations within a multidimensional embedding space. Some elements can map to a set of discrete locations in the embedding space. For instance, elements that correspond to discrete members of a predetermined vocabulary of tokens can map to discrete locations in the embedding space that are associated with those tokens. Other elements can be continuously distributed across the embedding space. For instance, some datatypes can be broken down into continuously defined portions (e.g., image patches) that can be described using continuously distributed locations within the embedding space.

[0217] In some implementations, the expressive power of the embedding space may not be limited to meanings associated with any particular set of tokens or other building blocks. For example, a continuous embedding space can encode a spectrum of high-order information. An individual piece of information (e.g., a token) can map to a particular point in that space: for instance, a token for the word “dog” can be projected to an embedded value that points to a particular location in the embedding space associated with canine-related information. Similarly, an image patch of an image of a dog on grass can also be projected into the embedding space. In some implementations, the projection of the image of the dog can be similar to the projection of the word “dog” while also having similarity to a projection of the word “grass,” while potentially being different from both. In some implementations, the projection of the image patch may not exactly align with any single projection of a single word. In some implementations, the projection of the image patch can align with a combination of the projections of the words “dog” and “grass.” In this manner, for instance, a high order embedding space can encode information that can be independent of data modalities in which the information is expressed.

[0218] Task indicator 9 can include a model or model component configured to identify a task being performed and inject, into input sequence 8, an input value represented by element 8-0 that signals which task is being performed. For instance, the input value can be provided as a data type associated with an input modality and projected along with that input modality (e.g., the input value can be a textual task label that is embedded along with other textual data in the input; the input value can be a pixel-based representation of a task that is embedded along with other image data in the input; etc.). The input value can be provided as a data type that differs from or is at least independent from other input(s). For instance, the input value represented by element 8-0 can be learned within a continuous embedding space.

[0219] Input modalities 10-1, 10-2, and 10-3 can be associated with various different data types (e.g., as described above with respect to input(s) 2 and output(s) 3).

[0220] Data-to-sequence models 11-1. 11-2. and 11-3 can be the same or different from each other. Data-to-sequence models 11-1, 11-2, and 11-3 can be adapted to each respective input modality 10-1, 10-2, and 10-3. For example, a textual data-to-sequence model can subdivide a portion of input text and project the subdivisions into element(s) in input sequence 8 (e.g., elements 8-1. 8-2, 8-3, etc.). An image data-to-sequence model can subdivide an input image and project the subdivisions into element(s) in input sequence 8 (e.g., elements 8-4, 8-5, 8-6, etc.). An arbitrary datatype data-to-sequence model cansubdivide an input of that arbitrary datatype and project the subdivisions into element(s) in input sequence 8 (e.g., elements 8-7. 8-8, 8-9, etc.).

[0221] Data-to-sequence models 11-1, 11-2, and 11-3 can form part of machine-learned sequence processing model (s) 4. Data-to-sequence models 11-1, 11-2, and 11-3 can be jointly trained with or trained independently from machine-learned sequence processing model(s) 4. Data-to-sequence models 11-1, 11-2, and 11-3 can be trained end-to-end with machine-learned sequence processing model(s) 4.Example Machine-learned Model Development Platform

[0222] FIG. 10 is a block diagram of an example model development platform 12 that can facilitate creation, adaptation, and refinement of example machine-learned models (e.g., machine-learned model(s) 1, sequence processing model(s) 4, etc.). Model development platform 12 can provide a number of different toolkits that developer systems can employ in the development of new or adapted machine-learned models.

[0223] Model development platform 12 can provide one or more model libraries 13 containing building blocks for new models. Model libraries 13 can include one or more pretrained foundational models 13-1, which can provide a backbone of processing power across various tasks. Model libraries 13 can include one or more pre-trained expert models 13-2, which can be focused on performance in particular domains of expertise. Model libraries 13 can include various model primitives 13-3, which can provide low-level architectures or components (optionally pre-trained), which can be assembled in various arrangements as desired.

[0224] Model development platform 12 can receive selections of various model components 14. Model development platform 12 can pass selected model components 14 to a workbench 15 that combines selected model components 14 into a development model 16.

[0225] Workbench 15 can facilitate further refinement and adaptation of development model 16 by leveraging a number of different toolkits integrated with model development platform 12. For example, workbench 15 can facilitate alignment of the development model 16 with a desired performance profile on various tasks using a model alignment toolkit 17.

[0226] Model alignment toolkit 17 can provide a number of tools for causing development model 16 to generate outputs aligned with desired behavioral characteristics. Alignment can include increasing an accuracy, precision, recall, etc. of model outputs.Alignment can include enforcing output styles, schema, or other preferential characteristics of model outputs. Alignment can be general or domain specific. For instance, a pre-trainedfoundational model 13-1 can begin with an initial level of performance across multiple domains. Alignment of the pre-trained foundational model 13-1 can include improving a performance in a particular domain of information or tasks (e.g., even at the expense of performance in another domain of information or tasks).

[0227] Model alignment toolkit 17 can integrate one or more dataset(s) 17-1 for aligning development model 16. Curated dataset(s) 17-1 can include labeled or unlabeled training data. Dataset(s) 17-1 can be obtained from public domain datasets. Dataset(s) 17-1 can be obtained from private datasets associated with one or more developer system(s) for the alignment of bespoke machine-learned model(s) customized for private use-cases.

[0228] Pre-training pipelines 17-2 can include a machine-learned model training workflow configured to update development model 16 over large-scale, potentially noisy datasets. For example, pre-training can leverage unsupervised learning techniques (e.g., denoising, etc.) to process large numbers of training instances to update model parameters from an initialized state and achieve a desired baseline performance. Pre- training pipelines 17-2 can leverage unlabeled datasets in dataset(s) 17-1 to perform pre-training. Workbench 15 can implement a pre-training pipeline 17-2 to pre-train development model 16.

[0229] Fine-tuning pipelines 17-3 can include a machine-learned model training workflow configured to refine the model parameters of development model 16 with higher-quality data. Fine-tuning pipelines 17-3 can update development model 16 by conducting supervised training with labeled dataset(s) in dataset(s) 17-1. Fine-tuning pipelines 17-3 can update development model 16 by conducting reinforcement learning using reward signals from user feedback signals. Workbench 15 can implement a fine-tuning pipeline 17-3 to finetune development model 16.

[0230] Prompt libraries 17-4 can include sets of inputs configured to induce behavior aligned with desired performance criteria. Prompt libraries 17-4 can include few-shot prompts (e.g., inputs providing examples of desired model outputs for prepending to a desired runtime query), chain-of-thought prompts (e.g., inputs providing step-by-step reasoning within the exemplars to facilitate thorough reasoning by the model), and the like.

[0231] Example prompts can be retrieved from an available repository of prompt libraries 17-4. Example prompts can be contributed by one or more developer systems using workbench 15.

[0232] In some implementations, pre-trained or fine-tuned models can achieve satisfactory performance without exemplars in the inputs. For instance, zero-shot prompts caninclude inputs that lack exemplars. Zero-shot prompts can be within a domain within a training dataset or outside of the training domain(s).

[0233] Prompt libraries 17-4 can include one or more prompt engineering tools. Prompt engineering tools can provide workflows for retrieving or learning optimized prompt values. Prompt engineering tools can facilitate directly learning prompt values (e.g., input element values) based on one or more training iterations. Workbench 15 can implement prompt engineering tools in development model 16.

[0234] Prompt libraries 17-4 can include pipelines for prompt generation. For example, inputs can be generated using development model 16 itself or other machine-learned models. In this manner, for instance, a first model can process information about a task and output an input for a second model to process in order to perform a step of the task. The second model can be the same as or different from the first model. Workbench 15 can implement prompt generation pipelines in development model 16.

[0235] Prompt libraries 17-4 can include pipelines for context injection. For instance, a performance of development model 16 on a particular task can improve if provided with additional context for performing the task. Prompt libraries 17-4 can include software components configured to identify desired context, retrieve the context from an external source (e.g., a database, a sensor, etc.), and add the context to the input prompt. Workbench 15 can implement context injection pipelines in development model 16.

[0236] Although various training examples described herein with respect to model development platform 12 refer to ‘'pre-training” and “fine-tuning,” it is to be understood that model alignment toolkit 17 can generally support a wide variety of training techniques adapted for training a wide variety of machine-learned models. Example training techniques can correspond to the example training method 1000 described above.

[0237] Model development platform 12 can include a model plugin toolkit 18. Model plugin toolkit 18 can include a variety of tools configured for augmenting the functionality' of a machine-learned model by integrating the machine-learned model with other systems, devices, and software components. For instance, a machine-learned model can use tools to increase performance quality where appropriate. For instance, deterministic tasks can be offloaded to dedicated tools in lieu of probabilistically performing the task with an increased risk of error. For instance, instead of autoregressively predicting the solution to a system of equations, a machine-learned model can recognize a tool to call for obtaining the solution and pass the system of equations to the appropriate tool. The tool can be a traditional system of equations solver that can operate deterministically to resolve the system of equations. Asanother example, a tool can have a machine-learned model with reduced model overhead compared to a larger machine-learned model. For instance, the model of the tool can be a less sophisticated model than the calling model that is specialized to a particular task or subset of tasks and can require fewer computing resources to produce a usable output. The output of the tool can be returned in response to the original query'. In this manner, tool use can allow some example models to focus on the strengths of machine-learned models — e g., understanding an intent in an unstructured request for a task — while augmenting the performance of the model by offloading certain tasks to a more focused tool for rote application of deterministic algorithms to a well-defined problem or for evaluation of simpler tasks that can be adequately performed by a less sophisticated model.

[0238] Model plugin toolkit 18 can include validation tools 18-1. Validation tools 18-1 can include tools that can parse and confirm output(s) of a machine-learned model.Validation tools 18-1 can include engineered heuristics that establish certain thresholds applied to model outputs. For example, validation tools 18-1 can ground the outputs of machine-learned models to structured data sources (e.g., to mitigate "hallucinations"). One example tool that can be included in validation tools 18-1 is a routing tool for routing a query from a user to a user-specific memory layer or a public data interface.

[0239] Model plugin toolkit 18 can include tooling packages 18-2 for implementing one or more tools that can include scripts or other executable code that can be executed alongside development model 16. Tooling packages 18-2 can include one or more inputs configured to cause machine-learned model(s) to implement the tools (e g., few-shot prompts that induce a model to output tool calls in the proper syntax, etc.). Tooling packages 18-2 can include, for instance, fine-tuning training data for training a model to use a tool.

[0240] Model plugin toolkit 18 can include interfaces for calling external application programming interfaces (APIs) 18-3. For instance, in addition to or in lieu of implementing tool calls or tool code directly with development model 16, development model 16 can be aligned to output instructions that initiate API calls to send or obtain data via external systems. As an example, the development model 16 can initiate API calls to one or more public data interface(s) to send or obtain data from one or more public data sources, such as webpages, databases, and so on.

[0241] Model plugin toolkit 18 can integrate with prompt libraries 17-4 to build a catalog of available tools for use with development model 16. For instance, a model can receive, in an input, a catalog of available tools, and the model can generate an output that selects a tool from the available tools and initiates a tool call for using the tool.

[0242] Model development platform 12 can include a computational optimization toolkit 19 for optimizing a computational performance of development model 16. For instance, tools for model compression 19-1 can allow development model 16 to be reduced in size while maintaining a desired level of performance. For instance, model compression 19-1 can include quantization workflows, weight pruning and sparsification techniques, etc. Tools for hardware acceleration 19-2 can facilitate the configuration of the model storage and execution formats to operate optimally on different hardware resources. For instance, hardware acceleration 19-2 can include tools for optimally sharding models for distributed processing over multiple processing units for increased bandwidth, lower unified memory requirements, etc. Tools for distillation 19-3 can provide for the training of lighter-weight models based on the knowledge encoded in development model 16. For instance, development model 16 can be a highly performant, large machine-learned model optimized using model development platform 12. To obtain a lightweight model for running in resource-constrained environments, a smaller model can be a “student model” that learns to imitate development model 16 as a “teacher model.” In this manner, for instance, the investment in learning the parameters and configurations of development model 16 can be efficiently transferred to a smaller model for more efficient inference.

[0243] Workbench 15 can implement one, multiple, or none of the toolkits implemented in model development platform 12. Workbench 15 can output an output model 20 based on development model 16. Output model 20 can be a deployment version of development model 16. Output model 20 can be a development or training checkpoint of development model 16. Output model 20 can be a distilled, compressed, or otherwise optimized version of development model 16.

[0244] FIG. 11 is a block diagram of an example training flow for training a machine-learned development model 16. One or more portion(s) of the example training flow can be implemented by a computing system that includes one or more computing devices such as, for example, computing systems described with reference to the other figures. Each respective portion of the example training flow can be performed by any (or any combination) of one or more computing devices. Moreover, one or more portion(s) of the example training flow can be implemented on the hardware components of the device(s) described herein, for example, to train one or more systems or models. FIG. 11 depicts elements performed in a particular order for purposes of illustration and discussion. Those of ordinary skill in the art, using the disclosures provided herein, will understand that the elements of any of the methods discussed herein can be adapted, rearranged, expanded.omitted, combined, or modified in various ways without deviating from the scope of the present disclosure. FIG. 11 is described with reference to elements / terms described with respect to other systems and figures for exemplar}' illustrated purposes and is not meant to be limiting. One or more portions of the example training flow can be performed additionally, or alternatively, by other systems.

[0245] Initially, development model 16 can persist in an initial state as an initialized model 21. Development model 16 can be initialized with weight values. Initial weight values can be random or based on an initialization schema. Initial weight values can be based on prior pre-training for the same or for a different model.

[0246] Initialized model 21 can undergo pre-training in a pre-training stage 22. Pretraining stage 22 can be implemented using one or more pre-training pipelines 17-2 over data from dataset(s) 17-1. Pre-training can be omitted, for example, if initialized model 21 is already pre-trained (e.g., development model 16 contains, is, or is based on a pre-trained foundational model or an expert model).

[0247] Pre-trained model 23 can then be a new version of development model 16, which can persist as development model 16 or as a new development model. Pre-trained model 23 can be the initial state if development model 16 was already pre-trained. Pre-trained model 23 can undergo fine-tuning in a fine-tuning stage 24. Fine-tuning stage 24 can be implemented using one or more fine-tuning pipelines 17-3 over data from dataset(s) 17-1. Fine-tuning can be omitted, for example, if a pre-trained model has satisfactory performance, if the model was already fine-tuned, or if other tuning approaches are preferred.

[0248] Fine-tuned model 29 can then be a new version of development model 16, which can persist as development model 16 or as a new development model. Fine-tuned model 29 can be the initial state if development model 16 was already fine-tuned. Fine-tuned model 29 can undergo refinement with user feedback 26. For instance, refinement with user feedback 26 can include reinforcement learning, optionally based on human feedback from human users of fine-tuned model 25. As reinforcement learning can be a form of fine-tuning, it is to be understood that fine-tuning stage 24 can subsume the stage for refining with user feedback 26. Refinement with user feedback 26 can produce a refined model 27. Refined model 27 can be output to downstream system(s) 28 for deployment or further development.

[0249] In some implementations, computational optimization operations can be applied before, during, or after each stage. For instance, initialized model 21 can undergo computational optimization 29-1 (e.g.. using computational optimization toolkit 19) before pre-training stage 22. Pre-trained model 23 can undergo computational optimization 29-2(e.g., using computational optimization toolkit 19) before fine-tuning stage 24. Fine-tuned model 25 can undergo computational optimization 29-3 (e.g., using computational optimization toolkit 19) before refinement with user feedback 26. Refined model 27 can undergo computational optimization 29-4 (e.g., using computational optimization toolkit 19) before output to downstream system(s) 28. Computational optimization(s) 29-1, . . . . 29-4 can all be the same, all be different, or include at least some different optimization techniques.Example Machine-learned Model Inference System

[0250] FIG. 12 is a block diagram of an inference system for operating one or more machine-learned model(s) 1 to perform inference (e.g., for training, for deployment, etc.). A model host 31 can receive machine-learned model(s) 1. Model host 31 can host one or more model instance(s) 31-1, which can be one or multiple instances of one or multiple models. Model host 31 can host model instance(s) 31-1 using available compute resources 31-2 associated with model host 31.

[0251] Model host 31 can perform inference on behalf of one or more client(s) 32. Client(s) 32 can transmit an input request 33 to model host 31. Using input request 33, model host 31 can obtain input(s) 2 for input to machine-learned model(s) 1. Machine-learned model(s) 1 can process input(s) 2 to generate output(s) 3. Using output(s) 3, model host 31 can return an output payload 34 for responding to input request 33 from client(s) 32. Output payload 34 can include or be based on output(s) 3.

[0252] Model host 31 can leverage various other resources and tools to augment the inference task. For instance, model host 31 can communicate with tool interfaces 35 to facilitate tool use by model instance(s) 31-1. Tool interfaces 35 can include local or remote APIs. Tool interfaces 35 can include integrated scripts or other software functionality. Model host 31 can engage online learning interface(s) 36 to facilitate ongoing improvements to machine-learned model(s) 1. For instance, online learning interface(s) 36 can be used within reinforcement learning loops to retrieve user feedback on inferences served by model host 31. Model host 31 can access runtime data source(s) 37 for augmenting input(s) 2 with additional contextual information. For instance, runtime data source(s) 37 can include a knowledge graph 37-1 that facilitates structured information retrieval for information associated with input request(s) 33 (e.g., a search engine service). Runtime data source(s) 37 can include public or private, external or local database(s) 37-2 that can store information associated with input request(s) 33 for augmenting input(s) 2. Runtime data source(s) 37 can include accountdata 37-3 which can be retrieved in association with a user account corresponding to a client 32 for customizing the behavior of model host 31 accordingly.

[0253] Model host 31 can be implemented by one or multiple computing devices or systems. Client(s) 2 can be implemented by one or multiple computing devices or systems, which can include computing devices or systems shared with model host 31.

[0254] For example, model host 31 can operate on a server system that provides a machine-learning service to client device(s) that operate client(s) 32 (e.g., over a local or wide-area network). Client device(s) can be end-user devices used by individuals. Client device(s) can be server systems that operate client(s) 32 to provide various functionality as a sen-ice to downstream end-user devices.

[0255] In some implementations, model host 31 can operate on a same device or system as client(s) 32. Model host 31 can be a machine-learning service that runs on-device to provide machine-learning functionality to one or multiple applications operating on a client device, which can include an application implementing client(s) 32. Model host 31 can be a part of a same application as client(s) 32. For instance, model host 31 can be a subroutine or method implemented by one part of an application, and client(s) 32 can be another subroutine or method that engages model host 31 to perform inference functions within the application. It is to be understood that model host 31 and client(s) 32 can have various different configurations.

[0256] Model instance(s) 31-1 can include one or more machine-learned models that are available for performing inference. Model instance(s) 31-1 can include weights or other model components that are stored on or in persistent storage, temporarily cached, or loaded into high-speed memory. Model instance(s) 31-1 can include multiple instance(s) of the same model (e.g., for parallel execution of more requests on the same model). Model instance(s) 31-1 can include instance(s) of different model(s). Model instance(s) 31-1 can include cached intermediate states of active or inactive model(s) used to accelerate inference of those models. For instance, an inference session with a particular model can generate significant amounts of computational results that can be re-used for future inference runs (e.g., using a KV cache for transformer-based models). These computational results can be saved in association with that inference session so that session can be executed more efficiently w hen resumed.

[0257] Compute resource(s) 31-2 can include one or more processors (central processing units, graphical processing units, tensor processing units, machine-learning accelerators, etc.) connected to one or more memory devices. Compute resource(s) 31-2 caninclude a dynamic pool of available resources shared with other processes. Compute resource(s) 31-2 can include memory devices large enough to fit an entire model instance in a single memory instance. Compute resource(s) 31-2 can also shard model instance(s) across multiple memoiy devices (e.g., using data parallelization or tensor parallelization, etc.). This can be done to increase parallelization or to execute a large model using multiple memoiy7devices which individually might not be able to fit the entire model into memory.

[0258] Input request 33 can include data for input(s) 2. Model host 31 can process input request 33 to obtain input(s) 2. Input(s) 2 can be obtained directly from input request 33 or can be retrieved using input request 33. Input request 33 can be submitted to model host 31 via an API.

[0259] Model host 31 can perform inference over batches of input requests 33 in parallel. For instance, a model instance 31-1 can be configured with an input structure that has a batch dimension. Separate input(s) 2 can be distributed across the batch dimension (e.g., rows of an array). The separate input(s) 2 can include completely different contexts. The separate input(s) 2 can be multiple inference steps of the same task. The separate input(s) 2 can be staggered in an input structure, such that any given inference cycle can be operating on different portions of the respective input(s) 2. In this manner, for instance, model host 31 can perform inference on the batch in parallel, such that output(s) 3 can also contain the batch dimension and return the inference results for the batched input(s) 2 in parallel. In this manner, for instance, batches of input request(s) 33 can be processed in parallel for higher throughput of output payload(s) 34.

[0260] Output payload 34 can include or be based on output(s) 3 from machine-learned model(s) 1. Model host 31 can process output(s) 3 to obtain output payload 34. This can include chaining multiple rounds of inference (e.g., iteratively, recursively, across the same model(s) or different model(s)) to arrive at a final output for a task to be returned in output payload 34. Output payload 34 can be transmitted to client(s) 32 via an API.

[0261] Online learning interface(s) 36 can facilitate reinforcement learning of machine-learned model(s) 1. Online learning interface(s) 36 can facilitate reinforcement learning with human feedback (RLHF). Online learning interface(s) 36 can facilitate federated learning of machine-learned model(s) 1.

[0262] Model host 31 can execute machine-learned model(s) 1 to perform inference for various tasks using various types of data. For example, various different input(s) 2 and output(s) 3 can be used for various different tasks. In some implementations, input(s) 2 can be or otherwise represent image data. Machine-learned model(s) 1 can process the image datato generate an output. As an example, machine-learned model(s) 1 can process the image data to generate an image recognition output (e.g., a recognition of the image data, a latent embedding of the image data, an encoded representation of the image data, a hash of the image data, etc.). As another example, machine-learned model(s) 1 can process the image data to generate an image segmentation output. As another example, machine-learned model(s) 1 can process the image data to generate an image classification output. As another example, machine-learned model(s) 1 can process the image data to generate an image data modification output (e.g., an alteration of the image data, etc.). As another example, machine-learned model(s) 1 can process the image data to generate an encoded image data output (e.g., an encoded and / or compressed representation of the image data, etc.). As another example, machine-learned model(s) 1 can process the image data to generate an upscaled image data output. As another example, machine-learned model(s) 1 can process the image data to generate a prediction output.

[0263] In some implementations, the task is a computer vision task. In some cases, input(s) 2 includes pixel data for one or more images and the task is an image processing task. For example, the image processing task can be image classification, where the output is a set of scores, each score corresponding to a different object class and representing the likelihood that the one or more images depict an object belonging to the object class. The image processing task can be object detection, where the image processing output identifies one or more regions in the one or more images and. for each region, a likelihood that region depicts an object of interest. As another example, the image processing task can be image segmentation, where the image processing output defines, for each pixel in the one or more images, a respective likelihood for each category' in a predetermined set of categories. For example, the set of categories can be foreground and background. As another example, the set of categories can be object classes. As another example, the image processing task can be depth estimation, where the image processing output defines, for each pixel in the one or more images, a respective depth value. As another example, the image processing task can be motion estimation, where the network input includes multiple images, and the image processing output defines, for each pixel of one of the input images, a motion of the scene depicted at the pixel between the images in the network input.

[0264] In some implementations, input(s) 2 can be or otherwise represent natural language data. Machine-learned model(s) 1 can process the natural language data to generate an output. As an example, machine-learned model(s) 1 can process the natural language data to generate a language encoding output. As another example, machine-learned model(s) 1 canprocess the natural language data to generate a latent text embedding output. As another example, machine-learned model(s) 1 can process the natural language data to generate a translation output. As another example, machine-learned model(s) 1 can process the natural language data to generate a classification output. As another example, machine-learned model(s) 1 can process the natural language data to generate a textual segmentation output. As another example, machine-learned model(s) 1 can process the natural language data to generate a semantic intent output. As another example, machine-learned model(s) 1 can process the natural language data to generate an upscaled text or natural language output (e.g., text or natural language data that is higher quality than the input text or natural language, etc.). As another example, machine-learned model(s) 1 can process the natural language data to generate a prediction output (e.g., one or more predicted next portions of natural language content).

[0265] In some implementations, input(s) 2 can be or otherwise represent speech data (e.g., data describing spoken natural language, such as audio data, textual data, etc.).Machine-learned model(s) 1 can process the speech data to generate an output. As an example, machine-learned model(s) 1 can process the speech data to generate a speech recognition output. As another example, machine-learned model(s) 1 can process the speech data to generate a speech translation output. As another example, machine-learned model(s) 1 can process the speech data to generate a latent embedding output. As another example, machine-learned model(s) 1 can process the speech data to generate an encoded speech output (e g., an encoded and / or compressed representation of the speech data, etc.). As another example, machine-learned model(s) 1 can process the speech data to generate an upscaled speech output (e.g.. speech data that is higher quality' than the input speech data, etc.). As another example, machine-learned model(s) 1 can process the speech data to generate a textual representation output (e.g., a textual representation of the input speech data, etc.). As another example, machine-learned model(s) 1 can process the speech data to generate a prediction output.

[0266] In some implementations, input(s) 2 can be or otherwise represent latent encoding data (e.g., a latent space representation of an input, etc.). Machine-learned model(s) 1 can process the latent encoding data to generate an output. As an example, machine-learned model(s) 1 can process the latent encoding data to generate a recognition output. As another example, machine-learned model(s) 1 can process the latent encoding data to generate a reconstruction output. As another example, machine-learned model(s) 1 can process the latent encoding data to generate a search output. As another example, machine-learned model(s) 1can process the latent encoding data to generate a reclustering output. As another example, machine-learned model(s) 1 can process the latent encoding data to generate a prediction output.

[0267] In some implementations, input(s) 2 can be or otherwise represent statistical data. Statistical data can be, represent, or otherwise include data computed and / or calculated from some other data source. Machine-learned model(s) 1 can process the statistical data to generate an output. As an example, machine-learned model(s) 1 can process the statistical data to generate a recognition output. As another example, machine-learned model(s) 1 can process the statistical data to generate a prediction output. As another example, machine-learned model(s) 1 can process the statistical data to generate a classification output. As another example, machine-learned model(s) 1 can process the statistical data to generate a segmentation output. As another example, machine-learned model(s) 1 can process the statistical data to generate a visualization output. As another example, machine-learned model(s) 1 can process the statistical data to generate a diagnostic output.

[0268] In some implementations, input(s) 2 can be or otherwise represent sensor data. Machine-learned model(s) 1 can process the sensor data to generate an output. As an example, machine-learned model(s) 1 can process the sensor data to generate a recognition output. As another example, machine-learned model(s) 1 can process the sensor data to generate a prediction output. As another example, machine-learned model(s) 1 can process the sensor data to generate a classification output. As another example, machine-learned model(s) 1 can process the sensor data to generate a segmentation output. As another example, machine-learned model(s) 1 can process the sensor data to generate a visualization output. As another example, machine-learned model(s) 1 can process the sensor data to generate a diagnostic output. As another example, machine-learned model(s) 1 can process the sensor data to generate a detection output.

[0269] In some implementations, machine-learned model(s) 1 can be configured to perform a task that includes encoding input data for reliable and / or efficient transmission or storage (and / or corresponding decoding). For example, the task can be an audio compression task. The input can include audio data and the output can be or can include compressed audio data. In another example, the input includes visual data (e.g. one or more images or videos), the output includes compressed visual data, and the task is a visual data compression task. In another example, the task can include generating an embedding for input data (e.g. input audio or visual data). In some cases, the input includes audio data representing a spoken utterance and the task is a speech recognition task. The output can be or can include a textoutput which is mapped to the spoken utterance. In some cases, the task includes encrypting or decrypting input data. In some cases, the task includes a microprocessor performance task, such as branch prediction or memory address translation.

[0270] In some implementations, the task is a generative task, and machine-learned model(s) 1 can be configured to output content generated in view of input(s) 2. For instance, input(s) 2 can be or otherwise represent data of one or more modalities that encodes context for generating additional content.

[0271] In some implementations, the task can be a text completion task. Machine-learned model(s) 1 can be configured to process input(s) 2 that represent textual data and to generate output(s) 3 that represent additional textual data that completes a textual sequence that includes input(s) 2. For instance, machine-learned model(s) 1 can be configured to generate output(s) 3 to complete a sentence, paragraph, or portion of text that follows from a portion of text represented by input(s) 2.

[0272] In some implementations, the task can be an instruction following task.Machine-learned model(s) 1 can be configured to process input(s) 2 that represent instructions to perform a function and to generate output(s) 3 that advance a goal of satisfying the instruction function (e.g., at least a step of a multi-step procedure to perform the function). Output(s) 3 can represent data of the same or of a different modality as input(s) 2. For instance, input(s) 2 can represent textual data (e.g., natural language instructions for a task to be performed) and machine-learned model(s) 1 can process input(s) 2 to generate output(s) 3 that represent textual data responsive to the instructions (e.g., natural language responses, programming language responses, machine language responses, etc.). Input(s) 2 can represent image data (e.g., image-based instructions for a task to be performed, optionally accompanied by textual instructions) and machine-learned model(s) 1 can process input(s) 2 to generate output(s) 3 that represent textual data responsive to the instructions (e.g., natural language responses, programming language responses, machine language responses, etc.). One or more output(s) 3 can be iteratively or recursively generated to sequentially process and accomplish steps toward accomplishing the requested functionality. For instance, an initial output can be executed by an external system or be processed by machine-learned model(s) 1 to complete an initial step of performing a function. Multiple steps can be performed, with a final output being obtained that is responsive to the initial instructions.

[0273] In some implementations, the task can be a question answering task. Machine-learned model(s) 1 can be configured to process input(s) 2 that represent a question to answer and to generate output(s) 3 that advance a goal of returning an answer to the question (e.g., atleast a step of a multi-step procedure to perform the function). Output(s) 3 can represent data of the same or of a different modality as input(s) 2. For instance, input(s) 2 can represent textual data (e.g., natural language instructions for a task to be performed) and machine-learned model(s) 1 can process input(s) 2 to generate output(s) 3 that represent textual data responsive to the question (e.g., natural language responses, programming language responses, machine language responses, etc.). Input(s) 2 can represent image data (e.g., image-based instructions for a task to be performed, optionally accompanied by textual instructions) and machine-learned model(s) 1 can process input(s) 2 to generate output(s) 3 that represent textual data responsive to the question (e.g., natural language responses, programming language responses, machine language responses, etc.). One or more output(s) 3 can be iteratively or recursively generated to sequentially process and accomplish steps toward answering the question. For instance, an initial output can be executed by an external system or be processed by machine-learned model(s) 1 to complete an initial step of obtaining an answer to the question (e.g., querying a database, performing a computation, executing a script, etc.). Multiple steps can be performed, with a final output being obtained that is responsive to the question.

[0274] In some implementations, the task can be an image generation task. Machine-learned model(s) 1 can be configured to process input(s) 2 that represent context regarding a desired portion of image content. The context can include text data, image data, audio data, etc. Machine-learned model(s) 1 can be configured to generate output(s) 3 that represent image data that depicts imagery related to the context. For instance, machine-learned model(s) 1 can be configured to generate pixel data of an image. Values for channel (s) associated with the pixels in the pixel data can be selected based on the context (e.g., based on a probability determined based on the context).

[0275] In some implementations, the task can be an audio generation task. Machine-learned model(s) 1 can be configured to process input(s) 2 that represent context regarding a desired portion of audio content. The context can include text data, image data, audio data, etc. Machine-learned model(s) 1 can be configured to generate output(s) 3 that represent audio data related to the context. For instance, machine-learned model(s) 1 can be configured to generate waveform data in the form of an image (e.g., a spectrogram). Values for channel(s) associated with pixels of the image can be selected based on the context. Machine-learned model(s) 1 can be configured to generate waveform data in the form of a sequence of discrete samples of a continuous waveform. Values of the sequence can be selected based on the context (e.g., based on a probability determined based on the context).

[0276] In some implementations, the task can be a data generation task. Machine-learned model(s) 1 can be configured to process input(s) 2 that represent context regarding a desired portion of data (e.g., data from various data domains, such as sensor data, image data, multimodal data, statistical data, etc.). The desired data can be, for instance, synthetic data for training other machine-learned models. The context can include arbitrary data type(s).Machine-learned model(s) 1 can be configured to generate output(s) 3 that represent data that aligns with the desired data. For instance, machine-learned model(s) 1 can be configured to generate data values for populating a dataset. Values for the data object(s) can be selected based on the context (e.g., based on a probability determined based on the context).Additional Disclosure

[0277] The technology discussed herein makes reference to servers, databases, software applications, and other computer-based systems, as well as actions taken and information sent to and from such systems. The inherent flexibility of computer-based systems allows for a great variety of possible configurations, combinations, and divisions of tasks and functionality between and among components. For instance, processes discussed herein can be implemented using a single device or component or multiple devices or components working in combination. Databases and applications can be implemented on a single system or distributed across multiple systems. Distributed components can operate sequentially or in parallel.

[0278] While the present subject matter has been described in detail with respect to various specific example embodiments thereof, each example is provided by way of explanation, not limitation of the disclosure. Those skilled in the art, upon attaining an understanding of the foregoing, can readily produce alterations to, variations of, and equivalents to such embodiments. Accordingly, the subject disclosure does not preclude inclusion of such modifications, variations or additions to the present subject matter as would be readily apparent to one of ordinary skill in the art. For instance, features illustrated or described as part of one embodiment can be used with another embodiment to yield a still further embodiment. Thus, it is intended that the present disclosure cover such alterations, variations, and equivalents.

[0279] Aspects of the disclosure have been described in terms of illustrative embodiments thereof. Any and all features in the following claims can be combined or rearranged in any way possible, including combinations of claims not explicitly enumerated in combination together, as the example claim dependencies listed herein should not be readas limiting the scope of possible combinations of features disclosed herein. Accordingly, the scope of the present disclosure is by way of example rather than by way of limitation, and the subject disclosure does not preclude inclusion of such modifications, variations or additions to the present subject matter as would be readily apparent to one of ordinary skill in the art. Moreover, terms are described herein using lists of example elements joined by conjunctions such as “and.” “or,” “but,” etc. It should be understood that such conjunctions are provided for explanatory purposes only. Clauses and other sequences of items joined by a particular conjunction such as “or,” for example, can refer to “and / or,” “at least one of’, “any combination of’ example elements listed therein, etc. Terms such as “based on” should be understood as “based at least in part on.”

[0280] The term “can” should be understood as referring to a possibility of a feature in various implementations and not as prescribing an ability7that is necessarily present in every7implementation. For example, the phrase “X can perform Y” should be understood as indicating that, in various implementations, X has the potential to be configured to perform Y, and not as indicating that in every’ instance X must always be able to perform Y. It should be understood that, in vanous implementations, X might be unable to perform Y and remain within the scope of the present disclosure.

[0281] The term “may” should be understood as referring to a possibility' of a feature in various implementations and not as prescribing an ability that is necessarily present in every implementation. For example, the phrase “X may perform Y” should be understood as indicating that, in various implementations, X has the potential to be configured to perform Y, and not as indicating that in every' instance X must always be able to perform Y. It should be understood that, in various implementations, X might be unable to perform Y and remain within the scope of the present disclosure.

Claims

WHAT IS CLAIMED IS:

1. A computer-implemented method, comprising:obtaining sequential source data comprising a plurality of data elements; generating a first representation of one or more first data elements of the plurality of data elements by a first element encoder;generating a second representation of a second data element of the plurality of data elements by a second element encoder;generating a latent projection based on the second representation;generating a predicted second representation based on the first representation and the latent projection by a representation predictor model;determining a loss between the second representation and the predicted second representation; andtraining at least the first element encoder based on the loss.

2. The computer-implemented method of any claim (e.g., claim 1), wherein the sequential source data comprises video data, and wherein the plurality of data elements comprise a plurality of video patches, each of the video patches comprising at least a portion of one or more frames of the video data.

3. The computer-implemented method of any claim (e.g., claim 1), wherein the first representation of the one or more first data elements comprises a first embedding of the one or more first data elements, and wherein the second representation of the second data element comprises a second embedding of the second data element.

4. The computer-implemented method of any claim (e.g., claim 1), wherein the one or more first data elements is ordered earlier in the sequential source data than the second data element.

5. The computer-implemented method of any claim (e.g., claim 1), wherein the first element encoder and the second element encoder are initialized from a common encoder model.

6. The computer-implemented method of any claim (e.g., claim 1), wherein the second element encoder is frozen during the training of the first element encoder.

7. The computer-implemented method of any claim (e.g., claim 1), wherein at least one of the first element encoder or the second element encoder comprises a masked autoencoder.

8. The computer-implemented method of any claim (e.g., claim 1), wherein at least one of the first element encoder or the second element encoder comprises smoothed parameters of another of the first element encoder or the second element encoder.

9. The computer-implemented method of any claim (e.g., claim 1), wherein generating a latent projection based on the second representation comprises generating, using a projection generation model, the latent projection based on the second representation.

10. The computer-implemented method of any claim (e.g., claim 9), wherein the projection generation model comprises a neural network.

11. The computer-implemented method of any claim (e.g., claim 1), wherein the latent projection comprises alatent vector.

12. The computer-implemented method of any claim (e.g., claim 1), wherein the latent projection comprises alatent description of the second element.

13. The computer-implemented method of any claim (e.g., claim 12), wherein the latent description comprises text data.

14. The computer-implemented method of any claim (e.g., claim 1), wherein the representation predictor model comprises a transformer model.

15. The computer-implemented method of any claim (e.g., claim 14), wherein the transformer model comprises a vision transformer (ViT) model.

16. The computer-implemented method of any claim (e.g., claim 1), further comprising:obtaining an inference data element; andsubsequent to training at least the first element encoder based on the loss, generating a first inference representation of the inference data element using the first element encoder.

17. The computer-implemented method of any claim (e.g., claim 16), further comprising:generating a second inference representation using the representation predictor model based on the first inference representation; andgenerating a second inference data element based on the second inference representation.

18. The computer-implemented method of any claim (e.g., claim 17), wherein generating the second inference representation is based on a guidance input.

19. The computer-implemented method of any claim (e.g., claim 18), wherein the guidance input comprises one or more of: noise data or a description of the second inference data element.

20. A computer-implemented method, comprising:obtaining video data comprising a plurality of video data elements, each video data element respectively comprising at least a portion of one or more video frames;generating a first representation of a first video data element of the plurality of video data elements by a first element encoder;generating a second representation of a second video data element of the plurality of data elements by a second element encoder, the second data element comprising data ordered later in the video data than the one or more first data elements;generating a latent projection based on the second representation, the latent projection being descriptive of the second video data element;generating a predicted second representation based on the first representation and the latent projection by a representation predictor model, the predicted second representation approximating the second representation;determining a loss between the second representation and the predicted second representation; andtraining at least the first element encoder based on the loss.

21. A computing system, comprising:one or more processors; andone or more non-transitory, computer-readable media storing instructions that, when implemented, cause the one or more processors to perform operations, the operations comprising:obtaining sequential source data comprising a plurality of data elements; generating a first representation of one or more first data elements of the plurality of data elements by a first element encoder;generating a second representation of a second data element of the plurality of data elements by a second element encoder;generating a latent projection based on the second representation; generating a predicted second representation based on the first representation and the latent projection by a representation predictor model;determining a loss between the second representation and the predicted second representation; andtraining at least the first element encoder based on the loss.

22. The computing system of any claim (e.g., claim 21), wherein the sequential source data comprises video data, and wherein the plurality of data elements comprise a plurality of video patches, each of the video patches comprising at least a portion of one or more frames of the video data.

23. The computing system of any claim (e.g., claim 21), wherein the first representation of the one or more first data elements comprises a first embedding of the one or more firstdata elements, and wherein the second representation of the second data element comprises a second embedding of the second data element.

24. The computing system of any claim (e.g., claim 21), wherein the one or more first data elements is ordered earlier in the sequential source data than the second data element.

25. The computing system of any claim (e.g., claim 21), wherein the first element encoder and the second element encoder are initialized from a common encoder model.

26. The computing system of any claim (e.g., claim 21), wherein the second element encoder is frozen during the training of the first element encoder.

27. The computing system of any claim (e.g., claim 21), wherein at least one of the first element encoder or the second element encoder comprises a masked autoencoder.

28. The computing system of any claim (e.g., claim 21), wherein at least one of the first element encoder or the second element encoder comprises smoothed parameters of another of the first element encoder or the second element encoder.

29. The computing system of any claim (e.g., claim 21), wherein generating a latent projection based on the second representation comprises generating, using a projection generation model, the latent projection based on the second representation.

30. The computing system of any claim (e.g., claim 29), wherein the projection generation model comprises a neural network.

31. The computing system of any claim (e.g., claim 21), wherein the latent projection comprises a latent vector.

32. The computing system of any claim (e.g., claim 21), wherein the latent projection comprises a latent description of the second element.

33. The computing system of any claim (e.g., claim 32), wherein the latent description comprises text data.

34. The computing system of any claim (e.g., claim 21), wherein the representation predictor model comprises a transformer model.

35. The computing system of any claim (e.g., claim 34), wherein the transformer model comprises a vision transformer (ViT) model.

36. The computing system of any claim (e.g., claim 21), wherein the operations further comprise:obtaining an inference data element; andsubsequent to training at least the first element encoder based on the loss, generating a first inference representation of the inference data element using the first element encoder.

37. The computing system of any claim (e.g., claim 36), wherein the operations further comprise:generating a second inference representation using the representation predictor model based on the first inference representation; andgenerating a second inference data element based on the second inference representation.

38. The computing system of any claim (e.g., claim 37), wherein generating the second inference representation is based on a guidance input.

39. The computing system of any claim (e.g., claim 38), wherein the guidance input comprises one or more of: noise data or a description of the second inference data element.

40. A computing system, comprising:one or more processors; andone or more non-transitory, computer-readable media storing instructions that, when implemented, cause the one or more processors to perform operations, the operations comprising:obtaining video data comprising a plurality of video data elements, each video data element respectively comprising at least a portion of one or more video frames;generating a first representation of a first video data element of the plurality of video data elements by a first element encoder;generating a second representation of a second video data element of the plurality of data elements by a second element encoder, the second data element comprising data ordered later in the video data than the one or more first data elements;generating a latent projection based on the second representation, the latent projection being descriptive of the second video data element;generating a predicted second representation based on the first representation and the latent projection by a representation predictor model, the predicted second representation approximating the second representation;determining a loss between the second representation and the predicted second representation; andtraining at least the first element encoder based on the loss.

41. One or more non-transitory. computer-readable media storing instructions that, when implemented, cause one or more processors to perform operations, the operations comprising:obtaining sequential source data comprising a plurality of data elements; generating a first representation of one or more first data elements of the plurality of data elements by a first element encoder;generating a second representation of a second data element of the plurality of data elements by a second element encoder;generating a latent projection based on the second representation; generating a predicted second representation based on the first representation and the latent projection by a representation predictor model;determining a loss between the second representation and the predicted second representation; andtraining at least the first element encoder based on the loss.

42. The one or more non-transitory, computer-readable media of any claim (e.g., claim 41), wherein the sequential source data comprises video data, and wherein the plurality of data elements comprise a plurality of video patches, each of the video patches comprising at least a portion of one or more frames of the video data.

43. The one or more non-transitory, computer-readable media of any claim (e.g., claim 41), wherein the first representation of the one or more first data elements comprises a first embedding of the one or more first data elements, and wherein the second representation of the second data element comprises a second embedding of the second data element.

44. The one or more non-transitory, computer-readable media of any claim (e.g., claim 41), wherein the one or more first data elements is ordered earlier in the sequential source data than the second data element.

45. The one or more non-transitory, computer-readable media of any claim (e.g., claim 41), wherein the first element encoder and the second element encoder are initialized from a common encoder model.

46. The one or more non-transitory, computer-readable media of any claim (e.g., claim 41), wherein the second element encoder is frozen during the training of the first element encoder.

47. The one or more non-transitory, computer-readable media of any claim (e.g., claim 41), wherein at least one of the first element encoder or the second element encoder comprises a masked autoencoder.

48. The one or more non-transitory, computer-readable media of any claim (e.g., claim 41), wherein at least one of the first element encoder or the second element encodercomprises smoothed parameters of another of the first element encoder or the second element encoder.

49. The one or more non- transitory, computer-readable media of any claim (e.g., claim 41), wherein generating a latent projection based on the second representation comprises generating, using a projection generation model, the latent projection based on the second representation.

50. The one or more non-transitory. computer-readable media of any claim (e.g., claim 49), wherein the projection generation model comprises a neural network.

51. The one or more non-transitory, computer-readable media of any claim (e.g., claim 41), wherein the latent projection comprises a latent vector.

52. The one or more non-transitory, computer-readable media of any claim (e.g., claim 41), wherein the latent projection comprises a latent description of the second element.

53. The one or more non-transitory, computer-readable media of any claim (e.g., claim 52), wherein the latent description comprises text data.

54. The one or more non-transitory, computer-readable media of any claim (e.g., claim 41), wherein the representation predictor model comprises a transformer model.

55. The one or more non-transitory, computer-readable media of any claim (e.g., claim 54), wherein the transformer model comprises a vision transformer (ViT) model.

56. The one or more non-transitory, computer-readable media of any claim (e.g., claim 41), wherein the operations further comprise:obtaining an inference data element; andsubsequent to training at least the first element encoder based on the loss, generating a first inference representation of the inference data element using the first element encoder.

57. The one or more non-transitory, computer-readable media of any claim (e.g., claim 56), wherein the operations further comprise:generating a second inference representation using the representation predictor model based on the first inference representation; andgenerating a second inference data element based on the second inference representation.

58. The one or more non-transitory, computer-readable media of any claim (e.g., claim 57), wherein generating the second inference representation is based on a guidance input.

59. The one or more non-transitory, computer-readable media of any claim (e.g., claim 58), wherein the guidance input comprises one or more of: noise data or a description of the second inference data element.

60. One or more non-transitory, computer-readable media storing instructions that, when implemented, cause one or more processors to perform operations, the operations comprising:obtaining video data comprising a plurality of video data elements, each video data element respectively comprising at least a portion of one or more video frames;generating a first representation of a first video data element of the plurality of video data elements by a first element encoder;generating a second representation of a second video data element of the plurality' of data elements by a second element encoder, the second data element comprising data ordered later in the video data than the one or more first data elements;generating a latent projection based on the second representation, the latent projection being descriptive of the second video data element;generating a predicted second representation based on the first representation and the latent projection by a representation predictor model, the predicted second representation approximating the second representation;determining a loss between the second representation and the predicted second representation; andtraining at least the first element encoder based on the loss.