Conditional object-centric slot attention learning for videos and other sequential data

The slot attention model tracks and updates entity attributes in time series by using a slot vector model and a predictor model, which solves the problems of wasted computational resources and insufficient prediction accuracy of existing models, and achieves more efficient video and sequential data processing.

CN115018067BActive Publication Date: 2026-01-02GOOGLE LLC
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202210585538.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2022-04-21
Filing Date
2022-05-26
Publication Date
2026-01-02
Estimated Expiration
2042-05-26

AI Technical Summary

Technical Problem

Existing machine learning models struggle to effectively utilize temporal correlations when processing video and other sequential data, leading to wasted computational resources and insufficient prediction accuracy.

Method used

By employing a slot attention model, and through a slot vector model and a predictor model, the slot vector is used to track and update the attribute changes of entities in the time series, thereby reducing the number of computation iterations and improving prediction accuracy.

Benefits of technology

By reducing computational resource usage and improving prediction accuracy, slot attention models can process video and other sequential data more efficiently and are suitable for a variety of tasks, including image reconstruction, text translation, and object attribute detection.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115018067B_ABST
    Figure CN115018067B_ABST
Patent Text Reader

Abstract

This application relates to methods, systems, and non-transitory computer-readable media associated with conditional object-centric slot attention learning for video and other sequential data. One method includes obtaining a first feature vector and a second feature vector respectively representing content of first and second image frames of an input video. The method can also include generating, based on the first feature vector, first slot vectors, where each slot vector represents an attribute of a corresponding entity as represented in the first image frame, and generating, based on the first slot vectors, predicted slot vectors including corresponding predicted slot vectors representing a transition of the attribute of the corresponding entity from the first image frame to the second image frame. The method can also include generating, based on the predicted slot vectors and the second feature vector, second slot vectors including corresponding slot vectors representing the attribute of the corresponding entity as represented in the second image frame, and determining an output based on the predicted slot vectors or the second slot vectors.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] Cross-references to related applications

[0002] This application claims priority to U.S. Provisional Patent Application No. 63 / 194,425, filed May 28, 2021, which is incorporated herein by reference as if fully set forth in this specification. Background Technology

[0003] Machine learning models can be used to process various types of data, including images, videos, time series, text, and / or point clouds. Improvements in machine learning models allow them to perform data processing faster and / or with less computational resources. Summary of the Invention

[0004] In a first example embodiment, a method may include obtaining a first plurality of feature vectors representing the content of a first input frame of an input data sequence and a second plurality of feature vectors representing the content of a second input frame of the input data sequence. The second input frame may follow the first input frame. The method may further include generating a first plurality of slot vectors based on the first plurality of feature vectors and through a slot vector model. Each corresponding slot vector in the first plurality of slot vectors may represent an attribute of a corresponding entity as represented in the first input frame. The method may further include generating a plurality of predicted slot vectors based on the first plurality of slot vectors and through a predictor model, the plurality of predicted slot vectors including a corresponding predicted slot vector for each corresponding slot vector in the first plurality of slot vectors. The corresponding predicted slot vector may represent a transition of the attribute of the corresponding entity from the first input frame to the second input frame. The method may further include generating a second plurality of slot vectors based on the plurality of predicted slot vectors and the second plurality of feature vectors and through a slot vector model, the second plurality of slot vectors including a corresponding slot vector for each corresponding slot vector in the first plurality of slot vectors. The corresponding slot vector may represent an attribute of the corresponding entity as represented in the second input frame. The method may further include determining the output based on one or more of the following: (i) at least one of the plurality of predicted slot vectors or (ii) at least one of the second plurality of slot vectors.

[0005] In a second example embodiment, a system can include a processor and a non-transitory computer-readable medium having stored thereon instructions that, when executed by the processor, cause the processor to perform operations. The operations can include obtaining a first plurality of feature vectors representing content of a first input frame of an input data sequence and a second plurality of feature vectors representing content of a second input frame of the input data sequence. The second input frame can be subsequent to the first input frame. The operations can also include generating, based on the first plurality of feature vectors and by a slot vector model, a first plurality of slot vectors. Each respective slot vector of the first plurality of slot vectors can represent an attribute of a corresponding entity as represented in the first input frame. The operations can also include generating, based on the first plurality of slot vectors and by a predictor model, a plurality of predicted slot vectors, the plurality of predicted slot vectors including, for each respective slot vector of the first plurality of slot vectors, a corresponding predicted slot vector. The corresponding predicted slot vector can represent a transition of the attribute of the corresponding entity from the first input frame to the second input frame. The method can also include generating, based on the plurality of predicted slot vectors and the second plurality of feature vectors and by the slot vector model, a second plurality of slot vectors, the second plurality of slot vectors including, for each respective slot vector of the first plurality of slot vectors, a corresponding slot vector. The corresponding slot vector can represent the attribute of the corresponding entity as represented in the second input frame. The operations can further include determining an output based on one or more of (i) at least one predicted slot vector of the plurality of predicted slot vectors or (ii) at least one slot vector of the second plurality of slot vectors.

[0006] In a third example embodiment, a non-transitory computer-readable medium can have stored thereon instructions that, when executed by a computing device, cause the computing device to perform operations. The operations can include obtaining a first plurality of feature vectors representing content of a first input frame of an input data sequence and a second plurality of feature vectors representing content of a second input frame of the input data sequence. The second input frame can be subsequent to the first input frame. The operations can also include generating, based on the first plurality of feature vectors and by a slot vector model, a first plurality of slot vectors. Each respective slot vector of the first plurality of slot vectors can represent an attribute of a corresponding entity as represented in the first input frame. The operations can also include generating, based on the first plurality of slot vectors and by a predictor model, a plurality of predicted slot vectors that includes, for each respective slot vector of the first plurality of slot vectors, a corresponding predicted slot vector. The corresponding predicted slot vector can represent a transition of the attribute of the corresponding entity from the first input frame to the second input frame. The operations can also include generating, based on the plurality of predicted slot vectors and the second plurality of feature vectors and by the slot vector model, a second plurality of slot vectors that includes, for each respective slot vector of the first plurality of slot vectors, a corresponding slot vector. The corresponding slot vector can represent the attribute of the corresponding entity as represented in the second input frame. The operations can also further include determining an output based on one or more of (i) at least one predicted slot vector of the plurality of predicted slot vectors or (ii) at least one slot vector of the second plurality of slot vectors.

[0007] In a fourth example embodiment, a system can include means for obtaining (i) a first plurality of feature vectors representing content of a first input frame of an input data sequence and (ii) a second plurality of feature vectors representing content of a second input frame of the input data sequence. The second input frame can be subsequent to the first input frame. The system can also include means for generating, based on the first plurality of feature vectors, a first plurality of slot vectors. Each respective slot vector of the first plurality of slot vectors can represent an attribute of a corresponding entity as represented in the first input frame. The system can also include means for generating, based on the first plurality of slot vectors, a plurality of predicted slot vectors that includes, for each respective slot vector of the first plurality of slot vectors, a corresponding predicted slot vector. The corresponding predicted slot vector can represent a transition of the attribute of the corresponding entity from the first input frame to the second input frame. The system can also include means for generating, based on the plurality of predicted slot vectors and the second plurality of feature vectors, a second plurality of slot vectors that includes, for each respective slot vector of the first plurality of slot vectors, a corresponding slot vector. The corresponding slot vector can represent the attribute of the corresponding entity as represented in the second input frame. The system can also further include means for determining an output based on one or more of (i) at least one predicted slot vector of the plurality of predicted slot vectors or (ii) at least one slot vector of the second plurality of slot vectors.

[0008] These and other embodiments, aspects, advantages, and alternatives will become apparent to those of ordinary skill in the art upon reading the following detailed description, with appropriate reference to the drawings. Further, the included description and drawings are intended to illustrate embodiments and are not intended to be limiting. For example, structural elements and process steps have been shown, described, and claimed in various combinations, it will be understood that the various combinations are not mutually exclusive and can be combined with each other or with other structural elements and process steps to create other embodiments. BRIEF DESCRIPTION OF DRAWINGS

[0009] Figure 1 A computing system is shown in accordance with examples described herein.

[0010] Figure 2 A computing device is shown in accordance with examples described herein.

[0011] Figure 3 A slot attention model is shown in accordance with examples described herein.

[0012] Figure 4 A slot vector is shown in accordance with examples described herein.

[0013] Figure 5 A sequence of slot vectors is shown in accordance with examples described herein.

[0014] Figure 6 A slot vector is shown in accordance with examples described herein.

[0015] Figure 7 A flowchart is shown in accordance with examples described herein.

[0016] Figure 8 A flowchart is shown in accordance with examples described herein. DETAILED DESCRIPTION

[0017] Example methods, devices, and systems are described herein. It should be understood that the words “example” and “exemplary” are used herein to mean “serving as an example, instance, or illustration.” Any implementation or feature described herein as “example,” “exemplary,” and / or “illustrative” is not necessarily to be construed as preferred or advantageous over other implementations or features. Thus, other implementations and

[0018] Accordingly, the example embodiments described herein are not meant to be limiting. It will be readily understood to those skilled in the art that the aspects of the present disclosure as generally described herein, and illustrated in the drawings, can be arranged, substituted, combined, separated, and designed in various different configurations.

[0019] Further, the features illustrated in the drawings, which comprise one or more embodiments, can be used alone or in any combination(s). The drawings are intended to be generally schematic and are not exact depictions of actual products or systems. Therefore, specific dimensions, shapes, and other parameters of the components shown in the drawings are not intended to be limiting and are provided for clarity of discussion only. The same or similar reference numbers in different drawings can represent the same or similar components.

[0020] Also, any listing of elements, blocks, or steps in a method or claim is for clarity only. Thus, such listing should not be interpreted as requiring or implying that those elements, blocks, or steps be performed in the particular order or sequence in which they are recited. The drawings are not drawn to scale.

[0021] I. SUMMARY

[0022] A slot attention model can be configured to determine, based on a distributed representation of a perceptual representation, an entity-centric (e.g., object-centric) representation of an entity contained in the perceptual representation. An example perceptual representation can take the form of an image containing one or more entities such as objects, surfaces, regions, backgrounds, or other environmental features. A machine learning model can be configured to produce a distributed representation of the image. For example, one or more convolutional neural networks can be configured to process the image and produce one or more convolutional feature maps, which can represent the output of various feature filters implemented by the one or more convolutional neural networks.

[0023] These convolutional feature maps can be considered a distributed representation of the entities in the image, as the features represented by the feature maps are related to different parts of the image region but are not directly / explicitly associated with any of the entities represented in the image data. On the other hand, an object-centric representation can associate one or more features with individual entities represented in the image data. Thus, for example, each feature in the distributed representation can be associated with a corresponding portion of the perceptual representation, while each feature in the entity-centric representation can be associated with a corresponding entity contained in the perceptual representation.

[0024] Accordingly, a slot attention model can be configured to produce a plurality of entity-centric representations (referred to herein as slot vectors) based on a plurality of distributed representations (referred to herein as feature vectors). Each slot vector can be an entity-specific semantic embedding representing one or more properties or characteristics of a corresponding entity. Thus, a slot attention model can be considered an interface between a perceptual representation and a structured set of variables represented by the slot vectors.

[0025] The slot attention model can include a neural network memory unit, such as a Gated Recurrent Unit (GRU) or a Long Short-Term Memory (LSTM) neural network, configured to store and update the plurality of slot vectors during iterations in which the slot attention model processes the feature vector. Parameters of the neural network memory unit can be learned during training of the slot attention model. The plurality of slot vectors can be initialized with values prior to a first iteration in which the slot attention model is processed. In one example, the plurality of slot vectors can be initialized with random values, e.g., selected from a normal distribution. In other implementations, the plurality of slot vectors can be initialized with values that bind and / or attach one or more of these slot vectors to a particular entity (e.g., based on previously determined values of slot vectors). Thus, a particular slot vector can be caused to represent the same entity in successive perceptual representations (e.g., successive segments of a video frame or audio waveform) by initializing the particular slot vector with a value of a particular slot vector previously determined with respect to one or more previous perceptual representations.

[0026] The slot attention model can further include a learnable value, key, and query function. The slot attention model can be configured to compute an attention matrix based on (i) dot products of the plurality of feature vectors transformed by the key function and (ii) the plurality of slot vectors transformed by the query function. Entries of the attention matrix can be normalized by a softmax function along a dimension corresponding to a slot vector, such that the slot vectors compete with each other to represent entities contained in the perceptual representation. For example, when the number of slot vectors corresponds to the number of columns of the attention matrix, each value within each respective row of the attention matrix can be normalized by a softmax function with respect to values within the respective row. On the other hand, when the number of slot vectors corresponds to the number of rows of the attention matrix, each value within each respective column of the attention matrix can be normalized by a softmax function with respect to values within the respective column.

[0027] The slot attention model can further be configured to compute an update matrix based on (i) the attention matrix and (ii) the plurality of feature vectors transformed by the value function. The neural network memory unit can be configured to update the plurality of slot vectors based on the update matrix and previous values of the slot vectors, allowing the values of the slot vectors to be refined in one or more iterations of the slot attention model to improve their representation accuracy.

[0028] The slot vectors produced by the slot attention model can be permutation invariant with respect to the feature vectors and permutation equivariant with respect to each other. Thus, for a given initialization of the slot vectors, the order of the feature vectors of a given perceptual representation can not affect the values of the slot vectors and / or the order of the values of the slot vectors. For a given perceptual representation, different initializations of the slot vectors can change the order of the slot vectors while the set of values of the slot vectors remains approximately constant. Thus, permuting the order of the slot vectors after their initialization can be equivalent to permuting the output of the slot attention model. The permutation equivariance of the slot vectors with respect to each other allows the slot vectors to be completely interchangeable and allows each slot vector to represent a variety of entities independent of the type, class, and / or semantics of the entity.

[0029] The plurality of slot vectors can be used by one or more machine learning models to perform a particular task, such as image reconstruction, text translation, object attribute / property detection, reward prediction, visual reasoning, question answering, control, and / or planning, among other possible tasks. Thus, the slot attention model can be trained jointly with the one or more machine learning models to produce slot vectors that are useful for performing the particular task of the one or more machine learning models. That is, the slot attention model can be trained to produce slot vectors in a task-specific manner such that the slot vectors represent information that is important for the particular task and omit information that is not important and / or relevant for the particular task.

[0030] Although the slot attention model can be trained for a particular task, the architecture of the slot attention model is not task-specific, thus allowing the slot attention model to be used for a variety of tasks. The slot attention model can be used for both supervised and unsupervised training tasks. Additionally, the slot attention model does not assume, expect, or rely on feature vectors representing a particular type of data (e.g., image data, waveform data, text data, etc.). Thus, the slot attention model can be used with any type of data that can be represented by one or more feature vectors, and the type of data can be based on the task for which the slot attention model is used.

[0031] Further, the slot vectors themselves can not be specialized for a particular entity type and / or class. Thus, when multiple entity classes are contained within a perceptual representation, each slot vector can be able to represent each entity regardless of its class. Each of the slot vectors can be bound or attached to a particular entity to represent its features, but this binding / attention does not depend on the entity type, class, and / or semantics. Having the slot vectors bound / attend to an entity can be driven by a downstream task that uses the slot vectors - the slot attention model can not“know” the object itself and can not distinguish, for example, between a cluster object, a color, and / or a spatial region.

[0032] The entity-centric representations (e.g., slot vectors) of features present within an input data sequence can be generated, tracked, and / or updated based on multiple input frames of the input data sequence. For example, slot vectors representing objects present in a video can be generated, tracked, and / or updated across different image frames of the video. In particular, rather than processing each input frame independently, the input frames can be processed as a sequence, with previous slot vectors providing information that can be used to generate subsequent slot vectors. Thus, the slot vectors generated for an input data sequence can be temporally coherent, with a given slot representing the same entity and / or feature across multiple input frames of the input data sequence.

[0033] Such temporally coherent slot vectors can be generated using a system that includes an encoder model, a slot vector model (alternatively referred to as a corrector), and a predictor model. The encoder model can be configured to generate feature vectors (e.g., distributed representations) representing the content of different input frames of an input data sequence. For example, the encoder model can be configured to generate, for each respective input frame of multiple input frames of an input data sequence, a corresponding set of feature vectors representing the content of the respective input frame. The slot vector model can be configured to process the feature vectors of a corresponding input frame and generate, based on the feature vectors, multiple slot vectors for the corresponding input frame. The predictor model can be configured to process the multiple slot vectors of a given input frame and generate, based on the multiple slot vectors, multiple predicted slot vectors for a subsequent frame.

[0034] In particular, the predictor model can be trained to model the dynamics and / or interactions of entities represented by the multiple slot vectors and, from this, predict the states of these entities in subsequent input frames (e.g., at future times). For example, in the context of a video, the predictor model can be configured to predict future states of objects in a scene (as represented by predicted slot vectors) based on current states of the objects in the scene (as represented by slot vectors). The slot vector model can also be configured to process the predicted slot vectors generated for a particular (e.g., future) input frame and the feature vectors corresponding to the particular input frame and, based on this, generate additional multiple slot vectors corresponding to the particular input frame. That is, the slot vector model can be configured to refine, correct, and / or update the information contained in the predicted slot vectors based on the information contained in the feature vectors, thereby allowing information about entity attributes to be aggregated and updated over time.

[0035] Thus, by using a predictor model in conjunction with a slot vector model, the slot vectors produced for a given input frame can borrow / utilize information from slot vectors of previous input frames, thereby enabling more accurate slot vectors to be produced using as few iterations as one iteration of the slot vector model. Additionally, the predictor model can be used to produce predicted slot vectors for multiple time steps into the future, thereby allowing the slot attention model to produce a set of slot vectors every n >= 2 input frames, rather than a set of slot vectors for every input frame, which can allow the system to reduce its use of computational resources. Further, the predictor model can allow the system to handle the absence of a previously existing entity in some input frames by using predicted slot vectors to indicate the absence of the entity (such as due to occlusion and / or movement of the previously seen entity).

[0036] Additionally, the system can include an initializer model configured to produce a plurality of initialization slot vectors configured to deterministically bind a given slot to a target entity. In particular, the slot vector model can be configured to process the initialization slot vectors and feature vectors of an initial input frame of the input data sequence to produce slot vectors corresponding to the initial input frame. Each respective slot vector of the initial frame can be bound to an entity indicated by a corresponding initialization slot vector. The initializer model can be configured to obtain, for at least some of the initialization slot vectors, an indication of a corresponding entity to be represented by a corresponding slot vector (alternatively referred to as a conditioned input), with any remaining initialization vectors being randomly initialized. For example, when a video represents three objects, a first object can be assigned to a first slot, a second object can be assigned to a second slot, and a third object can be assigned to a third slot, such that the slot vector of the first slot represents the first object, the slot vector of the second slot represents the second object, and the slot vector of the third slot represents the third object.

[0037] The indication of the corresponding entity can be, among other things, a bounding box at least partially outlining an entity within at least the initial input frame of the input data sequence, coordinates (e.g., centroid coordinates) associated with an entity within at least the initial input frame, and / or a segmentation mask outlining a shape of an entity within at least the initial input frame. In general, the indication of the corresponding entity can be a region of space defined by the initial input frame, where the region is associated with the corresponding entity. Depending on the format of the indication of the corresponding entity, the initializer model can be implemented as, for example, a multi-layer perceptron or a convolutional neural network.

[0038] Alternatively, in some implementations, the plurality of slot vectors can be learned during training and remain fixed during inference. For example, the same learned plurality of slot vectors can be used independent of the input data sequence.

[0039] The system can be trained to produce useful slot vectors and predicted slot vectors using a decoder configured to reconstruct an input frame, e.g., of an input data sequence, or a flow vector field associated with the input frame. For example, in the case of video, the system can be trained using an image reconstruction loss and / or an optical flow reconstruction loss. The reconstructed optical flow for a given image can be represented as, e.g., a color image, where the magnitude and direction of the optical flow at a given pixel is represented by the color of that pixel. Using an optical flow reconstruction loss can facilitate training and produce a model that is able to more accurately distinguish between different objects and / or features, as the motion indicated by the optical flow informs entity-centric representations since related features typically move together, while unrelated features do not. In particular, the optical flow reconstruction loss can adapt the system to accurately distinguish scenes and / or features containing complex textures.

[0040] The decoder used in training can then be replaced with a different task-specific decoder or other machine learning model. The task-specific decoder can be trained to interpret slot vectors produced by the system in the context of a particular task and produce a task-specific output, thereby allowing the system to be used in a variety of contexts and applications. In one example, the particular task can include controlling a robotic device, so the task-specific decoder can be trained to use information represented by the slot vectors to help control the robotic device. In another example, the particular task can include operating an autonomous vehicle, so the task-specific decoder can be trained to use information represented by the slot vectors to help operate the autonomous vehicle. Additionally, since the system operates on feature vectors, any input data sequence from which a feature vector can be produced can be processed by the system. Thus, in addition to video, the system can be applied to and / or used to produce point cloud sequences, waveforms (e.g., audio waveforms represented as spectrograms), text, RADAR data sequences, and / or other computer-generated and / or human-generated sequential data as output.

[0041] II. Example Computing Device

[0042] Figure 1 An example form factor of a computing system 100 is shown. The computing system 100 can be, e.g., a mobile phone, a tablet computer, or a wearable computing device. However, other embodiments are possible. The computing system 100 can include various elements, such as a body 102, a display 106, and buttons 108 and 110. The computing system 100 can also include a front-facing camera 104, a rear-facing camera 112, a front-facing infrared camera 114, and an infrared pattern projector 116.

[0043] The forward-facing camera 104 can be positioned on a side of the body 102 that is generally facing the user in operation (e.g., on the same side as the display 106). The rear-facing camera 112 can be positioned on a side of the body 102 opposite the forward-facing camera 104. The designation of the cameras as forward-facing and rear-facing is arbitrary, and the computing system 100 can include multiple cameras positioned on various sides of the body 102. The forward-facing camera 104 and the rear-facing camera 112 can each be configured to capture images in the visible spectrum.

[0044] The display 106 can represent a cathode ray tube (CRT) display, a light emitting diode (LED) display, a liquid crystal (LCD) display, a plasma display, an organic light emitting diode (OLED) display, or any other type of display known in the art. In some embodiments, the display 106 can display a digital representation of a current image captured by the forward-facing camera 104, the rear-facing camera 112, and / or the infrared camera 114, and / or an image that can be captured or recently captured by one or more of these cameras. Thus, the display 106 can function as a viewfinder for the cameras. The display 106 can also support touchscreen functionality, which can enable adjustment of settings and / or configuration of any aspect of the computing system 100.

[0045] The forward-facing camera 104 can include an image sensor and associated optical elements such as a lens. The forward-facing camera 104 can provide zoom capability or can have a fixed focal length. In other embodiments, interchangeable lenses can be used with the forward-facing camera 104. The forward-facing camera 104 can have a variable mechanical aperture and mechanical and / or electronic shutter. The forward-facing camera 104 can also be configured to capture still images, video images, or both. Further, the forward-facing camera 104 can represent a single-view, stereo, or multi-view camera. The rear-facing camera 112 and / or the infrared camera 114 can be similarly or differently arranged. Additionally, one or more of the forward-facing camera 104, the rear-facing camera 112, or the infrared camera 114 can be an array of one or more cameras.

[0046] Either or both of the forward-facing camera 104 and the rear-facing camera 112 can include or be associated with an illumination assembly that provides a field of light in the visible spectrum to illuminate a target object. For example, the illumination assembly can provide flash or constant illumination of the target object. The illumination assembly can also be configured to provide a field of light that includes one or more of structured light, polarized light, and light having a particular spectral content. Other types of fields of light known and used to recover three-dimensional (3D) models from objects are also possible in the context of embodiments herein.

[0047] The infrared pattern projector 116 can be configured to project an infrared structured light pattern onto a target object. In one example, the infrared projector 116 can be configured to project a dot pattern and / or a flood pattern. Thus, the infrared projector 116 can be used in conjunction with the infrared camera 114 to determine a plurality of depth values corresponding to different physical features of the target object.

[0048] That is, the infrared projector 116 can project a known and / or predetermined dot pattern onto a target object, and the infrared camera 114 can capture an infrared image of the target object that includes the projected dot pattern. The computing system 100 can then determine a correspondence between regions in the captured infrared image and particular portions of the projected dot pattern. Given the location of the infrared projector 116, the location of the infrared camera 114, and the location of regions within the captured infrared image that correspond to particular portions of the projected dot pattern, the computing system 100 can then use triangulation to estimate a depth to a surface of the target object. By repeating this step for different regions corresponding to different portions of the projected dot pattern, the computing system 100 can estimate depths for various physical features or portions of the target object. In this way, the computing system 100 can be used to produce a three-dimensional (3D) model of the target object.

[0049] The computing system 100 can also include an ambient light sensor that can continuously or from time to time determine an ambient brightness of a scene (e.g., in terms of visible light and / or infrared light) that the cameras 104, 112, and / or 114 can capture. In some implementations, the ambient light sensor can be used to adjust a display brightness of the display 106. Additionally, the ambient light sensor can be used to determine, or assist in determining, an exposure length of one or more of the cameras 104, 112, or 114.

[0050] The computing system 100 can be configured to use the display 106, as well as the forward-facing camera 104, the rear-facing camera 112, and / or the forward-facing infrared camera 114 to capture images of a target object. The captured images can be a plurality of still images or a video stream. Image capture can be triggered by activating the button 108, pressing a soft key on the display 106, or by some other mechanism. Depending on the implementation, images can be automatically captured at particular time intervals, e.g., upon pressing the button 108, upon suitable lighting conditions of the target object, upon movement of the computing system 100 a predetermined distance, or according to a predetermined capture schedule.

[0051] As described above, the functionality of the computing system 100 can be integrated into a computing device such as a wireless computing device, a cellular telephone, a tablet computer, a laptop computer, etc. For purposes of example, Figure 2is a simplified block diagram illustrating some components of an example computing device 200 that can include a camera assembly 224.

[0052] By way of example, and not limitation, the computing device 200 can be a cellular telephone (e.g., a smartphone), a still camera, a video camera, a computer (such as a desktop computer, a notebook computer, a tablet computer, or a handheld computer), a personal digital assistant (PDA), a home automation component, a digital video recorder (DVR), a digital television, a remote control, a wearable computing device, a game console, a robotic device, or some other type of device. As shown, the computing device 200 can include a communication interface 202, a user interface 204, a processor 206, a data store 208, and a camera assembly 224, all of which can be communicatively coupled together by a system bus, network, or other connection mechanism 210. Figure 2

[0053] The communication interface 202 can allow the computing device 200 to communicate with other devices, access networks, and / or transport networks using analog or digital modulation. Thus, the communication interface 202 can facilitate circuit- switched and / or packet-switched communications, such as plain old telephone service (POTS) communications and / or Internet Protocol (IP) or other packetized communications. For example, the communication interface 202 can include a chipset and antenna arrangement for wireless communication with a radio access network or access point. Moreover, the communication interface 202 can take the form of or include a wired interface, such as an Ethernet, Universal Serial Bus (USB), or High-Definition Multimedia Interface (HDMI) port. The communication interface 202 can also take the form of or include a wireless interface, such as a Wi-Fi, Bluetooth, Global Positioning System (GPS), or wide-area wireless interface (e.g., WiMAX or 3GPP Long-Term Evolution (LTE)). However, other forms of physical layer interface and other types of standard or proprietary communication protocols can be used over the communication interface 202. Moreover, the communication interface 202 can include multiple physical communication interfaces (e.g., a Wi-Fi interface, a Bluetooth interface, and a wide-area wireless interface).

[0054] ​​The user interface 204 can be used to allow the computing device 200 to interact with a human or non-human user, such as receiving input from and providing output to the user. Thus, the user interface 204 can include input components such as a keypad, keyboard, touch-sensitive panel, computer mouse, trackball, joystick, microphone, etc. The user interface 204 can also include one or more output components, such as a display screen, which can be combined with a touch-sensitive panel, for example. The display screen can be based on CRT, LCD, and / or LED technology, or other technologies now known or later developed. The user interface 204 can also be configured to produce audible output via a speaker, speaker jack, audio output port, audio output device, headphones, and / or other similar devices. The user interface 204 can also be configured to receive and / or capture audible utterances, noises, and / or signals through a microphone and / or other similar devices.

[0055] In some embodiments, the user interface 204 can include a display that functions as a viewfinder for still and / or video camera functionality (e.g., in both the visible and infrared spectrum) supported by the computing device 200. Additionally, the user interface 204 can include one or more buttons, switches, knobs, and / or dials that facilitate configuration and focusing of the camera functionality, as well as capture of images. Some or all of these buttons, switches, knobs, and / or dials can be implemented through a touch-sensitive panel.

[0056] The processor 206 can include one or more general-purpose processors (e.g., microprocessors) and / or one or more special-purpose processors (e.g., digital signal processors (DSPs), graphics processing units (GPUs), floating point units (FPUs), network processors, or application-specific integrated circuits (ASICs)). In some cases, special-purpose processors can be capable of image processing, image alignment, and image merging, among other things. The data storage 208 can include one or more volatile and / or non-volatile storage components, such as magnetic, optical, flash, or organic memory, and can be integrated in whole or part with the processor 206. The data storage 208 can include removable and / or non-removable components.

[0057] The processor 206 can be capable of executing program instructions 218 (e.g., compiled or interpreted program logic and / or machine code) stored in the data storage 208 to perform various functions described herein. Thus, the data storage 208 can include a non-transitory computer-readable medium having program instructions stored therein that, when executed by the computing device 200, cause the computing device 200 to perform any of the methods, processes, or operations disclosed in the specification and / or drawings. Execution of the program instructions 218 by the processor 206 can result in the processor 206 using the data 212.

[0058] As an example, program instructions 218 can include an operating system 222 (e.g., an operating system kernel, device drivers, and / or other components) installed on computing device 200 and one or more application programs 220 (e.g., a camera function, an address book, an email, a web browser, a social network, an audio-to-text function, a text translation function, and / or a gaming application). Similarly, data 212 can include operating system data 216 and application data 214. Operating system data 216 can be primarily accessed by operating system 222, while application data 214 can be primarily accessed by one or more application programs 220. Application data 214 can be arranged in a file system, which can be visible or hidden to a user of computing device 200.

[0059] Application programs 220 can communicate with operating system 222 through one or more application programming interfaces (APIs). These APIs can facilitate, for example, application programs 220 reading and / or writing application data 214, transmitting or receiving information via communication interface 202, receiving and / or displaying information on user interface 204, and so on. In some dialects, application programs 220 can be referred to simply as "apps." Additionally, application programs 220 can be downloaded to computing device 200 through one or more online application stores or application markets. However, application programs can also be installed on computing device 200 in other ways, such as via a web browser or through a physical interface on computing device 200 (e.g., a USB port).

[0060] Camera component 224 can include, without limitation, an aperture, a shutter, a recording surface (e.g., photographic film and / or an image sensor), a lens, a shutter button, an infrared projector, and / or a visible light projector. Camera component 224 can include components configured to capture images in the visible spectrum (e.g., electromagnetic radiation having wavelengths of 400-700 nanometers) as well as components configured to capture images in the infrared spectrum (e.g., electromagnetic radiation having wavelengths of 701 nanometers-1 millimeter). Camera component 224 can be controlled at least in part by software executed by processor 206.

[0061] III. Example Slot Attention Model

[0062] Figure 3A block diagram of a slot attention model 300 is shown. The slot attention model 300 can include a value function 308, a key function 310, a query function 312, a slot attention calculator 314, a slot update calculator 316, a slot vector initializer 318, and a neural network memory unit 320. The slot attention model 300 can be configured to receive input data 302 as input, which can include feature vectors 304-306. The input data 302 can alternatively be referred to as a perceptual representation. The input data 302 can correspond to and / or represent an input frame of an input data sequence.

[0063] The slot attention model 300 can be configured to produce slot vectors 322-324 based on the input data 302. The feature vectors 304-306 can represent distributed representations of entities in the input data 302, while the slot vectors 322-324 can represent entity-centric representations of these entities. The slot attention model 300 and its components can represent a combination of hardware and / or software components configured to implement the functions described herein. The slot vectors 322-324 can collectively define a latent representation of the input data 302. In some cases, the latent representation can represent a compression of the information contained in the input data 302. Thus, in some implementations, the slot attention model 300 can be used as and / or considered a machine learning encoder. Accordingly, the slot attention model 300 can be used for image reconstruction, text translation, and / or other applications that utilize machine learning encoders. Unlike certain other latent representations, each slot vector of the latent representation can capture characteristics of a corresponding entity or entities in the input data 302, and can do so without relying on an assumption about an order in which the entities are described with respect to the input data 302.

[0064] The input data 302 can represent various types of data, including, for example, image data (e.g., red-green-blue image data or grayscale image data), depth image data, point cloud data, audio data, time series data, and / or textual data, among others. In some cases, the input data 302 can be captured and / or produced by one or more sensors such as a visible light camera (e.g., the camera 104), a near-infrared camera (e.g., the infrared camera 114), a thermal camera, a stereo camera, a time-of-flight (ToF) camera, a light detection and ranging (LIDAR) device, a radio detection and ranging (RADAR) device, and / or a microphone, among others. In other cases, the input data 302 can additionally or alternatively include data produced by one or more users (e.g., words, sentences, paragraphs, and / or documents) or computing devices (e.g., a rendered three-dimensional environment, a time series graph), among others.

[0065] Input data 302 can be processed by one or more machine learning models (e.g., by an encoder model) to produce feature vectors 304-306. Each of feature vectors 304-306 can include a plurality of values, where each value corresponds to a particular dimension of the feature vector. In some implementations, the plurality of values of each feature vector can collectively represent an embedding of at least a portion of input data 302 in a vector space defined by the one or more machine learning models. For example, when input data 302 is an image, each of feature vectors 304-306 can be associated with one or more pixels in the image, and can represent various visual features of the one or more pixels. In some cases, the one or more machine learning models used to process input data 302 can include a convolutional neural network. Accordingly, feature vectors 304-306 can represent a mapping of convolutional features of input data 302, and can thus include outputs of various convolutional filters.

[0066] Each respective one of feature vectors 304-306 can include a position embedding that indicates a portion of input data 302 represented by the respective feature vector. For example, feature vectors 304-306 can be determined by adding a position embedding to convolutional features extracted from input data 302. Encoding a position associated with each respective one of feature vectors 304-306 as part of the respective feature vector, rather than through an order in which the respective feature vector is provided to slot attention model 300, allows feature vectors 304-306 to be provided to slot attention model 300 in a variety of different orders. Accordingly, including a position embedding as part of feature vectors 304-306 makes slot vectors 322-324 produced by slot attention model 300 permutation-invariant with respect to feature vectors 304-306.

[0067] For example, in the case of images, the position embedding can be produced by constructing a W x H x 4 tensor, where W and H represent the width and height, respectively, of the map of convolutional features of the input data 302. Each of the four values associated with each respective pixel along the W x H map can represent the position of the respective pixel relative to the image’s boundaries, borders, and / or edges along the corresponding directions of the image (i.e., up, down, right, and left). In some cases, each of the four values can be normalized to a range of 0 to 1, inclusive. The W x H x 4 tensor can be projected by a learnable linear mapping to the same dimensionality as the convolutional features (i.e., the same dimensionality as the feature vectors 304-306). The projected W x H x 4 tensor can then be added to the convolutional features to produce the feature vectors 304-306, thereby embedding the feature vectors 304-306 with position information. In some implementations, the sum of the projected W x H x 4 tensor and the convolutional features can be processed by one or more machine learning models (e.g., one or more multi-layer perceptrons) to produce the feature vectors 304-306. Similar position embeddings can also be included in the feature vectors 304-306 for other types of perceptual representations.

[0068] The feature vectors 304-306 can be provided as input to the key function 310. The feature vectors 304-306 can include N vectors each having I dimensions. Thus, in some implementations, the feature vectors 304-306 can be represented by an input matrix X having N rows (each row corresponding to a particular feature vector) and I columns.

[0069] In some implementations, the key function 310 can include a linear transformation represented by a key weight matrix W KEY with I rows and D columns. For example, the key function 310 can include a multi-layer perceptron that includes one or more hidden layers and utilizes one or more non-linear activation functions. The key function 310 (e.g., the key weight matrix W KEY ) can be learned during training of the slot attention model 300. The input matrix X can be transformed by the key function 310 to produce a key input matrix X KEY (e.g., X KEY = XW KEY ), which can be provided as input to the slot attention calculator 314. The key input matrix X KEY may include N rows and D columns.

[0070] The feature vectors 304-306 can also be provided as input to the value function 308. In some implementations, the value function 308 can include a linear transformation represented by a value weight matrix W VALUElinear transformation represented by a matrix. For example, the value function 308 can include a multilayer perceptron that includes one or more hidden layers and utilizes one or more non-linear activation functions. The value function 308 (e.g., the value weight matrix W VALUE ) can be learned during training of the slot attention model 300. The input matrix X can be transformed by the value function 308 to produce a value input matrix X VALUE (e.g., X VALUE = XW VALUE ), which can be provided as input to the slot update calculator 316. The value input matrix X VALUE may include N rows and D columns.

[0071] Because the dimensions of the key weight matrix W KEY and the value weight matrix W VALUE do not depend on the number N of feature vectors 304-306, different values of N can be used during training and during testing / usage of the slot attention model 300. For example, the slot attention model 300 can be trained on perception inputs with N = 1024 feature vectors, but can be used with N = 512 feature vectors or N = 2048 feature vectors. However, because at least one dimension of the key weight matrix W KEY and the value weight matrix W VALUE does depend on the dimension I of the feature vectors 304-306, the same value of I can be used during training and during testing / usage of the slot attention model 300.

[0072] The slot vector initializer 318 can be configured to initialize each of the slot vectors 322-324 stored by the neural network memory unit 320. In one example, the slot vector initializer 318 can be configured to initialize each of the slot vectors 322-324 with, for example, random values selected from a normal (i.e., Gaussian) distribution. In other examples, the slot vector initializer 318 can be configured to initialize one or more respective slot vectors of the slot vectors 322-324 with a“seed” value that is configured to cause the one or more respective slot vectors to attend / bind to and thereby represent a particular entity contained within the input data 302. For example, when processing image frames of a video, the slot vector initializer 318 can be configured to initialize the slot vectors 322-324 of a second image frame based on values of the slot vectors 322-324 determined with respect to a first image frame that precedes the second image frame. Thus, particular ones of the slot vectors 322-324 can be caused to represent the same entity across image frames of a video. Other types of sequential data can be similarly“seeded” by the slot vector initializer 318.

[0073] The slot vectors 322-324 can include K vectors each having S dimensions. Thus, in some implementations, the slot vectors 322-324 can be represented by an output matrix Y having K rows (each row corresponding to a particular slot vector) and S columns.

[0074] In some implementations, the query function 312 can include a linear transformation represented by a query weight matrix W QUERY having S rows and D columns. For example, the query function 312 can include a multi-layer perceptron that includes one or more hidden layers and utilizes one or more non-linear activation functions. The query function 312 (e.g., the query weight matrix W QUERY ) can be learned during training of the slot attention model 300. The output matrix Y can be transformed by the query function 312 to produce a query input matrix Y QUERY (e.g., Y QUERY = YW QUERY ), which can be provided as input to the slot attention calculator 314. The query output matrix Y QUERY may include K rows and D columns. Thus, the dimension D can be shared by the value function 308, the key function 310, and the query function 312.

[0075] Further, since the dimension of the query weight matrix W QUERY does not depend on the number K of slot vectors 322-324, different values of K can be used during training and during testing / usage of the slot attention model 300. For example, the slot attention model 300 can be trained with K = 7 slot vectors, but can be used with K = 5 slot vectors or K = 11 slot vectors. Thus, the slot attention model 300 can be configured to generalize over different numbers of slot vectors 322-324 without explicit training, although training and using the slot attention model 300 with the same number of slot vectors 322-324 can improve performance. However, since at least one dimension of the query weight matrix W QUERY does depend on the dimension S of the slot vectors 322-324, the same value of S can be used during training and during testing / usage of the slot attention model 300.

[0076] The slot attention calculator 314 can be configured to determine the attention matrix 340 based on the key input matrix X KEY produced by the key function 310 and the query input matrix Y QUERY produced by the query function 312. In particular, the slot attention calculator 314 can be configured to compute the dot product between the key input matrix X KEY and the transpose of the query output matrix Y QUERY . In some implementations, the slot attention calculator 314 can also divide the dot product by D (i.e., WVALUE , W KEY and / or W QUERY the square root of the number of columns of the matrix. Thus, the slot attention calculator 314 can implement the function where M represents a non-normalized version of the attention matrix 340 and can comprise N rows and K columns.

[0077] The slot attention calculator 314 can be configured to determine the attention matrix 340 by normalizing the values of the matrix M with respect to the output axis (i.e., with respect to the slot vectors 322-324). Thus, the values of the matrix M can be normalized along its rows (i.e., along the dimension K corresponding to the number of slot vectors 322-324). Thus, each value in each respective row can be normalized with respect to the K values contained in the respective row.

[0078] Thus, the slot attention calculator 314 can be configured to determine the attention matrix 340 by normalizing each respective value of the plurality of values of each respective row of the matrix M with respect to the plurality of values of the respective row. In particular, the slot attention calculator 314 can determine the attention matrix 340 according to where A i,j indicates the value at the position corresponding to row i and column j of the attention matrix 340, which can alternatively be referred to as the attention matrix A. Normalizing the matrix M in this way can cause the slots to compete with each other to represent a particular entity. The function implemented by the slot attention calculator 314 for computing A i,j can be referred to as a softmax function. The attention matrix A (i.e., the attention matrix 340) can comprise N rows and K columns.

[0079] In other implementations, the matrix M can be transposed prior to normalization, and the values of the matrix M T may thus be normalized along its columns (i.e., along the dimension K corresponding to the number of slot vectors 322-324). Thus, each value in each respective column of the matrix M T may be normalized with respect to the K values contained in the respective column. The slot attention calculator 314 can determine a transposed version of the attention matrix 340 according to where A indicates the value at the position corresponding to row i and column j of the transposed attention matrix 340, which can alternatively be referred to as the transposed attention matrix A T . However, the transposed attention matrix 340 can still be determined by normalizing the values of the matrix M with respect to the output axis (i.e., with respect to the slot vectors 322-324).

[0080] The slot update calculator 316 can be configured to update the slot vectors 322-324 based on the value input matrix XVALUE and the attention matrix 340 to determine an update matrix 342. In one implementation, the slot update calculator 316 can be configured to determine the update matrix 342 by determining the dot product of the transpose of the attention matrix A and the value input matrix X VALUE Thus, the slot update calculator 316 can implement the function U WEIGHTED SUM = A T X VALUE where the attention matrix A can be viewed as specifying the weights for the weighted sum computation, and the value input matrix X VALUE may be viewed as specifying the values for the weighted sum computation. The update matrix 342 can thus be represented by U WEIGHTED SUM which can include K rows and D columns.

[0081] In another implementation, the slot update calculator 316 can be configured to determine the update matrix 342 by determining the dot product of the transpose of the attention weight matrix W ATTENTION and the value input matrix X VALUE The elements / entries of the attention weight matrix W ATTENTION may be defined as or as the transpose thereof Thus, the slot update calculator 316 can implement the function U WEIGHTED MEAN = (W ATTENTION ) T X VALUE where the matrix A can be viewed as specifying the weights for the weighted average computation, and the value input matrix X VALUE may be viewed as specifying the values for the weighted average computation. The update matrix 342 can thus be represented by U WEIGHTED MEAN which can include K rows and D columns.

[0082] The update matrix 342 can be provided as input to the neural network memory unit 320, which can be configured to update the slot vectors 322-324 based on previous values of the slot vectors 322-324 (or predicted slot vectors generated based on the slot vectors 322-324) and the update matrix 342. The neural network memory unit 320 can include gated recurrent units (GRUs) and / or long short-term memory (LSTM) networks, as well as other neural networks or machine learning-based memory units configured to store and / or update the slot vectors 322-324. For example, in addition to GRUs and / or LSTMs, the neural network memory unit 320 can include one or more feedforward neural network layers configured to further modify the values of the slot vectors 322-324 after modification by the GRUs and / or LSTMs (and before being provided to the task-specific machine learning model 330).

[0083] In some implementations, the neural network memory unit 320 can be configured to update each of the slot vectors 322-324 during each processing iteration, rather than only some of the slot vectors 322-324 during each processing iteration. Training the neural network memory unit 320 to update the values of the slot vectors 322-324 based on the previous values of the slot vectors 322-324 (or a predicted slot vector generated based on the slot vectors 322-324) and based on the update matrix 342, rather than using the update matrix 342 as the updated values of the slot vectors 322-324, can improve accuracy and / or speed up convergence of the slot vectors 322-324.

[0084] The slot attention model 300 can be configured to generate the slot vectors 322-324 in an iterative manner. That is, the slot vectors 322-324 can be updated one or more times before being passed as input to the task-specific machine learning model 330. For example, the slot vectors 322-324 can be updated three times before being considered “ready” for use by the task-specific machine learning model 330. In particular, initial values for the slot vectors 322-324 can be assigned to the slot vectors 322-324 by the slot vector initializer 318. When the initial values are random, they can not accurately represent the entities contained in the input data 302. Accordingly, the feature vectors 304-306 and the randomly initialized slot vectors 322-324 can be processed by the components of the slot attention model 300 to refine the values of the slot vectors 322-324, thereby generating updated slot vectors 322-324.

[0085] After this first iteration or pass through the slot attention model 300, each of the slot vectors 322-324 can begin to attend to and / or bind to one or more corresponding entities contained in the input data 302 and thus represent the one or more corresponding entities. The feature vectors 304-306 and the now updated slot vectors 322-324 can again be processed by the components of the slot attention model 300 to further refine the values of the slot vectors 322-324, thereby generating another update to the slot vectors 322-324. After this second iteration or pass through the slot attention model 300, each of the slot vectors 322-324 can continue to attend to and / or bind to the one or more corresponding entities with increased strength, thereby representing the one or more corresponding entities with increased accuracy.

[0086] Further iterations can be performed, and each additional iteration can produce some improvement in the accuracy of each slot vector 322-324 to represent its corresponding one or more entities. After a predetermined number of iterations, the slot vectors 322-324 can converge to an approximately stable set of values, resulting in no additional accuracy improvement. Thus, the number of iterations of the slot attention model 300 can be selected based on (i) a desired level of representation accuracy of the slot vectors 322-324 and / or (ii) a desired processing time before the slot vectors 322-324 can be used by the task-specific machine learning model 330.

[0087] The task-specific machine learning model 330 can represent a plurality of different tasks, including supervised and unsupervised learning tasks. In some implementations, the task-specific machine learning model 330 can be co-trained with the slot attention model 300. Thus, depending on the particular task associated with the task-specific machine learning model 330, the slot attention model 300 can be trained to produce slot vectors 322-324 that are suitable and provide values useful in performing the particular task. In particular, the learning parameters associated with one or more of the value function 308, the key function 310, the query function 312, and / or the neural network memory unit 320 can be changed as a result of training based on the particular task associated with the task-specific machine learning model 330. In some implementations, the slot attention model 300 can be trained using adversarial training and / or contrastive learning, among other training techniques.

[0088] Compared to alternative methods for determining entity-centered representations, the slot attention model 300 can take less time to train (e.g., 24 hours compared to 7 days for an alternative method performed on the same computing hardware) and consume less memory resources (e.g., allowing a batch size of 64 compared to a batch size of 4 for an alternative method performed on the same computing hardware). In some implementations, the slot attention model 300 can also include one or more layer normalization. For example, layer normalization can be applied to the feature vectors 304-306 before being transformed by the key function 310, to the slot vectors 322-324 before being transformed by the query function 312, and / or to the slot vectors 322-324 after being at least partially updated by the neural network memory unit 320. Layer normalization can improve stability and speed up convergence of the slot attention model 300.

[0089] IV. Example Slot Vectors

[0090] Figure 4An example of a plurality of slot vectors that change with respect to a particular perceptual representation during a process of processing iterations by the slot attention model 300 is shown graphically. In this example, the input data 302 is represented by an image 400 that includes three entities: entity 410 (i.e., a circular object); entity 412 (i.e., a square object); and entity 414 (i.e., a triangular object). The image 400 can be processed by one or more machine learning models to produce feature vectors 304-306, each of which is represented by a corresponding grid element of a grid overlaid above the image 400. Thus, the leftmost grid element in the top row of the grid can represent the feature vector 304, the rightmost grid element in the bottom row of the grid can represent the feature vector 306, and the grid elements in between can represent other feature vectors. Thus, each grid element can represent a plurality of vector values associated with a corresponding feature vector.

[0091] Figure 4 The plurality of slot vectors is shown to have four slot vectors. However, in general, the number of slot vectors can be modifiable. For example, the number of slot vectors can be selected to be at least equal to the number of entities expected to be present in the input data 302, such that each entity can be represented by a corresponding slot vector. Thus, in the example shown, the number of slot vectors can be selected to be equal to the number of entities 410, 412, and 414 included in the image 400. Figure 4 In the example shown, the four slot vectors provided exceed the number of entities (i.e., three entities 410, 412, and 414) included in the image 400. In cases where the number of entities exceeds the number of slot vectors, one or more slot vectors can represent two or more entities.

[0092] The slot attention model 300 can be configured to process the feature vector associated with the image 400 and the initial values (e.g., randomly initialized) of the four slot vectors to produce slot vectors having values 402A, 404A, 406A, and 408A. The slot vector values 402A, 404A, 406A, and 408A can represent the output of a first iteration (lx) of the slot attention model 300. The slot attention model 300 can also be configured to process the feature vector and the slot vectors having values 402A, 404A, 406A, and 408A to produce slot vectors having values 402B, 404B, 406B, and 408B. The slot vector values 402B, 404B, 406B, and 408B can represent the output of a second iteration (2x) of the slot attention model 300. The slot attention model 300 can also be configured to process the feature vector and the slot vectors having values 402B, 404B, 406B, and 408B to produce slot vectors having values 402C, 404C, 406C, and 408C. The slot vector values 402C, 404C, 406C, and 408C can represent the output of a third iteration (3x) of the slot attention model 300. Visualizations of the slot vector values 402A, 404A, 406A, 408A, 402B, 404B, 406B, 408B, 402C, 404C, 406C, 408C can represent visualizations of the attention masks based on the attention matrix 340 and / or visualizations of the reconstruction masks produced by the task-specific machine learning model 330 at each iteration, among other examples.

[0093] The first slot vector (associated with values 402A, 402B, and 402C) can be configured to attend to and / or bind to the entity 410, thereby representing the attributes, characteristics, and / or features of the entity 410. In particular, after the first iteration of the slot attention model 300, the first slot vector can represent aspects of the entity 410 and the entity 412, as shown by the black filled regions in the visualization of the slot vector value 402A. After the second iteration of the slot attention model 300, the first slot vector can represent a larger portion of the entity 410 and a smaller portion of the entity 412, as shown by the increased black filled region of the entity 410 and the decreased black filled region of the entity 412 in the visualization of the slot vector value 402B. After the third iteration of the slot attention model 300, the first slot vector can approximately exclusively represent the entity 410 and can no longer represent the entity 412, as shown by the entity 410 being completely filled with black and the entity 412 being completely filled with white in the visualization of the slot vector value 402C. Thus, as the slot attention model 300 updates and / or refines the values of the first slot vector, the first slot vector can converge and / or focus on representing the entity 410. This attention and / or convergence of a slot vector to one or more entities is a result of the mathematical structure of the components of the slot attention model 300 and the task-specific training of the slot attention model 300.

[0094] The second slot vector (associated with values 404A, 404B, and 404C) can be configured to attend to and / or bind to entity 412, thereby representing attributes, characteristics, and / or features of entity 412. In particular, after the first iteration of slot attention model 300, the second slot vector can represent aspects of entity 412 and entity 410, as shown by the black filled region in the visualization of slot vector value 404A. After the second iteration of slot attention model 300, the second slot vector can represent a greater portion of entity 412 and can no longer represent entity 410, as shown by the increased black filled region of entity 412 and the entity 410 illustrated as completely filled with white in the visualization of slot vector value 404B. After the third iteration of slot attention model 300, the second slot vector can approximately exclusively represent entity 412 and can continue to no longer represent entity 410, as shown by the entity 412 completely filled with black and the entity 410 completely filled with white in the visualization of slot vector value 404C. Thus, as the slot attention model updates and / or refines the values of the second slot vector, the second slot vector can converge and / or focus on representing entity 412.

[0095] The third slot vector (associated with values 406A, 406B, and 406C) can be configured to attend to and / or bind to entity 414, thereby representing attributes, characteristics, and / or features of entity 414. In particular, after the first iteration of slot attention model 300, the third slot vector can represent aspects of entity 414, as shown by the black filled region in the visualization of slot vector value 406A. After the second iteration of slot attention model 300, the third slot vector can represent a greater portion of entity 414, as shown by the increased black filled region of entity 414 in the visualization of slot vector value 404B. After the third iteration of slot attention model 300, the third slot vector can represent approximately the entire entity 414, as shown by the entity 412 completely filled with black in the visualization of slot vector value 406C. Thus, as the slot attention model updates and / or refines the values of the third slot vector, the third slot vector can converge and / or focus on representing entity 414.

[0096] The fourth slot vector (associated with values 408A, 408B, and 408C) can be configured to attend to and / or bind to the background features of image 400, thereby representing the attributes, characteristics, and / or features of the background. Specifically, after the first iteration of slot attention model 300, the fourth slot vector can represent approximately the entire background and the respective portions of entities 410 and 414 that have not yet been represented by slot vector values 402A, 404A, and / or 406A, as shown by the black filled regions in the visualization of slot vector value 408A. After the second iteration of slot attention model 300, the fourth slot vector can represent approximately the entire background and a smaller portion of entities 410 and 414 that have not yet been represented by slot vector values 402B, 404B, and / or 406B, as shown by the black filled regions of the background and the reduced black filled regions of entities 410 and 414 in the visualization of slot vector value 408B. After the third iteration of slot attention model 300, the fourth slot vector can approximately represent the entire background exclusively, as shown by the background that is completely filled with black and entities 410, 412, and 414 that are completely filled with white in the visualization of slot vector value 408C. Thus, as slot attention model updates and / or refines the values of the fourth slot vector, the fourth slot vector can converge and / or focus on representing the background of image 400.

[0097] In some implementations, rather than representing the background of image 400, the fourth slot vector can assume a predetermined value that indicates that the fourth slot vector is not used to represent entities. Thus, the background can not be represented. Alternatively or additionally, when additional slot vectors are provided (e.g., a fifth slot vector), the additional vectors can represent portions of the background or can not be utilized. Thus, in some cases, slot attention model 300 can distribute the representation of the background among multiple slot vectors. In some implementations, a slot vector can treat entities within a perceptual representation the same as their background. Specifically, any of the slot vectors can be used to represent the background and / or entities (e.g., the background can be treated as another entity). Alternatively, in other implementations, one or more of the slot vectors can be reserved to represent the background.

[0098] The plurality of slot vectors can be invariant with respect to the order of the feature vectors and equivariable with respect to each other. That is, for a given initialization of the slot vectors, the order in which the feature vectors are provided at the input of slot attention model 300 does not affect the order and / or values of the slot vectors. However, different initializations of the slot vectors can affect the order of the slot vectors independent of the order of the feature vectors. Further, for a given set of feature vectors, the set of values of the slot vectors can remain constant, but the order of the slot vectors can be different. Thus, different initializations of the slot vectors can affect the pairing of the slot vectors to the entities contained in the perceptual representation, but the entities can still be represented with approximately the same set of slot vector values.

[0099] V. Example Slot Vector Sequence System

[0100] Figure 5 An example slot vector sequence system 500 is shown, which can be configured to produce, track, and / or update slot vectors for a sequence of input data frames. In particular, the slot vector sequence system 500 can include an initializer model 518, an encoder model 508, a slot vector model 512, a predictor model 516, and a task-specific model 530. The slot vector sequence system 500 can be configured to produce an output 532 based on an input data sequence 502.

[0101] The input data sequence 502 can include an input frame 504 through an input frame 506 (i.e., input frames 504-506). The input data sequence 502 can represent, for example, a video (e.g., color, grayscale, depth, etc.), a sequence of point clouds, a sequence of RADAR images, and / or a waveform or other time series (represented in the time domain and / or frequency domain), among others. Thus, in some cases, the input data sequence 502 can be produced by one or more sensors and can represent a physical environment. The input data sequence 502 can be represented as t where t e {1,..., T}, xi represents the input frame 504, and x T represents the input frame 506.

[0102] The output 532 can include, for example, a reconstruction of the input data sequence 502 or portions thereof, feature detection and / or classification within the input data sequence 502, control commands for a device (e.g., a robotic device, a vehicle, and / or a computing device or aspects thereof), and / or other data that can be used to perform one or more tasks based on the input data sequence 502. The output 532 can depend on the task for which the slot vector sequence system 500 is used and / or the type of input data sequence 502.

[0103] The encoder model 508 can be configured to produce feature vectors 510 based on the input frames of the input data sequence 502. For example, the encoder model 508 can be configured to produce a corresponding instance of the feature vectors 510 for each respective input frame of the input frames 504-506. For example, the feature vectors 510 can be represented as where f() can include trainable parameters of the slot vector sequence system 500. In some cases, the feature vectors 510 can collectively form and / or be arranged into a feature map. In some cases, the feature vectors 510 can represent the feature vectors 304-306 discussed in connection with Figure 3 the feature vectors 304-306 discussed. Thus, the input matrix X can represent the feature vectors h t = X t for a particular time step (i.e., h t where I = D enc .

[0104] The initializer model 518 can be configured to produce initialization slot vectors 522. In some implementations, the initialization slot vectors 522 can be produced based on a selection of one or more entities in the input frame 504 (i.e., the first input frame of the input data sequence 502) to be represented by the selected slot vector (i.e., based on a desired entity-to-slot vector mapping). Thus, the initialization slot vectors 522 can be configured such that the slot vector model 512 uses the selected one of the slot vectors 514 to represent a particular object. The initializer model 518 can be configured to receive, for a particular slot vector of the slot vectors 514, an indication of a corresponding entity (included in the input frame 504) to be represented by the particular slot vector. The indication of the corresponding entity can include, for example, a bounding box around the corresponding entity, a centroid of the corresponding entity and / or coordinates of other points within the corresponding entity, and / or a segmentation mask associated with the corresponding entity.

[0105] The initializer model 518 can include, for each respective slot vector of the K slot vector positions in the slot vectors 514, a corresponding machine learning model (e.g., an MLP) configured to transform the indication of the corresponding entity (possibly in combination with the input frame 504) into an initial set of values for the respective slot vector.

[0106] In other implementations, the initializer model 518 can be configured to produce the initialization slot vectors 522 randomly (e.g., by selecting values from a normal distribution). Thus, in some cases, the initializer model 518 can be similar to and / or can include aspects of the slot vector initializer discussed in Figure 3 Figure 3 In other implementations, the initialization vectors 522 can be learned during training and held fixed during inference. Thus, the same set of initialization slot vectors 522 can be used for multiple different input data sequences.

[0107] The slot vector model 512 can be configured to produce the slot vectors 514 based on the feature vectors 510 and other slot vectors. The other slot vectors can include, for example, the initialization slot vectors 522 (when the slot vector model is processing the first frame xi of the input data sequence 502), instances of previously determined slot vectors 514, and / or the predicted slot vectors 520. The slot vectors 514 can be represented as

[0108] In one example implementation, the slot vector model 512 can include the slot attention model 300 and / or components thereof. Thus, the slot vectors 514 can be represented in conjunction with the slot vectors 322-324 discussed in Figure 3 Thus, the output matrix Y can represent the slot vector values S t t at a particular time step (i.e., S t ​​of sets, where S = D. The function implemented by the slot vector model 512 can alternatively be represented as where GRU() represents a gated recurrent unit (although other neural network memory units such as LSTM can alternatively be used), denotes a Hadamard product, and In some implementations, the output of GRU() can be further processed by an MLP before being produced as the output of the slot vector model 512. GRU(), v(), k(), q(), and / or MLP() can comprise trainable parameters of the slot vector sequence system 500.

[0109] In other example implementations, the slot vector model 512 can comprise a detection transformer model and / or a tracking transformer model. The detection transformer model can be implemented as detailed in the paper titled “End-to-End Object Detection with Transformers” authored by Nicolas Carion et al. and published as arXiv:2005.12872v3. The tracking transformer model can be implemented as detailed in the paper titled “TrackFormer: Multi-Object Tracking with Transformers” authored by Tim Meinhardt et al. and published as arXiv:2101.02702v2.

[0110] The predictor model 516 can be configured to produce the predicted slot vector 520 based on the slot vector 514 and / or previously determined instances of the predicted slot vector 520. The predicted slot vector 520 can represent a transition of the slot vector 514 from a time (e.g., t) associated with the slot vector 514 to a subsequent time (e.g., t+1). Thus, while the slot vector 514 represents properties of entities contained in a particular input frame of the input data sequence 502 based on processing of that particular input frame, the predicted slot vector 520 can represent an expected future state of the properties of these entities in future input frames of the input data sequence 502. The expected future state of the properties of these entities can be determined by the predictor model 516 without needing to process the future input frames of the input data sequence 502. Thus, the predictor model 516 can be configured to determine an expected future behavior of one or more entities based on properties of the one or more entities as represented by a most recent set of one or more slot vectors of the at least one entity. In some cases, the slot vector 514 can alternatively be referred to as an observed slot vector 514 because its values are based on observed input frames of the input data sequence 502.

[0111] The predicted slot vector can be represented as The predictor model 516 can be configured to implement a function in The function LN() represents layer normalization, the function MLP() represents a multilayer perceptron, and the function MultiHeadSelfAttn() represents multi-head dot product attention. For example, the function MultiHeadSelfAttn() can represent a multi-head self-attention block from a transformer model, as described in detail in the paper titled "Attention Is All You Need" by Ashish Vaswani et al., published as arXiv:1706.03762v5. MultiHeadSelfAttn() and / or MLP() can include trainable parameters of a slot vector sequence system of 500.

[0112] The task-specific model 530 can be configured to generate output 532 based on slot vector 514 and / or predicted slot vector 520. The task-specific model 530 may include and / or be similar to... Figure 3 As shown and combined Figure 3 The task-specific machine learning model 330 is discussed. The slot vector model 512, the predictor model 516, and / or the task-specific model 530 can be jointly trained to obtain slot vectors 514 and / or predicted slot vectors 520 that are useful to the task-specific model 530 when performing its corresponding task. That is, the joint training of the slot vector model 512, the predictor model 516, and / or the task-specific model 530 allows the task-specific model 530 to "understand" the values ​​of slot vectors 514 and / or predicted slot vectors 520.

[0113] By relying on the predicted slot vector 520, the task-specific model 530 can operate based on the expected future state of the environment. For example, since the predicted slot vector 520 represents the expected future state of the entity represented by the input data sequence 502, the task-specific model 530 may be able to produce the output 532 faster than it would be if it operated solely based on the slot vector 514. In the context of, for example, the operation of an autonomous vehicle, this expectation of the future state can improve vehicle responsiveness and / or safety, as well as other benefits.

[0114] Additionally, slot vector 514 and / or predicted slot vector 520 can assist task-specific model 530 in processing and / or responding to occlusion of one or more entities. For example, if input frames 504-506 are processed individually rather than as a sequence, the slot vectors of input frames where one or more entities are occluded may not represent the occluded entities. However, when input frames 504-506 are processed as a sequence, the temporary occlusion of said one or more entities can be explicitly represented by the values ​​of slot vector 514 and / or predicted slot vector 520, rather than completely omitting the representation of these entities.

[0115] In some implementations, the task-specific model 530 may be a slot decoder model, which may be configured to generate reconstructions of one or more input frames of the input data sequence 502 based on slot vector 514 and / or predicted slot vector 520. In the case that the input data sequence 502 is video, the slot decoder model may additionally or alternatively predict optical flow between two or more input frames of the input data sequence 502. Input frame reconstruction and / or optical flow prediction may be performed, for example, as part of the training of the slot vector sequence system 500.

[0116] Element-wise (e.g., pixel-wise) reconstruction loss functions can be used to quantify the quality of the slot vector sequence system 500 producing slot vector 514 and / or predicted slot vector 520 by measuring the extent to which the original input and / or its aspects can be reconstructed based on these vectors. For example, the pixel-wise reconstruction loss function can be configured to compare the pixel values ​​of the reconstructed input frame with those of the original input frame. In another example, the pixel-wise reconstruction loss function can be configured to compare the pixel values ​​of the predicted optical flow between two input frames with the ground truth optical flow measured between those two input frames. The trainable parameters of the slot vector sequence system 500 can be adjusted until, for example, the loss value produced by the loss function is reduced below a threshold loss value.

[0117] In embodiments where the input data sequence 502 includes image data, the task-specific model 530 can individually decode the slot vectors 514 and / or predict each of the slot vectors 520 using a spatial broadcast decoder. Specifically, each slot can be broadcast onto a two-dimensional grid that can be augmented with positional embeddings. Each grid can be decoded using a convolutional neural network (whose parameters can be shared across each slot vector in the slot vectors 514) to produce an output of size W x H x 4, where W and H represent the width and height of the reconstructed slot-specific image data, respectively, and the additional 4 dimensions represent the red, green, and blue channels and their unnormalized alpha masks. The alpha masks can be normalized on the slot-specific images using a softmax function and can be used as blending weights to recombine and / or blend the slot-specific images into a final reconstruction of the original image frames. In other examples, the slot decoder model can be and / or may include aspects of a patch-based decoder.

[0118] The slot vector sequence system 500 can operate on the input data sequence 502 in a sequential manner. For example, the slot vector sequence system 500 can be configured to use an initializer model 518 to determine the initialization slot vector 522 based on the input frame x1. The encoder model 508 can be configured to generate the feature vector h1 based on the input frame x1. The slot vector model 512 can be configured to generate the slot vector based on the initialization slot vector 522 and the feature vector h1. Since input frame xi is the first input frame of input data sequence 502, previous instances of slot vector 514 and / or predicted slot vector 520 can not be available for input data sequence 502. Thus, for input frame xi, slot vector model 512 can generate slot vector 514 independently of predicted slot vector 520 and / or previous instances of slot vector 514, as opposed to using initialization slot vector 522 to "prime" slot vector model 512. Predictor model 516 can be configured to generate predicted slot vector S2 based on slot vector 514 predicted slot vector S2.

[0119] For input frame x2, encoder model 508 can be configured to generate feature vector h2 based on input frame x2. Slot vector model 512 can be configured to generate slot vector 514 based on feature vector h2 and predicted slot vector S2. Predictor model 516 can be configured to generate predicted slot vector S3 based on slot vector 514 Since initialization is no longer required in view of the availability of predicted slot vector S2, slot vector 514 can be generated independently of initialization slot vector 522 at t = 2. Predictor model 516 can be configured to generate predicted slot vector S3 based on slot vector 514 predicted slot vector S3. This sequence can be repeated for all T input frames of input data sequence 502.

[0120] In some implementations, predictor model 516 can be autoregressive. For example, predictor model 516 can be configured to generate predicted slot vector S2 based on any input frames xi of input data sequence 502 that slot vector model 512 has skipped. t-1 corresponding predicted slot vector S2. t-1 Instead of generating predicted slot vector S2 based on slot vector 514 predicted slot vector S2. t For example, when slot vector model 512 is configured to generate slot vector 514 for every kth input frame (e.g., 1, k+1, 2k+1,...) of input data sequence 502, predictor model 516 can be configured to generate predicted slot vector 520 for every frame of input data sequence, where every kth instance of predicted slot vector 520 is based on slot vector 514, and all other instances of predicted slot vector 520 are based on the immediately preceding instance of predicted slot vector 520. Thus, slot vector model 512 can be configured to omit (e.g., periodically or aperiodically) processing of some input frames of input data sequence 502, and predictor model 516 can be configured to perform autoregressive operations for any input frames that slot vector model 512 omits. Such an arrangement can be beneficial, for example, when the size and / or amount of computational resources involved in using slot vector model 512 exceeds the size and / or amount of computational resources involved in using predictor model 516.

[0121] Figure 6An example sequence of slot vectors corresponding to an example sequence of images is shown. Specifically, images 610, 612-614, and 616 (i.e., images 610-616) provide one example of input frames 504-506, and can collectively form a video. Image 610 can correspond to time t = 1, image 612 can correspond to time t = 2, image 614 can correspond to time t = 5, and image 616 can correspond to time t = 6. Images 610-616 can depict motion of entities 410, 412, and 414 from time t = 1 to time t = 6. Specifically, entity 410 can remain in a fixed position, entity 412 can move diagonally to the lower left corner, and entity 414 can move horizontally to the right.

[0122] At each of times t = 1 to t = 6, slot vector sequence system 500 can be configured to produce a corresponding set of values for a plurality of slot vectors (e.g., produce a corresponding instance of slot vector 514 for each of times t = 1 to t = 6). Slot vector sequence system 500 can be configured to produce: (i) for time t = 1, a slot vector having values 602A, 604A, 606A, and 608A; (ii) for time t = 2, a slot vector having values 602B, 604B, 606B, and 608B; (iii) for time t = 5, a slot vector having values 602C, 604C, 606C, and 608C; (iv) for time t = 6, a slot vector having values 602D, 604D, 606D, and 608D; and (v) for other times not shown in FIG. 6, slot vectors (not shown) having other corresponding values. Figure 6 At each of times t = 1 to t = 6, slot vector sequence system 500 can be configured to produce a corresponding set of values for a plurality of slot vectors (e.g., produce a corresponding instance of slot vector 514 for each of times t = 1 to t = 6). Slot vector sequence system 500 can be configured to produce: (i) for time t = 1, a slot vector having values 602A, 604A, 606A, and 608A; (ii) for time t = 2, a slot vector having values 602B, 604B, 606B, and 608B; (iii) for time t = 5, a slot vector having values 602C, 604C, 606C, and 608C; (iv) for time t = 6, a slot vector having values 602D, 604D, 606D, and 608D; and (v) for other times not shown in FIG. 6, slot vectors (not shown) having other corresponding values.

[0123] The first slot vector having values 602A, 602B, 602C, and 602D can represent entity 410 across time, the second slot vector having values 604A, 604B, 604C, and 604D can represent entity 412 across time, the third slot vector having values 606A, 606B, 606C, and 606D can represent entity 414 across time, and the fourth slot vector having values 608A, 608B, 608C, and 608D can represent the background across time. The binding, attention, and / or mapping of different slot vectors to respective entities can be a result of deterministic selection or random initialization by initializer model 518, as discussed in connection with Figure 5 For example, by initializing the first slot vector based on a bounding box, a segmentation mask, a center point, and / or other indication of entity 410, the first slot vector can be caused to represent entity 410 over time.

[0124] At Figure 4 each set of values shown in FIG. 6 for slot vectors represents a result of processing a single input frame multiple times to cause the slot vector to converge to the corresponding entity, unlike Figure 4 the initialization of slot vectors discussed in connection withFigure 6 The set of slot vector values ​​obtained by processing sequences of different input frames is shown. Because (i) the predictor model 516 is configured to model and / or predict the temporal variations of the corresponding entities and / or (ii) the variations between successive input frames of the input sequence are relatively small, the value of each slot vector changes gradually rather than abruptly. Figure 6 At least some of the slot vectors shown can remain bound and / or converge to their corresponding entities over time.

[0125] The predicted slot vector corresponding to input frames 610-616 may include... Figure 6 The number of vectors shown is the same as the slot vector (i.e., 4 vectors). In some cases, when... Figure 6 When visualized in this way, the values ​​of the predicted slot vectors can be expected to appear as interpolations between consecutive values ​​of the corresponding slot vectors. For example, the predicted slot vectors corresponding to time t=2 can appear as interpolations between values ​​602A and 602B, 604A and 604B, 606A and 606B, and 608A and 608B. Therefore, one iteration of the predictor model 516 can be considered as providing at least part of the benefits of one iteration of the slot vector model 512, thereby allowing the resulting slot vectors to remain bound and / or converge to the corresponding entities after at least one iteration of the slot vector model 512.

[0126] When predictor model 516 is highly accurate in predicting the future attributes of the corresponding entity, the predicted slot vector corresponding to time t=2 can appear to be approximately values ​​602B, 604B, 606B, and 608B. That is, for a given time step, the accurately determined predicted slot vector value can be approximately equal to (e.g., differing from it by no more than a threshold) the corresponding slot vector value determined for the given time step based on the input frame of that given time step. When predictor model 516 is relatively inaccurate in predicting the future attributes of the corresponding entity, the predicted slot vector corresponding to time t=2 can appear to be approximately values ​​602A, 604A, 606A, and 608A, or any value different from 602A, 602B, 604A, 604B, 606A, 606B, 608A, and 608B.

[0127] When the predictor model 516 is used autoregressively to generate a predicted slot vector for time t = 6 based on (i) slot vector values 602A, 604A, 604A, and 606A produced by the slot vector model 512 for time t = 1 and (ii) predicted slot vector values produced by the predictor model 516 for times t = 2 to t = 5 (but without the slot vector model 512 producing slot vector values for times t = 2 to t = 5), the predicted slot vector values for time t = 6 can differ from slot vector values 602D, 604D, 606D, and 608D by more than these values would differ if the predictor model 516 were not used autoregressively (i.e., if the slot vector model 512 produced slot vector values for times t = 2 to t = 5 to be used in determining the predicted slot vector values for time t = 6). That is, the autoregressive operation of the predictor model 516 can sacrifice some accuracy to reduce the use of computational resources by the slot vector sequence system 500.

[0128] VI. Additional Example Operations

[0129] Figure 7 Flowcharts illustrating operations related to generating, tracking, and / or updating slot vectors based on input frames of an input data sequence are shown. The input data sequence can include an ordered sequence of multiple input frames, each of which can be time- tagged according to its position in the sequence. These operations can be performed by the computing system 100, the computing device 200, the slot attention model 300, and / or the slot vector sequence system 500, among others. Figure 7 Embodiments can be simplified by removing any one or more of the features shown therein. Further, these embodiments can be combined with any previous figure or feature, aspect, and / or implementation described herein in other manners.

[0130] Block 700 can involve obtaining a first plurality of feature vectors representing content of a first input frame of an input data sequence and a second plurality of feature vectors representing content of a second input frame of the input data sequence. The second input frame can be subsequent to the first input frame. For example, the second input frame can be consecutive to the first input frame, or can be separated from the first input frame by one or more other input frames.

[0131] Block 702 can involve generating a first plurality of slot vectors based on the first plurality of feature vectors and through a slot vector model. Each respective slot vector of the first plurality of slot vectors can represent an attribute of a corresponding entity as represented in the first input frame. The corresponding entity can be represented differently in different input frames, resulting in different values for the respective slot vector over time.

[0132] Block 704 can involve generating, based on the first plurality of slot vectors and by a predictor model, a plurality of predicted slot vectors, the plurality of predicted slot vectors including, for each respective slot vector of the first plurality of slot vectors, a corresponding predicted slot vector. The corresponding predicted slot vector can represent a transition in the attribute of the corresponding entity from the first input frame to the second input frame.

[0133] Block 706 can involve generating, based on the plurality of predicted slot vectors and the second plurality of feature vectors and by a slot vector model, a second plurality of slot vectors, the second plurality of slot vectors including, for each respective slot vector of the first plurality of slot vectors, a corresponding slot vector. The corresponding slot vector can represent the attribute of the corresponding entity as represented in the second input frame.

[0134] Block 708 can involve determining the output based on one or more of: (i) at least one predicted slot vector of the plurality of predicted slot vectors or (ii) at least one slot vector of the second plurality of slot vectors.

[0135] In some embodiments, generating the plurality of predicted slot vectors can include determining a plurality of self-attention scores based on the first plurality of slot vectors and generating the plurality of predicted slot vectors based on the plurality of self-attention scores.

[0136] In some embodiments, determining the plurality of self-attention scores can include determining the plurality of self-attention scores using two or more self-attention heads, where each respective self-attention head of the two or more self-attention heads includes a corresponding set of model parameters according to which the first plurality of slot vectors are processed. For example, the plurality of self-attention scores can be determined using a multi-headed self-attention block from a Transformer model as detailed in the paper titled “Attention Is All You Need” authored by Ashish Vaswani et al. and published as arXiv: 1706.03762v5.

[0137] In some embodiments, generating the plurality of predicted slot vectors based on the plurality of self-attention scores can include generating, by an artificial neural network, an intermediate output based on (i) the plurality of self-attention scores and (ii) a first sum of the first plurality of slot vectors and determining the plurality of predicted slot vectors based on the intermediate output and a second sum of the first sum.

[0138] In some embodiments, generating the plurality of predicted slot vectors can include processing the first plurality of slot vectors by a graph neural network. Each respective slot vector of the first plurality of slot vectors can correspond to a node modeled by the graph neural network, and relationships between the first plurality of slot vectors can correspond to edges modeled by the graph neural network. The plurality of predicted slot vectors can be generated based on an output of the graph neural network.

[0139] In some embodiments, the graph neural network can be an interaction neural network configured to compare multiple different ordered pairs of slot vectors in the first plurality of slot vectors. For example, the graph neural network can be an interaction neural network as detailed in the paper titled “Interaction Networks for Learning about Objects, Relations and Physics” authored by Peter W. Battaglia et al. and published as arXiv:1612.00222v1.

[0140] In some embodiments, a plurality of initialization slot vectors can be generated. The plurality of initialization slot vectors can include, for each respective slot vector in the first plurality of slot vectors, a corresponding initialization slot vector. The first plurality of slot vectors can be further generated based on the plurality of initialization slot vectors.

[0141] In some embodiments, generating the plurality of initialization slot vectors can include, for a particular slot vector in the first plurality of slot vectors, obtaining an indication of a corresponding entity to be represented by the particular slot vector, and generating, based on the indication of the corresponding entity and for the particular slot vector, a corresponding initialization slot vector configured to bind the particular slot vector to the corresponding entity.

[0142] In some embodiments, the indication of the corresponding entity to be represented by the particular slot vector can include one or more of: (i) a bounding box associated with the corresponding entity, (ii) coordinates of a centroid of the corresponding entity, or (iii) a segmentation mask associated with the corresponding entity.

[0143] In some embodiments, the plurality of initialization slot vectors can be generated by a machine learning model configured to process the indication of the corresponding entity and generate the corresponding initialization slot vector. For example, the machine learning model can include one or more of: (i) a multi-layer perceptron configured to process the indication of the corresponding entity and generate the corresponding initialization slot vector, or (ii) a convolutional neural network configured to process the indication of the corresponding entity and generate the corresponding initialization slot vector.

[0144] In some embodiments, generating the plurality of initialization slot vectors can include generating, for at least one slot vector in the first plurality of slot vectors, a corresponding initialization slot vector based on a value selected from a normal distribution.

[0145] In some embodiments, generating the plurality of initialization slot vectors can include learning the plurality of initialization slot vectors during training.

[0146] In some embodiments, determining the output can include processing the one or more of (i) the at least one of the plurality of predicted slot vectors or (ii) the at least one of the second plurality of slot vectors through a machine learning model to produce a task-specific output for the particular task. The machine learning model can have been trained to produce the task-specific output based on one or more of the slot vectors or the predicted slot vectors. Determining the output can further include providing instructions to perform the particular task in accordance with the task-specific output produced by the machine learning model.

[0147] In some embodiments, the input data sequence can include a video, the first input frame can include a first image, the second input frame can include a second image, and the corresponding entity can include an object represented in the video.

[0148] In some embodiments, the input data sequence can include a series of consecutive point clouds of an environment, the first input frame can include a first point cloud, the second input frame can include a second point cloud, and the corresponding entity can include an object represented in the series of consecutive point clouds.

[0149] In some embodiments, the input data sequence can include a waveform, the first input frame can include a first spectrogram, the second input frame can include a second spectrogram, and the corresponding entity can include a waveform pattern represented in the waveform.

[0150] In some embodiments, the predicted slot vectors can be permutation symmetric with respect to each other such that for a plurality of different initializations of the first plurality of slot vectors with respect to a given input data sequence, a set of values of the plurality of predicted slot vectors can be approximately constant and an order of the plurality of predicted slot vectors can be variable. The predicted slot vectors can be permutation symmetric with respect to the first plurality of feature vectors such that for a plurality of different permutations of the first plurality of feature vectors, a set of values of the plurality of predicted slot vectors can be approximately constant. Thus, a predictor model configured to produce the predicted slot vectors can be permutation equivariant.

[0151] In some embodiments, the slot vector model can include at least one of (i) a slot attention model, (ii) a detection transformer model, or (iii) a tracking transformer model.

[0152] In some embodiments, one or more of (i) the slot vector model or (ii) the predictor model have been trained through a training process that includes obtaining a first plurality of training feature vectors representing content of a first training input frame of a training input data sequence, a second plurality of training feature vectors representing content of a second training input frame of the training input data sequence, and a ground truth output corresponding to the second training input frame. The second training input frame can be subsequent to the first training input frame. The training process can further include generating a first plurality of training slot vectors based on the first plurality of training feature vectors and through the slot vector model. Each respective training slot vector of the first plurality of training slot vectors can represent an attribute of a corresponding training entity as represented in the first training input frame. The training process can further include generating a plurality of predicted training slot vectors based on the first plurality of training slot vectors and through the predictor model, the plurality of predicted training slot vectors including, for each respective training slot vector of the first plurality of training slot vectors, a corresponding predicted training slot vector. The corresponding predicted slot vector can represent a transition of the attribute of the corresponding training entity from the first training input frame to the second training input frame. The training process can further include generating a second plurality of training slot vectors based on the plurality of predicted training slot vectors and the second plurality of training feature vectors and through the slot vector model, the second plurality of training slot vectors including, for each respective training slot vector of the first plurality of training slot vectors, a corresponding training slot vector. The corresponding training slot vector can represent the attribute of the corresponding training entity as represented in the second training input frame. The training process can further include determining a training output based on decoding at least one of (i) the plurality of predicted training slot vectors or (ii) the second plurality of training slot vectors, determining a loss value based on a comparison of the ground truth output and the training output through a loss function, and adjusting one or more parameters of one or more of (i) the slot vector model or (ii) the predictor model based on the loss value.

[0153] In some embodiments, determining the training output can include generating a reconstruction of the second training input frame based on one or more of (i) the plurality of predicted training slot vectors or (ii) the second plurality of training slot vectors.

[0154] In some embodiments, determining the training output can include determining a flow vector field representing a change between the first training input frame and the second training input frame based on one or more of (i) the plurality of predicted training slot vectors or (ii) the second plurality of training slot vectors.

[0155] In some embodiments, the first input frame and the second input frame can be non-consecutive, and thus separated by one or more intermediate input frames. Accordingly, the plurality of prediction slot vectors can include a plurality of sequential subsets of prediction slot vectors, where some of the sequential subsets are generated for the one or more intermediate input frames. The intermediate input frames can not be processed by the slot vector model, thus allowing the slot vector model to operate at a varying frame rate that is different from a frame rate at which the predictor model operates.

[0156] For example, when the second input frame is separated from the first input frame by one intermediate frame, a first plurality of prediction slot vectors can be generated for the intermediate frame based on the first plurality of slot vectors, and a second plurality of prediction slot vectors can be generated for the second input frame based on the first plurality of prediction slot vectors. The slot vector model can not process the first plurality of prediction slot vectors, but can process the second plurality of prediction slot vectors. Accordingly, when the first input frame and the second input frame are non-consecutive, the predictor model can be configured to generate prediction slot vectors based on other previously generated prediction slot vectors. Thus, the predictor model can be considered to be autoregressive. Each respective prediction slot vector of the second plurality of prediction slot vectors can be considered to represent a transition of the attribute of the corresponding entity from the first input frame to the second input frame, as the second plurality of prediction slot vectors reflect a transition of the attribute from the first input frame to the intermediate input frame due to the second plurality of prediction slot vectors being based on the first plurality of prediction slot vectors.

[0157] Figure 8 A flowchart illustrating operations related to training of a machine learning model configured to generate, track, and / or update slot vectors based on input frames of an input data sequence is shown. These operations can be performed by the computing system 100, the computing device 200, and / or the slot attention model 300, among others. Figure 8 Embodiments can be simplified by the removal of any one or more of the features shown therein. Further, these embodiments can be combined with any previous figure or feature, aspect, and / or implementation described herein in other manners.

[0158] Block 800 can involve obtaining a first plurality of training feature vectors representing content of a first training input frame of a training input data sequence, a second plurality of training feature vectors representing content of a second training input frame of the training input data sequence, and a ground truth output corresponding to the second training input frame. The second training input frame can be subsequent to the first training input frame.

[0159] Block 802 can involve generating, based on the first plurality of training feature vectors and by a slot vector model, a first plurality of training slot vectors. Each respective training slot vector of the first plurality of training slot vectors can represent an attribute of a corresponding training entity as represented in the first training input frame.

[0160] Block 804 can involve generating, based on the first plurality of training slot vectors and by a predictor model, a plurality of predicted training slot vectors, the plurality of predicted training slot vectors including, for each respective training slot vector of the first plurality of training slot vectors, a corresponding predicted training slot vector. The corresponding predicted training slot vector can represent a transition in the attribute of the corresponding training entity from the first training input frame to the second training input frame.

[0161] Block 806 can involve generating, based on the plurality of predicted training slot vectors and the second plurality of training feature vectors and by a slot vector model, a second plurality of training slot vectors, the second plurality of training slot vectors including, for each respective training slot vector of the first plurality of training slot vectors, a corresponding training slot vector. The corresponding training slot vector can represent the attribute of the corresponding training entity as represented in the second training input frame.

[0162] Block 808 can involve determining the training output based on decoding at least one of (i) the plurality of predicted training slot vectors or (ii) the second plurality of training slot vectors.

[0163] Block 810 can involve determining, by a loss function, a loss value based on a comparison of a ground truth output to the training output.

[0164] Block 812 can involve adjusting, based on the loss value, one or more parameters of one or more of (i) a slot vector model configured to generate slot vectors or (ii) a predictor model configured to generate predicted slot vectors.

[0165] VII. CONCLUSION

[0166] The present disclosure is not limited in scope to the particular embodiments described in this application, which are intended as illustrations of various aspects. Numerous modifications and changes can be made with respect to the methods and apparatus described herein without departing from the scope of the disclosure. Methods and apparatuses that are functionally equivalent to those described herein but are not identical thereto in terms of structure or design are within the scope of the disclosure. Such modifications and changes are intended to be within the scope of the claims appended hereto.

[0167] The above detailed description describes various features and operations of the disclosed systems, devices, and methods with reference to the accompanying drawings, which illustrate exemplary embodiments. In the drawings, like reference numbers generally indicate identical, functionally similar, and / or structurally similar elements. The example embodiments described herein and in the drawings are not meant to be limiting. Other embodiments can be utilized, and other changes can be made without departing from the scope of the subject matter presented herein. It will be readily understood that the aspects of the present disclosure, as generally described herein, and illustrated in the Figures, can be arranged, substituted, combined, separated, and designed in a wide variety of different configurations.

[0168] With respect to any or all of the message flow diagrams, scenarios, and flowcharts in the drawings and as discussed herein, each step, block, and / or communication can represent an information processing and / or information transmission according to example embodiments. Alternative embodiments include those in which the steps, blocks, transmissions, communications, requests, responses, and / or messages described are performed in a different order than shown or discussed, including substantially concurrently or in reverse order, depending on the functionality involved. Further, more or fewer blocks and / or operations can be utilized in alternative embodiments, and the messages flow diagrams, scenarios, and flowcharts discussed herein can be combined or split into separate diagrams, scenarios, and flowcharts. Figure 1 The specific arrangements shown in the drawings should not be taken as limiting. It will be apparent to those of ordinary skill in the art that other embodiments that include fewer or more of each element shown in a given drawing can be practiced within the scope of the example embodiments. Further, some of the illustrated elements can be combined or omitted, and / or additional elements can be added. Moreover, example embodiments can include elements that are not shown in the drawings.

[0169] Steps or blocks representing information processing can correspond to a circuit that can be configured to perform a particular logical function of the methods or techniques described herein. Alternatively or additionally, a block representing information processing can correspond to a module, segment, or portion of program code (including related data). The program code can include one or more instructions executable by a processor for implementing specific logical operations or acts in the methods or techniques. Program code and / or related data can be stored on any type of computer readable medium, such as a storage device including a random access memory (RAM), a disk drive, a solid-state drive, or another storage medium.

[0170] Computer readable media can also include non-transitory computer readable media, such as computer readable media that store data for short periods of time, like register memory, processor cache, and RAM. Computer readable media can also include non-transitory computer readable media that store program code and / or data for periods of time, like secondary or persistent long term storage, for example, readonly memory (ROM), optical or magnetic disks, solid-state drives, compact disks, etc. Computer readable media can also be any other volatile or non-volatile storage systems. Computer readable media can be considered computer readable storage media, or tangible storage devices, for example.

[0171] Further, steps or blocks representing one or more information transmissions can correspond to information transmissions between software and / or hardware modules within the same physical device. However, other information transmissions can be between software modules and / or hardware modules in different physical devices.

[0172] The specific arrangements shown in the drawings should not be taken as limiting. It will be apparent to those of ordinary skill in the art that other embodiments that include fewer or more of each element shown in a given drawing can be practiced within the scope of the example embodiments. Further, some of the illustrated elements can be combined or omitted, and / or additional elements can be added. Moreover, example embodiments can include elements that are not shown in the drawings.

[0173] While various aspects and embodiments have been disclosed herein, other aspects and embodiments will be apparent to those skilled in the art. The various aspects and embodiments disclosed herein are for purposes of exemplification and are not intended to be limiting, with the true scope being indicated by the following claims.

Claims

1. A computer-implemented method comprising: obtaining, by a machine learning model, a first plurality of feature vectors representing content of a first input frame of an input data sequence and a second plurality of feature vectors representing content of a second input frame of the input data sequence, wherein the second input frame is subsequent to the first input frame; generating, based on the first plurality of feature vectors and by a slot vector model of the machine learning model, a first plurality of slot vectors, wherein each respective slot vector of the first plurality of slot vectors represents an attribute of a corresponding entity as represented in the first input frame; generating, based on the first plurality of slot vectors and by a predictor model of the machine learning model, a plurality of predicted slot vectors, the plurality of predicted slot vectors including, for each respective slot vector of the first plurality of slot vectors, a corresponding predicted slot vector representing a transition of the attribute of the corresponding entity from the first input frame to the second input frame, wherein generating the plurality of predicted slot vectors comprises: determining a plurality of self-attention scores based on the first plurality of slot vectors; generating an intermediate output based on (i) the plurality of self-attention scores and (ii) a first sum of the first plurality of slot vectors; and determining the plurality of predicted slot vectors based on a second sum of the intermediate output and the first sum; generating, based on the plurality of predicted slot vectors and the second plurality of feature vectors and by the slot vector model, a second plurality of slot vectors, the second plurality of slot vectors including, for each respective slot vector of the first plurality of slot vectors, a corresponding slot vector representing an attribute of the corresponding entity as represented in the second input frame; determining, by the machine learning model, a task-specific output based on one or more of (i) at least one predicted slot vector of the plurality of predicted slot vectors or (ii) at least one slot vector of the second plurality of slot vectors; and causing a device to perform a task based on the task-specific output, wherein the input data sequence comprises a video, the first input frame comprises a first image, the second input frame comprises a second image, and the corresponding entity comprises an object represented in the video; or the input data sequence comprises a series of consecutive point clouds of an environment, the first input frame comprises a first point cloud, the second input frame comprises a second point cloud, and the corresponding entity comprises an object represented in the series of consecutive point clouds; or the input data sequence comprises a waveform, the first input frame comprises a first spectrogram, the second input frame comprises a second spectrogram, and the corresponding entity comprises a waveform pattern represented in the waveform.

2. The computer-implemented method of claim 1, wherein determining the plurality of self-attention scores comprises: determining the plurality of self-attention scores using two or more self-attention heads, wherein each respective self-attention head of the two or more self-attention heads includes a corresponding set of model parameters according to which the first plurality of slot vectors are processed.

3. The computer-implemented method of claim 1, wherein generating the plurality of predicted slot vectors comprises: processing the first plurality of slot vectors through a graph neural network, wherein each respective slot vector of the first plurality of slot vectors corresponds to a node modeled by the graph neural network, and wherein relationships between the first plurality of slot vectors correspond to edges modeled by the graph neural network; and generating the plurality of predicted slot vectors based on an output of the graph neural network.

4. The computer-implemented method of claim 3, wherein the graph neural network is an interaction neural network configured to compare pairs of different ordered slot vectors of the first plurality of slot vectors.

5. The computer-implemented method of claim 1, further comprising: generating a plurality of initialization slot vectors, the plurality of initialization slot vectors including, for each respective slot vector of the first plurality of slot vectors, a corresponding initialization slot vector, wherein the first plurality of slot vectors are generated further based on the plurality of initialization slot vectors.

6. The computer-implemented method of claim 5, wherein generating the plurality of initialization slot vectors comprises: for a particular slot vector of the first plurality of slot vectors, obtaining an indication of the corresponding entity to be represented by the particular slot vector; and generating, based on the indication of the corresponding entity and for the particular slot vector, the corresponding initialization slot vector configured to bind the particular slot vector to the corresponding entity.

7. The computer-implemented method of claim 6, wherein the indication of the corresponding entity to be represented by the particular slot vector includes one or more of: (i) a bounding box associated with the corresponding entity, (ii) coordinates of a centroid of the corresponding entity, or (iii) a segmentation mask associated with the corresponding entity, and wherein the plurality of initialization slot vectors are generated by an initializer model configured to process the indication of the corresponding entity and generate the corresponding initialization slot vector.

8. The computer-implemented method of claim 5, wherein generating the plurality of initialization slot vectors comprises: for at least one slot vector of the first plurality of slot vectors, generating the corresponding initialization slot vector based on a value selected from a normal distribution.

9. The computer-implemented method of claim 1, wherein determining the task-specific output comprises: processing, through the machine learning model, the one or more of (i) the at least one predicted slot vector of the plurality of predicted slot vectors or (ii) the at least one slot vector of the second plurality of slot vectors to generate a task-specific output for a particular task, wherein the machine learning model has been trained to generate task-specific outputs based on one or more of slot vectors or predicted slot vectors.

10. The computer-implemented method of claim 1, wherein the prediction slot vectors are permutation symmetric with respect to each other such that, for a plurality of different initializations of the first plurality of slot vectors with respect to a given input data sequence, a set of values of the plurality of prediction slot vectors is approximately constant and an order of the plurality of prediction slot vectors is variable, and wherein the prediction slot vectors are permutation symmetric with respect to the first plurality of feature vectors such that, for a plurality of different permutations of the first plurality of feature vectors, the set of values of the plurality of prediction slot vectors is approximately constant.

11. The computer-implemented method of claim 1, wherein the slot vector model comprises at least one of: (i) a slot attention model, (ii) a detection transformer model, or (iii) a tracking transformer model.

12. The computer-implemented method of claim 1, wherein one or more of (i) the slot vector model or (ii) the predictor model have been trained by a training process, the training process comprising: obtaining a first plurality of training feature vectors representing content of a first training input frame of a training input data sequence, a second plurality of training feature vectors representing content of a second training input frame of the training input data sequence, and a ground truth output corresponding to the second training input frame, wherein the second training input frame is subsequent to the first training input frame; generating a first plurality of training slot vectors based on the first plurality of training feature vectors, wherein each respective training slot vector of the first plurality of training slot vectors represents an attribute of a corresponding training entity as represented in the first training input frame; generating a plurality of prediction training slot vectors based on the first plurality of training slot vectors, the plurality of prediction training slot vectors including, for each respective training slot vector of the first plurality of training slot vectors, a corresponding prediction training slot vector representing a transition of the attribute of the corresponding training entity from the first training input frame to the second training input frame; generating a second plurality of training slot vectors based on the plurality of prediction training slot vectors and the second plurality of training feature vectors, the second plurality of training slot vectors including, for each respective training slot vector of the first plurality of training slot vectors, a corresponding training slot vector representing the attribute of the corresponding training entity as represented in the second training input frame; determining a training output based on decoding at least one of: (i) the plurality of prediction training slot vectors or (ii) the second plurality of training slot vectors; determining a loss value based on a comparison of the ground truth output and the training output by a loss function; and based on the loss value, adjusting one or more parameters of the one or more of (i) the slot vector model or (ii) the predictor model.

13. The computer-implemented method of claim 12, wherein determining the training output comprises: The reconstruction of the second training input frame is generated based on one or more of: (i) the plurality of predicted training slot vectors or (ii) the second plurality of training slot vectors.

14. The computer-implemented method of claim 12, wherein determining the training output comprises: determining a flow vector field representing changes between the first training input frame and the second training input frame based on one or more of: (i) the plurality of predicted training slot vectors or (ii) the second plurality of training slot vectors.

15. The computer-implemented method of claim 1, wherein the second input frame is separated from the first input frame by an intermediate input frame, and wherein generating the plurality of predicted slot vectors comprises: generating, based on the first plurality of slot vectors and through the predictor model, a plurality of intermediate predicted slot vectors, the plurality of intermediate predicted slot vectors including, for each respective slot vector in the first plurality of slot vectors, a corresponding intermediate predicted slot vector representing a transition of the property of the corresponding entity from the first input frame to the intermediate input frame; and generating, based on the plurality of intermediate predicted slot vectors and through the predictor model, the plurality of predicted slot vectors, the plurality of predicted slot vectors including, for each respective slot vector in the first plurality of slot vectors, the corresponding predicted slot vector representing a transition of the property of the corresponding entity from the intermediate input frame to the second input frame.

16. The computer-implemented method of claim 1, wherein the first plurality of slot vectors is generated using a first number of iterations of the slot vector model, wherein the second plurality of slot vectors is generated using a second number of iterations of the slot vector model, and wherein the second number of iterations is less than the first number of iterations.

17. The computer-implemented method of claim 1, wherein, causing the device to perform the task comprises: causing the device to interact with an environment represented by the sequence of input data.

18. A computing system comprising: a processor; and a non-transitory computer-readable medium having stored thereon instructions that, when executed by the processor, cause the processor to perform operations comprising: obtaining, through a machine learning model, a first plurality of feature vectors representing content of a first input frame of a sequence of input data and a second plurality of feature vectors representing content of a second input frame of the sequence of input data, wherein the second input frame is subsequent to the first input frame; generating, based on the first plurality of feature vectors and through a slot vector model of the machine learning model, a first plurality of slot vectors, wherein each respective slot vector in the first plurality of slot vectors represents a property of a corresponding entity as represented in the first input frame; ​ producing, based on the first plurality of slot vectors and by a predictor model of the machine learning model, a plurality of predicted slot vectors, the plurality of predicted slot vectors including, for each respective slot vector of the first plurality of slot vectors, a corresponding predicted slot vector representing a transition of the attribute of the corresponding entity from the first input frame to the second input frame, wherein producing the plurality of predicted slot vectors comprises: determining, based on the first plurality of slot vectors, a plurality of self-attention scores; producing, based on (i) the plurality of self-attention scores and (ii) a first sum of the first plurality of slot vectors, an intermediate output; and determining, based on a second sum of the intermediate output and the first sum, the plurality of predicted slot vectors; producing, based on the plurality of predicted slot vectors and the second plurality of feature vectors and by the slot vector model, a second plurality of slot vectors, the second plurality of slot vectors including, for each respective slot vector of the first plurality of slot vectors, a corresponding slot vector representing the attribute of the corresponding entity as represented in the second input frame; determining, by the machine learning model, a task-specific output based on one or more of (i) at least one predicted slot vector of the plurality of predicted slot vectors or (ii) at least one slot vector of the second plurality of slot vectors; and causing a device to perform a task based on the task-specific output, wherein the input data sequence comprises a video, the first input frame comprises a first image, the second input frame comprises a second image, and the corresponding entity comprises an object represented in the video; or the input data sequence comprises a series of consecutive point clouds of an environment, the first input frame comprises a first point cloud, the second input frame comprises a second point cloud, and the corresponding entity comprises an object represented in the series of consecutive point clouds; or the input data sequence comprises a waveform, the first input frame comprises a first spectrogram, the second input frame comprises a second spectrogram, and the corresponding entity comprises a waveform pattern represented in the waveform.

19. A non-transitory computer-readable medium having stored thereon instructions, which, when executed by a computing device, cause the computing device to perform operations comprising: obtaining, by a machine learning model, a first plurality of feature vectors representing content of a first input frame of an input data sequence and a second plurality of feature vectors representing content of a second input frame of the input data sequence, wherein the second input frame is subsequent to the first input frame; producing, based on the first plurality of feature vectors and by a slot vector model of the machine learning model, a first plurality of slot vectors, wherein each respective slot vector of the first plurality of slot vectors represents an attribute of a corresponding entity as represented in the first input frame; producing, based on the first plurality of slot vectors and by a predictor model of the machine learning model, a plurality of predicted slot vectors, the plurality of predicted slot vectors including, for each respective slot vector of the first plurality of slot vectors, a corresponding predicted slot vector representing a transition of the attribute of the corresponding entity from the first input frame to the second input frame, wherein producing the plurality of predicted slot vectors includes: determining, based on the first plurality of slot vectors, a plurality of self-attention scores; producing an intermediate output based on (i) the plurality of self-attention scores and (ii) a first sum of the first plurality of slot vectors; and determining the plurality of predicted slot vectors based on a second sum of the intermediate output and the first sum; producing, based on the plurality of predicted slot vectors and the second plurality of feature vectors and by the slot vector model, a second plurality of slot vectors, the second plurality of slot vectors including, for each respective slot vector of the first plurality of slot vectors, a corresponding slot vector representing the attribute of the corresponding entity as represented in the second input frame; determining, by the machine learning model, a task-specific output based on one or more of (i) at least one predicted slot vector of the plurality of predicted slot vectors or (ii) at least one slot vector of the second plurality of slot vectors; and causing a device to perform a task based on the task-specific output, wherein the input data sequence includes a video, the first input frame includes a first image, the second input frame includes a second image, and the corresponding entity includes an object represented in the video; or the input data sequence includes a series of consecutive point clouds of an environment, the first input frame includes a first point cloud, the second input frame includes a second point cloud, and the corresponding entity includes an object represented in the series of consecutive point clouds; or the input data sequence includes a waveform, the first input frame includes a first spectrogram, the second input frame includes a second spectrogram, and the corresponding entity includes a waveform pattern represented in the waveform.