Video creation by demonstration
The simulation system addresses the inefficiencies of existing video generation by using a generator neural network to create realistic simulations of complex environments with reduced computational resources, eliminating the need for explicit control signals.
Patent Information
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- GDM HOLDING LLC
- Filing Date
- 2025-11-14
- Publication Date
- 2026-05-21
AI Technical Summary
Existing video generation systems require substantial computational resources and explicit control signals, which are time-consuming to generate and costly, and often fail to produce realistic simulations that adhere to physical laws.
A simulation system using a generator neural network that learns to simulate complex physics directly from training data, generating videos by processing action data and context images without explicit control signals, leveraging spatial-temporal encoding and hardware accelerators like GPUs or TPUs.
The system efficiently generates realistic simulations of complex environments with temporal continuity and physical plausibility, reducing computational requirements and eliminating the need for costly control signal generation.
Smart Images

Figure US2025055595_21052026_PF_FP_ABST
Abstract
Description
[0001] Attorney Docket No.: 45288-0578W01
[0002] Video Creation by Demonstration
[0003] CROSS-REFERENCE TO RELATED APPLICATIONS
[0004] [1] This specification claims priority to U. S. Provisional Application No. 63 / 720,753, filed on November 14, 2024. The disclosure of the prior application is considered part of and is incorporated by reference in the disclosure of this application.
[0005] BACKGROUND
[0006] [2] This specification relates to processing data using machine learning models.
[0007] [3] Machine learning models receive an input and generate an output, e.g., a predicted output, based on the received input. Some machine learning models are parametric models and generate the output based on the received input and on values of the parameters of the model.
[0008] [4] Some machine learning models are deep models that employ multiple layers of models to generate an output for a received input. For example, a deep neural network is a deep machine learning model that includes an output layer and one or more hidden layers that each apply a non-linear transformation to a received input to generate an output.
[0009] SUMMARY
[0010] [5] This specification describes a system implemented as computer programs on one or more computers in one or more locations that generates a video item depicting one or more first actors performing one or more actions.
[0011] [6] The actors may comprise humans and / or animals and / or inanimate actors such as electromechanical apparatus (e.g., robot(s)).
[0012] [7] In general terms, the disclosure suggests that another video item (a second video item termed an "‘action video item”, also referred to as a “demonstration video”), depicting one or more second actors performing the one or more actions, is processed to obtain action data defining the one or more actions, and the action data is processed by a trained generator neural network, together with a context visual data item comprising at least one context image. The trained generator network has been trained to generate the video item depicting one or more first actors performing one or more actions, with at least one visual aspect of the video item being defined by the at least one context image. In particular, at least one visual aspect of the video is the same as in the at least one context image.
[0013] [8] The action data may additionally be indicative of the timing of the actions within the demonstration video and / or the manner in which the actions were performed, e.g. the orientations of the second actors at the time(s) in the demonstration video that they performed Attorney Docket No.: 45288-0578W01
[0014] the action(s), or interaction between the second actors and physical object(s) depicted in the demonstration video.
[0015] [9] Specifically, the at least one aspect of the video item may be the appearance of at least one of the first actor(s). That is, at least one of the first actor(s) may be different (in appearance) from at least one of the second actor(s), and the at least one context image may define the appearance of the at least one of the first actors.
[0016]
[0010] Alternatively or additionally, the at least one context image may define the appearance of an environment, and the video item may depict the one or more first actors performing the one of more actions in the environment depicted in the at least one environment.
[0017]
[0011] Thus, the video data item provides a simulation of how the one or more actions (which may also be called action concepts) would appear when carried out by the first actor(s) depicted by the at least one context image (e.g. instead of the second actor(s) if the first and second actor(s) are different), and / or in an environment depicted by the at least one context image (rather than the, possibly different, environment depicted in the action video item).
[0018]
[0012] The term ‘image’" is used here to mean a dataset which is at least one respective intensity value for each pixel of a two (or higher) dimensional array of pixels. For example, in black-and-white images, the intensity of each pixel can be represented as an integer or floating point number (i.e., a vector with one component) representing the brightness of the pixel. As another example, if the images are red-green-blue (RGB) images, each pixel can be represented as a vector with three integer or floating point components, which respectively represent the intensity of the red, green, and blue color of the pixel. The pixels of a video frame can be indexed by (x, y) coordinates, where x and y are integer values. Similarly, the term “video item” means an ordered sequence (“video sequence”) of images (i.e. a plurality of images, “frames”) depicting an evolution of a situation in an environment, e.g. actors moving within the environment. A “video sequence” is an ordered plurality of images which can be played, in the order, to a viewer to give an impression of smooth movement in the environment. In some examples, YUV or HSV color channels may be employed.
[0019]
[0013] The video item may comprise the context image(s) as frame(s). For example, the context image(s) may be the first frame(s) of the video item, such that the video item is a video sequence continuing the context image. In this case, the video item would show how the results of performing the action(s) in a situation depicted in the context image(s). Alternatively, in principle, the context image(s) may be the last frame(s) of the video item, such that the video item is video sequence leading up to the context image. In this case, the video item would show how the action(s) result in a situation depicted in the context image(s). However, this possibility Attorney Docket No.: 45288-0578W01
[0020] is not described further. In principle, the context image(s) may be at an intermediate location(s) in the video item, i.e. with one or more frames of the video item before the context image (or the first of multiple context images) and one or more frames of the video after the context image (or the last of multiple context images).
[0021]
[0014] The context image(s) may be captured from the real world, for example by at least one camera (e.g. a still camera, or a video camera with the context image(s) being frame(s) from a video sequence captured by the video camera). Thus, the video item is a simulation of how real-world actors perform the actions and / or how actors perform the action(s) in a real-world environment.
[0022]
[0015] The environment may be a physical environment. In some implementations, the physical environment comprises a real -world environment including a physical object.
[0023]
[0016] Similarly, the action video item may be a video captured from a real-world environment (optionally, a different real-world environment) using a video camera in which the second actors perform the one or more actions.
[0024]
[0017] The term ’ camera" as used herein is to be understood broadly. The cameras can include lidar or radar sensors within the environment for example. Furthermore, the context image(s), the video item and / or the action video item may be three-dimensional images, e.g., the context image and / or action video item may be captured by a camera such as a depth camera.
[0025]
[0018] The one or more actions depicted in the action video item may be actions to perform a task in the (e.g. real-world) environment, such as a manipulation task. For example, some or all of the actions may be actions performed on a physical object in the environment. Thus, the video item depicts the first actors performing the action(s) to perform the task. Thus, the video item can be used for guidance in how the one or more first actors should perform the task.
[0026]
[0019] An example is a case in which at least one of the first actors is an (e.g. real-world) electro-mechanical apparatus, such as a robot. The apparatus may be operative to move (e.g. by reconfiguration and / or by translation through its environment). For example, if the action video item shows the action(s) being performed by human(s), the video item may show some or all of the actions being performed by robot(s), or by human(s) and robot(s) together. Thus, the video item is a video simulation of the actions being performed by the first agents.
[0027]
[0020] The video item may be used, for example, in a decision of whether to perform the action(s) using the electro-mechanical apparatus(es). For example, the video item may be used to assess whether performing the action(s) using the apparatus(es) is feasible or efficient, so that a decision can be made whether to perform the action(s) using the apparatus(es). Optionally, a respective video item can be generated for each of a plurality of sets of first actor(s) which Attorney Docket No.: 45288-0578W01
[0028] are different electro-mechanical apparatus(es), and the video items can be compared to select a set of first actors (i.e. to select electro-mechanical apparatus(es)) to perform the action(s). The selected electro-mechanical apparatuses may then be used to perform the action(s).
[0029]
[0021] In another use of the video item, the electro-mechanical apparatus(es) which constitute the first actor(s) may be controlled to perform the one or more actions with a timing and / or a manner depicted by the video item, e.g. with the first actor(s) having the same locations and / or orientations as depicted in the video item when the actions are performed.
[0030]
[0022] In some cases, the video item may be provided to a user of the simulation system. The system may provide a representation of the simulation for display on a user device. For example, the system can generate images, video, virtual reality environments, augmented reality environments, etc., that depict the simulation.
[0031]
[0023] High-quality simulators may require substantial computational resources, which makes scaling up prohibitive. The simulation system described in this specification can generate simulations of complex environments over large numbers of time steps using fewer computational resources (e.g., memory and computing power) than some conventional simulation systems. For example, the simulation system can predict the state of a physical environment at a next time step by a single pass through the generator neural network, while conventional simulation systems may be required to perform a separate optimization at each time step. In particular, the simulation system generates simulations using a generator neural network that can learn to simulate complex physics directly from training data, and can generalize implicitly learned physics principles to accurately simulate a broader range of physical environments under different conditions than are directly represented in the training data. This also allows the system to generalize to larger and more complex settings than those used in training. In contrast, some conventional simulation systems require physics principles to be explicitly programmed, and must be manually adapted for the specific characteristics of each environment being simulated.
[0032]
[0024] The above described systems and methods can be configured for hardware acceleration. In such implementations, the method is performed by data processing apparatus comprising one or more computers and including one or more hardware accelerators units, e.g.. one or more GPUs (Graphics Processing Units) or TPUs (Tensor Processing Units). Thus, there is provided a simulation method which is specifically adapted for implementation using hardware accelerators.
[0033]
[0025] The extraction of the action data from the action video item can be performed using spatial-temporal encoding items for respective frames of the action video item. The spatial- Attorney Docket No.: 45288-0578W01
[0034] temporal encoding items are spatial-temporal feature data which encodes both ‘‘static” information which is “context data” encoding the visual appearance of elements depicted in the video item (that is, elements which are the second actors and / or the environment in which the second actors are located) depicted in the action video item, and action data encoding the action(s) the second actors take. The spatial-temporal encoding items for the respective frames of the action video item collectively form a spatial-temporal encoding of the action video item. Here “encoding” refers to processing data to form a revised dataset (also called “an encoding item” or “an embedding”), e.g. of reduced data size compared to the processed data. For example, an encoding item may have a data size (e.g. as measured in a number of bits) which is no more than 10%. or no more than 30%, of the size of the data from which it was produced.
[0035]
[0026] Note that although the actors and environment may be in common between frames, the positions of the actors and / or the parts of the environment which are visible, differ, so the context data is different for different frames of the action video item. In other words, the context data for a given frame may represent data in a given frame of the action video item which is correlated (e.g. to a level above a statistical threshold) with data in other frames of the action video item, rather than being identical for all frames of the action video item.
[0036]
[0027] The data encoding the action(s) may be obtained by subtracting the context data, from the spatial-temporal encoding. Specifically, the action video item may be additionally processed using a trained feature predictor neural network to extract context data, and this is subtracted from the spatial-temporal encoding, leaving the action data.
[0037]
[0028] The feature predictor network can include any appropriate types of neural network layers (e.g., fully connected layers, attention layers, pooling layers, convolutional layers, attention layers, and so forth) in any appropriate number (e.g., 2 layers, or 5 layers, or 10 layers) and connected in any appropriate configuration (e.g., as a directed graph of layers).
[0038]
[0029] The spatial-temporal encoding of the action item may be generated using a video encoder (spatial-temporal encoder) e.g. of a known kind, which first encodes frames of a video item (e.g. the action video item) separately (this may be done using a known spatial feature encoder (spatial encoder)) to form respective spatial feature encodings of the frames, and then aggregates the spatial feature encodings using a temporal aggregator (temporal aggregation layer) to form the spatial-temporal encoding which has a respective portion (spatial -temporal encoding item) for each frame. The spatial feature encoder can include any appropriate type(s) of neural network layers (e.g., embedding layers, convolution layers, attention layers, and so forth) in any appropriate number (e.g.. 2 layers, or 5 layers, or 10 layers) and connected in any appropriate configuration (e.g., as a directed graph of layers). Attorney Docket No.: 45288-0578W01
[0039]
[0030] The feature predictor neural network may operate on spatial feature encodings of individual frames generated using a spatial feature encoder (spatial encoder), to generate respective portions of the context data. If, as mentioned in the preceding paragraph, the spatial-temporal encoding of the action item is obtained by forming respective feature encodings of corresponding frames with a spatial encoder, and then temporally aggregating them, then the feature predictor neural network may optionally employ the same spatial feature encodings of the individual frames generated by the spatial encoder, thereby reducing the computational burden of performing the method. Note that, unlike the temporal aggregator, the feature predictor neural network processes the spatial feature encodings separately, to form respective portions of the context data.
[0040]
[0031] The feature predictor neural network and the unit which performs subtraction of the features predicted by the feature predictor from the spatial -temporal encoding, may be called an “appearance bottleneck”, which generates the action data. The units preceding the appearance bottleneck (i.e. the spatial encoder and the temporal aggregator) are an example of a vision encoder.
[0041]
[0032] The feature predictor neural network may be obtained by an iterative process using a training database of one or more training videos. In the iterative process, an “initial” feature predictor neural network is iteratively modified until a termination criterion is satisfied (e.g. a certain number of computing operations have been performed). In each iteration, one or more of a set of numerical parameters which define the feature predictor neural network are modified to reduce a measure of a difference between (i) context data which is an output of the feature predictor neural network based on an individual frame of the training video (e.g. the result of processing the frame using the spatial feature encoder, to form a spatial feature encoding, and then using the spatial feature encoder as the input to the feature predictor neural network), and (ii) a portion of the spatial-temporal encoding corresponding to the individual frame (i.e. the corresponding spatial-temporal encoding item).
[0042]
[0033] The output of the feature predictor neural network may be considered a per-frame prediction of the corresponding spatial-temporal encoding item. Though the feature predictor network cannot predict actions (since it operates on single frames), it is trained to generate a portion of context data specifying the visual appearance of elements depicted in a frame of the action video item (the second actor(s) and / or the environment the second actor(s)). As noted, this is often a component of the context data which is at least partly in common with the context data for other frames of the action video item, and in common with the corresponding spatial-temporal encoding item for the frame. Attorney Docket No.: 45288-0578W01
[0043]
[0034] The generator neural network is also trained using a training database of one or more training videos (“training video items”). It may be trained after the training of the feature predictor neural network. Alternatively, the generator neural network and feature predictor neural network may be trained jointly (i.e. with updates to the generator neural network and feature predictor neural network being either substantially simultaneous or interleaved). The generator neural network (generator model) may be an auto-regressive generator, which at each of a series of time-steps generates a corresponding frame of the video item, conditioned on the context image, the action data (e.g. received from the vision encoder and the appearance bottleneck) and also on one or more of the frames it has previously generated.
[0044]
[0035] In the iterative process, an “initial” generator neural network is iteratively modified until a termination criterion is satisfied (e.g. a certain number of computing operations have been performed). In each iteration, one or more of a set of numerical parameters which define the generator neural network are modified to reduce a measure of a difference between (i) a video item generated by the current generator neural network upon processing a context visual item comprising at least one context image included in one of the training videos and action data defining the one or more actions in the one of the training video, and (ii) the one of the training videos itself.
[0045]
[0036] The action data for the training video may be obtained by the manner explained above. That is, the training video item is processed to extract spatial-temporal feature data encoding the one or more actions and the appearance of the one or more actors. The training video item is also processed using spatial encoder and then the (e g. trained) feature predictor neural network to extract the context data. The action data is a difference between the spatial-temporal feature data and the context data.
[0046]
[0037] The processing of the training video to form the action data may comprise an initial step of performing an augmentation operation (e.g. random cropping, re-scaling etc.) on the training video, before the spatial-temporal feature data and context data are extracted. The augmentation operation may be different in different iterations, but it is selected from augmentation operations which preserve actions in the training video (i.e. the modified training video depicts the same actions as the original). The use of the augmentation operations reduces the risk of the action data being contaminated with any context data (“appearance leakage”), i.e. such that the generator neural network learns to rely on the action data for more than just a definition of actions which are present in the training video.
[0047]
[0038] The subject matter described in this specification can be implemented in particular embodiments so as to realize one or more of the following advantages. Attorney Docket No.: 45288-0578W01
[0048]
[0039] The present system makes it possible to transfer specific information (information about actions) from one video item (the action video item, also called a demonstration video) into a new video item which is being generated. Thus, video items can be generated including desired actions which it is not necessary to define explicitly using control signals, e.g. using textual descriptors, moving keypoints, dense depth or segmentation maps. These control signals are either so abstract that they can give inadequate control of the video item (textual descriptors, moving keypoints) such that the resulting videos are not realistic or do not adhere to the laws of physics, or difficult and expensive to acquire (dense depth or segmentation maps). The present system may generate a video item that integrates the demonstrated action into the context provided by the context image, resulting in both temporal continuity' and physical plausibility. Thus, a realistic physical simulation system can be provided, e.g. for obtaining information about how agents (e.g. robots) should be chosen and / or controlled to operate to perform tasks associated with the actions in the demonstration video.
[0049]
[0040] The term “spatial-temporal’' is used in this document to be equivalent to “spatiotemporal”. For example, a spatial-temporal encoder, a spatial-temporal encoding or spatial-temporal feature data may equivalently be referred to respectively as a spatiotemporal encoder, a spatiotemporal encoding or spatiotemporal feature data.
[0050]
[0041] The details of one or more embodiments of the subject matter of this specification are set forth in the accompanying drawings and the description below. Other features, aspects, and advantages of the subject matter will become apparent from the description, the drawings, and the claims.
[0051] BRIEF DESCRIPTION OF THE DRAWINGS
[0052]
[0042] Figure 1 illustrates a system which is an implementation of the present disclosure.
[0053]
[0043] Figure 2 illustrates the operation of an action data generation unit of the system of Figure 1.
[0054]
[0044] Figure 3 illustrates a method of generating a video item which in an example implementation.
[0055]
[0045] Figure 4 illustrates sub-steps of a step of the method of Figure. 3.
[0056]
[0046] Figure 5 illustrates a method of training a feature predictor neural network in an implementation.
[0057]
[0047] Figure 6 illustrates a method of training a generator neural network in an implementation. Attorney Docket No.: 45288-0578W01
[0058]
[0048] Figure 7 illustrates experimental results from an implementation of the present disclosure.
[0059]
[0049] Like reference numbers and designations in the various drawings indicate like elements.
[0060] DETAILED DESCRIPTION
[0061]
[0050] Most existing approaches to generating controllable video use explicit control signals to specify the content of the video, such as text prompts, moving keypoints, dense depth maps or segmentation maps. The control signals are time-consuming to generate, and the video generation consumes significant video resources. Furthermore, the training of the system typically requires a large training database of training examples which are training video items and associated control signals, and this database too is time-consuming to create.
[0062]
[0051] Figure 1 illustrates a simulation system which is an implementation of the present disclosure, and which does not include explicit control signals to generate a video. Components of the simulation system can be trained without requiring a training database of training examples including associated control signals.
[0063]
[0052] The simulation system employs a context visual data item comprising (or optionally consisting of) at least one context image 11, also denoted 3. For simplicity it is assumed that only a single context image 11 is used, but in variations the context visual data item may comprise multiple context images, e g. arranged in a sequence. The context image may depict one or more “first” actors and an environment in which those actor(s) are located. In the example of Figure 1, the actor is a cat. The context image 11 shows the appearance of the cat, and depicts an environment in which the cat is located. In Figure 1 this is illustrated by the background of stars in the context image 11.
[0064]
[0053] The simulation system further employs an action video item 12, also denoted V which is a sequence of multiple images (“frames”) in which one or more “second” actors (humans, animals, robots, etc.) perform one or more actions. In the case of Figure 1, there is a single actor, with the appearance of a bear, which, as shown in Figure 1 is in an upright posture in a first of the images of the sequence. The actions may include the bear walking forwards. Frames of the action video item 12 are indexed by a variable t. Although only five frames are shown, it is to be understood that they may be any number of frames, such as at least 10 or at least 100.
[0065]
[0054] The action video item 12 is processed by an action data generation unit 13 which comprises one or more vision encoders 14, denoted ℱ, and a unit called an “appearance Attorney Docket No.: 45288-0578W01
[0066] bottleneck”. The action data generation unit generates action data (also called “action latents”) denoted δVwhich encodes the actions performed in the action video item 12.
[0067]
[0055] A trained generator network referred to as a generation model 16, also denoted by Q. processes the context image 11 and the action data by a and generates a video item 17, which may be denoted V. The video item 17 comprises a sequence of images showing any first actor(s) depicted in the context image 11, in the environment depicted in the context image 11. Specifically, in the video item 17 the first actor(s) are depicted carrying out the actions shown in the action video item 12, in the environment shown in the context image 11.
[0068]
[0056] Note that the first frame of the video item 17 may be identical to the context image 11, but later frames of the video item 17 show the first actors performing the actions depicted in the action video item 12. For example, a frame 18 from the video item 17, which is later in the sequence than the first frame of the video item 17 depicts the cat in sitting up in the same attitude as the bear in the action video item 12. In other words, the action video item 12 is a simulation of a process (actions) depicted in the action data item 12, with the second actors replaced by the first actors. Thus, the system of Figure 1 is a simulated video which shows a simulation of the first actors, depicted in the context image 11, performing the actions.
[0069]
[0057] In some implementations, the context image 11 does not depict any actor, or only depicts actors which are the same as the second actors. In either case, the video item 12 may be generated with the actions depicted in the action video item 12 being performed by first actors which are the same as the second actors.
[0070]
[0058] Note that for most realistic results, the actions in the action video item 12 should be compatible with the actors and / or environment shown in the context image 11. For example, if the context image 11 only depicts actors which are incapable of performing the actions depicted in the action video item 12, then the realism of video item 17 may be less.
[0071]
[0059] Turning to Figure 2, the structure of the action data generation unit 13 is depicted. The vision encoders 14 are implemented by a spatial-temporal encoder 21 and a per-frame spatial encoder 22. This follows the structure described in the VideoPrism model (L. Zhao et al., “VideoPrism: A foundational visual encoder for video understanding”, arXiv:2402.13217 2024.)
[0072]
[0060] Specifically, a spatial-temporal encoder 21 generates spatial-temporal encoding 23 z from the action video item 12. The spatial -temporal encoding 23 (visual token) is spatial-temporal feature data which is a temporally-aggregated spatial-temporal (semantic) representation of the action video item 12. The spatial-temporal feature data encodes both Attorney Docket No.: 45288-0578W01
[0073] action(s) depicted in the action video item 12 and also context data indicating the appearance of elements (actor(s) and / or environment) depicted in the action video item 12. z ∈ ℝT×N×Dwhere T is the number of frames in the action video item 12, indexed by a variable t which takes integer values from 0 to T-1, and N and D represent the spatial and feature dimension respectively. N may be H / s x W / w where each frame of the action video item 12 has a pixel array size of H x W. and s and w are factors (e.g., 18) which indicate compression in the height and width directions respectively. In other words, spatial locations in the spatial-temporal feature data correspond to respective s x w pixel patches of the H x W pixel array of each frame of the action video item 12.
[0074]
[0061] Optionally, the spatial-temporal encoder 21 may be implemented using the factorized encoder architecture design from ViViT (A. Amab et al., “ViViT: A video vision transformer’, arXiv:2103.15691,2021) which first computes spatial representations per frame using a spatial encoder and then temporally aggregates each spatial location using a temporal aggregator (temporal aggregation layer) to form the dataset z indicated by 23 in Figure 2.
[0075]
[0062] The per-frame spatial encoder 22 generates, for each frame t. a respective set of feature data (e.g. a two-dimensional array of feature data) referred to as a spatial feature encoding of the frame. The set of spatial feature encodings is denoted by 24 in Figure 2. Optionally, the per-frame spatial encoder 22 may be shared with the spatial encoder (if any) of the spatial-temporal encoder 21.
[0076]
[0063] The appearance bottleneck 15 includes a per-frame feature predictor 25 (feature predictor neural network) denoted P which processes the spatial feature encodings 24 generated by the per-frame spatial encoder 22 individually, to extract, for each frame t, a two-dimensional array of corresponding context data denoted by ht. The set 26 of these T arrays of context data is denoted h. Since the feature predictor 25 operates on individual spatial feature encodings, it extracts minimal action information, since actions are typically represented in the action video item 12 multiple consecutive frames. The per-frame feature predictor 25 aligns h with the spatial-temporal encodings z, in the sense that the per-frame feature predictor 25 generates, independently for each frame using the respective spatial feature encoding, a “besteffort” approximation of the corresponding portion of the temporally-aggregated representation z. The visual appearance of an element (e.g. a second actor and / or environment) of the action video item 12 will thus be represented by similar features in the datasets 23 and 26. Attorney Docket No.: 45288-0578W01
[0077]
[0064] Each set of context data htencodes the appearance in frame t of the second actor(s) and / or the appearance of the environment in which the second actors are located. Typically, htindicates, for a given frame t, content which is also depicted in one or more other frames of the action video item, but the context data h is not identical for each frame, because, for example, different ones of the second actors may be visible in different frames of the action video data item 12, or a second actor may be at different locations in the different frames so that feature data representing the second actor is at a different location in the two-dimensional array of context data, or a given second actor may be facing in different directions in different frames of the action video data item 12, or different parts of the environment may be visible in different frames of the action video data item 12.
[0078]
[0065] The action data (action latents) δVis obtained by subtracting the aligned spatial representations h from the spatial-temporal encoding z, to give the difference between the spatial-temporal feature data and the context data h. That is,
[0079] δv= z - [h0, h1,
[0080] where [■] denotes a concatenation operation.
[0081]
[0066] Turning to Fig. 3, a computer-implemented method 30 which is an implementation of the present disclosure is illustrated schematically. The method 30 can be performed by a system of one or more computers located in one or more locations. For example, the system illustrated in Fig. 1, appropriately programmed, can perform the process 30.
[0082]
[0067] The method 30 is for generating a video item, e.g. video item 17, depicting one or more “first” actors performing actions, such as actions to accomplish a task. The term “actor” is used here to refer to a human, animal or inanimate object (such as a mechanical object, e.g. a robot).
[0083]
[0068] In step 31, an action video item (such as action video item 12) is processed to obtain action data defining one or more actions depicted in the action video item. This step may for example be performed by the action data generation unit 13 of Figures 1 and 2. In the action video item the action is performed by one or more “second” actors, who may be different from the first actors. For example, at least one of the first actors may have a different appearance from at least one of the second actors (e.g. the actor may be a different human individual and / or may be dressed differently).
[0084]
[0069] In step 32, a context visual data item comprising at least one context image (such as context image 11) is processed, together with the action data, to obtain a video item. This may be done by the generation model 16 of Figure 1, for example, to generate the video item 17. Attorney Docket No.: 45288-0578W01
[0085] The context image(s) 11 depict one or more first actors in a certain environment. The video item 17 depicts those same one or more first actors performing, for example in the same environment, the actions depicted in the action video item. The first actors may be different from second actors who perform the actions in the action video item. For example, the first actors may be different human individuals and / or differently dressed. Alternatively or additionally, the first actors may be mechanical system(s) rather than human actors in the action video item, showing how the mechanical system(s) can perform a task performed by human actors in the action video item.
[0086]
[0070] The generation model may for example be a latent diffusion model (LDM), as disclosed in T. Brooks at al., “Video generation models as world simulators”, OpenAI Blog, 2024; A. Gupta et al., “Photorealistic video generation with diffusion models”, in ECCV 2024; or A. Polyak, et al., “Movie Gen: A cast of media foundation models”, arXiv: 2410.13720, 2024. Training of the generation model is described below with reference to Figure 6.
[0087]
[0071] Figure 4 shows sub-steps of step 31 of the method 30.
[0088]
[0072] In sub-step 41. the action video item is processed, such as by the spatial-temporal encoder 21, to extract spatial-temporal feature data, such as the spatial -temporal encoding z.
[0089] The spatial-temporal feature data encodes the visual appearance of elements depicted in the action video item (i.e., the appearance of the second actor(s) and the environment in which they are located), and also encodes the actions the second actors perform in the action video item.
[0090]
[0073] As noted, the spatial-temporal encoder may be implemented as a ViViT model, which first computes respective spatial representations (spatial feature encodings) of each frame of the action video item, and then processes the spatial feature encodings together to temporally aggregate them, thereby forming the spatial-temporal feature data. This is performed by, for each spatial location (e.g., corresponding to a respective pixel of the pixel array of each frame of the action video, or to a respective s x w patch of the pixel array), processing the respective portions of the spatial feature encodings for that location.
[0091]
[0074] In sub-step 42, the action video item is processed to extract context data indicating the visual appearance of elements depicted in the action video item. This may be done frame-by-frame of the action video item, to generate respective portions of the context data separately. That is, for each frame of the action video item, a respective portion of the context data is generated using (only) that frame of the action video item.
[0092]
[0075] In principle, sub-step 42 may be performed by a trained feature predictor network sequentially processing frames of the action video item. However, more efficiently, sub-step Attorney Docket No.: 45288-0578W01
[0093] 42 may be performed using a spatial feature encoder, such as the per-frame spatial encoder 22 of Figure 2. The spatial feature encoder processes each frame t of the action video item separately to generate a respective spatial encoding of the frame, and then the feature predictor neural network (e.g. the per-frame feature predictor 25 of Figure 2) generates the corresponding portion of context data (e.g. ht) for the frame. The respective portions ht of context data for the T frames collectively form the context data h for the action video item.
[0094]
[0076] As noted, the spatial feature encoder used in step 42 may be shared with a spatial feature encoder of the spatial-temporal encoder 21 used in step 41, so that a given frame of the action video item is only processed once to generate a spatial feature encoding which is (i) processed by the feature predictor neural network to generate the portion of the context data for the frame and (ii) processed, together with the spatial feature encodings of the other frames of the action video item, by the temporal aggregator to form the spatial -temporal encoding z.
[0095]
[0077] In sub-step 43, the action data is obtained as a difference between the spatial-temporal feature data (e.g. z) and the context data (e.g. h).
[0096]
[0078] Training of a feature predictor neural network of the action data generation unit is described below with reference to Figure 5, which shows a method 50 in an implementation of the present disclosure. The method 50 can be performed by a system of one or more computers located in one or more locations.
[0097]
[0079] The training uses a training dataset D of training videos, and a pre-trained video encoder, such as the ViViT video encoder model (Amab et al., 2021).
[0098]
[0080] In step 51, an “initial” feature predictor neural network is obtained. The feature predictor neural network is defined by a plurality of numerical parameters, which may initially be set randomly, or to default values, or be the result of a prior training procedure. The feature predictor neural network J5may for example be implemented as a stack of transformer blocks, e.g. of the kind used in the ViViT model.
[0099]
[0081] In step 52, the training videos of the training dataset are processed by the video encoder to generate a respective spatial -temporal encoding (e.g. z) of each training video which comprises a portion (“spatial -temporal encoding item”, e.g. zt) for each respective / -th frame of the training video.
[0100]
[0082] A plurality of training steps 53 are then carried out successively, to successively update the feature predictor neural network. At each step 53 the current state of the feature predictor neural network is referred to as a current feature predictor neural network. That is, in the first training step, the current feature predictor neural network is the initial feature predictor neural Attorney Docket No.: 45288-0578W01
[0101] network, and in each subsequent step, the current feature predictor neural network is the feature predictor neural network as updated in the preceding step 53.
[0102]
[0083] In each training step 53, an output ht of the feature predictor neural network T3is formed based on a Mh frame of one of the training videos (e.g. a randomly chosen frame of a randomly chosen one of the training videos). The output ht of the feature predictor neural network P is not based on other frames of the one training video. For example, the / -th frame of the training video may be processed by a spatial encoder to form a spatial feature encoding of the frame, and the spatial feature encoding may be processed by the feature predictor neural network P to obtain ht. The parameters of the feature predictor neural network are then updated to reduce the value of a difference between zt and ht.
[0103]
[0084] Although step 53 has been explained with reference to a single frame, alternatively updating step 53 may be performed in parallel for a batch of frames of one or more of the training video(s).
[0104]
[0085] Step 53 is repeated until a termination criterion is met. For example, the termination criterion may be that step 53 has been performed a pre-determined number of times, or that in the most recent performance of step 53 the parameters of the feature predictor neural network changed by less than a threshold amount.
[0105]
[0086] Thus, the repeated steps 53 train P by iteratively modify ing the numerical parameters defining P to obtain:
[0106] J’ = arg™«IveD 2T-i||Zt_ ^||i (1)
[0107]
[0108] where zt∈ ℝN×Ddenotes the / -th per-frame slice of z (i.e. the spatial-temporal encoding item ztfor frame t), htdenotes the portion of the context data for the / -th frame of the training video predicted by P, and ||·|| denotes the
[0109]
[0110] norm (Manhattan distance). Intuitively, Eqn. (1) means that P learns to reconstruct temporally-aggregated information as well as it can using only data from individual frames of the training videos. As mentioned, that data may the individual spatial feature encodings produced by the per-frame spatial encoder 22. P uses these spatial feature encodings instead of using raw frames of the action video item 12, which makes its training easier in practice.
[0111]
[0087] Note that optionally steps 52 and 53 may be performed in parallel. For example, in each training step, at least one training video may be selected from the training dataset (e.g. randomly), processed by the video encoder to form the corresponding spatial-temporal encoding, and then used to update the numerical parameters of the feature predictor neural network as explained above. Attorney Docket No.: 45288-0578W01
[0112]
[0088] Turning to Figure 6, a method 60 is shown for training a generator neural network in an implementation of the present disclosure. The generator neural network may be the generator model 16 of Figure 1. The method 60 can be performed by a system of one or more computers located in one or more locations.
[0113]
[0089] The method 60 uses a training dataset 풟 of training videos, which may be the same training dataset used in the method 50. and a trained action data generation unit, which may be the action data generation unit 13 of Figure 1, and may incorporate a feature predictor neural network trained by the method 50 of Figure 5.
[0114]
[0090] In step 61, an “initial’' generator neural network is obtained. The generator neural network is defined by a plurality of numerical parameters, which may initially be set randomly, or to default values, or be the result of a prior training procedure. The feature predictor neural network Cj may for example be implemented using the WALT model of Gupta et al., 2024.
[0115]
[0091] A plurality of training steps 62 are then carried out successively, to successively update the generator neural network. At each step 62 the current state of the generator neural network is referred to as a current generator neural network. That is, in the first training step 62, the current generator neural network is the initial generator neural network, and in each subsequent step 62, the current generator neural network is the generator neural network as updated in the preceding step 62.
[0116]
[0092] In each step 62, a training video is selected from the training dataset 풟, and T+1 frames v0:Tare selected from the selected training video, where T is an integer greater than one and the T+l frames vtare labelled by an index t which runs from 0 to T.
[0117]
[0093] The set of T frames v1:Tis a training video item V. Random augmentations are optionally performed on the T frames v1:T(i.e. the selected frames except the first one v0). The augmentations are ones which do not change the actions depicted by the T frames. For example, they may include random spatial cropping of the frames. The T resulting frames are denoted
[0118]
[0094] The T frames v'1:Tare processed by the action data generation unit to generate action data δV'.
[0119]
[0095] The generator neural network Q then processes the first frame v0as a context image and the action data δV'to generate a target video V which is Q(v0, δV') and which is composed of T frames corresponding to the T frames of the training video item V.
[0120]
[0096] An update is then made to the parameters of the generator neural network Q to reduce a measure ℒ(Q(v0, δV'), v1:T) of the difference between the target video V and the training Attorney Docket No.: 45288-0578W01
[0121] video item V. Here L is a denoising loss function which is a distance measure. In principle it may be defined as a sum, over the pixels of the T frames of the target video, of a Manhattan distance or Euclidean distance between an intensity value of the pixel in the frame of the target video V and the corresponding pixel of the corresponding frame the training video item V. However, in the experiments reported below, L is a velocity loss function, following the practice of Gupta et al., 2024 and T. Salimans et al., ''Progressive distillation for fast sampling of diffusion models”, in ICLR, 2022. A tokenizer is used to compress both v0and v1:Tinto a low dimensional latent space.
[0122]
[0097] The result of the optional augmentation is to cause misalignment and disparity between the training video item V and the target video V, making it harder for the generator neural network to copy the appearance information directly from the training video item V.
[0123]
[0098] Step 62 is repeated until a termination criterion is met. For example, the termination criterion may be that step 62 has been performed a pre-determined number of times, or that in the most recent performance of step 62 the parameters of the generator neural network changed by less than a threshold amount.
[0124]
[0099] Thus, the repeated steps 62 train Q by iteratively modifying the numerical parameters defining Q to obtain:
[0125] argmin
[0126] Q — g £(S(VO>8V')’V1: T)- (2)
[0127]
[0128]
[0100] Figure 7 shows the experimental results from an implementation of the techniques disclosed herein. The experiments used T=16 frames, with each frame of the action video item being a pixel array of 288x288 pixels. The spatial-temporal encoder was a VideoPrism encoder from L. Zhao et al., 2024, using 0.1B parameters.
[0129]
[0101] The per-frame feature predictor P was constructed from 4 transformer encoder blocks from ViViT (Amab et al., 2021) with 1024 hidden dimensions and 8 heads for multi-head attention. The per-frame feature predictor received as input the VideoPrism spatial encoder output for a particular frame. During the training of the per-frame feature predictor using method 50, an LI loss was used and the training involved 30k iterations with the Adam optimizer (D. P. Kingma, “Adam: A method for stochastic optimization”, in ICLR 2015) with 10’4base learning rate and 4.5×10-2weight decay.
[0130]
[0102] The generator model Q was the WALT model (Gupta et al., 2024) with 0.3B parameters. The generator model Q was trained using method 60 to generate videos of T=16 frames with 128x128 spatial resolution. The context image was natively passed to generator model as the context image for conditional generation. The latents δVwere projected into a sequence of Attorney Docket No.: 45288-0578W01
[0131] 20480-dimensional latent vectors in place of the text embeddings used in Gupta et al., 2024. The generator model was initialized with a pre-trained I+T2V checkpoint, and fine-tuned on downstream datasets in 500k iterations. In operation, a classifier-free guidance scale of 1.25 was used (J. Ho et al., “Classifier-free diffusion guidance, arXiv: 2207.12598, 2022).
[0132]
[0103] The experiments were conducted on three datasets, namely Epic Kitchens (D. Damen et al., “Scaling egocentric vision: the EPIC-KITECHENS dataset" in ECCV, 2018), Something-Something v2 (SSv2) (R. Goyal et al., “The 'something something’ video database for learning and evaluating visual common sense”, 2017) and Fractal (A Brohan et al., “RT-1: Robotics transformer for real-world control at scale’’, arXiv: 2212.06817, 2022).
[0133]
[0104] Figure 7(a) shows frames from an example action video item, showing a robot arm moving an object towards the left. Figure 7(b) shows an example context image showing a tabletop (environment) with an object. Figure 7(c) shows frames from the generated video item. This begins with a frame which is the same as the context image, and then shows a robot arm moving the object. Since in this case the context image did not include an actor, the video item has been generated with the first actor being identical to the actor (the robot arm) of the action video item. The visual appearance of the environment in the video item is according to the context image.
[0134]
[0105] Table 1 shows experimental results of human evaluation preferences comparing the present method to four baseline methods WALT (Gupta et al., 2024), MotionDirector (R. Zhao et al., “MotionDirector: Motion Customization of text-to-video diffusion models”, 2024), CogVideoX (Z. Yang et al, “CogvideoX: Text-to-video diffusion models with an expert transformer”, 2024, arXiv:2408.06072) and DynamicCrafter (Jinbo Xing, et al., “DynamiCrafter: Animating open-domain images with diffusion priors”, in European Conference on Computer Vision, pp 399-417, 2024). WALT (the “large” version of the model, with 0.3B parameters) was fine-tuned on each dataset individually with ground truth captions. CogVideoX (5B parameters) and DynamiCrafter (1.6B parameters) were used “out of the box” (i.e. without fine-tuning) for a system-level comparison. MotionDirector (1.8B parameters) used image, text and video as conditional signals. When generating videos, MotionDirector learns per-instance appearance and motion from an initial frame and demonstration via LoRA (low-rank adaptation) fine tuning, and generates videos based on a given text prompt. For all baselines, the groundtruth caption was used as the prompt during video generation.
[0135]
[0106] Ten human raters were asked to evaluate performance in terms of three rubrics (criteria): visual quality (VQ), action transferability between demonstration and generated videos (AT) and context image constancy (CC). 20 video examples were curated from each Attorney Docket No.: 45288-0578W01
[0136] dataset, totaling 60 examples for the evaluation set. For each example, an initial frame, a demonstration video and two generated videos (one generated using an implementation of the present disclosure (referred to as “ours” below), and one generated using one of the four baseline techniques) are provided to the rater, and the rater was asked to choose the better video from the two generated videos based on each of the three rubrics. The preference rates shown in Table 1 were then calculated, where each number shows the proportion of the raters (on a scale from 0 to 1) who preferred the implementation of the present disclosure, compared to the baseline under each dataset and rubric. The present method is clearly favored by the human raters across all the datasets, baselines and rubrics.
[0137]
[0138] Table 1
[0139]
[0107] Simulations generated by the simulation system described in this specification can be used for any of a variety of technical purposes.
[0140]
[0108] In some cases, an agent (e.g., a reinforcement learning agent) interacting with a physical environment may use the simulation system to generate one or more simulations of the environment that simulate the effects of the agent performing various actions in the environment. In these cases, the agent may use the simulations of the environment as part of determining whether to perform certain actions in the environment.
[0141]
[0109] In some cases, the system can use the simulated states of the environment to control an agent interacting with the environment. For example the system can simulate states of the environment when the agent performs a sequence of actions over a sequence of time, determining that a final state of the environment satisfies a success criterion, and then cause the agent to perform the sequence of actions in the environment when the success criterion is satisfied. The system can determine the success criterion based on whether respective locations, shapes, or configurations of the objects within the environment are within certain tolerances for the objects.
[0142]
[0110] The system can simulate the state of the environment based on data received from sensors of the agent. The agent can be any of a variety of agents interacting with the environment. For example, the agent can be a robot manipulating objects in the environment. Attorney Docket No.: 45288-0578W01
[0143] As another example, the agent can be an autonomous vehicle navigating through the environment.
[0144] [Hl] For example, the system can use the simulated states of the environment for real-world control of an agent or device, in particular for control tasks, e.g. to assist a robot in manipulating a deformable object. Thus, the physical environment may comprise a real-world environment including a physical object e.g. an object to be picked up or manipulated. Determining the state of the physical environment at the next time step may comprise determining a representation of a shape or configuration of the physical object e.g. by capturing an image of the object. Determining the state of the physical environment at the next time step may comprise determining a predicted representation of the shape or configuration of the physical object e.g. when subject to a force or deformation e.g. from an actuator of a robot. The method may further comprise controlling the robot using the predicted representation to manipulate the physical object, e.g. using an actuator, towards a target location, shape or configuration of the physical object by controlling the robot to optimize an objective function dependent upon a difference between the predicted representation and the target location, shape or configuration of the physical object. Controlling the robot may involve providing control signals to the robot based on the predicted representation to cause the robot to perform actions, e g. using an actuator of the robot, to manipulate the physical object to perform a task. For example, this may involve controlling the robot, e.g. the actuator, using a reinforcement learning process with a reward that is at least partly based on a value of the objective function, to leam to perform a task which involves manipulating the physical object.
[0145]
[0112] In some cases, the video item may be processed to determine that a feasibility criterion is satisfied, and a physical apparatus (e.g. article) or system may be designed and / or constructed in response to the feasibility criterion being satisfied. For example, as previously described, the context image may comprise data representing a shape or state of an object (e.g. the article), and determining the video item may comprise determining a representation of the shape or state of the object at subsequent times. The feasibility criterion may be that the video item realistically shows the object performing a certain role, e.g. moving along a certain path, or being manipulated to perform a certain task. That is, determining whether the feasibility criterion is satisfied may include a process (e.g. carried out by a human) of determining whether the video item appears realistic, e.g. does not include physically impossible actions. If the video item does not appear realistic, this may indicate that the object is not suitable for the role (i.e. the feasibility criterion is determined not to have been met). Attorney Docket No.: 45288-0578W01
[0146] [H3] In some cases, the method may comprise obtaining a respective value for one or more design parameters; and generating the context image using the respective values of the design parameters. In this case, designing the article in response to the feasibility criterion being satisfied comprises designing the article based on the respective values of the one or more design parameters.
[0147]
[0114] In some cases, the above described systems and methods may be used for design optimization. For example as previously described, the context image defining the state of the physical environment may comprise design parameters for an article. A method of designing the article may then comprise adjusting the design parameters according to one or more design criteria for the object, e.g. to minimize stress in the object when subject to a force or deformation e.g. by including a representation of the force or deformation in the data defining the state of the physical environment. The process may include making (i.e. fabricating in the real-world) a physical object with the optimized design parameters. The physical object may be e.g. for part of a mechanical structure.
[0148]
[0115] For example, if the design parameters represent a shape or structure of a physical object (e.g., an aircraft wing), then the design parameters can be provided for use in manufacturing an object having the design defined by the design parameters. The object can be manufactured using any appropriate manufacturing process, e.g., a machining process or an additive manufacturing process. In particular, the system can implement an appropriate manufacturing process to manufacture an object having the design defined by the design parameters. As another example, if the design parameters define the design of a process, e.g. a chemical process or a mechanical process, then the design parameters can be provided for use in implementing a process having the design defined by the design parameters. In particular, the system can implement (i.e. in the real-world) a process having the design defined by the design parameters. When the design parameters define the shape or configuration of physical object, the method can include making a physical obj ect to a design specified by the design parameters.
[0149] [H6] In some cases, the design parameters can define, e.g., a shape of an object, e.g., all or part of a vehicle, e.g.. a car, a truck, an aircraft, a watercraft, a rocket, etc. In particular examples, the design parameters can define the shape of a wing of an aircraft or the shape of a hull of a watercraft. The design parameters can define the shape of an object, e.g., by defining a respective position of each control point in a set of control points that parametrize the shape of the object, or by defining the vertices and edges of a mesh representing the shape of the object. The system may simulate, e.g., fluid (e.g.. air) dynamics in an environment. For example, the system can simulate a stress field or a pressure field in an environment, e.g., that defines a Attorney Docket No.: 45288-0578W01
[0150] respective stress or pressure at each position in a grid or mesh spanning the environment. A feasibility criterion may be a measure of one or more aerodynamic features of the object (e.g., a drag coefficient or a lift coefficient of the object), or a measure of physical stress or force exerted on the object under specified environment conditions (e.g., the maximum stress exerted on any part of the object). A design criterion may be a measure of one or more aerodynamic features of the object (e.g., a drag coefficient or a lift coefficient of the object), or a measure of physical stress or force exerted on the object under specified environment conditions (e.g., the maximum stress exerted on any part of the object).
[0151] [H7] In some cases, the design parameters can define, e.g., a structure of an object, e.g., of a vehicle, a bridge, or a building. In particular examples, the design parameters can define the structure of the chassis or frame of a vehicle, or the structure of supports within a bridge or building. The design parameters can define the structure of an object, e.g., by representing the positions, orientations, thicknesses, and connectivity of rods, beams, struts, and ties defining the structure of the object. The system may simulate, e.g., structural mechanics in an environment. For example, the system may simulate a force, stress, or pressure field, e g., that defines a respective force, stress, or pressure at each position in a grid or mesh spanning the structure. A feasibility criterion may represent e g., the behavior of the structure under a mechanical load, e.g., a maximum force, stress, or pressure on any part of the structure under the mechanical load. A design criterion may represent e.g., the behavior of the structure under a mechanical load, e.g., a maximum force, stress, or pressure on any part of the structure under the mechanical load.
[0152]
[0118] In some cases, the design parameters can define, e.g., a composition of a material, e.g., an alloy. In particular examples, the design parameters can define the composition of a material, e.g., by defining, for each of multiple possible constituent materials, a fraction of the material that is represented by the constituent material. The system may simulate, e.g.: changes in the chemical composition of the material over time resulting from specified environmental conditions; or a force, stress, or pressure field representing force, stress, or pressure at each position in a grid or mesh spanning an object made of the material. A feasibility criterion may characterize, e.g., corrosion of the material over time, or behavior of an object made of the material under a mechanical load. A design criterion may characterize, e g., corrosion of the material over time, or behavior of an object made of the material under a mechanical load.
[0153]
[0119] In some cases, the design parameters can define a design of a chemical process, e.g., defining when and how various chemicals should be combined in a chemical process. For example, the design parameters can define the speed of a mixer that agitates the contents of a Attorney Docket No.: 45288-0578W01
[0154] vat. and for each chemical in a set of chemicals, when the chemical should be added to the vat and in what amount. The system may simulate, e.g., chemical dynamics within an environment. For example, the simulation neural network can simulate a concentration field in an environment, e.g., that defines a respective concentration of each of one or more chemicals at each position in a grid or mesh spanning the environment. A feasibility criterion may measure, e.g.: a yield of the chemical process, e.g., an amount of a desired end product that is produced as a result of the chemical process; or a quality (e.g., purity) of the end product. A design criterion may measure, e.g.: a yield of the chemical process, e.g., an amount of a desired end product that is produced as a result of the chemical process; or a quality (e.g., purity) of the end product.
[0155]
[0120] In some cases, the design parameters can define a design of a mechanical process, e.g., defining, for each fan in an environment (e.g., a mine): (i) a rotational speed of the blades of the fan, and (ii) an orientation of the fan. The system may simulate a flow field in the environment, e.g., that defines a respective direction of airflow, strength of airflow, and concentration of gases at each position in a grid or mesh spanning the environment. A feasibility critenon may characterize, e.g., a distribution and concentration of one or gases (e.g., oxygen) in the environment, e.g., as a result of the operation of the fans. A design criterion may characterize, e.g., a distribution and concentration of one or gases (e.g., oxygen) in the environment, e.g., as a result of the operation of the fans.
[0156] [1211 Certain novel aspects of the subj ect matter of this specification are set forth in the claims below.
[0157]
[0122] In this specification, the term "configured" is used in relation to computing systems and environments, as well as computer program components. A computing system or environment is considered "configured" to perform specific operations or actions when it possesses the necessary software, firmware, hardware, or a combination thereof, enabling it to carry out those operations or actions during operation. For instance, configuring a system might involve installing a software library with specific algorithms, updating firmware with new instructions for handling data, or adding a hardware component for enhanced processing capabilities. Similarly, one or more computer programs are "configured" to perform particular operations or actions when they contain instructions that, upon execution by a computing device or hardware, cause the device to perform those intended operations or actions.
[0158]
[0123] The embodiments and functional operations described in this specification can be implemented in various forms, including digital electronic circuitry, software, firmware, computer hardware (encompassing the disclosed structures and their structural equivalents), or Attorney Docket No.: 45288-0578W01
[0159] any combination thereof. The subject matter can be realized as one or more computer programs, essentially modules of computer program instructions encoded on a tangible non-transitory storage medium for execution by or to control the operation of a computing device or hardware. The storage medium can be a storage device such as a hard drive or solid-state drive (SSD), a storage medium, a random or serial access memory device, or a combination of these. Additionally or alternatively, the program instructions can be encoded on a transmitted signal, such as a machine-generated electrical, optical, or electromagnetic signal, designed to carry information for transmission to a receiving device or system for execution by a computing device or hardware. Furthermore, implementations may leverage emerging technologies like quantum computing or neuromorphic computing for specific applications, and may be deployed in distributed or cloud-based environments where components reside on different machines or within a cloud infrastructure.
[0160]
[0124] The term "computing device or hardware" refers to the physical components involved in data processing and encompasses all types of devices and machines used for this purpose. Examples include processors or processing units, computers, multiple processors or computers working together, graphics processing units (GPUs), tensor processing units (TPUs), and specialized processing hardware such as field-programmable gate arrays (FPGAs) or application-specific integrated circuits (ASICs). In addition to hardware, a computing device or hardware may also include code that creates an execution environment for computer programs. This code can take the form of processor firmware, a protocol stack, a database management system, an operating system, or a combination of these elements. Embodiments may particularly benefit from utilizing the parallel processing capabilities of GPUs, in a General-Purpose computing on Graphics Processing Units (GPGPU) context, where code specifically designed for GPU execution, often called kernels or shaders, is employed. Similarly, TPUs excel at running optimized tensor operations crucial for many machine learning algorithms. By leveraging these accelerators and their specialized programming models, the system can achieve significant speedups and efficiency gains for tasks involving artificial intelligence and machine learning, particularly in areas such as computer vision, natural language processing, and robotics.
[0161]
[0125] A computer program, also referred to as software, an application, a module, a script, code, or simply a program, can be written in any programming language, including compiled or interpreted languages, and declarative or procedural languages. It can be deployed in various forms, such as a standalone program, a module, a component, a subroutine, or any other unit suitable for use within a computing environment. A program may or may not correspond to a Attorney Docket No.: 45288-0578W01
[0162] single file in a file system and can be stored in various ways. This includes being embedded within a file containing other programs or data (e.g., scripts within a markup language document), residing in a dedicated file, or distributed across multiple coordinated files (e.g., files storing modules, subprograms, or code segments). A computer program can be executed on a single computer or across multiple computers, whether located at a single site or distributed across multiple sites and interconnected through a data communication network. The specific implementation of the computer programs may involve a combination of traditional programming languages and specialized languages or libraries designed for GPGPU programming or TPU utilization, depending on the chosen hardware platform and desired performance characteristics.
[0163]
[0126] In this specification, the term "engine" broadly refers to a software-based system, subsystem, or process designed to perform one or more specific functions. An engine is typically implemented as one or more software modules or components installed on one or more computers, which can be located at a single site or distributed across multiple locations. In some instances, one or more dedicated computers may be used for a particular engine, while in other cases, multiple engines may operate concurrently on the same one or more computers. Examples of engine functions within the context of Al and machine learning could include data pre-processing and cleaning, feature engineering and extraction, model training and optimization, inference and prediction generation, and post-processing of results. The specific design and implementation of engines will depend on the overall architecture and the distribution of computational tasks across various hardware components, including CPUs, GPUs, TPUs, and other specialized processors.
[0164]
[0127] The processes and logic flows described in this specification can be executed by one or more programmable computers running one or more computer programs to perform functions by operating on input data and generating output. Additionally, graphics processing units (GPUs) and tensor processing units (TPUs) can be utilized to enable concurrent execution of aspects of these processes and logic flows, significantly accelerating performance. This approach offers significant advantages for computationally intensive tasks often found in Al and machine learning applications, such as matrix multiplications, convolutions, and other operations that exhibit a high degree of parallelism. By leveraging the parallel processing capabilities of GPUs and TPUs, significant speedups and efficiency gains compared to relying solely on CPUs can be achieved. Alternatively or in combination with programmable computers and specialized processors, these processes and logic flows can also be implemented using specialized processing hardware, such as field-programmable gate arrays (FPGAs) or Attorney Docket No.: 45288-0578W01
[0165] application-specific integrated circuits (ASICs), for even greater performance or energy efficiency in specific use cases.
[0166]
[0128] Computers capable of executing a computer program can be based on general -purpose microprocessors, special-purpose microprocessors, or a combination of both. They can also utilize any other type of central processing unit (CPU). Additionally, graphics processing units (GPUs), tensor processing units (TPUs), and other machine learning accelerators can be employed to enhance performance, particularly for tasks involving artificial intelligence and machine learning. These accelerators often work in conjunction with CPUs, handling specialized computations while the CPU manages overall system operations and other tasks. Typically, a CPU receives instructions and data from read-only memory (ROM), random access memory (RAM), or both. The essential elements of a computer include a CPU for executing instructions and one or more memory devices for storing instructions and data. The specific configuration of processing units and memory will depend on factors like the complexity of the Al model, the volume of data being processed, and the desired performance and latency requirements. Embodiments can be implemented on a wide range of computing platforms, from small embedded devices with limited resources to large-scale data center systems with high-performance computing capabilities. The system may include storage devices like hard drives, SSDs, or flash memory for persistent data storage.
[0167]
[0129] Computer-readable media suitable for storing computer program instructions and data encompass all forms of non-volatile memory, media, and memory devices. Examples include semiconductor memory devices such as read-only memory (ROM), solid-state drives (SSDs), and flash memory devices; hard disk drives (HDDs); optical media; and optical discs such as CDs, DVDs, and Blu-ray discs. The specific type of computer-readable media used will depend on factors such as the size of the data, access speed requirements, cost considerations, and the desired level of portability or permanence.
[0168]
[0130] To facilitate user interaction, embodiments of the subject matter described in this specification can be implemented on a computing device equipped with a display device, such as a liquid crystal display (LCD) or an organic light-emitting diode (OLED) display, for presenting information to the user. Input can be provided by the user through various means, including a keyboard,, touchscreens, voice commands, gesture recognition, or other input modalities depending on the specific device and application. Additional input methods can include acoustic, speech, or tactile input, while feedback to the user can take the form of visual, auditory, or tactile feedback. Furthermore, computers can interact with users by exchanging documents with a user's device or application. This can involve sending web content or data in Attorney Docket No.: 45288-0578W01
[0169] response to requests or sending and receiving text messages or other forms of messages through mobile devices or messaging platforms. The selection of input and output modalities will depend on the specific application and the desired form of user interaction.
[0170]
[0131] Machine learning models can be implemented and deployed using machine learning frameworks, such as TensorFlow or JAX. These frameworks offer comprehensive tools and libraries that facilitate the development, training, and deployment of machine learning models.
[0171]
[0132] Embodiments of the subject matter described in this specification can be implemented within a computing system comprising one or more components, depending on the specific application and requirements. These may include a back-end component, such as a back-end server or cloud-based infrastructure; an optional middleware component, such as a middleware server or application programming interface (API), to facilitate communication and data exchange; and a front-end component, such as a client device with a user interface, a web browser, or an app, through which a user can interact with the implemented subject matter. For instance, the described functionality could be implemented solely on a client device (e.g., for on-device machine learning) or deployed as a combination of front-end and back-end components for more complex applications. These components, when present, can be interconnected using any form or medium of digital data communication, such as a communication network like a local area network (LAN) or a wide area network (WAN) including the Internet. The specific system architecture and choice of components will depend on factors such as the scale of the application, the need for real-time processing, data security requirements, and the desired user experience.
[0172]
[0133] The computing system can include clients and servers that may be geographically separated and interact through a communication network. The specific type of network, such as a local area network (LAN), a wide area network (WAN), or the Internet, will depend on the reach and scale of the application. The client-server relationship is established through computer programs running on the respective computers and designed to communicate with each other using appropriate protocols. These protocols may include HTTP, TCP / IP, or other specialized protocols depending on the nature of the data being exchanged and the security requirements of the system. In certain embodiments, a server transmits data or instructions to a user's device, such as a computer, smartphone, or tablet, acting as a client. The client device can then process the received information, display results to the user, and potentially send data or feedback back to the server for further processing or storage. This allows for dynamic interactions between the user and the system, enabling a wide range of applications and functionalities. Attorney Docket No.: 45288-0578W01
[0173]
[0134] While this specification contains many specific implementation details, these should not be construed as limitations on the scope of any invention or on the scope of what may be claimed, but rather as descriptions of features that may be specific to particular embodiments of particular inventions. Certain features that are described in this specification in the context of separate embodiments can also be implemented in combination in a single embodiment. Conversely, various features that are described in the context of a single embodiment can also be implemented in multiple embodiments separately or in any suitable subcombination. Moreover, although features may be described above as acting in certain combinations and even initially be claimed as such, one or more features from a claimed combination can in some cases be excised from the combination, and the claimed combination may be directed to a subcombination or variation of a subcombination.
[0174]
[0135] Similarly, while operations are depicted in the drawings and recited in the claims in a particular order, this should not be understood as requiring that such operations be performed in the particular order shown or in sequential order, or that all illustrated operations be performed, to achieve desirable results. In certain circumstances, multitasking and parallel processing may be advantageous. Moreover, the separation of vanous system modules and components in the embodiments described above should not be understood as requiring such separation in all embodiments, and it should be understood that the described program components and systems can generally be integrated together in a single software product or packaged into multiple software products.
[0175]
[0136] This specification also provides the subject-matter of the following numbered clauses:
[0176]
[0137] 1. A computer-implemented method of generating a video item depicting one or more first actors performing one or more actions, based on a context visual data item and an action video item depicting one or more second actors performing the one or more actions. the method comprising:
[0177] processing the action video item to obtain action data defining the one or more actions; and
[0178] processing the context visual data item and the action data by a trained generator neural network, to obtain the video item;
[0179] wherein the context visual data item comprises at least one context image defining a visual aspect of the video item.
[0180]
[0138] 2. The method of clause 1 in which the one or more first actors have a different corresponding appearance from the one or more second actors, and the at least one context image defines the appearance of at least one of the first actors. Attorney Docket No.: 45288-0578W01
[0181]
[0139] 3. The method of clause 1 or clause 2 in which the context visual data item defines the appearance of an environment, and the video item depicts the one or more first actors performing the one or more actions in the at least one environment.
[0182]
[0140] 4. The method of any preceding clause in which processing the action video item to extract the action data defining the one or more actions comprises:
[0183] processing the action video item to extract spatial-temporal feature data encoding the one or more actions and the visual appearance of elements depicted in the action video item;
[0184] processing the action video item using a trained feature predictor neural network to extract context data indicating the visual appearance of elements depicted in the action video item; and
[0185] obtaining the action data as a difference between the spatial-temporal feature data and the context data.
[0186]
[0141] 5. The method of clause 4 in which the processing the action video item using a trained feature predictor neural network is performed on each frame of the action video item separately, to obtain respective portions of the context data for each frame.
[0187]
[0142] 6. The method of clause 5 in which processing the action video item using atrained feature predictor neural network comprises encoding frames of the action video item separately by a spatial encoder to form respective spatial feature encodings, and then processing the spatial feature encodings separately to extract the context data.
[0188] |143| 7. The method of any of clauses 4 to 6. in which processing the action video item to extract spatial-temporal feature data comprises encoding frames of the action video item separately by a spatial encoder to form respective spatial feature encodings, and temporally aggregating the spatial feature encodings to form the spatial-temporal feature data.
[0189]
[0144] 8. The method of any preceding clause in which at least one of the one or more first actors is a robot.
[0190]
[0145] 9. The method of clause 7 further comprising using the video item to control the robot to emulate the one or more actions depicted in the video item.
[0191]
[0146] 10. The method of any preceding clause in which the context image is an image of a real-world environment captured by a camera.
[0192]
[0147] 11. The method of any preceding clause in which the action video item is a video of a real-w orld environment captured by a video camera.
[0193]
[0148] 12. The method of any preceding clause in which the one or more actions are actions to perform a task, the first video item depicting how to perform the task in a context defined by the context visual data item. Attorney Docket No.: 45288-0578W01
[0194]
[0149] 13. A method according to any preceding clause in which the video item comprises the context image as a frame, the video item depicting at least one of:
[0195] a video sequence continuing the context image, and
[0196] a video sequence leading up to the context image.
[0197]
[0150] 14. A method of obtaining a feature predictor neural network operative to process a video item in which one or more actors perform actions, to extract context data indicating the visual appearance of elements depicted in the video item, the method comprising:
[0198] obtaining an initial feature predictor neural network;
[0199] for at least one training video, processing frames of the training video together using a video encoder to generate a spatial-temporal encoding of the training video comprising a spatial-temporal encoding item for each respective frame of the training video; and
[0200] a plurality of steps of:
[0201] modifying a current feature predictor neural network to form a corresponding modified feature predictor neural network, the modifying reducing a measure of a difference between an output of the feature predictor neural network based on an individual frame of the training video, and the spatial-temporal encoding item of the individual frame,
[0202] wherein in the first of the steps the current feature predictor neural network is the initial feature predictor neural network, and in each subsequent step the current feature predictor neural network is the modified feature predictor neural network formed in the preceding step.
[0203]
[0151] 15. A method according to clause 14 in which the processing frames of the training video together with the video encoder to generate the spatial-temporal encoding of the training video, is performed by:
[0204] processing the frames of the frames of the training video individually with a spatial feature encoder to form respective spatial feature encodings, and
[0205] aggregating the spatial feature encodings with a temporal aggregator, to generate for each frame of the training video a respective spatial-temporal encoding item.
[0206]
[0152] 16. A method according to clause 14 or clause 15, in which the output of the feature predictor neural network based on an individual frame of the training video is formed by: processing the individual frame with a spatial feature encoder to form a spatial feature encoding; and
[0207] processing the spatial feature encoding with the feature predictor neural network to obtain the output of the feature predictor neural network. Attorney Docket No.: 45288-0578W01
[0208]
[0153] 17. A method of training a generator neural network configured, upon receiving a context visual item comprising a context image, and action data defining one or more actions, to generate a video item depicting one or more actors performing the one or more actions, a visual aspect of the video item being defined by the context image, the method comprising: obtaining an initial generator neural network configured to process a network input to generate a video item;
[0209] a plurality of steps of modifying a current generator neural network to reduce a measure of a difference between:
[0210] a training video item depicting one or more actors performing one or more actions, and
[0211] a video item generated by the current generator neural network upon processing a context visual item comprising at least one context image included in the training video item and action data defining the one or more actions depicted in the training video item.
[0212]
[0154] 18. The method of clause 17 in which the context image defines the appearance of one or more actors depicted in the video item generated by the cunent generator neural network.
[0213]
[0155] 19. The method of clause 17 or clause 18 in which the context image defines the appearance of an environment depicted in the video item generated by the current generator neural network.
[0214]
[0156] 20. The method of any one of clauses 17 to 19 in which the action data is obtained by processing at least part of the training video item.
[0215]
[0157] 21. The method of clause 20 in which the action data is obtained by processing the at least part of the training video item to extract spatial-temporal feature data encoding the one or more actions and the visual appearance of elements depicted in the training video item;
[0216] processing the at least part of the training video item using a trained feature predictor neural network to extract context data indicating the visual appearance of elements depicted in the training video item; and
[0217] obtaining the action data as a difference between the spatial-temporal feature data and the context data.
[0218]
[0158] 22. The method of clause 21 in which the processing of the training video item to extract context data is performed by a feature predictor network obtained by the method of any of clauses 13 to 16.
[0219]
[0159] 23. The method of clause 21 or clause 22 in which the processing of the training video item comprises performing a selected augmentation operation on the training video item Attorney Docket No.: 45288-0578W01
[0220] to form an augmented video item, the feature data and the context data being extracted from the augmented video item.
[0221]
[0160] 24. A system comprising one or more computers and one or more storage devices storing instructions that when executed by the one or more computers cause the one or more computers to perform the operations of the method of any one of clauses 1-23.
[0222]
[0161] 25. One or more computer storage media storing instructions that when executed by one or more computers cause the one or more computers to perform the operations of the method of any one of clauses 1-23.
[0223]
[0162] Particular embodiments of the subject matter have been described. Other embodiments are within the scope of the following claims. For example, the actions recited in the claims can be performed in a different order and still achieve desirable results. As one example, the processes depicted in the accompanying figures do not necessarily require the particular order shown, or sequential order, to achieve desirable results. In some cases, multitasking and parallel processing may be advantageous.
[0224]
[0163] What is claimed is:
Claims
Attorney Docket No.: 45288-0578W01CLAIMS1. A computer-implemented method of generating a video item depicting one or more first actors performing one or more actions, based on a context visual data item and an action video item depicting one or more second actors performing the one or more actions,the method comprising:processing the action video item to obtain action data defining the one or more actions; andprocessing the context visual data item and the action data by a trained generator neural network, to obtain the video item;wherein the context visual data item comprises at least one context image defining a visual aspect of the video item.
2. The method of claim 1 in which the one or more first actors have a different corresponding appearance from the one or more second actors, and the at least one context image defines the appearance of at least one of the first actors.
3. The method of claim 1 or claim 2 in which the context visual data item defines the appearance of an environment, and the video item depicts the one or more first actors performing the one or more actions in the at least one environment.
4. The method of any preceding claim in which processing the action video item to extract the action data defining the one or more actions comprises:processing the action video item to extract spatial-temporal feature data encoding the one or more actions and the visual appearance of elements depicted in the action video item;processing the action video item using a trained feature predictor neural network to extract context data indicating the visual appearance of elements depicted in the action video item; andobtaining the action data as a difference between the spatial-temporal feature data and the context data.
5. The method of claim 4 in which the processing the action video item using a trained feature predictor neural network is performed on each frame of the action video item separately, to obtain respective portions of the context data for each frame.Attorney Docket No.: 45288-0578W016. The method of claim 5 in which processing the action video item using a trained feature predictor neural network comprises encoding frames of the action video item separately by a spatial encoder to form respective spatial feature encodings, and then processing the spatial feature encodings separately to extract the context data.
7. The method of any of claims 4 to 6, in which processing the action video item to extract spatial-temporal feature data comprises encoding frames of the action video item separately by a spatial encoder to form respective spatial feature encodings, and temporally aggregating the spatial feature encodings to form the spatial-temporal feature data.
8. The method of any preceding claim in which at least one of the one or more first actors is a robot.
9. The method of claim 7 further comprising using the video item to control the robot to emulate the one or more actions depicted in the video item.
10. The method of any preceding claim in which the context image is an image of a real-world environment captured by a camera.
11. The method of any preceding claim in which the action video item is a video of a real-world environment captured by a video camera.
12. The method of any preceding claim in which the one or more actions are actions to perform a task, the first video item depicting how to perform the task in a context defined by the context visual data item.
13. A method according to any preceding claim in which the video item comprises the context image as a frame, the video item depicting at least one of:a video sequence continuing the context image, anda video sequence leading up to the context image.Attorney Docket No.: 45288-0578W0114. A method of obtaining a feature predictor neural network operative to process a video item in which one or more actors perform actions, to extract context data indicating the visual appearance of elements depicted in the video item, the method comprising:obtaining an initial feature predictor neural network;for at least one training video, processing frames of the training video together using a video encoder to generate a spatial-temporal encoding of the training video comprising a spatial-temporal encoding item for each respective frame of the training video; anda plurality of steps of:modifying a current feature predictor neural network to form a corresponding modified feature predictor neural network, the modifying reducing a measure of a difference between an output of the feature predictor neural network based on an individual frame of the training video, and the spatial-temporal encoding item of the individual frame, wherein in the first of the steps the current feature predictor neural network is the initial feature predictor neural network, and in each subsequent step the current feature predictor neural network is the modified feature predictor neural network formed in the preceding step.
15. A method according to claim 14 in which the processing frames of the training video together with the video encoder to generate the spatial-temporal encoding of the training video, is performed by:processing the frames of the frames of the training video individually with a spatial feature encoder to form respective spatial feature encodings, andaggregating the spatial feature encodings with a temporal aggregator, to generate for each frame of the training video a respective spatial-temporal encoding item.
16. A method according to claim 14 or claim 15, in which the output of the feature predictor neural network based on an individual frame of the training video is formed by: processing the individual frame with a spatial feature encoder to form a spatial feature encoding; andprocessing the spatial feature encoding with the feature predictor neural network to obtain the output of the feature predictor neural network.
17. A method of training a generator neural network configured, upon receiving a context visual item comprising a context image, and action data defining one or more actions, toAttorney Docket No.: 45288-0578W01generate a video item depicting one or more actors performing the one or more actions, a visual aspect of the video item being defined by the context image, the method comprising:obtaining an initial generator neural network configured to process a network input to generate a video item;a plurality of steps of modifying a current generator neural network to reduce a measure of a difference between:a training video item depicting one or more actors performing one or more actions, anda video item generated by the current generator neural network upon processing a context visual item comprising at least one context image included in the training video item and action data defining the one or more actions depicted in the training video item.
18. The method of claim 17 in which the context image defines the appearance of one or more actors depicted in the video item generated by the current generator neural network.
19. The method of claim 17 or claim 18 in which the context image defines the appearance of an environment depicted in the video item generated by the current generator neural network.
20. The method of any one of claims 17 to 19 in which the action data is obtained by processing at least part of the training video item.
21. The method of claim 20 in which the action data is obtained by processing the at least part of the training video item to extract spatial-temporal feature data encoding the one or more actions and the visual appearance of elements depicted in the training video item;processing the at least part of the training video item using a trained feature predictor neural network to extract context data indicating the visual appearance of elements depicted in the training video item; andobtaining the action data as a difference between the spatial-temporal feature data and the context data.Attorney Docket No.: 45288-0578W0122. The method of claim 21 in which the processing of the training video item to extract context data is performed by a feature predictor network obtained by the method of any of claims 13 to 16.
23. The method of claim 21 or claim 22 in which the processing of the training video item comprises performing a selected augmentation operation on the training video item to form an augmented video item, the feature data and the context data being extracted from the augmented video item.
24. A system comprising one or more computers and one or more storage devices storing instructions that when executed by the one or more computers cause the one or more computers to perform the operations of the method of any one of claims 1-23.
25. One or more computer storage media storing instructions that when executed by one or more computers cause the one or more computers to perform the operations of the method of any one of claims 1-23.