ANIMATION OF IMAGES USING POINT TRAJECTORIES
By generating point trajectories and then video frames using diffusion neural networks, the system efficiently animates still images with realistic movements, addressing high computational costs and pose ambiguity.
Patent Information
- Application Number
- BR112025018927
- Authority / Receiving Office
- BR · BR
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-02-01
- Filing Date
- 2024-03-08
- Publication Date
- 2026-07-28
AI Technical Summary
Animating still images is challenging due to high computational cost and difficulty in predicting realistic movements, and it is poorly posed as there are many possible and realistic future trajectories for objects in a still image.
The system decomposes the task into two parts: generating a set of point trajectories from an input image and then generating video frames from these trajectories and the input image, using diffusion neural networks to ensure realistic and plausible movement.
This approach reduces computational complexity and ensures the generated videos depict realistic object movements, allowing multiple plausible animations from the same input image.
Smart Images

Figure 00000000_0000_ABST
Abstract
Description
1 / 62 ANIMATION OF IMAGES USING TRAJECTORIES OF CROSS-REFERENCE POINTS TO RELATED REQUESTS
[001] This application claims priority for U.S. Provisional Applications No. 63 / 450,951, filed March 8, 2023; No. 63 / 452,405, filed March 15, 2023; and No. 63 / 548,824, filed February 1, 2024. The disclosure of the prior applications is considered part of, and is incorporated herein by reference into, the disclosure of this application. BACKGROUND
[002] This descriptive report refers to processing inputs that include images using neural networks.
[003] Neural networks are machine learning models that employ one or more layers of nonlinear units to predict an output for a given input. Some neural networks include one or more hidden layers in addition to an output layer. The output of each hidden layer is used as input for the next layer in the network, that is, the next hidden layer or the output layer. Each layer of the network generates an output from a given input according to the current value inputs of a respective set of parameters. SUMMARY
[004] This descriptive report describes a system implemented as computer programs on one or more computers in one or more locations that animate images using a generative neural network system.
[005] Animating an input image refers to generating a video that represents an animation of the input image, for example, a video depicting how the scene depicted in the input image changes over time. Examples of changes over time include the movement of objects in the scene, as well as changes in Petition 870250079681, dated 05 / 09 / 2025, page 11 / 99 2 / 62 Lighting and other image properties.
[006] This descriptive report also describes techniques for training a point-tracking neural network.
[007] A point-tracking neural network is a neural network that processes an input video to generate a network output that includes, for each query point in a set of one or more query points in a given frame of the input video, a respective predicted spatial position of the query point in the other video frames in the sequence.
[008] After training the point-tracking neural network, the point-tracking neural network can be used for a variety of purposes.
[009] As an example, the point-tracking neural network can be used to generate training data to train generative neural networks in the generative neural network system that animates the images.
[0010] Other examples of uses for the point-tracking neural network will be described below.
[0011] The subject matter described in this descriptive report can be implemented in particular ways in order to achieve one or more of the following advantages.
[0012] Animating still images is generally a difficult problem, both because video modeling usually has an extremely high computational cost and because it is difficult to predict realistic movements from a single still image.
[0013] To address these issues, the described system decomposes the challenging task of animating still images into two parts: (i) first generating a set of point trajectories from an input image and (ii) then generating the video from the point trajectories and the input image. That is, the system first generates a Petition 870250079681, dated 05 / 09 / 2025, page 12 / 99 3 / 62 A set of point trajectories that are a dense and explicit representation of the surface motion of objects in the scene depicted in the input image, and then generates the frame pixels in the video from the point trajectories and the input image.
[0014] Thus, the described system can use point trajectories to meet the appropriate location in the input image when generating the frame pixels in order to produce an appearance that is consistent throughout the video. In particular, point trajectories dictate the movement that should be present in the video, ensuring the physical plausibility of the video because the video is generated conditioned by the point trajectories already generated.
[0015] In addition, by first generating point trajectories, the system decomposes the video generation problem into two computationally efficient steps and also ensures that the videos generated by the system show realistic movement of the objects depicted in the input image.
[0016] Furthermore, animating a still image without additional information is a poorly posed problem, because there may be many possible and realistic future trajectories for any object depicted in the still image. By using a diffusion neural network to generate point trajectories from the still image, the system can effectively sample from the space of realistic surface motion trajectories in order to ensure that the final generated video represents a realistic sample of the highly multimodal distribution of the object's motion. Moreover, since the system uses a diffusion neural network to generate the point trajectories, the system can effectively sample multiple plausible sets of point trajectories, given the same input image, allowing the system to generate multiple different plausible videos from the same input image, each representing a different sample of the distribution. Petition 870250079681, dated 05 / 09 / 2025, page 13 / 99 4 / 62 of the object's movement.
[0017] This descriptive report also describes techniques for training a point-tracking neural network that effectively leverages unlabeled data to improve neural network training through unsupervised learning. As a particular example, the system can use the described unsupervised learning techniques to fine-tune a pre-trained point-tracking neural network that was trained through supervised learning. That is, the system can use the described unsupervised learning technique to leverage unlabeled data to improve the performance of the pre-trained neural network.For example, the neural network may have been trained on a dataset that includes synthetic video sequences, for example, which includes only or mainly synthetic video sequences, and the system can then further train (fine-tune) the neural network through unsupervised learning on fine-tuned data that includes untagged real-world video sequences. This can improve the point-tracking neural network's ability to generalize to diverse real-world videos, that is, to perform point-tracking tasks that require processing real-world video sequences, even with no, or limited, tagged real-world data available. Particularly, while true pixel-level trajectories can be easily generated when generating a synthetic video, obtaining accurate pixel-level trajectory tags for real-world videos can be difficult or impossible.The techniques described allow the system to incorporate untagged real-world videos in order to improve the performance of the trained point-tracking neural network.
[0018] The details of one or more modalities of the subject matter of this descriptive report are set out in the attached drawings and in the description. Petition 870250079681, dated 05 / 09 / 2025, page 14 / 99 5 / 62 description below. Other features, aspects and advantages of the material will become apparent from the description, figures and claims. BRIEF DESCRIPTION OF THE FIGURES
[0019] FIG. 1A shows an example of generating a video that animates an input image.
[0020] FIG. 1B is a diagram of an example image animation system.
[0021] FIG. 2 is a flow diagram of an example process for generating a set of point paths from an input image.
[0022] FIG. 3 is a flow diagram of an example process for generating a video from an input image and a set of point paths.
[0023] FIG. 4 shows example architectures of the point-tracking neural network.
[0024] FIG. 5 shows an example of training the point-tracking neural network.
[0025] Similar reference numbers and designations in the various figures indicate similar elements. DETAILED DESCRIPTION
[0026] FIG. 1A shows an example of generating a video 102 that animates an input image 104.
[0027] As shown in the example in FIG. 1A, an image animation system implemented as computer programs on one or more computers in one or more locations receives the input image 104.
[0028] For example, the system may receive input image 104 from a system user.
[0029] The system then animates the input image 104 generating Petition 870250079681, dated 05 / 09 / 2025, page 15 / 99 6 / 62 a video 102 that represents an animation of input image 104 depicting how the scene depicted in input image 104 changes over time.
[0030] That is, even though the input image 104 is a static image at a single point in time, the system generates a video 102 that includes a corresponding video frame at each of the multiple time steps (starting with the input image 104 as the first frame in the first time step in the video) and represents a realistic estimate of how the scene depicted in the input image 104 would change over time.
[0031] Specifically, the system breaks down the task of animating the video into two steps.
[0032] First, the system processes the input image 104 to generate a set of point trajectories 106.
[0033] Each point trajectory 106 corresponds to a different point in the input image 104 and includes, for each of the time steps in the video, a predicted spatial position of the corresponding point in a video frame at the time step in the video (to be generated by the system).
[0034] Each point is a point in a corresponding video frame, that is, it specifies a respective spatial position, that is, a respective pixel, in a corresponding plurality of video frames. Thus, each point in a given trajectory 106 can be represented as a point (x, y, t), where x, y are the spatial coordinates of the point and t is the index of the corresponding video frame in the video 102.
[0035] The point trajectories are represented as dashed curves in FIG. 1A, with points at future time points represented as points on the curve and points that are closer to the tip of the curve being further away in time from the entry image. Petition 870250079681, dated 05 / 09 / 2025, page 16 / 99 7 / 62 of. Although only three point trajectories are shown in FIG. 1 for ease of illustration, in practice the system can generate many more point trajectories so that the trajectories represent a dense representation of the future movement of points in the input image 104. For example, the system can generate a respective point trajectory for each grid cell in a grid, for example, an 8-pixel by 8-pixel grid superimposed on the input image.
[0036] Optionally, the 106 point trajectory may also include a corresponding occlusion score for each of the frames representing the probability that the corresponding point will be occluded in the video frame at the time step.
[0037] For example, the system can generate the trajectories of 106 points from the input image 104 using a generative neural network.
[0038] This will be described in more detail below.
[0039] The system then processes the input image 104 and the point trajectories 106 to generate the video 102, that is, it generates the frames in the video 102 from the input image 104 and the point trajectories 106.
[0040] For example, the system can generate video 102 from input image 104 and point trajectories 106 using another generative neural network.
[0041] Although not shown in FIG. 1A, because in the example in FIG. 1A, point trajectories 106 indicated that the points on the person's right arm would likely move away from the person's body, video 102 may show the person raising their right arm in future frames. Similarly, as point trajectories 106 indicated that the points on the person's left leg would likely move upward and away from the body Petition 870250079681, dated 05 / 09 / 2025, page 17 / 99 8 / 62 person, the video may show the person kicking the other person's left leg in future frames.
[0042] Thus, the system breaks down the challenging task of animating still images by first generating a dense and explicit representation of the surface motion of objects in the scene depicted in input 104, that is, by virtue of point trajectories 106, and then generating the frame pixels in the video from point trajectories 106 and input image 104.
[0043] Thus, the system can use the trajectories 106 to meet the appropriate location in the input image 104 when generating the frame pixels in order to produce an appearance that is consistent throughout the video. In particular, the point trajectories 106 dictate the movement that should be present in the video 102, ensuring physical plausibility.
[0044] FIG. 1B is a diagram of an example image animation system 100. Image animation system 100 is an example of a system implemented as computer programs on one or more computers in one or more locations, in which the systems, components, and techniques described below can be implemented.
[0045] As described above, system 100 generates a video 102 that animates an input image 104, that is, it generates a video 102 that represents an animation of the input image 104 that depicts how the scene depicted in the input image 104 changes over time.
[0046] Specifically, system 100 receives an input image 104. For example, the system may receive the input image 104 as input from a user.
[0047] System 100 processes a first input derived from input image 104 using a first generative neural network Petition 870250079681, dated 05 / 09 / 2025, page 18 / 99 9 / 62 110 to generate respective point trajectories 106 for each of one or more points in the input image 104. The first generative neural network 110 will also be referred to in this descriptive report as the trajectory model.
[0048] For example, the points can be pixels randomly sampled from the input image 104, they can be pixels in a grid overlaid on the input image 104, or they can be points specified by the user who provided the input image 104 (or another user).
[0049] Each point trajectory 106 includes, for each of the plurality of time steps in the video 102, a predicted spatial position of the corresponding point in a video frame at the time step in the video 102 (to be generated by the system 100).
[0050] Optionally, the 106-point trajectory may also include a corresponding occlusion score for each of the frames. The occlusion score for a given frame estimates the probability that the corresponding point will be occluded in the video frame at the given time step.
[0051] For example, the first generative neural network 110 might be a diffusion neural network that generates each point path from a corresponding noisy path conditioned on the first input. This diffusion neural network will be referred to as a path diffusion neural network. As used in this document, noisy refers to values that have been sampled from a noise distribution, for example, a Gaussian distribution or other appropriate distribution. For a given point path, the noisy path therefore includes, for each spatial position in the given point path, a corresponding spatial position that is sampled from a corresponding noise distribution.
[0052] The path diffusion neural network can generally have Petition 870250079681, dated 05 / 09 / 2025, page 19 / 99 10 / 62 any suitable neural network architecture that allows the path diffusion neural network to map an input that includes a noisy path to a noise removal output that defines an update for the noisy path.
[0053] As an example, the path diffusion neural network can be a two-dimensional convolutional neural network, for example, one that has a U-Net or other convolutional architecture. Optionally, the path diffusion neural network can include, as part of the convolutional architecture, one or more self-attention layers.
[0054] The generation of the set of trajectories using the trajectory diffusion neural network will be described in more detail below with reference to FIG. 2.
[0055] System 100 then generates each of the video frames in video 102 using a second generative neural network 120 and based on the input image 104 and one or more point trajectories 106. The second generative neural network 120 will also be referred to in this descriptive report as the pixel model.
[0056] For example, the second generative neural network 120 may be a diffusion neural network that generates each video frame (that is, generates one or more intensity values for each of the pixels in the video frame) in the video from a corresponding noisy video frame conditioned on the input image 104 and the trajectories 106. This diffusion neural network will be referred to as a pixel diffusion neural network.
[0057] The pixel diffusion neural network can generally have any appropriate neural network architecture that allows the pixel diffusion neural network to map an input to an output image.
[0058] As an example, the pixel diffusion neural network can Petition 870250079681, dated 05 / 09 / 2025, page 20 / 99 11 / 62 being a convolutional neural network, for example, one that has a U-Net or other convolutional architecture. Optionally, the pixel diffusion neural network may include, as part of the convolutional architecture, one or more self-attention layers.
[0059] The generation of video 102 will be described in more detail below with reference to FIG. 3.
[0060] Thus, system 100 uses two different generative neural networks, one that generates the point trajectories 106 and one that generates the video 102, given the point trajectories 106. Thus, the outputs of the trajectory model dictate the movement that must be present in the video 102 generated by the pixel model, ensuring the physical plausibility of the generated video 102.
[0061] Once video 102 is generated, system 100 can use the video for any of a variety of purposes. For example, system 100 can store video 102 or provide video 102 for presentation to a user, for example, the user who submitted input image 104.
[0062] Before using the first and second generative neural networks 110 and 120 to generate videos, system 100 or another training system trains neural networks 110 and 120.
[0063] For example, the training system can train the first generative neural network 110 on a training dataset that includes (i) a plurality of video sequences, where a video sequence is a sequence of video frames and (ii) for each of the video sequences, a respective point trajectory for each of one or more points in the first frame within the video sequence.
[0064] In other words, the training system can generate training data to train the first generative neural network 110 from a set of video sequences. For example, the system can Petition 870250079681, dated 05 / 09 / 2025, page 21 / 99 12 / 62 generate training examples, each corresponding to one of the video sequences, and each including, as a training input image, the first frame within the corresponding video sequence and, as the target output, the respective point trajectories to the points in the first frame.
[0065] The training system can then train the first generative neural network 110 on the training examples using an objective that is appropriate for the type of neural network being used. For example, when the first generative neural network 110 is a diffusion neural network, the objective could be a score matching objective.
[0066] As another example, the training system can train the second generative neural network 120 on the same training dataset.
[0067] That is, the training system can generate training data to train the second generative neural network 120 from the set of video sequences. For example, the system can generate training examples that each correspond to one of the video sequences and include, each as a training input image, the first frame within the corresponding video sequence, as a target set of point trajectories, the respective point trajectories for the points in the first frame within the corresponding video sequence, and as the target output, the corresponding video sequence.
[0068] The training system can then train the second generative neural network 120 on the training examples using an objective that is appropriate for the type of neural network being used. For example, when the second generative neural network 120 is a diffusion neural network, the objective could be a score matching objective. Petition 870250079681, dated 05 / 09 / 2025, page 22 / 99 13 / 62
[0069] In some cases, the training system can generate point trajectories for at least some of the video sequences by processing the video sequence using a point-tracking neural network.
[0070] A point-tracking neural network is a neural network that processes an input video to generate a network output that includes, for each query point in a set of one or more query points in a given frame of the input video, a respective predicted spatial position of the query point in the other video frames in the sequence and, optionally, an occlusion estimate for the query point.
[0071] That is, because a large number of densely tagged videos may not be available for use in training generative neural networks, the training system can use the point-tracking neural network to predict point trajectories for at least some of the video sequences and then use the predictions of the point-tracking neural network to train the first and second generative neural networks 110 and 120.
[0072] Point-tracking neural networks can generally have any suitable architecture and can be trained using any appropriate technique.
[0073] Some examples of architectures for the point-tracking neural network are described below with reference to FIG. 4.
[0074] Another example of a neural network architecture for point tracking is described in Doersch, et al, TAP-Vid: A Benchmark for Tracking Any Point in a Video, arXiv:2211.03726.
[0075] An example of point-tracking neural network training is described below, with reference to FIG. 5.
[0076] Other examples of point-tracking neural network training are described in Doersch, et al, TAP-Vid: A Ben Petition 870250079681, dated 05 / 09 / 2025, p. 23 / 99 14 / 62 chmark for Tracking Any Point in a Video, arXiv: 2211.03726 and Doersch, et al, TAPIR: Tracking Any Point with per-frame Initialization and temporal Refinement, arXiv:2306.08637.
[0077] FIG. 2 is a flow diagram of an example process 200 for generating a set of point paths. For convenience, process 200 will be described as being performed by a system of one or more computers located in one or more locations. For example, an image animation system, for example, the image animation system 100 depicted in FIG. 1B, appropriately programmed according to this descriptive report, can perform process 200.
[0078] The system receives an input image (step 202).
[0079] The system processes the input image using an image encoding neural network to generate an encoded representation of the input image (step 204).
[0080] Generally, the encoded representation includes a respective feature vector for each of the multiple spatial regions in the input image. For example, when the input image is an H x W image, the encoded representation might be an H / k x W / k map of feature vectors. For example, k could be equal to 4, 8, or 16.
[0081] The image encoding neural network can generally have any neural network architecture suitable for encoding input images. For example, the image encoding neural network can be a convolutional neural network or a view transformer neural network.
[0082] In some cases, the system uses a pre-trained image encoding neural network that has already been trained to generate image representations in a representation learning task. In some other cases, the system trains the encoding neural network of Petition 870250079681, dated 05 / 09 / 2025, page 24 / 99 15 / 62 image along with the first generative neural network.
[0083] The system generates, for each point path in the set, a corresponding noise path (step 206).
[0084] Specifically, for a given point path, the noisy path includes a corresponding value for each value in that given point path. To generate the noisy path, the system collects samples of each of these values from a corresponding noise distribution, for example, a Gaussian distribution or another appropriate distribution.
[0085] Thus, as described above, each point trajectory includes, for each time step in the video, (i) the predicted spatial position of the corresponding point in the video frame at the time step and, optionally, (ii) an occlusion score that estimates a probability that the corresponding point will be occluded in the video frame at the time step.
[0086] The spatial positions predicted in the video frames can be absolute spatial positions, that is, represented as absolute coordinates in an image coordinate system, or relative spatial positions, that is, represented as coordinates relative to the corresponding point in the input image.
[0087] Therefore, the system samples the predicted spatial position and, when included, the occlusion score for each time step of the respective noise distributions, resulting in the trajectory including, for each time step, noisy coordinates, that is, a noisy predicted spatial position and a noisy occlusion estimate.
[0088] Optionally, the noisy path may also include additional values that help the first generative neural network to make effective use of the information contained in the noisy path.
[0089] For example, noisy coordinates can be increased Petition 870250079681, dated 05 / 09 / 2025, page 25 / 99 16 / 62 tadas with a positional encoding. That is, the noisy trajectory may include, for each set of noisy coordinates, a positional encoding of the noisy coordinates. As a particular example, the positional encodings may be Fourier positional encodings that encode the noisy coordinates using a fixed number of Fourier features.
[0090] As another example, the noisy trajectory may also include, in addition to absolute noisy coordinates, relative noisy coordinates, for example, coordinates of the predicted noisy spatial position in a coordinate system centered on the corresponding point in the input image.
[0091] The system then generates each point path in the set from the corresponding noisy path and the encoded representation of the input image (step 208) using the first generative neural network.
[0092] In particular, in the example in FIG. 2, the first generative neural network is a diffusion neural network (path diffusion neural network).
[0093] To generate point trajectories, the system performs a sequence of back-diffusion iterations using the path diffusion neural network.
[0094] In each reverse diffusion iteration, the system processes an input to the reverse diffusion iteration that includes the noisy point trajectories from the reverse diffusion iteration using the trajectory diffusion neural network and conditioned to the encoded representation of the input image to generate a noise removal output that defines an update for the noisy point trajectories.
[0095] For example, the noise removal output could be a prediction of, for each noisy point path, the point path. Petition 870250079681, dated 05 / 09 / 2025, page 26 / 99 17 / 62 tos real (unknown) corresponding. As another example, the noise removal output might be a prediction of, for each noisy point trajectory, the noise that was added to the corresponding real point trajectory to generate the noisy point trajectory. As described above, when the input includes both absolute and relative coordinates, in some implementations the noise removal output predicts the relative coordinates while, in some other implementations, the noise removal output predicts the absolute coordinates.
[0096] The system can condition the path diffusion neural network on the encoded representation of the input image in any one of a variety of ways.
[0097] For example, the system may include the encoded representation as part of the input to the iteration, for example, concatenated with noisy point trajectories.
[0098] As another example, the path diffusion neural network may include one or more conditioning layers that receive as input the encoded representation of the input image.
[0099] As an example, each conditioning layer can be a cross-attention layer that cross-attends in the encoded representation of the input image.
[00100] As another example, each conditioning layer can be a conditional group normalization layer. That is, after the group normalization layer means subtraction and normalization of variance within each group to produce a normalized output Z, a group normalization layer would typically apply a scaling and shifting operation. To generate a conditional group normalization layer, these are replaced by linear projections of the conditioning input. For example, the system might resize the encoded representation so that its spatial dimensions are the same size as Z and, Petition 870250079681, dated 05 / 09 / 2025, page 27 / 99 18 / 62 Next, apply the respective transformations learned, for example, two 1 χ 1 convolution layers, to create a scale and a multiplier that are the same size as Z. These scales and multipliers can then be applied in place of the scaling and shifting operation.
[00101] For each reverse diffusion iteration, the system then uses the noise removal output to update the noisy trajectories. For example, the system can generate an estimate of the updated noisy trajectories from the noise removal output and then apply a diffusion sampler, for example, the DDPM (Probabilistic Diffusion Model Noise Removal) sampler, the DDIM (Implicit Diffusion Model Noise Removal) sampler, or another appropriate sampler, to the noise removal output to generate the updated noisy trajectories. When the noise removal output is a prediction of, for each noisy point trajectory, the corresponding actual (unknown) point trajectory, the system can directly use the noise removal output as the estimate.When the noise removal output is a prediction of, for each noisy point path, the noise that was added to the corresponding actual point path to generate the noisy point path, the system can determine the estimate of the actual noisy point paths, the noise removal output, and a noise level for the current reverse diffusion iteration. Optionally, after the last reverse diffusion iteration, the system can refrain from using the diffusion sampler and can instead use the estimate as the updated noisy paths.
[00102] The system then uses the updated noisy trajectories after the last reverse diffusion iteration to generate the endpoint trajectories.
[00103] When point trajectories include estimates of Petition 870250079681, dated 05 / 09 / 2025, page 28 / 99 19 / 62 occlusion, because the diffusion neural network operates in a continuous space, the system can apply a smoothing operation to make the occlusion continuous during the backdiffusion iterations. The system can then backtrack from a binary occlusion estimate to the noisy occlusion estimate after the last backdiffusion iteration.
[00104] For example, let õfser be an occlusion indicator at time t. For example, the occlusion indicator might be equal to 1 if the point is occluded at time t -1 otherwise. For each point t, let f be the nearest time such that õ? Φ õf. The system can calculate õt= Qt* (L — φ1^1) to generate a value that decays exponentially towards the extreme values 1 and -1 as the distance from a transition increases, but preserves the sign of õf, making decoding straightforward. The system can use õf as the occlusion estimate for each time step, scaled down to the range (0; 0.2). This scaling encourages the model to reconstruct the motion first and then reconstruct the occlusion value based on the reconstructed motion. The system can then backtrack from the transitions of the õf estimates in order to reconstruct the occlusion indicators õf for each of the time steps.
[00105] Thus, the system iteratively removes noise from noisy trajectories to generate endpoint trajectories.
[00106] As described above, the path diffusion neural network can generally have any appropriate neural network architecture that allows the path diffusion neural network to map an input that includes a noisy path to a noise removal output that defines an update for the noisy path.
[00107] As an example, the path diffusion neural network can be a two-dimensional convolutional neural network, for example, Petition 870250079681, dated 05 / 09 / 2025, page 29 / 99 20 / 62 one that has a U-Net or other convolutional architecture. Optionally, the path diffusion neural network may include, as part of the convolutional architecture, one or more self-attention layers.
[00108] FIG. 3 is a flow diagram of an example process 300 for generating a video from an input image and a set of point paths. For convenience, process 300 will be described as being performed by a system of one or more computers located in one or more locations. For example, an image animation system, for example, the image animation system 100 depicted in FIG. 1B, appropriately programmed according to this descriptive report, can perform process 300.
[00109] The system receives an input image and a set of point trajectories (step 302).
[00110] The system generates, for each video frame in the video, a corresponding noisy video frame (step 304).
[00111] Specifically, for a given video frame, the noisy video frame includes a corresponding intensity value for each intensity value in that video frame. To generate the noisy video frame, the system samples each of these values from a corresponding noise distribution, for example, a Gaussian distribution or another appropriate distribution.
[00112] The system then generates each video frame from the corresponding noisy video frame, the input image, and the point trajectories (step 306) using the second generative neural network.
[00113] In particular, in the example in FIG. 3, the second generative neural network is a diffusion neural network (pixel diffusion neural network).
[00114] To generate the video frames, the system performs a se Petition 870250079681, dated 05 / 09 / 2025, page 30 / 99 21 / 62 sequence of reverse diffusion iterations using the pixel diffusion neural network. That is, in each reverse diffusion iteration, the system updates the corresponding noisy video frame for each of the video frames in each of the time steps.
[00115] In each reverse diffusion iteration and for each video frame, the system processes an input for the reverse diffusion iteration that includes the corresponding noisy video frame from the reverse diffusion iteration using the pixel diffusion neural network and conditioned to the input image and point trajectories to generate a noise removal output that defines an update for the noisy video frame.
[00116] For example, the noise removal output may be a prediction of the corresponding (unknown) actual video frame.
[00117] As another example, the noise removal output can be a prediction of the noise that was added to the corresponding actual video frame to generate the noisy video frame.
[00118] The system can condition the pixel diffusion neural network on the input image and point trajectories in any of a variety of ways.
[00119] As an example, the system may include, in the input to the diffusion neural network for a given video frame at a given point in time, a version of the input image that has been distorted according to one or more point trajectories to represent the image. That is, the system may generate a distorted version of the input image that has been distorted according to one or more point trajectories, i.e., it is a representation of how the input image appears at the point in time if the points move according to the point trajectories.
[00120] The system can generate a distorted version of the input image in any of a variety of ways. Petition 870250079681, dated 05 / 09 / 2025, page 31 / 99 22 / 62
[00121] As an example, the system can use a stretch-based distortion. For a given frame t, the trajectories at time t identify where each stretch in the input image should appear. Therefore, the system can construct a new image where each local stretch is placed in its correct location, using bilinear interpolation (for example) to achieve subpixel precision. However, this can produce gaps between stretches under certain circumstances. In some implementations, to account for this, the system may actually distort a larger stretch around each point. When multiple stretches appear covering the same output pixel, the system may weight them inversely proportional to their distance from the center of the strip.
[00122] Optionally, to address the fact that aliasing can occur when multiple segments overlap, the system can perform the distortion by distorting each segment multiple times, with the difference between the distortions being the way the system calculates the blending weights. The system can then include all distorted versions of the input image in the input to the diffusion neural network. For example, let pÉJ.fser be the position of the trajectory starting at point i, j in the original image at time t according to the point trajectory to point i, j. In the original distortion, the weight for the segment carried by the trajectory to any particular pixel pfe (assuming p'f is close enough to be within the segment) is proportional to ι / άφ^ρ'^, where d is the Euclidean distance. The system can instead use different weightings for different distortions. For example, the weighting can be i / d^ (v+q,pff) where η is different for different distortions.For example, when there are five distortions, the system can have η and {(0, 0), (-2, 0), (0,-2), (2, 0), 0, 2)}. This allows the model to use the differences in intensities between them. Petition 870250079681, dated 05 / 09 / 2025, page 32 / 99 23 / 62 different distortions to infer the original values for different segments and then use that information to 'undo' any aliasing.
[00123] The system can then include this distorted version of the input image with the noisy version of the video frame, for example, by concatenating the two images.
[00124] As another example, the system can condition the diffusion neural network on a distorted version of features extracted from the input image, that is, on features of the input image that have been distorted according to one or more point trajectories to represent the image.
[00125] For example, the features can be the feature vectors in an encoded representation of the input image. For example, the encoded representation can be the same encoded representation described above with reference to the step, or it can be a different encoded representation generated by a separately trained image encoding neural network.
[00126] To distort these features by a time step t, the system can use the position of each feature at time step t according to the set of point paths. The system can then use bilinear interpolation (for example) to place the feature in the appropriate location within a grid of distorted features. Optionally, the system can monitor the number of features that have been placed within any particular grid cell (specifically, the sum of bilinear interpolation weights) and normalize by that sum, unless the sum is less than 0.5, in which case the system can divide by 0.5.
[00127] The system can condition the diffusion neural network on the distorted feature version in any appropriate way, for example, using one of the conditioning techniques described. Petition 870250079681, dated 05 / 09 / 2025, page 33 / 99 24 / 62 above with reference to FIG. 2.
[00128] In some implementations, for at least a subset of the frames, the system also includes in the input to the neural network for the frame, temporal context from other frames in the video. For example, for a given frame, the input to a given reverse diffusion iteration might include the current (noisy) version of one or more previous frames in the video from the previous reverse diffusion iteration. Instead, or in addition, for a given frame, the input to a given reverse diffusion iteration might include the current (noisy) version of one or more subsequent frames in the video from the previous reverse diffusion iteration.
[00129] For each reverse diffusion iteration, the system then uses, for each video frame, the noise removal output for the video frame to update the corresponding noisy video frame. For example, the system can generate an estimate of the updated noisy video frame from the noise removal output and then apply a diffusion sampler, for example, the DDPM (Diffusion Probabilistic Noise Removal Model) sampler, the DDIM (Diffusion Implicit Noise Removal Model) sampler, or another appropriate sampler, to the estimate to generate the updated noisy video frame. When the noise removal output is a prediction of the corresponding (unknown) actual frame, the system can directly use the noise removal output as the estimate.When the noise removal output is a prediction of the noise that was added to the corresponding actual frame to generate the noisy frame, the system can determine the estimate of the actual noisy frame, the noise removal output, and a noise level for the current reverse diffusion iteration. Optionally, after the last reverse diffusion iteration, the system can refrain from using the diffusion sampler and can instead use the estimate as the noisy frame. Petition 870250079681, dated 05 / 09 / 2025, page 34 / 99 25 / 62 updated.
[00130] The system then uses the noisy video frames updated after the last reverse diffusion iteration to generate the video frames in the video.
[00131] FIG. 4 shows example architectures 410, 420 and 430 of a point-tracking neural network.
[00132] A point-tracking neural network is a neural network that processes an input that includes (i) a video that includes a plurality of video frames and (ii) a set of one or more lookup points.
[00133] Each lookup point is a point in a corresponding set of video frames, that is, it specifies a respective spatial position, that is, a respective pixel, in a corresponding set of video frames.
[00134] The 400 point tracking neural network processes the set of one or more query points and the video sequence, i.e., the pixel intensity values of the video frames in the video sequence, to generate a network output that includes, for each query point, a respective predicted spatial position of the query point in the other video frames in the sequence, i.e., in the video frames different from the video frame corresponding to the query point.
[00135] Because of the neural network architecture or the processing path design, the network output may also include a predicted spatial position for the corresponding video frame. However, in some of these cases, the system may disregard the predicted spatial position for the corresponding video frame (because the actual position in the corresponding video frame is provided as input to the system).
[00136] That is, given a lookup point in one of the frames of Petition 870250079681, dated 05 / 09 / 2025, page 35 / 99 In a 26 / 62 video, the point-tracking neural network can generate a prediction of the spatial position of the query point in the other video frames within the video.
[00137] The predicted position of a given lookup point in a given other video frame is a prediction of the location of the portion of the scene that was depicted at the given lookup point in the corresponding video frame. For example, if, at the given lookup point, the corresponding video frame depicted a particular point on the surface of an object in the scene, the predicted position of the given lookup point identifies the predicted position of the same particular point on the surface of the object in the given other video frame.
[00138] In some implementations, the point-tracking neural network also generates, for each query point, a corresponding occlusion score for the query point for each of the other video frames in the sequence. The occlusion score for a given query point in a given video frame represents the probability that the query point will be occluded in that given video frame, that is, that the portion of the scene that was depicted at the query point in the corresponding video frame will be occluded in that given video frame.
[00139] Specifically, in the example architecture 410, the point-tracking neural network is configured to, for a given video sequence and corresponding point in a given frame in the video sequence, generate a query feature for the corresponding point.
[00140] The neural network is then configured to generate, using the query feature, a cost volume that includes a corresponding cost map for each of the plurality of frames in the video sequence. Petition 870250079681, dated 05 / 09 / 2025, page 36 / 99 27 / 62
[00141] For example, the neural network can process a sequence of video frames and receive a lookup point in a video frame.
[00142] The neural network processes the video frames in the video sequence using a visual principal structure neural network to generate a feature grid that includes a respective visual feature, i.e., a feature vector, for each of a plurality of spatial locations in each of the video frames. Generally, each of the spatial locations corresponds to a different region of the video frame. For example, the feature grid might be a w / 8 x h / 8 grid of d-dimensional visual features, with each visual feature corresponding to an 8 x 8 pixel grid of the corresponding video frame.
[00143] The visual principal structure neural network can have any appropriate architecture that allows the neural network to map the video sequence to a feature grid. In the example in FIG. 3, the visual principal structure neural network is a 3D convolutional neural network (ConvNet), for example, a TSM-ResNet-18 or other appropriate convolutional neural network. In other examples, the visual principal structure neural network can be a different type of neural network, for example, a Vision Transformer neural network.
[00144] The neural network can then generate an extracted feature for the query point from the spatial position of the query point in the corresponding video frame and the respective visual features for one or more of the plurality of spatial locations in the corresponding video frame.
[00145] For example, the neural network can generate the extracted feature (query features) by performing an interpolation, for example, a bilinear interpolation, of the visual features of a set of spatial locations that are within a neighborhood. Petition 870250079681, dated 05 / 09 / 2025, page 37 / 99 28 / 62 local location of the respective spatial position (íg,;5) of the consultation in the corresponding video frame
[00146] The neural network then generates a cost volume from the feature grid and the feature extracted for the query point. For example, the cost volume might have a corresponding cost value for each spatial location in each of the video frames. That is, the cost volume includes an h'xw' x 1 grid of cost values for each video frame in the sequence.
[00147] To calculate the cost value for a given spatial location in a given video frame, the system calculates a scalar product between the extracted feature and the visual feature for the given spatial location in the given video frame.
[00148] The neural network is then configured to generate, for each of the plurality of frames, an initial position of the corresponding point in the frame and an initial occlusion estimate for the corresponding point in the frame using the cost map for the frame.
[00149] Generally, to generate the predicted positions, the neural network can process the cost volume using a decoding neural network to generate, for each video frame, a respective score for each spatial location in the video frame.
[00150] For each of the plurality of different video frames from the corresponding video frame, the neural network can then generate the initial predicted position from the respective scores for the spatial locations in the video frame.
[00151] When the neural network also predicts occlusion, as in the example in FIG. 4, the neural network can process the cost volume using the decoding neural network to generate, for each video frame other than the corresponding video frame, the respective predicted position of the query point in the video frame and the respective Petition 870250079681, dated 05 / 09 / 2025, page 38 / 99 29 / 62 occlusion score for the point of reference in the video frame.
[00152] For example, the neural network can perform this processing independently for each of the video frames. That is, for a given video frame, the neural network processes the h' x W x 1 portion of the cost volume for that given video frame using the decoding neural network to generate the predicted position of the query point within that given video frame and the occlusion score for the query point in that video frame.
[00153] For example, the decoder neural network may include a set of shared layers and respective branches for position inference and occlusion inference.
[00154] The decoder neural network processes the cost volume portion for the given video frame using shared layers, for example, using a convolutional layer followed by a rectified linear unit activation function (ReLU) layer, to generate a shared output.
[00155] For the occlusion inference branch, the decoding neural network processes the shared output using the layers in the occlusion inference branch to generate a single occlusion score (logit) for the given video frame. For example, the occlusion inference branch includes a first set of layers that collapse the shared output into a vector, for example, using spatial mean clustering, followed by a second set of layers that regress the occlusion logit for the given video frame from the single vector. For example, the second set of layers might include a linear layer, a Leaky ReLU, and another linear layer that produces a single logit. A Leaky ReLU might be a ReLU activation function that has a slope (e.g., a relatively small slope, such as a slope of 0.01) for negative values (e.g., opposite to a slope). Petition 870250079681, dated 05 / 09 / 2025, page 39 / 99 30 / 62 flat for negative values).
[00156] For position inference branching, the neural network can apply a set of layers, for example, a convolutional (Conv) layer with a single output, followed by a spatial softmax, to generate a respective score for each spatial location in the video frame.
[00157] The neural network can then calculate a smooth argmax to calculate the position from the respective scores.
[00158] To calculate the smooth argmax, the neural network can identify the spatial location with the highest score, that is, the argmax location, according to the respective scores for the spatial locations. The neural network can then identify each spatial location that is within a fixed-size window B of the spatial location with the highest score, that is, the argmax location, and then determine the predicted position by calculating a weighted average.
[00159] In other words, the neural network determines the predicted position by calculating a weighted sum of the spatial locations within the fixed-size window of the spatial location argmax, with the weight for each spatial location being calculated based on, for example, being equal to or directly proportional to the ratio between the score for the spatial location and a sum of the scores for the spatial locations within the fixed-size window.
[00160] In example architecture 410, the neural network uses the initial positions and initial occlusion estimates as the final outputs of the neural network.
[00161] In this example, a training system can train the neural network on the following loss function for a given query point for each frame t in a given training video: Petition 870250079681, dated 05 / 09 / 2025, page 40 / 99 31 / 62 (1 (ρ^,ρ^ + ÃBCE^dt, of) where õfé is the fundamental truth occupancy for the video frame t, p is the fundamental truth position of the query point in the video frame í, ofé is the predicted occupancy for the video frame t, pfé is the predicted position of the query point in the video frame t, is the Huber loss between pfe and pf, λ is a weight value that can be provided as input to the system or determined through a hyperparameter sweep, and BCE(õt,ot) is the binary cross-entropy loss between õfe or
[00162] In example architectures 420 and 430, once the initial estimate is generated, neural network 400 generates the point path to the corresponding point by refining the initial positions, refining the initial occlusion estimates for the plurality of frames, or both, using an initial point path that includes the initial positions and the initial occlusion estimates for the plurality of frames.
[00163] In particular, in the example in FIG. 4, the neural network refines the initial positions and initial occlusion estimates for the plurality of frames by performing one or more refinement iterations.
[00164] In each refinement iteration, the 400 neural network generates, from a current point trajectory from the refinement iterations, a set of local scoring maps that, for each frame, capture similarity between features in a neighborhood of the predicted position in the current point trajectory in the frame and in the query feature. For the first refinement iteration, the current point trajectory is the initial point trajectory. For any subsequent refinement iterations, the current point trajectory is the updated point trajectory after the previous refinement iteration. Petition 870250079681, dated 05 / 09 / 2025, page 41 / 99 32 / 62
[00165] The neural network then processes an input that includes the set of local score maps using a refinement neural network to generate an update for the current point trajectory. As a particular example, the input to the refinement neural network might include the set of local score maps, the query feature, and the current point trajectory.
[00166] For example, in example architecture 430, the refinement neural network is a depth-mixing neural network that propagates information across frames using depth-convolutional layers. Each depth-convolutional layer may, for example, have a temporal receptive field that spans a plurality (e.g., all or a subset) of the frames.
[00167] In this example, a training system can train the neural network on a loss that is a sum of the loss calculated above for the initial predictions and the predictions after each refinement iteration.
[00168] In some implementations, the neural network can also generate an uncertainty estimate for each predicted spatial position that represents the neural network's uncertainty in the prediction. For example, the neural network can output an additional logit within the occlusion path and process the additional logit to generate an uncertainty estimate ut. When the neural network makes use of refinement, the neural network also refines the uncertainty estimates in each refinement iteration.
[00169] In this example, the loss above may include an additional term Q — ô^bceCü^u^ where üté is an estimate of the fundamental truth uncertainty for the forecast that is equal to one if the Euclidean distance between pte and pt is greater than a threshold and zero otherwise.
[00170] What follows is a description of a specific example of Petition 870250079681, dated 05 / 09 / 2025, page 42 / 99 33 / 62 architecture 430.
[00171] Given an estimated position, occlusion, and uncertainty for each frame, the goal of each iteration / refinement procedure is to calculate an update that adjusts the estimate to be closer to the fundamental truth by integrating information over time. The update is based on a set of local score maps that capture the similarity of the query point (i.e., scalar products) to features in the vicinity of the trajectory.
[00172] For example, these can be computed in a pyramid of different resolutions, therefore, for a given trajectory, they have the form (H' x W' χ L), where H' = W' = 7, the size of the local neighborhood, and L is the number of levels of the spatial pyramid. For example, the different pyramid levels can be calculated by spatially grouping the volume of features F that includes features for spatial locations within video frames. This set of similarities is post-processed with a refinement network to predict the refined position, occlusion, and uncertainty estimate.
[00173] Unlike initialization, however, the neural network includes local scoring maps for multiple frames simultaneously as input for post-processing. As a particular example, the neural network might include the current position estimate, the raw query features, and the (flattened) local scoring maps in a tensor of shape T χ (C + K + 4), where C is the number of channels in the query feature, K = H' W' L is the number of values in the flattened local scoring map, and 4 extra dimensions for position, occlusion, and uncertainty.
[00174] The output of this refinement network in the iteration is a residual (jpyoyiY) that is added to the position, occlusion, uncertainty estimate and query feature, respectively. Petition 870250079681, dated 05 / 09 / 2025, page 43 / 99 34 / 62
[00175] is of T x C format; thus, after the first iteration, slightly different query features are used in each frame when calculating new local score maps.
[00176] These positions and score maps are fed into a convolutional network to compute an update. For example, each block might include a 1x1 convolution block and a depth-in convolution block. For example, the 1x1 convolution blocks might be cross-channel layers and the depth-in convolutions might be layers within the channel. Because the architecture is convolutional and uses convolutions to incorporate convolutions over time, this convolutional architecture can run on sequences of any length.
[00177] FIG. 5 shows an example 500 of training a point-tracking neural network, for example, a point-tracking neural network that has one of the 410, 420 or 430 architectures or that has a different network architecture.
[00178] In the example in FIG. 5, the system fine-tunes a pre-trained point-tracking neural network through unsupervised learning.
[00179] As a particular example, the system can fine-tune a pre-trained point-tracking neural network that has been trained through supervised learning on a dataset that includes synthetic video sequences, for example, that includes only or mainly synthetic video sequences, through unsupervised learning on fine-tuned data that includes untagged real-world video sequences. This can improve the point-tracking neural network's ability to generalize to diverse real-world videos, i.e., to perform point-tracking tasks that require sequence processing. Petition 870250079681, dated 05 / 09 / 2025, page 44 / 99 35 / 62 real-world video companies.
[00180] Specifically, the system receives data specifying a point-tracking neural network professor 510.
[00181] The teacher neural network 510 is configured to receive a teacher point-tracking input that includes (i) a video sequence comprising a plurality of video frames and (ii) a query point in one of the plurality of video frames and to generate a teacher point-tracking output comprising, for each other video frame of the plurality of video frames, a position of the query point in the other video frame and an occlusion estimate for the query point in the other video frame.
[00182] The system then trains an aluna 520 point-tracking neural network through unsupervised learning.
[00183] Like the teacher neural network 510, the student point-tracking neural network 520 is configured to receive a student point-tracking input that includes (i) the video sequence comprising the plurality of video frames and (ii) the query point in one of the plurality of video frames and to generate a student point-tracking output that includes, for each other video frame of the plurality of video frames, a position of the query point in the other video frame and an occlusion estimate for the query point in the other video frame.
[00184] In particular, prior to training, the teacher point-tracking neural network 510 was pre-trained. In the example in FIG. 5, the teacher point-tracking neural network was pre-trained by means of supervised learning, for example, on a dataset of synthetic video sequences to produce pre-trained network parameters 502.
[00185] For example, the system or other training system Petition 870250079681, dated 05 / 09 / 2025, page 45 / 99 36 / 62 may have pre-trained the Professor 510 point-tracking neural network on one of the loss functions above.
[00186] For example, the teacher neural network may have a 410, 420, or 430 architecture, or a different point-tracking neural network architecture.
[00187] Furthermore, in the example in FIG. 5, the teacher and student 510 and 520 point-tracking neural networks have the same architecture or, more generally, the student 520 point-tracking neural network includes all the same parameters as the teacher 510 neural network and, optionally, additional parameters. For example, student 520 may include, as part of the main structure, one or more additional convolutional residual layers that are each initialized to represent an identity transformation.
[00188] To take advantage of this, before training the student 520 point-tracking neural network, the system initializes student 520 point-tracking neural network parameters using the teacher 510 point-tracking neural network, i.e., to be the pre-trained network parameters 502.
[00189] The system then trains the aluna 520 point-tracking neural network in multiple training steps.
[00190] In each training step, the system receives a set of one or more training video sequences from a training dataset that includes a plurality of training video sequences. As shown in the example in FIG. 5, the videos in the training dataset can be real-world videos that include real-world video frames.
[00191] For each training video sequence, the system applies one or more transformations to the training video sequence to generate a transformed video sequence (re-frames) Petition 870250079681, dated 05 / 09 / 2025, page 46 / 99 37 / 62 corrupted images). For example, one or more transformations include spatial transformations, image corruption, or both.
[00192] As a specific example, given an input video, the system can create a second view of the input video by resizing each frame to a lower resolution (varying linearly over time, for example, over the course of training) and overlaying them onto a black background at a random position (also varying linearly over time). This can be computed in an affine transformation aligned to the frame axis Φ in coordinates, applied to the pixels.
[00193] Optionally, the system can further degrade this visualization by applying random JPEG degradation to make the task more difficult, before pasting it onto the black background. Both operations lose texture information; therefore, the network must learn higher-level—and possibly semantic—cues (e.g., the tip of the cat's upper left ear), rather than lower-level texture matching, in order to track points correctly.
[00194] The system then selects one or more teacher lookup points within the training video sequence. For example, the system can sample each teacher lookup point by sampling a position and a time step index within the training video sequence, both uniformly random.
[00195] For each teacher lookup point, the system generates a teacher point tracking output for the teacher lookup point by processing the video sequence using the teacher point tracking neural network 510.
[00196] The system identifies a corresponding student lookup point in the transformed video sequence, that is, it identifies a student lookup point in the transformed video sequence that corresponds to the teacher lookup point. Petition 870250079681, dated 05 / 09 / 2025, page 47 / 99 38 / 62
[00197] Specifically, the system can identify an initial student query point using the teacher point tracking output and then transform the initial student query point consistent with one or more transformations to generate the final student query point. That is, the system can also apply Φ to the student query coordinates to generate the coordinates of the final student query point.
[00198] For example, to identify an initial student lookup point, the system can sample a point randomly from the trajectory generated for the teacher lookup point. This can enforce that each track forms an equivalence class by training the student 520 neural network to produce the same track, regardless of which point is used as a lookup.
[00199] The system generates a student point-tracking output for the student query point by processing the transformed video sequence using the student 520 point-tracking neural network.
[00200] The system then trains the student point-tracking neural network using a loss function that, for each training video sequence and for each teacher lookup point, measures a difference between (i) the student point-tracking output for the corresponding student lookup point and (ii) the teacher point-tracking output for the teacher lookup point.
[00201] For example, the system can, for each training video sequence and for each teacher query point, generate a pseudo-markup from the teacher point tracking output. The loss function can then measure an error between the pseudo-markup and the student point tracking output.
[00202] For example, when training using the loss above, the system can generate a pseudo-markup for a given frame of Petition 870250079681, dated 05 / 09 / 2025, page 48 / 99 39 / 62 video (i) defining the location of the fundamental truth in the given video frame for spatial prediction in the teacher point tracking output, (ii) defining the fundamental truth occlusion score to 1 only if the teacher occlusion estimate exceeds a threshold value, for example, zero, and optionally, (iii) defining the fundamental truth uncertainty to 1 only if the distance between the student prediction and the teacher prediction for the frame exceeds the threshold distance.
[00203] Next, the system can train the aluna point-tracking neural network on the loss function above, but with pseudo-tagging in place of the fundamental truth outputs.
[00204] Note, however, that if the teacher has not tracked the point correctly, the student query may be a different real-world point than the teacher's, leading to an erroneous training signal. In some implementations, to account for this, the system may determine, for each teacher query point, whether to mask the teacher query point from the loss function by applying cycle consistency between the student point tracking output and the teacher point tracking output, i.e., and then determine that it should mask any teacher query point that is not cycle-consistent. Masking a teacher query point refers to removing or setting to zero all quantities that depend on the teacher query point when calculating the loss function.
[00205] Furthermore, it may be that the teacher's predictions are less accurate than the student's for points that are closer in time to the student query frame.In some implementations, to take this into account, the system can determine, for each teacher query point, whether to mask the predictions for the teacher query point to one of the loss function's training video frames by applying a mask. Petition 870250079681, dated 05 / 09 / 2025, page 49 / 99 40 / 62 proximity. That is, the system can apply the proximity mask to mask, from the loss function, any training video frames that are closer to the student lookup frame than to the teacher lookup frame. Masking a training video frame to a teacher lookup point refers to removing or setting to zero all quantities that depend on both the training video frame and the teacher lookup point when calculating the loss function.
[00206] Optionally, during unsupervised training, the system can continue training the point-tracking neural network through supervised learning, for example, to help prevent catastrophic forgetting. In such cases, in some or all training steps, the system also trains the student's neural network through supervised learning on a set of tagged video sequences.
[00207] In some or all training steps, the system can update the teacher 510 point-tracking neural network using the student 520 point-tracking neural network. For example, as shown in FIG. 5, the system can, at specified points during training, update the parameters of the teacher 510 point-tracking neural network to be an exponential moving average (EMA) of parameters of the student 520 point-tracking neural network. That is, instead of freezing the teacher 510 point-tracking neural network during training, the system continues to update the teacher 510 neural network during training to improve the quality of the training process.
[00208] The point-tracking neural network described with reference to FIGS. 4 or 5 can be used for any of a variety of purposes.
[00209] For example, the point-tracking neural network po Petition 870250079681, dated 05 / 09 / 2025, page 50 / 99 41 / 62 of this can be used to generate training data for the generative neural networks described above.
[00210] As another example, the predicted positions generated by the point-tracking neural network can be used to generate a reward signal to train a robot or other agent through reinforcement learning, for example, if the task being performed by the agent requires moving a point in the scene from one location to another, the distance between the predicted position of the point in the last video frame in the sequence and the destination location can be used to generate a reward.
[00211] As another example, the predicted positions (and optionally the occlusion scores) generated by the point-tracking neural network can be provided as an additional input to a policy neural network to control a robot or other agent interacting with an environment. In this example, the lookup points could be points of interest in the last frame of the video sequence, and the predictions for the previous frames in the sequence could be provided as input to the policy neural network, for example, to provide a signal as to the recent movement of the agent or other objects in the environment.
[00212] As another example, the predicted positions (and optionally the occlusion scores) generated by the point-tracking neural network can be provided, along with the corresponding video sequence, as input to a video comprehension neural network, for example, an action classification neural network or a topic classification neural network, to provide additional information to the video comprehension neural network regarding movement in the scene.
[00213] As another example, the predicted positions (and optionally the occlusion scores) generated by the tracking neural network Petition 870250079681, dated 05 / 09 / 2025, page 51 / 99 42 / 62 point increments can be used for imitation learning, that is, to allow the imitation of movement rather than appearance.
[00214] As described above, the first generative neural network and / or the second generative neural network may be a diffusion neural network. In general, a diffusion neural network is a neural network that, at any given time step, is configured to process a diffusion input that includes (i) a current noisy data item (such as a noisy trajectory, noisy coordinate, noisy occlusion estimate, or noisy video frame) and, optionally, (ii) data specifying the given time step to generate a noise removal output that defines an estimate of a noise component of the current noisy data item, given the current time step. The noise component estimate is an estimate of the noise that was added to an original data item (such as a point trajectory or a video frame) to generate the current noisy data item. The noise removal output may be, for example, the noise component estimate or an estimate of the original data item.
[00215] After training, the system or another inference system uses the trained diffusion neural network to generate an output data item (such as a point trajectory or a video frame) in multiple time steps, performing a backscattering process to gradually remove noise from an initial data item until the final output data item is reached. Some or all of the values in the initial data item are noisy values, that is, they are sampled from an appropriate noise distribution.
[00216] In other words, the initialized data item has the same dimensionality as the final data item, but has noisy values. For example, the system can initialize the data item, that is, it can generate the first instance of the data item, sampling each value in the data item from a corresponding noise distribution, for example, Petition 870250079681, dated 05 / 09 / 2025, page 52 / 99 43 / 62 a Gaussian distribution or a different noise distribution. That is, the output data item includes multiple values and the initial data item includes the same number of values, with some or all of the values being sampled from a corresponding noise distribution.
[00217] As noted above, the pixel diffusion neural network can generally have any appropriate neural network architecture that allows the pixel diffusion neural network to map an input to an output image.
[00218] In some implementations, the pixel diffusion neural network architecture may be similar to the U-Net neural network architecture described by O. Ronneberger et al. in “U-Net: Convolutional Networks for Biomedical Image Segmentation”, arXiv:1505.04597. In one particular example, the pixel diffusion neural network may be implemented as convolutional neural networks including a downward analysis path and an upward synthesis path, where each path includes multiple neural network layers. The analysis path may include multiple subsampling layers, for example, convolutional, and the synthesis path may include multiple suprasampling layers, for example, supra-convolutional. In addition to the convolutional layers, upward and / or downward sampling may be partially or fully implemented by interpolation. Neural networks may include shortcut or residual hop connections between layers of equal resolution in the analysis and synthesis paths.In some implementations, at least one of a set of one or more layers between the analysis and synthesis paths includes a fully connected set of layers.
[00219] In some implementations, the first or second generative neural network may comprise one or more layers of self-attention. In general, a layer of self-attention may be one that Petition 870250079681, dated 05 / 09 / 2025, page 53 / 99 44 / 62 applies an attention mechanism to elements of an embedding (e.g., input data) to update each element of the embedding. For example, an input embedding is used to determine a query vector and a set of key-value vector pairs (key-value query attention), and the updated embedding comprises a weighted sum of the values, weighted by a similarity function of the query for each respective key. There are many different attention mechanisms that can be used. For example, the attention mechanism could be a scalar product attention mechanism applying a query vector to the respective key vector to determine the respective weights for each value vector, then combining a plurality of value vectors using the respective weights to determine the output of the attention layer for each element of the input sequence.
[00220] The system and method described in this document can be used to predict how a physical system or environment at a particular time will evolve over one or more subsequent time steps. That is, the input image may depict a scene comprising a physical environment, and the corresponding video generated from the input image may comprise video frames that predict the physical environment at one or more time steps after the particular time. For example, the input image may comprise one or more objects, and each video frame may be a prediction of the location and / or configuration of each of the objects in the physical environment at the corresponding time step.
[00221] Video can be used as an input for a control task. For example, video can be used for predictive model control of an agent, such as a robot or mechanical agent. In such cases, the noisy trajectory can be conditioned on one or more actions that can be performed by the robot or mechanical agent, of Petition 870250079681, dated 05 / 09 / 2025, page 54 / 99 45 / 62 so that the video is a prediction of the physical environment that would be obtained if the robot or mechanical agent performed one or more actions. Thus, the video can be used by a control system to control a mechanical agent, such as a robot, to perform a particular task by processing the predicted video using the control system to generate a control signal to direct the mechanical agent, according to the video, to perform the task. The input image can be captured by one or more sensors in the physical environment, for example, one or more sensors of the robot or mechanical agent. The input image may depict a part of the robot or mechanical agent, such as a gripper hand or other device for manipulating objects.As another example, video frames, or an embedding of video frames, can be provided as an additional input to a policy neural network that is used to select actions to be performed by a robot or other agent interacting with a physical environment to accomplish a specific task, such as navigating within the physical environment or identifying or manipulating one or more objects in the physical environment.
[00222] In some implementations, the system and method described in this document can be used for estimation or prediction of human pose. For example, the input image may depict a person in a pose at a particular moment, and each of the video frames may then be a prediction of a pose that the person will adopt at a corresponding time after that particular moment.
[00223] This descriptive report uses the term configured in connection with computer systems and program components. A system of one or more computers being configured to perform particular operations or actions means that the system has installed on it software, firmware, hardware, or a combination thereof that, Petition 870250079681, dated 05 / 09 / 2025, page 55 / 99 46 / 62 in operation, cause the system to perform operations or actions. One or more computer programs being configured to perform particular operations or actions means that the one or more programs include instructions that, when executed by the data processing device, cause the device to perform the operations or actions. With reference to specific neural networks, a neural network can be configured to perform a specific action by being trained to perform that particular action.
[00224] The material modalities and functional operations described in this descriptive report may be implemented in digital electronic circuits, in tangibly embedded computer software or firmware, in computer hardware, including the structures disclosed in this descriptive report and their structural equivalents, or in combinations of one or more thereof. The material modalities described in this descriptive report may be implemented as one or more computer programs, for example, one or more computer program instruction modules encoded in a tangible non-transient storage medium for execution by, or to control the operation of, data processing apparatus. The computer storage medium may be a machine-readable storage device, a machine-readable storage substrate, a random-access or serial memory device, or a combination of one or more thereof.Alternatively or additionally, the program instructions may be encoded in an artificially generated propagated signal, for example, a machine-generated electrical, optical, or electromagnetic signal that is generated to encode information for transmission to a suitable receiving device for execution by a data processing device.
[00225] The term data processing device refers to Petition 870250079681, dated 05 / 09 / 2025, page 56 / 99 47 / 62 Data processing hardware encompasses all types of apparatus, devices, and machines for data processing, including, for example, a programmable processor, a computer, or multiple processors or computers. The apparatus may also be, or include, special-purpose logic circuits, for example, an FPGA (field-programmable gate array) or an ASIC (application-specific integrated circuit). The apparatus may optionally include, in addition to hardware, code that creates an execution environment for computer programs, for example, code that constitutes processor firmware, a protocol stack, a database management system, an operating system, or a combination of one or more of these.
[00226] A computer program, which may also be referred to or described as a program, software, a software application, an app, a module, a software module, a script, or code, may be written in any form of programming language, including compiled or interpreted languages, or declarative or procedural languages; and may be implemented in any form, including as a stand-alone program or as a module, component, subroutine, or other unit suitable for use in a computing environment. A program may, but does not have to, correspond to a file in a file system. A program may be stored in a portion of a file that contains other programs or data, for example, one or more scripts stored in a markup language document, in a single file dedicated to the program in question, or in multiple coordinated files, for example, files that store one or more modules, subprograms, or portions of code.A computer program can be implemented to run on one computer or on multiple computers located on the same site. Petition 870250079681, dated 05 / 09 / 2025, page 57 / 99 48 / 62 distributed across multiple sites and interconnected by a data communication network.
[00227] In this descriptive report, the term database is used broadly to refer to any collection of data: the data need not be structured in any particular way, nor structured in any way, and may be stored on storage devices in one or more locations. Thus, for example, an index database may include multiple collections of data, each of which may be organized and accessed differently.
[00228] Similarly, in this descriptive report, the term tool is used broadly to refer to a software-based system, subsystem, or process that is programmed to perform one or more specific functions. Generally, a tool will be implemented as one or more software modules or components, installed on one or more computers in one or more locations. In some cases, one or more computers will be dedicated to a specific tool; in other cases, multiple tools may be installed and run on the same computer or computers.
[00229] The processes and logical flows described in this descriptive report can be performed by one or more programmable computers executing one or more computer programs to perform functions operating on input data and generating output. The processes and logical flows can also be performed by special-purpose logic circuits, for example, an FPGA or an ASIC, or by a combination of special-purpose logic circuits and one or more programmed computers.
[00230] Computers suitable for running a computer program may be based on general-purpose or special-purpose microprocessors, or both, or any other type of unit. Petition 870250079681, dated 05 / 09 / 2025, pp. 58 / 99 49 / 62 of central processing. Generally, a central processing unit will receive instructions and data from either read-only memory or random-access memory, or both. The essential elements of a computer are a central processing unit to perform or execute instructions and one or more memory devices to store instructions and data. The central processing unit and memory may be supplemented by, or incorporated into, logic circuits for special purposes. In general, a computer will also include, or be operationally coupled to receive data from, or transfer data to, or both, one or more mass storage devices to store data, for example, magnetic disks, magnetic optical disks, or optical disks. However, a computer does not need to have such devices.Furthermore, a computer can be incorporated into another device, for example, a mobile phone, a personal digital assistant (PDA), a mobile audio or video player, a game console, a Global Positioning System (GPS) receiver, or a portable storage device, for example, a Universal Serial Bus (USB) flash drive, to name just a few.
[00231] Computer-readable media suitable for storing computer program instructions and data include all forms of non-volatile memory, memory media and devices, including, by way of example, semiconductor memory devices, for example, EPROM, EEPROM and flash memory devices; magnetic disks, for example, internal hard disks or removable disks; magneto-optical disks; and CD-ROM and DVD-ROM disks.
[00232] To provide interaction with a user, the modalities of the subject matter described in this descriptive report may be implemented on a computer with a display device, for example, a CRT (cathode ray tube) or LCD (crystal glass display) monitor. Petition 870250079681, dated 05 / 09 / 2025, page 59 / 99 50 / 62 liquid), to display information to the user, and a keyboard and a pointing device, for example, a mouse or a trackball, by which the user can provide input to the computer. Other types of devices can be used to provide interaction with a user as well; for example, feedback provided to the user can be any form of sensory feedback, for example, visual feedback, auditory feedback, or tactile feedback; and user input can be received in any form, including acoustic, speech, or tactile input. Furthermore, a computer can interact with a user by sending documents to and receiving documents from a device that is used by the user; for example, by sending web pages to a web browser on a user's device in response to requests received from the web browser.Also, a computer can interact with a user by sending text messages or other forms of messaging to a personal device, for example, a smartphone, which is running a messaging application and receiving responsive messages from the user in return.
[00233] The data processing apparatus for implementing machine learning models may also include, for example, special-purpose hardware accelerator units for processing common and computationally intensive parts of machine learning training or production, for example, inference, workloads.
[00234] Machine learning models can be implemented and deployed using a machine learning framework, for example, a TensorFlow framework or a Jax framework.
[00235] The modalities of the subject matter described in this descriptive report can be implemented in a computer system that in Petition 870250079681, dated 05 / 09 / 2025, p. 60 / 99 51 / 62 includes a back-end component, for example, a data server, or includes a middleware component, for example, an application server, or includes a front-end component, for example, a client computer with a graphical user interface, a web browser, or an application through which a user can interact with an implementation of the subject matter described in this descriptive report, or any combination of one or more of these back-end, middleware, or front-end components. The system components may be interconnected by any form or means of digital data communication, for example, a communication network. Examples of communication networks include a local area network (LAN) and a wide area network (WAN), for example, the Internet.
[00236] The computer system may include clients and servers. A client and a server are generally remote from each other and typically interact through a communication network.The client-server relationship emerges by virtue of computer programs running on the respective computers and through a client / server relationship with each other. In some modalities, a server transmits data, for example, an HTML page, to a user's device, for example, for the purpose of displaying data and receiving user input from a user interacting with the device, which acts as a client. The data generated on the user's device, for example, a result of the user's interaction, can be received by the server from the device.
[00237] Although this descriptive report contains many specific implementation details, these should not be interpreted as limitations on the scope of any invention or on the scope of what may be claimed, but rather as descriptions of features that may be specific to particular embodiments of particular inventions. Certain features that are described in this report Petition 870250079681, dated 05 / 09 / 2025, p. 61 / 99 52 / 62 Descriptive features in the context of separate embodiments can also be implemented in combination in a single embodiment. Conversely, several features that are described in the context of a single embodiment can also be implemented in multiple embodiments separately or in any suitable subcombination. Furthermore, although features may be described above as acting in certain combinations and even initially claimed as such, one or more features of a claimed combination may, in some cases, be removed from the combination and the claimed combination may be directed to a subcombination or variation of a subcombination.
[00238] Similarly, although operations are depicted in the figures and cited in the claims in a particular order, this should not be understood as requiring that such operations be performed in the particular order shown or in sequential order, or that all illustrated operations be performed to achieve desirable results. In certain circumstances, multitasking and parallel processing may be advantageous. Furthermore, the separation of various system modules and components in the embodiments described above should not be understood as requiring such separation in all embodiments, and it should be understood that the described program components and systems can generally be integrated together into a single software product or packaged into multiple software products.
[00239] The particular embodiments of the matter have been described. Other embodiments are within the scope of the following claims. For example, the actions mentioned in the claims can be performed in a different order and still achieve desirable results. As an example, the processes depicted in the accompanying figures do not necessarily require the particular order shown, or Petition 870250079681, dated 05 / 09 / 2025, page 62 / 99 53 / 62 sequential processing is necessary to achieve desired results. In some cases, multitasking and parallel processing may be advantageous.
[00240] The aspects of this disclosure may be as set out in the following clauses:
[00241] Clause 1. A method performed by one or more computers, the method comprising: receive data specifying a teacher point-tracking neural network that is configured to receive a teacher point-tracking input comprising (i) a video sequence comprising a plurality of video frames and (ii) a query point in one of the plurality of video frames and to generate a teacher point-tracking output comprising, for each other video frame of the plurality of video frames, a position of the query point in the other video frame and an occlusion estimate for the query point in the other video frame;and to train a student point-tracking neural network via unsupervised learning, wherein the student point-tracking neural network is configured to receive a student point-tracking input comprising (i) the video sequence comprising the plurality of video frames and (ii) the query point in one of the plurality of video frames and to generate a student point-tracking output comprising, for each other video frame of the plurality of video frames, a position of the query point in the other video frame and an occlusion estimate for the query point in the other video frame, the training comprising, in each of a plurality of training steps:; receive a set of one or more training video sequences from a training dataset comprising Petition 870250079681, dated 05 / 09 / 2025, page 63 / 99 54 / 62 providing a plurality of training video sequences; for each training video sequence: Apply one or more transformations to the training video sequence to generate a transformed video sequence; Select one or more instructor reference points within the training video sequence; for each point of reference, professor: Generate a teacher point-tracking output for the teacher query point by processing the video sequence using the teacher point-tracking neural network; Identify a corresponding student lookup point in the transformed video sequence; Generate a student point-tracing output for the student lookup point by processing the transformed video sequence using the student point-tracing neural network; and train the student point-tracing neural network using a loss function that, for each training video sequence and for each teacher lookup point, measures a difference between the student point-tracing output for the corresponding student lookup point and the teacher point-tracing output for the teacher lookup point.
[00242] Clause 2. The method of clause 1, in which, prior to training, the teacher point-tracking neural network was pre-trained.
[00243] Clause 3. The method of clause 2, in which the teacher point-tracking neural network was pre-trained through supervised learning.
[00244] Clause 4. The method of clause 3, in which the point-tracking neural network was pre-trained through supervised learning on a dataset of synthetic video sequences Petition 870250079681, dated 05 / 09 / 2025, p. 64 / 99 55 / 62 ticas.
[00245] Clause 5. The method of any previous clause, where the teacher and student point-tracking neural networks have the same architecture.
[00246] Clause 6. The method of clause 5, comprising, even before training the student point-tracking neural network, initializing parameters of the student point-tracking neural network using the teacher point-tracking neural network.
[00247] Clause 7. The method of either clause 5 or 6, the training further comprising, in each of at least one subset of the training steps:
[00248] update the teacher point-tracking neural network using the student point-tracking neural network.
[00249] Clause 8. The method in clause 7, where the updating of the teacher point-tracking neural network using the student point-tracking neural network comprises:
[00250] update teacher point-tracking neural network parameters to be an exponential moving average (EMA) of student point-tracking neural network parameters.
[00251] Clause 9. The method of any preceding clause, wherein one or more transformations comprise spatial transformations.
[00252] Clause 10. The method of any previous clause, in which one or more transformations comprise image corruptions.
[00253] Clause 11. The method of any preceding clause, wherein the identification of a corresponding student lookup point in the transformed video sequence comprises: Identify an initial student lookup point using the teacher point tracking output; and Petition 870250079681, dated 05 / 09 / 2025, pp. 65 / 99 56 / 62 transform the initial student lookup point consistent with one or more transformations to generate the student lookup point.
[00254] Clause 12. The method of any previous clause, wherein the teacher point tracking output and the student point tracking output further comprise the respective uncertainty scores for each predicted position.
[00255] Clause 13. The method of any preceding clause, wherein the training of the student point-tracking neural network using a loss function that, for each training video sequence and for each teacher lookup point, measures a difference between the student point-tracking output for the corresponding student lookup point and the teacher point-tracking output for the teacher lookup point, comprises: For each training video sequence and for each teacher lookup point, generate a pseudo-markup from the teacher point tracking output, where the loss function measures an error between the pseudo-markup and the student point tracking output.
[00256] Clause 14. The method of any preceding clause, wherein the training of the student point-tracking neural network using a loss function that, for each training video sequence and for each teacher lookup point, measures a difference between the student point-tracking output for the corresponding student lookup point and the teacher point-tracking output for the teacher lookup point comprises: Determine whether to mask the teacher lookup point from the loss function by applying loop consistency between the student point tracking output and the teacher point tracking output. Petition 870250079681, dated 05 / 09 / 2025, pp. 66 / 99 57 / 62
[00257] Clause 15. The method of any preceding clause, wherein the training of the student point-tracking neural network using a loss function that, for each training video sequence and for each teacher lookup point, measures a difference between the student point-tracking output for the corresponding student lookup point and the teacher point-tracking output for the teacher lookup point comprises: Determine whether to mask predictions for the teacher query point for any loss function training videos by applying a proximity mask.
[00258] Clause 16. The method of any preceding clause, wherein the training further comprises, in each of a plurality of training stages: Train the learner neural network through supervised learning on a set of tagged video sequences.
[00259] Clause 17. A system comprising: one or more computers; and one or more storage devices communicatively coupled to one or more computers, wherein the one or more storage devices store instructions which, when executed by the one or more computers, cause the one or more computers to perform operations of the respective method of any of clauses 1-16.
[00260] Clause 18. One or more non-transient computer storage media storing instructions that, when executed by one or more computers, cause the one or more computers to perform operations of the respective method of any of clauses 1-16.
[00261] Clause 19. A method performed by one or more computers to generate a video that animates an input image. Petition 870250079681, dated 05 / 09 / 2025, pp. 67 / 99 58 / 62 through a plurality of time stages, the method comprising: Receive the input image; To process a first input derived from the input image using a first generative neural network to generate respective point trajectories for each of one or more points in the input image, where each point trajectory comprises, for each of the plurality of time steps in the video, a predicted spatial position of the corresponding point in a video frame at the time step in the video; and to generate each of the video frames in the video that animates the input image using a second generative neural network and based on the input image and the one or more point trajectories.
[00262] Clause 20. The method of clause 19, wherein each point trajectory comprises, for each of the time steps, (i) the predicted spatial position of the corresponding point in the video frame at the time step and (ii) an occlusion score that estimates a probability that the corresponding point will be occluded in the video frame at the time step.
[00263] Clause 21. The method of clause 19 or clause 20, further comprising: To process the input image using an image encoding neural network to generate an encoded representation of the input image, where the first input comprises the encoded representation of the input image.
[00264] Clause 22. The method of any of the 1921 clauses, wherein the first generative neural network is a diffusion neural network that generates each point path from a corresponding noisy path conditioned on the first input. Petition 870250079681, dated 05 / 09 / 2025, pp. 68 / 99 59 / 62
[00265] Clause 23. The method of clause 22, wherein the first generative neural network was trained on a training dataset comprising (i) a plurality of video sequences and (ii) for each of the video sequences, a respective point trajectory for each of one or more points in a first frame within the video sequence.
[00266] Clause 24. Clause 22 or Clause 23 method, where the first generative neural network comprises a two-dimensional convolutional neural network. Clause 25. The method of clause 24, where the two-dimensional convolutional neural network is a U-Net.
[00267] Clause 26. The method of clause 24 or clause 25, in which the two-dimensional convolutional neural network comprises one or more layers of self-attention.
[00268] Clause 27. The method according to any of clauses 22-26, where each corresponding noisy path comprises a concatenation of noisy coordinates and an estimate of noisy occlusion.
[00269] Clause 28. Method, from clause 27, in which noisy coordinates are augmented with a positional coding.
[00270] Clause 29. The method of clause 23, wherein, for at least a subset of the video sequences, one or more of the point trajectories were generated by processing an input comprising the corresponding point in the first frame using a point-tracking neural network.
[00271] Clause 30. The method of clause 29, in which the point-tracking neural network is configured for, for each video sequence in the subset and for a corresponding point in the first frame of the video: generate a query feature for the matching point Petition 870250079681, dated 05 / 09 / 2025, pp. 69 / 99 60 / 62 corresponding; Using the query feature, generate a cost volume comprising a corresponding cost map for each of a plurality of frames in the video sequence; and generate, for each of the plurality of frames, an initial position of the corresponding point in the frame and an initial occlusion estimate for the corresponding point in the frame using the cost map for the frame.
[00272] Clause 31. The method of clause 30, in which the point-tracking neural network is further configured for: Generate the point trajectory to the corresponding point by refining the initial positions, refining the initial occlusion estimates for the plurality of frames, or both, using an initial point trajectory that comprises the initial positions and the initial occlusion estimates for the plurality of frames.
[00273] Clause 32. The method of clause 31, in which refining the initial positions and initial occlusion estimates for the plurality of frames using an initial point trajectory comprising the initial positions and initial occlusion estimates for the plurality of frames comprises, in each of one or more refinement iterations: To generate, from a current point trajectory through refinement iterations, a set of local scoring maps that, for each frame, capture similarity between features in a neighborhood of the predicted position in the current point trajectory in the frame and the query feature; and to process an input comprising the set of local scoring maps using a refinement neural network to generate an update to the current point trajectory.
[00274] Clause 33. The method of clause 32, in which the neu network Petition 870250079681, dated 05 / 09 / 2025, pp. 70 / 99 61 / 62 refinement neural network is a depth-mix neural network that propagates information across frames using deep convolutional layers.
[00275] Clause 34. The method of clause 32 or clause 33, wherein the input comprises the set of local scoring maps, the query feature and the current point trajectory.
[00276] Clause 35. The method of any of the clauses 1934, wherein the second generative neural network is a diffusion neural network that generates each frame from a corresponding noisy frame conditioned on a conditioning input derived from the input image and one or more point trajectories.
[00277] Clause 36. The method of clause 35, wherein, for each image, the conditioning input comprises a distorted version of the input image that has been distorted according to one or more point trajectories to represent the image.
[00278] Clause 37. The method of clause 36, wherein, for each image, the conditioning input comprises a distorted version of features extracted from the input image that have been distorted according to one or more point trajectories to represent the image.
[00279] Clause 38. The method of clause 36 or clause 37, wherein the diffusion neural network generates each frame through a plurality of back-diffusion iterations and wherein, for at least a subset of the frames, the input to a given back-diffusion iteration comprises a current version of one or more previous frames in the video from the back-diffusion iteration.
[00280] Clause 39. A system comprising:
[00281] one or more computers; and one or more storage devices communicatively coupled to one or more computers, wherein the one or more Petition 870250079681, dated 05 / 09 / 2025, pp. 71 / 99 62 / 62 storage devices store instructions that, when executed by one or more computers, cause the one or more computers to perform operations of the respective method of any of clauses 19-38.
[00282] Clause 40. One or more non-transient computer storage media storing instructions which, when executed by one or more computers, cause the one or more computers to perform operations of the respective method of any of clauses 19-38. Petition 870250079681, dated 05 / 09 / 2025, pp. 72 / 99
Claims
1 / 7 CLAIMS 1. A method performed by one or more computers for generating a video that animates an input image through a plurality of time steps, the method characterized in that it comprises: receiving the input image; processing a first input derived from the input image using a first generative neural network to generate respective point trajectories for each of one or more points in the input image, wherein each point trajectory comprises, for each of the plurality of time steps in the video, a predicted spatial position of the corresponding point in a video frame at the time step in the video; and generating each of the video frames in the video that animates the input image using a second generative neural network and based on the input image and the one or more point trajectories.
2. Method, according to claim 1, characterized in that each point trajectory comprises, for each of the time steps, (i) the predicted spatial position of the corresponding point in the video frame at the time step and (ii) an occlusion score that estimates a probability that the corresponding point will be occluded in the video frame at the time step.
3. A method according to claim 1 or 2, characterized in that it further comprises: processing the input image using an image encoding neural network to generate an encoded representation of the input image, wherein the first input comprises the encoded representation of the input image.
4. Method, according to any of the preceding claims Petition 870250079681, dated 05 / 09 / 2025, page 73 / 99 2 / 7, characterized in that the first generative neural network is a diffusion neural network that generates each point path from a corresponding noisy path conditioned on the first input.
5. Method according to claim 4, characterized in that the first generative neural network was trained on a training dataset comprising (i) a plurality of video sequences and (ii) for each of the video sequences, a respective point trajectory for each of one or more points in a first frame within the video sequence.
6. Method, according to claim 4 or 5, characterized in that the first generative neural network comprises a two-dimensional convolutional neural network.
7. Method according to claim 6, characterized in that the two-dimensional convolutional neural network is a UNet.
8. Method, according to claim 6 or 7, characterized in that the two-dimensional convolutional neural network comprises one or more layers of self-attention.
9. A method, according to any one of claims 4 to 8, characterized in that each corresponding noisy trajectory comprises a concatenation of noisy coordinates and an estimate of noisy occlusion.
10. Method, according to claim 9, characterized in that the noisy coordinates are augmented with a positional encoding.
11. Method, according to claim 5, characterized in that, for at least a subset of the video sequences, one or more of the point trajectories were generated by processing an input comprising the corresponding point in the first frame using a point-tracking neural network.
12. Method, according to claim 11, characterized in that the point-tracking neural network is configured to, for each video sequence in the subset and for a corresponding point in the first frame of the video: generate a query feature for the corresponding point; generate, using the query feature, a cost volume comprising a respective cost map for each of a plurality of frames in the video sequence; and generate, for each of the plurality of frames, an initial position of the corresponding point in the frame and an initial occlusion estimate for the corresponding point in the frame using the cost map for the frame.
13. Method, according to claim 12, characterized in that the point-tracking neural network is further configured to: generate the point trajectory for the corresponding point by refining the initial positions, refining the initial occlusion estimates for the plurality of frames, or both by using an initial point trajectory comprising the initial positions and the initial occlusion estimates for the plurality of frames.
14. Method according to claim 13, characterized in that refining the initial positions and initial occlusion estimates for the plurality of frames using an initial point trajectory comprising the initial positions and initial occlusion estimates for the plurality of frames comprises, in each of one or more refinement iterations: generating, from a current point trajectory across the refinement iterations, a set of local scoring maps that, for each frame, capture similarity between features in a neighborhood of the predicted position in the current point trajectory in the frame and the query feature; and processing an input comprising the set of local scoring maps using a refinement neural network to generate an update to the current point trajectory.
15. Method, according to claim 14, characterized in that the refinement neural network is a depth-mixing neural network that propagates information through frames using depth-convolutional layers.
16. Method, according to claim 14 or 15, characterized in that the input comprises the set of local scoring maps, the query feature and the current point trajectory.
17. A method, according to any of the preceding claims, characterized in that the second generative neural network is a diffusion neural network that generates each frame from a corresponding noisy frame conditioned on a conditioning input derived from the input image and one or more point trajectories.
18. Method according to claim 17, characterized in that, for each image, the conditioning input comprises a distorted version of the input image that has been distorted according to one or more point trajectories to represent the image.
19. Method according to claim 18, characterized in that, for each image, the conditioning input comprises a distorted version of features extracted from the input image that have been distorted according to one or more point trajectories to represent the image. Petition 870250079681, dated 05 / 09 / 2025, p. 76 / 99 5 / 7 20. A method according to claim 18 or 19, characterized in that the diffusion neural network generates each frame through a plurality of back-diffusion iterations and wherein, for at least a subset of the frames, the input to a given back-diffusion iteration comprises a current version of one or more previous frames in the video from the back-diffusion iteration.
21. A method, according to any of the preceding claims, characterized in that the input image depicts a real-world environment at a particular moment and each of the video frames is a prediction of the real-world environment at a corresponding time after that particular moment.
22. A method according to claim 21, characterized in that it further comprises using a control system to control a robot or mechanical agent to perform a particular task by processing video and using the control system to generate one or more control signals to control the robot or mechanical agent, in accordance with the video, to perform the task.
23. A method, according to any of the preceding claims, characterized in that the input image depicts a person in a pose at a particular moment, and each of the video frames is a prediction of a pose that the person will adopt at a corresponding moment after that particular moment.
24. A method performed by one or more computers to generate, for each of one or more video sequences, a respective point path for each of one or more points in a first frame within the video sequence, the method characterized in that it comprises: generating one or more of the point paths by processing Petition 870250079681, dated 05 / 09 / 2025, page 77 / 99 6 / 7 an input comprising the corresponding point in the first frame using a point-tracking neural network, wherein the point-tracking neural network is configured to: for each video sequence and for the corresponding point in the first frame of the video: generate a query feature for the corresponding point; generate, using the query feature, a cost volume comprising a respective cost map for each of a plurality of frames in the video sequence;and generate, for each of the plurality of frames, an initial position of the corresponding point in the frame and an initial occlusion estimate for the corresponding point in the frame using the cost map for the frame; and generate the point path to the corresponding point by refining the initial positions, refining the initial occlusion estimates for the plurality of frames, or both, using an initial point path that comprises the initial positions and the initial occlusion estimates for the plurality of frames.
25. A method according to claim 24, characterized in that refining the initial positions and initial occlusion estimates for a plurality of frames using an initial point trajectory comprising the initial positions and initial occlusion estimates for a plurality of frames comprises, in each of one or more refinement iterations: generating, from a current point trajectory across the refinement iterations, a set of local scoring maps that, for each frame, capture similarity between features in a neighborhood of the predicted position in the current point trajectory in the frame and the query feature; and processing an input comprising the set of local scoring maps using a refinement neural network to generate an update to the current point trajectory.
26. Method according to claim 25, characterized in that the refinement neural network is a depth-mixing neural network that propagates information through frames using depth-convolutional layers.
27. Method, according to claim 25 or 26, characterized in that the input comprises the set of local scoring maps, the query feature and the current point trajectory.
28. Method, according to any one of claims 24 to 27, characterized in that the method is for generating a training dataset comprising (i) one or more video sequences and (ii) for each of the video sequences, the respective point trajectory for each of one or more points in the first frame within the video sequence.
29. A system characterized in that it comprises: one or more computers; and one or more storage devices communicatively coupled to one or more computers, wherein the one or more storage devices store instructions which, when executed by the one or more computers, cause the one or more computers to perform operations of the respective method, as defined in any one of claims 1 to 28.
30. Non-transient computer storage media characterized in that they store instructions which, when executed by one or more computers, cause the one or more computers to perform operations of the respective method, as defined in any one of claims 1 to 28. Petition 870250079681, dated 05 / 09 / 2025, pp. 79 / 99